EDBT 2026 Demo / reviewers in the wild / expert
Linfeng Song
dblp:136/3610
· DBLP profile ↗
86ranked-venue papers
19as first author
60since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 72 · 13 first-author · 46 since 2021Graphics, computer vision, multimedia, augmented reality and games · 12 · 1 first-author · 5 since 2021Computer networks · 11 · 5 first-author · 11 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | EconProver: Towards More Economical Test-Time Scaling for Automated Theorem ProvingabstractMukai Li, Linfeng Song, Zhenwen Liang, Jiahao Xu, Shansan Gong, Qi Liu, Haitao Mi, Dong Yu. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Mukai Li, Linfeng Song, Zhenwen Liang, Shansan Gong, Qi Liu 0049, Haitao Mi, Dong Yu 0001 |
ACL (1) | 2 |
| 2026 | Your Reasoning Model is Secretly a Reward Model - Optimization-Free Verification from ExperienceabstractZhenwen Liang, Ruosen Li, Yujun Zhou, Linfeng Song, Dian Yu, Xinya Du, Haitao Mi, Dong Yu. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Zhenwen Liang, Ruosen Li, Yujun Zhou 0002, Linfeng Song, Dian Yu 0001, Xinya Du, Haitao Mi, Dong Yu 0001 |
ACL (1) | 4 |
| 2026 | Crossing the Reward Bridge: Expanding Reinforcement Learning with Verifiable Rewards Across Diverse DomainsabstractYi Su, Dian Yu, Linfeng Song, Juntao Li, Haitao Mi, Zhaopeng Tu, Min Zhang, Dong Yu. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Yi Su 0006, Dian Yu 0001, Linfeng Song, Juntao Li 0005, Haitao Mi, Zhaopeng Tu, Min Zhang 0005, Dong Yu 0001 |
ACL (1) | 3 |
| 2026 | HF Skywave Massive MIMO Communications with Interference Sparsity-Aware Turbo Receiver
Linfeng Song, Rui Sun 0017, Ding Shi, Yanbo Yu, Anan Lu, Xiqi Gao 0001, Geoffrey Ye Li, Xiang-Gen Xia 0001 |
WCNC | 1 |
| 2026 | Signal Detection for User-Centric Network Massive MIMO SystemabstractIn this paper, we investigate the signal detection for user-centric network (UCN) massive multi-input multi-output (mMIMO) system. We consider that the users are divided into multiple user groups (UGs). For each UG, leveraging the interference sparsity, we reveal that the performance of the minimum-mean-square-error (MMSE) detector can be guaranteed in the network mMIMO system by using the matched filtering (MF) outputs of the intra-group and interfering users. Then, with the base station (BS) connection sparsity, we reveal that the detection performance of each UG is primarily determined by a limited number of associated BSs. To facilitate practical application, we propose a straightforward user grouping method and outline the process for determining interfering users and associated BSs for each UG in UCN mMIMO systems. Then, we propose a user-centric detection method that decouples the detection process for each UG into two stages. In the first stage, local MF is performed at each associated BSs using local information. In the second stage, group-wise interference cancellation (IC) is carried out to obtain detection results at the primary serving BS (PSBS) of each UG, with information exchanged from auxiliary serving BSs (ASBSs). Simulation results confirm the effectiveness and computational efficiency of our proposed user-centric detection for the UCN mMIMO system. Rui Sun 0017, Linfeng Song, Chen Sun 0004, Ding Shi, Xiqi Gao 0001, Xiang-Gen Xia 0001 |
IEEE Trans. Commun. | 2 |
| 2026 | Beam Structured Turbo Receiver for HF Skywave Massive MIMOabstractIn this paper, we investigate receiver design for high frequency (HF) skywave massive multiple-input multiple-output (MIMO) communications. We first establish a modified beam based channel model (BBCM) by performing uniform sampling for directional cosine with deterministic sampling interval, where the beam matrix is constructed using a phase-shifted discrete Fourier transform (DFT) matrix. Based on the modified BBCM, we propose a beam structured turbo receiver (BSTR) involving low-dimensional beam domain signal detection for grouped user terminals (UTs), which is proved to be asymptotically optimal in terms of minimizing mean-squared error (MSE). Moreover, we extend it to windowed BSTR by introducing a windowing approach for interference suppression and complexity reduction, and propose a well-designed energy-focusing window. We also present an efficient implementation of the windowed BSTR by exploiting the structure properties of the beam matrix and the beam domain channel sparsity. Simulation results validate the superior performance of the proposed receivers but with remarkably low complexity. Linfeng Song, Ding Shi, Xiqi Gao 0001, Geoffrey Ye Li, Xiang-Gen Xia 0001 |
IEEE Trans. Wirel. Commun. | 1 |
| 2026 | Interference Sparsity-Aware Turbo Receiver for HF Skywave Massive MIMOabstractIn this paper, we propose a low complexity turbo receiver for high frequency (HF) skywave massive multiple-input multiple-output (MIMO) systems. We first introduce the beam based channel model (BBCM) with uniform sampling for directional cosine. By leveraging the BBCM, we reveal the interference sparsity of HF skywave massive MIMO systems, which is defined as the asymptotic sparsity of the channel Gram matrix. Exploiting the interference sparsity, we provide a condition of extracting sufficient observation for signal detection. Motivated by this condition, we construct the interference user terminal (UT) set (IUS) and extract the observation vector from the received signal after matched filtering (MF) for each UT. Then, a low-dimensional interference sparsity-aware detector (ISD) is separately designed for each UT by minimizing the mean-squared error (MSE), and the interference sparsity-aware turbo receiver (ISTR) is subsequently formulated using ISDs. Under a relaxed version of the condition for sufficient observation selection, we prove the optimality of the ISTR. Further, we develop an efficient implementation of the ISTR, involving approximate computation of the ISD, the signal reconstructed by ISD and the channel Gram matrix. Moreover, an efficient construction of IUS using the statistical channel state information (CSI) is also proposed. Simulation results confirm that the proposed ISTR achieves excellent performance with relatively low complexity. Linfeng Song, Rui Sun 0017, Ding Shi, Yanbo Yu, Anan Lu, Xiqi Gao 0001, Geoffrey Ye Li, Xiang-Gen Xia 0001 |
IEEE Trans. Wirel. Commun. | 1 |
| 2025 | LiteSearch: Efficient Tree Search with Dynamic Exploration Budget for Math ReasoningabstractRecent research suggests that tree search algorithms (e.g. Monte Carlo Tree Search) can dramatically boost LLM performance on complex mathematical reasoning tasks. However, they often require more than 10 times the computational resources of greedy decoding due to wasteful search strategies, making them difficult to be deployed in practical applications. This study introduces a novel guided tree search algorithm with a goal-directed heuristic function and node-level exploration budget (maximum number of children) calculation to tackle this issue. By considering the search progress towards the final answer (history) and the guidance from a value network (future) trained without any step-wise annotations, our algorithm iteratively selects the most promising tree node before expanding it within the boundaries of the allocated computational budget. Experiments conducted on the GSM8K, TabMWP, and MATH datasets demonstrate that our method not only offers competitive performance but also enjoys significantly lower computational costs compared to baseline methods. Ante Wang, Linfeng Song, Baolin Peng, Dian Yu 0001, Haitao Mi, Jinsong Su, Dong Yu 0001 |
AAAI | 2 |
| 2025 | Don't Get Lost in the Trees: Streamlining LLM Reasoning by Overcoming Tree Search Exploration PitfallsabstractRecent advancements in tree search algorithms guided by verifiers have significantly enhanced the reasoning capabilities of large language models (LLMs), but at the cost of increased computational resources. In this work, we identify two key challenges contributing to this inefficiency: \textit{over-exploration} due to redundant states with semantically equivalent content, and \textit{under-exploration} caused by high variance in verifier scoring leading to frequent trajectory switching. To address these issues, we propose FETCH – an e{\bf f}fici{\bf e}nt {\bf t}ree sear{\bf ch} framework, which is a flexible, plug-and-play system compatible with various tree search algorithms.Our framework mitigates over-exploration by merging semantically similar states using agglomerative clustering of text embeddings obtained from a fine-tuned SimCSE model. To tackle under-exploration, we enhance verifiers by incorporating temporal difference learning with adjusted \lambda-returns during training to reduce variance, and employing a verifier ensemble to aggregate scores during inference. Experiments on GSM8K, GSM-Plus, and MATH datasets demonstrate that our methods significantly improve reasoning accuracy and computational efficiency across four different tree search algorithms, paving the way for more practical applications of LLM-based reasoning. The code is available at https://github.com/DeepLearnXMU/Fetch. Ante Wang, Linfeng Song, Dian Yu 0001, Haitao Mi, Xiangyu Duan, Zhaopeng Tu, Jinsong Su, Dong Yu 0001 |
ACL (1) | 2 |
| 2025 | SR-LLM: Rethinking the Structured Representation in Large Language ModelabstractJiahuan Zhang, Tianheng Wang, Ziyi Huang, Yulong Wu, Hanqing Wu, DongbaiChen DongbaiChen, Linfeng Song, Yue Zhang, Guozheng Rao, Kaicheng Yu. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Tianheng Wang, DongbaiChen DongbaiChen, Linfeng Song, Yue Zhang 0031, Guozheng Rao, Kaicheng Yu |
ACL (1) | 7 |
| 2025 | Entropy Guided Extrapolative Decoding to Improve Factuality in Large Language ModelsabstractLarge language models (LLMs) exhibit impressive natural language capabilities but suffer from hallucination – generating content ungrounded in the realities of training data. Recent work has focused on decoding techniques to improve factuality in decoding by leveraging LLMs’ hierarchical representation of factual knowledge, manipulating the predicted distributions at inference time. Current state-of-the-art approaches refine decoding by contrasting logits from a lower layer with the final layer to exploit information related factuality within the model forward procedure. However, such methods often assume the final layer is most reliable one and the lower layer selection process depends on it. In this work, we first propose logit extrapolation of critical token probabilities beyond the last layer for more accurate contrasting. We additionally employ layer-wise entropy-guided lower layer selection, decoupling the selection process from the final layer. Experiments demonstrate strong performance - surpassing state-of-the-art on multiple different datasets by large margins. Analyses show different kinds of prompts respond to different selection strategies. Lifeng Jin, Linfeng Song, Haitao Mi, Baolin Peng, Dong Yu 0001 |
COLING | 3 |
| 2025 | Beam Structured Turbo Receiver for HF Skywave Massive MIMO CommunicationsabstractIn this paper, we investigate receiver design for high frequency (HF) skywave massive multiple-input multipleoutput (MIMO) communications. We first establish a modified beam based channel model by performing uniform sampling for directional cosine with deterministic sampling interval, where the beam matrix is constructed as discrete Fourier transform (DFT)based structure. Rooted in the modified beam based channel model, we propose a beam structured turbo receiver (BSTR) involving low-dimensional beam structured signal detection for grouped user terminals (UTs). Then we present efficient implementation of the BSTR by exploiting the structure properties of the beam matrix and the beam domain channel sparsity. Simulation results validate the superior performance of the proposed BSTR with low complexity. Linfeng Song, Ding Shi, Xiqi Gao 0001, Geoffrey Ye Li, Xiang-Gen Xia 0001 |
ICC | 1 |
| 2025 | Iterative Nash Policy Optimization: Aligning LLMs with General Preferences via No-Regret LearningabstractReinforcement Learning with Human Feedback (RLHF) has achieved great success
in aligning large language models (LLMs) with human preferences. Prevalent
RLHF approaches are reward-based, following the Bradley-Terry (BT) model assumption, which may not fully capture the complexity of human preferences. In
this paper, we explore RLHF under a general preference framework and approach
it from a game-theoretic perspective. Specifically, we formulate the problem as
a two-player game and propose a novel online algorithm, iterative Nash policy
optimization (INPO). The key idea is to let the policy play against itself via no-
regret learning, thereby approximating the Nash policy. Unlike previous methods,
INPO bypasses the need for estimating the expected win rate for individual responses, which typically incurs high computational or annotation costs. Instead,
we introduce a new loss objective that is directly minimized over a preference
dataset. We provide theoretical analysis for our approach and demonstrate its
effectiveness through experiments on various representative benchmarks. With an
LLaMA-3-8B-based SFT model, INPO achieves a 42.6% length-controlled win
rate on AlpacaEval 2.0 and a 37.8% win rate on Arena-Hard, showing substantial
improvement over the state-of-the-art online RLHF algorithms. Dian Yu 0001, Baolin Peng, Linfeng Song, Mingyue Huo, Nan Jiang 0008, Haitao Mi, Dong Yu 0001 |
ICLR | 4 |
| 2025 | Do NOT Think That Much for 2+3=? On the Overthinking of Long Reasoning ModelsabstractThe remarkable performance of long reasoning models can be attributed to their ability to emulate human-like long-time thinking during inference. These models employ extended chain-of-thought (CoT) processes, exploring multiple strategies to enhance problem-solving capabilities. However, a critical question remains: How to intelligently and efficiently scale computational resources during testing. This paper presents the first comprehensive study on the prevalent issue of overthinking in these models, where long reasoning models generate redundant solutions that contribute minimally to accuracy and diversity, thereby wasting computational resources on simple problems with minimal benefit. We introduce novel efficiency metrics from both outcome and process perspectives to evaluate the rational use of computational resources by long reasoning models. Using a self-training paradigm, we propose strategies to mitigate overthinking, simplifying reasoning processes without compromising accuracy. Experimental results show that our approach successfully reduces computational overhead while preserving model performance across a range of testsets with varying difficulty levels, such as GSM8K, MATH500, GPQA, and AIME. Our code is open-source and available at https://github.com/galaxyChen/overthinking. Zhiwei He 0002, Jianhui Pang, Dian Yu 0001, Linfeng Song, Qiuzhi Liu, Mengfei Zhou, Zhuosheng Zhang 0001, Rui Wang 0015, Zhaopeng Tu, Haitao Mi, Dong Yu 0001 |
ICML | 7 |
| 2025 | MPS-Prover: Advancing Stepwise Theorem Proving by Multi-Perspective Search and Data CurationabstractAutomated Theorem Proving (ATP) in formal languages remains a formidable challenge in AI, demanding rigorous logical deduction and navigating vast search spaces. While large language models (LLMs) have shown promising performance, existing stepwise provers often suffer from biased search guidance, leading to inefficiencies and suboptimal proof strategies. This paper introduces the Multi-Perspective Search Prover (MPS-Prover), a novel stepwise ATP system designed to overcome these limitations. MPS-Prover incorporates two key innovations: a highly effective post-training data curation strategy that prunes approximately 40\% of redundant training data without sacrificing performance, and a multi-perspective tree search mechanism. This search integrates a learned critic model with strategically designed heuristic rules to diversify tactic selection, prevent getting trapped in unproductive states, and enhance search robustness. Extensive evaluations demonstrate that MPS-Prover achieves state-of-the-art performance on multiple challenging benchmarks, including miniF2F and ProofNet, outperforming prior 7B parameter models. Furthermore, our analyses reveal that MPS-Prover generates significantly shorter and more diverse proofs compared to existing stepwise and whole-proof methods, highlighting its efficiency and efficacy. Our work advances the capabilities of LLM-based formal reasoning and offers a robust framework and a comprehensive analysis for developing more powerful theorem provers. Zhenwen Liang, Linfeng Song, Tao Yang 0033, Haitao Mi, Dong Yu 0001 |
NeurIPS | 2 |
| 2025 | Thoughts Are All Over the Place: On the Underthinking of Long Reasoning ModelsabstractLong reasoning models (LRMs) such as OpenAI's o1 and DeepSeek's R1 have demonstrated remarkable abilities in complex reasoning tasks by scaling test-time compute and exhibiting human-like deep thinking. However, we identify a phenomenon we term underthinking, where LRMs frequently switch between different reasoning thoughts without sufficiently exploring promising paths to reach a correct solution. This behavior leads to inadequate depth of reasoning and decreased performance, particularly on challenging mathematical problems. To systematically analyze this issue, we conduct experiments on three challenging test sets and two representative open-source LRMs, revealing that frequent thought switching correlates with incorrect responses. We introduce a novel metric to quantify underthinking by measuring token efficiency in incorrect answers. To address underthinking, we propose a decoding strategy with thought switching penalty (Tip) that discourages premature transitions between thoughts, encouraging deeper exploration of each reasoning path. Experimental results demonstrate that our approach improves accuracy across challenging datasets without requiring model fine-tuning. Our findings contribute to understanding reasoning inefficiencies in LRMs and offer a practical solution to enhance their problem-solving capabilities. Our code is open-source and available at https://github.com/wangyuenlp/underthinking. Yue Wang 0039, Qiuzhi Liu, Zhiwei He 0002, Linfeng Song, Dian Yu 0001, Juntao Li 0005, Zhuosheng Zhang 0001, Rui Wang 0015, Zhaopeng Tu, Haitao Mi, Dong Yu 0001 |
NeurIPS | 7 |
| 2025 | Improving LLM General Preference Alignment via Optimistic Online Mirror DescentabstractReinforcement learning from human feedback (RLHF) has demonstrated remarkable effectiveness in aligning large language models (LLMs) with human preferences. Many existing alignment approaches rely on the Bradley-Terry (BT) model assumption, which assumes the existence of a ground-truth reward for each prompt-response pair. However, this assumption can be overly restrictive when modeling complex human preferences. In this paper, we drop the BT model assumption and study LLM alignment under general preferences, formulated as a two-player game. Drawing on theoretical insights from learning in games, we integrate optimistic online mirror descent into our alignment framework to approximate the Nash policy. Theoretically, we demonstrate that our approach achieves an $\mathcal{O}(T^{-1})$ bound on the duality gap, improving upon the previous $\mathcal{O}(T^{-1/2})$ result. Meanwhile, it enjoys a linear convergence rate in the last iterate, a property not achieved by previous methods. More importantly, we implement our method and show through experiments that it outperforms state-of-the-art RLHF algorithms across multiple representative benchmarks. Dian Yu 0001, Tao Ge 0001, Linfeng Song, Zhichen Zeng 0001, Haitao Mi, Nan Jiang 0008, Dong Yu 0001 |
NeurIPS | 4 |
| 2025 | Mitigating the negative impact of over-association for conversational query production
Ante Wang, Linfeng Song, Zijun Min, Xiaoli Wang 0002, Junfeng Yao, Jinsong Su |
Inf. Process. Manag. | 2 |
| 2025 | Beam Structured Precoder for HF Skywave Massive MIMO-OFDM Communications With Channel Smoothness ConstraintabstractIn this paper, we investigate precoder design for high frequency (HF) skywave massive multiple-input multiple-output (MIMO) communications with orthogonal frequency division multiplexing (OFDM) modulation. We first reveal the effect of the precoder on the effective channel at receivers and formulate the precoder design for a group of subcarriers as a sum-rate maximization problem, where the delay spread of the effective channel is constrained to maintain its smoothness. Then with the beam based channel model and beam domain channel sparsity, the design of space domain precoders for a group of subcarriers are transformed into that of a space-frequency (SF) beam domain vector and the resulting space domain precoder at each subcarrier is beam structured. Efficient calculation for design and implementation of the beam structured precoder (BSP) is proposed. Moreover, effective channel estimation with the BSP is discussed. Simulation results show that the proposed BSP can enhance the effective channel estimation performance and significantly improve the system performance. Ding Shi, Linfeng Song, Xuzhong Zhang, Xiqi Gao 0001, Jiaheng Wang 0001, Xiaohu You 0001, Geoffrey Ye Li, Xiang-Gen Xia 0001 |
IEEE Trans. Commun. | 2 |
| 2024 | Response Enhanced Semi-supervised Dialogue Query GenerationabstractLeveraging vast and continually updated knowledge from the Internet has been considered an important ability for a dialogue system. Therefore, the dialogue query generation task is proposed for generating search queries from dialogue histories, which will be submitted to a search engine for retrieving relevant websites on the Internet. In this regard, previous efforts were devoted to collecting conversations with annotated queries and training a query producer (QP) via standard supervised learning. However, these studies still face the challenges of data scarcity and domain adaptation. To address these issues, in this paper, we propose a semi-supervised learning framework -- SemiDQG, to improve model performance with unlabeled conversations. Based on the observation that the search query is typically related to the topic of dialogue response, we train a response-augmented query producer (RA) to provide rich and effective training signals for QP. We first apply a similarity-based query selection strategy to select high-quality RA-generated pseudo queries, which are used to construct pseudo instances for training QP and RA. Then, we adopt the REINFORCE algorithm to further enhance QP, with RA-provided rewards as fine-grained training signals. Experimental results and in-depth analysis of three benchmarks show the effectiveness of our framework in cross-domain and low-resource scenarios. Particularly, SemiDQG significantly surpasses ChatGPT and competitive baselines. Our code is available at \url{https://github.com/DeepLearnXMU/SemiDQG}. Jianheng Huang, Ante Wang, Linfeng Gao, Linfeng Song, Jinsong Su |
AAAI | 4 |
| 2024 | Mitigating Catastrophic Forgetting in Large Language Models with Self-Synthesized RehearsalabstractJianheng Huang, Leyang Cui, Ante Wang, Chengyi Yang, Xinting Liao, Linfeng Song, Junfeng Yao, Jinsong Su. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Jianheng Huang, Leyang Cui, Ante Wang, Xinting Liao, Linfeng Song, Junfeng Yao, Jinsong Su |
ACL (1) | 6 |
| 2024 | Self-Alignment for Factuality: Mitigating Hallucinations in LLMs via Self-EvaluationabstractXiaoying Zhang, Baolin Peng, Ye Tian, Jingyan Zhou, Lifeng Jin, Linfeng Song, Haitao Mi, Helen Meng. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Baolin Peng, Jingyan Zhou, Lifeng Jin, Linfeng Song, Haitao Mi, Helen M. Meng |
ACL (1) | 6 |
| 2024 | A Knowledge Plug-and-Play Test Bed for Open-domain Dialogue GenerationabstractKnowledge-based, open-domain dialogue generation aims to build chit-chat systems that talk to humans using mined support knowledge. Many types and sources of knowledge have previously been shown to be useful as support knowledge. Even in the era of large language models, response generation grounded in knowledge retrieved from additional up-to-date sources remains a practically important approach. While prior work using single-source knowledge has shown a clear positive correlation between the performances of knowledge selection and response generation, there are no existing multi-source datasets for evaluating support knowledge retrieval. Further, prior work has assumed that the knowledge sources available at test time are the same as during training. This unrealistic assumption unnecessarily handicaps models, as new knowledge sources can become available after a model is trained. In this paper, we present a high-quality benchmark named multi-source Wizard of Wikipedia (Ms.WoW) for evaluating multi-source dialogue knowledge selection and response generation. Unlike existing datasets, it contains clean support knowledge, grounded at the utterance level and partitioned into multiple knowledge sources. We further propose a new challenge, dialogue knowledge plug-and-play, which aims to test an already trained dialogue model on using new support knowledge from previously unseen sources in a zero-shot fashion. Xiangci Li, Linfeng Song, Lifeng Jin, Haitao Mi, Jessica Ouyang 0001, Dong Yu 0001 |
LREC/COLING | 2 |
| 2024 | The Trickle-down Impact of Reward Inconsistency on RLHFabstractStandard practice within Reinforcement Learning from Human Feedback (RLHF) involves optimizing against a Reward Model (RM), which itself is trained to reflect human preferences for desirable generations. A notable subject that is understudied is the (in-)consistency of RMs --- whether they can recognize the semantic changes to different prompts and
appropriately adapt their reward assignments
--- and their impact on the downstream RLHF model.
In this paper, we visit a series of research questions relevant to RM inconsistency:
(1) How can we measure the consistency of reward models?
(2) How consistent are the existing RMs and how can we improve them?
(3) In what ways does reward inconsistency influence the chatbots resulting from the RLHF model training?
We propose **Contrast Instruction** -- a benchmarking strategy for the consistency of RM.
Each example in **Contrast Instruction** features a pair of lexically similar instructions with different ground truth responses. A consistent RM is expected to rank the corresponding instruction and response higher than other combinations. We observe that current RMs trained with the standard ranking objective fail miserably on \contrast{} compared to average humans. To show that RM consistency can be improved efficiently without using extra training budget, we propose two techniques **ConvexDA** and **RewardFusion**, which enhance reward consistency
through extrapolation during the RM training and inference stage, respectively.
We show that RLHF models trained with a more consistent RM yield more useful responses, suggesting that reward inconsistency exhibits a trickle-down effect on the downstream RLHF process. Lingfeng Shen, Linfeng Song, Lifeng Jin, Baolin Peng, Haitao Mi, Daniel Khashabi, Dong Yu 0001 |
ICLR | 3 |
| 2024 | Toward Self-Improvement of LLMs via Imagination, Searching, and CriticizingabstractDespite the impressive capabilities of Large Language Models (LLMs) on various tasks, they still struggle with scenarios that involves complex reasoning and planning. Self-correction and self-learning emerge as viable solutions, employing strategies that allow LLMs to refine their outputs and learn from self-assessed rewards. Yet, the efficacy of LLMs in self-refining its response, particularly in complex reasoning and planning task, remains dubious. In this paper, we introduce AlphaLLM for the self-improvements of LLMs, which integrates Monte Carlo Tree Search (MCTS) with LLMs to establish a self-improving loop, thereby enhancing the capabilities of LLMs without additional annotations. Drawing inspiration from the success of AlphaGo, AlphaLLM addresses the unique challenges of combining MCTS with LLM for self-improvement, including data scarcity, the vastness search spaces of language tasks, and the subjective nature of feedback in language tasks. AlphaLLM is comprised of prompt synthesis component, an efficient MCTS approach tailored for language tasks, and a trio of critic models for precise feedback. Our experimental results in mathematical reasoning tasks demonstrate that AlphaLLM significantly enhances the performance of LLMs without additional annotations, showing the potential for self-improvement in LLMs. Baolin Peng, Linfeng Song, Lifeng Jin, Dian Yu 0001, Haitao Mi, Dong Yu 0001 |
NeurIPS | 3 |
| 2024 | Robust Precoding for HF Skywave Massive MIMO With Slepian TransformabstractIn this paper, we address robust precoding in high-frequency (HF) skywave massive multiple-input multiple-output (MIMO) systems with imperfect channel state information (CSI). We first employ a sparse beam baseda posteriorichannel model and demonstrate that robust precoding can be efficiently solved in the Slepian transform domain with a large number of base station (BS) antennas. Next, we introduce two Slepian transform based robust precoding methods, including a joint approach that leverages inverse fast Fourier transform (IFFT) for reduced complexity with a large number of user terminals (UTs). We then establish a local optimum for the Slepian transform domain robust precoder (STRP) design using the majorization minimization (MM) algorithm, taking advantages of HF skywave massive MIMO channel sparsity and Slepian sequence properties. Further, two distinct designs are presented: separate STRP (SSTRP) and joint STRP (JSTRP). Simulation results confirm the effectiveness of proposed robust precoders, showcasing their excellent ergodic sum-rate performance and low complexity. Linfeng Song, Ding Shi, Lu Gan 0002, Xiqi Gao 0001 |
IEEE Trans. Commun. | 1 |
| 2024 | Beam Structured Signal Detector for HF Skywave Massive MIMO-OFDM CommunicationsabstractIn this paper, we investigate signal detection for HF skywave massive multiple-input multiple-output (MIMO) communications with orthogonal frequency division multiplexing (OFDM) modulation. We first introduce beam based channel models (BBCM) in the space domain at each subcarrier and in the space-frequency domain for all subcarriers. Based on the BBCM in the space domain, we propose a beam structured detector (BSD) for each subcarrier. Specifically, we prove that the space domain detector design can be transformed into that of a beam domain detector without sacrificing optimality, and the asymptotically optimal space domain detector is beam structured with a low-dimensional beam domain detector, thus significantly reducing the design and implementation complexities. Furthermore, we extend the BSD to the space-frequency domain based on the BBCM jointly for all subcarriers. The design of space-frequency domain detector is also converted to that of a low-dimensional beam domain detector, which enables a very efficient design and implementation of BSD. Simulation results demonstrate the low complexity and satisfactory performance of the proposed detectors. Ding Shi, Linfeng Song, Xiqi Gao 0001, Jiaheng Wang 0001, Mats Bengtsson, Geoffrey Ye Li |
IEEE Trans. Wirel. Commun. | 2 |
| 2024 | Beam Structured Channel Estimation for HF Skywave Massive MIMO-OFDM CommunicationsabstractIn this paper, we investigate high frequency (HF) skywave massive multiple-input multiple-output (MIMO) communications with orthogonal frequency division multiplexing (OFDM) modulation. Based on the triple-beam (TB) based channel model and the channel sparsity in the TB domain, we propose a beam structured channel estimation (BSCE) approach. Specifically, we show that the space-frequency-time (SFT) domain estimator design for each TB domain channel element can be transformed into that of a low-dimensional TB domain estimator and the resulting SFT domain estimator is beam structured. We also present a method to select the TBs used for BSCE. Then we generalize the proposed BSCE by introducing window functions and a turbo principle to achieve a superior trade-off between complexity and performance. Furthermore, we present a low-complexity design and implementation of BSCE by exploiting the characteristics of the TB matrix. Simulation results validate the proposed theory and methods. Ding Shi, Linfeng Song, Xiqi Gao 0001, Jiaheng Wang 0001, Mats Bengtsson, Geoffrey Ye Li, Xiang-Gen Xia 0001 |
IEEE Trans. Wirel. Commun. | 2 |
| 2023 | A Survey on Zero Pronoun TranslationabstractZero pronouns (ZPs) are frequently omitted in pro-drop languages (e.g.Chinese, Hungarian, and Hindi), but should be recalled in nonpro-drop languages (e.g.English).This phenomenon has been studied extensively in machine translation (MT), as it poses a significant challenge for MT systems due to the difficulty in determining the correct antecedent for the pronoun.This survey paper highlights the major works that have been undertaken in zero pronoun translation (ZPT) after the neural revolution so that researchers can recognize the current state and future directions of this field.We provide an organization of the literature based on evolution, dataset, method, and evaluation.In addition, we compare and analyze competing models and evaluation metrics on different benchmarks.We uncover a number of insightful findings such as: 1) ZPT is in line with the development trend of large language model; 2) data limitation causes learning bias in languages and domains; 3) performance improvements are often reported on single benchmarks, but advanced methods are still far from realworld use; 4) general-purpose metrics are not reliable on nuances and complexities of ZPT, emphasizing the necessity of targeted metrics; 5) apart from commonly-cited errors, ZPs will cause risks of gender bias. Longyue Wang, Siyou Liu, Mingzhou Xu, Linfeng Song, Shuming Shi 0001, Zhaopeng Tu |
ACL (1) | 4 |
| 2023 | SafeConv: Explaining and Correcting Conversational Unsafe BehaviorabstractOne of the main challenges open-domain endto-end dialogue systems, or chatbots, face is the prevalence of unsafe behavior, such as toxic languages and harmful suggestions.However, existing dialogue datasets do not provide enough annotation to explain and correct such unsafe behavior.In this work, we construct a new dataset called SAFECONV for the research of conversational safety: (1) Besides the utterancelevel safety labels, SAFECONV also provides unsafe spans in an utterance, information able to indicate which words contribute to the detected unsafe behavior; (2) SAFECONV provides safe alternative responses to continue the conversation when unsafe behavior detected, guiding the conversation to a gentle trajectory.By virtue of the comprehensive annotation of SAFECONV, we benchmark three powerful models for the mitigation of conversational unsafe behavior, including a checker to detect unsafe utterances, a tagger to extract unsafe spans, and a rewriter to convert an unsafe response to a safe version.Moreover, we explore the huge benefits brought by combining the models for explaining the emergence of unsafe behavior and detoxifying chatbots.Experiments show that the detected unsafe behavior could be well explained with unsafe spans and popular chatbots could be detoxified by a huge extent.The dataset is available at https://github.com/mianzhang/SafeConv. Lifeng Jin, Linfeng Song, Haitao Mi, Wenliang Chen, Dong Yu 0001 |
ACL (1) | 3 |
| 2023 | Friend-training: Learning from Models of Different but Related TasksabstractCurrent self-training methods such as standard self-training, co-training, tri-training, and others often focus on improving model performance on a single task, utilizing differences in input features, model architectures, and training processes.However, many tasks in natural language processing are about different but related aspects of language, and models trained for one task can be great teachers for other related tasks.In this work, we propose friendtraining, a cross-task self-training framework, where models trained to do different tasks are used in an iterative training, pseudo-labeling, and retraining process to help each other for better selection of pseudo-labels.With two dialogue understanding tasks, conversational semantic role labeling and dialogue rewriting, chosen for a case study, we show that the models trained with the friend-training framework achieve the best performance compared to strong baselines. Lifeng Jin, Linfeng Song, Haitao Mi, Xiabing Zhou, Dong Yu 0001 |
EACL | 3 |
| 2023 | An Empirical Study of Retrieval-Enhanced Graph Neural NetworksabstractGraph Neural Networks (GNNs) are effective tools for graph representation learning. Most GNNs rely on a recursive neighborhood aggregation scheme, named message passing, thereby their theoretical expressive power is limited to the first-order Weisfeiler-Lehman test (1-WL). An effective approach to this challenge is to explicitly retrieve some annotated examples used to enhance GNN models. While retrieval-enhanced models have been proved to be effective in many language and vision domains, it remains an open question how effective retrieval-enhanced GNNs are when applied to graph datasets. Motivated by this, we want to explore how the retrieval idea can help augment the useful information learned in the graph neural networks, and we design a retrieval-enhanced scheme called GRAPHRETRIEVAL, which is agnostic to the choice of graph neural network models. In GRAPHRETRIEVAL, for each input graph, similar graphs together with their ground-true labels are retrieved from an existing database. Thus they can act as a potential enhancement to complete various graph property predictive tasks. We conduct comprehensive experiments over 13 datasets, and we observe that GRAPHRETRIEVAL is able to reach substantial improvements over existing GNNs. Moreover, our empirical study also illustrates that retrieval enhancement is a promising remedy for alleviating the long-tailed label distribution problem. Dingmin Wang, Shengchao Liu, Hanchen Wang 0002, Bernardo Cuenca Grau, Linfeng Song, Jian Tang 0005, Qi Liu 0049 |
ECAI | 5 |
| 2023 | Tree based Progressive Regression Model for Watch-Time Prediction in Short-video RecommendationabstractAn accurate prediction of watch time has been of vital importance to enhance user engagement in video recommender systems. To achieve this, there are four properties that a watch time prediction framework should satisfy: first, despite its continuous value, watch time is also an ordinal variable and the relative ordering between its values reflects the differences in user preferences. Therefore the ordinal relations should be reflected in watch time predictions. Second, the conditional dependence between the video-watching behaviors should be captured in the model. For instance, one has to watch half of the video before he/she finishes watching the whole video. Third, modeling watch time with a point estimation ignores the fact that models might give results with high uncertainty and this could cause bad cases in recommender systems. Therefore the framework should be aware of prediction uncertainty. Forth, the real-life recommender systems suffer from severe bias amplifications thus an estimation without bias amplification is expected. Xiao Lin 0002, Xiaokai Chen, Linfeng Song, Biao Li 0002, Peng Jiang 0002 |
KDD | 3 |
| 2023 | Beam Structured Signal Detection for HF Skywave Massive MIMO CommunicationsabstractIn this paper, we investigate signal detection for HF skywave massive multiple-input multiple-output (MIMO) communications with orthogonal frequency division multiplexing (OFDM) modulation. We first introduce beam based channel model (BBCM) in the space domain and reveal the sparsity of the channel in the space-beam domain. Based on the BBCM in the space domain, we propose a beam structured detector (BSD) for each subcarrier. Specifically, we prove that the space domain detector design can be transformed to that of a beam domain detector without sacrificing optimality, and the asymptotically optimal space domain detector is beam structured with a low-dimensional beam domain detector, thus significantly reducing the design and implementation complexities. Furthermore, we provide a beam selection criterion to choose the beams that are used for the BSD. Simulation results demonstrate the low complexity and satisfactory performance of the proposed detector. Ding Shi, Linfeng Song, Xiqi Gao 0001, Jiaheng Wang 0001, Mats Bengtsson, Geoffrey Ye Li |
VTC Fall | 2 |
| 2023 | Search-engine-augmented dialogue response generation with cheaply supervised query production
Ante Wang, Linfeng Song, Qi Liu 0049, Haitao Mi, Longyue Wang, Zhaopeng Tu, Jinsong Su, Dong Yu 0001 |
Artif. Intell. | 2 |
| 2023 | Discover, Explain, Improve: An Automatic Slice Detection Benchmark for Natural Language ProcessingabstractAbstract Pretrained natural language processing (NLP) models have achieved high overall performance, but they still make systematic errors. Instead of manual error analysis, research on slice detection models (SDMs), which automatically identify underperforming groups of datapoints, has caught escalated attention in Computer Vision for both understanding model behaviors and providing insights for future model training and designing. However, little research on SDMs and quantitative evaluation of their effectiveness have been conducted on NLP tasks. Our paper fills the gap by proposing a benchmark named “Discover, Explain, Improve (DEIm)” for classification NLP tasks along with a new SDM Edisa. Edisa discovers coherent and underperforming groups of datapoints; DEIm then unites them under human-understandable concepts and provides comprehensive evaluation tasks and corresponding quantitative metrics. The evaluation in DEIm shows that Edisa can accurately select error-prone datapoints with informative semantic features that summarize error patterns. Detecting difficult datapoints directly boosts model performance without tuning any original model parameters, showing that discovered slices are actionable for users.1 Wenyue Hua, Lifeng Jin, Linfeng Song, Haitao Mi, Dong Yu 0001 |
Trans. Assoc. Comput. Linguistics | 3 |
| 2023 | OpenFact: Factuality Enhanced Open Knowledge ExtractionabstractAbstract We focus on the factuality property during the extraction of an OpenIE corpus named OpenFact, which contains more than 12 million high-quality knowledge triplets. We break down the factuality property into two important aspects—expressiveness and groundedness—and we propose a comprehensive framework to handle both aspects. To enhance expressiveness, we formulate each knowledge piece in OpenFact based on a semantic frame. We also design templates, extra constraints, and adopt human efforts so that most OpenFact triplets contain enough details. For groundedness, we require the main arguments of each triplet to contain linked Wikidata1 entities. A human evaluation suggests that the OpenFact triplets are much more accurate and contain denser information compared to OPIEC-Linked (Gashteovski et al., 2019), one recent high-quality OpenIE corpus grounded to Wikidata. Further experiments on knowledge base completion and knowledge base question answering show the effectiveness of OpenFact over OPIEC-Linked as supplementary knowledge to Wikidata as the major KG. Linfeng Song, Ante Wang, Xiaoman Pan, Hongming Zhang 0009, Dian Yu 0001, Lifeng Jin, Haitao Mi, Jinsong Su, Yue Zhang 0004, Dong Yu 0001 |
Trans. Assoc. Comput. Linguistics | 1 |
| 2023 | D$^{2}$PSG: Multi-Party Dialogue Discourse Parsing as Sequence GenerationabstractConversational discourse analysis aims to extract the interactions between dialogue turns, which is crucial for modeling complex multi-party dialogues. As the benchmarks are still limited in size and human annotations are costly, the current standard approaches apply pretrained language models, but they still require randomly initialized classifiers to make predictions. These classifiers usually require massive data to work smoothly with the pretrained encoder, causing severe data hunger issue. We propose two convenient strategies to formulate this task as a sequence generation problem, where classifier decisions are carefully converted into sequence of tokens. We then adopt a pretrained T5 1 model to solve this task so that no parameters are randomly initialized. We also leverage the descriptions of the discourse relations to help model understand their meanings. Experiments on two popular benchmarks show that our approach outperforms previous state-of-the-art models by a large margin, and it is also more robust in zero-shot and few-shot settings. Ante Wang, Linfeng Song, Lifeng Jin, Junfeng Yao, Haitao Mi, Chen Lin 0001, Jinsong Su, Dong Yu 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2023 | Channel Acquisition for HF Skywave Massive MIMO-OFDM CommunicationsabstractIn this paper, we investigate channel acquisition for high frequency (HF) skywave massive multiple-input multiple-output (MIMO) communications with orthogonal frequency division multiplexing (OFDM) modulation. We first introduce the concept of triple beams (TBs) in the space-frequency-time (SFT) domain and establish a TB based channel model using sampled triple steering vectors. With the established channel model, we then investigate the optimal channel estimation and pilot design for pilot segments. Specifically, we find the conditions that allow pilot reuse among multiple user terminals (UTs), which significantly reduces pilot overhead and increases the number of UTs that can be served. Moreover, we propose a channel prediction method for data segments based on the estimated TB domain channel. To reduce the complexity, we formulate the channel estimation as a statistical inference problem and then obtain the channel by the proposed constrained Bethe free energy minimization (CBFEM) based channel estimation algorithm, which can be implemented with low complexity by exploiting the structure of the TB matrix together with the chirp z-transform (CZT). Simulation results demonstrate the superior performance of the proposed channel acquisition approach. Ding Shi, Linfeng Song, Xiqi Gao 0001, Cheng-Xiang Wang 0001, Geoffrey Ye Li |
IEEE Trans. Wirel. Commun. | 2 |
| 2022 | Hierarchical Context Tagging for Utterance RewritingabstractUtterance rewriting aims to recover coreferences and omitted information from the latest turn of a multi-turn dialogue. Recently, methods that tag rather than linearly generate sequences have proven stronger in both in- and out-of-domain rewriting settings. This is due to a tagger's smaller search space as it can only copy tokens from the dialogue context. However, these methods may suffer from low coverage when phrases that must be added to a source utterance cannot be covered by a single context span. This can occur in languages like English that introduce tokens such as prepositions into the rewrite for grammaticality. We propose a hierarchical context tagger (HCT) that mitigates this issue by predicting slotted rules (e.g., "besides _") whose slots are later filled with context spans. HCT (i) tags the source string with token-level edit actions and slotted rules and (ii) fills in the resulting rule slots with spans from the dialogue context. This rule tagging allows HCT to add out-of-context tokens and multiple spans at once; we further cluster the rules to truncate the long tail of the rule distribution. Experiments on several benchmarks show that HCT can outperform state-of-the-art rewriting systems by ~2 BLEU points. Lisa Jin, Linfeng Song, Lifeng Jin, Dong Yu 0001, Daniel Gildea |
AAAI | 2 |
| 2022 | Variational Graph Autoencoding as Cheap Supervision for AMR Coreference ResolutionabstractCoreference resolution over semantic graphs like AMRs aims to group the graph nodes that represent the same entity.This is a crucial step for making document-level formal semantic representations.With annotated data on AMR coreference resolution, deep learning approaches have recently shown great potential for this task, yet they are usually data hungry and annotating data is costly.We propose a general pretraining method using variational graph autoencoder (VGAE) for AMR coreference resolution, which can leverage any general AMR corpus and even automatically parsed AMR data.Experiments on benchmarks show that the pretraining approach achieves performance gains of up to 6% absolute F1 points.Moreover, our model significantly improves on the previous state-of-theart model by up to 11% F1 points. Irene Li, Linfeng Song, Kun Xu 0005, Dong Yu 0001 |
ACL (1) | 2 |
| 2022 | Semantic-based Pre-training for Dialogue UnderstandingabstractPre-trained language models have made great progress on dialogue tasks. However, these models are typically trained on surface dialogue text, thus are proven to be weak in understanding the main semantic meaning of a dialogue context. We investigate Abstract Meaning Representation (AMR) as explicit semantic knowledge for pre-training models to capture the core semantic information in dialogues during pre-training. In particular, we propose a semantic-based pre-training framework that extends the standard pre-training framework (Devlin et al.,2019) by three tasks for learning 1) core semantic units, 2) semantic relations and 3) the overall semantic representation according to AMR graphs. Experiments on the understanding of both chit-chats and task-oriented dialogues show the superiority of our model. To our knowledge, we are the first to leverage a deep semantic representation for dialogue pre-training. Xuefeng Bai 0001, Linfeng Song, Yue Zhang 0004 |
COLING | 2 |
| 2022 | Cross-domain Generalization for AMR ParsingabstractMeaning Representation (AMR) parsing aims to predict an AMR graph from textual input.Recently, there has been notable growth in AMR parsing performance.However, most existing work focuses on improving the performance in the specific domain, ignoring the potential domain dependence of AMR parsing systems.To address this, we extensively evaluate five representative AMR parsers on five domains and analyze challenges to cross-domain AMR parsing.We observe that challenges to cross-domain AMR parsing mainly arise from the distribution shift of words and AMR concepts.Based on our observation, we investigate two approaches to reduce the domain distribution divergence of text and AMR features, respectively.Experimental results on two out-of-domain test sets show the superiority of our method. Xuefeng Bai 0001, Sen Yang 0005, Leyang Cui, Linfeng Song, Yue Zhang 0004 |
EMNLP | 4 |
| 2022 | GuoFeng: A Benchmark for Zero Pronoun Recovery and TranslationabstractMingzhou Xu, Longyue Wang, Derek F. Wong, Hongye Liu, Linfeng Song, Lidia S. Chao, Shuming Shi, Zhaopeng Tu. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. 2022. Mingzhou Xu, Longyue Wang, Derek F. Wong, Hongye Liu, Linfeng Song, Lidia S. Chao, Shuming Shi 0001, Zhaopeng Tu |
EMNLP | 5 |
| 2022 | Learning a Grammar Inducer from Massive Uncurated Instructional VideosabstractVideo-aided grammar induction aims to leverage video information for finding more accurate syntactic grammars for accompanying text.While previous work focuses on building systems for inducing grammars on text that are well-aligned with video content, we investigate the scenario, in which text and video are only in loose correspondence.Such data can be found in abundance online, and the weak correspondence is similar to the indeterminacy problem studied in language acquisition.Furthermore, we build a new model that can better learn video-span correlation without manually designed features adopted by previous work.Experiments show that our model trained only on large-scale YouTube data with no textvideo alignment reports strong and robust performances across three unseen datasets, despite domain shift and noisy label issues.Furthermore our model yields higher F1 scores than the previous state-of-the-art systems trained on in-domain data. Songyang Zhang 0004, Linfeng Song, Lifeng Jin, Haitao Mi, Kun Xu 0005, Dong Yu 0001, Jiebo Luo 0001 |
EMNLP | 2 |
| 2022 | Channel Estimation for HF Skywave Massive MIMO-OFDM with Triple-Beam Based Channel ModelabstractIn this paper, we investigate channel estimation for high frequency (HF) skywave massive multiple-input multiple-output (MIMO) communications with orthogonal frequency division multiplexing (OFDM) modulation. We first introduce the concept of triple beams (TBs) in the space-frequency-time (SFT) domain and establish a TB based channel model using sampled triple steering vectors. With the established channel model, we then investigate the optimal channel estimation and pilot design for pilot segments. Specifically, we find the conditions that allow pilot reuse among multiple user terminals (UTs), which significantly reduces pilot overhead. To reduce the complexity, we are able to formulate the channel estimation as a sparse signal recovery problem due to the channel sparsity in the TB domain and then obtain the channel by the proposed constrained Bethe free energy minimization (CBFEM) based channel estimation algorithm. Simulation results demonstrate the superior performance of the proposed channel estimation approach. Ding Shi, Linfeng Song, Xiqi Gao 0001, Cheng-Xiang Wang 0001, Geoffrey Ye Li |
GLOBECOM | 2 |
| 2022 | CASA: Conversational Aspect Sentiment Analysis for Dialogue UnderstandingabstractDialogue understanding has always been a bottleneck for many conversational tasks, such as dialogue response generation and conversational question answering. To expedite the progress in this area, we introduce the task of conversational aspect sentiment analysis (CASA) that can provide useful fine-grained sentiment information for dialogue understanding and planning. Overall, this task extends the standard aspect-based sentiment analysis to the conversational scenario with several major adaptations. To aid the training and evaluation of data-driven methods, we annotate 3,000 chit-chat dialogues (27,198 sentences) with fine-grained sentiment information, including all sentiment expressions, their polarities and the corresponding target mentions. We also annotate an out-of-domain test set of 200 dialogues for robustness evaluation. Besides, we develop multiple baselines based on either pretrained BERT or self-attention for preliminary study. Experimental results show that our BERT-based model has strong performances for both in-domain and out-of-domain datasets, and thorough analysis indicates several potential directions for further improvements. Linfeng Song, Chunlei Xin, Shaopeng Lai, Ante Wang, Jinsong Su, Kun Xu 0005 |
J. Artif. Intell. Res. | 1 |
| 2022 | A Multi-Level Optimization Framework for End-to-End Text AugmentationabstractAbstract Text augmentation is an effective technique in alleviating overfitting in NLP tasks. In existing methods, text augmentation and downstream tasks are mostly performed separately. As a result, the augmented texts may not be optimal to train the downstream model. To address this problem, we propose a three-level optimization framework to perform text augmentation and the downstream task end-to- end. The augmentation model is trained in a way tailored to the downstream task. Our framework consists of three learning stages. A text summarization model is trained to perform data augmentation at the first stage. Each summarization example is associated with a weight to account for its domain difference with the text classification data. At the second stage, we use the model trained at the first stage to perform text augmentation and train a text classification model on the augmented texts. At the third stage, we evaluate the text classification model trained at the second stage and update weights of summarization examples by minimizing the validation loss. These three stages are performed end-to-end. We evaluate our method on several text classification datasets where the results demonstrate the effectiveness of our method. Code is available at https://github.com/Sai-Ashish/End-to-End-Text-Augmentation. Sai Ashish Somayajula, Linfeng Song, Pengtao Xie |
Trans. Assoc. Comput. Linguistics | 2 |
| 2022 | An AST Structure Enhanced Decoder for Code GenerationabstractCurrently, the most dominant neural code generation modelsare often equipped with a tree-structured LSTM decoder, which outputs a sequence of actions to construct an Abstract Syntax Tree (AST) via pre-order traversal. However, such a decoder has two obvious drawbacks. First, except for the parent action, other faraway and important history actions rarely contribute to the current decision. Second, it also neglects future actions, which may be crucial for the prediction of the current action. To deal with these issues, in this paper, we propose a novel AST structure enhanced decoder for code generation, which significantly extends the decoder with respect to the above two aspects. First, we introduce an AST information enhanced attention mechanism to fully exploit history actions, of which impacts are further distinguished according to their syntactic distances, action types and relative positions; Second, we jointly model the predictions of current action and its important future action via multi-task learning, where the learned hidden state of the latter can be further leveraged to improve the former. Experimental results on commonly-used datasets demonstrate the effectiveness of our proposed decoder.1 Linfeng Song, Yubin Ge, Fandong Meng, Junfeng Yao, Jinsong Su |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2022 | Evidence Integration for Multi-Hop Reading Comprehension With Graph Neural NetworksabstractMulti-hop reading comprehension focuses on one type of factoid question, where a system needs to properly integrate multiple pieces of evidence to correctly answer a question. Previous work approximates global evidence with local coreference information, encoding coreference chains with DAG-styled GRU layers within a gated-attention reader. However, coreference is limited in providing information for rich inference. We introduce a new method for better connecting global evidence, which forms more complex graphs compared to DAGs. To perform evidence integration on our graphs, we investigate two recent graph neural networks, namely graph convolutional network (GCN) and graph recurrent network (GRN). Experiments on two standard datasets show that richer global information leads to better answers. Our approach shows highly competitive performances on these datasets without deep language models (such as ELMo). Linfeng Song, Zhiguo Wang 0006, Mo Yu, Yue Zhang 0004, Radu Florian, Daniel Gildea |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2021 | Semantic Representation for Dialogue ModelingabstractXuefeng Bai, Yulong Chen, Linfeng Song, Yue Zhang. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Xuefeng Bai 0001, Yulong Chen 0001, Linfeng Song, Yue Zhang 0004 |
ACL/IJCNLP (1) | 3 |
| 2021 | End-to-End AMR Corefencence ResolutionabstractQiankun Fu, Linfeng Song, Wenyu Du, Yue Zhang. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Qiankun Fu, Linfeng Song, Wenyu Du, Yue Zhang 0004 |
ACL/IJCNLP (1) | 2 |
| 2021 | RAST: Domain-Robust Dialogue Rewriting as Sequence TaggingabstractThe task of dialogue rewriting aims to reconstruct the latest dialogue utterance by copying the missing content from the dialogue context.Until now, the existing models for this task suffer from the robustness issue, i.e., performances drop dramatically when testing on a different dataset.We address this robustness issue by proposing a novel sequence-taggingbased model so that the search space is significantly reduced, yet the core of this task is still well covered.As a common issue of most tagging models for text generation, the model's outputs may lack fluency.To alleviate this issue, we inject the loss signal from BLEU or GPT-2 under a REINFORCE framework.Experiments show huge improvements of our model over the current state-of-the-art systems when transferring to another dataset. Linfeng Song, Liwei Wang 0009, Kun Xu 0005, Zhaopeng Tu, Dong Yu 0001 |
EMNLP (1) | 2 |
| 2021 | Instance-adaptive training with noise-robust losses against noisy labelsabstractIn order to alleviate the huge demand for annotated datasets for different tasks, many recent natural language processing datasets have adopted automated pipelines for fast-tracking usable data.However, model training with such datasets poses a challenge because popular optimization objectives are not robust to label noise induced in the annotation generation process.Several noise-robust losses have been proposed and evaluated on tasks in computer vision, but they generally use a single dataset-wise hyperparamter to control the strength of noise resistance.This work proposes novel instance-adaptive training frameworks to change dataset-wise hyperparameters of noise resistance in such losses to be instance-specific.Such instance-specific noise resistance hyperparameters are predicted by special instance-level label quality predictors, which are trained along with the main models.Experiments on noisy and corrupted NLP datasets show that proposed instance-adaptive training frameworks help increase the noiserobustness provided by such losses, promoting the use of the frameworks and associated losses in training NLP models with noisy data. Lifeng Jin, Linfeng Song, Kun Xu 0005, Dong Yu 0001 |
EMNLP (1) | 2 |
| 2021 | A Structure Self-Aware Model for Discourse Parsing on Multi-Party DialoguesabstractConversational discourse structures aim to describe how a dialogue is organized, thus they are helpful for dialogue understanding and response generation. This paper focuses on predicting discourse dependency structures for multi-party dialogues. Previous work adopts incremental methods that take the features from the already predicted discourse relations to help generate the next one. Although the inter-correlations among predictions considered, we find that the error propagation is also very serious and hurts the overall performance. To alleviate error propagation, we propose a Structure Self-Aware (SSA) model, which adopts a novel edge-centric Graph Neural Network (GNN) to update the information between each Elementary Discourse Unit (EDU) pair layer by layer, so that expressive representations can be learned without historical predictions. In addition, we take auxiliary training signals (e.g. structure distillation) for better representation learning. Our model achieves the new state-of-the-art performances on two conversational discourse parsing benchmarks, largely outperforming the previous methods. Ante Wang, Linfeng Song, Shaopeng Lai, Junfeng Yao, Jinsong Su |
IJCAI | 2 |
| 2021 | Video-aided Unsupervised Grammar InductionabstractSongyang Zhang, Linfeng Song, Lifeng Jin, Kun Xu, Dong Yu, Jiebo Luo. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Songyang Zhang 0004, Linfeng Song, Lifeng Jin, Kun Xu 0005, Dong Yu 0001, Jiebo Luo 0001 |
NAACL-HLT | 2 |
| 2021 | Distant Finetuning with Discourse Relations for Stance Classification
Lifeng Jin, Kun Xu 0005, Linfeng Song, Dong Yu 0001 |
NLPCC (2) | 3 |
| 2021 | Enhanced aspect-based sentiment analysis models with progressive self-supervised attention learning
Jinsong Su, Jialong Tang, Ziyao Lu, Yubin Ge, Linfeng Song, Deyi Xiong, Le Sun 0001, Jiebo Luo 0001 |
Artif. Intell. | 6 |
| 2021 | An External Knowledge Enhanced Graph-based Neural Network for Sentence OrderingabstractAs an important text coherence modeling task, sentence ordering aims to coherently organize a given set of unordered sentences. To achieve this goal, the most important step is to effectively capture and exploit global dependencies among these sentences. In this paper, we propose a novel and flexible external knowledge enhanced graph-based neural network for sentence ordering. Specifically, we first represent the input sentences as a graph, where various kinds of relations (i.e., entity-entity, sentence-sentence and entity-sentence) are exploited to make the graph representation more expressive and less noisy. Then, we introduce graph recurrent network to learn semantic representations of the sentences. To demonstrate the effectiveness of our model, we conduct experiments on several benchmark datasets. The experimental results and in-depth analysis show our model significantly outperforms the existing state-of-the-art models. Yongjing Yin, Shaopeng Lai, Linfeng Song, Chulun Zhou, Xianpei Han, Junfeng Yao, Jinsong Su |
J. Artif. Intell. Res. | 3 |
| 2021 | Conversational Semantic Role LabelingabstractSemantic role labeling (SRL) aims to extract the arguments for each predicate in an input sentence. Traditional SRL can fail to analyze dialogues because it only works on every single sentence, while ellipsis and anaphora frequently occur in dialogues. To address this problem, we propose the conversational SRL task, where an argument can be the dialogue participants, a phrase in the dialogue history or the current sentence. As the existing SRL datasets are in the sentence level, we manually annotate semantic roles for 3000 chit-chat dialogues (27198 sentences) to boost the research in this direction. Experiments show that while traditional SRL systems (even with the help of coreference resolution or rewriting) perform poorly for analyzing dialogues, modeling dialogue histories and participants greatly helps the performance, indicating that adapting SRL to conversations is very promising for universal dialogue understanding. Our initial study by applying CSRL to two mainstream conversational tasks, dialogue response generation and dialogue context rewriting, also confirms the usefulness of CSRL. Kun Xu 0005, Han Wu 0004, Linfeng Song, Haisong Zhang, Linqi Song, Dong Yu 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2020 | Relation Extraction Exploiting Full Dependency ForestsabstractDependency syntax has long been recognized as a crucial source of features for relation extraction. Previous work considers 1-best trees produced by a parser during preprocessing. However, error propagation from the out-of-domain parser may impact the relation extraction performance. We propose to leverage full dependency forests for this task, where a full dependency forest encodes all possible trees. Such representations of full dependency forests provide a differentiable connection between a parser and a relation extraction model, and thus we are also able to study adjusting the parser parameters based on end-task loss. Experiments on three datasets show that full dependency forests and parser adjustment give significant improvements over carefully designed baselines, showing state-of-the-art or competitive performances on biomedical or newswire benchmarks. Lifeng Jin, Linfeng Song, Yue Zhang 0004, Kun Xu 0005, Wei-Yun Ma, Dong Yu 0001 |
AAAI | 2 |
| 2020 | Coordinated Reasoning for Cross-Lingual Knowledge Graph AlignmentabstractExisting entity alignment methods mainly vary on the choices of encoding the knowledge graph, but they typically use the same decoding method, which independently chooses the local optimal match for each source entity. This decoding method may not only cause the “many-to-one” problem but also neglect the coordinated nature of this task, that is, each alignment decision may highly correlate to the other decisions. In this paper, we introduce two coordinated reasoning methods, i.e., the Easy-to-Hard decoding strategy and joint entity alignment algorithm. Specifically, the Easy-to-Hard strategy first retrieves the model-confident alignments from the predicted results and then incorporates them as additional knowledge to resolve the remaining model-uncertain alignments. To achieve this, we further propose an enhanced alignment model that is built on the current state-of-the-art baseline. In addition, to address the many-to-one problem, we propose to jointly predict entity alignments so that the one-to-one constraint can be naturally incorporated into the alignment prediction. Experimental results show that our model achieves the state-of-the-art performance and our reasoning methods can also significantly improve existing baselines. Kun Xu 0005, Linfeng Song, Yansong Feng 0002, Yan Song 0003, Dong Yu 0001 |
AAAI | 2 |
| 2020 | Enhancing Pointer Network for Sentence Ordering with Pairwise Ordering PredictionsabstractDominant sentence ordering models use a pointer network decoder to generate ordering sequences in a left-to-right fashion. However, such a decoder only exploits the noisy left-side encoded context, which is insufficient to ensure correct sentence ordering. To address this deficiency, we propose to enhance the pointer network decoder by using two pairwise ordering prediction modules: The FUTURE module predicts the relative orientations of other unordered sentences with respect to the candidate sentence, and the HISTORY module measures the local coherence between several (e.g., 2) previously ordered sentences and the candidate sentence, without the influence of noisy left-side context. Using the pointer mechanism, we then incorporate this dynamically generated information into the decoder as a supplement to the left-side context for better predictions. On several commonly-used datasets, our model significantly outperforms other baselines, achieving the state-of-the-art performance. Further analyses verify that pairwise ordering predictions indeed provide extra useful context as expected, leading to better sentence ordering. We also evaluate our sentence ordering models on a downstream task, multi-document summarization, and the summaries reordered by our model achieve the best coherence scores. Our code is available at https://github.com/DeepLearnXMU/Pairwise.git. Yongjing Yin, Fandong Meng, Jinsong Su, Yubin Ge, Linfeng Song, Jie Zhou 0016, Jiebo Luo 0001 |
AAAI | 5 |
| 2020 | Neural Simile Recognition with Cyclic Multitask Learning and Local AttentionabstractSimile recognition is to detect simile sentences and to extract simile components, i.e., tenors and vehicles. It involves two subtasks: simile sentence classification and simile component extraction. Recent work has shown that standard multitask learning is effective for Chinese simile recognition, but it is still uncertain whether the mutual effects between the subtasks have been well captured by simple parameter sharing. We propose a novel cyclic multitask learning framework for neural simile recognition, which stacks the subtasks and makes them into a loop by connecting the last to the first. It iteratively performs each subtask, taking the outputs of the previous subtask as additional inputs to the current one, so that the interdependence between the subtasks can be better explored. Extensive experiments show that our framework significantly outperforms the current state-of-the-art model and our carefully designed baselines, and the gains are still remarkable using BERT. Source Code of this paper are available on https://github.com/DeepLearnXMU/Cyclic. Jiali Zeng, Linfeng Song, Jinsong Su, Jiebo Luo 0001 |
AAAI | 2 |
| 2020 | Structural Information Preserving for Graph-to-Text GenerationabstractThe task of graph-to-text generation aims at producing sentences that preserve the meaning of input graphs.As a crucial defect, the current state-of-the-art models may mess up or even drop the core structural information of input graphs when generating outputs.We propose to tackle this problem by leveraging richer training signals that can guide our model for preserving input information.In particular, we introduce two types of autoencoding losses, each individually focusing on different aspects (a.k.a.views) of input graphs.The losses are then back-propagated to better calibrate our model via multi-task training.Experiments on two benchmarks for graph-to-text generation show the effectiveness of our approach over a state-of-the-art baseline.Our code is available at http://github.com/ Soistesimmer/AMR-multiview. Linfeng Song, Ante Wang, Jinsong Su, Yue Zhang 0004, Kun Xu 0005, Yubin Ge, Dong Yu 0001 |
ACL | 1 |
| 2020 | ZPR2: Joint Zero Pronoun Recovery and Resolution using Multi-Task Learning and BERTabstractZero pronoun recovery and resolution aim at recovering the dropped pronoun and pointing out its anaphoric mentions, respectively.We propose to better explore their interaction by solving both tasks together, while the previous work treats them separately.For zero pronoun resolution, we study this task in a more realistic setting, where no parsing trees or only automatic trees are available, while most previous work assumes gold trees.Experiments on two benchmarks show that joint modeling significantly outperforms our baseline that already beats the previous state of the arts. Linfeng Song, Kun Xu 0005, Yue Zhang 0004, Jianshu Chen, Dong Yu 0001 |
ACL | 1 |
| 2020 | Online Back-Parsing for AMR-to-Text GenerationabstractAMR-to-text generation aims to recover a text containing the same meaning as an input AMR graph.Current research develops increasingly powerful graph encoders to better represent AMR graphs, with decoders based on standard language modeling being used to generate outputs.We propose a decoder that back predicts projected AMR graphs on the target sentence during text generation.As the result, our outputs can better preserve the input meaning than standard decoders.Experiments on two AMR benchmarks show the superiority of our model over the previous state-of-the-art system based on graph Transformer. Xuefeng Bai 0001, Linfeng Song, Yue Zhang 0004 |
EMNLP (1) | 2 |
| 2020 | Semantic Role Labeling Guided Multi-turn Dialogue ReWriterabstractFor multi-turn dialogue rewriting, the capacity of effectively modeling the linguistic knowledge in dialog context and getting rid of the noises is essential to improve its performance.Existing attentive models attend to all words without prior focus, which results in inaccurate concentration on some dispensable words.In this paper, we propose to use semantic role labeling (SRL), which highlights the core semantic information of who did what to whom, to provide additional guidance for the rewriter model.Experiments show that this information significantly improves a RoBERTa-based model that already outperforms previous stateof-the-art systems. Kun Xu 0005, Haochen Tan, Linfeng Song, Han Wu 0004, Haisong Zhang, Linqi Song, Dong Yu 0001 |
EMNLP (1) | 3 |
| 2020 | Rich Syntactic and Semantic Information Helps Unsupervised Text Style TransferabstractText style transfer aims to change an input sentence to an output sentence by changing its text style while preserving the content.Previous efforts on unsupervised text style transfer only use the surface features of words and sentences.As a result, the transferred sentences may either have inaccurate or missing information compared to the inputs.We address this issue by explicitly enriching the inputs via syntactic and semantic structures, from which richer features are then extracted to better capture the original information.Experiments on two text-style-transfer tasks show that our approach improves the content preservation of a strong unsupervised baseline model thereby demonstrating improved transfer performance. Hongyu Gong, Linfeng Song, Suma Bhat |
INLG | 2 |
| 2019 | SemBleu: A Robust Metric for AMR Parsing EvaluationabstractEvaluating AMR parsing accuracy involves comparing pairs of AMR graphs. The major evaluation metric, SMATCH (Cai and Knight, 2013), searches for one-to-one mappings between the nodes of two AMRs with a greedy hill-climbing algorithm, which leads to search errors. We propose SEMBLEU, a robust metric that extends BLEU (Papineni et al., 2002) to AMRs. It does not suffer from search errors and considers non-local correspondences in addition to local ones. SEMBLEU is fully content-driven and punishes situations where a system's output does not preserve most information from the input. Preliminary experiments on both sentence and corpus levels show that SEMBLEU has slightly higher consistency with human judgments than SMATCH. Our code is available at http://github.com/ freesunshine0316/sembleu. Linfeng Song, Daniel Gildea |
ACL (1) | 1 |
| 2019 | Progressive Self-Supervised Attention Learning for Aspect-Level Sentiment AnalysisabstractIn aspect-level sentiment classification (ASC), it is prevalent to equip dominant neural models with attention mechanisms, for the sake of acquiring the importance of each context word on the given aspect. However, such a mechanism tends to excessively focus on a few frequent words with sentiment polarities, while ignoring infrequent ones. In this paper, we propose a progressive self-supervised attention learning approach for neural ASC models, which automatically mines useful attention supervision information from a training corpus to refine attention mechanisms. Specifically, we iteratively conduct sentiment predictions on all training instances. Particularly, at each iteration, the context word with the maximum attention weight is extracted as the one with active/misleading influence on the correct/incorrect prediction of every instance, and then the word itself is masked for subsequent iterations. Finally, we augment the conventional training objective with a regularization term, which enables ASC models to continue equally focusing on the extracted active context words while decreasing weights of those misleading ones. Experimental results on multiple datasets show that our proposed approach yields better attention mechanisms, leading to substantial improvements over the two state-of-the-art neural ASC models. Source code and trained models are available at https://github.com/DeepLearnXMU/PSSAttention. Jialong Tang, Ziyao Lu, Jinsong Su, Yubin Ge, Linfeng Song, Le Sun 0001, Jiebo Luo 0001 |
ACL (1) | 5 |
| 2019 | Leveraging Dependency Forest for Neural Medical Relation ExtractionabstractLinfeng Song, Yue Zhang, Daniel Gildea, Mo Yu, Zhiguo Wang, Jinsong Su. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Linfeng Song, Yue Zhang 0004, Daniel Gildea, Mo Yu, Zhiguo Wang 0006, Jinsong Su |
EMNLP/IJCNLP (1) | 1 |
| 2019 | Neural Collective Entity Linking Based on Recurrent Random Walk Network LearningabstractBenefiting from the excellent ability of neural networks on learning semantic representations, existing studies for entity linking (EL) have resorted to neural networks to exploit both the local mention-to-entity compatibility and the global interdependence between different EL decisions for target entity disambiguation. However, most neural collective EL methods depend entirely upon neural networks to automatically model the semantic dependencies between different EL decisions, which lack of the guidance from external knowledge. In this paper, we propose a novel end-to-end neural network with recurrent random-walk layers for collective EL, which introduces external knowledge to model the semantic interdependence between different EL decisions. Specifically, we first establish a model based on local context features, and then stack random-walk layers to reinforce the evidence for related EL decisions into high-probability decisions, where the semantic interdependence between candidate entities is mainly induced from an external knowledge base. Finally, a semantic regularizer that preserves the collective EL decisions consistency is incorporated into the conventional objective function, so that the external knowledge base can be fully exploited in collective EL decisions. Experimental results and in-depth analysis on various datasets show that our model achieves better performance than other state-of-the-art models. Our code and data are released at https://github.com/DeepLearnXMU/RRWEL. Mengge Xue, Weiming Cai, Jinsong Su, Linfeng Song, Yubin Ge, Bin Wang 0004 |
IJCAI | 4 |
| 2019 | Graph-based Neural Sentence OrderingabstractSentence ordering is to restore the original paragraph from a set of sentences. It involves capturing global dependencies among sentences regardless of their input order. In this paper, we propose a novel and flexible graph-based neural sentence ordering model, which adopts graph recurrent network \citep{Zhang:acl18} to accurately learn semantic representations of the sentences. Instead of assuming connections between all pairs of input sentences, we use entities that are shared among multiple sentences to make more expressive graph representations with less noise. Experimental results show that our proposed model outperforms the existing state-of-the-art systems on several benchmark datasets, demonstrating the effectiveness of our model. We also conduct a thorough analysis on how entities help the performance. Our code is available at https://github.com/DeepLearnXMU/NSEG.git. Yongjing Yin, Linfeng Song, Jinsong Su, Jiali Zeng, Chulun Zhou, Jiebo Luo 0001 |
IJCAI | 2 |
| 2019 | Semantic Neural Machine Translation using AMRabstractAbstract It is intuitive that semantic representations can be useful for machine translation, mainly because they can help in enforcing meaning preservation and handling data sparsity (many sentences correspond to one meaning) of machine translation models. On the other hand, little work has been done on leveraging semantics for neural machine translation (NMT). In this work, we study the usefulness of AMR (abstract meaning representation) on NMT. Experiments on a standard English-to-German dataset show that incorporating AMR as additional knowledge can significantly improve a strong attention-based sequence-to-sequence neural translation model. Linfeng Song, Daniel Gildea, Yue Zhang 0004, Zhiguo Wang 0006, Jinsong Su |
Trans. Assoc. Comput. Linguistics | 1 |
| 2018 | A Graph-to-Sequence Model for AMR-to-Text GenerationabstractThe problem of AMR-to-text generation is to recover a text representing the same meaning as an input AMR graph.The current state-of-the-art method uses a sequence-to-sequence model, leveraging LSTM for encoding a linearized AMR structure.Although it is able to model non-local semantic information, a sequence LSTM can lose information from the AMR graph structure, and thus faces challenges with large graphs, which result in long sequences.We introduce a neural graph-to-sequence model, using a novel LSTM structure for directly encoding graph-level semantics.On a standard benchmark, our model shows superior results to existing methods in the literature. Linfeng Song, Yue Zhang 0004, Zhiguo Wang 0006, Daniel Gildea |
ACL (1) | 1 |
| 2018 | Sequence-to-sequence Models for Cache Transition SystemsabstractIn this paper, we present a sequenceto-sequence based approach for mapping natural language sentences to AMR semantic graphs.We transform the sequence to graph mapping problem to a word sequence to transition action sequence problem using a special transition system called a cache transition system.To address the sparsity issue of neural AMR parsing, we feed feature embeddings from the transition state to provide relevant local information for each decoder state.We present a monotonic hard attention model for the transition framework to handle the strictly left-to-right alignment between each transition state and the current buffer input focus.We evaluate our neural transition model on the AMR parsing task, and our parser outperforms other sequence-to-sequence approaches and achieves competitive results in comparison with the best-performing models. 1 Xiaochang Peng, Linfeng Song, Daniel Gildea, Giorgio Satta |
ACL (1) | 2 |
| 2018 | Sentence-State LSTM for Text RepresentationabstractBi-directional LSTMs are a powerful tool for text representation.On the other hand, they have been shown to suffer various limitations due to their sequential nature.We investigate an alternative LSTM structure for encoding text, which consists of a parallel state for each word.Recurrent steps are used to perform local and global information exchange between words simultaneously, rather than incremental reading of a sequence of words.Results on various classification and sequence labelling benchmarks show that the proposed model has strong representation power, giving highly competitive performances compared to stacked BiLSTM models with similar parameter numbers. Yue Zhang 0004, Qi Liu 0049, Linfeng Song |
ACL (1) | 3 |
| 2018 | N-ary Relation Extraction using Graph-State LSTMabstractCross-sentence n-ary relation extraction detects relations among n entities across multiple sentences.Typical methods formulate an input as a document graph, integrating various intra-sentential and inter-sentential dependencies.The current state-of-the-art method splits the input graph into two DAGs, adopting a DAG-structured LSTM for each.Though being able to model rich linguistic knowledge by leveraging graph edges, important information can be lost in the splitting procedure.We propose a graph-state LSTM model, which uses a parallel state to model each word, recurrently enriching state values via message passing.Compared with DAG LSTMs, our graph LSTM keeps the original graph structure, and speeds up computation by allowing more parallelization.On a standard benchmark, our model shows the best result in the literature. Linfeng Song, Yue Zhang 0004, Zhiguo Wang 0006, Daniel Gildea |
EMNLP | 1 |
| 2018 | Neural Transition-based Syntactic LinearizationabstractThe task of linearization is to find a grammatical order given a set of words.Traditional models use statistical methods.Syntactic linearization systems, which generate a sentence along with its syntactic tree, have shown state-of-the-art performance.Recent work shows that a multilayer LSTM language model outperforms competitive statistical syntactic linearization systems without using syntax.In this paper, we study neural syntactic linearization, building a transition-based syntactic linearizer leveraging a feed forward neural network, observing significantly better results compared to LSTM language models on this task. Linfeng Song, Yue Zhang 0004, Daniel Gildea |
INLG | 1 |
| 2016 | AMR-to-text generation as a Traveling Salesman ProblemabstractThe task of AMR-to-text generation is to generate grammatical text that sustains the semantic meaning for a given AMR graph.We attack the task by first partitioning the AMR graph into smaller fragments, and then generating the translation for each fragment, before finally deciding the order by solving an asymmetric generalized traveling salesman problem (AGTSP).A Maximum Entropy classifier is trained to estimate the traveling costs, and a TSP solver is used to find the optimized solution.The final model reports a BLEU score of 22.44 on the SemEval-2016 Task8 dataset. Linfeng Song, Yue Zhang 0004, Xiaochang Peng, Zhiguo Wang 0006, Daniel Gildea |
EMNLP | 1 |
| 2015 | A Synchronous Hyperedge Replacement Grammar based approach for AMR parsingabstractThis paper presents a synchronous-graphgrammar-based approach for string-to-AMR parsing.We apply Markov Chain Monte Carlo (MCMC) algorithms to learn Synchronous Hyperedge Replacement Grammar (SHRG) rules from a forest that represents likely derivations consistent with a fixed string-to-graph alignment.We make an analogy of string-to-AMR parsing to the task of phrase-based machine translation and come up with an efficient algorithm to learn graph grammars from string-graph pairs.We propose an effective approximation strategy to resolve the complexity issue of graph compositions.We also show some useful strategies to overcome existing problems in an SHRG-based parser and present preliminary results of a graph-grammar-based approach. Xiaochang Peng, Linfeng Song, Daniel Gildea |
CoNLL | 2 |
| 2014 | Joint Morphological Generation and Syntactic LinearizationabstractThere has been growing interest in stochastic methods to natural language generation (NLG). While most NLG pipelines separate morphological generation and syntactic linearization, the two tasks are closely related. In this paper, we study joint morphological generation and linearization, making use of word order and inflections information for both tasks and reducing error propagation. Experiments show that the joint method significantly outperforms a strong pipelined baseline (by 1.1 BLEU points). It also achieves the best reported result on the Generation Challenge 2011 shared task. Linfeng Song, Yue Zhang 0004, Qun Liu 0001 |
AAAI | 1 |
| 2014 | Syntactic SMT Using a Discriminative Text Generation ModelabstractWe study a novel architecture for syntactic SMT. In contrast to the dominant approach in the literature, the system does not rely on translation rules, but treat translation as an unconstrained target sentence gen-eration task, using soft features to cap-ture lexical and syntactic correspondences between the source and target languages. Target syntax features and bilingual trans-lation features are trained consistently in a discriminative model. Experiments us-ing the IWSLT 2010 dataset show that the system achieves BLEU comparable to the state-of-the-art syntactic SMT systems. 1 Yue Zhang 0004, Linfeng Song, Qun Liu 0001 |
EMNLP | 3 |
| 2013 | Translation with Source Constituency and Dependency TreesabstractWe present a novel translation model, which simultaneously exploits the constituency and dependency trees on the source side, to combine the advantages of two types of trees.We take head-dependents relations of dependency trees as backbone and incorporate phrasal nodes of constituency trees as the source side of our translation rules, and the target side as strings.Our rules hold the property of long distance reorderings and the compatibility with phrases.Large-scale experimental results show that our model achieves significantly improvements over the constituency-to-string (+2.45 BLEU on average) and dependencyto-string (+0.91 BLEU on average) models, which only employ single type of trees, and significantly outperforms the state-of-theart hierarchical phrase-based model (+1.12BLEU on average), on three Chinese-English NIST test sets. Fandong Meng, Linfeng Song, Yajuan Lü, Qun Liu 0001 |
EMNLP | 3 |
| 2011 | Bagging-based System Combination for Domain Adaption
Linfeng Song, Haitao Mi, Yajuan Lü, Qun Liu 0001 |
MTSummit | 1 |