EDBT 2026 Demo / reviewers in the wild / expert
Kehai Chen
dblp:78/9623
· DBLP profile ↗
82ranked-venue papers
14as first author
50since 2021 · last 2026
0000-0002-4346-7618ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 72 · 14 first-author · 46 since 2021Graphics, computer vision, multimedia, augmented reality and games · 10 · 3 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 9 · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Long-form RewardBench: Evaluating Reward Models for Long-form GenerationabstractThe widespread adoption of reinforcement learning-based alignment highlights the growing importance of reward models. Various benchmarks have been built to evaluate reward models in various domains and scenarios. However, a significant gap remains in assessing reward models for long-form generation, despite its critical role in real-world applications. To bridge this, we introduce Long-form RewardBench, the first reward modeling testbed specifically designed for long-form generation. Our benchmark encompasses five key subtasks: QA, RAG, Chat, Writing, and Reasoning. We collected instruction and preference data through a meticulously designed multi-stage data collection process, and conducted extensive experiments on 20+ mainstream reward models, including both classifiers and generative models. Our findings reveal that current models still lack long-form reward modeling capabilities. Furthermore, we designed a novel Long-form Needle-in-a-Haystack Test, which revealed a correlation between reward modeling performance and the error's position within a response, as well as the overall response length, with distinct characteristics observed between classification and generative models. Finally, we demonstrate that classifier exhibit better generalizability compared to generative models trained on the same data. As the first benchmark for long-form reward modeling, this work aims to offer a robust platform for visualizing progress in this crucial area. Hui Huang 0021, Yancheng He, Muyun Yang, Kehai Chen, Conghui Zhu, Hailong Cao, Tiejun Zhao |
AAAI | 6 |
| 2026 | Lost in Benchmarks? Rethinking Large Language Model Benchmarking with Item Response TheoryabstractThe evaluation of large language models (LLMs) via benchmarks is widespread, yet inconsistencies between different leaderboards and poor separability among top models raise concerns about their ability to accurately reflect authentic model capabilities. This paper provides a critical analysis of benchmark effectiveness, examining mainstream prominent LLM benchmarks using results from diverse models. We first propose Pseudo-Siamese Network for Item Response Theory (PSN-IRT), an enhanced Item Response Theory framework that incorporates a rich set of item parameters within an IRT-grounded architecture. PSN-IRT can be utilized for accurate and reliable estimations of item characteristics and model abilities. Based on PSN-IRT, we conduct extensive analysis on 11 LLM benchmarks comprising 41,871 items, revealing significant and varied shortcomings in their measurement quality. Furthermore, we demonstrate that leveraging PSN-IRT is able to construct smaller benchmarks while maintaining stronger alignment with human preference. Hongli Zhou 0001, Hui Huang 0021, Ziqing Zhao, Lvyuan Han, Huicheng Wang, Kehai Chen, Muyun Yang, Conghui Zhu, Hailong Cao, Tiejun Zhao |
AAAI | 6 |
| 2026 | SAT: Balancing Reasoning Accuracy and Efficiency with Stepwise Adaptive ThinkingabstractWeiyang Huang, Xuefeng Bai, Kehai Chen, Xinyang Chen, Yibin Chen, Weili Guan, Min Zhang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Weiyang Huang, Xuefeng Bai 0001, Kehai Chen, Xinyang Chen 0001, Yibin Chen, Weili Guan, Min Zhang 0005 |
ACL (1) | 3 |
| 2026 | WSDPO: A Generative Word Sense Disambiguation Framework with Chain-of-Thought and Preference OptimizationabstractKunpeng Kang, Shuaimin Li, Kaiyuan Zhang, Luyang Zhang, Jiasheng Si, Bing Xu, Kehai Chen, Muyun Yang, Wenpeng Lu. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Kunpeng Kang, Shuaimin Li, Jiasheng Si, Kehai Chen, Muyun Yang, Wenpeng Lu |
ACL (1) | 7 |
| 2026 | DUAL RM: Beyond Rule-based Preference Reward Modeling via Meta-RewardabstractXiaobo Liang, Wanfu Wang, Qipeng Huang, Yuyang Ding, Zecheng Tang, Yixin Ji, Qianben Chen, Zhe Zhao, Kehai Chen, Juntao Li, Min Zhang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Xiaobo Liang, Wanfu Wang, Qipeng Huang, Yuyang Ding, Zecheng Tang, Yixin Ji, Qianben Chen, Kehai Chen, Juntao Li 0005, Min Zhang 0005 |
ACL (1) | 9 |
| 2026 | Diagnosing and Remedying Representation Deficiencies for Deterministic Reasoning in KGQAabstractGewen Liang, Mufan Xu, Kehai Chen, Wei Wang, Yuwei Wang, Muyun Yang, Tiejun Zhao, Min Zhang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Gewen Liang, Mufan Xu, Kehai Chen, Wei Wang 0164, Muyun Yang, Tiejun Zhao, Min Zhang 0005 |
ACL (1) | 3 |
| 2026 | LLM With Relation Classifier for Document-Level Relation ExtractionabstractLarge language models (LLMs) have created a new paradigm for natural language processing. Despite their advancement, LLM-based methods still lag behind traditional approaches in document-level relation extraction (DocRE), a critical task for understanding complex entity relations within long context. This paper investigates the causes of this performance gap, identifying the dispersion of attention by LLMs due to entity pairs without relations as a key factor. We then introduce a novel classifier-LLM approach to DocRE. Particularly, the proposed approach begins with a classifier designed to select entity pair candidates that exhibit potential relations and then feed them to LLM for final relation classification. This method ensures that the LLM's attention is directed at relation-expressing entity pairs instead of those without relations during inference. Experiments on DocRE benchmarks reveal that our method significantly outperforms recent LLM-based DocRE models and narrows the performance gap with state-of-the-art BERT-based models. Xingzuo Li, Kehai Chen, Min Zhang 0005 |
IEEE Trans. Big Data | 2 |
| 2025 | Look Before You Leap: Enhance Attention and Vigilance Regarding Harmful Content with GuidelineLLMabstractDespite being empowered with alignment mechanisms, large language models (LLMs) are increasingly vulnerable to emerging jailbreak attacks that can compromise their alignment mechanisms. This vulnerability poses significant risks to real-world applications. Existing work faces challenges in both training efficiency and generalization capabilities (i.e., Reinforcement Learning from Human Feedback and Red-Teaming). Developing effective strategies to enable LLMs to resist continuously evolving jailbreak attempts represents a significant challenge. To address this challenge, we propose a novel defensive paradigm called GuidelineLLM, which assists LLMs in recognizing queries that may have harmful content. Before LLMs respond to a query, GuidelineLLM first identifies potential risks associated with the query, summarizes these risks into guideline suggestions, and then feeds these guidelines to the responding LLMs. Importantly, our approach eliminates the necessity for additional safety fine-tuning of the LLMs themselves; only the GuidelineLLM requires fine-tuning. This characteristic enhances the general applicability of GuidelineLLM across various LLMs. Experimental results demonstrate that GuidelineLLM can significantly reduce the attack success rate (ASR) against LLM (an average reduction of 34.17% ASR) while maintaining the usefulness of LLM in handling benign queries. Shaoqing Zhang, Zhuosheng Zhang 0001, Kehai Chen, Rongxiang Weng, Muyun Yang, Tiejun Zhao, Min Zhang 0005 |
AAAI | 3 |
| 2025 | Make Imagination Clearer! Stable Diffusion-based Visual Imagination for Multimodal Machine TranslationabstractAndong Chen, Yuchen Song, Kehai Chen, Xuefeng Bai, Muyun Yang, Liqiang Nie, Jie Liu, Tiejun Zhao, Min Zhang. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Andong Chen 0001, Kehai Chen, Xuefeng Bai 0001, Muyun Yang, Liqiang Nie, Jie Liu 0001, Tiejun Zhao, Min Zhang 0005 |
ACL (1) | 3 |
| 2025 | Efficient Safety Alignment of Large Language Models via Preference Re-ranking and Representation-based Reward ModelingabstractReinforcement Learning (RL) algorithms for safety alignment of Large Language Models (LLMs), such as Direct Preference Optimization (DPO), encounter the challenge of distribution shift.Current approaches typically address this issue through online sampling from the target policy, which requires significant computational resources.In this paper, we hypothesize that during off-policy training, while the ranking order of output generated by policy changes, their overall distribution remains relatively stable.This stability allows the conversion of the sampling process from the target policy into a computationally efficient reranking of preference data.Building on this hypothesis, we propose a new framework that leverages the model's intrinsic safety judgment capability to extract reward signals, which are then used to calculate label confidence for preference reordering.Extensive experiments and theoretical analysis demonstrate that the proposed method effectively addresses the distribution shift issue, remarkably enhancing the safety performance while avoiding about 300x computational overheads. Qiyuan Deng, Xuefeng Bai 0001, Kehai Chen, Yaowei Wang 0001, Liqiang Nie, Min Zhang 0005 |
ACL (1) | 3 |
| 2025 | Generative Reward Modeling via Synthetic Criteria Preference LearningabstractGenerative Reward Models (GenRMs) leverage synthesized Chains of Thought (CoT) to reduce the need for massive labeled data, but this approach introduces risks of overoptimization due to the inability to guarantee the correctness of the CoTs.Identifying and optimizing unexpected behaviors within these synthesized CoT remains a challenge, as it heavily depends on precise annotations of intermediate behavior, similar to process supervision.In this work, we introduce a criteria-based preference tree for reward modeling, where each path in the tree represents a reasoning trajectory based on synthesized criteria.Crucially, each reasoning trajectory can be independently optimized through RL algorithm.These fine-grained process reward signals are derived from the inferencetime computations and predefined rules, eliminating the need for human supervision.In experiments, SyncPL 1 showed significant improvements over baselines on multiple human preference benchmarks.We further demonstrate that synthesized data can be learned using a long CoT format, analogous to an o1-like model, further enhancing performance while keeping stability and efficiency during training. Xiaobo Liang, Haoke Zhang, Juntao Li 0005, Kehai Chen, Qiaoming Zhu, Min Zhang 0005 |
ACL (1) | 4 |
| 2025 | Benchmarking and Improving Large Vision-Language Models for Fundamental Visual Graph Understanding and ReasoningabstractLarge Vision-Language Models (LVLMs) have demonstrated remarkable performance across diverse tasks.Despite great success, recent studies show that LVLMs encounter substantial limitations when engaging with visual graphs.To study the reason behind these limitations, we propose VGCURE, a comprehensive benchmark covering 22 tasks for examining the fundamental graph understanding and reasoning capacities of LVLMs.Extensive evaluations conducted on 14 LVLMs reveal that LVLMs are weak in basic graph understanding and reasoning tasks, particularly those concerning relational or structurally complex information.Based on this observation, we propose a structure-aware fine-tuning framework to enhance LVLMs with structure learning abilities through three self-supervised learning tasks.Experiments validate the effectiveness of our method in improving LVLMs' performance on fundamental and downstream graph learning tasks, as well as enhancing their robustness against complex visual graphs. Yingjie Zhu, Xuefeng Bai 0001, Kehai Chen, Yang Xiang 0003, Jun Yu 0002, Min Zhang 0005 |
ACL (1) | 3 |
| 2025 | ZigZagKV: Dynamic KV Cache Compression for Long-context Modeling based on Layer UncertaintyabstractLarge Language models (LLMs) have become a research hotspot. To accelerate the inference of LLMs, storing computed caches in memory has become the standard technique. However, as the inference length increases, growing KV caches might lead to out-of-memory issues. Many existing methods address this issue through KV cache compression, primarily by preserving key tokens throughout all layers to reduce information loss. Most of them allocate a uniform budget size for each layer to retain. However, we observe that the minimum budget sizes needed to retain essential information vary across layers and models based on the perspectives of attention and hidden state output. Building on this observation, this paper proposes a simple yet effective KV cache compression method that leverages layer uncertainty to allocate budget size for each layer. Experimental results show that the proposed method can reduce memory usage of the KV caches to only ~20% when compared to full KV inference while achieving nearly lossless performance. Meizhi Zhong, Xikai Liu, Chen Zhang 0020, Yikun Lei, Yan Gao 0017, Yao Hu 0002, Kehai Chen, Min Zhang 0005 |
COLING | 7 |
| 2025 | Understanding the RoPE Extensions of Long-Context LLMs: An Attention PerspectiveabstractEnabling LLMs to handle lengthy context is currently a research hotspot. Most LLMs are built upon rotary position embedding (RoPE), a popular position encoding method. Therefore, a prominent path is to extrapolate the RoPE trained on comparably short texts to far longer texts. A heavy bunch of efforts have been dedicated to boosting the extrapolation via extending the formulations of the RoPE, however, few of them have attempted to showcase their inner workings comprehensively. In this paper, we are driven to offer a straightforward yet in-depth understanding of RoPE extensions from an attention perspective and on two benchmarking tasks. A broad array of experiments reveals several valuable findings: 1) Maintaining attention patterns to those at the pretrained length improves extrapolation; 2) Large attention uncertainty leads to retrieval errors; 3) Using longer continual pretraining lengths for RoPE extensions could reduce attention uncertainty and significantly enhance extrapolation. Meizhi Zhong, Chen Zhang 0020, Yikun Lei, Xikai Liu, Yan Gao 0017, Yao Hu 0002, Kehai Chen, Min Zhang 0005 |
COLING | 7 |
| 2025 | Benchmarking LLMs for Translating Classical Chinese Poetry: Evaluating Adequacy, Fluency, and EleganceabstractLarge language models (LLMs) have shown remarkable performance in general translation tasks.However, the increasing demand for high-quality translations that are not only adequate but also fluent and elegant.To assess the extent to which current LLMs can meet these demands, we introduce a suitable benchmark (PoetMT) for translating classical Chinese poetry into English.This task requires not only adequacy in translating culturally and historically significant content but also a strict adherence to linguistic fluency and poetic elegance.Our study reveals that existing LLMs fall short of this task.To address these issues, we propose RAT, a Retrieval-Augmented machine Translation method that enhances the translation process by incorporating knowledge related to classical poetry.Additionally, we propose an automatic evaluation metric based on GPT-4, which better assesses translation quality in terms of adequacy, fluency, and elegance, overcoming the limitations of traditional metrics.Our dataset and code will be made available 1 . Andong Chen 0001, Lianzhang Lou, Kehai Chen, Xuefeng Bai 0001, Yang Xiang 0003, Muyun Yang, Tiejun Zhao, Min Zhang 0005 |
EMNLP | 3 |
| 2025 | ORPP: Self-Optimizing Role-playing Prompts to Enhance Language Model CapabilitiesabstractHigh-quality prompts are crucial for eliciting outstanding performance from large language models (LLMs) on complex tasks.Existing research has explored model-driven strategies for prompt optimization.However, these methods often suffer from high computational overhead or require strong optimization capabilities from the model itself, which limits their broad applicability.To address these challenges, we propose ORPP, a framework that enhances model performance by optimizing and generating roleplaying prompts.The core idea of ORPP is to confine the prompt search space to role-playing scenarios, thereby fully activating the model's intrinsic capabilities through carefully crafted, high-quality role-playing prompts.Specifically, ORPP first performs iterative optimization on a small subset of training samples to generate high-quality role-playing prompts.Then, leveraging the model's few-shot learning capability, it transfers the optimization experience to efficiently generate suitable prompts for the remaining samples.Our experimental results show that ORPP not only matches but in most cases surpasses existing mainstream prompt optimization methods in terms of performance.Notably, ORPP suggests great "plug-and-play" capability.In most cases, it can be integrated with various other prompt methods and further enhance their effectiveness. Yifan Duan, Yihong Tang, Kehai Chen, Liqiang Nie, Min Zhang 0005 |
EMNLP | 3 |
| 2025 | Generator-Assistant Stepwise Rollback Framework for Large Language Model AgentabstractLarge language model (LLM) agents typically adopt a step-by-step reasoning framework, in which they interleave the processes of thinking and acting to accomplish the given task.However, this paradigm faces a deeprooted one-pass issue whereby each generated intermediate thought is plugged into the trajectory regardless of its correctness, which can cause irreversible error propagation.To address the issue, this paper proposes a novel framework called Generator-Assistant Stepwise Rollback (GA-Rollback) to induce better decision-making for LLM agents.Particularly, GA-Rollback utilizes a generator to interact with the environment and an assistant to examine each action produced by the generator, where the assistant triggers a rollback operation upon detection of incorrect actions.Moreover, we introduce two additional strategies tailored for the rollback scenario to further improve its effectiveness.Extensive experiments show that GA-Rollback achieves significant improvements over several strong baselines on three widely used benchmarks.Our analysis further reveals that GA-Rollback can function as a robust plug-and-play module, integrating seamlessly with other methods. 1 Xingzuo Li, Kehai Chen, Xuefeng Bai 0001, Yong Xu 0001, Min Zhang 0005 |
EMNLP | 2 |
| 2025 | Merge then Realign: Simple and Effective Modality-Incremental Continual Learning for Multimodal LLMsabstractRecent advances in Multimodal Large Language Models (MLLMs) have enhanced their versatility as they integrate a growing number of modalities.Considering the heavy cost of training MLLMs, it is efficient to reuse the existing ones and extend them to more modalities through Modality-incremental Continual Learning (MCL).The exploration of MCL is in its early stages.In this work, we dive into the causes of performance degradation in MCL.We uncover that it suffers not only from forgetting as in traditional continual learning, but also from misalignment between the modality-agnostic and modality-specific components.To this end, we propose an elegantly simple MCL paradigm called "MErge then ReAlign" (MERA) to address both forgetting and misalignment.MERA avoids introducing heavy model budgets or modifying model architectures, hence is easy to deploy and highly reusable in the MLLM community.Extensive experiments demonstrate the impressive performance of MERA, holding an average of 99.84% Backward Relative Gain when extending to four modalities, achieving nearly lossless MCL performance.Our findings underscore the misalignment issue in MCL.More broadly, our work showcases how to adjust different components of MLLMs during continual learning. Dingkun Zhang, Shuhan Qi, Xinyu Xiao, Kehai Chen, Xuan Wang 0002 |
EMNLP | 4 |
| 2025 | A Survey on the Feedback Mechanism of LLM-based AI AgentsabstractLarge language models (LLMs) are increasingly being adopted to develop general-purpose AI agents. However, it remains challenging for these LLM-based AI agents to efficiently learn from feedback and iteratively optimize their strategies. To address this challenge, tremendous efforts have been dedicated to designing diverse feedback mechanisms for LLM-based AI agents. To provide a comprehensive overview of this rapidly evolving field, this paper presents a systematic review of these studies, offering a holistic perspective on the feedback mechanisms in LLM-based AI agents. We begin by discussing the construction of LLM-based AI agents, introducing a generalized framework that encapsulates much of the existing work. Next, we delve into the exploration of feedback mechanisms, categorizing them into four distinct types: internal feedback, external feedback, multi-agent feedback, and human feedback. Additionally, we provide an overview of evaluation protocols and benchmarks specifically tailored for LLM-based AI agents. Finally, we highlight the significant challenges and identify potential directions for future studies. The relevant papers are summarized and will be consistently updated at https://github.com/kevinson7515/Agents-Feedback-Mechanisms. Xuefeng Bai 0001, Kehai Chen, Xinyang Chen 0001, Xiucheng Li, Yang Xiang 0003, Jin Liu 0012, Hong-Dong Li, Yaowei Wang 0001, Liqiang Nie, Min Zhang 0005 |
IJCAI | 3 |
| 2025 | BIMCompNet: Multimodal Dataset for Geometric Deep Learning in Building Information ModelabstractBuilding Information Model (BIM) has become a significantly digital platform for representing buildings in the Architecture, Engineering, and Construction (AEC) industry. However, the absence of extensive, class- diverse, and balanced datasets at the BIM component level has limited the development of AI-driven BIM analysis. In this study, BIMCompNet is proposed as a large-scale multimodal dataset from Industry Foundation Classes (IFC), which can learn BIM component geometry features from multiple representation methods, including rendered views, point clouds, mesh structures, voxel grids, and semantic graphs. BIMCompNet is constructed by a standardized two-stage processing pipeline: (1) At the model level, geometry units are normalized to the SI units, models are converted to the IFC format, metadata is anonymized, and components are automatically extracted into individual IFC files. (2) At the component level, semantic labels are corrected, geometry and positioning are aligned, duplicates at model and project levels are removed, and five synchronized modalities (OBJ meshes, multi-view images, point clouds, voxel grids, and heterogeneous IFC graphs) are generated. BIMCompNet comprises 1,304,206 cleaned and labeled components across 87 IFC classes, collected from 1,607 real-world BIM models spanning 14 building types. To mitigate class imbalance, underrepresented classes are merged, and dominant classes are down-sampled to create balanced subsets suitable for robust AI model training and benchmarking. Benchmarking is performed on classification tasks by different models with multiple data modalities. Both the dataset and the processing pipeline will be publicly released to support reproducibility and private dataset extension. Mingsong Yang, Xinhong Hei 0001, Kehai Chen, Haining Meng, Haoyang Dong |
ACM Multimedia | 3 |
| 2025 | MoDification: Mixture of Depths Made EasyabstractChen Zhang, Meizhi Zhong, Qimeng Wang, Xuantao Lu, Zheyu Ye, Chengqiang Lu, Yan Gao, Yao Hu, Kehai Chen, Min Zhang, Dawei Song. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Chen Zhang 0020, Meizhi Zhong, Qimeng Wang, Xuantao Lu, Zheyu Ye, Chengqiang Lu, Yan Gao 0017, Yao Hu 0002, Kehai Chen, Min Zhang 0005, Dawei Song 0001 |
NAACL (Long Papers) | 9 |
| 2025 | XIFBench: Evaluating Large Language Models on Multilingual Instruction FollowingabstractLarge Language Models (LLMs) have demonstrated remarkable instruction-following capabilities across various applications. However, their performance in multilingual settings lacks systematic investigation, with existing evaluations lacking fine-grained constraint analysis across diverse linguistic contexts. We introduce XIFBench, a comprehensive constraint-based benchmark for evaluating multilingual instruction-following abilities of LLMs, comprising 558 instructions with 0-5 additional constraints across five categories (Content, Style, Situation, Format, and Numerical) in six languages spanning different resource levels. To support reliable and consistent cross-lingual evaluation, we implement three methodological innovations: cultural accessibility annotation, constraint-level translation validation, and requirement-based evaluation using English requirements as semantic anchors across languages. Extensive experiments with various LLMs not only quantify performance disparities across resource levels but also provide detailed insights into how language resources, constraint categories, instruction complexity, and cultural specificity influence multilingual instruction-following. Our code and data are available at https://github.com/zhenyuli801/XIFBench. Kehai Chen, Xuefeng Bai 0001, Yaoyin Zhang, Xuchen Wei, Juntao Li 0005, Min Zhang 0005 |
NeurIPS | 2 |
| 2025 | Thinking in Character: Advancing Role-Playing Agents with Role-Aware ReasoningabstractThe advancement of Large Language Models (LLMs) has spurred significant interest in Role-Playing Agents (RPAs) for applications such as emotional companionship and virtual interaction. However, recent RPAs are often built on explicit dialogue data, lacking deep, human-like internal thought processes, resulting in superficial knowledge and style expression. While Large Reasoning Models (LRMs) can be employed to simulate character thought, their direct application is hindered by attention diversion (i.e., RPAs forget their role) and style drift (i.e., overly formal and rigid reasoning rather than character-consistent reasoning). To address these challenges, this paper introduces a novel Role-Aware Reasoning (RAR) method, which consists of two important stages: Role Identity Activation (RIA) and Reasoning Style Optimization (RSO). RIA explicitly guides the model with character profiles during reasoning to counteract attention diversion, and then RSO aligns reasoning style with the character and scene via LRM distillation to mitigate style drift. Extensive experiments demonstrate that the proposed RAR significantly enhances the performance of RPAs by effectively addressing attention diversion and style drift. Yihong Tang, Kehai Chen, Muyun Yang, Zhengyu Niu, Tiejun Zhao, Min Zhang 0005 |
NeurIPS | 2 |
| 2025 | MASTER: Enhancing Large Language Model via Multi-Agent Simulated TeachingabstractInstruction fine-tuning is crucial in NLP tasks, enhancing pretrained models' instruction-following capabilities and task-specific performance. However, obtaining high-quality fine-tuning data for large models is challenging due to data collection difficulties and high production costs. To address this, we propose MASTER, a novel data augmentation method that enriches original data through interactions among multiple agents with varying cognitive levels. We simulate three pedagogically grounded teaching scenarios, leveraging multi-agent conversations to generate high-quality teacher-student interaction data. Utilizing MASTER, we construct BOOST-QA, a fine-tuning dataset augmented from existing datasets like Orca-Math-200k, ProcQA, and OpenHermes2.5. Experiments show that models fine-tuned with BOOST-QA perform excellently across multiple benchmarks, demonstrating strong multitask generalization. Notably, MASTER significantly improves models' reasoning abilities in complex tasks, providing valuable insights for future research. Yihong Tang, Kehai Chen, Jie Liu 0001, Min Zhang 0005 |
NeurIPS | 3 |
| 2025 | Exploring the Translation Mechanism of Large Language ModelsabstractWhile large language models (LLMs) demonstrate remarkable success in multilingual translation, their internal core translation mechanisms, even at the fundamental word level, remain insufficiently understood.
To address this critical gap, this work introduces a systematic framework for interpreting the mechanism behind LLM translation from the perspective of computational components.
This paper first proposes subspace-intervened path patching for precise, fine-grained causal analysis, enabling the detection of components crucial to translation tasks and subsequently characterizing their behavioral patterns in human-interpretable terms.
Comprehensive experiments reveal that translation is predominantly driven by a sparse subset of components: specialized attention heads serve critical roles in extracting source language, translation indicators, and positional features, which are then integrated and processed by specific multi-layer perceptrons (MLPs) into intermediary English-centric latent representations before ultimately yielding the final translation.
The significance of these findings is underscored by the empirical demonstration that targeted fine-tuning a minimal parameter subset (<5%) enhances translation performance while preserving general capabilities. This result further indicates that these crucial components generalize effectively to sentence-level translation and are instrumental in elucidating more intricate translation tasks. Kehai Chen, Xuefeng Bai 0001, Xiucheng Li, Yang Xiang 0003, Min Zhang 0005 |
NeurIPS | 2 |
| 2025 | Unified Transferability Metrics for Time Series Foundation ModelsabstractWith the increasing number of time series pre-trained models, designing transferability evaluation metrics for time series has become an urgent problem to address.
While transferability evaluation has been extensively studied in computer vision, we aim to address a critical gap by developing tailored metrics for time series analysis.
In this paper, we introduce TEMPLATE, a transferability estimation framework specifically tailored for versatile time series analysis, comprising three complementary metrics: (1) Dependency Learning Score quantifies a model’s capacity to capture temporal dependencies. (2) Pattern Learning Score evaluates the representation quality in extracting discriminative temporal patterns. (3) Task Adaptation Score assesses cross-task generalization capability, enabling versatile time series analysis. TEMPLATE presents a versatile framework compatible with both classification and regression paradigms. Through comprehensive benchmarking across 5 distinct downstream tasks, our method demonstrates superior capability in identifying optimal pre-trained models from heterogeneous model pools for transfer learning. Compared to the state-of-the-art method ETran, our approach improves the weighted Kendall's $\tau_w$ across 5 downstream tasks by 35\%. The code is available at https://github.com/ooooooover/TEMPLATE. Weiyang Zhang, Xinyang Chen 0001, Xiucheng Li, Kehai Chen, Weili Guan, Liqiang Nie |
NeurIPS | 4 |
| 2025 | Adaptive Inner Speech Text Alignment for LLM-Based Speech Translation
Henglyu Liu, Andong Chen 0001, Kehai Chen, Xuefeng Bai 0001, Meizhi Zhong, Yuan Qiu 0001, Min Zhang 0005 |
NLPCC (3) | 3 |
| 2025 | TianWen: A Comprehensive Benchmark for Evaluating LLMs in Chinese Classical Poetry Understanding and Reasoning
Zhenwu Pei, Rongbo Chen, Xuefeng Bai 0001, Kehai Chen, Yingjie Zhu, Andong Chen 0001, Min Zhang 0005 |
NLPCC (1) | 4 |
| 2025 | TF-Attack: Transferable and fast adversarial attacks on large language models
Kehai Chen, Lemao Liu, Xuefeng Bai 0001, Yang Xiang 0003, Min Zhang 0005 |
Knowl. Based Syst. | 2 |
| 2024 | Speech Sense Disambiguation: Tackling Homophone Ambiguity in End-to-End Speech TranslationabstractEnd-to-end speech translation (ST) presents notable disambiguation challenges as it necessitates simultaneous cross-modal and crosslingual transformations.While word sense disambiguation is an extensively investigated topic in textual machine translation, the exploration of disambiguation strategies for ST models remains limited.Addressing this gap, this paper introduces the concept of speech sense disambiguation (SSD), specifically emphasizing homophones -words pronounced identically but with different meanings.To facilitate this, we first create a comprehensive homophone dictionary and an annotated dataset rich with homophone information established based on speech-text alignment.Building on this unique dictionary, we introduce AmbigST, an innovative homophone-aware contrastive learning approach that integrates a homophone-aware masking strategy.Our experiments on different MuST-C and CoVoST ST benchmarks demonstrate that AmbigST sets new performance standards.Specifically, it achieves SOTA results on BLEU scores for English to German, Spanish, and French ST tasks, underlining its effectiveness in reducing speech sense ambiguity. Tengfei Yu, Xuebo Liu 0002, Liang Ding 0006, Kehai Chen, Dacheng Tao, Min Zhang 0005 |
ACL (1) | 4 |
| 2024 | Context Consistency between Training and Inference in Simultaneous Machine TranslationabstractSimultaneous Machine Translation (SiMT) aims to yield a real-time partial translation with a monotonically growing source-side context.However, there is a counterintuitive phenomenon about the context usage between training and inference: e.g., in wait-k inference, model consistently trained with wait-k is much worse than that model inconsistently trained with wait-k ′ (k ′ ̸ = k) in terms of translation quality.To this end, we first investigate the underlying reasons behind this phenomenon and uncover the following two factors: 1) the limited correlation between translation quality and training loss; 2) exposure bias between training and inference.Based on both reasons, we then propose an effective training approach called context consistency training accordingly, which encourages consistent context usage between training and inference by optimizing translation quality and latency as bi-objectives and exposing the predictions to the model during the training.The experiments on three language pairs demonstrate that our SiMT system encouraging context consistency outperforms existing SiMT systems with context inconsistency for the first time.1 Meizhi Zhong, Lemao Liu, Kehai Chen, Min Zhang 0005 |
ACL (1) | 3 |
| 2024 | EmoCRT: An Emotion-Cause Relation Enhanced Model for Causal Emotion Entailment
Zhilong Zhao, Bufan Xu, Muyun Yang, Kehai Chen, Tiejun Zhao |
NLPCC (5) | 5 |
| 2024 | Modeling Inter-Aspect Relations With Clause and Contrastive Learning for Aspect-Based Sentiment AnalysisabstractAspect-based sentiment analysis (ABSA) is a fine-grained sentiment analysis task that aims to identify the sentiment polarity of the given aspect. Recent studies fail to establish the relation among multiple aspects in one sentence. To address this issue, a clause-level relational graph attention network with contrastive learning (CLRCL) model is proposed. Specifically, the given sentence is segmented into clauses to obtain the relation between two aspects based on clause-level interaction. Then, to integrate multiple-aspect information, a clause-level relational graph which contains all aspects and inter-aspect relations is developed. Notably, to precisely learn the inter-aspect relations, the supervised contrastive learning strategy is used. Experimental results reveal that the proposed model is a competitive alternative compared with the state-of-the-art methods. Zhixun Qiu, Kehai Chen, Yun Xue 0002, Zhengxuan Zhang |
IEEE Trans. Comput. Soc. Syst. | 2 |
| 2023 | Improving Translation Quality Estimation with Bias MitigationabstractState-of-the-art translation Quality Estimation (QE) models are proven to be biased.More specifically, they over-rely on monolingual features while ignoring the bilingual semantic alignment.In this work, we propose a novel method to mitigate the bias of the QE model and improve estimation performance.Our method is based on the contrastive learning between clean and noisy sentence pairs.We first introduce noise to the target side of the parallel sentence pair, forming the negative samples.With the original parallel pairs as the positive sample, the QE model is contrastively trained to distinguish the positive samples from the negative ones.This objective is jointly trained with the regression-style quality estimation, so as to prevent the QE model from overfitting to monolingual features.Experiments on WMT QE evaluation datasets demonstrate that our method improves the estimation performance by a large margin while mitigating the bias 1 . Hui Huang 0021, Shuangzhi Wu, Kehai Chen, Hui Di, Muyun Yang, Tiejun Zhao |
ACL (1) | 3 |
| 2023 | PromptST: Abstract Prompt Learning for End-to-End Speech TranslationabstractAn end-to-end speech-to-text (S2T) translation model is usually initialized from a pretrained speech recognition encoder and a pretrained text-to-text (T2T) translation decoder.Although this straightforward setting has been shown empirically successful, there do not exist clear answers to the research questions: 1) how are speech and text modalities fused in S2T model and 2) how to better fuse the two modalities?In this paper, we take the first step toward understanding the fusion of speech and text features in S2T model.We first design and release a 10GB linguistic probing benchmark, namely Speech-Senteval, to investigate the acoustic and linguistic behaviors of S2T models.Preliminary analysis reveals that the uppermost encoder layers of the S2T model can not learn linguistic knowledge efficiently, which is crucial for accurate translation.Based on the finding, we further propose a simple plug-in prompt-learning strategy on the uppermost encoder layers to broaden the abstract representation power of the encoder of S2T models.We call such a promptenhanced S2T model PromptST.Experimental results on four widely-used S2T datasets show that PromptST can deliver significant improvements over a strong baseline by capturing richer linguistic knowledge.Benchmarks, Tengfei Yu, Liang Ding 0006, Xuebo Liu 0002, Kehai Chen, Meishan Zhang, Dacheng Tao, Min Zhang 0005 |
EMNLP | 4 |
| 2023 | INFORM : Information eNtropy based multi-step reasoning FOR large language ModelsabstractLarge language models (LLMs) have demonstrated exceptional performance in reasoning tasks with dedicated Chain-of-Thought (CoT) prompts.Further enhancing CoT prompts with exquisite exemplars can significantly improve reasoning performance.However, the effectiveness of CoT prompts may fluctuate dramatically with different choices of in-context examples.Additionally, manual construction of rationale steps can be time-consuming, presenting challenges for the widespread adoption of CoT prompting.In this work, we propose a novel approach by introducing information entropy (IE) as a criteria on for CoT prompt selection.We extend this criterion to the CoT generation and inference stages, automatically generating CoT prompts with higher information entropy scores and adaptively determining the number of samples.These three stages together form our proposed information entropy based multi-step reasoning for large language models, named INFORM.Our experiments across seven reasoning benchmarks utilizing two language models(GPT-3.5-Turboand text-davinci-003) demonstrate the superiority of INFORM both in performance and efficiency. 1 Chuyue Zhou, Wangjie You, Juntao Li 0005, Kehai Chen, Min Zhang 0005 |
EMNLP | 5 |
| 2023 | Universal Multimodal Representation for Language UnderstandingabstractRepresentation learning is the foundation of natural language processing (NLP). This work presents new methods to employ visual information as assistant signals to general NLP tasks. For each sentence, we first retrieve a flexible number of images either from a light topic-image lookup table extracted over the existing sentence-image pairs or a shared cross-modal embedding space that is pre-trained on out-of-shelf text-image pairs. Then, the text and images are encoded by a Transformer encoder and convolutional neural network, respectively. The two sequences of representations are further fused by an attention layer for the interaction of the two modalities. In this study, the retrieval process is controllable and flexible. The universal visual representation overcomes the lack of large-scale bilingual sentence-image pairs. Our method can be easily applied to text-only tasks without manually annotated multimodal parallel corpora. We apply the proposed method to a wide range of natural language generation and understanding tasks, including neural machine translation, natural language inference, and semantic similarity. Experimental results show that our method is generally effective for different tasks and languages. Analysis indicates that the visual signals enrich textual representations of content words, provide fine-grained grounding information about the relationship between concepts and events, and potentially conduce to disambiguation. Zhuosheng Zhang 0001, Kehai Chen, Rui Wang 0015, Masao Utiyama, Eiichiro Sumita, Zuchao Li, Hai Zhao 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2023 | Document-Level Relation Extraction with Path ReasoningabstractDocument-level relation extraction (DocRE) aims to extract relations among entities across multiple sentences within a document by using reasoning skills (i.e., pattern recognition, logical reasoning, coreference reasoning, etc.) related to the reasoning paths between two entities. However, most of the advanced DocRE models only attend to the feature representations of two entities to determine their relation, and do not consider one complete reasoning path from one entity to another entity, which may hinder the accuracy of relation extraction. To address this issue, this article proposes a novel method to capture this reasoning path from one entity to another entity, thereby better simulating reasoning skills to classify relation between two entities. Furthermore, we introduce an additional attention layer to summarize multiple reasoning paths for further enhancing the performance of the DocRE model. Experimental results on a large-scale document-level dataset show that the proposed approach achieved a significant performance improvement on a strong heterogeneous graph-based baseline. Kehai Chen, Tiejun Zhao |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 2 |
| 2022 | Effective Graph Context Representation for Document-level Machine TranslationabstractDocument-level neural machine translation (DocNMT) universally encodes several local sentences or the entire document. Thus, DocNMT does not consider the relevance of document-level contextual information, for example, some context (i.e., content words, logical order, and co-occurrence relation) is more effective than another auxiliary context (i.e., functional and auxiliary words). To address this issue, we first utilize the word frequency information to recognize content words in the input document, and then use heuristical relations to summarize content words and sentences as a graph structure without relying on external syntactic knowledge. Furthermore, we apply graph attention networks to this graph structure to learn its feature representation, which allows DocNMT to more effectively capture the document-level context. Experimental results on several widely-used document-level benchmarks demonstrated the effectiveness of the proposed approach. Kehai Chen, Muyun Yang, Masao Utiyama, Eiichiro Sumita, Rui Wang 0015, Min Zhang 0005 |
IJCAI | 1 |
| 2022 | Document-Level Relation Extraction with Sentences Importance Estimation and FocusingabstractDocument-level relation extraction (DocRE) aims to determine the relation between two entities from a document of multiple sentences.Recent studies typically represent the entire document by sequence-or graph-based models to predict the relations of all entity pairs.However, we find that such a model is not robust and exhibits bizarre behaviors: it predicts correctly when an entire test document is fed as input, but errs when non-evidence sentences are removed.To this end, we propose a Sentence Importance Estimation and Focusing (SIEF) framework for DocRE, where we design a sentence importance score and a sentence focusing loss, encouraging DocRE models to focus on evidence sentences.Experimental results on two domains show that our SIEF not only improves overall performance, but also makes DocRE models more robust.Moreover, SIEF is a general framework, shown to be effective when combined with a variety of base DocRE models. 1 Kehai Chen, Lili Mou, Tiejun Zhao |
NAACL-HLT | 2 |
| 2022 | Text Compression-Aided Transformer EncodingabstractText encoding is one of the most important steps in Natural Language Processing (NLP). It has been done well by the self-attention mechanism in the current state-of-the-art Transformer encoder, which has brought about significant improvements in the performance of many NLP tasks. Though the Transformer encoder may effectively capture general information in its resulting representations, the backbone information, meaning the gist of the input text, is not specifically focused on. In this paper, we propose explicit and implicit text compression approaches to enhance the Transformer encoding and evaluate models using this approach on several typical downstream tasks that rely on the encoding heavily. Our explicit text compression approaches use dedicated models to compress text, while our implicit text compression approach simply adds an additional module to the main model to handle text compression. We propose three ways of integration, namely backbone source-side fusion, target-side fusion, and both-side fusion, to integrate the backbone information into Transformer-based models for various downstream tasks. Our evaluation on benchmark datasets shows that the proposed explicit and implicit text compression approaches improve results in comparison to strong baselines. We therefore conclude, when comparing the encodings to the baseline models, text compression helps the encoders to learn better language representations. Zuchao Li, Zhuosheng Zhang 0001, Hai Zhao 0001, Rui Wang 0015, Kehai Chen, Masao Utiyama, Eiichiro Sumita |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2022 | Advancing Chinese Event Detection via Revisiting Character InformationabstractRecently, character information has been successfully introduced into the encoder-decoder event detection model to relieve the trigger-word mismatch problem, thus achieving impressive results in the languages without natural delimiters (i.e., Chinese). However, it is introduced into the encoder or the decoder separately, which makes the advantage of character information not be captured and represented adequately for event detection. In this article, we proposed a novel method to model character information in both the encoding and decoding stages to advance the neural event detection model. In particular, the proposed method can encode both words and characters and predict their event types jointly and further leverage interactions between word and its characters to optimize the inference. Experimental results show that the proposed model outperforms previous event detection methods on the ACE2005 Chinese benchmark. We release our code at Github. 1 Yanxia Qin, Yue Zhang 0004, Kehai Chen, Min Zhang 0005 |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 4 |
| 2022 | Integrating Prior Translation Knowledge Into Neural Machine TranslationabstractNeural machine translation (NMT), which is an encoder-decoder joint neural language model with an attention mechanism, has achieved impressive results on various machine translation tasks in the past several years. However, the language model attribute of NMT tends to produce fluent yet sometimes unfaithful translations, which hinders the improvement of translation capacity. In response to this problem, we propose a simple and efficient method to integrate prior translation knowledge into NMT in a universal manner that is compatible with neural networks. Meanwhile, it enables NMT to consider the crossing language translation knowledge from the source-side of the training pipeline of NMT, thereby making full use of the prior translation knowledge to enhance the performance of NMT. The experimental results on two large-scale benchmark translation tasks demonstrated that our approach achieved a significant improvement over a strong baseline. Kehai Chen, Rui Wang 0015, Masao Utiyama, Eiichiro Sumita |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2022 | Data-Driven Fuzzy Target-Side Representation for Intelligent Translation SystemabstractThe encoder–decoder framework has been widely used in various practical artificial intelligence cyber-physical systems, including intelligent translation systems. The decoding process in such a framework usually demands the target-side representation, which is often learned by an autoaggressive decoder to simulate the target context information at the current time-step. However, the autoaggressive decoder only captures the previously generated partial target fragment and fails in simulating the global contextual information. In this article, we propose a new data-driven fuzzy context representation strategy to simulate the global target information. Specifically, we design two fuzzy methods to the global target contextual information, which are bag-of-words of target language generated via a softmax layer from the source-side representation and whole target sentence retrieved from the translation memory according to the source-side representation. Both methods facilitate the autoaggressive decoder to handle the global target context at the current time-step, thereby learning a more effective context vector for the generation of target translation. Extensive experiments on two machine translation tasks demonstrated that the proposed method achieved 3% improvement of BLEU score over a strong baseline. Kehai Chen, Muyun Yang, Tiejun Zhao, Min Zhang 0005 |
IEEE Trans. Fuzzy Syst. | 1 |
| 2022 | A Pattern Driven Graph Ranking Approach to Attribute Extraction for Knowledge GraphabstractAttribution extraction refers to find the attributes for the instances of a given semantic class, which is essential to enhance the schema of a knowledge graph. To facilitate the attribution extraction from the query log, this article proposes a pattern driven graph ranking approach to jointly employ the pattern and context distribution information. First, a simple pattern on query text is applied to automatically acquire seed attributes. Then, a graph-based weight propagation is designed to rank the patterns by context distribution algorithm information. Experimental results show that, on a Chinese query log collected by Baidu, the automatically acquired seeds are more representative than the classical manually assembled seeds, achieving an improvement of 11.6% in MAP as compared to the baseline approach. And the graph-based ranking algorithm manipulates the two types of evidence more effectively, outperforming both the distributional similarity based baseline and the HITS algorithm by 29.2% and 11.3%, respectively. Muyun Yang, Kehai Chen, Shu-Qi Sun, Zhongyuan Han, Leilei Kong, Qingye Meng |
IEEE Trans. Ind. Informatics | 2 |
| 2021 | Document-Level Relation Extraction with ReconstructionabstractIn document-level relation extraction (DocRE), graph structure is generally used to encode relation information in the input document to classify the relation category between each entity pair, and has greatly advanced the DocRE task over the past several years. However, the learned graph representation universally models relation information between all entity pairs regardless of whether there are relationships between these entity pairs. Thus, those entity pairs without relationships disperse the attention of the encoder-classifier DocRE for ones with relationships, which may further hind the improvement of DocRE. To alleviate this issue, we propose a novel encoder-classifier-reconstructor model for DocRE. The reconstructor manages to reconstruct the ground-truth path dependencies from the graph representation, to ensure that the proposed DocRE model pays more attention to encode entity pairs with relationships in the training. Furthermore, the reconstructor is regarded as a relationship indicator to assist relation classification in the inference, which can further improve the performance of DocRE model. Experimental results on a large-scale DocRE dataset show that the proposed model can significantly improve the accuracy of relation extraction on a strong heterogeneous graph-based baseline. The code is publicly available at https://github.com/xwjim/DocRE-Rec. Kehai Chen, Tiejun Zhao |
AAAI | 2 |
| 2021 | Self-Training for Unsupervised Neural Machine Translation in Unbalanced Training Data ScenariosabstractHaipeng Sun, Rui Wang, Kehai Chen, Masao Utiyama, Eiichiro Sumita, Tiejun Zhao. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Haipeng Sun, Rui Wang 0015, Kehai Chen, Masao Utiyama, Eiichiro Sumita, Tiejun Zhao |
NAACL-HLT | 3 |
| 2021 | Context-aware positional representation for self-attention networks
Kehai Chen, Rui Wang 0015, Masao Utiyama, Eiichiro Sumita |
Neurocomputing | 1 |
| 2021 | Unsupervised Neural Machine Translation for Similar and Distant Language Pairs: An Empirical StudyabstractUnsupervised neural machine translation (UNMT) has achieved remarkable results for several language pairs, such as French–English and German–English. Most previous studies have focused on modeling UNMT systems; few studies have investigated the effect of UNMT on specific languages. In this article, we first empirically investigate UNMT for four diverse language pairs (French/German/Chinese/Japanese–English). We confirm that the performance of UNMT in translation tasks for similar language pairs (French/German–English) is dramatically better than for distant language pairs (Chinese/Japanese–English). We empirically show that the lack of shared words and different word orderings are the main reasons that lead UNMT to underperform in Chinese/Japanese–English. Based on these findings, we propose several methods, including artificial shared words and pre-ordering, to improve the performance of UNMT for distant language pairs. Moreover, we propose a simple general method to improve translation performance for all these four language pairs. The existing UNMT model can generate a translation of a reasonable quality after a few training epochs owing to a denoising mechanism and shared latent representations. However, learning shared latent representations restricts the performance of translation in both directions, particularly for distant language pairs, while denoising dramatically delays convergence by continuously modifying the training data. To avoid these problems, we propose a simple, yet effective and efficient, approach that (like UNMT) relies solely on monolingual corpora: pseudo-data-based unsupervised neural machine translation. Experimental results for these four language pairs show that our proposed methods significantly outperform UNMT baselines. Haipeng Sun, Rui Wang 0015, Masao Utiyama, Benjamin Marie, Kehai Chen, Eiichiro Sumita, Tiejun Zhao |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 5 |
| 2021 | Modeling Future Cost for Neural Machine TranslationabstractExisting neural machine translation (NMT) systems utilize sequence-to-sequence neural networks to generate target translation word by word, and then make the generated word at each time-step and the counterpart in the references as consistent as possible. However, the trained translation model tends to focus on ensuring the accuracy of the generated target word at the current time-step and does not consider its future cost which means the expected cost of generating the subsequent target translation (i.e., the next target word). To respond to this issue, in this article, we propose a simple and effective method to model the future cost of each target word for NMT systems. In detail, a future cost representation is learned based on the current generated target word and its contextual information to compute an additional loss to guide the training of the NMT model. Furthermore, the learned future cost representation at the current time-step is used to help the generation of the next target word in the decoding. Experimental results on three widely-used translation datasets, including the WMT14 English-to-German, WMT14 English-to-French, and WMT17 Chinese-to-English, show that the proposed approach achieves significant improvements over strong Transformer-based NMT baseline. Chaoqun Duan, Kehai Chen, Rui Wang 0015, Masao Utiyama, Eiichiro Sumita, Conghui Zhu, Tiejun Zhao |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2020 | Explicit Sentence Compression for Neural Machine TranslationabstractState-of-the-art Transformer-based neural machine translation (NMT) systems still follow a standard encoder-decoder framework, in which source sentence representation can be well done by an encoder with self-attention mechanism. Though Transformer-based encoder may effectively capture general information in its resulting source sentence representation, the backbone information, which stands for the gist of a sentence, is not specifically focused on. In this paper, we propose an explicit sentence compression method to enhance the source sentence representation for NMT. In practice, an explicit sentence compression goal used to learn the backbone information in a sentence. We propose three ways, including backbone source-side fusion, target-side fusion, and both-side fusion, to integrate the compressed sentence into NMT. Our empirical tests on the WMT English-to-French and English-to-German translation tasks show that the proposed sentence compression method significantly improves the translation performances over strong baselines. Zuchao Li, Rui Wang 0015, Kehai Chen, Masao Utiyama, Eiichiro Sumita, Zhuosheng Zhang 0001, Hai Zhao 0001 |
AAAI | 3 |
| 2020 | Content Word Aware Neural Machine TranslationabstractNeural machine translation (NMT) encodes the source sentence in a universal way to generate the target sentence word-byword.However, NMT does not consider the importance of word in the sentence meaning, for example, some words (i.e., content words) express more important meaning than others (i.e., function words).To address this limitation, we first utilize word frequency information to distinguish between content and function words in a sentence, and then design a content word-aware NMT to improve translation performance.Empirical results on the WMT14 English-to-German, WMT14 English-to-French, and WMT17 Chineseto-English translation tasks show that the proposed methods can significantly improve the performance of Transformer-based NMT. Kehai Chen, Rui Wang 0015, Masao Utiyama, Eiichiro Sumita |
ACL | 1 |
| 2020 | Knowledge Distillation for Multilingual Unsupervised Neural Machine TranslationabstractUnsupervised neural machine translation (UNMT) has recently achieved remarkable results for several language pairs. However, it can only translate between a single language pair and cannot produce translation results for multiple language pairs at the same time. That is, research on multilingual UNMT has been limited. In this paper, we empirically introduce a simple method to translate between thirteen languages using a single encoder and a single decoder, making use of multilingual data to improve UNMT for all language pairs. On the basis of the empirical findings, we propose two knowledge distillation methods to further enhance multilingual UNMT performance. Our experiments on a dataset with English translated to and from twelve other languages (including three language families and six language branches) show remarkable results, surpassing strong unsupervised individual baselines while achieving promising performance between non-English language pairs in zero-shot translation scenarios and alleviating poor performance in low-resource language pairs. Haipeng Sun, Rui Wang 0015, Kehai Chen, Masao Utiyama, Eiichiro Sumita, Tiejun Zhao |
ACL | 3 |
| 2020 | Robust Unsupervised Neural Machine Translation with Adversarial Denoising TrainingabstractUnsupervised neural machine translation (UNMT) has recently attracted great interest in the machine translation community.The main advantage of the UNMT lies in its easy collection of required large training text sentences while with only a slightly worse performance than supervised neural machine translation which requires expensive annotated translation pairs on some translation tasks.In most studies, the UMNT is trained with clean data without considering its robustness to the noisy data.However, in real-world scenarios, there usually exists noise in the collected input sentences which degrades the performance of the translation system since the UNMT is sensitive to the small perturbations of the input sentences.In this paper, we first time explicitly take the noisy data into consideration to improve the robustness of the UNMT based systems.First of all, we clearly defined two types of noises in training sentences, i.e., word noise and word order noise, and empirically investigate its effect in the UNMT, then we propose adversarial training methods with denoising process in the UNMT.Experimental results on several language pairs show that our proposed methods substantially improved the robustness of the conventional UNMT systems in noisy scenarios. Haipeng Sun, Rui Wang 0015, Kehai Chen, Xugang Lu, Masao Utiyama, Eiichiro Sumita, Tiejun Zhao |
COLING | 3 |
| 2020 | Robust Machine Reading Comprehension by Learning Soft labelsabstractNeural models have achieved great success on the task of machine reading comprehension (MRC), which are typically trained on hard labels.We argue that hard labels limit the model capability on generalization due to the label sparseness problem.In this paper, we propose a robust training method for MRC models to address this problem.Our method consists of three strategies, 1) label smoothing, 2) word overlapping, 3) distribution prediction.All of them help to train models on soft labels.We validate our approach on the representative architecture -ALBERT.Experimental results show that our method can greatly boost the baseline with 1% improvement in average, and achieve state-of-the-art performance on NewsQA and QUOREF. Shuangzhi Wu, Muyun Yang, Kehai Chen, Tiejun Zhao |
COLING | 4 |
| 2020 | Neural Machine Translation with Universal Visual Representation
Zhuosheng Zhang 0001, Kehai Chen, Rui Wang 0015, Masao Utiyama, Eiichiro Sumita, Zuchao Li, Hai Zhao 0001 |
ICLR | 2 |
| 2020 | Data-dependent Gaussian Prior Objective for Language Generation
Zuchao Li, Rui Wang 0015, Kehai Chen, Masao Utiyama, Eiichiro Sumita, Zhuosheng Zhang 0001, Hai Zhao 0001 |
ICLR | 3 |
| 2020 | Towards More Diverse Input Representation for Neural Machine TranslationabstractSource input information plays a very important role in the Transformer-based translation system. In practice, word embedding and positional embedding of each word are added as the input representation. Then self-attention networks are used to encode the global dependencies in the input representation to generate a source representation. However, this processing on the source representation only adopts a single source feature and excludes richer and more diverse features such as recurrence features, local features, and syntactic features, which results in tedious representation and thereby hinders the further translation performance improvement. In this paper, we introduce a simple and efficient method to encode more diverse source features into the input representation simultaneously, and thereby learning an effective source representation by self-attention networks. In particular, the proposed grouped strategy is only applied to the input representation layer, to keep the diversity of translation information and the efficiency of the self-attention networks at the same time. Experimental results show that our approach improves the translation performance over the state-of-the-art baselines of Transformer in regard to WMT14 English-to-German and NIST Chinese-to-English machine translation tasks. Kehai Chen, Rui Wang 0015, Masao Utiyama, Eiichiro Sumita, Tiejun Zhao, Muyun Yang, Hai Zhao 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2020 | Unsupervised Neural Machine Translation With Cross-Lingual Language Representation AgreementabstractUnsupervised cross-lingual language representation initialization methods such as unsupervised bilingual word embedding (UBWE) pre-training and cross-lingual masked language model (CMLM) pre-training, together with mechanisms such as denoising and back-translation, have advanced unsupervised neural machine translation (UNMT), which has achieved impressive results on several language pairs, particularly French-English and German-English. Typically, UBWE focuses on initializing the word embedding layer in the encoder and decoder of UNMT, whereas the CMLM focuses on initializing the entire encoder and decoder of UNMT. However, UBWE/CMLM training and UNMT training are independent, which makes it difficult to assess how the quality of UBWE/CMLM affects the performance of UNMT during UNMT training. In this paper, we first empirically explore relationships between UNMT and UBWE/CMLM. The empirical results demonstrate that the performance of UBWE and CMLM has a significant influence on the performance of UNMT. Motivated by this, we propose a novel UNMT structure with cross-lingual language representation agreement to capture the interaction between UBWE/CMLM and UNMT during UNMT training. Experimental results on several language pairs demonstrate that the proposed UNMT models improve significantly over the corresponding state-of-the-art UNMT baselines. Haipeng Sun, Rui Wang 0015, Kehai Chen, Masao Utiyama, Eiichiro Sumita, Tiejun Zhao |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2020 | A Novel Sentence-Level Agreement Architecture for Neural Machine TranslationabstractIn neural machine translation (NMT), there is a natural correspondence between source and target sentences. The traditional NMT method does not explicitly model the translation agreement on sentence-level. In this article, we propose a comprehensive and novel sentence-level agreement architecture to alleviate this problem. It directly minimizes the difference between the representations of the source-side and target-side sentence on sentence-level. First, we compare a variety of sentence representation strategies and propose a “Gated Sum” sentence representation to achieve better sentence semantic information. Then, rather than a single-layer sentence-level agreement architecture, we further propose a multi-layer sentence agreement architecture to make the source and target semantic spaces closer layer by layer. The proposed agreement module can be integrated into NMT as an additional training objective function, and can also be used to enhance the representation of the source-side sentences. Experiments on the NIST Chinese-to-English and the WMT English-to-German translation tasks show that the proposed agreement architecture achieves significant improvements over state-of-the-art baselines, demonstrating the effectiveness and necessity of exploiting sentence-level agreement for NMT. Rui Wang 0015, Kehai Chen, Xing Wang 0007, Tiejun Zhao, Min Zhang 0005 |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2020 | A Hierarchical Clustering Approach to Fuzzy Semantic Representation of Rare Words in Neural Machine TranslationabstractRare words are usually replaced with a singletoken in the current encoder-decoder style of neural machine translation, challenging the translation modeling by an obscured context. In this article, we propose to build a fuzzy semantic representation (FSR) method for rare words through a hierarchical clustering method to group rare words together, and integrate it into the encoder-decoder framework. This hierarchical structure can compensate for the semantic information in both source and target sides, and providing fuzzy context information to capture the semantic of rare words. The introduced FSR can also alleviate the data sparseness, which is the bottleneck in dealing with rare words in neural machine translation. In particular, our method is easily extended to the transformer-based neural machine translation model and learns the FSRs of all in-vocabulary words to enhance the sentence representations in addition to rare words. Our experiments on Chinese-to-English translation tasks confirm a significant improvement in the translation quality brought by the proposed method. Muyun Yang, Shujie Liu 0001, Kehai Chen, Enbo Zhao, Tiejun Zhao |
IEEE Trans. Fuzzy Syst. | 3 |
| 2019 | Neural Machine Translation with Reordering EmbeddingsabstractThe reordering model plays an important role in phrase-based statistical machine translation.However, there are few works that exploit the reordering information in neural machine translation.In this paper, we propose a reordering mechanism to learn the reordering embedding of a word based on its contextual information.These reordering embeddings are stacked together with self-attention networks to learn sentence representation for machine translation.The reordering mechanism can be easily integrated into both the encoder and the decoder in the Transformer translation system.Experimental results on WMT'14 English-to-German, NIST Chinese-to-English, and WAT ASPEC Japanese-to-English translation tasks demonstrate that the proposed methods can significantly improve the performance of the Transformer translation system. Kehai Chen, Rui Wang 0015, Masao Utiyama, Eiichiro Sumita |
ACL (1) | 1 |
| 2019 | Unsupervised Bilingual Word Embedding Agreement for Unsupervised Neural Machine TranslationabstractUnsupervised bilingual word embedding (UBWE), together with other technologies such as back-translation and denoising, has helped unsupervised neural machine translation (UNMT) achieve remarkable results in several language pairs.In previous methods, UBWE is first trained using nonparallel monolingual corpora and then this pre-trained UBWE is used to initialize the word embedding in the encoder and decoder of UNMT.That is, the training of UBWE and UNMT are separate.In this paper, we first empirically investigate the relationship between UBWE and UNMT.The empirical findings show that the performance of UNMT is significantly affected by the performance of UBWE.Thus, we propose two methods that train UNMT with UBWE agreement.Empirical results on several language pairs show that the proposed methods significantly outperform conventional UNMT. Haipeng Sun, Rui Wang 0015, Kehai Chen, Masao Utiyama, Eiichiro Sumita, Tiejun Zhao |
ACL (1) | 3 |
| 2019 | Lattice-Based Transformer Encoder for Neural Machine TranslationabstractNeural machine translation (NMT) takes deterministic sequences for source representations.However, either wordlevel or subword-level segmentations have multiple choices to split a source sequence with different word segmentors or different subword vocabulary sizes.We hypothesize that the diversity in segmentations may affect the NMT performance.To integrate different segmentations with the state-of-the-art NMT model, Transformer, we propose lattice-based encoders to explore effective word or subword representation in an automatic way during training.We propose two methods: 1) lattice positional encoding and 2) lattice-aware self-attention.These two methods can be used together and show complementary to each other to further improve translation performance.Experiment results show superiorities of lattice-based encoders in word-level and subword-level representations over conventional Transformer encoder. Fengshun Xiao, Jiangtong Li, Hai Zhao 0001, Rui Wang 0015, Kehai Chen |
ACL (1) | 5 |
| 2019 | Sentence-Level Agreement for Neural Machine TranslationabstractThe training objective of neural machine translation (NMT) is to minimize the loss between the words in the translated sentences and those in the references. In NMT, there is a natural correspondence between the source sentence and the target sentence. However, this relationship has only been represented using the entire neural network and the training objective is computed in word-level. In this paper, we propose a sentence-level agreement module to directly minimize the difference between the representation of source and target sentence. The proposed agreement module can be integrated into NMT as an additional training objective function and can also be used to enhance the representation of the source sentences. Empirical results on the NIST Chinese-to-English and WMT English-to-German tasks show the proposed agreement module can significantly improve the NMT performance. Rui Wang 0015, Kehai Chen, Masao Utiyama, Eiichiro Sumita, Min Zhang 0005, Tiejun Zhao |
ACL (1) | 3 |
| 2019 | Recurrent Positional Embedding for Neural Machine TranslationabstractKehai Chen, Rui Wang, Masao Utiyama, Eiichiro Sumita. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Kehai Chen, Rui Wang 0015, Masao Utiyama, Eiichiro Sumita |
EMNLP/IJCNLP (1) | 1 |
| 2019 | A SST-Dependent Geophysical Model Function for HY-2A Microwave ScatterometerabstractSatellite scatterometer provides an effective approach to obtaining the ocean surface winds on a global scale. However, none of the geophysical model functions(GMF) currently used in wind retrieval takes the effect of sea surface temperature(SST) into account. Many experiments demonstrate that SST indeed have an impact on the behavior of the microwave backscatter of the ocean surface. This paper utilized the co-located backscatter coefficients from HY-2A scatterometer, re-analysis wind speed and direction data of ECMWF(European Center for Medium-Range Weather Forecast), and SST(sea surface temperature) data acquired by WindSat to examine the dependence of radar backscatter coefficient on SST and then derive a SST-dependent geophysical model function(SST-GMF). The validation experiment results indicate that the SST-GMF leads to a smaller wind speed bias at the low(0-10oC) and the high(20-30oC) SST ranges compared to the NSCAT-2 model, resulting in a higher wind speed retrieval precision. Xuetong Xie, Dongxuan Tian, Kehai Chen, Zhifeng Wu, Songhong Tan |
IGARSS | 3 |
| 2019 | A Bilingual Adversarial Autoencoder for Unsupervised Bilingual Lexicon InductionabstractUnsupervised bilingual lexicon induction aims to generate bilingual lexicons without any cross-lingual signals. Successfully solving this problem would benefit many downstream tasks, such as unsupervised machine translation and transfer learning. In this work, we propose an unsupervised framework, named bilingual adversarial autoencoder, which automatically generates bilingual lexicon for a pair of languages from their monolingual word embeddings. In contrast to existing frameworks which learn a direct cross-lingual mapping of word embeddings from the source language to the target language, we train two autoencoders jointly to transform the source and the target monolingual word embeddings into a shared embedding space, where a word and its translation are close to each other. In this way, we capture the cross-lingual features of word embeddings from different languages and use them to induce bilingual lexicons. By conducting extensive experiments across eight language pairs, we demonstrate that the proposed method significantly outperforms the existing adversarial methods and even achieves best-published results across most language pairs. Xuefeng Bai 0001, Hailong Cao, Kehai Chen, Tiejun Zhao |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2019 | Neural Machine Translation With Sentence-Level Topic ContextabstractTraditional neural machine translation (NMT) methods use the word-level context to predict target language translation while neglecting the sentence-level context, which has been shown to be beneficial for translation prediction in statistical machine translation. This paper represents the sentence-level context as latent topic representations by using a convolution neural network, and designs a topic attention to integrate source sentence-level topic context information into both attention-based and Transformer-based NMT. In particular, our method can improve the performance of NMT by modeling source topics and translations jointly. Experiments on the large-scale LDC Chinese-to-English translation tasks and WMT'14 English-to-German translation tasks show that the proposed approach can achieve significant improvements compared with baseline systems. Kehai Chen, Rui Wang 0015, Masao Utiyama, Eiichiro Sumita, Tiejun Zhao |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2018 | Syntax-Directed Attention for Neural Machine TranslationabstractAttention mechanism, including global attention and local attention, plays a key role in neural machine translation (NMT). Global attention attends to all source words for word prediction. In comparison, local attention selectively looks at fixed-window source words. However, alignment weights for the current target word often decrease to the left and right by linear distance centering on the aligned source position and neglect syntax distance constraints. In this paper, we extend the local attention with syntax-distance constraint, which focuses on syntactically related source words with the predicted target word to learning a more effective context vector for predicting translation. Moreover, we further propose a double context NMT architecture, which consists of a global context vector and a syntax-directed context vector from the global attention, to provide more translation performance for NMT from source representation. The experiments on the large-scale Chinese-to-English and English-to-German translation tasks show that the proposed approach achieves a substantial and significant improvement over the baseline system. Kehai Chen, Rui Wang 0015, Masao Utiyama, Eiichiro Sumita, Tiejun Zhao |
AAAI | 1 |
| 2018 | A Neural Approach to Source Dependence Based Context Model for Statistical Machine TranslationabstractIn statistical machine translation, translation prediction considers not only the aligned source word itself but also its source contextual information. Learning context representation is a promising method for improving translation results, particularly through neural networks. Most of the existing methods process context words sequentially and neglect source long-distance dependencies. In this paper, we propose a novel neural approach to source dependence-based context representation for translation prediction. The proposed model is capable of not only encoding source long-distance dependencies but also capturing functional similarities to better predict translations (i.e., word form translations and ambiguous word translations). To verify our method, the proposed mode is incorporated into phrase-based and hierarchical phrase-based translation models, respectively. Experiments on large-scale Chinese-to-English and English-to-German translation tasks show that the proposed approach achieves significant improvement over the baseline systems and outperforms several existing context-enhanced methods. Kehai Chen, Tiejun Zhao, Muyun Yang, Lemao Liu, Akihiro Tamura, Rui Wang 0015, Masao Utiyama, Eiichiro Sumita |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2018 | Sentence Selection and Weighting for Neural Machine Translation Domain AdaptationabstractNeural machine translation (NMT) has been prominent in many machine translation tasks. However, in some domain-specific tasks, only the corpora from similar domains can improve translation performance. If out-of-domain corpora are directly added into the in-domain corpus, the translation performance may even degrade. Therefore, domain adaptation techniques are essential to solve the NMT domain problem. Most existing methods for domain adaptation are designed for the conventional phrase-based machine translation. For NMT domain adaptation, there have been only a few studies on topics such as fine tuning, domain tags, and domain features. In this paper, we have four goals for sentence level NMT domain adaptation. First, the NMT's internal sentence embedding is exploited and the sentence embedding similarity is used to select out-of-domain sentences that are close to the in-domain corpus. Second, we propose three sentence weighting methods, i.e., sentence weighting, domain weighting, and batch weighting, to balance the data distribution during NMT training. Third, in addition, we propose dynamic training methods to adjust the sentence selection and weighting during NMT training. Fourth, to solve the multidomain problem in a real-world NMT scenario where the domain distributions of training and testing data often mismatch, we proposed a multidomain sentence weighting method to balance the domain distributions of training data and match the domain distributions of training and testing data. The proposed methods are evaluated in international workshop on spoken language translation (IWSLT) English-to-French/German tasks and a multidomain English-to-French task. Empirical results show that the sentence selection and weighting methods can significantly improve the NMT performance, outperforming the existing baselines. Rui Wang 0015, Masao Utiyama, Andrew M. Finch, Lemao Liu, Kehai Chen, Eiichiro Sumita |
IEEE ACM Trans. Audio Speech Lang. Process. | 5 |
| 2017 | Translation Prediction with Source Dependency-Based Context RepresentationabstractLearning context representations is very promising to improve translation results, particularly through neural networks. Previous efforts process the context words sequentially and neglect their internal syntactic structure. In this paper, we propose a novel neural network based on bi-convolutional architecture to represent the source dependency-based context for translation prediction. The proposed model is able to not only encode the long-distance dependencies but also capture the functional similarities for better translation prediction (i.e., ambiguous words translation and word forms translation). Examined by a large-scale Chinese-English translation task, the proposed approach achieves a significant improvement (of up to +1.9 BLEU points) over the baseline system, and meanwhile outperforms a number of context-enhanced comparison system. Kehai Chen, Tiejun Zhao, Muyun Yang, Lemao Liu |
AAAI | 1 |
| 2017 | Neural Machine Translation with Source Dependency RepresentationabstractSource dependency information has been successfully introduced into statistical machine translation.However, there are only a few preliminary attempts for Neural Machine Translation (NMT), such as concatenating representations of source word and its dependency label together.In this paper, we propose a novel attentional NMT with source dependency representation to improve translation performance of NMT, especially on long sentences.Empirical results on NIST Chinese-to-English translation task show that our method achieves 1.6 BLEU improvements on average over a strong NMT system. Kehai Chen, Rui Wang 0015, Masao Utiyama, Lemao Liu, Akihiro Tamura, Eiichiro Sumita, Tiejun Zhao |
EMNLP | 1 |
| 2017 | Instance Weighting for Neural Machine Translation Domain AdaptationabstractInstance weighting has been widely applied to phrase-based machine translation domain adaptation.However, it is challenging to be applied to Neural Machine Translation (NMT) directly, because NMT is not a linear model.In this paper, two instance weighting technologies, i.e., sentence weighting and domain weighting with a dynamic weight learning strategy, are proposed for NMT domain adaptation.Empirical results on the IWSLT English-German/French tasks show that the proposed methods can substantially improve NMT performance by up to 2.7-6.7 BLEU points, outperforming the existing baselines by up to 1.6-3.6BLEU points. Rui Wang 0015, Masao Utiyama, Lemao Liu, Kehai Chen, Eiichiro Sumita |
EMNLP | 4 |
| 2017 | Context-Aware Smoothing for Neural Machine TranslationabstractIn Neural Machine Translation (NMT), each word is represented as a low-dimension, real-value vector for encoding its syntax and semantic information. This means that even if the word is in a different sentence context, it is represented as the fixed vector to learn source representation. Moreover, a large number of Out-Of-Vocabulary (OOV) words, which have different syntax and semantic information, are represented as the same vector representation of “unk”. To alleviate this problem, we propose a novel context-aware smoothing method to dynamically learn a sentence-specific vector for each word (including OOV words) depending on its local context words in a sentence. The learned context-aware representation is integrated into the NMT to improve the translation performance. Empirical results on NIST Chinese-to-English translation task show that the proposed approach achieves 1.78 BLEU improvements on average over a strong attentional NMT, and outperforms some existing systems. Kehai Chen, Rui Wang 0015, Masao Utiyama, Eiichiro Sumita, Tiejun Zhao |
IJCNLP(1) | 1 |
| 2016 | Improving Dependency Parsing on Clinical Text with Syntactic Clusters from Web Text
Xiuming Qiao, Hailong Cao, Tiejun Zhao, Kehai Chen |
ICONIP (1) | 4 |
| 2012 | A study on wind vector retrieval algorithm for rotating fan-beam scatterometerabstractRotating fan-beam scatterometer (RFSCAT) is a new type of satellite scatterometer that was proposed about one decade ago.However, just as other rotating scatterometers, relatively larger wind retrieval errors occur in the nadir and outer regions than in the middle regions of the swath. In order to address this problem, a modified wind vector retrieval algorithm for RFSCAT is presented in this paper. The new algorithm is featured with adaptively extending the range of wind direction for each wind vector cell position across the whole swath according to the distribution histogram of the retrieved wind direction bias. Simulation experiments demonstrated that the new established algorithm can effectively improve the wind direction retrieval accuracy in the nadir and outer regions of the RFSCAT swath. Xuetong Xie, Shi Huan, Jianqiang Liu 0001, Shuyan Lang, Youguang Zhang, Di Zhu 0001, Kehai Chen, Juhong Zou, Zhou Huang 0002, Weijun Tao |
IGARSS | 7 |
| 2012 | A wind direction extension based algorithm for scatterometer wind vector retrievalabstractAccording to the lower efficiency and larger wind direction errors in the nadir region of the swath, a new combined wind retrieval algorithm is proposed for conically scanning scatterometer in this paper. The presented algorithm has the dual advantages of both higher efficiency and higher wind direction retrieval accuracy by combining the wind speed standard deviation algorithm and the wind direction interval retrieval(DIR) algorithm. It adopts wind speed standard deviation as criterion for searching possible wind vector solutions and retrieves potential wind direction interval for the first and second ambiguities based on the change rate of the wind speed standard deviation. Some SeaWinds L2A data and collocated buoy data were used to validate the algorithm. Retrieval experiments indicated that the algorithm can significantly reduce the wind direction retrieval errors in the nadir region. Xuetong Xie, Mingsen Lin, Kehai Chen, Zhou Huang 0002, Dongxuan Tian, Rongrong He, Juhong Zou |
IGARSS | 3 |
| 2008 | A Study on Geophysical Model Function Modeling with Water Surface Temperature as One of the Input ParametersabstractGeophysical model function is the basis for the wind vector retrieval with scatterometer and a number of models have been developed to operationally retrieve the ocean surface wind in the past three decades. However, none of the operational models ever took the water surface temperature into account in its modeling, which is considered to have some effect on the ocean backscattering, and in turn on the model accuracy. Taking Sea Winds as an example, this paper attempts to develop new geophysical model functions with surface temperature to be taken into account by using its level 2A data and corresponding buoy data. For contrast, two independent models are established for the ocean water and fresh water respectively. The modeling results and analysis indicate that some effect of the surface temperature on backscatter were found for both types of water, but with a larger extent of the temperature effect for fresh water. Xuetong Xie, Kehai Chen, Wenxian Yu, Weidong Hu, Qiming Zeng, Yu Fang 0001 |
IGARSS (1) | 2 |
| 2008 | Validation of QSCAT-1 Geophysical Model Function Using Seawinds Level 2 and Buoy DataabstractGeophysical model function(GMF) is the basis and prerequisite for the ocean surface wind vector retrieval with scatterometers. Among many operational models, the Qscat-1 model was specifically developed for SeaWinds scatterometer and is being applied to its operational wind retrieval. This paper is to validate the accuracy of the Qscat-1 model by using some SeaWinds Level 2 data and corresponding buoy data. First, a comparison between L2B and co-located buoy wind speed was made to analyze the systematic bias between them, and then a new geophysical model function was established using the match-ups of the L2A and Buoy data to further valuate the accuracy of the Qscat-1 model. The analytical and modeling results indicate that there may be some systematic error in the Qscat-1 model. Xuetong Xie, Qiming Zeng, Weidong Hu, Wenxian Yu, Kehai Chen, Yu Fang 0001 |
IGARSS (1) | 5 |
| 2005 | A new fast wind vector retrieval algorithm for seawinds scatterometerabstractAccording to the main shortcoming of the traditional Maximum Likelihood Estimation(MLE) based algorithm for its high computational complexity, a new fast wind vector retrieval algorithm for SeaWinds Scatterometer is derived in this paper. The new fast algorithm adopts wind speed standard deviation instead of objective function as its criterion for searching possible wind vector solutions, which leads to lower complexity than traditional MLE algorithm. In order to further reduce retrieval computations, the new algorithm is implemented by a two-step method. First step accomplishes a coarse searching for likely wind vector solutions, while the second step functions as a fine adjustment for each coarse solution. Using some SeaWinds L2A and corresponding L2B data, the new algorithm is validated. The results indicate good performance and high retrieval accuracy of the new algorithm for experiment data. Xuetong Xie, Yu Fang 0001, Xiaoxiang Chen, Kehai Chen |
IGARSS | 4 |