Tiejun Zhao

dblp:35/1787 · also Tie-Jun Zhao · DBLP profile ↗
← Back
158ranked-venue papers
2as first author
59since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 134 · 50 since 2021Graphics, computer vision, multimedia, augmented reality and games · 29 · 15 since 2021Databases, data management, data science and information retrieval · 8 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 5 · 1 first-authorComputer networks · 3 · 2 since 2021
YearPublicationVenuePosition
2026 Long-form RewardBench: Evaluating Reward Models for Long-form Generation
abstract
The widespread adoption of reinforcement learning-based alignment highlights the growing importance of reward models. Various benchmarks have been built to evaluate reward models in various domains and scenarios. However, a significant gap remains in assessing reward models for long-form generation, despite its critical role in real-world applications. To bridge this, we introduce Long-form RewardBench, the first reward modeling testbed specifically designed for long-form generation. Our benchmark encompasses five key subtasks: QA, RAG, Chat, Writing, and Reasoning. We collected instruction and preference data through a meticulously designed multi-stage data collection process, and conducted extensive experiments on 20+ mainstream reward models, including both classifiers and generative models. Our findings reveal that current models still lack long-form reward modeling capabilities. Furthermore, we designed a novel Long-form Needle-in-a-Haystack Test, which revealed a correlation between reward modeling performance and the error's position within a response, as well as the overall response length, with distinct characteristics observed between classification and generative models. Finally, we demonstrate that classifier exhibit better generalizability compared to generative models trained on the same data. As the first benchmark for long-form reward modeling, this work aims to offer a robust platform for visualizing progress in this crucial area.
Hui Huang 0021, Yancheng He, Muyun Yang, Kehai Chen, Conghui Zhu, Hailong Cao, Tiejun Zhao
AAAI10
2026 Lost in Benchmarks? Rethinking Large Language Model Benchmarking with Item Response Theory
abstract
The evaluation of large language models (LLMs) via benchmarks is widespread, yet inconsistencies between different leaderboards and poor separability among top models raise concerns about their ability to accurately reflect authentic model capabilities. This paper provides a critical analysis of benchmark effectiveness, examining mainstream prominent LLM benchmarks using results from diverse models. We first propose Pseudo-Siamese Network for Item Response Theory (PSN-IRT), an enhanced Item Response Theory framework that incorporates a rich set of item parameters within an IRT-grounded architecture. PSN-IRT can be utilized for accurate and reliable estimations of item characteristics and model abilities. Based on PSN-IRT, we conduct extensive analysis on 11 LLM benchmarks comprising 41,871 items, revealing significant and varied shortcomings in their measurement quality. Furthermore, we demonstrate that leveraging PSN-IRT is able to construct smaller benchmarks while maintaining stronger alignment with human preference.
Hongli Zhou 0001, Hui Huang 0021, Ziqing Zhao, Lvyuan Han, Huicheng Wang, Kehai Chen, Muyun Yang, Conghui Zhu, Hailong Cao, Tiejun Zhao
AAAI13
2026 Diagnosing and Remedying Representation Deficiencies for Deterministic Reasoning in KGQA
abstract
Gewen Liang, Mufan Xu, Kehai Chen, Wei Wang, Yuwei Wang, Muyun Yang, Tiejun Zhao, Min Zhang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Gewen Liang, Mufan Xu, Kehai Chen, Wei Wang 0164, Muyun Yang, Tiejun Zhao, Min Zhang 0005
ACL (1)7
2026 A multi-task learning framework for integrated assessment in agricultural applications
Yuexing Han, Jiahao Ge, Tiejun Zhao, Qiaochuan Chen
Inf. Sci.4
2026 ASMem: Anchor sparse memory for multi-domain knowledge editing of large language models
Guanyu Zheng, Xv Wang, Haochang Wang, Tiejun Zhao, Chengqing Zong
Neural Networks7
2025 Look Before You Leap: Enhance Attention and Vigilance Regarding Harmful Content with GuidelineLLM
abstract
Despite being empowered with alignment mechanisms, large language models (LLMs) are increasingly vulnerable to emerging jailbreak attacks that can compromise their alignment mechanisms. This vulnerability poses significant risks to real-world applications. Existing work faces challenges in both training efficiency and generalization capabilities (i.e., Reinforcement Learning from Human Feedback and Red-Teaming). Developing effective strategies to enable LLMs to resist continuously evolving jailbreak attempts represents a significant challenge. To address this challenge, we propose a novel defensive paradigm called GuidelineLLM, which assists LLMs in recognizing queries that may have harmful content. Before LLMs respond to a query, GuidelineLLM first identifies potential risks associated with the query, summarizes these risks into guideline suggestions, and then feeds these guidelines to the responding LLMs. Importantly, our approach eliminates the necessity for additional safety fine-tuning of the LLMs themselves; only the GuidelineLLM requires fine-tuning. This characteristic enhances the general applicability of GuidelineLLM across various LLMs. Experimental results demonstrate that GuidelineLLM can significantly reduce the attack success rate (ASR) against LLM (an average reduction of 34.17% ASR) while maintaining the usefulness of LLM in handling benign queries.
Shaoqing Zhang, Zhuosheng Zhang 0001, Kehai Chen, Rongxiang Weng, Muyun Yang, Tiejun Zhao, Min Zhang 0005
AAAI6
2025 Make Imagination Clearer! Stable Diffusion-based Visual Imagination for Multimodal Machine Translation
abstract
Andong Chen, Yuchen Song, Kehai Chen, Xuefeng Bai, Muyun Yang, Liqiang Nie, Jie Liu, Tiejun Zhao, Min Zhang. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Andong Chen 0001, Kehai Chen, Xuefeng Bai 0001, Muyun Yang, Liqiang Nie, Jie Liu 0001, Tiejun Zhao, Min Zhang 0005
ACL (1)8
2025 MuSC: Improving Complex Instruction Following with Multi-granularity Self-Contrastive Training
abstract
Hui Huang, Jiaheng Liu, Yancheng He, Shilong Li, Bing Xu, Conghui Zhu, Muyun Yang, Tiejun Zhao. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Hui Huang 0021, Yancheng He, Conghui Zhu, Muyun Yang, Tiejun Zhao
ACL (1)8
2025 Word-level Cross-lingual Structure in Large Language Models
abstract
Large Language Models (LLMs) have demonstrated exceptional performance across a broad spectrum of cross-lingual Natural Language Processing (NLP) tasks. However, previous methods predominantly focus on leveraging parallel corpus to conduct instruction data for continuing pre-training or fine-tuning. They ignored the state of parallel data on the hidden layers of LLMs. In this paper, we demonstrate Word-level Cross-lingual Structure (WCS) of LLM which proves that the word-level embedding on the hidden layers are isomorphic between languages. We find that the hidden states of different languages’ input on the LLMs hidden layers can be aligned with an orthogonal matrix on word-level. We prove this conclusion in both mathematical and downstream task ways on two representative LLM foundations, LLaMA2 and BLOOM. Besides, we propose an Isomorphism-based Data Augmentation (IDA) method to apply the WCS on a downstream cross-lingual task, Bilingual Lexicon Induction (BLI), in both supervised and unsupervised ways. The experiment shows the significant improvement of our proposed method over all the baselines, especially on low-resource languages.
Hailong Cao, Tiejun Zhao
COLING4
2025 A Chain-of-Task Framework for Instruction Tuning of LLMs Based on Chinese Grammatical Error Correction
abstract
Over-correction is a critical issue for large language models (LLMs) to address Grammatical Error Correction (GEC) task, esp. for Chinese. This paper proposes a Chain-of-Task (CoTask) framework to reduce over-correction. The CoTask framework is applied as multi-task instruction tuning of LLMs by decomposing the process of grammatical error analysis to design auxiliary tasks and adjusting the types and combinations of training tasks. A supervised fine-tuning (SFT) strategy is also presented to enhance the performance of LLMs, together with an algorithm for automatic dataset annotation to avoid additional manual costs. Experimental results demonstrate that our method achieves new state-of-the-art results on both FCGEC (in-domain) and NaCGEC (out-of-domain) test sets.
Xinpeng Liu 0008, Muyun Yang, Hailong Cao, Conghui Zhu, Tiejun Zhao, Wenpeng Lu
COLING6
2025 LoRA-drop: Efficient LoRA Parameter Pruning based on Output Evaluation
abstract
Low-Rank Adaptation (LoRA) is currently the most commonly used Parameter-efficient fine-tuning (PEFT) method. However, it still faces high computational and storage costs to models with billions of parameters. Most previous studies have tackled this issue by using pruning techniques. Nonetheless, these efforts only analyze LoRA parameter features to evaluate their importance, such as parameter count, size, and gradient. In fact, the output of LoRA directly impacts the fine-tuned model. Preliminary experiments indicate that a fraction of LoRA possesses significantly high output values, substantially influencing the layer output. Motivated by the observation, we propose LoRA-drop. Concretely, LoRA-drop evaluates the importance of LoRA based on the LoRA output. Then we retain LoRA for important layers and the other layers share the same LoRA. We conduct abundant experiments with models of different scales on NLU and NLG tasks. Results demonstrate that LoRA-drop can achieve performance comparable to full fine-tuning and LoRA while retaining 50% of the LoRA parameters on average.
Hongyun Zhou, Conghui Zhu, Tiejun Zhao, Muyun Yang
COLING5
2025 Benchmarking LLMs for Translating Classical Chinese Poetry: Evaluating Adequacy, Fluency, and Elegance
abstract
Large language models (LLMs) have shown remarkable performance in general translation tasks.However, the increasing demand for high-quality translations that are not only adequate but also fluent and elegant.To assess the extent to which current LLMs can meet these demands, we introduce a suitable benchmark (PoetMT) for translating classical Chinese poetry into English.This task requires not only adequacy in translating culturally and historically significant content but also a strict adherence to linguistic fluency and poetic elegance.Our study reveals that existing LLMs fall short of this task.To address these issues, we propose RAT, a Retrieval-Augmented machine Translation method that enhances the translation process by incorporating knowledge related to classical poetry.Additionally, we propose an automatic evaluation metric based on GPT-4, which better assesses translation quality in terms of adequacy, fluency, and elegance, overcoming the limitations of traditional metrics.Our dataset and code will be made available 1 .
Andong Chen 0001, Lianzhang Lou, Kehai Chen, Xuefeng Bai 0001, Yang Xiang 0003, Muyun Yang, Tiejun Zhao, Min Zhang 0005
EMNLP7
2025 Self-Relevance-Based Multimodal In-Context Learning for Multimodal Named Entity Recognition
abstract
Recently, Multimodal Named Entity Recognition (MNER) has attracted significant attention. Although MNER utilizing in-context learning has shown improved performance, modality retrieval bias often diminishes the relevance of in-context examples. To address this issue, we propose a self-relevance-based multimodal in-context learning method to mitigate modality retrieval bias by dynamically adjusting the weight of each modality. Specifically, we first measure the self-relevance of the query by calculating the similarity between textual and visual modalities, which helps to assess how much visual information contributes to the textual context. Then, we rank the similarity of different modalities, adjust the image rankings based on self-relevance to reduce modality retrieval bias, and integrate them to select the k most relevant examples. Finally, we use task definition and retrieved examples as effective guidance provided to the Multimodal Large Language Models to obtain feedback. Experimental results demonstrate that our method achieves SOTA performance on two benchmark datasets.
Muyun Yang, Hailong Cao, Conghui Zhu, Wenpeng Lu, Tiejun Zhao
ICME7
2025 MAGE: Multimodal Alignment and Generation Enhancement via Bridging Visual and Semantic Spaces
abstract
In the latest advancements in multimodal learning, effectively addressing the spatial and semantic losses of visual data after encoding remains a critical challenge. This is because the performance of large multimodal models is positively correlated with the coupling between visual encoders and large language models. Existing approaches often face issues such as vector gaps or semantic disparities, resulting in information loss during the propagation process. To address these issues, we propose MAGE (Multimodal Alignment and Generation Enhancement), a novel framework that bridges the semantic spaces of vision and text through an innovative alignment mechanism. By introducing the Intelligent Alignment Network (IAN), MAGE achieves dimensional and semantic alignment. To reduce the gap between synonymous heterogeneous data, we employ a training strategy that combines cross-entropy and mean squared error, significantly enhancing the alignment effect. Moreover, to enhance MAGE’s “Any-to-Any” capability, we developed a fine-tuning dataset for multimodal tool-calling instructions to expand the model’s output capability boundaries. Finally, our proposed multimodal large model architecture, MAGE, achieved significantly better performance compared to similar works across various evaluation benchmarks, including MME, MMBench, and SEED. Complete code and appendix are available at: https://github.com/GTCOM-NLP/MAGE
Shaojun E, Jiaheng Wu, Tiejun Zhao
IJCAI5
2025 Thinking in Character: Advancing Role-Playing Agents with Role-Aware Reasoning
abstract
The advancement of Large Language Models (LLMs) has spurred significant interest in Role-Playing Agents (RPAs) for applications such as emotional companionship and virtual interaction. However, recent RPAs are often built on explicit dialogue data, lacking deep, human-like internal thought processes, resulting in superficial knowledge and style expression. While Large Reasoning Models (LRMs) can be employed to simulate character thought, their direct application is hindered by attention diversion (i.e., RPAs forget their role) and style drift (i.e., overly formal and rigid reasoning rather than character-consistent reasoning). To address these challenges, this paper introduces a novel Role-Aware Reasoning (RAR) method, which consists of two important stages: Role Identity Activation (RIA) and Reasoning Style Optimization (RSO). RIA explicitly guides the model with character profiles during reasoning to counteract attention diversion, and then RSO aligns reasoning style with the character and scene via LRM distillation to mitigate style drift. Extensive experiments demonstrate that the proposed RAR significantly enhances the performance of RPAs by effectively addressing attention diversion and style drift.
Yihong Tang, Kehai Chen, Muyun Yang, Zhengyu Niu, Tiejun Zhao, Min Zhang 0005
NeurIPS6
2025 DuplexMamba: Enhancing Real-Time Speech Conversations with Duplex and Streaming Capabilities
Hongyun Zhou, Conghui Zhu, Tiejun Zhao, Muyun Yang
NLPCC (2)7
2025 Thoughts Behind Attack: Enhancing Security Against Jailbreak Attacks Using Chain-of-Thought
Zhe Tao, Muyun Yang, Hongjiao Guan, Wenpeng Lu, Hailong Cao, Conghui Zhu, Tiejun Zhao
NLPCC (4)8
2025 UCFA-Net: A U-shaped cross-fusion network with attention mechanism for enhanced polyp segmentation
abstract
Abstract Enhancing the precision of computer‐assisted polyp segmentation and delineation during colonoscopies assists in the removal of potentially precancerous tissue, thus reducing the risk of malignant transformation. Most of the current medical segmentation models use the traditional U‐shaped network structure, but they suffer from the problem of information loss during the encoding and decoding of images. To advance towards an autonomous model for detailed polyp segmentation, the authors propose a new framework for polyp segmentation called U‐shaped cross‐fusion network with attention mechanism (UCFA‐Net), which employs a pyramid vision transformer as encoder to extract image features at multiple scales. Furthermore, the multi‐scale cross‐fusion module cross‐fuses the different scale features and then goes through the multi‐scale convolutional parallel feedforward transformer module for modelling the global and local information. Finally, progressive attentional up‐sampling module acts as a decoder for up‐sampling with progressive attention to get the final polyp segmentation result. The authors comprehensive testing demonstrates that their network achieves superior average scores across the five datasets and exhibits greater robustness in the face of diverse and demanding scenarios, when compared to current state‐of‐the‐art approaches.
Shuai Wang 0051, Tiejun Zhao, Guocun Wang
IET Image Process.2
2025 Enhancing bilingual lexicon induction via harnessing polysemous words
Qiuyu Ding, Hailong Cao, Muyun Yang, Tiejun Zhao
Neurocomputing5
2025 Enhancing word distinction for bilingual lexicon induction with generalized antonym knowledge
Qiuyu Ding, Hailong Cao, Tiejun Zhao
Knowl. Based Syst.4
2025 Interactive Conversational Head Generation
abstract
We introduce a new conversation head generation benchmark for synthesizing behaviors of a single interlocutor in a face-to-face conversation. The capability to automatically synthesize interlocutors which can participate in long and multi-turn conversations is vital and offer benefits for various applications, including digital humans, virtual agents, and social robots. While existing research primarily focuses on talking head generation (one-way interaction), hindering the ability to create a digital human for conversation (two-way) interaction due to the absence of listening and interaction parts. In this work, we construct two datasets to address this issue, "ViCo" for independent talking and listening head generation tasks at the sentence level, and "ViCo-X", for synthesizing interlocutors in multi-turn conversational scenarios. Based on ViCo and ViCo-X, we define three novel tasks targeting the interaction modeling during the face-to-face conversation: 1) responsive listening head generation making listeners respond actively to the speaker with non-verbal signals, 2) expressive talking head generation guiding speakers to be aware of listeners' behaviors, and 3) conversational head generation to integrate the talking/listening ability in one interlocutor. Along with the datasets, we also propose corresponding baseline solutions to the three aforementioned tasks. Experimental results show that our baseline method could generate responsive and vivid agents that can collaborate with real person to fulfil the whole conversation.
Mohan Zhou, Yalong Bai, Wei Zhang 0031, Ting Yao 0003, Tiejun Zhao
IEEE Trans. Pattern Anal. Mach. Intell.5
2025 StyleInject: Parameter Efficient Tuning of Text-to-Image Diffusion Models
abstract
The ability to fine-tune generative models for text-to-image generation tasks is crucial, particularly when facing the complexity involved in accurately interpreting and visualizing textual inputs. While LoRA is efficient for language model adaptation, it often falls short in text-to-image tasks due to the intricate demands of image generation, such as accommodating a broad spectrum of styles and nuances. To bridge this gap, we introduce StyleInject, a specialized fine-tuning approach tailored for text-to-image models. StyleInject comprises multiple parallel low-rank parameter matrices, maintaining the diversity of visual features. It dynamically adapts to varying styles by adjusting the variance of visual features based on the characteristics of the input signal. This approach significantly minimizes the impact on the original model’s text-image alignment capabilities while adeptly adapting to various styles in transfer learning. StyleInject proves particularly effective in learning from and enhancing a range of advanced, community-fine-tuned generative models. Our comprehensive experiments, including both small-sample and large-scale data fine-tuning as well as base model distillation, show that StyleInject surpasses traditional LoRA in both text-image semantic consistency and human preference evaluation, all while ensuring greater parameter efficiency.
Mohan Zhou, Yalong Bai, Qing Yang 0033, Tiejun Zhao
ACM Trans. Multim. Comput. Commun. Appl.4
2024 Enhancing Bilingual Lexicon Induction via Bi-directional Translation Pair Retrieving
abstract
Most Bilingual Lexicon Induction (BLI) methods retrieve word translation pairs by finding the closest target word for a given source word based on cross-lingual word embeddings (WEs). However, we find that solely retrieving translation from the source-to-target perspective leads to some false positive translation pairs, which significantly harm the precision of BLI. To address this problem, we propose a novel and effective method to improve translation pair retrieval in cross-lingual WEs. Specifically, we consider both source-side and target-side perspectives throughout the retrieval process to alleviate false positive word pairings that emanate from a single perspective. On a benchmark dataset of BLI, our proposed method achieves competitive performance compared to existing state-of-the-art (SOTA) methods. It demonstrates effectiveness and robustness across six experimental languages, including similar language pairs and distant language pairs, under both supervised and unsupervised settings.
Qiuyu Ding, Hailong Cao, Tiejun Zhao
AAAI3
2024 Spot the Error: Non-autoregressive Graphic Layout Generation with Wireframe Locator
abstract
Layout generation is a critical step in graphic design to achieve meaningful compositions of elements. Most previous works view it as a sequence generation problem by concatenating element attribute tokens (i.e., category, size, position). So far the autoregressive approach (AR) has achieved promising results, but is still limited in global context modeling and suffers from error propagation since it can only attend to the previously generated tokens. Recent non-autoregressive attempts (NAR) have shown competitive results, which provides a wider context range and the flexibility to refine with iterative decoding. However, current works only use simple heuristics to recognize erroneous tokens for refinement which is inaccurate. This paper first conducts an in-depth analysis to better understand the difference between the AR and NAR framework. Furthermore, based on our observation that pixel space is more sensitive in capturing spatial patterns of graphic layouts (e.g., overlap, alignment), we propose a learning-based locator to detect erroneous tokens which takes the wireframe image rendered from the generated layout sequence as input. We show that it serves as a complementary modality to the element sequence in object space and contributes greatly to the overall performance. Experiments on two public datasets show that our approach outperforms both AR and NAR baselines. Extensive studies further prove the effectiveness of different modules with interesting findings. Our code will be available at https://github.com/ffffatgoose/SpotError.
Jieru Lin, Danqing Huang, Tiejun Zhao, Dechen Zhan, Chin-Yew Lin
AAAI3
2024 Exploring the Neural Dynamics in Temporal Lobe Epilepsy: A Study using Transformer and Hidden Markov Models
abstract
Advancing the understanding of temporal lobe epilepsy (TLE) requires sophisticated analytical tools. In this study, we introduce a hybrid model, namely the HMM-Wavformer, aiming at identifying the phasic brain activity patterns during seizures on Stereo-electroencephalography (SEEG) records. The model is composed of a wavelet packet decomposition (WPD) based signal processing module, an embedding module for spatial feature extraction, and a Transformer module to weigh the time-series frequency importance. The model is trained with a downstream seizure detection task on the HUP-iEEG dataset, demonstrating an accuracy of 92.75%. Frequency analysis identifies the most sensitive bands in TLE seizure detection. The Hidden Markov Model (HMM) is applied for the time-series analysis, categorizing the seizures into three ictal phases. Complementary analyses using power spectra and brain networks pinpoint biomarkers for each phase. The analysis results indicate that, the HMM-Wavformer model is able to effectively depict the neural dynamics of TLE seizures, aligning with prior medical studies, and provide a more detailed description of the staged characteristics of these seizures.
Zhiguo Lin, Shihang Ding, Chunying Fang, Hongjian Bo, Cong Xu 0004, Shengkun Yu, Yifei Gu, Tiejun Zhao, Haifeng Li 0001
BIBM11
2024 EmoCRT: An Emotion-Cause Relation Enhanced Model for Causal Emotion Entailment
Zhilong Zhao, Bufan Xu, Muyun Yang, Kehai Chen, Tiejun Zhao
NLPCC (5)6
2024 Enhancing isomorphism between word embedding spaces for distant languages bilingual lexicon induction
Qiuyu Ding, Hailong Cao, Tiejun Zhao
Neural Comput. Appl.4
2024 An efficient confusing choices decoupling framework for multi-choice tasks over texts
Yingyao Wang, Junwei Bao 0001, Chaoqun Duan, Youzheng Wu, Xiaodong He 0001, Conghui Zhu, Tiejun Zhao
Neural Comput. Appl.7
2024 CLUE: Contrastive language-guided learning for referring video object segmentation
Wanjun Zhong, Jie Li 0055, Tiejun Zhao
Pattern Recognit. Lett.4
2024 Decomposed Meta-Learning for Few-Shot Sequence Labeling
abstract
Few-shot sequence labeling is a general problem formulation for many natural language understanding tasks in data-scarcity scenarios, which require models to generalize to new types via only a few labeled examples. Recent advances mostly adopt metric-based meta-learning and thus face the challenges of modeling the miscellaneousOtherprototype and the inability to generalize to classes with large domain gaps. To overcome these challenges, we propose a decomposed meta-learning framework for few-shot sequence labeling that breaks down the task into few-shot mention detection and few-shot type classification, and sequentially tackles them via meta-learning. Specifically, we employ model-agnostic meta-learning (MAML) to prompt the mention detection model to learn boundary knowledge shared across types. With the detected mention spans, we further leverage the MAML-enhanced span-level prototypical network for few-shot type classification. In this way, the decomposition framework bypasses the requirement of modeling the miscellaneousOtherprototype. Meanwhile, the adoption of the MAML algorithm enables us to explore the knowledge contained in support examples more efficiently, so that our model can quickly adapt to new types using only a few labeled examples. Under our framework, we explore a basic implementation that uses two separate models for the two subtasks. We further propose a joint model to reduce model size and inference time, making our framework more applicable for scenarios with limited resources. Extensive experiments on nine benchmark datasets, including named entity recognition, slot tagging, event detection, and part-of-speech tagging, show that the proposed approach achieves start-of-the-art performance across various few-shot sequence labeling tasks.
Qianhui Wu, Huiqiang Jiang, Jieru Lin, Börje Karlsson 0001, Tiejun Zhao, Chin-Yew Lin
IEEE ACM Trans. Audio Speech Lang. Process.6
2024 Operation-Augmented Numerical Reasoning for Question Answering
abstract
Question answering requiring numerical reasoning, which generally involves symbolic operations such as sorting, counting, and addition, is a challenging task. To address such a problem, existing mixture-of-experts (MoE)-based methods design several specific answer predictors to handle different types of questions and achieve promising performance. However, they ignore the modeling and exploitation of fine-grained reasoning-related operations to support numerical reasoning, encountering the inadequacy in reasoning capability and interpretability. To alleviate this issue, we propose OPERA, an operation-augmented numerical reasoning framework. Concretely, we systematically define a scalable operation set to model numerical reasoning. We first identify reasoning-related operations based on context and then softly execute them to imitate the answer reasoning procedure via an operation-aware cross-attention mechanism. Finally, we utilize the operation-augmented semantic representation of execution results to support answer prediction. We verify the effectiveness and generalization of OPERA in two scenarios with different knowledge sources and reasoning capabilities. Specifically, we conduct extensive experiments on two textual datasets, DROP and RACENum, and a table-text hybrid dataset TAT-QA. Experiment results show that OPERA outperforms previous strong methods on the DROP, RACENum, and TAT-QA datasets. Further, we statistically and visually analyze its interpretability.
Yongwei Zhou, Junwei Bao 0001, Youzheng Wu, Xiaodong He 0001, Tiejun Zhao
IEEE ACM Trans. Audio Speech Lang. Process.5
2023 Improving Translation Quality Estimation with Bias Mitigation
abstract
State-of-the-art translation Quality Estimation (QE) models are proven to be biased.More specifically, they over-rely on monolingual features while ignoring the bilingual semantic alignment.In this work, we propose a novel method to mitigate the bias of the QE model and improve estimation performance.Our method is based on the contrastive learning between clean and noisy sentence pairs.We first introduce noise to the target side of the parallel sentence pair, forming the negative samples.With the original parallel pairs as the positive sample, the QE model is contrastively trained to distinguish the positive samples from the negative ones.This objective is jointly trained with the regression-style quality estimation, so as to prevent the QE model from overfitting to monolingual features.Experiments on WMT QE evaluation datasets demonstrate that our method improves the estimation performance by a large margin while mitigating the bias 1 .
Hui Huang 0021, Shuangzhi Wu, Kehai Chen, Hui Di, Muyun Yang, Tiejun Zhao
ACL (1)6
2023 CoLaDa: A Collaborative Label Denoising Framework for Cross-lingual Named Entity Recognition
abstract
Tingting Ma, Qianhui Wu, Huiqiang Jiang, Börje Karlsson, Tiejun Zhao, Chin-Yew Lin. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Qianhui Wu, Huiqiang Jiang, Börje Karlsson 0001, Tiejun Zhao, Chin-Yew Lin
ACL (1)5
2023 Visual-Aware Text-to-Speech*
abstract
Dynamically synthesizing talking speech that actively responds to a listening head is critical during the face-to-face interaction. For example, the speaker could take advantage of the listener’s facial expression to adjust the tones, stressed syllables, or pauses. In this work, we present a new visual-aware text-to-speech (VA-TTS) task to synthesize speech conditioned on both textual inputs and sequential visual feedback (e.g., nod, smile) of the listener in face-to-face communication. Different from traditional text-to-speech, VA-TTS highlights the impact of visual modality. On this newly-minted task, we devise a baseline model to fuse phoneme linguistic information and listener visual signals for speech synthesis. Extensive experiments on multimodal conversation dataset ViCo-X verify our proposal for generating more natural audio with scenario-appropriate rhythm and prosody.
Mohan Zhou, Yalong Bai, Wei Zhang 0031, Ting Yao 0003, Tiejun Zhao, Tao Mei 0001
ICASSP5
2023 Leveraging Visual Prompts To Guide Language Modeling for Referring Video Object Segmentation
abstract
Referring Video Object Segmentation (R-VOS) aims to segment object masks in a target video given a language query describing the object. It is a challenging task that requires modeling the semantics of a natural language query and its correspondence to the target video. Previous works directly use visual-agnostic language features from uni-modal language models, and only interact with visual features in late decoding stages. We propose to encode visual-enriched language features by using visual prompts as guidance in the early encoding stage. The proposed visual prompt is constructed by modulating visual features of key frames with alignment scores to text inputs. The alignment score is computed with a pre-trained visual-language contrastive model. We concatenate visual prompts with text inputs to encode visual-enriched language features, which serve as queries for target object segmentation in a Transformer-based decoder. Our method outperforms the previous state-of-the-art method (+2.3) on Refer-Youtube-VOS benchmark.
Wanjun Zhong, Jie Li 0055, Tiejun Zhao
ICIP4
2023 Learning and Evaluating Human Preferences for Conversational Head Generation
abstract
A reliable and comprehensive evaluation metric that aligns with manual preference assessments is crucial for conversational head video synthesis methods development. Existing quantitative evaluations often fail to capture the full complexity of human preference, as they only consider limited evaluation dimensions. Qualitative evaluations and user studies offer a solution but are time-consuming and labor-intensive. This limitation hinders the advancement of conversational head generation algorithms and systems. In this paper, we propose a novel learning-based evaluation metric named Preference Score (PS) for fitting human preference according to the quantitative evaluations across different dimensions. PS can serve as a quantitative evaluation without the need for human annotation. Experimental results validate the superiority of Preference Score in aligning with human perception, and also demonstrate robustness and generalizability to unseen data, making it a valuable tool for advancing conversation head generation. We expect this metric could facilitate new advances in conversational head generation. Project page: https://github.com/dc3ea9f/PreferenceScore.
Mohan Zhou, Yalong Bai, Wei Zhang 0031, Ting Yao 0003, Tiejun Zhao, Tao Mei 0001
ACM Multimedia5
2023 An Explicit-Memory Few-Shot Joint Learning Model
Fanfan Du, Tiejun Zhao, Shafqat Ail
NLPCC (1)3
2023 Towards Making the Most of LLM for Translation Quality Estimation
Hui Huang 0021, Shuangzhi Wu, Xinnian Liang, Yanrui Shi, Peihao Wu, Muyun Yang, Tiejun Zhao
NLPCC (1)8
2023 Unsupervised Clustering for Negative Sampling to Optimize Open-Domain Question Answering Retrieval
Feiqing Zhuang, Conghui Zhu, Tiejun Zhao
NLPCC (2)3
2023 Document-Level Relation Extraction with Path Reasoning
abstract
Document-level relation extraction (DocRE) aims to extract relations among entities across multiple sentences within a document by using reasoning skills (i.e., pattern recognition, logical reasoning, coreference reasoning, etc.) related to the reasoning paths between two entities. However, most of the advanced DocRE models only attend to the feature representations of two entities to determine their relation, and do not consider one complete reasoning path from one entity to another entity, which may hinder the accuracy of relation extraction. To address this issue, this article proposes a novel method to capture this reasoning path from one entity to another entity, thereby better simulating reasoning skills to classify relation between two entities. Furthermore, we introduce an additional attention layer to summarize multiple reasoning paths for further enhancing the performance of the DocRE model. Experimental results on a large-scale document-level dataset show that the proposed approach achieved a significant performance improvement on a strong heterogeneous graph-based baseline.
Kehai Chen, Tiejun Zhao
ACM Trans. Asian Low Resour. Lang. Inf. Process.3
2023 Dual Word Embedding for Robust Unsupervised Bilingual Lexicon Induction
abstract
The word embedding models such as Word2vec and FastText simultaneously learn dual representations of input vectors and output vectors. In contrast, almost all existing unsupervised bilingual lexicon induction (UBLI) methods use only input vectors without utilizing output vectors. In this paper, we propose a novel approach to making full use of both input and output vectors for more robust and strong UBLI. We discover the Common Difference Property that one orthogonal transformation can connect not only the input vectors of two languages but also the output vectors. Therefore, we can learn just one transformation to induce two different dictionaries from the input and output vectors, respectively. Between these two quite different dictionaries, a more accurate lexicon with less noise can be induced by taking the intersection of them in UBLI procedure. Extensive experiments show that our method achieves much more robust and strong results than state-of-the-art methods in distant language pairs, while reserving comparable performances in similar language pairs.
Hailong Cao, Liguo Li, Conghui Zhu, Muyun Yang, Tiejun Zhao
IEEE ACM Trans. Audio Speech Lang. Process.5
2023 Real-time Image Enhancement with Attention Aggregation
abstract
Image enhancement has stimulated significant research works over the past years for its great application potential in video conferencing scenarios. Nevertheless, most existing image enhancement approaches are still struggling to find a good tradeoff that reduces the computational cost as much as possible while maintaining plausible result quality. Recently, curve-based mapping methods are proposed and have shown great potential for real-time and high-quality image enhancement of arbitrary resolutions. In this article, we take advantage of the curve-based mapping representation and focus on further improving the enhancement quality and robustness, while minimizing additional computational costs. Specifically, we (1) carefully re-formulate the curve function to improve learning stability, and (2) aggregate different semantic attention into the curve regression process, which can overcome the major problems of curve-based methods that generate moderate results with low contrast. The semantic attention is jointly learned with the supervision from class activation mapping of pre-trained feature extractors, thus reducing the manual annotation cost of semantic labels. Experiments have shown that our proposed method significantly improves curve-based methods both qualitatively and quantitatively, achieving visually plausible results compared with other deep neural network-based enhancement methods, and maintains a very low computational cost, i.e., taking 18.7 ms for a 360p image on a single P40 GPU. Extensive experiments demonstrate that our method is also capable of video enhancement tasks.
Jie Li 0055, Tiejun Zhao, Yadong Wang 0001
ACM Trans. Multim. Comput. Commun. Appl.3
2022 Cross-lingual Feature Extraction from Monolingual Corpora for Low-resource Unsupervised Bilingual Lexicon Induction
abstract
Despite their progress in high-resource language settings, unsupervised bilingual lexicon induction (UBLI) models often fail on corpora with low-resource distant language pairs due to insufficient initialization. In this work, we propose a cross-lingual feature extraction (CFE) method to learn the cross-lingual features from monolingual corpora for low-resource UBLI, enabling representations of words with the same meaning leveraged by the initialization step. By integrating cross-lingual representations with pre-trained word embeddings in a fully unsupervised initialization on UBLI, the proposed method outperforms existing state-of-the-art methods on low-resource language pairs (EN-VI, EN-TH, EN-ZH, EN-JA). The ablation study also proves that the learned cross-lingual features can enhance the representational ability and robustness of the existing embedding model.
Hailong Cao, Tiejun Zhao, Wei Peng 0011
COLING3
2022 Responsive Listening Head Generation: A Benchmark Dataset and Baseline
Mohan Zhou, Yalong Bai, Wei Zhang 0031, Ting Yao 0003, Tiejun Zhao, Tao Mei 0001
ECCV (38)5
2022 UniRPG: Unified Discrete Reasoning over Table and Text as Program Generation
abstract
Question answering requiring discrete reasoning, e.g., arithmetic computing, comparison, and counting, over knowledge is a challenging task.In this paper, we propose UniRPG, a semantic-parsing-based approach advanced in interpretability and scalability, to perform Unified discrete Reasoning over heterogeneous knowledge resources, i.e., table and text, as Program Generation.Concretely, UniRPG consists of a neural programmer and a symbolic program executor, where a program is the composition of a set of pre-defined general atomic and higher-order operations and arguments extracted from table and text.First, the programmer parses a question into a program by generating operations and copying arguments, and then, the executor derives answers from table and text based on the program.To alleviate the costly program annotation issue, we design a distant supervision approach for programmer learning, where pseudo programs are automatically constructed without annotated derivations.Extensive experiments on the TAT-QA dataset show that UniRPG achieves tremendous improvements and enhances interpretability and scalability compared with previous state-of-theart methods, even without derivation annotation.Moreover, it achieves promising performance on the textual dataset DROP without derivation annotation. 1
Yongwei Zhou, Junwei Bao 0001, Chaoqun Duan, Youzheng Wu, Xiaodong He 0001, Tiejun Zhao
EMNLP6
2022 On the Effectiveness of Sentence Encoding for Intent Detection Meta-Learning
abstract
Tingting Ma, Qianhui Wu, Zhiwei Yu, Tiejun Zhao, Chin-Yew Lin. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022.
Qianhui Wu, Zhiwei Yu 0001, Tiejun Zhao, Chin-Yew Lin
NAACL-HLT4
2022 Document-Level Relation Extraction with Sentences Importance Estimation and Focusing
abstract
Document-level relation extraction (DocRE) aims to determine the relation between two entities from a document of multiple sentences.Recent studies typically represent the entire document by sequence-or graph-based models to predict the relations of all entity pairs.However, we find that such a model is not robust and exhibits bizarre behaviors: it predicts correctly when an entire test document is fed as input, but errs when non-evidence sentences are removed.To this end, we propose a Sentence Importance Estimation and Focusing (SIEF) framework for DocRE, where we design a sentence importance score and a sentence focusing loss, encouraging DocRE models to focus on evidence sentences.Experimental results on two domains show that our SIEF not only improves overall performance, but also makes DocRE models more robust.Moreover, SIEF is a general framework, shown to be effective when combined with a variety of base DocRE models. 1
Kehai Chen, Lili Mou, Tiejun Zhao
NAACL-HLT4
2022 OPERA: Operation-Pivoted Discrete Reasoning over Text
abstract
Yongwei Zhou, Junwei Bao, Chaoqun Duan, Haipeng Sun, Jiahui Liang, Yifan Wang, Jing Zhao, Youzheng Wu, Xiaodong He, Tiejun Zhao. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022.
Yongwei Zhou, Junwei Bao 0001, Chaoqun Duan, Haipeng Sun, Jiahui Liang, Yifan Wang 0016, Youzheng Wu, Xiaodong He 0001, Tiejun Zhao
NAACL-HLT10
2022 Toward automatic support for leading court debates: a novel task proposal & effective approach of judicial question generation
Changzhen Ji, Xiaozhong Liu 0001, Adam Jatowt, Sourav S. Bhowmick, Changlong Sun, Conghui Zhu, Tiejun Zhao
Neural Comput. Appl.8
2022 Data-Driven Fuzzy Target-Side Representation for Intelligent Translation System
abstract
The encoder–decoder framework has been widely used in various practical artificial intelligence cyber-physical systems, including intelligent translation systems. The decoding process in such a framework usually demands the target-side representation, which is often learned by an autoaggressive decoder to simulate the target context information at the current time-step. However, the autoaggressive decoder only captures the previously generated partial target fragment and fails in simulating the global contextual information. In this article, we propose a new data-driven fuzzy context representation strategy to simulate the global target information. Specifically, we design two fuzzy methods to the global target contextual information, which are bag-of-words of target language generated via a softmax layer from the source-side representation and whole target sentence retrieved from the translation memory according to the source-side representation. Both methods facilitate the autoaggressive decoder to handle the global target context at the current time-step, thereby learning a more effective context vector for the generation of target translation. Extensive experiments on two machine translation tasks demonstrated that the proposed method achieved 3% improvement of BLEU score over a strong baseline.
Kehai Chen, Muyun Yang, Tiejun Zhao, Min Zhang 0005
IEEE Trans. Fuzzy Syst.3
2021 Document-Level Relation Extraction with Reconstruction
abstract
In document-level relation extraction (DocRE), graph structure is generally used to encode relation information in the input document to classify the relation category between each entity pair, and has greatly advanced the DocRE task over the past several years. However, the learned graph representation universally models relation information between all entity pairs regardless of whether there are relationships between these entity pairs. Thus, those entity pairs without relationships disperse the attention of the encoder-classifier DocRE for ones with relationships, which may further hind the improvement of DocRE. To alleviate this issue, we propose a novel encoder-classifier-reconstructor model for DocRE. The reconstructor manages to reconstruct the ground-truth path dependencies from the graph representation, to ensure that the proposed DocRE model pays more attention to encode entity pairs with relationships in the training. Furthermore, the reconstructor is regarded as a relationship indicator to assist relation classification in the inference, which can further improve the performance of DocRE model. Experimental results on a large-scale DocRE dataset show that the proposed model can significantly improve the accuracy of relation extraction on a strong heterogeneous graph-based baseline. The code is publicly available at https://github.com/xwjim/DocRE-Rec.
Kehai Chen, Tiejun Zhao
AAAI3
2021 A Neural Conversation Generation Model via Equivalent Shared Memory Investigation
abstract
Conversation generation as a challenging task in Natural Language Generation (NLG) has been increasingly attracting attention over the last years. A number of recent works adopted sequence-to-sequence structures along with external knowledge, which successfully enhanced the quality of generated conversations. Nevertheless, few works utilized the knowledge extracted from similar conversations for utterance generation. Taking conversations in customer service and court debate domains as examples, it is evident that essential entities/phrases, as well as their associated logic and inter-relationships, can be extracted and borrowed from similar conversation instances. Such information could provide useful signals for improving conversation generation. In this paper, we propose a novel reading and memory framework called Deep Reading Memory Network (DRMN) which is capable of remembering useful information of similar conversations for improving utterance generation. We apply our model to two large-scale conversation datasets of justice and e-commerce fields. Experiments prove that the proposed model outperforms the state-of-the-art approaches.
Changzhen Ji, Xiaozhong Liu 0001, Adam Jatowt, Changlong Sun, Conghui Zhu, Tiejun Zhao
CIKM7
2021 T-Bert: A Spam Review Detection Model Combining Group Intelligence and Personalized Sentiment Information
Tiejun Zhao, Jiyun Zhou
ICANN (5)3
2021 PTWA: Pre-Training with Word Attention for Chinese Named Entity Recognition
abstract
Recently, the character-based model that incorporates potential word information has proven effective for Chinese named entity recognition (NER). However, due to the independence of the pre-trained character model and the lexicon, it will cause the embedding space to be misaligned and cannot be combined well. Chinese pre-trained encoders usually process text as characters. It ignores the information carried by the larger granular information, so the encoder cannot easily adapt to certain character combinations. Because large-grained information is ignored and Chinese does not have clear character boundaries, this will lead to the loss of important semantic information, which is an important problem for Chinese. In this paper, we propose PTWA: pre-training with word attention for Chinese named entity recognition. PTWA uses multi-head word attention to form a word vector from multiple word vectors, and proposes a word length prediction task to better integrate the word vector into pre-training. With the powerful capabilities of the transformer, PTWA can explicitly make full use of potential word information without adding an external lexicon, and can coexist with pre-trained models that implicitly use word information (such as BERT-WWM, and ERNIE). Experiments conducted on four Chinese NER datasets show that the performance of PTWA is better than other word-word models and Chinese pre-training models.
Kaixin Ma, Tiejun Zhao, Jiyun Zhou
IJCNN3
2021 Self-Training for Unsupervised Neural Machine Translation in Unbalanced Training Data Scenarios
abstract
Haipeng Sun, Rui Wang, Kehai Chen, Masao Utiyama, Eiichiro Sumita, Tiejun Zhao. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021.
Haipeng Sun, Rui Wang 0015, Kehai Chen, Masao Utiyama, Eiichiro Sumita, Tiejun Zhao
NAACL-HLT6
2021 EviDR: Evidence-Emphasized Discrete Reasoning for Reasoning Machine Reading Comprehension
Yongwei Zhou, Junwei Bao 0001, Haipeng Sun, Jiahui Liang, Youzheng Wu, Xiaodong He 0001, Bowen Zhou 0001, Tiejun Zhao
NLPCC (1)8
2021 Unsupervised Neural Machine Translation for Similar and Distant Language Pairs: An Empirical Study
abstract
Unsupervised neural machine translation (UNMT) has achieved remarkable results for several language pairs, such as French–English and German–English. Most previous studies have focused on modeling UNMT systems; few studies have investigated the effect of UNMT on specific languages. In this article, we first empirically investigate UNMT for four diverse language pairs (French/German/Chinese/Japanese–English). We confirm that the performance of UNMT in translation tasks for similar language pairs (French/German–English) is dramatically better than for distant language pairs (Chinese/Japanese–English). We empirically show that the lack of shared words and different word orderings are the main reasons that lead UNMT to underperform in Chinese/Japanese–English. Based on these findings, we propose several methods, including artificial shared words and pre-ordering, to improve the performance of UNMT for distant language pairs. Moreover, we propose a simple general method to improve translation performance for all these four language pairs. The existing UNMT model can generate a translation of a reasonable quality after a few training epochs owing to a denoising mechanism and shared latent representations. However, learning shared latent representations restricts the performance of translation in both directions, particularly for distant language pairs, while denoising dramatically delays convergence by continuously modifying the training data. To avoid these problems, we propose a simple, yet effective and efficient, approach that (like UNMT) relies solely on monolingual corpora: pseudo-data-based unsupervised neural machine translation. Experimental results for these four language pairs show that our proposed methods significantly outperform UNMT baselines.
Haipeng Sun, Rui Wang 0015, Masao Utiyama, Benjamin Marie, Kehai Chen, Eiichiro Sumita, Tiejun Zhao
ACM Trans. Asian Low Resour. Lang. Inf. Process.7
2021 Modeling Future Cost for Neural Machine Translation
abstract
Existing neural machine translation (NMT) systems utilize sequence-to-sequence neural networks to generate target translation word by word, and then make the generated word at each time-step and the counterpart in the references as consistent as possible. However, the trained translation model tends to focus on ensuring the accuracy of the generated target word at the current time-step and does not consider its future cost which means the expected cost of generating the subsequent target translation (i.e., the next target word). To respond to this issue, in this article, we propose a simple and effective method to model the future cost of each target word for NMT systems. In detail, a future cost representation is learned based on the current generated target word and its contextual information to compute an additional loss to guide the training of the NMT model. Furthermore, the learned future cost representation at the current time-step is used to help the generation of the next target word in the decoding. Experimental results on three widely-used translation datasets, including the WMT14 English-to-German, WMT14 English-to-French, and WMT17 Chinese-to-English, show that the proposed approach achieves significant improvements over strong Transformer-based NMT baseline.
Chaoqun Duan, Kehai Chen, Rui Wang 0015, Masao Utiyama, Eiichiro Sumita, Conghui Zhu, Tiejun Zhao
IEEE ACM Trans. Audio Speech Lang. Process.7
2021 Detecting Source Contextual Barriers for Understanding Neural Machine Translation
abstract
In machine translation evaluation, the traditional wisdom measures model's generalization ability in an average sense, for example by using corpus BLEU. However, the statistics of corpus BLEU cannot provide comprehensive understanding and fine-grained analysis on model's generalization ability. As a remedy, this paper attempts to understand NMT at fine-grained level, by detecting contextual barriers within an unseen input sentence that \textit{cause} the degradation in model's translation quality. It proposes a principled definition of source contextual barriers as well as its modified version which is tractable in computation and operates at word-level. Based on the modified one, three simple methods are proposed for barrier detection by search-aware risk estimation through counterfactual generation. Extensive analyses are conducted on those detected contextual barrier words on both Zh$\Leftrightarrow$En NIST benchmarks. Potential usages motivated from barrier words are also discussed.
Lemao Liu, Conghui Zhu, Rui Wang 0015, Tiejun Zhao, Shuming Shi 0001
IEEE ACM Trans. Audio Speech Lang. Process.5
2020 Knowledge Distillation for Multilingual Unsupervised Neural Machine Translation
abstract
Unsupervised neural machine translation (UNMT) has recently achieved remarkable results for several language pairs. However, it can only translate between a single language pair and cannot produce translation results for multiple language pairs at the same time. That is, research on multilingual UNMT has been limited. In this paper, we empirically introduce a simple method to translate between thirteen languages using a single encoder and a single decoder, making use of multilingual data to improve UNMT for all language pairs. On the basis of the empirical findings, we propose two knowledge distillation methods to further enhance multilingual UNMT performance. Our experiments on a dataset with English translated to and from twelve other languages (including three language families and six language branches) show remarkable results, surpassing strong unsupervised individual baselines while achieving promising performance between non-English language pairs in zero-shot translation scenarios and alleviating poor performance in low-resource language pairs.
Haipeng Sun, Rui Wang 0015, Kehai Chen, Masao Utiyama, Eiichiro Sumita, Tiejun Zhao
ACL6
2020 Demographics Should Not Be the Reason of Toxicity: Mitigating Discrimination in Text Classifications with Instance Weighting
abstract
With the recent proliferation of the use of text classifications, researchers have found that there are certain unintended biases in text classification datasets.For example, texts containing some demographic identity-terms (e.g., "gay", "black") are more likely to be abusive in existing abusive language detection datasets.As a result, models trained with these datasets may consider sentences like "She makes me happy to be gay" as abusive simply because of the word "gay."In this paper, we formalize the unintended biases in text classification datasets as a kind of selection bias from the non-discrimination distribution to the discrimination distribution.Based on this formalization, we further propose a model-agnostic debiasing training framework by recovering the non-discrimination distribution using instance weighting, which does not require any extra resources or annotations apart from a pre-defined set of demographic identity-terms.Experiments demonstrate that our method can effectively alleviate the impacts of the unintended biases without significantly hurting models' generalization ability.
Conghui Zhu, Tiejun Zhao
ACL6
2020 Robust Unsupervised Neural Machine Translation with Adversarial Denoising Training
abstract
Unsupervised neural machine translation (UNMT) has recently attracted great interest in the machine translation community.The main advantage of the UNMT lies in its easy collection of required large training text sentences while with only a slightly worse performance than supervised neural machine translation which requires expensive annotated translation pairs on some translation tasks.In most studies, the UMNT is trained with clean data without considering its robustness to the noisy data.However, in real-world scenarios, there usually exists noise in the collected input sentences which degrades the performance of the translation system since the UNMT is sensitive to the small perturbations of the input sentences.In this paper, we first time explicitly take the noisy data into consideration to improve the robustness of the UNMT based systems.First of all, we clearly defined two types of noises in training sentences, i.e., word noise and word order noise, and empirically investigate its effect in the UNMT, then we propose adversarial training methods with denoising process in the UNMT.Experimental results on several language pairs show that our proposed methods substantially improved the robustness of the conventional UNMT systems in noisy scenarios.
Haipeng Sun, Rui Wang 0015, Kehai Chen, Xugang Lu, Masao Utiyama, Eiichiro Sumita, Tiejun Zhao
COLING7
2020 Learning to Decouple Relations: Few-Shot Relation Classification with Entity-Guided Attention and Confusion-Aware Training
abstract
This paper aims to enhance the few-shot relation classification especially for sentences that jointly describe multiple relations.Due to the fact that some relations usually keep high cooccurrence in the same context, previous few-shot relation classifiers struggle to distinguish them with few annotated instances.To alleviate the above relation confusion problem, we propose CTEG, a model equipped with two mechanisms to learn to decouple these easily-confused relations.On the one hand, an Entity-Guided Attention (EGA) mechanism, which leverages the syntactic relations and relative positions between each word and the specified entity pair, is introduced to guide the attention to filter out information causing confusion.On the other hand, a Confusion-Aware Training (CAT) method is proposed to explicitly learn to distinguish relations by playing a pushing-away game between classifying a sentence into a true relation and its confusing relation.Extensive experiments are conducted on the FewRel dataset, and the results show that our proposed model achieves comparable and even much better results to strong baselines in terms of accuracy.Furthermore, the ablation test and case study verify the effectiveness of our proposed EGA and CAT, especially in addressing the relation confusion problem.
Yingyao Wang, Junwei Bao 0001, Guangyi Liu 0005, Youzheng Wu, Xiaodong He 0001, Bowen Zhou 0001, Tiejun Zhao
COLING7
2020 Robust Machine Reading Comprehension by Learning Soft labels
abstract
Neural models have achieved great success on the task of machine reading comprehension (MRC), which are typically trained on hard labels.We argue that hard labels limit the model capability on generalization due to the label sparseness problem.In this paper, we propose a robust training method for MRC models to address this problem.Our method consists of three strategies, 1) label smoothing, 2) word overlapping, 3) distribution prediction.All of them help to train models on soft labels.We validate our approach on the representative architecture -ALBERT.Experimental results show that our method can greatly boost the baseline with 1% improvement in average, and achieve state-of-the-art performance on NewsQA and QUOREF.
Shuangzhi Wu, Muyun Yang, Kehai Chen, Tiejun Zhao
COLING5
2020 Look-Into-Object: Self-Supervised Structure Modeling for Object Recognition
abstract
Most object recognition approaches predominantly focus on learning discriminative visual patterns, while overlooking the holistic object structure. Though important, structure modeling usually requires significant manual annotations and therefore is labor-intensive. In this paper, we propose to ``look into object" (explicitly yet intrinsically model the object structure) through incorporating self-supervisions into the traditional framework. We show the recognition backbone can be substantially enhanced for more robust representation learning, without any cost of extra annotation and inference speed. Specifically, we first propose an object-extent learning module for localizing the object according to the visual patterns shared among the instances in the same category. We then design a spatial context learning module for modeling the internal structures of the object, through predicting the relative positions within the extent. These two modules can be easily plugged into any backbone networks during training and detached at inference time. Extensive experiments show that our look-into-object approach (LIO) achieves large performance gain on a number of benchmarks, including generic object recognition (ImageNet) and fine-grained object recognition tasks (CUB, Cars, Aircraft). We also show that this learning paradigm is highly generalizable to other tasks such as object detection and segmentation (MS COCO). Project page: https://github.com/JDAI-CV/LIO.
Mohan Zhou, Yalong Bai, Wei Zhang 0031, Tiejun Zhao, Tao Mei 0001
CVPR4
2020 Multimodal Matching Transformer for Live Commenting
abstract
Automatic live commenting aims to provide real-time comments on videos for viewers. It encourages users engagement on online video sites, and is also a good benchmark for video-to-text generation. Recent work on this task adopts encoder-decoder models to generate comments. However, these methods do not model the interaction between videos and comments explicitly, so they tend to generate popular comments that are often irrelevant to the videos. In this work, we aim to improve the relevance between live comments and videos by modeling the cross-modal interactions among different modalities. To this end, we propose a multimodal matching transformer to capture the relationships among comments, vision, and audio. The proposed model is based on the transformer framework and can iteratively learn the attention-aware representations for each modality. We evaluate the model on a publicly available live commenting dataset. Experiments show that the multimodal matching transformer model outperforms the state-of-the-art methods.
Chaoqun Duan, Lei Cui 0001, Shuming Ma, Furu Wei, Conghui Zhu, Tiejun Zhao
ECAI6
2020 Cross Copy Network for Dialogue Generation
abstract
07 2 8 22 0 2 02 !7" 2 #$% &' () (* + ' + ,+ -./0-12(.3 .456#$%&' (672' ($ 83 ' &$&$9% .,:6#$(4;2.,672' ($ ) (<' $($=(' >-% * ' + 5?3 ..@' (4+ .(6=A8BCDEFEGHIJGKIILMBINOPQREBMCSORTURTUMCVGWHTKEXTXTYEUBMBIN KEJZ[\HEU]ETUTMQ]JO BFTU^KIU^M_BKHGTIXTIMBINOPBIU^FJEOGDCFTIWHFEGMQ]JMBU `a3 8 1b8 ) (+ 2-:$* +/ -c5-$% * 6$,<' -(1-*/ % .@<'/ / -% -(+ d-3 <*c' + (-* *+ 2-$12' ->-@-(+ *./* -e,-(1-f + .f* -e,-(1-@.<-3*g -h 4h 6iA0jk$+ + -(+ ' .(6l.' (+ -%9-(-% $+ .%m-+c.% n*$(<0% $(* / .%@-% o + .-(2$(1-<'$3 .4,-1.(+ -(+4-(-% $+ ' .(hp2' 3 -1.(+-(+q,-(15$(<$11,% $15./ + -(* -% >-$*+ 2-@$r .%'(<' 1$+ .%*/ .%@.<-3+ % $' (' (46<' $3 .4,-3.4'1* 61$% % 5' (41% ' + ' 1$3' (/ .%@$+ ' .(/.%*.@-:$%+ ' 1,3 $%<.@$' (* 6$% -./ + -(' 4(.% - ' 1-$(<1.,%+<-&$+ -<' $3 .4,-$*-s$@:3 -* 61.@:$+ ' &3 -3 .4'1*1$(&-.&*-% >-< $1% .**<' / / -% -(+<' $3 .4,-'(* + $(1-* 6$(<+ 2' *' (f / .%@$+ ' .(1$(:%.>'<->' + $3->' <-(1-/ .%,++ -% f $(1-4-(-% $+ ' .(h)(+ 2' *:$:-% 6c-:% .:.* -$ (.>-3(-+ c.% n$% 12' + -1+ ,% -f7% .**7.:5m-+ f c.% n*g 006o+ .-s:3.%-+ 2-1,% % -(+<' $3 .41.(+ -s+$(<* ' @' 3 $%<' $3 .4,-'(* + $(1-* t3 .4'1$3 * + % ,1+ ,% -* ' @,3 + $(-.,* 3 5h us:-% ' @-(+ *c' + 2 + c.+ $* n* 61.,% +<-&$+ -$(<1,* + .@-%*-% >' 1-1.(+ -(+4-(-% $+ ' .(6:%.>-<+2$++ 2-:% .:.* -< $3 4.% ' + 2@ ' ** ,:-% ' .%+.-s'* + ' (4* + $+ -f ./f $% + 1.(+ -(+4-(-% $+ ' .(@.<-3 * h v w 8 12 b8 2 8*$(' @:.% + $(++ $* n' (m$+ ,% $3i$(4,$4-9-(-% $f + ' .(x6y 6<' $3 .4,-4-(-%$+ ' .(-@:.c-% *$c' <-* :-1+ % ,@./$::3 ' 1$+ ' .(*6* ,12$*12$+ &.+$(<1,* f + .@-%*-% >' 1-$,+ .@$+' .(h)(+ 2-:$* +/ -c5-$% * 6 &% -$n+ 2% .,42*'(<' $3 .4,-4-(-%$+ ' .(+-12(.3 .45/ .1,*-<.($* -% ' -*./* -e,-(1-f + .f* -e,-(1-@. -%-+$3 h 6z{|}o hj.% -% -1-(+ 3 56-sf + -% ($3n(.c3-<4-' *-@:3 .5-<+.-(2$(1-@.<-3:-% / .%@$(1-hp,-+$3 hg z{|~o i' ,-+$3 hg z{|o 1$($* * ' * +<' $3 .4,-4-(-%$+ ' .(&5,*' (4n(.c3-<4-+ % ' :3 -* hA' @' 3 $% 3 56i'-+$3 hg z{|~o $r :,% n$%-+$3 h g z{|o #,$(4-+$3 hg z{|o -<<5-+$3 hg z{|~o -s:3 .%-<.1,@-(+$*n(.c3-<4-<' * 1.>-% 5/ .%<'f $3 .4,-4-(-%$+ ' .(6$(<'$-+$3 hg z{|o --+$3 h g z{z{o 92$;>' (' (-r $<-+$3 hg z{|o l$% + 2$* $% $+ 2' $( -% 6,($/ / .%<$&3 -n(.c3 -<4-1.(*+ % ,1f + ' .($(<<-/-1+ ' >-<.@$' ($<$:+ $+ ' .(%-* + % ' 1++ 2-' % ,+ ' 3 ' ;$+ ' .(h7.:5f &$* -<4-(-% $+ ' .(@.<-3 *g ' (5$3 *-+$3 h 6 z{|9,-+$3 h 6z{|o2$>-&--(c' <-3 5$<.:+ -< ' (1.(+ -(+4-(-% $+ ' .(+$* n*$(<* 2.c&-+ + -%% -* ,3 + * 1.@:$% -<+ .*-e,-(1-f + .f* -e,-(1-@.<-3*c2-( / $1- .1$&,3 $% 5:% .&3-@h02$(n*+ .+ 2-' %($+ ,% -./3 ->-% $4' (4>.1$&,3 $% 5$(<1.(+-s+ <' * + % ' &,+ ' .(*/.%1.(+ -(+1.:56'+-($&3 -*+ .1.:5e,' % ' -*/ % .@+2-1,* + .@-%*c' 3 34-+* ' @' 3 $%% -* :.(* -*/ % .@+2-* + $/ / h ) +@.+ ' >$+ -*,*+ .&,' 3 <$@.<-3+2$+1$((.+.(3 5 1.:5+ 2-1.(+ -(+c' + 2' (+ 2-,::-%1.(+-s+./+ 2-+ $% 4-+<' $3 .4,-'(* + $(1-6&,+$3 * .3-$% (+ 2-* ' @' 3 $% :$+ + -% (*$1% .**<' / / -% -(+* ' @' 3 $%1$* -*./+ 2-+ $% 4-+ ' (*+ $(1-hA,12-s+ -% ($31.:51$(&-1%' + ' 1$3' (*.@-* 1-($% ' .*h 8** 2.c(' (' 4,% -h |6c-:% .:.* -+ c.<' / / -% f -(+n' (<*./1.:5@-12$(' * @*' (+ 2' ** + ,<571 8 bb2451.(+-s+ f <-:-(<-(+' (/ .%@$+ ' .(c'+ 2' ( + 2-+ $% 4-+<' $3 .4,-'(* + $(1-6$(<21 28 b245 3 .4'1f <-:-(<-(+1.(+-(+$1% .**<' / / -% -(+t A' @' 3 $% 7$* -* tx 0y h02' */ % $@-c.%n' *3 $&-3 -<$*7% .** f 7.:5m-+ c.% n*x 006y h8*-s-@:3 $%<' $3 .4,-<-:'1+ -<6r ,<4-*@$5% -:-$+g 2.% ' ;.(+ $31.:' - $3' <$+ -+ 2-:% .:.* -<@.<-3 6c--@f :3 .5+c.<' / / -% -(+<' $3 .4,-<$+$* -+ */ % .@+c..% f + 2.4.($3<.@$'(*f $(<
Changzhen Ji, Xiaozhong Liu 0001, Changlong Sun, Conghui Zhu, Tiejun Zhao
EMNLP (1)7
2020 Incorporating Phrase-Level Agreement into Neural Machine Translation
Xing Wang 0007, Min Zhang 0005, Tiejun Zhao
NLPCC (1)4
2020 Towards More Diverse Input Representation for Neural Machine Translation
abstract
Source input information plays a very important role in the Transformer-based translation system. In practice, word embedding and positional embedding of each word are added as the input representation. Then self-attention networks are used to encode the global dependencies in the input representation to generate a source representation. However, this processing on the source representation only adopts a single source feature and excludes richer and more diverse features such as recurrence features, local features, and syntactic features, which results in tedious representation and thereby hinders the further translation performance improvement. In this paper, we introduce a simple and efficient method to encode more diverse source features into the input representation simultaneously, and thereby learning an effective source representation by self-attention networks. In particular, the proposed grouped strategy is only applied to the input representation layer, to keep the diversity of translation information and the efficiency of the self-attention networks at the same time. Experimental results show that our approach improves the translation performance over the state-of-the-art baselines of Transformer in regard to WMT14 English-to-German and NIST Chinese-to-English machine translation tasks.
Kehai Chen, Rui Wang 0015, Masao Utiyama, Eiichiro Sumita, Tiejun Zhao, Muyun Yang, Hai Zhao 0001
IEEE ACM Trans. Audio Speech Lang. Process.5
2020 Unsupervised Neural Machine Translation With Cross-Lingual Language Representation Agreement
abstract
Unsupervised cross-lingual language representation initialization methods such as unsupervised bilingual word embedding (UBWE) pre-training and cross-lingual masked language model (CMLM) pre-training, together with mechanisms such as denoising and back-translation, have advanced unsupervised neural machine translation (UNMT), which has achieved impressive results on several language pairs, particularly French-English and German-English. Typically, UBWE focuses on initializing the word embedding layer in the encoder and decoder of UNMT, whereas the CMLM focuses on initializing the entire encoder and decoder of UNMT. However, UBWE/CMLM training and UNMT training are independent, which makes it difficult to assess how the quality of UBWE/CMLM affects the performance of UNMT during UNMT training. In this paper, we first empirically explore relationships between UNMT and UBWE/CMLM. The empirical results demonstrate that the performance of UBWE and CMLM has a significant influence on the performance of UNMT. Motivated by this, we propose a novel UNMT structure with cross-lingual language representation agreement to capture the interaction between UBWE/CMLM and UNMT during UNMT training. Experimental results on several language pairs demonstrate that the proposed UNMT models improve significantly over the corresponding state-of-the-art UNMT baselines.
Haipeng Sun, Rui Wang 0015, Kehai Chen, Masao Utiyama, Eiichiro Sumita, Tiejun Zhao
IEEE ACM Trans. Audio Speech Lang. Process.6
2020 A Novel Sentence-Level Agreement Architecture for Neural Machine Translation
abstract
In neural machine translation (NMT), there is a natural correspondence between source and target sentences. The traditional NMT method does not explicitly model the translation agreement on sentence-level. In this article, we propose a comprehensive and novel sentence-level agreement architecture to alleviate this problem. It directly minimizes the difference between the representations of the source-side and target-side sentence on sentence-level. First, we compare a variety of sentence representation strategies and propose a “Gated Sum” sentence representation to achieve better sentence semantic information. Then, rather than a single-layer sentence-level agreement architecture, we further propose a multi-layer sentence agreement architecture to make the source and target semantic spaces closer layer by layer. The proposed agreement module can be integrated into NMT as an additional training objective function, and can also be used to enhance the representation of the source-side sentences. Experiments on the NIST Chinese-to-English and the WMT English-to-German translation tasks show that the proposed agreement architecture achieves significant improvements over state-of-the-art baselines, demonstrating the effectiveness and necessity of exploiting sentence-level agreement for NMT.
Rui Wang 0015, Kehai Chen, Xing Wang 0007, Tiejun Zhao, Min Zhang 0005
IEEE ACM Trans. Audio Speech Lang. Process.5
2020 A Joint Sentence Scoring and Selection Framework for Neural Extractive Document Summarization
abstract
Extractive document summarization methods aim to extract important sentences to form a summary. Previous works perform this task by first scoring all sentences in the document then selecting most informative ones; while we propose to jointly learn the two steps with a novel end-to-end neural network framework. Specifically, the sentences in the input document are represented as real-valued vectors through a neural document encoder. Then the method builds the output summary by extracting important sentences one by one. Different from previous works, the proposed joint sentence scoring and selection framework directly predicts the relative sentence importance score according to both sentence content and previously selected sentences. We evaluate the proposed framework with two realizations: a hierarchical recurrent neural network based model; and a pre-training based model that uses BERT as the document encoder. Experiments on two datasets show that the proposed joint framework outperforms the state-of-the-art extractive summarization models which treat sentence scoring and selection as two subtasks.
Qingyu Zhou, Nan Yang 0002, Furu Wei, Shaohan Huang, Ming Zhou 0001, Tiejun Zhao
IEEE ACM Trans. Audio Speech Lang. Process.6
2020 A Hierarchical Clustering Approach to Fuzzy Semantic Representation of Rare Words in Neural Machine Translation
abstract
Rare words are usually replaced with a singletoken in the current encoder-decoder style of neural machine translation, challenging the translation modeling by an obscured context. In this article, we propose to build a fuzzy semantic representation (FSR) method for rare words through a hierarchical clustering method to group rare words together, and integrate it into the encoder-decoder framework. This hierarchical structure can compensate for the semantic information in both source and target sides, and providing fuzzy context information to capture the semantic of rare words. The introduced FSR can also alleviate the data sparseness, which is the bottleneck in dealing with rare words in neural machine translation. In particular, our method is easily extended to the transformer-based neural machine translation model and learns the FSRs of all in-vocabulary words to enhance the sentence representations in addition to rare words. Our experiments on Chinese-to-English translation tasks confirm a significant improvement in the translation quality brought by the proposed method.
Muyun Yang, Shujie Liu 0001, Kehai Chen, Enbo Zhao, Tiejun Zhao
IEEE Trans. Fuzzy Syst.6
2019 Unsupervised Bilingual Word Embedding Agreement for Unsupervised Neural Machine Translation
abstract
Unsupervised bilingual word embedding (UBWE), together with other technologies such as back-translation and denoising, has helped unsupervised neural machine translation (UNMT) achieve remarkable results in several language pairs.In previous methods, UBWE is first trained using nonparallel monolingual corpora and then this pre-trained UBWE is used to initialize the word embedding in the encoder and decoder of UNMT.That is, the training of UBWE and UNMT are separate.In this paper, we first empirically investigate the relationship between UBWE and UNMT.The empirical findings show that the performance of UNMT is significantly affected by the performance of UBWE.Thus, we propose two methods that train UNMT with UBWE agreement.Empirical results on several language pairs show that the proposed methods significantly outperform conventional UNMT.
Haipeng Sun, Rui Wang 0015, Kehai Chen, Masao Utiyama, Eiichiro Sumita, Tiejun Zhao
ACL (1)6
2019 Sentence-Level Agreement for Neural Machine Translation
abstract
The training objective of neural machine translation (NMT) is to minimize the loss between the words in the translated sentences and those in the references. In NMT, there is a natural correspondence between the source sentence and the target sentence. However, this relationship has only been represented using the entire neural network and the training objective is computed in word-level. In this paper, we propose a sentence-level agreement module to directly minimize the difference between the representation of source and target sentence. The proposed agreement module can be integrated into NMT as an additional training objective function and can also be used to enhance the representation of the source sentences. Empirical results on the NIST Chinese-to-English and WMT English-to-German tasks show the proposed agreement module can significantly improve the NMT performance.
Rui Wang 0015, Kehai Chen, Masao Utiyama, Eiichiro Sumita, Min Zhang 0005, Tiejun Zhao
ACL (1)7
2019 Selection Bias Explorations and Debias Methods for Natural Language Sentence Matching Datasets
abstract
Natural Language Sentence Matching (NLSM) has gained substantial attention from both academics and the industry, and rich public datasets contribute a lot to this process.However, biased datasets can also hurt the generalization performance of trained models and give untrustworthy evaluation results.For many NLSM datasets, the providers select some pairs of sentences into the datasets, and this sampling procedure can easily bring unintended pattern, i.e., selection bias.One example is the QuoraQP dataset, where some content-independent naïve features are unreasonably predictive.Such features are the reflection of the selection bias and termed as the "leakage features."In this paper, we investigate the problem of selection bias on six NLSM datasets and find that four out of them are significantly biased.We further propose a training and evaluation framework to alleviate the bias.Experimental results on QuoraQP suggest that the proposed framework can improve the generalization ability of trained models, and give more trustworthy evaluation results for real-world adoptions.
Jian Liang 0002, Shiyu Chang, Mo Yu, Conghui Zhu, Tiejun Zhao
ACL (1)8
2019 Understanding Data Augmentation in Neural Machine Translation: Two Perspectives towards Generalization
abstract
Guanlin Li, Lemao Liu, Guoping Huang, Conghui Zhu, Tiejun Zhao. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Lemao Liu, Guoping Huang, Conghui Zhu, Tiejun Zhao
EMNLP/IJCNLP (1)5
2019 Transfer Learning Methods for Spoken Language Understanding
abstract
In this paper, we present a series of methods to improve the performance of spoken language understanding in the 1st Chinese Audio-Textual Spoken Language Understanding Challenge (CATSLU 2019) which is aimed to improve the robustness for automatic speech recognition (ASR) errors and to solve the problem of not enough labeled data in new domains. We combine word information and char information to improve the performance of the semantic parser. We also use some transfer learning methods like correlation alignments to improve the robustness of the spoken language understanding system. Then we merge the rule method and the neural network method to raise system output performance. In video and weather domains with few training data, we use both the transfer learning model trained on multi-domain data and the rule-based approach. Our approaches achieve F1 scores of 86.83%, 92.84%, 94.16%, and 93.04% on the test sets of map, music, video and weather domains.
Chengda Tang, Xiaotian Zhao, Xuancai Li, Zhuolin Jin, Dequan Zheng, Tiejun Zhao
ICMI7
2019 CATSLU: The 1st Chinese Audio-Textual Spoken Language Understanding Challenge
abstract
Spoken language understanding (SLU) is a key component of conversational dialogue systems, which converts user utterances into semantic representations. The previous works almost focus on parsing semantic from textual inputs (top hypothesis of speech recognition and even manual transcripts) while losing information hidden in the audio. We herein describe the 1st Chinese Audio-Textual Spoken Language Understanding Challenge (CATSLU) which introduces a new dataset with audio-textual information, multiple domains and domain knowledge. We introduce two scenarios of audio-textual SLU in which participants are encouraged to utilize data of other domains or not. In this paper, we will describe the challenge and results.
Su Zhu, Tiejun Zhao, Chengqing Zong, Kai Yu 0004
ICMI3
2019 Learning Domain Invariant Word Representations for Parsing Domain Adaptation
Xiuming Qiao, Yue Zhang 0004, Tiejun Zhao
NLPCC (1)3
2019 Target Oriented Data Generation for Quality Estimation of Machine Translation
Huanqin Wu, Muyun Yang, Junguo Zhu, Tiejun Zhao
NLPCC (1)5
2019 Frequent Pattern-based Graph Exploration
abstract
Visual graph exploration can help the users have an intuitive impression about a dataset for the first time. However, it is better to provide some guidance so that the users can quickly locate the "interesting" or "informative" area in the graph. Then they can start the next stage of sense-making. We propose a visual exploration system which provides some frequent sub-graphs extracted from the whole graph, and the relationship between them. Thus, users can raise their study questions more quickly when they first confront a graph dataset. We evaluate our exploration system with two datasets. Specifically, we demonstrate how to raise a question and find the answer with our system, which validates the effectiveness of this project.
Kai Yan 0003, Tiejun Zhao
VINCI3
2019 A Bilingual Adversarial Autoencoder for Unsupervised Bilingual Lexicon Induction
abstract
Unsupervised bilingual lexicon induction aims to generate bilingual lexicons without any cross-lingual signals. Successfully solving this problem would benefit many downstream tasks, such as unsupervised machine translation and transfer learning. In this work, we propose an unsupervised framework, named bilingual adversarial autoencoder, which automatically generates bilingual lexicon for a pair of languages from their monolingual word embeddings. In contrast to existing frameworks which learn a direct cross-lingual mapping of word embeddings from the source language to the target language, we train two autoencoders jointly to transform the source and the target monolingual word embeddings into a shared embedding space, where a word and its translation are close to each other. In this way, we capture the cross-lingual features of word embeddings from different languages and use them to induce bilingual lexicons. By conducting extensive experiments across eight language pairs, we demonstrate that the proposed method significantly outperforms the existing adversarial methods and even achieves best-published results across most language pairs.
Xuefeng Bai 0001, Hailong Cao, Kehai Chen, Tiejun Zhao
IEEE ACM Trans. Audio Speech Lang. Process.4
2019 Text Generation From Tables
abstract
This paper proposes a neural generative model, namely Table2Seq, to generate a natural language sentence based on a table. Specifically, the model maps a table to continuous vectors and then generates a natural language sentence by leveraging the semantics of a table. Since rare words, e.g., entities and values, usually appear in a table, we develop a flexible copying mechanism that selectively replicates contents from the table to the output sequence. We conduct extensive experiments to demonstrate the effectiveness of our Table2Seq model and the utility of the designed copying mechanism. On the WIKIBIO and SIMPLEQUESTIONS datasets, the Table2Seq model improves the state-of-the-art results from 34.70 to 40.26 and from 33.32 to 39.12 in terms of BLEU-4 scores, respectively. Moreover, we construct an open-domain dataset WIKITABLETEXT that includes 13 318 descriptive sentences for 4962 tables. Our Table2Seq model achieves a BLEU-4 score of 38.23 on WIKITABLETEXT outperforming template-based and language model based approaches. Furthermore, through experiments on 1 M table-query pairs from a search engine, our Table2Seq model considering the structured part of a table, i.e., table attributes and table cells, as additional information outperforms a sequence-to-sequence model considering only the sequential part of a table, i.e., table caption.
Junwei Bao 0001, Duyu Tang, Nan Duan 0001, Ming Zhou 0001, Tiejun Zhao
IEEE ACM Trans. Audio Speech Lang. Process.6
2019 Neural Machine Translation With Sentence-Level Topic Context
abstract
Traditional neural machine translation (NMT) methods use the word-level context to predict target language translation while neglecting the sentence-level context, which has been shown to be beneficial for translation prediction in statistical machine translation. This paper represents the sentence-level context as latent topic representations by using a convolution neural network, and designs a topic attention to integrate source sentence-level topic context information into both attention-based and Transformer-based NMT. In particular, our method can improve the performance of NMT by modeling source topics and translations jointly. Experiments on the large-scale LDC Chinese-to-English translation tasks and WMT'14 English-to-German translation tasks show that the proposed approach can achieve significant improvements compared with baseline systems.
Kehai Chen, Rui Wang 0015, Masao Utiyama, Eiichiro Sumita, Tiejun Zhao
IEEE ACM Trans. Audio Speech Lang. Process.5
2018 Table-to-Text: Describing Table Region With Natural Language
abstract
In this paper, we present a generative model to generate a natural language sentence describing a table region, e.g., a row. The model maps a row from a table to a continuous vector and then generates a natural language sentence by leveraging the semantics of a table. To deal with rare words appearing in a table, we develop a flexible copying mechanism that selectively replicates contents from the table in the output sequence. Extensive experiments demonstrate the accuracy of the model and the power of the copying mechanism. On two synthetic datasets, WIKIBIO and SIMPLEQUESTIONS, our model improves the current state-of-the-art BLEU-4 score from 34.70 to 40.26 and from 33.32 to 39.12, respectively. Furthermore, we introduce an open-domain dataset WIKITABLETEXT including 13,318 explanatory sentences for 4,962 tables. Our model achieves a BLEU-4 score of 38.23, which outperforms template based and language model based approaches.
Junwei Bao 0001, Duyu Tang, Nan Duan 0001, Yuanhua Lv, Ming Zhou 0001, Tiejun Zhao
AAAI7
2018 Syntax-Directed Attention for Neural Machine Translation
abstract
Attention mechanism, including global attention and local attention, plays a key role in neural machine translation (NMT). Global attention attends to all source words for word prediction. In comparison, local attention selectively looks at fixed-window source words. However, alignment weights for the current target word often decrease to the left and right by linear distance centering on the aligned source position and neglect syntax distance constraints. In this paper, we extend the local attention with syntax-distance constraint, which focuses on syntactically related source words with the predicted target word to learning a more effective context vector for predicting translation. Moreover, we further propose a double context NMT architecture, which consists of a global context vector and a syntax-directed context vector from the global attention, to provide more translation performance for NMT from source representation. The experiments on the large-scale Chinese-to-English and English-to-German translation tasks show that the proposed approach achieves a substantial and significant improvement over the baseline system.
Kehai Chen, Rui Wang 0015, Masao Utiyama, Eiichiro Sumita, Tiejun Zhao
AAAI5
2018 Forest-Based Neural Machine Translation
abstract
Tree-based neural machine translation (NMT) approaches, although achieved impressive performance, suffer from a major drawback: they only use the 1best parse tree to direct the translation, which potentially introduces translation mistakes due to parsing errors.For statistical machine translation (SMT), forestbased methods have been proven to be effective for solving this problem, while for NMT this kind of approach has not been attempted.This paper proposes a forest-based NMT method that translates a linearized packed forest under a simple sequence-to-sequence framework (i.e., a forest-to-string NMT model).The BLEU score of the proposed method is higher than that of the string-to-string NMT, treebased NMT, and forest-based SMT systems.
Chunpeng Ma, Akihiro Tamura, Masao Utiyama, Tiejun Zhao, Eiichiro Sumita
ACL (1)4
2018 Neural Document Summarization by Jointly Learning to Score and Select Sentences
abstract
Sentence scoring and sentence selection are two main steps in extractive document summarization systems.However, previous works treat them as two separated subtasks.In this paper, we present a novel end-to-end neural network framework for extractive document summarization by jointly learning to score and select sentences.It first reads the document sentences with a hierarchical encoder to obtain the representation of sentences.Then it builds the output summary by extracting sentences one by one.Different from previous methods, our approach integrates the selection strategy into the scoring model, which directly predicts the relative importance given previously selected sentences.Experiments on the CNN/Daily Mail dataset show that the proposed framework significantly outperforms the state-of-the-art extractive summarization models.
Qingyu Zhou, Nan Yang 0002, Furu Wei, Shaohan Huang, Ming Zhou 0001, Tiejun Zhao
ACL (1)6
2018 Deep Attention Neural Tensor Network for Visual Question Answering
Yalong Bai, Jianlong Fu, Tiejun Zhao, Tao Mei 0001
ECCV (12)3
2018 Point Set Registration for Unsupervised Bilingual Lexicon Induction
abstract
Inspired by the observation that word embeddings exhibit isomorphic structure across languages, we propose a novel method to induce a bilingual lexicon from only two sets of word embeddings, which are trained on monolingual source and target data respectively. This is achieved by formulating the task as point set registration which is a more general problem. We show that a transformation from the source to the target embedding space can be learned automatically without any form of cross-lingual supervision. By properly adapting a traditional point set registration model to make it be suitable for processing word embeddings, we achieved state-of-the-art performance on the unsupervised bilingual lexicon induction task. The point set registration problem has been well-studied and can be solved by many elegant models, we thus opened up a new opportunity to capture the universal lexical semantic structure across languages.
Hailong Cao, Tiejun Zhao
IJCAI2
2018 Attention-Fused Deep Matching Network for Natural Language Inference
abstract
Natural language inference aims to predict whether a premise sentence can infer another hypothesis sentence. Recent progress on this task only relies on a shallow interaction between sentence pairs, which is insufficient for modeling complex relations. In this paper, we present an attention-fused deep matching network (AF-DMN) for natural language inference. Unlike existing models, AF-DMN takes two sentences as input and iteratively learns the attention-aware representations for each side by multi-level interactions. Moreover, we add a self-attention mechanism to fully exploit local context information within each sentence. Experiment results show that AF-DMN achieves state-of-the-art performance and outperforms strong baselines on Stanford natural language inference (SNLI), multi-genre natural language inference (MultiNLI), and Quora duplicate questions datasets.
Chaoqun Duan, Lei Cui 0001, Xinchi Chen, Furu Wei, Conghui Zhu, Tiejun Zhao
IJCAI6
2018 Improving Vector Space Word Representations Via Kernel Canonical Correlation Analysis
abstract
Cross-lingual word embeddings are representations for vocabularies of two or more languages in one common continuous vector space and are widely used in various natural language processing tasks. A state-of-the-art way to generate cross-lingual word embeddings is to learn a linear mapping, with an assumption that the vector representations of similar words in different languages are related by a linear relationship. However, this assumption does not always hold true, especially for substantially different languages. We therefore propose to use kernel canonical correlation analysis to capture a non-linear relationship between word embeddings of two languages. By extensively evaluating the learned word embeddings on three tasks (word similarity, cross-lingual dictionary induction, and cross-lingual document classification) across five language pairs, we demonstrate that our proposed approach achieves essentially better performances than previous linear methods on all of the three tasks, especially for language pairs with substantial typological difference.
Xuefeng Bai 0001, Hailong Cao, Tiejun Zhao
ACM Trans. Asian Low Resour. Lang. Inf. Process.3
2018 Question Generation With Doubly Adversarial Nets
abstract
We study the problem of question generation on a specific domain, where there are no labeled data. To address this problem, we propose a novel neural question generation approach called DoubAN, or doubly adversarial nets, which fully utilizes labeled data from other domains (source domains) and unlabeled data from the target domain. Learning a DoubAN involves two adversarial procedures between a question generator and two adversaries. One adversary is a domain-classification discriminator (DC-Dis), which is designed to help the generator learn domain-general representations of the input text. The other is a question-answering discriminator (QA-Dis), which provides more training data with estimated reward scores for generated text-question pairs. We conduct experiments on the SQuAD dataset as target-domain unlabeled data and the NewsQA dataset as source-domain labeled data. Experiment results show that our DoubAN achieves better results than baselines. Compared to model variants, which adopt only DC-Dis or QA-Dis, we find that the DC-Dis and QA-Dis indirectly interact with each other and jointly improve the quality of generated questions on the target domain. Moreover, extensive analysis and discussion prove the reasonableness and effectiveness of our proposed approach.
Junwei Bao 0001, Yeyun Gong, Nan Duan 0001, Ming Zhou 0001, Tiejun Zhao
IEEE ACM Trans. Audio Speech Lang. Process.5
2018 A Neural Approach to Source Dependence Based Context Model for Statistical Machine Translation
abstract
In statistical machine translation, translation prediction considers not only the aligned source word itself but also its source contextual information. Learning context representation is a promising method for improving translation results, particularly through neural networks. Most of the existing methods process context words sequentially and neglect source long-distance dependencies. In this paper, we propose a novel neural approach to source dependence-based context representation for translation prediction. The proposed model is capable of not only encoding source long-distance dependencies but also capturing functional similarities to better predict translations (i.e., word form translations and ambiguous word translations). To verify our method, the proposed mode is incorporated into phrase-based and hierarchical phrase-based translation models, respectively. Experiments on large-scale Chinese-to-English and English-to-German translation tasks show that the proposed approach achieves significant improvement over the baseline systems and outperforms several existing context-enhanced methods.
Kehai Chen, Tiejun Zhao, Muyun Yang, Lemao Liu, Akihiro Tamura, Rui Wang 0015, Masao Utiyama, Eiichiro Sumita
IEEE ACM Trans. Audio Speech Lang. Process.2
2018 Automatic Data Augmentation from Massive Web Images for Deep Visual Recognition
abstract
Large-scale image datasets and deep convolutional neural networks (DCNNs) are the two primary driving forces for the rapid progress in generic object recognition tasks in recent years. While lots of network architectures have been continuously designed to pursue lower error rates, few efforts are devoted to enlarging existing datasets due to high labeling costs and unfair comparison issues. In this article, we aim to achieve lower error rates by augmenting existing datasets in an automatic manner. Our method leverages both the web and DCNN, where the web provides massive images with rich contextual information, and DCNN replaces humans to automatically label images under the guidance of web contextual information. Experiments show that our method can automatically scale up existing datasets significantly from billions of web pages with high accuracy. The performance on object recognition tasks and transfer learning tasks have been significantly improved by using the automatically augmented datasets, which demonstrates that more supervisory information has been automatically gathered from the web. Both the dataset and models trained on the dataset have been made publicly available.
Yalong Bai, Kuiyuan Yang, Tao Mei 0001, Wei-Ying Ma, Tiejun Zhao
ACM Trans. Multim. Comput. Commun. Appl.5
2017 Translation Prediction with Source Dependency-Based Context Representation
abstract
Learning context representations is very promising to improve translation results, particularly through neural networks. Previous efforts process the context words sequentially and neglect their internal syntactic structure. In this paper, we propose a novel neural network based on bi-convolutional architecture to represent the source dependency-based context for translation prediction. The proposed model is able to not only encode the long-distance dependencies but also capture the functional similarities for better translation prediction (i.e., ambiguous words translation and word forms translation). Examined by a large-scale Chinese-English translation task, the proposed approach achieves a significant improvement (of up to +1.9 BLEU points) over the baseline system, and meanwhile outperforms a number of context-enhanced comparison system.
Kehai Chen, Tiejun Zhao, Muyun Yang, Lemao Liu
AAAI2
2017 Deterministic Attention for Sequence-to-Sequence Constituent Parsing
abstract
The sequence-to-sequence model is proven to be extremely successful in constituent parsing. It relies on one key technique, the probabilistic attention mechanism, to automatically select the context for prediction. Despite its successes, the probabilistic attention model does not always select the most important context. For example, the headword and boundary words of a subtree have been shown to be critical when predicting the constituent label of the subtree, but this contextual information becomes increasingly difficult to learn as the length of the sequence increases. In this study, we proposed a deterministic attention mechanism that deterministically selects the important context and is not affected by the sequence length. We implemented two different instances of this framework. When combined with a novel bottom-up linearization method, our parser demonstrated better performance than that achieved by the sequence-to-sequence parser with probabilistic attention mechanism.
Chunpeng Ma, Lemao Liu, Akihiro Tamura, Tiejun Zhao, Eiichiro Sumita
AAAI4
2017 Neural Machine Translation with Source Dependency Representation
abstract
Source dependency information has been successfully introduced into statistical machine translation.However, there are only a few preliminary attempts for Neural Machine Translation (NMT), such as concatenating representations of source word and its dependency label together.In this paper, we propose a novel attentional NMT with source dependency representation to improve translation performance of NMT, especially on long sentences.Empirical results on NIST Chinese-to-English translation task show that our method achieves 1.6 BLEU improvements on average over a strong NMT system.
Kehai Chen, Rui Wang 0015, Masao Utiyama, Lemao Liu, Akihiro Tamura, Eiichiro Sumita, Tiejun Zhao
EMNLP7
2017 Context-Aware Smoothing for Neural Machine Translation
abstract
In Neural Machine Translation (NMT), each word is represented as a low-dimension, real-value vector for encoding its syntax and semantic information. This means that even if the word is in a different sentence context, it is represented as the fixed vector to learn source representation. Moreover, a large number of Out-Of-Vocabulary (OOV) words, which have different syntax and semantic information, are represented as the same vector representation of “unk”. To alleviate this problem, we propose a novel context-aware smoothing method to dynamically learn a sentence-specific vector for each word (including OOV words) depending on its local context words in a sentence. The learned context-aware representation is integrated into the NMT to improve the translation performance. Empirical results on NIST Chinese-to-English translation task show that the proposed approach achieves 1.78 BLEU improvements on average over a strong attentional NMT, and outperforms some existing systems.
Kehai Chen, Rui Wang 0015, Masao Utiyama, Eiichiro Sumita, Tiejun Zhao
IJCNLP(1)5
2017 An Information Retrieval-Based Approach to Table-Based Question Answering
Junwei Bao 0001, Nan Duan 0001, Ming Zhou 0001, Tiejun Zhao
NLPCC4
2017 An Empirical Study on Incorporating Prior Knowledge into BLSTM Framework in Answer Selection
Muyun Yang, Tiejun Zhao, Dequan Zheng, Sheng Li 0003
NLPCC3
2016 Constraint-Based Question Answering with Knowledge Graph
abstract
WebQuestions and SimpleQuestions are two benchmark data-sets commonly used in recent knowledge-based question answering (KBQA) work. Most questions in them are ‘simple’ questions which can be answered based on a single relation in the knowledge base. Such data-sets lack the capability of evaluating KBQA systems on complicated questions. Motivated by this issue, we release a new data-set, namely ComplexQuestions, aiming to measure the quality of KBQA systems on ‘multi-constraint’ questions which require multiple knowledge base relations to get the answer. Beside, we propose a novel systematic KBQA approach to solve multi-constraint questions. Compared to state-of-the-art methods, our approach not only obtains comparable results on the two existing benchmark data-sets, but also achieves significant improvements on the ComplexQuestions.
Junwei Bao 0001, Nan Duan 0001, Ming Zhou 0001, Tiejun Zhao
COLING5
2016 A Distribution-based Model to Learn Bilingual Word Embeddings
abstract
We introduce a distribution based model to learn bilingual word embeddings from monolingual data. It is simple, effective and does not require any parallel data or any seed lexicon. We take advantage of the fact that word embeddings are usually in form of dense real-valued low-dimensional vector and therefore the distribution of them can be accurately estimated. A novel cross-lingual learning objective is proposed which directly matches the distributions of word embeddings in one language with that in the other language. During the joint learning process, we dynamically estimate the distributions of word embeddings in two languages respectively and minimize the dissimilarity between them through standard back propagation algorithm. Our learned bilingual word embeddings allow to group each word and its translations together in the shared vector space. We demonstrate the utility of the learned embeddings on the task of finding word-to-word translations from monolingual corpora. Our model achieved encouraging performance on data in both related languages and substantially different languages.
Hailong Cao, Tiejun Zhao
COLING2
2016 Improving Dependency Parsing on Clinical Text with Syntactic Clusters from Web Text
Xiuming Qiao, Hailong Cao, Tiejun Zhao, Kehai Chen
ICONIP (1)3
2016 Building A Case-based Semantic English-Chinese Parallel Treebank
Huaxing Shi, Tiejun Zhao, Keh-Yih Su
LREC2
2016 Self-adaptive statistical process control for anomaly detection in time series
Dequan Zheng, Fenghuan Li, Tiejun Zhao
Expert Syst. Appl.3
2016 Corrigendum to "Self-adaptive statistical process control for anomaly detection in time series" [Expert Systems With Applications 57 (2016) 324-336]
Dequan Zheng, Fenghuan Li, Tiejun Zhao
Expert Syst. Appl.3
2016 Improving Unsupervised Dependency Parsing with Knowledge from Query Logs
abstract
Unsupervised dependency parsing becomes more and more popular in recent years because it does not need expensive annotations, such as treebanks, which are required for supervised and semi-supervised dependency parsing. However, its accuracy is still far below that of supervised dependency parsers, partly due to the fact that their parsing model is insufficient to capture linguistic phenomena underlying texts. The performance for unsupervised dependency parsing can be improved by mining knowledge from the texts and by incorporating it into the model. In this article, syntactic knowledge is acquired from query logs to help estimate better probabilities in dependency models with valence. The proposed method is language independent and obtains an improvement of 4.1% unlabeled accuracy on the Penn Chinese Treebank by utilizing additional dependency relations from the Sogou query logs and Baidu query logs. Morever, experiments show that the proposed model achieves improvements of 8.07% on CoNLL 2007 English using the AOL query logs. We believe query logs are useful sources of syntactic knowledge for many natural language processing (NLP) tasks.
Xiuming Qiao, Hailong Cao, Tiejun Zhao
ACM Trans. Asian Low Resour. Lang. Inf. Process.3
2015 Efficient Disfluency Detection with Transition-based Parsing
abstract
Shuangzhi Wu, Dongdong Zhang, Ming Zhou, Tiejun Zhao. Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2015.
Shuangzhi Wu, Dongdong Zhang 0001, Ming Zhou 0001, Tiejun Zhao
ACL (1)4
2015 Bilingual Lexicon Extraction with Forced Correlation from Comparable Corpora
Chunyue Zhang, Tiejun Zhao
ICONIP (2)2
2015 Automatic Image Dataset Construction from Click-through Logs Using Deep Neural Network
abstract
Labelled image datasets are the backbone for high-level image understanding tasks with wide application scenarios, and continuously drive and evaluate the progress of feature designing and supervised learning models. Recently, the million scale labelled image dataset further contributes to the rebirth of deep convolutional neural network and bypass manual designing handcraft features. However, the construction process of image dataset is mainly manual-based and quite labor intensive, which often take years' efforts to construct a million scale dataset with high quality. In this paper, we propose a deep learning based method to construct large scale image dataset in an automatic way. Specifically, word representation and image representation are learned in a deep neural network from large amount of click-through logs, and further used to define word-word similarity and image-word similarity. These two similarities are used to automatize the two labor intensive steps in manual-based image dataset construction: query formation and noisy image removal. With a new proposed cross convolutional filter regularizer, we can construct a million scale image dataset in one week. Finally, two image datasets are constructed to verify the effectiveness of the method. In addition to scale, the automatically constructed dataset has comparable accuracy, diversity and cross-dataset generalization with manually labelled image datasets.
Yalong Bai, Kuiyuan Yang, Wei Yu 0004, Chang Xu 0008, Wei-Ying Ma, Tiejun Zhao
ACM Multimedia6
2015 A Maximum Entropy Approach to Discourse Coherence Modeling
abstract
This paper introduces a maximum entropy method to Discourse Coherence Modeling (DCM). Different from the state-of-art supervised entity-grid model and unsupervised cohesion-driven model, the model we proposed only takes as input lexicon features, which increases the training speed and decoding speed significantly. We conduct an evaluation on two publicly available benchmark data sets via sentence ordering tasks, and the results confirm the effectiveness of our maximum entropy based approach in DCM.
Muyun Yang, Shujie Liu 0001, Sheng Li 0003, Tiejun Zhao
NLPCC5
2015 Bilingual Lexicon Extraction with Temporal Distributed Word Representation from Comparable Corpora
abstract
Distributed word representation has been found to be highly effective to extract a bilingual lexicon from comparable corpora by a simple linear transformation. However, polysemous words often vary their meanings at different time points in the corresponding corpora. A single word representation which is learned from the whole corpora can’t express the temporal change of the word meaning very well. This paper proposes a simple solution which exploits the temporal distributed word representation for polysemous words . The experimental results confirm that the proposed solution can offer better performance on the English-to-Chinese bilingual lexicon extraction task.
Chunyue Zhang, Tiejun Zhao
NLPCC2
2014 Knowledge-Based Question Answering as Machine Translation
abstract
A typical knowledge-based question answering (KB-QA) system faces two challenges: one is to transform natural language questions into their meaning representations (MRs); the other is to retrieve answers from knowledge bases (KBs) using generated MRs.Unlike previous methods which treat them in a cascaded manner, we present a translation-based approach to solve these two tasks in one unified framework.We translate questions to answers based on CYK parsing.Answers as translations of the span covered by each CYK cell are obtained by a question translation method, which first generates formal triple queries as MRs for the span based on question patterns and relation expressions, and then retrieves answers from a given KB based on triple queries generated.A linear model is defined over derivations, and minimum error rate training is used to tune feature weights based on a set of question-answer pairs.Compared to a KB-QA system using a state-of-the-art semantic parser, our method achieves better results.
Junwei Bao 0001, Nan Duan 0001, Ming Zhou 0001, Tiejun Zhao
ACL (1)4
2014 A Lexicalized Reordering Model for Hierarchical Phrase-based Translation
Hailong Cao, Dongdong Zhang 0001, Mu Li 0001, Ming Zhou 0001, Tiejun Zhao
COLING5
2014 Soft Dependency Matching for Hierarchical Phrase-based Machine Translation
Hailong Cao, Dongdong Zhang 0001, Ming Zhou 0001, Tiejun Zhao
COLING4
2014 Improving Pivot-Based Statistical Machine Translation by Pivoting the Co-occurrence Count of Phrase Pairs
abstract
To overcome the scarceness of bilingual corpora for some language pairs in machine translation, pivot-based SMT uses pivot language as a "bridge" to generate source-target translation from sourcepivot and pivot-target translation.One of the key issues is to estimate the probabilities for the generated phrase pairs.In this paper, we present a novel approach to calculate the translation probability by pivoting the co-occurrence count of source-pivot and pivot-target phrase pairs.Experimental results on Europarl data and web data show that our method leads to significant improvements over the baseline systems.
Zhongjun He, Hua Wu 0003, Conghui Zhu, Haifeng Wang 0001, Tiejun Zhao
EMNLP6
2014 Bag-of-Words Based Deep Neural Network for Image Retrieval
abstract
This work targets image retrieval task hold by MSR-Bing Grand Challenge. Image retrieval is considered as a challenge task because of the gap between low-level image representation and high-level textual query representation. Recently further developed deep neural network sheds light on narrowing the gap by learning high-level image representation from raw pixels. In this paper, we proposed a bag-of-words based deep neural network for image retrieval task, which learns high-level image representation and maps images into bag-of-words space. The DNN model is trained on the large scale clickthrough data, and the relevance between query and image is measured by the cosine similarity of query's bag-of-words representation and image's bag-of-words representation predicted by DNN, the visual similarity of images is computed by high-level image representation extracted via the DNN model too. Finally, PageRank algorithm is used to further improve the ranking list by considering visual similarity of images for each query. The experimental results achieved state-of-the-art performance and verified the effectiveness of our proposed method.
Yalong Bai, Wei Yu 0004, Tianjun Xiao, Chang Xu 0008, Kuiyuan Yang, Wei-Ying Ma, Tiejun Zhao
ACM Multimedia7
2014 Discriminative Training for Log-Linear Based SMT: Global or Local Methods
abstract
In statistical machine translation, the standard methods such as MERT tune a single weight with regard to a given development data. However, these methods suffer from two problems due to the diversity and uneven distribution of source sentences. First, their performance is highly dependent on the choice of a development set, which may lead to an unstable performance for testing. Second, the sentence level translation quality is not assured since tuning is performed on the document level rather than on sentence level. In contrast with the standard global training in which a single weight is learned, we propose novel local training methods to address these two problems. We perform training and testing in one step by locally learning the sentence-wise weight for each input sentence. Since the time of each tuning step is unnegligible and learning sentence-wise weights for the entire test set means many passes of tuning, it is a great challenge for the efficiency of local training. We propose an efficient two-phase method to put the local training into practice by employing the ultraconservative update. On NIST Chinese-to-English translation tasks with both medium and large scales of training data, our local training methods significantly outperform standard methods with the maximal improvements up to 2.0 BLEU points, meanwhile their efficiency is comparable to that of the standard methods.
Lemao Liu, Tiejun Zhao, Taro Watanabe, Hailong Cao, Conghui Zhu
ACM Trans. Asian Lang. Inf. Process.2
2013 Additive Neural Networks for Statistical Machine Translation
Lemao Liu, Taro Watanabe, Eiichiro Sumita, Tiejun Zhao
ACL (1)4
2013 Hierarchical Phrase Table Combination for Machine Translation
Conghui Zhu, Taro Watanabe, Eiichiro Sumita, Tiejun Zhao
ACL (1)4
2013 Exploring Deep Belief Nets to Detect and Categorize Chinese Entities
Yu Chen 0022, Dequan Zheng, Tiejun Zhao
ADMA (1)3
2013 Chinese Terminology Extraction Using EM-Based Transfer Learning Method
Yanxia Qin, Dequan Zheng, Tiejun Zhao, Min Zhang 0005
CICLing (1)3
2013 Improving Pivot-Based Statistical Machine Translation Using Random Walk
abstract
This paper proposes a novel approach that utilizes a machine learning method to improve pivot-based statistical machine translation (SMT).For language pairs with few bilingual data, a possible solution in pivot-based SMT using another language as a "bridge" to generate source-target translation.However, one of the weaknesses is that some useful sourcetarget translations cannot be generated if the corresponding source phrase and target phrase connect to different pivot phrases.To alleviate the problem, we utilize Markov random walks to connect possible translation phrases between source and target language.Experimental results on European Parliament data, spoken language data and web data show that our method leads to significant improvements on all the tasks over the baseline system.
Zhongjun He, Hua Wu 0003, Haifeng Wang 0001, Conghui Zhu, Tiejun Zhao
EMNLP6
2013 Fusion of Word and Letter Based Metrics for Automatic MT Evaluation
Muyun Yang, Junguo Zhu, Sheng Li 0003, Tiejun Zhao
IJCAI4
2013 Learning Domain Differences Automatically for Dependency Parsing Adaptation
Mo Yu, Tiejun Zhao, Yalong Bai
IJCAI2
2013 Tuning SMT with a Large Number of Features via Online Feature Grouping
Lemao Liu, Tiejun Zhao, Taro Watanabe, Eiichiro Sumita
IJCNLP2
2013 Repairing Incorrect Translation with Examples
Junguo Zhu, Muyun Yang, Sheng Li 0003, Tiejun Zhao
IJCNLP4
2013 Compound Embedding Features for Semi-supervised Learning
Mo Yu, Tiejun Zhao, Daxiang Dong, Dianhai Yu
HLT-NAACL2
2013 Phrase Table Combination Deficiency Analyses in Pivot-Based SMT
Yiming Cui 0001, Conghui Zhu, Tiejun Zhao, Dequan Zheng
NLDB4
2012 Phrasal Syntactic Category Sequence Model for Phrase-Based MT
Hailong Cao, Eiichiro Sumita, Tiejun Zhao, Sheng Li 0003
CICLing (2)3
2012 Research on Text Categorization Based on a Weakly-Supervised Transfer Learning Method
Dequan Zheng, Chenghe Zhang, Geli Fei, Tiejun Zhao
CICLing (2)4
2012 Locally Training the Log-Linear Model for SMT
Lemao Liu, Hailong Cao, Taro Watanabe, Tiejun Zhao, Mo Yu, Conghui Zhu
EMNLP-CoNLL4
2011 Target-dependent Twitter Sentiment Classification
Long Jiang, Mo Yu, Ming Zhou 0001, Tiejun Zhao
ACL5
2011 Hypergraph Training and Decoding of System Combination in SMT
Tiejun Zhao, Sheng Li 0003
MTSummit2
2011 A Unified and Discriminative Soft Syntactic Constraint Model for Hierarchical Phrase-based Translation
Lemao Liu, Tiejun Zhao, Hailong Cao
MTSummit2
2011 Improvement of Machine Translation Evaluation by Simple Linguistically Motivated Features
Muyun Yang, Shu-Qi Sun, Junguo Zhu, Sheng Li 0003, Tiejun Zhao
J. Comput. Sci. Technol.5
2010 A Deterministic Method to Predict Phrase Boundaries of a Syntactic Tree
Zhaoxia Dong, Tiejun Zhao
ICIC (2)2
2010 Chinese Named Entity Recognition with a Sequence Labeling Approach: Based on Characters, or Based on Words?
Zhangxun Liu, Conghui Zhu, Tiejun Zhao
ICIC (2)3
2010 Predicting query potential for personalization, classification or regression?
abstract
The goal of predicting query potential for personalization is to determine which queries can benefit from personalization. In this paper, we investigate which kind of strategy is better for this task: classification or regression. We quantify the potential benefits of personalizing search results using two implicit click-based measures: Click entropy and Potential@N. Meanwhile, queries are characterized by query features and history features. Then we build C-SVM classification model and epsilon-SVM regression model respectively according to these two measures. The experimental results show that the classification model is a better choice for predicting query potential for personalization.
Muyun Yang, Sheng Li 0003, Tiejun Zhao, Haoliang Qi
SIGIR4
2010 A delimiter-based general approach for Chinese term extraction
abstract
Abstract This article addresses a two‐step approach for term extraction. In the first step on term candidate extraction, a new delimiter‐based approach is proposed to identify features of the delimiters of term candidates rather than those of the term candidates themselves. This delimiter‐based method is much more stable and domain independent than the previous approaches. In the second step on term verification, an algorithm using link analysis is applied to calculate the relevance between term candidates and the sentences from which the terms are extracted. All information is obtained from the working domain corpus without the need for prior domain knowledge. The approach is not targeted at any specific domain and there is no need for extensive training when applying it to new domains. In other words, the method is not domain dependent and it is especially useful for resource‐limited domains. Evaluations of Chinese text in two different domains show quite significant improvements over existing techniques and also verify its efficiency and its relatively domain‐independent nature. The proposed method is also very effective for extracting new terms so that it can serve as an efficient tool for updating domain knowledge, especially for expanding lexicons.
Qin Lu 0001, Tiejun Zhao
J. Assoc. Inf. Sci. Technol.3
2008 Chinese Term Extraction Using Minimal Resources
Qin Lu 0001, Tiejun Zhao
COLING3
2008 Diagnostic Evaluation of Machine Translation Systems Using Automatically Constructed Linguistic Check-Points
Ming Zhou 0001, Shujie Liu 0001, Mu Li 0001, Dongdong Zhang 0001, Tiejun Zhao
COLING6
2008 Chinese Term Extraction Based on Delimiters
Qin Lu 0001, Tiejun Zhao
LREC3
2007 A Unified Tagging Approach to Text Normalization
Conghui Zhu, Jie Tang 0001, Hang Li 0001, Hwee Tou Ng, Tiejun Zhao
ACL5
2007 Incorporating Passage Feature Within Language Model Framework for Information Retrieval
Ke Dang, Tiejun Zhao, Haoliang Qi, Dequan Zheng
CICLing2
2007 Research on a Novel Word Co-occurrence Model and Its Application
Dequan Zheng, Tiejun Zhao, Sheng Li 0003, Hao Yu 0005
KSEM2
2007 Recent advances on NLP research in Harbin Institute of Technology
Tiejun Zhao, Yi Guan, Ting Liu 0001, Qiang Wang 0001
Frontiers Comput. Sci. China1
2006 Improving English Subcategorization Acquisition with Diathesis Alternations as Heuristic Information
Xiwu Han, Tiejun Zhao, Xingshang Fu
ACL2
2006 Linguistic Knowledge Representation and Automatic Acquisition Based on a Combination of Ontology with Statistical Method
Dequan Zheng, Tiejun Zhao, Sheng Li 0003, Hao Yu 0005
KSEM2
2006 Chinese Information Processing and Its Prospects
Sheng Li 0003, Tiejun Zhao
J. Comput. Sci. Technol.2
2004 Subcategorization Acquisition and Evaluation for Chinese Verbs
Xiwu Han, Tiejun Zhao, Haoliang Qi, Hao Yu 0005
COLING2
2004 FML-Based SCF Predefinition Learning for Chinese Verbs
Xiwu Han, Tiejun Zhao, Muyun Yang
IJCNLP2
2002 Learning Chinese Bracketing Knowledge Based on a Bilingual Language Model
Yajuan Lü, Sheng Li 0003, Tiejun Zhao, Muyun Yang
COLING3
2002 An Automatic Evaluation Method for Localization Oriented Lexicalised EBMT System
Ming Zhou 0001, Tiejun Zhao, Hao Yu 0005, Sheng Li 0003
COLING3
2002 Research of Machine Learning Method for Specific Information Recognition on the Internet
abstract
With the available resources on the Internet becoming plentiful, a large amount of harmful information is permeating in and has been seriously affecting people's normal work and living. Therefore, harmful data streams must be recognized and filtered out effectively. After analyzing some harmful contents in Internet information streams, we present a new method, which recognizes specific information by machine learning (ML). We extracted key information from a number of corpuses through the ML method to obtain the part of speech (POS) transfer-form for key information by learning from corpuses, which is based on the same pronunciation matching of key information. Furthermore, the testing value of key information will be obtained in a real corpus to examine the likelihood between matching rules from information streams and those learnt from corpuses through the average value of POS transfer probability of key information. Therefore, the testing value for the whole real data stream will be obtained The experiment proved that the method was efficient for recognizing certain Internet harmful information.
Dequan Zheng, Tiejun Zhao, Hao Yu 0005, Sheng Li 0003
ICMI3
2000 Bilingual Dictionary Based Sentence Alignment for Chinese English Bitext
Tiejun Zhao, Muyun Yang, Li Ping Qian 0001, Gaolin Fang
ICMI1