VLDB 2026 Research / reviewers in the wild / expert
Shujian Huang
dblp:57/8451
· DBLP profile ↗
118ranked-venue papers
3as first author
61since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 107 · 3 first-author · 53 since 2021Graphics, computer vision, multimedia, augmented reality and games · 24 · 1 first-author · 11 since 2021Databases, data management, data science and information retrieval · 6 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | How Does Alignment Enhance LLMs' Multilingual Capabilities? A Language Neurons PerspectiveabstractMultilingual Alignment is an effective and representative paradigm to enhance LLMs' multilingual capabilities, which transfers the capabilities from the high-resource languages to the low-resource languages. Meanwhile, some research on language-specific neurons provides a new perspective to analyze and understand LLMs' mechanisms. However, we find that there are many neurons that are shared by multiple but not all languages and cannot be correctly classified. In this work, we propose a ternary classification methodology that categorizes neurons into three types, including language-specific neurons, language-related neurons, and general neurons. And we propose a corresponding identification algorithm to distinguish these different types of neurons. Furthermore, based on the distributional characteristics of different types of neurons, we divide the LLMs' internal process for multilingual inference into four parts: (1) multilingual understanding, (2) shared semantic space reasoning, (3) multilingual output space transformation, and (4) vocabulary space outputting. Additionally, we systematically analyze the models before and after alignment with a focus on different types of neurons. We also analyze the phenomenon of ''Spontaneous Multilingual Alignment''. Overall, our work conducts a comprehensive investigation based on different types of neurons, providing empirical results and valuable insights to better understand multilingual alignment and multilingual capabilities of LLMs. Shimao Zhang, Zhejian Lai, Xiang Liu 0023, Shuaijie She, Yeyun Gong, Shujian Huang, Jiajun Chen 0001 |
AAAI | 7 |
| 2026 | Reasoning While Asking: Transforming Reasoning Large Language Models from Passive Solvers to Proactive InquirersabstractXin Chen, Feng Jiang, Yiqian Zhang, Hardy Chen, Shuo Yan, Wenya Xie, Min Yang, Shujian Huang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Xin Chen 0032, Feng Jiang 0007, Hardy Chen, Wenya Xie, Min Yang 0007, Shujian Huang |
ACL (1) | 8 |
| 2026 | Improving Long-Context Translation via Self-Supervised Dual LearningabstractLarge language models (LLMs) with long context windows offer the potential to translate entire documents in a single pass, yet they frequently suffer from catastrophic information distortion, undermining the strict faithfulness required for translation.This challenge is compounded by the scarcity of documentlevel parallel data, which makes both supervised fine-tuning and reliable evaluation prohibitively expensive.We propose LongDu, a self-supervised post-training framework that improves long-document translation reliability via round-trip consistency.Given monolingual documents, LongDu samples multiple candidate translations, back-translates each candidate, and optimizes the model to prefer translations that best reconstruct the source.To make this signal robust for long-form generation, we design a reward that filters trivial failure modes (e.g., copying and local language drift) before applying a reconstruction and fluency score, enabling stable reinforcement learning without human annotations.We additionally introduce Long-CIRT, an automatic evaluation protocol that quantifies information distortion by measuring how much a LLM's performance degrades after a translation cycle.Across multiple base models, LongDu substantially improves information retention and translation quality, with gains that generalize beyond the training length range and to unseen target languages. Shanbo Cheng, Shuaijie She, Jiajun Chen 0001, Shujian Huang |
ACL (1) | 6 |
| 2026 | A Data-Efficient Path to Multilingual LLMs: Language Expansion via Post-training PARAMΔ Integration into Upcycled MoEabstractHao Zhou, Tianhao Li, Zhijun Wang, Shuaijie She, Linjuan Wu, Hao-Ran Wei, Baosong Yang, Jiajun Chen, Shujian Huang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Hao Zhou 0012, Shuaijie She, Linjuan Wu, Baosong Yang, Jiajun Chen 0001, Shujian Huang |
ACL (1) | 9 |
| 2026 | Why not transform chat large language models to non-English?
Xiang Geng, Ming Zhu 0010, Jiahuan Li, Zhejian Lai, Shuaijie She, Yinglu Li, Yuang Li, Chang Su 0001, Xinglin Lyu, Min Zhang 0042, Jiajun Chen 0001, Hao Yang 0006, Shujian Huang |
Frontiers Comput. Sci. | 17 |
| 2026 | Hacking reference-free image captioning metrics
Zheng Ma 0012, Changxin Wang, Yawen Ouyang, Fei Zhao 0012, Shujian Huang, Jiajun Chen 0001 |
Frontiers Comput. Sci. | 6 |
| 2026 | Improved paraphrase generation via controllable latent diffusion
Ziyuan Zhuang, Xiang Geng, Shujian Huang, Jiajun Chen 0001 |
Frontiers Comput. Sci. | 4 |
| 2025 | MoE-LPR: Multilingual Extension of Large Language Models Through Mixture-of-Experts with Language Priors RoutingabstractLarge Language Models (LLMs) are often English-centric due to the disproportionate distribution of languages in their pre-training data. Enhancing non-English language capabilities through post-pretraining often results in catastrophic forgetting of high-resource languages. Previous methods either achieve good expansion with severe forgetting or slight forgetting with poor expansion, indicating the challenge of balancing language expansion while preventing forgetting. In this paper, we propose a method called MoE-LPR (Mixture-of-Experts with Language Priors Routing) to alleviate this problem. MoE-LPR employs a two-stage training approach to enhance the multilingual capability. First, the model is post-pretrained into a Mixture-of-Experts(MoE) architecture by upcycling, where all the original parameters are frozen and new experts are added. In this stage, we focus improving the ability on expanded languages, without using any original language data. Then, the model reviews the knowledge of the original languages with replay data amounting to less than 1% of post-pretraining, where we incorporate language priors routing to better recover the abilities of the original languages. Evaluations on multiple benchmarks show that MoE-LPR outperforms other post-pretraining methods. Freezing original parameters preserves original language knowledge while adding new experts preserves the learning ability. Reviewing with LPR enables effective utilization of multilingual knowledge within the parameters. Additionally, the MoE architecture maintains the same inference overhead while increasing total model parameters. Extensive experiments demonstrate MoE-LPR’s effectiveness in improving expanded languages and preserving original language proficiency with superior scalability. Hao Zhou 0012, Shujian Huang, Xue Han 0018, Junlan Feng, Chao Deng 0002, Weihua Luo, Jiajun Chen 0001 |
AAAI | 3 |
| 2025 | Alleviating Distribution Shift in Synthetic Data for Machine Translation Quality EstimationabstractQuality Estimation (QE) models evaluate the quality of machine translations without reference translations, serving as the reward models for the translation task.Due to the data scarcity, synthetic data generation has emerged as a promising solution.However, synthetic QE data often suffers from distribution shift, which can manifest as discrepancies between pseudo and real translations, or in pseudo labels that do not align with human preferences.To tackle this issue, we introduce DCSQE, a novel framework for alleviating distribution shift in synthetic QE data.To reduce the difference between pseudo and real translations, we employ the constrained beam search algorithm and enhance translation diversity through the use of distinct generation models.DCSQE uses references—i.e., translation supervision signals—to guide both the generation and annotation processes, enhancing the quality of token-level labels.DCSQE further identifies the shortest phrase covering consecutive error tokens, mimicking human annotation behavior, to assign the final phrase-level labels.Specially, we underscore that the translation model can not annotate translations of itself accurately.Extensive experiments demonstrate that DCSQE outperforms SOTA baselines like CometKiwi in both supervised and unsupervised settings.Further analysis offers insights into synthetic data generation that could benefit reward models for other tasks.The code is available at https://github.com/NJUNLP/njuqe. Xiang Geng, Zhejian Lai, Jiajun Chen 0001, Hao Yang 0006, Shujian Huang |
ACL (1) | 5 |
| 2025 | SLAM: Towards Efficient Multilingual Reasoning via Selective Language AlignmentabstractDespite the significant improvements achieved by large language models (LLMs) in English reasoning tasks, these models continue to struggle with multilingual reasoning. Recent studies leverage a full-parameter and two-stage training paradigm to teach models to first understand non-English questions and then reason. However, this method suffers from both substantial computational resource computing and catastrophic forgetting. The fundamental cause is that, with the primary goal of enhancing multilingual comprehension, an excessive number of irrelevant layers and parameters are tuned during the first stage. Given our findings that the representation learning of languages is merely conducted in lower-level layers, we propose an efficient multilingual reasoning alignment approach that precisely identifies and fine-tunes the layers responsible for handling multilingualism. Experimental results show that our method, SLAM, only tunes 6 layers’ feed-forward sub-layers including 6.5-8% of all parameters within 7B and 13B LLMs, achieving superior average performance than all strong baselines across 10 languages. Meanwhile, SLAM only involves one training stage, reducing training time by 4.1-11.9× compared to the two-stage method. Yuchun Fan, Yongyu Mu, Lei Huang 0021, Junhao Ruan, Tong Xiao 0001, Shujian Huang |
COLING | 8 |
| 2025 | Self-Evolution Knowledge Distillation for LLM-based Machine TranslationabstractKnowledge distillation (KD) has shown great promise in transferring knowledge from larger teacher models to smaller student models. However, existing KD strategies for large language models often minimize output distributions between student and teacher models indiscriminately for each token. This overlooks the imbalanced nature of tokens and their varying transfer difficulties. In response, we propose a distillation strategy called Self-Evolution KD. The core of this approach involves dynamically integrating teacher distribution and one-hot distribution of ground truth into the student distribution as prior knowledge, which promotes the distillation process. It adjusts the ratio of prior knowledge based on token learning difficulty, fully leveraging the teacher model’s potential. Experimental results show our method brings an average improvement of approximately 1.4 SacreBLEU points across four translation directions in the WMT22 test sets. Further analysis indicates that the improvement comes from better knowledge transfer from teachers, confirming our hypothesis. Yuncheng Song, Liang Ding 0006, Changtong Zan, Shujian Huang |
COLING | 4 |
| 2025 | SDGO: Self-Discrimination-Guided Optimization for Consistent Safety in Large Language ModelsabstractLarge Language Models (LLMs) excel at various natural language processing tasks but remain vulnerable to jailbreaking attacks that induce harmful content generation.In this paper, we reveal a critical safety inconsistency: LLMs can more effectively identify harmful requests as discriminators than defend against them as generators.This insight inspires us to explore aligning the model's inherent discrimination and generation capabilities.To this end, we propose SDGO (Self-Discrimination-Guided Optimization), a reinforcement learning framework that leverages the model's own discrimination capabilities as a reward signal to enhance generation safety through iterative selfimprovement.Our method does not require any additional annotated data or external models during the training phase.Extensive experiments demonstrate that SDGO significantly improves model safety compared to both promptbased and training-based baselines while maintaining helpfulness on general benchmarks.By aligning LLMs' discrimination and generation capabilities, SDGO brings robust performance against out-of-distribution (OOD) jailbreaking attacks.This alignment achieves tighter coupling between these two capabilities, enabling the model's generation capability to be further enhanced with only a small amount of discriminative samples. Peng Ding 0001, Dailin Li, Jiajun Chen 0001, Shujian Huang |
EMNLP | 7 |
| 2025 | Understanding LLMs' Cross-Lingual Context Retrieval: How Good It Is And Where It Comes FromabstractCross-lingual context retrieval (extracting contextual information in one language based on requests in another) is a fundamental aspect of cross-lingual alignment, but the performance and mechanism of it for large language models (LLMs) remains unclear.In this paper, we evaluate the cross-lingual context retrieval of over 40 LLMs across 12 languages, using cross-lingual machine reading comprehension (xMRC) as a representative scenario.Our results show that post-trained open LLMs show strong cross-lingual context retrieval ability, comparable to closed-source LLMs such as GPT-4o, and their estimated oracle performances greatly improve after posttraining.Our mechanism analysis shows that the cross-lingual context retrieval process can be divided into two main phases: question encoding and answer retrieval, which are formed in pre-training and post-training respectively.The phasing stability correlates with xMRC performance, and the xMRC bottleneck lies at the last model layers in the second phase, where the effect of post-training can be evidently observed.Our results also indicate that largerscale pretraining cannot improve the xMRC performance.Instead, larger LLMs need further multilingual post-training to fully unlock their cross-lingual context retrieval potential. Changjiang Gao, Hankun Lin, Xue Han 0018, Junlan Feng, Chao Deng 0002, Jiajun Chen 0001, Shujian Huang |
EMNLP | 8 |
| 2025 | R-PRM: Reasoning-Driven Process Reward ModelingabstractProcess Reward Models (PRMs) have emerged as a promising solution to address the reasoning mistakes of large language models (LLMs).However, existing PRMs typically output evaluation scores directly, limiting both learning efficiency and evaluation accuracy.This limitation is further compounded by the scarcity of annotated data.To address these issues, we propose Reasoning-Driven Process Reward Modeling (R-PRM), which activates inherent reasoning to enhance process-level evaluation.First, we leverage stronger LLMs to generate seed data from limited annotations, effectively activating reasoning capabilities and enabling comprehensive step-by-step evaluation.Second, we explore self-improvement of our PRM through preference optimization, without requiring additional annotated data.Third, we introduce inference time scaling to fully harness our model's reasoning potential.Extensive experiments demonstrate R-PRM's effectiveness: on ProcessBench and PRMBench, it surpasses strong baselines by 13.9 and 8.5 F1 scores.When applied to guide mathematical reasoning, R-PRM achieves consistent accuracy improvements of over 8.6 points across six challenging datasets.Further analysis reveals that R-PRM exhibits more comprehensive evaluation and robust generalization, indicating its broader potential.Problem Previous Steps Previous Steps Analysis: This step starts by ... ...... Shuaijie She, Junxiao Liu, Jiajun Chen 0001, Shujian Huang |
EMNLP | 6 |
| 2025 | EnAnchored-X2X: English-Anchored Optimization for Many-to-Many TranslationabstractLarge language models (LLMs) have demonstrated strong machine translation capabilities for English-centric language pairs but underperform in direct non-English (x2x) translation.This work addresses this limitation through a synthetic data generation framework that leverages models' established English-to-x (en2x) capabilities.By extending English parallel corpora into omnidirectional datasets and developing an English-referenced quality evaluation proxy, we enable effective collection of high-quality x2x training data.Combined with preference-based optimization, our method achieves significant improvement across 72 x2x directions for widely used LLMs, while generalizing to enhance en2x performance.The results demonstrate that strategic exploitation of English-centric strengths can bootstrap comprehensive multilingual translation capabilities in LLMs.We release codes, datasets, and model checkpoints at Sen Yang 0015, Jiajun Chen 0001, Shujian Huang, Shanbo Cheng |
EMNLP | 5 |
| 2025 | "I've Heard of You!": Generate Spoken Named Entity Recognition Data for Unseen EntitiesabstractSpoken named entity recognition (NER) aims to identify named entities from speech, playing an important role in speech processing. New named entities appear every day, however, annotating their Spoken NER data is costly. In this paper, we demonstrate that existing Spoken NER systems perform poorly when dealing with previously unseen named entities. To tackle this challenge, we propose a method for generating Spoken NER data based on a named entity dictionary (NED) to reduce costs. Specifically, we first use a large language model (LLM) to generate sentences from the sampled named entities and then use a text-to-speech (TTS) system to generate the speech. Furthermore, we introduce a noise metric to filter out noisy data. To evaluate our approach, we release a novel Spoken NER benchmark along with a corresponding NED containing 8,853 entities. Experiment results show that our method achieves state-of-the-art (SOTA) performance in the in-domain, zero-shot domain adaptation, and fully zero-shot settings. Our data will be available at https://github.com/DeepLearnXMU/HeardU. Xiang Geng, Yuang Li, Mengxin Ren, Wei Tang 0013, Jiahuan Li, Zhibin Lan, Min Zhang 0042, Hao Yang 0006, Shujian Huang, Jinsong Su |
ICASSP | 10 |
| 2025 | DPLM-2: A Multimodal Diffusion Protein Language ModelabstractProteins are essential macromolecules defined by their amino acid sequences, which determine their three-dimensional structures and, consequently, their functions in all living organisms. Therefore, generative protein modeling necessitates a multimodal approach to simultaneously model, understand, and generate both sequences and structures. However, existing methods typically use separate models for each modality, limiting their ability to capture the intricate relationships between sequence and structure. This results in suboptimal performance in tasks that requires joint understanding and generation of both modalities.
In this paper, we introduce DPLM-2, a multimodal protein foundation model that extends discrete diffusion protein language model (DPLM) to accommodate both sequences and structures.
To enable structural learning with the language model, 3D coordinates are converted to discrete tokens using a lookup-free quantization-based tokenizer.
By training on both experimental and high-quality synthetic structures, DPLM-2 learns the joint distribution of sequence and structure, as well as their marginals and conditionals.
We also implement an efficient warm-up strategy to exploit the connection between large-scale evolutionary data and structural inductive biases from pre-trained sequence-based protein language models.
Empirical evaluation shows that DPLM-2 can simultaneously generate highly compatible amino acid sequences and their corresponding 3D structures eliminating the need for a two-stage generation approach.
Moreover, DPLM-2 demonstrates competitive performance in various conditional generation tasks, including folding, inverse folding, and scaffolding with multimodal motif inputs. Xinyou Wang, Zaixiang Zheng, Dongyu Xue, Shujian Huang, Quanquan Gu |
ICLR | 5 |
| 2025 | Elucidating the Design Space of Multimodal Protein Language ModelsabstractMultimodal protein language models (PLMs) integrate sequence and token-based structural information, serving as a powerful foundation for protein modeling, generation, and design. However, the reliance on tokenizing 3D structures into discrete tokens causes substantial loss of fidelity about fine-grained structural details and correlations. In this paper, we systematically elucidate the design space of multimodal PLMs to overcome their limitations. We identify tokenization loss and inaccurate structure token predictions by the PLMs as major bottlenecks. To address these, our proposed design space covers improved generative modeling, structure-aware architectures and representation learning, and data exploration. Our advancements approach finer-grained supervision, demonstrating that token-based multimodal PLMs can achieve robust structural modeling. The effective design methods dramatically improve the structure generation diversity, and notably, folding abilities of our 650M model by reducing the RMSD from 5.52 to 2.36 on PDB testset, even outperforming 3B baselines and on par with the specialized folding models. Project page and code: https://bytedance.github.io/dplm/dplm-2.1. Cheng-Yen Hsieh, Xinyou Wang, Daiheng Zhang, Dongyu Xue, Shujian Huang, Zaixiang Zheng, Quanquan Gu |
ICML | 6 |
| 2025 | Large Language Models Are Cross-Lingual Knowledge-Free ReasonersabstractPeng Hu, Sizhe Liu, Changjiang Gao, Xin Huang, Xue Han, Junlan Feng, Chao Deng, Shujian Huang. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Sizhe Liu, Changjiang Gao, Xue Han 0018, Junlan Feng, Chao Deng 0002, Shujian Huang |
NAACL (Long Papers) | 8 |
| 2025 | Multi-candidate Speculative Decoding
Sen Yang 0015, Shujian Huang, Xinyu Dai, Jiajun Chen 0001 |
NLPCC (1) | 2 |
| 2024 | Measuring Meaning Composition in the Human Brain with Composition Scores from Large Language ModelsabstractThe process of meaning composition, wherein smaller units like morphemes or words combine to form the meaning of phrases and sentences, is essential for human sentence comprehension.Despite extensive neurolinguistic research into the brain regions involved in meaning composition, a computational metric to quantify the extent of composition is still lacking.Drawing on the key-value memory interpretation of transformer feed-forward network blocks, we introduce the Composition Score, a novel model-based metric designed to quantify the degree of meaning composition during sentence comprehension.Experimental findings show that this metric correlates with brain clusters associated with word frequency, structural processing, and general sensitivity to words, suggesting the multifaceted nature of meaning composition during human sentence comprehension. 1 * Corresponding authors, equal contribution.1 Our code and data are released on GitHub. Changjiang Gao, Jixing Li, Jiajun Chen 0001, Shujian Huang |
ACL (1) | 4 |
| 2024 | MAPO: Advancing Multilingual Reasoning through Multilingual-Alignment-as-Preference OptimizationabstractShuaijie She, Wei Zou, Shujian Huang, Wenhao Zhu, Xiang Liu, Xiang Geng, Jiajun Chen. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Shuaijie She, Shujian Huang, Xiang Liu 0023, Xiang Geng, Jiajun Chen 0001 |
ACL (1) | 3 |
| 2024 | Formality is Favored: Unraveling the Learning Preferences of Large Language Models on Data with Conflicting KnowledgeabstractHaving been trained on massive pretraining data, large language models have shown excellent performance on many knowledge-intensive tasks.However, pretraining data tends to contain misleading and even conflicting information, and it is intriguing to understand how LLMs handle these noisy data during training.In this study, we systematically analyze LLMs' learning preferences for data with conflicting knowledge.We find that pretrained LLMs establish learning preferences similar to humans, i.e., preferences towards formal texts and texts with fewer spelling errors, resulting in faster learning and more favorable treatment of knowledge in data with such features when facing conflicts.This finding is generalizable across models and languages and is more evident in larger models.An in-depth analysis reveals that LLMs tend to trust data with features that signify consistency with the majority of data, and it is possible to instill new preferences and erase old ones by manipulating the degree of consistency with the majority data. Jiahuan Li, Yiqing Cao, Shujian Huang, Jiajun Chen 0001 |
EMNLP | 3 |
| 2024 | PreAlign: Boosting Cross-Lingual Transfer by Early Establishment of Multilingual AlignmentabstractLarge language models demonstrate reasonable multilingual abilities, despite predominantly English-centric pretraining.However, the spontaneous multilingual alignment in these models is shown to be weak, leading to unsatisfactory cross-lingual transfer and knowledge sharing.Previous works attempt to address this issue by explicitly injecting multilingual alignment information during or after pretraining.Thus for the early stage in pretraining, the alignment is weak for sharing information or knowledge across languages.In this paper, we propose PREALIGN, a framework that establishes multilingual alignment prior to language model pretraining.PREALIGN injects multilingual alignment by initializing the model to generate similar representations of aligned words and preserves this alignment using a code-switching strategy during pretraining.Extensive experiments in a synthetic English to English-Clone setting demonstrate that PREALIGN significantly outperforms standard multilingual joint training in language modeling, zero-shot crosslingual transfer, and cross-lingual knowledge application.Further experiments in real-world scenarios further validate PREALIGN's effectiveness across various languages and model sizes. Jiahuan Li, Shujian Huang, Aarron Ching, Xinyu Dai, Jiajun Chen 0001 |
EMNLP | 2 |
| 2024 | Getting More from Less: Large Language Models are Good Spontaneous Multilingual LearnersabstractShimao Zhang, Changjiang Gao, Wenhao Zhu, Jiajun Chen, Xin Huang, Xue Han, Junlan Feng, Chao Deng, Shujian Huang. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024. Shimao Zhang, Changjiang Gao, Jiajun Chen 0001, Xue Han 0018, Junlan Feng, Chao Deng 0002, Shujian Huang |
EMNLP | 9 |
| 2024 | EfficientRAG: Efficient Retriever for Multi-Hop Question AnsweringabstractZiyuan Zhuang, Zhiyang Zhang, Sitao Cheng, Fangkai Yang, Jia Liu, Shujian Huang, Qingwei Lin, Saravan Rajmohan, Dongmei Zhang, Qi Zhang. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024. Ziyuan Zhuang, Sitao Cheng, Fangkai Yang, Shujian Huang, Qingwei Lin, Saravan Rajmohan, Dongmei Zhang 0001, Qi Zhang 0066 |
EMNLP | 6 |
| 2024 | Diffusion Language Models Are Versatile Protein LearnersabstractThis paper introduces diffusion protein language model (DPLM), a versatile protein language model that demonstrates strong generative and predictive capabilities for protein sequences. We first pre-train scalable DPLMs from evolutionary-scale protein sequences within a generative self-supervised discrete diffusion probabilistic framework, which generalizes language modeling for proteins in a principled way. After pre-training, DPLM exhibits the ability to generate structurally plausible, novel and diverse protein sequences for unconditional generation. We further demonstrate the proposed diffusion generative pre-training make DPLM possess a better understanding of proteins, making it a superior representation learner, which can be fine-tuned for various predictive tasks, comparing favorably to ESM2. Moreover, DPLM can be tailored for various needs, which showcases its prowess of conditional generation in several ways: (1) conditioning on partial peptide sequences, e.g., generating scaffolds for functional motifs with high success rate; (2) incorporating other modalities as conditioners, e.g., structure-conditioned generation for inverse folding; and (3) steering sequence generation towards desired properties, e.g., satisfying specified secondary structures, through a plug-and-play classifier guidance. Xinyou Wang, Zaixiang Zheng, Dongyu Xue, Shujian Huang, Quanquan Gu |
ICML | 5 |
| 2024 | Hallu-PI: Evaluating Hallucination in Multi-modal Large Language Models within Perturbed InputsabstractMulti-modal Large Language Models (MLLMs) have demonstrated remarkable performance on various visual-language understanding and generation tasks. However, MLLMs occasionally generate content inconsistent with the given images, which is known as "hallucination". Prior works primarily center on evaluating hallucination using standard, unperturbed benchmarks, which overlook the prevalent occurrence of perturbed inputs in real-world scenarios-such as image cropping or blurring-that are critical for a comprehensive assessment of MLLMs' hallucination. In this paper, to bridge this gap, we propose Hallu-PI, the first benchmark designed to evaluate Hallucination in MLLMs within Perturbed Inputs. Specifically, Hallu-PI consists of seven perturbed scenarios, containing 1,260 perturbed images from 11 object types. Each image is accompanied by detailed annotations, which include fine-grained hallucination types, such as existence, attribute, and relation. We equip these annotations with a rich set of questions, making Hallu-PI suitable for both discriminative and generative tasks. Extensive experiments on 12 mainstream MLLMs, such as GPT-4V and Gemini-Pro Vision, demonstrate that these models exhibit significant hallucinations on Hallu-PI, which is not observed in unperturbed scenarios. Furthermore, our research reveals a severe bias in MLLMs' ability to handle different types of hallucinations. We also design two baselines specifically for perturbed scenarios, namely Perturbed-Reminder and Perturbed-ICL. We hope that our study will bring researchers' attention to the limitations of MLLMs when dealing with perturbed inputs, and spur further investigations to address this issue. Our code and datasets are publicly available at https://github.com/NJUNLP/Hallu-PI. Peng Ding 0001, Jingyu Wu, Jun Kuang, Dan Ma 0008, Xuezhi Cao, Shi Chen 0005, Jiajun Chen 0001, Shujian Huang |
ACM Multimedia | 9 |
| 2024 | A Wolf in Sheep's Clothing: Generalized Nested Jailbreak Prompts can Fool Large Language Models EasilyabstractPeng Ding, Jun Kuang, Dan Ma, Xuezhi Cao, Yunsen Xian, Jiajun Chen, Shujian Huang. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Peng Ding 0001, Jun Kuang, Dan Ma 0008, Xuezhi Cao, Yunsen Xian, Jiajun Chen 0001, Shujian Huang |
NAACL-HLT | 7 |
| 2024 | Multilingual Pretraining and Instruction Tuning Improve Cross-Lingual Knowledge Alignment, But Only ShallowlyabstractChangjiang Gao, Hongda Hu, Peng Hu, Jiajun Chen, Jixing Li, Shujian Huang. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Changjiang Gao, Hongda Hu, Jiajun Chen 0001, Jixing Li, Shujian Huang |
NAACL-HLT | 6 |
| 2024 | MT-PATCHER: Selective and Extendable Knowledge Distillation from Large Language Models for Machine TranslationabstractJiahuan Li, Shanbo Cheng, Shujian Huang, Jiajun Chen. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Jiahuan Li, Shanbo Cheng, Shujian Huang, Jiajun Chen 0001 |
NAACL-HLT | 3 |
| 2024 | Exploring the Factual Consistency in Dialogue Comprehension of Large Language ModelsabstractShuaijie She, Shujian Huang, Xingyun Wang, Yanke Zhou, Jiajun Chen. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Shuaijie She, Shujian Huang, Xingyun Wang, Yanke Zhou, Jiajun Chen 0001 |
NAACL-HLT | 2 |
| 2024 | Eliciting the Translation Ability of Large Language Models via Multilingual Finetuning with Translation InstructionsabstractAbstract Large-scale pretrained language models (LLMs), such as ChatGPT and GPT4, have shown strong abilities in multilingual translation, without being explicitly trained on parallel corpora. It is intriguing how the LLMs obtain their ability to carry out translation instructions for different languages. In this paper, we present a detailed analysis by finetuning a multilingual pretrained language model, XGLM-7.5B, to perform multilingual translation following given instructions. Firstly, we show that multilingual LLMs have stronger translation abilities than previously demonstrated. For a certain language, the translation performance depends on its similarity to English and the amount of data used in the pretraining phase. Secondly, we find that LLMs’ ability to carry out translation instructions relies on the understanding of translation instructions and the alignment among different languages. With multilingual finetuning with translation instructions, LLMs could learn to perform the translation task well even for those language pairs unseen during the instruction tuning phase. Jiahuan Li, Hao Zhou 0012, Shujian Huang, Shanbo Cheng, Jiajun Chen 0001 |
Trans. Assoc. Comput. Linguistics | 3 |
| 2023 | Denoising Pre-training for Machine Translation Quality Estimation with Curriculum LearningabstractQuality estimation (QE) aims to assess the quality of machine translations when reference translations are unavailable. QE plays a crucial role in many real-world applications of machine translation. Because labeled QE data are usually limited in scale, recent research, such as DirectQE, pre-trains QE models with pseudo QE data and obtains remarkable performance. However, there tends to be inevitable noise in the pseudo data, hindering models from learning QE accurately. Our study shows that the noise mainly comes from the differences between pseudo and real translation outputs. To handle this problem, we propose CLQE, a denoising pre-training framework for QE based on curriculum learning. More specifically, we propose to measure the degree of noise in the pseudo QE data with some metrics based on statistical or distributional features. With the guidance of these metrics, CLQE gradually pre-trains the QE model using data from cleaner to noisier. Experiments on various benchmarks reveal that CLQE outperforms DirectQE and other strong baselines. We also show that with our framework, pre-training converges faster than directly using the pseudo data. We make our CLQE code available (https://github.com/NJUNLP/njuqe). Xiang Geng, Jiahuan Li, Shujian Huang, Hao Yang 0006, Shimin Tao, Jiajun Chen 0001 |
AAAI | 4 |
| 2023 | Selective Knowledge Distillation for Non-Autoregressive Neural Machine TranslationabstractBenefiting from the sequence-level knowledge distillation, the Non-Autoregressive Transformer (NAT) achieves great success in neural machine translation tasks. However, existing knowledge distillation has side effects, such as propagating errors from the teacher to NAT students, which may limit further improvements of NAT models and are rarely discussed in existing research. In this paper, we introduce selective knowledge distillation by introducing an NAT evaluator to select NAT-friendly targets that are of high quality and easy to learn. In addition, we introduce a simple yet effective progressive distillation method to boost NAT performance. Experiment results on multiple WMT language directions and several representative NAT models show that our approach can realize a flexible trade-off between the quality and complexity of training data for NAT models, achieving strong performances. Further analysis shows that distilling only 5% of the raw translations can help an NAT outperform its counterpart trained on raw data by about 2.4 BLEU. Chengqi Zhao, Shujian Huang |
AAAI | 4 |
| 2023 | CoP: Factual Inconsistency Detection by Controlling the PreferenceabstractAbstractive summarization is the process of generating a summary given a document as input. Although significant progress has been made, the factual inconsistency between the document and the generated summary still limits its practical applications. Previous work found that the probabilities assigned by the generation model reflect its preferences for the generated summary, including the preference for factual consistency, and the preference for the language or knowledge prior as well. To separate the preference for factual consistency, we propose an unsupervised framework named CoP by controlling the preference of the generation model with the help of prompt. More specifically, the framework performs an extra inference step in which a text prompt is introduced as an additional input. In this way, another preference is described by the generation probability of this extra inference process. The difference between the above two preferences, i.e. the difference between the probabilities, could be used as measurements for detecting factual inconsistencies. Interestingly, we found that with the properly designed prompt, our framework could evaluate specific preferences and serve as measurements for fine-grained categories of inconsistency, such as entity-related inconsistency, coreference-related inconsistency, etc. Moreover, our framework could also be extended to the supervised setting to learn better prompt from the labeled data as well. Experiments show that our framework achieves new SOTA results on three factual inconsisency detection tasks. Shuaijie She, Xiang Geng, Shujian Huang, Jiajun Chen 0001 |
AAAI | 3 |
| 2023 | BLEURT Has Universal Translations: An Analysis of Automatic Metrics by Minimum Risk TrainingabstractAutomatic metrics play a crucial role in machine translation.Despite the widespread use of n-gram-based metrics, there has been a recent surge in the development of pre-trained model-based metrics that focus on measuring sentence semantics.However, these neural metrics, while achieving higher correlations with human evaluations, are often considered to be black boxes with potential biases that are difficult to detect.In this study, we systematically analyze and compare various mainstream and cutting-edge automatic metrics from the perspective of their guidance for training machine translation systems.Through Minimum Risk Training (MRT), we find that certain metrics exhibit robustness defects, such as the presence of universal adversarial translations in BLEURT and BARTScore.In-depth analysis suggests two main causes of these robustness deficits: distribution biases in the training datasets, and the tendency of the metric paradigm.By incorporating token-level constraints, we enhance the robustness of evaluation metrics, which in turn leads to an improvement in the performance of machine translation systems.Codes are available at https://github.com/ powerpuffpomelo/fairseq_mrt. Tao Wang 0086, Chengqi Zhao, Shujian Huang, Jiajun Chen 0001, Mingxuan Wang |
ACL (1) | 4 |
| 2023 | Local Interpretation of Transformer Based on Linear DecompositionabstractIn recent years, deep neural networks (DNNs) have achieved state-of-the-art performance on a wide range of tasks.However, limitations in interpretability have hindered their applications in the real world.This work proposes to interpret neural networks by linear decomposition and finds that the ReLU-activated Transformer can be considered as a linear model on a single input.We further leverage the linearity of the model and propose a linear decomposition of the model output to generate local explanations.Our evaluation of sentiment classification and machine translation shows that our method achieves competitive performance in efficiency and fidelity of explanation.In addition, we demonstrate the potential of our approach in applications with examples of error analysis on multiple tasks. Sen Yang 0015, Shujian Huang, Xinyu Dai, Jiajun Chen 0001 |
ACL (1) | 2 |
| 2023 | INK: Injecting kNN Knowledge in Nearest Neighbor Machine TranslationabstractNeural machine translation has achieved promising results on many translation tasks.However, previous studies have shown that neural models induce a non-smooth representation space, which harms its generalization results.Recently, kNN-MT has provided an effective paradigm to smooth the prediction based on neighbor representations during inference.Despite promising results, kNN-MT usually requires large inference overhead.We propose an effective training framework INK to directly smooth the representation space via adjusting representations of kNN neighbors with a small number of new parameters.The new parameters are then used to refresh the whole representation datastore to get new kNN knowledge asynchronously.This loop keeps running until convergence.Experiments on four benchmark datasets show that INK achieves average gains of 1.99 COMET and 1.0 BLEU, outperforming the state-of-the-art kNN-MT system with 0.02× memory space and 1.9× inference speedup 1 . Jingjing Xu 0001, Shujian Huang, Lingpeng Kong, Jiajun Chen 0001 |
ACL (1) | 3 |
| 2023 | Improved Pseudo Data for Machine Translation Quality Estimation with Constrained Beam SearchabstractXiang Geng, Yu Zhang, Zhejian Lai, Shuaijie She, Wei Zou, Shimin Tao, Hao Yang, Jiajun Chen, Shujian Huang. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 2023. Xiang Geng, Zhejian Lai, Shuaijie She, Shimin Tao, Hao Yang 0006, Jiajun Chen 0001, Shujian Huang |
EMNLP | 9 |
| 2023 | IMTLab: An Open-Source Platform for Building, Evaluating, and Diagnosing Interactive Machine Translation SystemsabstractXu Huang, Zhirui Zhang, Ruize Gao, Yichao Du, Lemao Liu, Guoping Huang, Shuming Shi, Jiajun Chen, Shujian Huang. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 2023. Zhirui Zhang, Yichao Du, Lemao Liu, Guoping Huang, Shuming Shi 0001, Jiajun Chen 0001, Shujian Huang |
EMNLP | 9 |
| 2023 | Addressing Linguistic Bias through a Contrastive Analysis of Academic Writing in the NLP DomainabstractIt has been well documented that a reviewer's opinion of the nativeness of expression in an academic paper affects the likelihood of it being accepted for publication.Previous works have also shone a light on the stress and anxiety authors who are non-native English speakers experience when attempting to publish in international venues.We explore how this might be a concern in the field of Natural Language Processing (NLP) through conducting a comprehensive statistical analysis of NLP paper abstracts, identifying how authors of different linguistic backgrounds differ in the lexical, morphological, syntactic and cohesive aspects of their writing.Through our analysis, we identify that there are a number of characteristics that are highly variable across the different corpora examined in this paper.This indicates potential for the presence of linguistic bias.Therefore, we outline a set of recommendations to publishers of academic journals and conferences regarding their guidelines and resources for prospective authors in order to help enhance inclusivity and fairness. Robert Ridley, Zhen Wu 0002, Shujian Huang, Xinyu Dai |
EMNLP | 4 |
| 2023 | Only 5% Attention Is All You Need: Efficient Long-range Document-level Neural Machine TranslationabstractZihan Liu, Zewei Sun, Shanbo Cheng, Shujian Huang, Mingxuan Wang. Proceedings of the 13th International Joint Conference on Natural Language Processing and the 3rd Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Zewei Sun, Shanbo Cheng, Shujian Huang, Mingxuan Wang |
IJCNLP (1) | 4 |
| 2023 | Food-500 Cap: A Fine-Grained Food Caption Benchmark for Evaluating Vision-Language ModelsabstractVision-language models (VLMs) have shown impressive performance in substantial downstream multi-modal tasks. However, only comparing the fine-tuned performance on downstream tasks leads to the poor interpretability of VLMs, which is adverse to their future improvement. Several prior works have identified this issue and used various probing methods under a zero-shot setting to detect VLMs' limitations, but they all examine VLMs using general datasets instead of specialized ones. In practical applications, VLMs are usually applied to specific scenarios, such as e-commerce and news fields, so the generalization of VLMs in specific domains should be given more attention. In this paper, we comprehensively investigate the capabilities of popular VLMs in a specific field, the food domain. To this end, we build a food caption dataset, Food-500 Cap, which contains 24,700 food images with 494 categories. Each image is accompanied by a detailed caption, including fine-grained attributes of food, such as the ingredient, shape, and color. We also provide a culinary culture taxonomy that classifies each food category based on its geographic origin in order to better analyze the performance differences of VLM in different regions. Experiments on our proposed datasets demonstrate that popular VLMs underperform in the food domain compared with their performance in the general domain. Furthermore, our research reveals severe bias in VLMs' ability to handle food items from different geographic regions. We adopt diverse probing methods and evaluate nine VLMs belonging to different architectures to verify the aforementioned observations. We hope that our study will bring researchers' attention to VLM's limitations when applying them to the domain of food or culinary cultures, and spur further investigations to address this issue. Zheng Ma 0012, Mianzhi Pan, Kanzhi Cheng, Shujian Huang, Jiajun Chen 0001 |
ACM Multimedia | 6 |
| 2022 | Non-parametric Online Learning from Human Feedback for Neural Machine TranslationabstractWe study the problem of online learning with human feedback in the human-in-the-loop machine translation, in which the human translators revise the machine-generated translations and then the corrected translations are used to improve the neural machine translation (NMT) system. However, previous methods require online model updating or additional translation memory networks to achieve high-quality performance, making them inflexible and inefficient in practice. In this paper, we propose a novel non-parametric online learning method without changing the model structure. This approach introduces two k-nearest-neighbor (KNN) modules: one module memorizes the human feedback, which is the correct sentences provided by human translators, while the other balances the usage of the history human feedback and original NMT models adaptively. Experiments conducted on EMEA and JRC-Acquis benchmarks demonstrate that our proposed method obtains substantial improvements on translation accuracy and achieves better adaptation performance with less repeating human correction operations. Dongqi Wang 0005, Zhirui Zhang, Shujian Huang, Jiajun Chen 0001 |
AAAI | 4 |
| 2022 | latent-GLAT: Glancing at Latent Variables for Parallel Text GenerationabstractYu Bao, Hao Zhou, Shujian Huang, Dongqi Wang, Lihua Qian, Xinyu Dai, Jiajun Chen, Lei Li. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022. Hao Zhou 0012, Shujian Huang, Dongqi Wang 0005, Lihua Qian, Xinyu Dai, Jiajun Chen 0001, Lei Li 0005 |
ACL (1) | 3 |
| 2022 | BiTIIMT: A Bilingual Text-infilling Method for Interactive Machine TranslationabstractYanling Xiao, Lemao Liu, Guoping Huang, Qu Cui, Shujian Huang, Shuming Shi, Jiajun Chen. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022. Yanling Xiao, Lemao Liu, Guoping Huang, Qu Cui, Shujian Huang, Shuming Shi 0001, Jiajun Chen 0001 |
ACL (1) | 5 |
| 2022 | Towards Multi-label Unknown Intent DetectionabstractMulti-class unknown intent detection has made remarkable progress recently. However, it has a strong assumption that each utterance has only one intent, which does not conform to reality because utterances often have multiple intents. In this paper, we propose a more desirable task, multi-label unknown intent detection, to detect whether the utterance contains the unknown intent, in which each utterance may contain multiple intents. In this task, the unique utterances simultaneously containing known and unknown intents make existing multi-class methods easy to fail. To address this issue, we propose an intuitive and effective method to recognize whether All Intents contained in the utterance are Known (AIK). Our high-level idea is to predict the utterance’s intent number, then check whether the utterance contains the same number of known intents. If the number of known intents is less than the number of intents, it implies that the utterance also contains unknown intents. We benchmark AIK over existing methods, and empirical results suggest that our method obtains state-of-the-art performances. For example, on the MultiWOZ 2.3 dataset, AIK significantly reduces the FPR95 by 12.25% compared to the best baseline. Yawen Ouyang, Zhen Wu 0002, Xinyu Dai, Shujian Huang, Jiajun Chen 0001 |
COLING | 4 |
| 2022 | Alleviating the Inequality of Attention Heads for Neural Machine TranslationabstractRecent studies show that the attention heads in Transformer are not equal. We relate this phenomenon to the imbalance training of multi-head attention and the model dependence on specific heads. To tackle this problem, we propose a simple masking method: HeadMask, in two specific ways. Experiments show that translation improvements are achieved on multiple language pairs. Subsequent empirical analyses also support our assumption and confirm the effectiveness of the method. Zewei Sun, Shujian Huang, Xinyu Dai, Jiajun Chen 0001 |
COLING | 2 |
| 2022 | Learning from Adjective-Noun Pairs: A Knowledge-enhanced Framework for Target-Oriented Multimodal Sentiment ClassificationabstractTarget-oriented multimodal sentiment classification (TMSC) is a new subtask of aspect-based sentiment analysis, which aims to determine the sentiment polarity of the opinion target mentioned in a (sentence, image) pair. Recently, dominant works employ the attention mechanism to capture the corresponding visual representations of the opinion target, and then aggregate them as evidence to make sentiment predictions. However, they still suffer from two problems: (1) The granularity of the opinion target in two modalities is inconsistent, which causes visual attention sometimes fail to capture the corresponding visual representations of the target; (2) Even though it is captured, there are still significant differences between the visual representations expressing the same mood, which brings great difficulty to sentiment prediction. To this end, we propose a novel Knowledge-enhanced Framework (KEF) in this paper, which can successfully exploit adjective-noun pairs extracted from the image to improve the visual attention capability and sentiment prediction capability of the TMSC task. Extensive experimental results show that our framework consistently outperforms state-of-the-art works on two public datasets. Fei Zhao 0012, Zhen Wu 0002, Siyu Long, Xinyu Dai, Shujian Huang, Jiajun Chen 0001 |
COLING | 5 |
| 2022 | Structure-Unified M-Tree Coding Solver for Math Word ProblemabstractAs one of the challenging NLP tasks, designing math word problem (MWP) solvers has attracted increasing research attention for the past few years.In previous work, models designed by taking into account the properties of the binary tree structure of mathematical expressions at the output side have achieved better performance.However, the expressions corresponding to a MWP are often diverse (e.g., n 1 +n 2 ×n 3 -n 4 , n 3 ×n 2 -n 4 +n 1 , etc.), and so are the corresponding binary trees, which creates difficulties in model learning due to the non-deterministic output space.In this paper, we propose the Structure-Unified M-Tree Coding Solver (SUMC-Solver), which applies a tree with any M branches (M-tree) to unify the output structures.To learn the M-tree, we use a mapping to convert the M-tree into the M-tree codes, where codes store the information of the paths from tree root to leaf nodes and the information of leaf nodes themselves, and then devise a Sequence-to-Code (seq2code) model to generate the codes.Experimental results on the widely used MAWPS and Math23K datasets have demonstrated that SUMC-Solver not only outperforms several state-of-the-art models under similar experimental settings but also performs much better under low-resource conditions 1 . Bin Wang 0016, Jiangzhou Ju, Xinyu Dai, Shujian Huang, Jiajun Chen 0001 |
EMNLP | 5 |
| 2022 | Helping the Weak Makes You Strong: Simple Multi-Task Learning Improves Non-Autoregressive TranslatorsabstractRecently, non-autoregressive (NAR) neural machine translation models have received increasing attention due to their efficient parallel decoding.However, the probabilistic framework of NAR models necessitates conditional independence assumption on target sequences, falling short of characterizing human language data.This drawback results in less informative learning signals for NAR models under conventional MLE training, thereby yielding unsatisfactory accuracy compared to their autoregressive (AR) counterparts.In this paper, we propose a simple and model-agnostic multitask learning framework to provide more informative learning signals.During training stage, we introduce a set of sufficiently weak AR decoders that solely rely on the information provided by NAR decoder to make prediction, forcing the NAR decoder to become stronger or else it will be unable to support its weak AR partners.Experiments on WMT and IWSLT datasets show that our approach can consistently improve accuracy of multiple NAR baselines without adding any additional decoding overhead. Xinyou Wang, Zaixiang Zheng, Shujian Huang |
EMNLP | 3 |
| 2022 | FGraDA: A Dataset and Benchmark for Fine-Grained Domain Adaptation in Machine TranslationabstractPrevious research for adapting a general neural machine translation (NMT) model into a specific domain usually neglects the diversity in translation within the same domain, which is a core problem for domain adaptation in real-world scenarios. One representative of such challenging scenarios is to deploy a translation system for a conference with a specific topic, e.g., global warming or coronavirus, where there are usually extremely less resources due to the limited schedule. To motivate wider investigation in such a scenario, we present a real-world fine-grained domain adaptation task in machine translation (FGraDA). The FGraDA dataset consists of Chinese-English translation task for four sub-domains of information technology: autonomous vehicles, AI education, real-time networks, and smart phone. Each sub-domain is equipped with a development set and test set for evaluation purposes. To be closer to reality, FGraDA does not employ any in-domain bilingual training data but provides bilingual dictionaries and wiki knowledge base, which can be easier obtained within a short time. We benchmark the fine-grained domain adaptation task and present in-depth analyses showing that there are still challenging problems to further improve the performance with heterogeneous resources. Shujian Huang, Tong Pu, Pingxuan Huang, Wei Chen 0071, Jiajun Chen 0001 |
LREC | 2 |
| 2021 | DirectQE: Direct Pretraining for Machine Translation Quality EstimationabstractMachine Translation Quality Estimation (QE) is a task of predicting the quality of machine translations without relying on any reference. Recently, the predictor-estimator framework trains the predictor as a feature extractor, which leverages the extra parallel corpora without QE labels, achieving promising QE performance. However, we argue that there are gaps between the predictor and the estimator in both data quality and training objectives, which preclude QE models from benefiting from a large number of parallel corpora more directly. We propose a novel framework called DirectQE that provides a direct pretraining for QE tasks. In DirectQE, a generator is trained to produce pseudo data that is closer to the real QE data, and a detector is pretrained on these data with novel objectives that are akin to the QE task. Experiments on widely used benchmarks show that DirectQE outperforms existing methods, without using any pretraining models such as BERT. We also give extensive analyses showing how fixing the two gaps contributes to our improvements. Qu Cui, Shujian Huang, Jiahuan Li, Xiang Geng, Zaixiang Zheng, Guoping Huang, Jiajun Chen 0001 |
AAAI | 2 |
| 2021 | Automated Cross-prompt Scoring of Essay TraitsabstractThe majority of current research in Automated Essay Scoring (AES) focuses on prompt-specific scoring of either the overall quality of an essay or the quality with regards to certain traits. In real-world applications obtaining labelled data for a target essay prompt is often expensive or unfeasible, requiring the AES system to be able to perform well when predicting scores for essays from unseen prompts. As a result, some recent research has been dedicated to cross-prompt AES. However, this line of research has thus far only been concerned with holistic, overall scoring, with no exploration into the scoring of different traits. As users of AES systems often require feedback with regards to different aspects of their writing, trait scoring is a necessary component of an effective AES system. Therefore, to address this need, we introduce a new task named Automated Cross-prompt Scoring of Essay Traits, which requires the model to be trained solely on non-target-prompt essays and to predict the holistic, overall score as well as scores for a number of specific traits for target-prompt essays. This task challenges the model's ability to generalize in order to score essays from a novel domain as well as its ability to represent the quality of essays from multiple different aspects. In addition, we introduce a new, innovative approach which builds on top of a state-of-the-art method for cross-prompt AES. Our method utilizes a trait-attention mechanism and a multi-task architecture that leverages the relationships between each trait to simultaneously predict the overall score and the score of each individual trait. We conduct extensive experiments on the widely used ASAP and ASAP++ datasets and demonstrate that our approach is able to outperform leading prompt-specific trait scoring and cross-prompt AES methods. Robert Ridley, Liang He 0009, Xinyu Dai, Shujian Huang, Jiajun Chen 0001 |
AAAI | 4 |
| 2021 | Learning Kernel-Smoothed Machine Translation with Retrieved ExamplesabstractHow to effectively adapt neural machine translation (NMT) models according to emerging cases without retraining?Despite the great success of neural machine translation, updating the deployed models online remains a challenge.Existing non-parametric approaches that retrieve similar examples from a database to guide the translation process are promising but are prone to overfit the retrieved examples.However, non-parametric methods are prone to overfit the retrieved examples.In this work, we propose to learn Kernel-Smoothed Translation with Example Retrieval (KSTER), an effective approach to adapt neural machine translation models online.Experiments on domain adaptation and multi-domain machine translation datasets show that even without expensive retraining, KSTER is able to achieve improvement of 1.1 to 1.5 BLEU scores over the best existing online adaptation methods.The code and trained models are released at https://github.com/jiangqn/KSTER. Qingnan Jiang, Mingxuan Wang, Shanbo Cheng, Shujian Huang, Lei Li 0005 |
EMNLP (1) | 5 |
| 2021 | Meta-LMTC: Meta-Learning for Large-Scale Multi-Label Text ClassificationabstractLarge-scale multi-label text classification (LMTC) tasks often face long-tailed label distributions, where many labels have few or even no training instances.Although current methods can exploit prior knowledge to handle these few/zero-shot labels, they neglect the metaknowledge contained in the dataset that can guide models to learn with few samples.In this paper, for the first time, this problem is addressed from a meta-learning perspective.However, the simple extension of meta-learning approaches to multi-label classification is suboptimal for LMTC tasks due to long-tailed label distribution and coexisting of few-and zeroshot scenarios.We propose a meta-learning approach named META-LMTC.Specifically, it constructs more faithful and more diverse tasks according to well-designed sampling strategies and directly incorporates the objective of adapting to new low-resource tasks into the metalearning phase.Extensive experiments show that META-LMTC achieves state-of-the-art performance against strong baselines and can still enhance powerful BERTlike models. Ran Wang 0010, Xi'ao Su, Siyu Long, Xinyu Dai, Shujian Huang, Jiajun Chen 0001 |
EMNLP (1) | 5 |
| 2021 | Non-Autoregressive Translation by Learning Target Categorical CodesabstractYu Bao, Shujian Huang, Tong Xiao, Dongqi Wang, Xinyu Dai, Jiajun Chen. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Shujian Huang, Tong Xiao 0001, Dongqi Wang 0005, Xinyu Dai, Jiajun Chen 0001 |
NAACL-HLT | 2 |
| 2021 | Duplex Sequence-to-Sequence Learning for Reversible Machine TranslationabstractSequence-to-sequence learning naturally has two directions. How to effectively utilize supervision signals from both directions? Existing approaches either require two separate models, or a multitask-learned model but with inferior performance. In this paper, we propose REDER (Reversible Duplex Transformer), a parameter-efficient model and apply it to machine translation. Either end of REDER can simultaneously input and output a distinct language. Thus REDER enables {\em reversible machine translation} by simply flipping the input and output ends. Experiments verify that REDER achieves the first success of reversible machine translation, which helps outperform its multitask-trained baselines by up to 1.3 BLEU. Zaixiang Zheng, Hao Zhou 0012, Shujian Huang, Jiajun Chen 0001, Jingjing Xu 0001, Lei Li 0005 |
NeurIPS | 3 |
| 2021 | Dual Side Deep Context-aware Modulation for Social RecommendationabstractSocial recommendation is effective in improving the recommendation performance by leveraging social relations from online social networking platforms. Social relations among users provide friends’ information for modeling users’ interest in candidate items and help items expose to potential consumers (i.e., item attraction). However, there are two issues haven’t been well-studied: Firstly, for the user interests, existing methods typically aggregate friends’ information contextualized on the candidate item only, and this shallow context-aware aggregation makes them suffer from the limited friends’ information. Secondly, for the item attraction, if the item’s past consumers are the friends of or have a similar consumption habit to the targeted user, the item may be more attractive to the targeted user, but most existing methods neglect the relation enhanced context-aware item attraction. Bairan Fu, Wenming Zhang, Guang-Neng Hu, Xinyu Dai, Shujian Huang, Jiajun Chen 0001 |
WWW | 5 |
| 2021 | Integrating heterogeneous thesauruses for Chinese synonyms
Peng Wu 0037, Yingjie Zhang 0002, Shujian Huang, Xinyu Dai, Jiajun Chen 0001 |
Frontiers Comput. Sci. | 4 |
| 2020 | Generating Diverse Translation by Manipulating Multi-Head AttentionabstractTransformer model (Vaswani et al. 2017) has been widely used in machine translation tasks and obtained state-of-the-art results. In this paper, we report an interesting phenomenon in its encoder-decoder multi-head attention: different attention heads of the final decoder layer align to different word translation candidates. We empirically verify this discovery and propose a method to generate diverse translations by manipulating heads. Furthermore, we make use of these diverse translations with the back-translation technique for better data augmentation. Experiment results show that our method generates diverse translations without a severe drop in translation quality. Experiments also show that back-translation with these diverse translations could bring a significant improvement in performance on translation tasks. An auxiliary experiment of conversation response generation task proves the effect of diversity as well. Zewei Sun, Shujian Huang, Xinyu Dai, Jiajun Chen 0001 |
AAAI | 2 |
| 2020 | GRET: Global Representation Enhanced TransformerabstractTransformer, based on the encoder-decoder framework, has achieved state-of-the-art performance on several natural language generation tasks. The encoder maps the words in the input sentence into a sequence of hidden states, which are then fed into the decoder to generate the output sentence. These hidden states usually correspond to the input words and focus on capturing local information. However, the global (sentence level) information is seldom explored, leaving room for the improvement of generation quality. In this paper, we propose a novel global representation enhanced Transformer (GRET) to explicitly model global representation in the Transformer network. Specifically, in the proposed model, an external state is generated for the global representation from the encoder. The global representation is then fused into the decoder during the decoding process to improve generation quality. We conduct experiments in two text generation tasks: machine translation and text summarization. Experimental results on four WMT machine translation tasks and LCSTS text summarization task demonstrate the effectiveness of the proposed approach on natural language generation1. Rongxiang Weng, Shujian Huang, Heng Yu 0006, Lidong Bing, Weihua Luo, Jiajun Chen 0001 |
AAAI | 3 |
| 2020 | Acquiring Knowledge from Pre-Trained Model to Neural Machine TranslationabstractPre-training and fine-tuning have achieved great success in natural language process field. The standard paradigm of exploiting them includes two steps: first, pre-training a model, e.g. BERT, with a large scale unlabeled monolingual data. Then, fine-tuning the pre-trained model with labeled data from downstream tasks. However, in neural machine translation (NMT), we address the problem that the training objective of the bilingual task is far different from the monolingual pre-trained model. This gap leads that only using fine-tuning in NMT can not fully utilize prior language knowledge. In this paper, we propose an Apt framework for acquiring knowledge from pre-trained model to NMT. The proposed approach includes two modules: 1). a dynamic fusion mechanism to fuse task-specific features adapted from general knowledge into NMT network, 2). a knowledge distillation paradigm to learn language knowledge continuously during the NMT training process. The proposed approach could integrate suitable knowledge from pre-trained models to improve the NMT. Experimental results on WMT English to German, German to English and Chinese to English machine translation tasks show that our model outperforms strong baselines and the fine-tuning counterparts. Rongxiang Weng, Heng Yu 0006, Shujian Huang, Shanbo Cheng, Weihua Luo |
AAAI | 3 |
| 2020 | Latent Opinions Transfer Network for Target-Oriented Opinion Words ExtractionabstractTarget-oriented opinion words extraction (TOWE) is a new subtask of ABSA, which aims to extract the corresponding opinion words for a given opinion target in a sentence. Recently, neural network methods have been applied to this task and achieve promising results. However, the difficulty of annotation causes the datasets of TOWE to be insufficient, which heavily limits the performance of neural models. By contrast, abundant review sentiment classification data are easily available at online review sites. These reviews contain substantial latent opinions information and semantic patterns. In this paper, we propose a novel model to transfer these opinions knowledge from resource-rich review sentiment classification datasets to low-resource task TOWE. To address the challenges in the transfer process, we design an effective transformation method to obtain latent opinions, then integrate them into TOWE. Extensive experimental results show that our model achieves better performance compared to other state-of-the-art methods and significantly outperforms the base model without transferring opinions knowledge. Further analysis validates the effectiveness of our model. Zhen Wu 0002, Fei Zhao 0012, Xinyu Dai, Shujian Huang, Jiajun Chen 0001 |
AAAI | 4 |
| 2020 | Explicit Semantic Decomposition for Definition GenerationabstractDefinition generation, which aims to automatically generate dictionary definitions for words, has recently been proposed to assist the construction of dictionaries and help people understand unfamiliar texts.However, previous works hardly consider explicitly modeling the "components" of definitions, leading to under-specific generation results.In this paper, we propose ESD, namely Explicit Semantic Decomposition for definition generation, which explicitly decomposes meaning of words into semantic components, and models them with discrete latent variables for definition generation.Experimental results show that ESD achieves substantial improvements on WordNet and Oxford benchmarks over strong previous baselines. Jiahuan Li, Shujian Huang, Xinyu Dai, Jiajun Chen 0001 |
ACL | 3 |
| 2020 | Dialogue State Tracking with Explicit Slot Connection ModelingabstractRecent proposed approaches have made promising progress in dialogue state tracking (DST). However, in multi-domain scenarios, ellipsis and reference are frequently adopted by users to express values that have been mentioned by slots from other domains. To handle these phenomena, we propose a Dialogue State Tracking with Slot Connections (DST-SC) model to explicitly consider slot correlations across different domains. Given a target slot, the slot connecting mechanism in DST-SC can infer its source slot and copy the source slot value directly, thus significantly reducing the difficulty of learning and reasoning. Experimental results verify the benefits of explicit slot connection modeling, and our model achieves state-of-the-art performance on MultiWOZ 2.0 and MultiWOZ 2.1 datasets. Yawen Ouyang, Moxin Chen, Xinyu Dai, Yinggong Zhao, Shujian Huang, Jiajun Chen 0001 |
ACL | 5 |
| 2020 | A Reinforced Generation of Adversarial Examples for Neural Machine TranslationabstractNeural machine translation systems tend to fail on less decent inputs despite its significant efficacy, which may significantly harm the credibility of these systems-fathoming how and when neural-based systems fail in such cases is critical for industrial maintenance.Instead of collecting and analyzing bad cases using limited handcrafted error features, here we investigate this issue by generating adversarial examples via a new paradigm based on reinforcement learning.Our paradigm could expose pitfalls for a given performance metric, e.g., BLEU, and could target any given neural machine translation architecture.We conduct experiments of adversarial attacks on two mainstream neural machine translation architectures, RNN-search, and Transformer.The results show that our method efficiently produces stable attacks with meaning-preserving adversarial examples.We also present a qualitative and quantitative analysis for the preference pattern of the attack, demonstrating its capability of pitfall exposure. Shujian Huang, Xinyu Dai, Jiajun Chen 0001 |
ACL | 2 |
| 2020 | Enhance Prototypical Network with Text Descriptions for Few-shot Relation ClassificationabstractRecently few-shot relation classification has drawn much attention. It devotes to addressing the long-tail relation problem by recognizing the relations from few instances. The existing metric learning methods aim to learn the prototype of classes and make prediction according to distances between query and prototypes. However, it is likely to make unreliable predictions due to the text diversity. It is intuitive that the text descriptions of relation and entity can provide auxiliary support evidence for relation classification. In this paper, we propose TD-Proto, which enhances prototypical network with relation and entity descriptions. We design a collaborative attention module to extract beneficial and instructional information of sentence and entity respectively. A gate mechanism is proposed to fuse both information dynamically so as to obtain a knowledge-aware instance. Experimental results demonstrate that our method achieves excellent performance. Kaijia Yang, Nantao Zheng, Xinyu Dai, Liang He 0009, Shujian Huang, Jiajun Chen 0001 |
CIKM | 5 |
| 2020 | A Simple and Effective Approach to Robust Unsupervised Bilingual Dictionary InductionabstractUnsupervised Bilingual Dictionary Induction methods based on the initialization and the selflearning have achieved great success in similar language pairs, e.g., English-Spanish.But they still fail and have an accuracy of 0% in many distant language pairs, e.g., English-Japanese.In this work, we show that this failure results from the gap between the actual initialization performance and the minimum initialization performance for the self-learning to succeed.We propose Iterative Dimension Reduction to bridge this gap.Our experiments show that this simple method does not hamper the performance of similar language pairs and achieves an accuracy of 13.64∼55.53%between English and four distant languages, i.e., Chinese, Japanese, Vietnamese and Thai. Yanyang Li, Yingfeng Luo, Quan Du, Huizhen Wang, Shujian Huang, Tong Xiao 0001 |
COLING | 6 |
| 2020 | Mirror-Generative Neural Machine Translation
Zaixiang Zheng, Hao Zhou 0012, Shujian Huang, Lei Li 0005, Xinyu Dai, Jiajun Chen 0001 |
ICLR | 3 |
| 2020 | Towards Making the Most of Context in Neural Machine TranslationabstractDocument-level machine translation manages to outperform sentence level models by a small margin, but have failed to be widely adopted. We argue that previous research did not make a clear use of the global context, and propose a new document-level NMT framework that deliberately models the local context of each sentence with the awareness of the global context of the document in both source and target languages. We specifically design the model to be able to deal with documents containing any number of sentences, including single sentences. This unified approach allows our model to be trained elegantly on standard datasets without needing to train on sentence and document level data separately. Experimental results demonstrate that our model outperforms Transformer baselines and previous document-level NMT models with substantial margins of up to 2.1 BLEU on state-of-the-art baselines. We also provide analyses which show the benefit of context far beyond the neighboring two or three sentences, which previous studies have typically incorporated. Zaixiang Zheng, Xiang Yue, Shujian Huang, Jiajun Chen 0001, Alexandra Birch |
IJCAI | 3 |
| 2020 | Learning to Generate Personalized Query Auto-Completions via a Multi-View Multi-Task Attentive ApproachabstractIn this paper, we study the task of Query Auto-Completion (QAC), which is a very significant feature of modern search engines. In real industrial application, there always exist two major problems of QAC - weak personalization and unseen queries. To address these problems, we propose M2A, a multi-view multi-task attentive framework to learn personalized query auto-completion models. We propose a new Transformer-based hierarchical encoder to model different kinds of sequential behaviors, which can be seen as multiple distinct views of the user's searching history, and then a prefix-to-history attention mechanism is used to select the most relevant information to compose the final intention representation. To learn more informative representations, we propose to incorporate multi-task learning into the model training. Two different kinds of supervisory information provided by query logs are utilized at the same time by jointly training a CTR prediction model and a query generation model. Jiwei Tan, Hongbo Deng, Shujian Huang, Jiajun Chen 0001 |
KDD | 5 |
| 2020 | Transformer-Based Multi-aspect Modeling for Multi-aspect Multi-sentiment Analysis
Zhen Wu 0002, Chengcan Ying, Xinyu Dai, Shujian Huang, Jiajun Chen 0001 |
NLPCC (2) | 4 |
| 2020 | Opinion Transmission Network for Jointly Improving Aspect-Oriented Opinion Words Extraction and Sentiment Classification
Chengcan Ying, Zhen Wu 0002, Xinyu Dai, Shujian Huang, Jiajun Chen 0001 |
NLPCC (1) | 4 |
| 2020 | MSGE: A Multi-step Gated Model for Knowledge Graph Completion
Chunyang Tan, Kaijia Yang, Xinyu Dai, Shujian Huang, Jiajun Chen 0001 |
PAKDD (1) | 4 |
| 2020 | Improving Self-Attention Networks With Sequential RelationsabstractRecently, self-attention networks show strong advantages of sentence modeling in many NLP tasks. However, self-attention mechanism computes the interactions of every pair of words independently regardless of their positions, which makes it not able to capture the sequential relations between words in different positions in a sentence. In this paper, we improve the self-attention networks by better integrating sequential relations, which is essential for modeling natural languages. Specifically, we 1) propose a position-based attention to model the interaction between two words regarding positions; 2) perform separated attention for the context before and after the current position, respectively; and 3) merge the above two parts with a position-aware gated fusion mechanism. Experiments in natural language inference, machine translation and sentiment analysis tasks show that our sequential relation modeling helps self-attention networks outperform existing approaches. We also provide extensive analyses to shed light on what the models have learned about the sequential relations. Zaixiang Zheng, Shujian Huang, Rongxiang Weng, Xinyu Dai, Jiajun Chen 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2019 | Generating Sentences from Disentangled Syntactic and Semantic SpacesabstractVariational auto-encoders (VAEs) are widely used in natural language generation due to the regularization of the latent space.However, generating sentences from the continuous latent space does not explicitly model the syntactic information.In this paper, we propose to generate sentences from disentangled syntactic and semantic spaces.Our proposed method explicitly models syntactic information in the VAE's latent space by using the linearized tree sequence, leading to better performance of language generation.Additionally, the advantage of sampling in the disentangled syntactic and semantic latent spaces enables us to perform novel applications, such as the unsupervised paraphrase generation and syntaxtransfer generation.Experimental results show that our proposed model achieves similar or better performance in various tasks, compared with state-of-the-art related work. Hao Zhou 0012, Shujian Huang, Lei Li 0005, Lili Mou, Olga Vechtomova, Xinyu Dai, Jiajun Chen 0001 |
ACL (1) | 3 |
| 2019 | Learning Representation Mapping for Relation Detection in Knowledge Base Question AnsweringabstractRelation detection is a core step in many natural language process applications including knowledge base question answering.Previous efforts show that single-fact questions could be answered with high accuracy.However, one critical problem is that current approaches only get high accuracy for questions whose relations have been seen in the training data.But for unseen relations, the performance will drop rapidly.The main reason for this problem is that the representations for unseen relations are missing.In this paper, we propose a simple mapping method, named representation adapter, to learn the representation mapping for both seen and unseen relations based on previously learned relation embedding.We employ the adversarial objective and the reconstruction objective to improve the mapping performance.We re-organize the popular Sim-pleQuestion dataset to reveal and evaluate the problem of detecting unseen relations.Experiments show that our method can greatly improve the performance of unseen relations while the performance for those seen part is kept comparable to the state-of-the-art. 1 Peng Wu 0037, Shujian Huang, Rongxiang Weng, Zaixiang Zheng, Jiajun Chen 0001 |
ACL (1) | 2 |
| 2019 | Fine-grained Knowledge Fusion for Sequence Labeling Domain AdaptationabstractHuiyun Yang, Shujian Huang, Xin-Yu Dai, Jiajun Chen. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Huiyun Yang, Shujian Huang, Xinyu Dai, Jiajun Chen 0001 |
EMNLP/IJCNLP (1) | 2 |
| 2019 | Dynamic Past and Future for Neural Machine TranslationabstractZaixiang Zheng, Shujian Huang, Zhaopeng Tu, Xin-Yu Dai, Jiajun Chen. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Zaixiang Zheng, Shujian Huang, Zhaopeng Tu, Xinyu Dai, Jiajun Chen 0001 |
EMNLP/IJCNLP (1) | 2 |
| 2019 | Correct-and-Memorize: Learning to Translate from Interactive RevisionsabstractState-of-the-art machine translation models are still not on a par with human translators. Previous work takes human interactions into the neural machine translation process to obtain improved results in target languages. However, not all model--translation errors are equal -- some are critical while others are minor. In the meanwhile, same translation mistakes occur repeatedly in similar context. To solve both issues, we propose CAMIT, a novel method for translating in an interactive environment. Our proposed method works with critical revision instructions, therefore allows human to correct arbitrary words in model-translated sentences. In addition, CAMIT learns from and softly memorizes revision actions based on the context, alleviating the issue of repeating mistakes. Experiments in both ideal and real interactive translation settings demonstrate that our proposed CAMIT enhances machine translation results significantly while requires fewer revision instructions from human compared to previous methods. Rongxiang Weng, Hao Zhou 0012, Shujian Huang, Lei Li 0005, Jiajun Chen 0001 |
IJCAI | 3 |
| 2019 | Utilizing Non-Parallel Text for Style Transfer by Making Partial ComparisonsabstractText style transfer aims to rephrase a given sentence into a different style without changing its original content. Since parallel corpora (i.e. sentence pairs with the same content but different styles) are usually unavailable, most previous works solely guide the transfer process with distributional information, i.e. using style-related classifiers or language models, which neglect the correspondence of instances, leading to poor transfer performance, especially for the content preservation. In this paper, we propose making partial comparisons to explicitly model the content and style correspondence of instances, respectively. To train the partial comparators, we propose methods to extract partial-parallel training instances automatically from the non-parallel data, and to further enhance the training process by using data augmentation. We perform experiments that compare our method to other existing approaches on two review datasets. Both automatic and manual evaluations show that our approach can significantly improve the performance of existing adversarial methods, and outperforms most state-of-the-art models. Our code and data will be available on Github. Shujian Huang, Xinyu Dai, Jiajun Chen 0001 |
IJCAI | 2 |
| 2018 | Improving Review Representations With User Attention and Product Attention for Sentiment ClassificationabstractNeural network methods have achieved great success in reviews sentiment classification. Recently, some works achieved improvement by incorporating user and product information to generate a review representation. However, in reviews, we observe that some words or sentences show strong user's preference, and some others tend to indicate product's characteristic. The two kinds of information play different roles in determining the sentiment label of a review. Therefore, it is not reasonable to encode user and product information together into one representation. In this paper, we propose a novel framework to encode user and product information. Firstly, we apply two individual hierarchical neural networks to generate two representations, with user attention or with product attention. Then, we design a combined strategy to make full use of the two representations for training and final prediction. The experimental results show that our model obviously outperforms other state-of-the-art methods on IMDB and Yelp datasets. Through the visualization of attention over words related to user or product, we validate our observation mentioned above. Zhen Wu 0002, Xinyu Dai, Cunyan Yin, Shujian Huang, Jiajun Chen 0001 |
AAAI | 4 |
| 2018 | Unsupervised Bilingual Lexicon Induction via Latent Variable ModelsabstractBilingual lexicon extraction has been studied for decades and most previous methods have relied on parallel corpora or bilingual dictionaries.Recent studies have shown that it is possible to build a bilingual dictionary by aligning monolingual word embedding spaces in an unsupervised way.With the recent advances in generative models, we propose a novel approach which builds cross-lingual dictionaries via latent variable models and adversarial training with no parallel corpora.To demonstrate the effectiveness of our approach, we evaluate our approach on several language pairs and the experimental results show that our model could achieve competitive and even superior performance compared with several state-of-the-art models. Zi-Yi Dou, Zhi-Hao Zhou, Shujian Huang |
EMNLP | 3 |
| 2018 | Dynamic Oracle for Neural Machine Translation in Decoding Phase
Zi-Yi Dou, Hao Zhou 0012, Shujian Huang, Xinyu Dai, Jiajun Chen 0001 |
LREC | 3 |
| 2018 | Combining Character and Word Information in Neural Machine Translation Using a Multi-Level AttentionabstractHuadong Chen, Shujian Huang, David Chiang, Xinyu Dai, Jiajun Chen. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). 2018. Huadong Chen, Shujian Huang, David Chiang 0001, Xinyu Dai, Jiajun Chen 0001 |
NAACL-HLT | 2 |
| 2018 | Improving Aspect Identification with Reviews Segmentation
Tianhao Ning, Zhen Wu 0002, Xinyu Dai, Shujian Huang, Jiajun Chen 0001 |
NLPCC (1) | 5 |
| 2018 | Modeling Past and Future for Neural Machine TranslationabstractExisting neural machine translation systems do not explicitly model what has been translated and what has not during the decoding phase. To address this problem, we propose a novel mechanism that separates the source information into two parts: translated Past contents and untranslated Future contents, which are modeled by two additional recurrent layers. The Past and Future contents are fed to both the attention model and the decoder states, which provides Neural Machine Translation (NMT) systems with the knowledge of translated and untranslated contents. Experimental results show that the proposed approach significantly improves the performance in Chinese-English, German-English, and English-German translation tasks. Specifically, the proposed model outperforms the conventional coverage model in terms of both the translation quality and the alignment error rate. Zaixiang Zheng, Hao Zhou 0012, Shujian Huang, Lili Mou, Xinyu Dai, Jiajun Chen 0001, Zhaopeng Tu |
Trans. Assoc. Comput. Linguistics | 3 |
| 2018 | Collaborative Filtering with Topic and Social Latent Factors Incorporating Implicit FeedbackabstractRecommender systems (RSs) provide an effective way of alleviating the information overload problem by selecting personalized items for different users. Latent factors-based collaborative filtering (CF) has become the popular approaches for RSs due to its accuracy and scalability. Recently, online social networks and user-generated content provide diverse sources for recommendation beyond ratings. Although social matrix factorization (Social MF) and topic matrix factorization (Topic MF) successfully exploit social relations and item reviews, respectively; both of them ignore some useful information. In this article, we investigate the effective data fusion by combining the aforementioned approaches. First, we propose a novel model MR3 to jointly model three sources of information (i.e., ratings, item reviews, and social relations) effectively for rating prediction by aligning the latent factors and hidden topics. Second, we incorporate the implicit feedback from ratings into the proposed model to enhance its capability and to demonstrate its flexibility. We achieve more accurate rating prediction on real-life datasets over various state-of-the-art methods. Furthermore, we measure the contribution from each of the three data sources and the impact of implicit feedback from ratings, followed by the sensitivity analysis of hyperparameters. Empirical studies demonstrate the effectiveness and efficacy of our proposed model and its extension. Guang-Neng Hu, Xinyu Dai, Feng-Yu Qiu, Tao Li 0001, Shujian Huang, Jiajun Chen 0001 |
ACM Trans. Knowl. Discov. Data | 6 |
| 2017 | Improved Neural Machine Translation with a Syntax-Aware Encoder and DecoderabstractMost neural machine translation (NMT) models are based on the sequential encoder-decoder framework, which makes no use of syntactic information.In this paper, we improve this model by explicitly incorporating source-side syntactic trees.More specifically, we propose (1) a bidirectional tree encoder which learns both sequential and tree structured representations; (2) a tree-coverage model that lets the attention depend on the source-side syntax.Experiments on Chinese-English translation demonstrate that our proposed models outperform the sequential attentional model as well as a stronger baseline with a bottom-up tree encoder and word coverage.1 Huadong Chen, Shujian Huang, David Chiang 0001, Jiajun Chen 0001 |
ACL (1) | 2 |
| 2017 | A Multi-view Clustering Model for Event Detection in Twitter
Richard D. Shang, Xinyu Dai, Weiyi Ge, Shujian Huang, Jiajun Chen 0001 |
CICLing (2) | 4 |
| 2017 | Top-Rank Enhanced Listwise Optimization for Statistical Machine TranslationabstractPairwise ranking methods are the basis of many widely used discriminative training approaches for structure prediction problems in natural language processing (NLP).Decomposing the problem of ranking hypotheses into pairwise comparisons enables simple and efficient solutions.However, neglecting the global ordering of the hypothesis list may hinder learning.We propose a listwise learning framework for structure prediction problems such as machine translation.Our framework directly models the entire translation list's ordering to learn parameters which may better fit the given listwise samples.Furthermore, we propose top-rank enhanced loss functions, which are more sensitive to ranking errors at higher positions.Experiments on a large-scale Chinese-English translation task show that both our listwise learning framework and top-rank enhanced listwise losses lead to significant improvements in translation quality. Huadong Chen, Shujian Huang, David Chiang 0001, Xinyu Dai, Jiajun Chen 0001 |
CoNLL | 2 |
| 2017 | Neural Machine Translation with Word PredictionsabstractIn the encoder-decoder architecture for neural machine translation (NMT), the hidden states of the recurrent structures in the encoder and decoder carry the crucial information about the sentence.These vectors are generated by parameters which are updated by back-propagation of translation errors through time.We argue that propagating errors through the end-to-end recurrent structures are not a direct way of control the hidden vectors.In this paper, we propose to use word predictions as a mechanism for direct supervision.More specifically, we require these vectors to be able to predict the vocabulary in target sentence.Our simple mechanism ensures better representations in the encoder and decoder without using any extra data or annotation.It is also helpful in reducing the target side vocabulary and improving the decoding efficiency.Experiments on Chinese-English and German-English machine translation tasks show BLEU improvements by 4.53 and 1.3, respectively. Rongxiang Weng, Shujian Huang, Zaixiang Zheng, Xinyu Dai, Jiajun Chen 0001 |
EMNLP | 2 |
| 2017 | Word-Context Character Embeddings for Chinese Word SegmentationabstractNeural parsers have benefited from automatically labeled data via dependencycontext word embeddings.We investigate training character embeddings on a word-based context in a similar way, showing that the simple method significantly improves state-of-the-art neural word segmentation models, beating tritraining baselines for leveraging autosegmented data. Hao Zhou 0012, Zhenting Yu, Yue Zhang 0004, Shujian Huang, Xinyu Dai, Jiajun Chen 0001 |
EMNLP | 4 |
| 2017 | Compressing Neural Networks by Applying Frequent Item-Set Mining
Zi-Yi Dou, Shujian Huang, Yi-Fan Su |
ICANN (2) | 2 |
| 2017 | Deep Matrix Factorization Models for Recommender SystemsabstractRecommender systems usually make personalized recommendation with user-item interaction ratings, implicit feedback and auxiliary information. Matrix factorization is the basic idea to predict a personalized ranking over a set of items for an individual user with the similarities among users and items. In this paper, we propose a novel matrix factorization model with neural network architecture. Firstly, we construct a user-item matrix with explicit ratings and non-preference implicit feedback. With this matrix as the input, we present a deep structure learning architecture to learn a common low dimensional space for the representations of users and items. Secondly, we design a new loss function based on binary cross entropy, in which we consider both explicit ratings and implicit feedback for a better optimization. The experimental results show the effectiveness of both our proposed model and the loss function. On several benchmark datasets, our model outperformed other state-of-the-art methods. We also conduct extensive experiments to evaluate the performance within different experimental settings. Hong-Jian Xue, Xinyu Dai, Shujian Huang, Jiajun Chen 0001 |
IJCAI | 4 |
| 2017 | AGRA: An Analysis-Generation-Ranking Framework for Automatic Abbreviation from Paper TitlesabstractPeople sometimes choose word-like abbreviations to refer to items with a long description. These abbreviations usually come from the descriptive text of the item and are easy to remember and pronounce, while preserving the key idea of the item. Coming up with a nice abbreviation is not an easy job, even for human. Previous assistant naming systems compose names by applying hand-written rules, which may not perform well. In this paper, we propose to view the naming task as an artificial intelligence problem and create a data set in the domain of academic naming. To generate more delicate names, we propose a three-step framework, including description analysis, candidate generation and abbreviation ranking, each of which is parameterized and optimizable. We conduct experiments to compare different settings of our framework with several analysis approaches from different perspectives. Compared to online or baseline systems, our framework could achieve the best results. Shujian Huang, Cam-Tu Nguyen, Xiaoliang Wang 0001, Xinyu Dai, Jiajun Chen 0001, Yang Yu 0001 |
IJCAI | 3 |
| 2017 | A Neural Probabilistic Structured-Prediction Method for Transition-Based Natural Language ProcessingabstractWe propose a neural probabilistic structured-prediction method for transition-based natural language processing, which integrates beam search and contrastive learning. The method uses a global optimization model, which can leverage arbitrary features over non-local context. Beam search is used for efficient heuristic decoding, and contrastive learning is performed for adjusting the model according to search errors. When evaluated on both chunking and dependency parsing tasks, the proposed method achieves significant accuracy improvements over the locally normalized greedy baseline on the two tasks, respectively. Hao Zhou 0012, Yue Zhang 0004, Chuan Cheng, Shujian Huang, Xinyu Dai, Jiajun Chen 0001 |
J. Artif. Intell. Res. | 4 |
| 2016 | A Search-Based Dynamic Reranking Model for Dependency ParsingabstractWe propose a novel reranking method to extend a deterministic neural dependency parser.Different to conventional k-best reranking, the proposed model integrates search and learning by utilizing a dynamic action revising process, using the reranking model to guide modification for the base outputs and to rerank the candidates.The dynamic reranking model achieves an absolute 1.78% accuracy improvement over the deterministic baseline parser on PTB, which is the highest improvement by neural rerankers in the literature. Hao Zhou 0012, Yue Zhang 0004, Shujian Huang, Junsheng Zhou, Xinyu Dai, Jiajun Chen 0001 |
ACL (1) | 3 |
| 2016 | Tree-State Based Rule Selection Models for Hierarchical Phrase-Based Machine Translation
Shujian Huang, Huifeng Sun, Chengqi Zhao, Jinsong Su, Xinyu Dai, Jiajun Chen 0001 |
IJCAI | 1 |
| 2016 | Tagging Chinese microblogger via sparse feature selectionabstractIn new media era, users post messages to record their daily lives and express their opinions via social media platforms, such as microblog. Recently, it is an attractive topic to tag users from the users generation contents. Tags for a microblog user, as the description for his/her interests, concerns or occupational characteristics, are playing an important role in user indexing, personalized recommendation, and so on. Previous works apply keyword extraction methods to present the interests of users. However, it is hard for keyword extraction to give accurate results when the data is deficient and noisy. In this paper, we propose a novel method to tag the users. Firstly, we apply feature selection via sparse classifier to generate preliminary tags for users. Then we also apply feature selection method to extend the tags. Finally, we refine the tags with a reranking strategy. We conduct our experiments on the data of the most popular Chinese microblog (Sina Weibo). The experimental results show that our method improves the performance significantly over other methods. Richard D. Shang, Xinyu Dai, Shujian Huang, Jiajun Chen 0001 |
IJCNN | 3 |
| 2016 | Evaluating a Deterministic Shift-Reduce Neural Parser for Constituent Parsing
Hao Zhou 0012, Yue Zhang 0004, Shujian Huang, Xinyu Dai, Jiajun Chen 0001 |
LREC | 3 |
| 2016 | PRIMT: A Pick-Revise Framework for Interactive Machine TranslationabstractShanbo Cheng, Shujian Huang, Huadong Chen, Xin-Yu Dai, Jiajun Chen. Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2016. Shanbo Cheng, Shujian Huang, Huadong Chen, Xinyu Dai, Jiajun Chen 0001 |
HLT-NAACL | 2 |
| 2016 | Adaptation of Language Models for SMT Using Neural Networks with Topic InformationabstractNeural network language models (LMs) are shown to be effective in improving the performance of statistical machine translation (SMT) systems. However, state-of-the-art neural network LMs usually use words before the current position as context and neglect global topic information, which can help machine translation (MT) systems to select better translation candidates from a higher perspective. In this work, we propose improvement of the state-of-the-art feedforward neural language model with topic information. Two main issues need to be tackled when adding topics into neural network LMs for SMT: one is how to incorporate topics to the neural network; the other is how to get target-side topic distribution before translation. We incorporate topics by appending topic distribution to the input layer of a feedforward LM. We adopt a multinomial logistic-regression (MLR) model to predict the target-side topic distribution based on source side information. Moreover, we propose a feedforward neural network model to learn joint representations on the source side for topic prediction. LM experiments demonstrate that the perplexity on validation set can be greatly reduced by the topic-enhanced feedforward LM, and the prediction of target-side topics can be improved dramatically with the MLR model equipped with the joint source representations. A final MT experiment, conducted on a large-scale Chinese--English dataset, shows that our feedforward LM with predicted topics improves the translation performance against a strong baseline. Yinggong Zhao, Shujian Huang, Xinyu Dai, Jiajun Chen 0001 |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 2 |
| 2016 | Enhancing Shift-Reduce Constituent Parsing with Action N-Gram ModelabstractCurrent shift-reduce parsers “understand” the context by embodying a large number of binary indicator features with a discriminative model. In this article, we propose the action n-gram model, which utilizes the action sequence to help parsing disambiguation. The action n-gram model is trained on action sequences produced by parsers with the n-gram estimation method, which gives a smoothed maximum likelihood estimation of the action probability given a specific action history. We show that incorporating action n-gram models into a state-of-the-art parsing framework could achieve parsing accuracy improvements on three datasets across two languages. Hao Zhou 0012, Shujian Huang, Junsheng Zhou, Yue Zhang 0004, Huadong Chen, Xinyu Dai, Chuan Cheng, Jiajun Chen 0001 |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 2 |
| 2015 | Structured Sparsity with Group-Graph RegularizationabstractIn many learning tasks with structural properties, structural sparsity methods help induce sparse models, usually leading to better interpretability and higher generalization performance. One popular approach is to use group sparsity regularization that enforces sparsity on the clustered groups of features, while another popular approach is to adopt graph sparsity regularization that considers sparsity on the link structure of graph embedded features. Both the group and graph structural properties co-exist in many applications. However, group sparsity and graph sparsity have not been considered simultaneously yet. In this paper, we propose a g2-regularization that takes group and graph sparsity into joint consideration, and present an effective approach for its optimization. Experiments on both synthetic and real data show that, enforcing group-graph sparsity lead to better performance than using group sparsity or graph sparsity only. Xinyu Dai, Shujian Huang, Jiajun Chen 0001, Zhi-Hua Zhou |
AAAI | 3 |
| 2015 | Non-linear Learning for Statistical Machine TranslationabstractShujian Huang, Huadong Chen, Xin-Yu Dai, Jiajun Chen. Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2015. Shujian Huang, Huadong Chen, Xinyu Dai, Jiajun Chen 0001 |
ACL (1) | 1 |
| 2015 | A Neural Probabilistic Structured-Prediction Model for Transition-Based Dependency ParsingabstractHao Zhou, Yue Zhang, Shujian Huang, Jiajun Chen. Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2015. Hao Zhou 0012, Yue Zhang 0004, Shujian Huang, Jiajun Chen 0001 |
ACL (1) | 3 |
| 2015 | A Unified Framework for Jointly Learning Distributed Representations of Word and Attributes
Liqiang Niu, Xinyu Dai, Shujian Huang, Jiajun Chen 0001 |
ACML | 3 |
| 2015 | Sentiment Classification with Graph Sparsity Regularization
Xinyu Dai, Chuan Cheng, Shujian Huang, Jiajun Chen 0001 |
CICLing (2) | 3 |
| 2015 | Graph-Based Collective Lexical Selection for Statistical Machine TranslationabstractLexical selection is of great importance to statistical machine translation. In this paper, we propose a graph-based frame-work for collective lexical selection. The framework is established on a translation graph that captures not only local associ-ations between source-side content words and their target translations but also target-side global dependencies in terms of relat-edness among target items. We also in-troduce a random walk style algorithm to collectively identify translations of source-side content words that are strongly related in translation graph. We validate the ef-fectiveness of our lexical selection frame-work on Chinese-English translation. Ex-periment results with large-scale training data show that our approach significantly improves lexical selection. 1 Jinsong Su, Deyi Xiong, Shujian Huang, Xianpei Han, Junfeng Yao |
EMNLP | 3 |
| 2015 | A Synthetic Approach for Recommendation: Combining Ratings, Social Relations, and Reviews
Guang-Neng Hu, Xinyu Dai, Yunya Song, Shujian Huang, Jiajun Chen 0001 |
IJCAI | 4 |
| 2015 | Word Segmentation of Micro Blogs with BaggingabstractThis paper describes the model we designed for the Chinese word segmentation Task of NLPCC 2015. We firstly apply a word-based perceptron algorithm to build the base segmenter. Then, we use a Bootstrap Aggregating model of bagging which improves the segmentation results consistently on the three tracks of closed, semi-open and open test. Considering the characteristics of Weibo text, we also perform rule-based adaptation before decoding. Finally, our model achieves F-score 95.12% on closed track, 95.3% on semi-open track and 96.09% on open track. Zhenting Yu, Xinyu Dai, Si Shen, Shujian Huang, Jiajun Chen 0001 |
NLPCC | 4 |
| 2015 | Resolving Coordinate Structures for Chinese Constituent ParsingabstractCoordinate structures are linguistic structures consisting of two or more conjuncts, which usually compose into larger constituent as a whole unit. However, the boundary of each conjunct is difficult to identify, which makes it difficult to parse the whole coordinate and larger structures. In labeled data, such as the Penn Chinese Tree Bank (CTB), coordinate structures are not labeled explicitly, which makes solving the problem more complicated. In this paper, we treat resolving coordinate structures as an independent sub-problem of parsing. We first define coordinate structures explicitly and design rules to extract the coordinate structures from labeled CTB data. Then a specifically designed grammar is proposed for automatic parsing of coordinate structures. We propose two groups of new features to better model coordinate structures in a shift-reduce parsing framework. Our approach can achieve a $$15\%$$ improvement in F-1 score on resolving coordinate structures. Yichu Zhou, Shujian Huang, Xinyu Dai, Jiajun Chen 0001 |
NLPCC | 2 |
| 2013 | Forgetting Word Segmentation in Chinese Text Classification with L1-Regularized Logistic Regression
Xinyu Dai, Shujian Huang, Jiajun Chen 0001 |
PAKDD (2) | 3 |
| 2011 | Language Model Weight Adaptation Based on Cross-entropy for Statistical Machine Translation
Yinggong Zhao, Yangsheng Ji, Ning Xi 0003, Shujian Huang, Jiajun Chen 0001 |
PACLIC | 4 |
| 2010 | Improving Word Alignment by Semi-Supervised Ensemble
Shujian Huang, Kangxi Li, Xinyu Dai, Jiajun Chen 0001 |
CoNLL | 1 |