Chulun Zhou

dblp:246/2903 · DBLP profile ↗
← Back
17ranked-venue papers
5as first author
14since 2021 · last 2025
0000-0002-1544-817XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 16 · 5 first-author · 13 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1
YearPublicationVenuePosition
2025 The Essence of Contextual Understanding in Theory of Mind: A Study on Question Answering with Story Characters
abstract
Theory-of-Mind (ToM) is a fundamental psychological capability that allows humans to understand and interpret the mental states of others. Humans infer others’ thoughts by integrating causal cues and indirect clues from broad contextual information, often derived from past interactions. In other words, human ToM heavily relies on the understanding about the backgrounds and life stories of others. Unfortunately, this aspect is largely overlooked in existing benchmarks for evaluating machines’ ToM capabilities, due to their usage of short narratives without global context, especially personal background of characters. In this paper, we verify the importance of comprehensive contextual understanding about personal backgrounds in ToM and assess the performance of LLMs in such complex scenarios. To achieve this, we introduce CharToM-QA benchmark, comprising 1,035 ToM questions based on characters from classic novels. Our human study reveals a significant disparity in performance: the same group of educated participants performs dramatically better when they have read the novels compared to when they have not. In parallel, our experiments on state-of-the-art LLMs, including the very recent o1 and DeepSeek-R1 models, show that LLMs still perform notably worse than humans, despite that they have seen these stories during pre-training. This highlights the limitations of current LLMs in capturing the nuanced contextual information required for ToM reasoning.
Chulun Zhou, Qiujing Wang, Mo Yu, Xiaoqian Yue, Shunchi Zhang, Jie Zhou 0016, Wai Lam
ACL (1)1
2023 HyperNetwork-based Decoupling to Improve Model Generalization for Few-Shot Relation Extraction
abstract
Few-shot relation extraction (FSRE) aims to train a model that can deal with new relations using only a few labeled examples.Most existing studies employ Prototypical Networks for FSRE, which usually overfits the relation classes in the training set and cannot generalize well to unseen relations.By investigating the class separation of an FSRE model, we find that model upper layers are prone to learn relation-specific knowledge.Therefore, in this paper, we propose a HyperNetworkbased Decoupling approach to improve the generalization of FSRE models.Specifically, our model consists of an encoder, a network generator (for producing relation classifiers) and the generated-then-finetuned classifiers for every N -way-K-shot episode.Meanwhile, we design a two-step training strategy along with a class-agnostic aligner, by which the generated classifiers focus on acquiring relation-specific knowledge and the encoder is encouraged to learn more general relation knowledge.In this way, the roles of upper and lower layers in our FSRE model are explicitly decoupled, thus enhancing its generalizing capability during testing.Experiments on two public datasets demonstrate the effectiveness of our method.Our source code is available at https: //github.com/DeepLearnXMU/FSRE-HDN.
Chulun Zhou, Fandong Meng, Jinsong Su, Yidong Chen 0001, Jie Zhou 0016
EMNLP2
2023 Multi-modal graph contrastive encoding for neural machine translation
Yongjing Yin, Jiali Zeng, Jinsong Su, Chulun Zhou, Fandong Meng, Jie Zhou 0016, Degen Huang, Jiebo Luo 0001
Artif. Intell.4
2023 A Multi-Task Multi-Stage Transitional Training Framework for Neural Chat Translation
abstract
Neural chat translation (NCT) aims to translate a cross-lingual chat between speakers of different languages. Existing context-aware NMT models cannot achieve satisfactory performances due to the following inherent problems: 1) limited resources of annotated bilingual dialogues; 2) the neglect of modelling conversational properties; 3) training discrepancy between different stages. To address these issues, in this paper, we propose a multi-task multi-stage transitional (MMT) training framework, where an NCT model is trained using the bilingual chat translation dataset and additional monolingual dialogues. We elaborately design two auxiliary tasks, namely utterance discrimination and speaker discrimination, to introduce the modelling of dialogue coherence and speaker characteristic into the NCT model. The training process consists of three stages: 1) sentence-level pre-training on large-scale parallel corpus; 2) intermediate training with auxiliary tasks using additional monolingual dialogues; 3) context-aware fine-tuning with gradual transition. Particularly, the second stage serves as an intermediate phase that alleviates the training discrepancy between the pre-training and fine-tuning stages. Moreover, to make the stage transition smoother, we train the NCT model using a gradual transition strategy, i.e., gradually transiting from using monolingual to bilingual dialogues. Extensive experiments on two language pairs demonstrate the effectiveness and superiority of our proposed training framework.
Chulun Zhou, Yunlong Liang, Fandong Meng, Jie Zhou 0016, Jin An Xu, Min Zhang 0005, Jinsong Su
IEEE Trans. Pattern Anal. Mach. Intell.1
2022 A Variational Hierarchical Model for Neural Cross-Lingual Summarization
abstract
Yunlong Liang, Fandong Meng, Chulun Zhou, Jinan Xu, Yufeng Chen, Jinsong Su, Jie Zhou. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022.
Yunlong Liang, Fandong Meng, Chulun Zhou, Jin An Xu, Yufeng Chen 0005, Jinsong Su, Jie Zhou 0016
ACL (1)3
2022 Confidence Based Bidirectional Global Context Aware Training Framework for Neural Machine Translation
abstract
Most dominant neural machine translation (NMT) models are restricted to make predictions only according to the local context of preceding words in a left-to-right manner.Although many previous studies try to incorporate global information into NMT models, there still exist limitations on how to effectively exploit bidirectional global context.In this paper, we propose a Confidence Based Bidirectional Global Context Aware (CB-BGCA) training framework for NMT, where the NMT model is jointly trained with an auxiliary conditional masked language model (CMLM).The training consists of two stages:(1) multi-task joint training; (2) confidence based knowledge distillation.At the first stage, by sharing encoder parameters, the NMT model is additionally supervised by the signal from the CMLM decoder that contains bidirectional global contexts.Moreover, at the second stage, using the CMLM as teacher, we further pertinently incorporate bidirectional global context to the NMT model on its unconfidently-predicted target words via knowledge distillation.Experimental results show that our proposed CB-BGCA training framework significantly improves the NMT model by +1.02, +1.30 and +0.57BLEU scores on three large-scale translation datasets, namely WMT'14 Englishto-German, WMT'19 Chinese-to-English and WMT'14 English-to-French, respectively.
Chulun Zhou, Fandong Meng, Jie Zhou 0016, Min Zhang 0005, Jinsong Su
ACL (1)1
2022 Towards Robust Neural Machine Translation with Iterative Scheduled Data-Switch Training
abstract
Most existing methods on robust neural machine translation (NMT) construct adversarial examples by injecting noise into authentic examples and indiscriminately exploit two types of examples. They require the model to translate both the authentic source sentence and its adversarial counterpart into the identical target sentence within the same training stage, which may be a suboptimal choice to achieve robust NMT. In this paper, we first conduct a preliminary study to confirm this claim and further propose an Iterative Scheduled Data-switch Training Framework to mitigate this problem. Specifically, we introduce two training stages, iteratively switching between authentic and adversarial examples. Compared with previous studies, our model focuses more on just one type of examples at each single stage, which can better exploit authentic and adversarial examples, and thus obtaining a better robust NMT model. Moreover, we introduce an improved curriculum learning method with a sampling strategy to better schedule the process of noise injection. Experimental results show that our model significantly surpasses several competitive baselines on four translation benchmarks. Our source code is available at https://github.com/DeepLearnXMU/RobustNMT-ISDST.
Zhongjian Miao, Xiang Li 0104, Liyan Kang, Wen Zhang 0015, Chulun Zhou, Yidong Chen 0001, Bin Wang 0004, Min Zhang 0005, Jinsong Su
COLING5
2022 Towards Robust k-Nearest-Neighbor Machine Translation
abstract
k-Nearest-Neighbor Machine Translation (kNN-MT) becomes an important research direction of NMT in recent years.Its main idea is to retrieve useful key-value pairs from an additional datastore to modify translations without updating the NMT model.However, the underlying retrieved noisy pairs will dramatically deteriorate the model performance.In this paper, we conduct a preliminary study and find that this problem results from not fully exploiting the prediction of the NMT model.To alleviate the impact of noise, we propose a confidence-enhanced kNN-MT model with robust training.Concretely, we introduce the NMT confidence to refine the modeling of two important components of kNN-MT: kNN distribution and the interpolation weight.Meanwhile we inject two types of perturbations into the retrieved pairs for robust training.Experimental results on four benchmark datasets demonstrate that our model not only achieves significant improvements over current kNN-MT models, but also exhibits better robustness.Our code is available at https://github.com/ DeepLearnXMU/Robust-knn-mt.
Ziyao Lu, Fandong Meng, Chulun Zhou, Jie Zhou 0016, Degen Huang, Jinsong Su
EMNLP4
2022 A novel multi-domain machine reading comprehension model with domain interference mitigation
Chulun Zhou, Shaojie He, Jinsong Su
Neurocomputing1
2022 Exploring Multi-Stage Information Interactions for Multi-Source Neural Machine Translation
abstract
Existing studies for multi-source neural machine translation (NMT) either separately model different source sentences or resort to the conventional single-source NMT by simply concatenating all source sentences. However, there exist two drawbacks in these approaches. First, they ignore the explicit word-level semantic interactions between source sentences, which have been shown effective in the embeddings of multilingual texts. Second, multiple source sentences are simultaneously encoded by an NMT model, which is unable to fully exploit the semantic information of each source sentence. In this paper, we explore multi-stage information interactions for multi-source NMT. Specifically, we first propose a multi-source NMT model that performs information interactions at the encoding stage. Its encoder contains multiple semantic interaction layers, each of which sequentially consists of (1) monolingual semantic interaction sub-layer, which is based on the self-attention mechanism and used to learn word-level monolingual contextual representations of source sentences, and (2) cross-lingual semantic interaction sub-layer, which leverages word alignments to perform fine-grained semantic transitions among hidden states of different source sentences. Furthermore, at the training stage, we introduce a mutual distillation based training framework, where single-source models and ours perform information interactions. Such framework can fully exploit the semantic information of each source sentence to enhance our model. Extensive experimental results on the WMT14 English-German-French dataset show our method exhibits significant improvements upon competitive baselines.
Ziyao Lu, Xiang Li 0104, Yang Liu 0005, Chulun Zhou, Jianwei Cui 0002, Bin Wang 0004, Min Zhang 0005, Jinsong Su
IEEE ACM Trans. Audio Speech Lang. Process.4
2021 Exploring Dynamic Selection of Branch Expansion Orders for Code Generation
abstract
Hui Jiang, Chulun Zhou, Fandong Meng, Biao Zhang, Jie Zhou, Degen Huang, Qingqiang Wu, Jinsong Su. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Chulun Zhou, Fandong Meng, Biao Zhang 0002, Jie Zhou 0016, Degen Huang, Qingqiang Wu 0001, Jinsong Su
ACL/IJCNLP (1)2
2021 Towards Making the Most of Dialogue Characteristics for Neural Chat Translation
abstract
Neural Chat Translation (NCT) aims to translate conversational text between speakers of different languages.Despite the promising performance of sentence-level and context-aware neural machine translation models, there still remain limitations in current NCT models because the inherent dialogue characteristics of chat, such as dialogue coherence and speaker personality, are neglected.In this paper, we propose to promote the chat translation by introducing the modeling of dialogue characteristics into the NCT model.To this end, we design four auxiliary tasks including monolingual response generation, cross-lingual response generation, next utterance discrimination, and speaker identification.Together with the main chat translation task, we optimize the NCT model through the training objectives of all these tasks.By this means, the NCT model can be enhanced by capturing the inherent dialogue characteristics, thus generating more coherent and speaker-relevant translations.Comprehensive experiments on four language directions (English⇔German and English⇔Chinese) verify the effectiveness and superiority of the proposed approach.
Yunlong Liang, Chulun Zhou, Fandong Meng, Jin An Xu, Yufeng Chen 0005, Jinsong Su, Jie Zhou 0016
EMNLP (1)2
2021 Multi-modal neural machine translation with deep semantic interactions
Jinsong Su, Jinchang Chen, Chulun Zhou, Yubin Ge, Qingqiang Wu 0001, Yongxuan Lai
Inf. Sci.4
2021 An External Knowledge Enhanced Graph-based Neural Network for Sentence Ordering
abstract
As an important text coherence modeling task, sentence ordering aims to coherently organize a given set of unordered sentences. To achieve this goal, the most important step is to effectively capture and exploit global dependencies among these sentences. In this paper, we propose a novel and flexible external knowledge enhanced graph-based neural network for sentence ordering. Specifically, we first represent the input sentences as a graph, where various kinds of relations (i.e., entity-entity, sentence-sentence and entity-sentence) are exploited to make the graph representation more expressive and less noisy. Then, we introduce graph recurrent network to learn semantic representations of the sentences. To demonstrate the effectiveness of our model, we conduct experiments on several benchmark datasets. The experimental results and in-depth analysis show our model significantly outperforms the existing state-of-the-art models.
Yongjing Yin, Shaopeng Lai, Linfeng Song, Chulun Zhou, Xianpei Han, Junfeng Yao, Jinsong Su
J. Artif. Intell. Res.4
2020 A Novel Graph-based Multi-modal Fusion Encoder for Neural Machine Translation
abstract
Multi-modal neural machine translation (NMT) aims to translate source sentences into a target language paired with images. However, dominant multi-modal NMT models do not fully exploit fine-grained semantic correspondences between semantic units of different modalities, which have potential to refine multi-modal representation learning. To deal with this issue, in this paper, we propose a novel graph-based multi-modal fusion encoder for NMT. Specifically, we first represent the input sentence and image using a unified multi-modal graph, which captures various semantic relationships between multi-modal semantic units (words and visual objects). We then stack multiple graph-based multi-modal fusion layers that iteratively perform semantic interactions to learn node representations. Finally, these representations provide an attention-based context vector for the decoder. We evaluate our proposed encoder on the Multi30K datasets. Experimental results and in-depth analysis show the superiority of our multi-modal NMT model.
Yongjing Yin, Fandong Meng, Jinsong Su, Chulun Zhou, Zhengyuan Yang, Jie Zhou 0016, Jiebo Luo 0001
ACL4
2020 Exploring Contextual Word-level Style Relevance for Unsupervised Style Transfer
abstract
Unsupervised style transfer aims to change the style of an input sentence while preserving its original content without using parallel training data. In current dominant approaches, owing to the lack of fine-grained control on the influence from the target style, they are unable to yield desirable output sentences. In this paper, we propose a novel attentional sequence-to-sequence (Seq2seq) model that dynamically exploits the relevance of each output word to the target style for unsupervised style transfer. Specifically, we first pretrain a style classifier, where the relevance of each input word to the original style can be quantified via layer-wise relevance propagation. In a denoising auto-encoding manner, we train an attentional Seq2seq model to reconstruct input sentences and repredict word-level previously-quantified style relevance simultaneously. In this way, this model is endowed with the ability to automatically predict the style relevance of each output word. Then, we equip the decoder of this model with a neural style component to exploit the predicted wordlevel style relevance for better style transfer. Particularly, we fine-tune this model using a carefully-designed objective function involving style transfer, style relevance consistency, content preservation and fluency modeling loss terms. Experimental results show that our proposed model achieves state-of-the-art performance in terms of both transfer accuracy and content preservation.
Chulun Zhou, Xinyan Xiao, Jinsong Su, Hua Wu 0003
ACL1
2019 Graph-based Neural Sentence Ordering
abstract
Sentence ordering is to restore the original paragraph from a set of sentences. It involves capturing global dependencies among sentences regardless of their input order. In this paper, we propose a novel and flexible graph-based neural sentence ordering model, which adopts graph recurrent network \citep{Zhang:acl18} to accurately learn semantic representations of the sentences. Instead of assuming connections between all pairs of input sentences, we use entities that are shared among multiple sentences to make more expressive graph representations with less noise. Experimental results show that our proposed model outperforms the existing state-of-the-art systems on several benchmark datasets, demonstrating the effectiveness of our model. We also conduct a thorough analysis on how entities help the performance. Our code is available at https://github.com/DeepLearnXMU/NSEG.git.
Yongjing Yin, Linfeng Song, Jinsong Su, Jiali Zeng, Chulun Zhou, Jiebo Luo 0001
IJCAI5