VLDB 2026 Research / reviewers in the wild / expert
Haipeng Sun
dblp:155/8355
· DBLP profile ↗
14ranked-venue papers
9as first author
9since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 11 · 8 first-author · 7 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SCC3: a novel structure-connected cognition cube network for Manchu word recognition
Xiaojun Bi 0002, Wenhao Tao, Haipeng Sun |
Expert Syst. Appl. | 4 |
| 2025 | Comet: Dialog Context Fusion Mechanism for End-to-End Task-Oriented Dialog with Multi-task LearningabstractExisting end-to-end task-oriented dialog systems often encounter challenges arising from implicit information, coreference, and the presence of noisy and irrelevant data within the dialog context. These issues hinder the system’s ability to fully comprehend critical information and lead to inaccurate responses. To address these concerns, we propose Comet, a dialog context fusion mechanism for end-to-end task-oriented dialog, augmented with three supplementary tasks: dialog summarization, domain prediction, and slot detection. Dialog summarization facilitates a more comprehensive understanding of important dialog context information by Comet. Domain prediction enables Comet to concentrate on domain-specific information, thus reducing interference from irrelevant information. Slot detection empowers Comet to accurately identify and comprehend essential dialog context information. Additionally, we introduce a data refinement strategy to enhance the comprehensiveness and recommendability of the generated responses. Experimental results demonstrate the superior performance of our proposed methods compared to existing end-to-end task-oriented dialog systems, achieving state-of-the-art results on the MultiWOZ and CrossWOZ datasets. Haipeng Sun, Junwei Bao 0001, Youzheng Wu, Xiaodong He 0001 |
COLING | 1 |
| 2023 | Deep reinforce learning for joint optimization of condition-based maintenance and spare ordering
Shen-Gang Hao, Jun Zheng 0007, Haipeng Sun, Quanxin Zhang 0001, Li Zhang 0099, Nan Jiang 0021, Yuanzhang Li 0001 |
Inf. Sci. | 4 |
| 2023 | Improving the invisibility of adversarial examples with perceptually adaptive perturbation
Yu-an Tan 0001, Haipeng Sun, Yuhang Zhao 0003, Quanxin Zhang 0001, Yuanzhang Li 0001 |
Inf. Sci. | 3 |
| 2022 | OPERA: Operation-Pivoted Discrete Reasoning over TextabstractYongwei Zhou, Junwei Bao, Chaoqun Duan, Haipeng Sun, Jiahui Liang, Yifan Wang, Jing Zhao, Youzheng Wu, Xiaodong He, Tiejun Zhao. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Yongwei Zhou, Junwei Bao 0001, Chaoqun Duan, Haipeng Sun, Jiahui Liang, Yifan Wang 0016, Youzheng Wu, Xiaodong He 0001, Tiejun Zhao |
NAACL-HLT | 4 |
| 2022 | A fine-grained and traceable multidomain secure data-sharing model for intelligent terminals in edge-cloud collaboration scenariosabstractSecure data-sharing technology is a bridge for various collaborative operations among intelligent terminals in the edge-cloud collaborative application scenario. For the shared data involves different levels of confidentiality, intelligent terminals for collaborative operations may be distributed in multiple management domains, and the private information of intelligent terminals is easy to be leaked in edge-cloud collaboration scenarios, the security of data sharing is severely threatened. To solve these problems, this paper proposed a fine-grained and traceable multidomain secure data-sharing model for intelligent terminals. In this model, a key self-certification algorithm is proposed, which avoids potential security threats of key leakage during the key distribution process. The model combines attribute encryption and threshold function to achieve more fine-grained and more flexible secure data sharing; it uses blockchain technology to achieve integrity verification of stored data and traceability of shared data, and it combines on-chain and off-chain databases to achieve rapid retrieval and positioning of shared data distributed among multiple domains, which improves the efficiency of data sharing among domains. The security of the model proposed by us is proved, and compared with the cited literature, it is shown that the proposed model has certain advantages in terms of computational complexity and time consumption. Haipeng Sun, Yu-an Tan 0001, Qikun Zhang, Yuanzhang Li 0001, Shangbo Wu |
Int. J. Intell. Syst. | 1 |
| 2021 | Self-Training for Unsupervised Neural Machine Translation in Unbalanced Training Data ScenariosabstractHaipeng Sun, Rui Wang, Kehai Chen, Masao Utiyama, Eiichiro Sumita, Tiejun Zhao. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Haipeng Sun, Rui Wang 0015, Kehai Chen, Masao Utiyama, Eiichiro Sumita, Tiejun Zhao |
NAACL-HLT | 1 |
| 2021 | EviDR: Evidence-Emphasized Discrete Reasoning for Reasoning Machine Reading Comprehension
Yongwei Zhou, Junwei Bao 0001, Haipeng Sun, Jiahui Liang, Youzheng Wu, Xiaodong He 0001, Bowen Zhou 0001, Tiejun Zhao |
NLPCC (1) | 3 |
| 2021 | Unsupervised Neural Machine Translation for Similar and Distant Language Pairs: An Empirical StudyabstractUnsupervised neural machine translation (UNMT) has achieved remarkable results for several language pairs, such as French–English and German–English. Most previous studies have focused on modeling UNMT systems; few studies have investigated the effect of UNMT on specific languages. In this article, we first empirically investigate UNMT for four diverse language pairs (French/German/Chinese/Japanese–English). We confirm that the performance of UNMT in translation tasks for similar language pairs (French/German–English) is dramatically better than for distant language pairs (Chinese/Japanese–English). We empirically show that the lack of shared words and different word orderings are the main reasons that lead UNMT to underperform in Chinese/Japanese–English. Based on these findings, we propose several methods, including artificial shared words and pre-ordering, to improve the performance of UNMT for distant language pairs. Moreover, we propose a simple general method to improve translation performance for all these four language pairs. The existing UNMT model can generate a translation of a reasonable quality after a few training epochs owing to a denoising mechanism and shared latent representations. However, learning shared latent representations restricts the performance of translation in both directions, particularly for distant language pairs, while denoising dramatically delays convergence by continuously modifying the training data. To avoid these problems, we propose a simple, yet effective and efficient, approach that (like UNMT) relies solely on monolingual corpora: pseudo-data-based unsupervised neural machine translation. Experimental results for these four language pairs show that our proposed methods significantly outperform UNMT baselines. Haipeng Sun, Rui Wang 0015, Masao Utiyama, Benjamin Marie, Kehai Chen, Eiichiro Sumita, Tiejun Zhao |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 1 |
| 2020 | Knowledge Distillation for Multilingual Unsupervised Neural Machine TranslationabstractUnsupervised neural machine translation (UNMT) has recently achieved remarkable results for several language pairs. However, it can only translate between a single language pair and cannot produce translation results for multiple language pairs at the same time. That is, research on multilingual UNMT has been limited. In this paper, we empirically introduce a simple method to translate between thirteen languages using a single encoder and a single decoder, making use of multilingual data to improve UNMT for all language pairs. On the basis of the empirical findings, we propose two knowledge distillation methods to further enhance multilingual UNMT performance. Our experiments on a dataset with English translated to and from twelve other languages (including three language families and six language branches) show remarkable results, surpassing strong unsupervised individual baselines while achieving promising performance between non-English language pairs in zero-shot translation scenarios and alleviating poor performance in low-resource language pairs. Haipeng Sun, Rui Wang 0015, Kehai Chen, Masao Utiyama, Eiichiro Sumita, Tiejun Zhao |
ACL | 1 |
| 2020 | Robust Unsupervised Neural Machine Translation with Adversarial Denoising TrainingabstractUnsupervised neural machine translation (UNMT) has recently attracted great interest in the machine translation community.The main advantage of the UNMT lies in its easy collection of required large training text sentences while with only a slightly worse performance than supervised neural machine translation which requires expensive annotated translation pairs on some translation tasks.In most studies, the UMNT is trained with clean data without considering its robustness to the noisy data.However, in real-world scenarios, there usually exists noise in the collected input sentences which degrades the performance of the translation system since the UNMT is sensitive to the small perturbations of the input sentences.In this paper, we first time explicitly take the noisy data into consideration to improve the robustness of the UNMT based systems.First of all, we clearly defined two types of noises in training sentences, i.e., word noise and word order noise, and empirically investigate its effect in the UNMT, then we propose adversarial training methods with denoising process in the UNMT.Experimental results on several language pairs show that our proposed methods substantially improved the robustness of the conventional UNMT systems in noisy scenarios. Haipeng Sun, Rui Wang 0015, Kehai Chen, Xugang Lu, Masao Utiyama, Eiichiro Sumita, Tiejun Zhao |
COLING | 1 |
| 2020 | Unsupervised Neural Machine Translation With Cross-Lingual Language Representation AgreementabstractUnsupervised cross-lingual language representation initialization methods such as unsupervised bilingual word embedding (UBWE) pre-training and cross-lingual masked language model (CMLM) pre-training, together with mechanisms such as denoising and back-translation, have advanced unsupervised neural machine translation (UNMT), which has achieved impressive results on several language pairs, particularly French-English and German-English. Typically, UBWE focuses on initializing the word embedding layer in the encoder and decoder of UNMT, whereas the CMLM focuses on initializing the entire encoder and decoder of UNMT. However, UBWE/CMLM training and UNMT training are independent, which makes it difficult to assess how the quality of UBWE/CMLM affects the performance of UNMT during UNMT training. In this paper, we first empirically explore relationships between UNMT and UBWE/CMLM. The empirical results demonstrate that the performance of UBWE and CMLM has a significant influence on the performance of UNMT. Motivated by this, we propose a novel UNMT structure with cross-lingual language representation agreement to capture the interaction between UBWE/CMLM and UNMT during UNMT training. Experimental results on several language pairs demonstrate that the proposed UNMT models improve significantly over the corresponding state-of-the-art UNMT baselines. Haipeng Sun, Rui Wang 0015, Kehai Chen, Masao Utiyama, Eiichiro Sumita, Tiejun Zhao |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2019 | Unsupervised Bilingual Word Embedding Agreement for Unsupervised Neural Machine TranslationabstractUnsupervised bilingual word embedding (UBWE), together with other technologies such as back-translation and denoising, has helped unsupervised neural machine translation (UNMT) achieve remarkable results in several language pairs.In previous methods, UBWE is first trained using nonparallel monolingual corpora and then this pre-trained UBWE is used to initialize the word embedding in the encoder and decoder of UNMT.That is, the training of UBWE and UNMT are separate.In this paper, we first empirically investigate the relationship between UBWE and UNMT.The empirical findings show that the performance of UNMT is significantly affected by the performance of UBWE.Thus, we propose two methods that train UNMT with UBWE agreement.Empirical results on several language pairs show that the proposed methods significantly outperform conventional UNMT. Haipeng Sun, Rui Wang 0015, Kehai Chen, Masao Utiyama, Eiichiro Sumita, Tiejun Zhao |
ACL (1) | 1 |
| 2010 | Super Baozi vs Sushi ManabstractSuper Baozi, being tired of being in Catering, is longing for developing in Recreation. With intense enthusiasm and strong perseverance, he has learned to sing songs and to play nunchakus. In this video, Super baozi is fighting with Sushi man, in which most fighting moves were adopted from Bruce Lee's move. Here, Super Baozi wants to take this opportunity to pay a public tribute to Bruce Lee. Haipeng Sun |
SIGGRAPH ASIA (Computer Animation Festival) | 1 |