Qiuhui Liu

dblp:211/9473 · DBLP profile ↗
← Back
9ranked-venue papers
0as first author
4since 2021 · last 2025
0000-0002-6936-4107ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 9 · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
5 papers
Deep learning architectures and training · 47% Machine translation · 36% Optimization for machine learning · 9%

Topics — the 9 heaviest of 9, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Machine translation
neural machine translation
1.432021
Multi-Head Highly Parallelized LSTM Decoder for Neural Machine Translation · ACL/IJCNLP (1) 2021
Lipschitz Constrained Parameter Initialization for Deep Transformers · ACL 2020
Learning Source Phrase Representations for Neural Machine Translation · ACL 2020
Machine learning › Deep learning architectures and training
transformer
0.622020
Lipschitz Constrained Parameter Initialization for Deep Transformers · ACL 2020
Efficient Context-Aware Neural Machine Translation with Layer-Wise Weighting and Input-Aware Gating · IJCAI 2020
Machine learning › Deep learning architectures and training
recurrent neural network
0.512021
Multi-Head Highly Parallelized LSTM Decoder for Neural Machine Translation · ACL/IJCNLP (1) 2021
Machine learning › Optimization for machine learning › hyperparameter optimization
batch size selection
0.412020
Dynamically Adjusting Transformer Batch Size by Monitoring Gradient Direction Change · ACL 2020
Natural language and speech › Machine translation › neural machine translation
context-aware neural machine translation
0.412020
Efficient Context-Aware Neural Machine Translation with Layer-Wise Weighting and Input-Aware Gating · IJCAI 2020
Machine learning › Deep learning architectures and training › transformer › transformer training
deep transformer training
0.412020
Lipschitz Constrained Parameter Initialization for Deep Transformers · ACL 2020
Machine learning › Representation and self-supervised learning › text embedding
phrase representation learning
0.412020
Learning Source Phrase Representations for Neural Machine Translation · ACL 2020
Machine learning › Deep learning architectures and training › transformer
transformer training
0.412020
Dynamically Adjusting Transformer Batch Size by Monitoring Gradient Direction Change · ACL 2020
Machine learning › Deep learning architectures and training
weight initialization
0.412020
Lipschitz Constrained Parameter Initialization for Deep Transformers · ACL 2020

Methods — techniques the papers use, named apart from their topics

multi-head attention · 0.9parallel decoding · 0.5residual connections · 0.4lipschitz constraint · 0.4layer normalization · 0.4input-aware gating · 0.4gradient sampling · 0.4gradient accumulation · 0.4cross-attention · 0.4attentive phrase representation · 0.4
YearPublicationVenuePosition
2025 Very-Long-Distance Dependency Capturing Evaluation via Language Modeling Based on Gender Consistency
Hongfei Xu, Zhuofei Liang, Josef van Genabith, Deyi Xiong, Hongying Zan, Qiuhui Liu, Tengxun Zhang
NLPCC (4)6
2024 Rewiring the Transformer with Depth-Wise LSTMs
abstract
Stacking non-linear layers allows deep neural networks to model complicated functions, and including residual connections in Transformer layers is beneficial for convergence and performance. However, residual connections may make the model “forget” distant layers and fail to fuse information from previous layers effectively. Selectively managing the representation aggregation of Transformer layers may lead to better performance. In this paper, we present a Transformer with depth-wise LSTMs connecting cascading Transformer layers and sub-layers. We show that layer normalization and feed-forward computation within a Transformer layer can be absorbed into depth-wise LSTMs connecting pure Transformer attention layers. Our experiments with the 6-layer Transformer show significant BLEU improvements in both WMT 14 English-German / French tasks and the OPUS-100 many-to-many multilingual NMT task, and our deep Transformer experiments demonstrate the effectiveness of depth-wise LSTM on the convergence and performance of deep Transformers.
Hongfei Xu, Qiuhui Liu, Josef van Genabith, Deyi Xiong
LREC/COLING3
2021 Multi-Head Highly Parallelized LSTM Decoder for Neural Machine Translation
abstract
Hongfei Xu, Qiuhui Liu, Josef van Genabith, Deyi Xiong, Meng Zhang. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Hongfei Xu, Qiuhui Liu, Josef van Genabith, Deyi Xiong, Meng Zhang 0019
ACL/IJCNLP (1)2
2021 Probing Word Translations in the Transformer and Trading Decoder for Encoder Layers
abstract
Hongfei Xu, Josef van Genabith, Qiuhui Liu, Deyi Xiong. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021.
Hongfei Xu, Josef van Genabith, Qiuhui Liu, Deyi Xiong
NAACL-HLT3
2020 Dynamically Adjusting Transformer Batch Size by Monitoring Gradient Direction Change
abstract
The choice of hyper-parameters affects the performance of neural models.While much previous research (Sutskever et al., 2013;Duchi et al., 2011;Kingma and Ba, 2015) focuses on accelerating convergence and reducing the effects of the learning rate, comparatively few papers concentrate on the effect of batch size.In this paper, we analyze how increasing batch size affects gradient direction, and propose to evaluate the stability of gradients with their angle change.Based on our observations, the angle change of gradient direction first tends to stabilize (i.e.gradually decrease) while accumulating mini-batches, and then starts to fluctuate.We propose to automatically and dynamically determine batch sizes by accumulating gradients of mini-batches and performing an optimization step at just the time when the direction of gradients starts to fluctuate.To improve the efficiency of our approach for large models, we propose a sampling approach to select gradients of parameters sensitive to the batch size.Our approach dynamically determines proper and efficient batch sizes during training.In our experiments on the WMT 14 English to German and English to French tasks, our approach improves the Transformer with a fixed 25k batch size by +0.73 and +0.82 BLEU respectively.
Hongfei Xu, Josef van Genabith, Deyi Xiong, Qiuhui Liu
ACL4
2020 Learning Source Phrase Representations for Neural Machine Translation
abstract
The Transformer translation model (Vaswani et al., 2017) based on a multi-head attention mechanism can be computed effectively in parallel and has significantly pushed forward the performance of Neural Machine Translation (NMT).Though intuitively the attentional network can connect distant words via shorter network paths than RNNs, empirical analysis demonstrates that it still has difficulty in fully capturing long-distance dependencies (Tang et al., 2018).Considering that modeling phrases instead of words has significantly improved the Statistical Machine Translation (SMT) approach through the use of larger translation blocks ("phrases") and its reordering ability, modeling NMT at phrase level is an intuitive proposal to help the model capture long-distance relationships.In this paper, we first propose an attentive phrase representation generation mechanism which is able to generate phrase representations from corresponding token representations.In addition, we incorporate the generated phrase representations into the Transformer translation model to enhance its ability to capture long-distance relationships.In our experiments, we obtain significant improvements on the WMT 14 English-German and English-French tasks on top of the strong Transformer baseline, which shows the effectiveness of our approach.Our approach helps Transformer Base models perform at the level of Transformer Big models, and even significantly better for long sentences, but with substantially fewer parameters and training steps.The fact that phrase representations help even in the big setting further supports our conjecture that they make a valuable contribution to long-distance relations.
Hongfei Xu, Josef van Genabith, Deyi Xiong, Qiuhui Liu, Jingyi Zhang 0002
ACL4
2020 Lipschitz Constrained Parameter Initialization for Deep Transformers
abstract
The Transformer translation model employs residual connection and layer normalization to ease the optimization difficulties caused by its multi-layer encoder/decoder structure. Previous research shows that even with residual connection and layer normalization, deep Transformers still have difficulty in training, and particularly Transformer models with more than 12 encoder/decoder layers fail to converge. In this paper, we first empirically demonstrate that a simple modification made in the official implementation, which changes the computation order of residual connection and layer normalization, can significantly ease the optimization of deep Transformers. We then compare the subtle differences in computation order in considerable detail, and present a parameter initialization method that leverages the Lipschitz constraint on the initialization of Transformer parameters that effectively ensures training convergence. In contrast to findings in previous research we further demonstrate that with Lipschitz parameter initialization, deep Transformers with the original computation order can converge, and obtain significant BLEU improvements with up to 24 layers. In contrast to previous research which focuses on deep encoders, our approach additionally enables Transformers to also benefit from deep decoders.
Hongfei Xu, Qiuhui Liu, Josef van Genabith, Deyi Xiong, Jingyi Zhang 0002
ACL2
2020 Efficient Context-Aware Neural Machine Translation with Layer-Wise Weighting and Input-Aware Gating
abstract
Existing Neural Machine Translation (NMT) systems are generally trained on a large amount of sentence-level parallel data, and during prediction sentences are independently translated, ignoring cross-sentence contextual information. This leads to inconsistency between translated sentences. In order to address this issue, context-aware models have been proposed. However, document-level parallel data constitutes only a small part of the parallel data available, and many approaches build context-aware models based on a pre-trained frozen sentence-level translation model in a two-step training manner. The computational cost of these approaches is usually high. In this paper, we propose to make the most of layers pre-trained on sentence-level data in contextual representation learning, reusing representations from the sentence-level Transformer and significantly reducing the cost of incorporating contexts in translation. We find that representations from shallow layers of a pre-trained sentence-level encoder play a vital role in source context encoding, and propose to perform source context encoding upon weighted combinations of pre-trained encoder layers' outputs. Instead of separately performing source context and input encoding, we propose to iteratively and jointly encode the source input and its contexts and to generate input-aware context representations with a cross-attention layer and a gating mechanism, which resets irrelevant information in context encoding. Our context-aware Transformer model outperforms the recent CADec [Voita et al., 2019c] on the English-Russian subtitle data and is about twice as fast in training and decoding.
Hongfei Xu, Deyi Xiong, Josef van Genabith, Qiuhui Liu
IJCAI4
2017 Improving Chinese-English Neural Machine Translation with Detected Usages of Function Words
Kunli Zhang, Hongfei Xu, Deyi Xiong, Qiuhui Liu, Hongying Zan
NLPCC4