EDBT 2026 Demo / reviewers in the wild / expert
Biao Zhang 0002
dblp:83/3266-2
· DBLP profile ↗
22ranked-venue papers
15as first author
2since 2021 · last 2022
0000-0002-4865-7090ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 21 · 15 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 first-authorDatabases, data management, data science and information retrieval · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
12 papers |
Machine translation · 36% Deep learning architectures and training · 22% Generative modeling · 15% | |
| Software engineering, system software, and programming languages
1 paper |
Compilers and program optimization · 100% |
Topics — the 24 heaviest of 24, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Machine translation
neural machine translation |
2.0 | 6 | 2020 | Neural Machine Translation with Deep Attention · IEEE Trans. Pattern Anal. Mach. Intell. 2020 Future-Aware Knowledge Distillation for Neural Machine Translation · IEEE ACM Trans. Audio Speech Lang. Process. 2019 Simplifying Neural Machine Translation with Addition-Subtraction Twin-Gated Recurrent Networks · EMNLP 2018 |
Machine learning › Generative modeling
variational autoencoder |
0.8 | 3 | 2018 | Variational Recurrent Neural Machine Translation · AAAI 2018 Variational Neural Discourse Relation Recognizer · EMNLP 2016 Variational Neural Machine Translation · EMNLP 2016 |
Machine learning › Deep learning architectures and training
attention mechanism |
0.8 | 2 | 2020 | Neural Machine Translation with Deep Attention · IEEE Trans. Pattern Anal. Mach. Intell. 2020 Accelerating Neural Transformer via an Average Attention Network · ACL (1) 2018 |
Natural language and speech › Machine translation
statistical machine translation |
0.5 | 2 | 2017 | BattRAE: Bidimensional Attention-Based Recursive Autoencoders for Learning Bilingual Phrase Embeddings · AAAI 2017 Bilingual Correspondence Recursive Autoencoder for Statistical Machine Translation · EMNLP 2015 |
Compilers and program optimization
code generation |
0.5 | 1 | 2021 | Exploring Dynamic Selection of Branch Expansion Orders for Code Generation · ACL/IJCNLP (1) 2021 |
Natural language and speech › Information extraction and text analysis › discourse analysis › discourse relation recognition
implicit discourse relation recognition |
0.5 | 2 | 2016 | Variational Neural Discourse Relation Recognizer · EMNLP 2016 Shallow Convolutional Neural Network for Implicit Discourse Relation Recognition · EMNLP 2015 |
Machine learning › Efficient and distributed learning › model compression
knowledge distillation |
0.4 | 1 | 2019 | Future-Aware Knowledge Distillation for Neural Machine Translation · IEEE ACM Trans. Audio Speech Lang. Process. 2019 |
Natural language and speech › Language models and text generation › decoding
efficient decoding |
0.3 | 1 | 2018 | Accelerating Neural Transformer via an Average Attention Network · ACL (1) 2018 |
Machine learning › Deep learning architectures and training › recurrent neural network › gated recurrent network
gated recurrent unit |
0.3 | 1 | 2018 | Simplifying Neural Machine Translation with Addition-Subtraction Twin-Gated Recurrent Networks · EMNLP 2018 |
Machine learning › Deep learning architectures and training
recurrent neural network |
0.3 | 1 | 2018 | Simplifying Neural Machine Translation with Addition-Subtraction Twin-Gated Recurrent Networks · EMNLP 2018 |
Natural language and speech › Machine translation › neural machine translation
variational neural machine translation |
0.3 | 1 | 2018 | Variational Recurrent Neural Machine Translation · AAAI 2018 |
Machine learning › Deep learning architectures and training › recurrent neural network
recurrent neural network encoder |
0.3 | 1 | 2017 | A Context-Aware Recurrent Encoder for Neural Machine Translation · IEEE ACM Trans. Audio Speech Lang. Process. 2017 |
Machine learning › Generative modeling › diffusion model
conditional generation |
0.2 | 1 | 2016 | Variational Neural Machine Translation · EMNLP 2016 |
Natural language and speech › Information extraction and text analysis › discourse analysis
discourse relation recognition |
0.2 | 1 | 2016 | Variational Neural Discourse Relation Recognizer · EMNLP 2016 |
Machine learning › Probabilistic and Bayesian machine learning › structured models
latent variable model |
0.2 | 1 | 2016 | Variational Neural Discourse Relation Recognizer · EMNLP 2016 |
Machine learning › Generative modeling
variational encoder-decoder |
0.2 | 1 | 2016 | Variational Neural Machine Translation · EMNLP 2016 |
Natural language and speech › Information extraction and text analysis
discourse analysis |
0.2 | 1 | 2015 | Shallow Convolutional Neural Network for Implicit Discourse Relation Recognition · EMNLP 2015 |
Machine learning › Transfer learning and domain adaptation
domain adaptation |
0.2 | 1 | 2015 | Discriminative Reordering Model Adaptation via Structural Learning · IJCAI 2015 |
Machine learning › Representation and self-supervised learning › hierarchical representation
recursive autoencoder |
0.2 | 1 | 2015 | Bilingual Correspondence Recursive Autoencoder for Statistical Machine Translation · EMNLP 2015 |
Natural language and speech › Machine translation › statistical machine translation
reordering model |
0.2 | 1 | 2015 | Discriminative Reordering Model Adaptation via Structural Learning · IJCAI 2015 |
Machine learning › Deep learning architectures and training
encoder-decoder architecture |
0.1 | 1 | 2020 | Neural Machine Translation with Deep Attention · IEEE Trans. Pattern Anal. Mach. Intell. 2020 |
Machine learning › Representation and self-supervised learning › text embedding
phrase representation learning |
0.1 | 1 | 2017 | BattRAE: Bidimensional Attention-Based Recursive Autoencoders for Learning Bilingual Phrase Embeddings · AAAI 2017 |
Natural language and speech › Machine translation
chinese-english translation |
0.1 | 1 | 2015 | Bilingual Correspondence Recursive Autoencoder for Statistical Machine Translation · EMNLP 2015 |
Machine learning › Deep learning architectures and training
convolutional neural network |
0.1 | 1 | 2015 | Shallow Convolutional Neural Network for Implicit Discourse Relation Recognition · EMNLP 2015 |
Methods — techniques the papers use, named apart from their topics
reparameterization · 0.8recursive autoencoder · 0.5variational inference · 0.5sequence-to-tree decoding · 0.5neural posterior approximation · 0.5dynamic selection · 0.5multi-layer attention · 0.4deep attention · 0.4neural language model · 0.4knowledge distillation · 0.4attention mechanism · 0.4neural posterior approximator · 0.3
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2022 | AAN+: Generalized Average Attention Network for Accelerating Neural TransformerabstractTransformer benefits from the high parallelization of attention networks in fast training, but it still suffers from slow decoding partially due to the linear dependency O(m) of the decoder self-attention on previous target words at inference. In this paper, we propose a generalized average attention network (AAN+) aiming at speeding up decoding by reducing the dependency from O(m) to O(1). We find that the learned self-attention weights in the decoder follow some patterns which can be approximated via a dynamic structure. Based on this insight, we develop AAN+, extending our previously proposed average attention (Zhang et al., 2018a, AAN) to support more general position- and content-based attention patterns. AAN+ only requires to maintain a small constant number of hidden states during decoding, ensuring its O(1) dependency. We apply AAN+ as a drop-in replacement of the decoder selfattention and conduct experiments on machine translation (with diverse language pairs), table-to-text generation and document summarization. With masking tricks and dynamic programming, AAN+ enables Transformer to decode sentences around 20% faster without largely compromising in the training speed and the generation performance. Our results further reveal the importance of the localness (neighboring words) in AAN+ and its capability in modeling long-range dependency. Biao Zhang 0002, Deyi Xiong, Yubin Ge, Junfeng Yao, Jinsong Su |
J. Artif. Intell. Res. | 1 |
| 2021 | Exploring Dynamic Selection of Branch Expansion Orders for Code GenerationabstractHui Jiang, Chulun Zhou, Fandong Meng, Biao Zhang, Jie Zhou, Degen Huang, Qingqiang Wu, Jinsong Su. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Chulun Zhou, Fandong Meng, Biao Zhang 0002, Jie Zhou 0016, Degen Huang, Qingqiang Wu 0001, Jinsong Su |
ACL/IJCNLP (1) | 4 |
| 2020 | Neural Machine Translation with Deep AttentionabstractDeepening neural models has been proven very successful in improving the model's capacity when solving complex learning tasks, such as the machine translation task. Previous efforts on deep neural machine translation mainly focus on the encoder and the decoder, while little on the attention mechanism. However, the attention mechanism is of vital importance to induce the translation correspondence between different languages where shallow neural networks are relatively insufficient, especially when the encoder and decoder are deep. In this paper, we propose a deep attention model (DeepAtt). Based on the low-level attention information, DeepAtt is capable of automatically determining what should be passed or suppressed from the corresponding encoder layer so as to make the distributed representation appropriate for high-level attention and translation. We conduct experiments on NIST Chinese-English, WMT English-German, and WMT English-French translation tasks, where, with five attention layers, DeepAtt yields very competitive performance against the state-of-the-art results. We empirically find that with an adequate increase of attention layers, DeepAtt tends to produce more accurate attention weights. An in-depth analysis on the translation of important context words further reveals that DeepAtt significantly improves the faithfulness of system translations. Biao Zhang 0002, Deyi Xiong, Jinsong Su |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2020 | Alignment-Supervised Bidimensional Attention-Based Recursive Autoencoders for Bilingual Phrase RepresentationabstractExploiting semantic interactions between the source and target linguistic items at different levels of granularity is crucial for generating compact vector representations for bilingual phrases. To achieve this, we propose alignment-supervised bidimensional attention-based recursive autoencoders (ABattRAE) in this paper. ABattRAE first individually employs two recursive autoencoders to recover hierarchical tree structures of bilingual phrase, and treats the subphrase covered by each node on the tree as a linguistic item. Unlike previous methods, ABattRAE introduces a bidimensional attention network to measure the semantic matching degree between linguistic items of different languages, which enables our model to integrate information from all nodes by dynamically assigning varying weights to their corresponding embeddings. To ensure the accuracy of the generated attention weights in the attention network, ABattRAE incorporates word alignments as supervision signals to guide the learning procedure. Using the general stochastic gradient descent algorithm, we train our model in an end-to-end fashion, where the semantic similarity of translation equivalents is maximized while the semantic similarity of nontranslation pairs is minimized. Finally, we incorporate a semantic feature based on the learned bilingual phrase representations into a machine translation system for better translation selection. Experimental results on NIST Chinese-English and WMT English-German test sets show that our model achieves substantial improvements of up to 2.86 and 1.09 BLEU points over the baseline, respectively. Extensive in-depth analyses demonstrate the superiority of our model in learning bilingual phrase embeddings. Biao Zhang 0002, Deyi Xiong, Jinsong Su |
IEEE Trans. Cybern. | 1 |
| 2020 | Neural Machine Translation With GRU-Gated Attention ModelabstractNeural machine translation (NMT) heavily relies on context vectors generated by an attention network to predict target words. In practice, we observe that the context vectors for different target words are quite similar to one another and translations with such nondiscriminatory context vectors tend to be degenerative. We ascribe this similarity to the invariant source representations that lack dynamics across decoding steps. In this article, we propose a novel gated recurrent unit (GRU)-gated attention model (GAtt) for NMT. By updating the source representations with the previous decoder state via a GRU, GAtt enables translation-sensitive source representations that then contribute to discriminative context vectors. We further propose a variant of GAtt by swapping the input order of the source representations and the previous decoder state to the GRU. Experiments on the NIST Chinese-English, WMT14 English-German, and WMT17 English-German translation tasks show that the two GAtt models achieve significant improvements over the vanilla attention-based NMT. Further analyses on the attention weights and context vectors demonstrate the effectiveness of GAtt in enhancing the discriminating capacity of representations and handling the challenging issue of overtranslation. Biao Zhang 0002, Deyi Xiong, Jinsong Su |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2019 | Future-Aware Knowledge Distillation for Neural Machine TranslationabstractAlthough future context is widely regarded useful for word prediction in machine translation, it is quite difficult in practice to incorporate it into neural machine translation. In this paper, we propose a future-aware knowledge distillation framework (FKD) to address this issue. In the FKD framework, we learn to distill future knowledge from a backward neural language model (teacher) to future-aware vectors (student) during the training phase. The future-aware vector for each word position is computed in a bridge network and optimized towards the corresponding hidden state in the backward neural language model via a knowledge distillation mechanism. We further propose an algorithm to jointly train the neural machine translation model, neural language model and knowledge distillation module end-to-end. The learned future-aware vectors are incorporated into the attention layer of the decoder to provide full-range context information during the decoding phase. Experiments on the NIST Chinese-English and WMT English-German translation tasks show that the proposed method significantly improves translation quality and word alignment. Biao Zhang 0002, Deyi Xiong, Jinsong Su, Jiebo Luo 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2018 | Variational Recurrent Neural Machine TranslationabstractPartially inspired by successful applications of variational recurrent neural networks, we propose a novel variational recurrent neural machine translation (VRNMT) model in this paper. Different from the variational NMT, VRNMT introduces a series of latent random variables to model the translation procedure of a sentence in a generative way, instead of a single latent variable. Specifically, the latent random variables are included into the hidden states of the NMT decoder with elements from the variational autoencoder. In this way, these variables are recurrently generated, which enables them to further capture strong and complex dependencies among the output translations at different timesteps. In order to deal with the challenges in performing efficient posterior inference and large-scale training during the incorporation of latent variables, we build a neural posterior approximator, and equip it with a reparameterization technique to estimate the variational lower bound. Experiments on Chinese-English and English-German translation tasks demonstrate that the proposed model achieves significant improvements over both the conventional and variational NMT models. Jinsong Su, Deyi Xiong, Yaojie Lu 0001, Xianpei Han, Biao Zhang 0002 |
AAAI | 6 |
| 2018 | Accelerating Neural Transformer via an Average Attention NetworkabstractWith parallelizable attention networks, the neural Transformer is very fast to train.However, due to the auto-regressive architecture and self-attention in the decoder, the decoding procedure becomes slow.To alleviate this issue, we propose an average attention network as an alternative to the self-attention network in the decoder of the neural Transformer.The average attention network consists of two layers, with an average layer that models dependencies on previous positions and a gating layer that is stacked over the average layer to enhance the expressiveness of the proposed attention network.We apply this network on the decoder part of the neural Transformer to replace the original target-side self-attention model.With masking tricks and dynamic programming, our model enables the neural Transformer to decode sentences over four times faster than its original version with almost no loss in training time and translation performance.We conduct a series of experiments on WMT17 translation tasks, where on 6 different language pairs, we obtain robust and consistent speed-ups in decoding.1 Biao Zhang 0002, Deyi Xiong, Jinsong Su |
ACL (1) | 1 |
| 2018 | Simplifying Neural Machine Translation with Addition-Subtraction Twin-Gated Recurrent NetworksabstractIn this paper, we propose an additionsubtraction twin-gated recurrent network (ATR) to simplify neural machine translation.The recurrent units of ATR are heavily simplified to have the smallest number of weight matrices among units of all existing gated RNNs.With the simple addition and subtraction operation, we introduce a twin-gated mechanism to build input and forget gates which are highly correlated.Despite this simplification, the essential non-linearities and capability of modeling long-distance dependencies are preserved.Additionally, the proposed ATR is more transparent than LSTM/GRU due to the simplification.Forward self-attention can be easily established in ATR, which makes the proposed network interpretable.Experiments on WMT14 translation tasks demonstrate that ATR-based neural machine translation can yield competitive performance on English-German and English-French language pairs in terms of both translation quality and speed.Further experiments on NIST Chinese-English translation, natural language inference and Chinese word segmentation verify the generality and applicability of ATR on different natural language processing tasks. Biao Zhang 0002, Deyi Xiong, Jinsong Su, Huiji Zhang |
EMNLP | 1 |
| 2018 | Otem&Utem: Over- and Under-Translation Evaluation Metric for NMT
Biao Zhang 0002, Xiangwen Zhang, Jinsong Su |
NLPCC (1) | 2 |
| 2018 | Learning better discourse representation for implicit discourse relation recognition via attention networks
Biao Zhang 0002, Deyi Xiong, Jinsong Su, Min Zhang 0005 |
Neurocomputing | 1 |
| 2018 | A neural generative autoencoder for bilingual word embeddings
Jinsong Su, Biao Zhang 0002, Changxing Wu, Deyi Xiong |
Inf. Sci. | 3 |
| 2018 | Alignment-consistent recursive neural networks for bilingual phrase embeddings
Jinsong Su, Biao Zhang 0002, Deyi Xiong, Yang Liu 0005, Min Zhang 0005 |
Knowl. Based Syst. | 2 |
| 2017 | BattRAE: Bidimensional Attention-Based Recursive Autoencoders for Learning Bilingual Phrase EmbeddingsabstractIn this paper, we propose a bidimensional attention based recursiveautoencoder (BattRAE) to integrate clues and sourcetargetinteractions at multiple levels of granularity into bilingualphrase representations. We employ recursive autoencodersto generate tree structures of phrases with embeddingsat different levels of granularity (e.g., words, sub-phrases andphrases). Over these embeddings on the source and targetside, we introduce a bidimensional attention network to learntheir interactions encoded in a bidimensional attention matrix,from which we extract two soft attention weight distributionssimultaneously. These weight distributions enableBattRAE to generate compositive phrase representations viaconvolution. Based on the learned phrase representations, wefurther use a bilinear neural model, trained via a max-marginmethod, to measure bilingual semantic similarity. To evaluatethe effectiveness of BattRAE, we incorporate this semanticsimilarity as an additional feature into a state-of-the-art SMTsystem. Extensive experiments on NIST Chinese-English testsets show that our model achieves a substantial improvementof up to 1.63 BLEU points on average over the baseline. Biao Zhang 0002, Deyi Xiong, Jinsong Su |
AAAI | 1 |
| 2017 | A Context-Aware Recurrent Encoder for Neural Machine TranslationabstractNeural machine translation (NMT) heavily relies on its encoder to capture the underlying meaning of a source sentence so as to generate a faithful translation. However, most NMT encoders are built upon either unidirectional or bidirectional recurrent neural networks, which either do not deal with future context or simply concatenate the history and future context to form context-dependent word representations, implicitly assuming the independence of the two types of contextual information. In this paper, we propose a novel context-aware recurrent encoder (CAEncoder), as an alternative to the widely-used bidirectional encoder, such that the future and history contexts can be fully incorporated into the learned source representations. Our CAEncoder involves a two-level hierarchy: The bottom level summarizes the history information, whereas the upper level assembles the summarized history and future context into source representations. Additionally, CAEncoder is as efficient as the bidirectional RNN encoder in terms of both training and decoding. Experiments on both Chinese-English and English-German translation tasks show that CAEncoder achieves significant improvements over the bidirectional RNN encoder on a widely-used NMT system. Biao Zhang 0002, Deyi Xiong, Jinsong Su, Hong Duan |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2016 | Convolution-Enhanced Bilingual Recursive Neural Network for Bilingual Semantic ModelingabstractEstimating similarities at different levels of linguistic units, such as words, sub-phrases and phrases, is helpful for measuring semantic similarity of an entire bilingual phrase. In this paper, we propose a convolution-enhanced bilingual recursive neural network (ConvBRNN), which not only exploits word alignments to guide the generation of phrase structures but also integrates multiple-level information of the generated phrase structures into bilingual semantic modeling. In order to accurately learn the semantic hierarchy of a bilingual phrase, we develop a recursive neural network to constrain the learned bilingual phrase structures to be consistent with word alignments. Upon the generated source and target phrase structures, we stack a convolutional neural network to integrate vector representations of linguistic units on the structures into bilingual phrase embeddings. After that, we fully incorporate information of different linguistic units into a bilinear semantic similarity model. We introduce two max-margin losses to train the ConvBRNN model: one for the phrase structure inference and the other for the semantic similarity model. Experiments on NIST Chinese-English translation tasks demonstrate the high quality of the generated bilingual phrase structures with respect to word alignments and the effectiveness of learned semantic similarities on machine translation. Jinsong Su, Biao Zhang 0002, Deyi Xiong, Jianmin Yin |
COLING | 2 |
| 2016 | Bilingual Autoencoders with Global Descriptors for Modeling Parallel SentencesabstractParallel sentence representations are important for bilingual and cross-lingual tasks in natural language processing. In this paper, we explore a bilingual autoencoder approach to model parallel sentences. We extract sentence-level global descriptors (e.g. min, max) from word embeddings, and construct two monolingual autoencoders over these descriptors on the source and target language. In order to tightly connect the two autoencoders with bilingual correspondences, we force them to share the same decoding parameters and minimize a corpus-level semantic distance between the two languages. Being optimized towards a joint objective function of reconstruction and semantic errors, our bilingual antoencoder is able to learn continuous-valued latent representations for parallel sentences. Experiments on both intrinsic and extrinsic evaluations on statistical machine translation tasks show that our autoencoder achieves substantial improvements over the baselines. Biao Zhang 0002, Deyi Xiong, Jinsong Su, Hong Duan, Min Zhang 0005 |
COLING | 1 |
| 2016 | Variational Neural Machine TranslationabstractModels of neural machine translation are often from a discriminative family of encoderdecoders that learn a conditional distribution of a target sentence given a source sentence.In this paper, we propose a variational model to learn this conditional distribution for neural machine translation: a variational encoderdecoder model that can be trained end-to-end.Different from the vanilla encoder-decoder model that generates target translations from hidden representations of source sentences alone, the variational model introduces a continuous latent variable to explicitly model underlying semantics of source sentences and to guide the generation of target translations.In order to perform efficient posterior inference and large-scale training, we build a neural posterior approximator conditioned on both the source and the target sides, and equip it with a reparameterization technique to estimate the variational lower bound.Experiments on both Chinese-English and English-German translation tasks show that the proposed variational neural machine translation achieves significant improvements over the vanilla neural machine translation baselines. Biao Zhang 0002, Deyi Xiong, Jinsong Su, Hong Duan, Min Zhang 0005 |
EMNLP | 1 |
| 2016 | Variational Neural Discourse Relation RecognizerabstractImplicit discourse relation recognition is a crucial component for automatic discourselevel analysis and nature language understanding.Previous studies exploit discriminative models that are built on either powerful manual features or deep discourse representations.In this paper, instead, we explore generative models and propose a variational neural discourse relation recognizer.We refer to this model as VarNDRR.VarNDRR establishes a directed probabilistic model with a latent continuous variable that generates both a discourse and the relation between the two arguments of the discourse.In order to perform efficient inference and learning, we introduce neural discourse relation models to approximate the prior and posterior distributions of the latent variable, and employ these approximated distributions to optimize a reparameterized variational lower bound.This allows VarNDRR to be trained with standard stochastic gradient methods.Experiments on the benchmark data set show that VarNDRR can achieve comparable results against stateof-the-art baselines without using any manual features. Biao Zhang 0002, Deyi Xiong, Jinsong Su, Qun Liu 0001, Rongrong Ji, Hong Duan, Min Zhang 0005 |
EMNLP | 1 |
| 2015 | Bilingual Correspondence Recursive Autoencoder for Statistical Machine TranslationabstractLearning semantic representations and tree structures of bilingual phrases is beneficial for statistical machine translation.In this paper, we propose a new neural network model called Bilingual Correspondence Recursive Autoencoder (BCor-rRAE) to model bilingual phrases in translation.We incorporate word alignments into BCorrRAE to allow it freely access bilingual constraints at different levels.BCorrRAE minimizes a joint objective on the combination of a recursive autoencoder reconstruction error, a structural alignment consistency error and a crosslingual reconstruction error so as to not only generate alignment-consistent phrase structures, but also capture different levels of semantic relations within bilingual phrases.In order to examine the effectiveness of BCorrRAE, we incorporate both semantic and structural similarity features built on bilingual phrase representations and tree structures learned by BCorrRAE into a state-of-the-art SMT system.Experiments on NIST Chinese-English test sets show that our model achieves a substantial improvement of up to 1.55 BLEU points over the baseline. Jinsong Su, Deyi Xiong, Biao Zhang 0002, Yang Liu 0005, Junfeng Yao, Min Zhang 0005 |
EMNLP | 3 |
| 2015 | Shallow Convolutional Neural Network for Implicit Discourse Relation RecognitionabstractImplicit discourse relation recognition remains a serious challenge due to the absence of discourse connectives.In this paper, we propose a Shallow Convolutional Neural Network (SCNN) for implicit discourse relation recognition, which contains only one hidden layer but is effective in relation recognition.The shallow structure alleviates the overfitting problem, while the convolution and nonlinear operations help preserve the recognition and generalization ability of our model.Experiments on the benchmark data set show that our model achieves comparable and even better performance when comparing against current state-of-the-art systems. Biao Zhang 0002, Jinsong Su, Deyi Xiong, Yaojie Lu 0001, Hong Duan, Junfeng Yao |
EMNLP | 1 |
| 2015 | Discriminative Reordering Model Adaptation via Structural Learning
Biao Zhang 0002, Jinsong Su, Deyi Xiong, Hong Duan, Junfeng Yao |
IJCAI | 1 |