Zi-Yi Dou

dblp:205/8985 · DBLP profile ↗
← Back
28ranked-venue papers
19as first author
17since 2021 · last 2025
0000-0002-8647-0200ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 28 · 19 first-author · 17 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 4 first-author · 4 since 2021
YearPublicationVenuePosition
2025 MRAG-Bench: Vision-Centric Evaluation for Retrieval-Augmented Multimodal Models
abstract
Existing multimodal retrieval benchmarks primarily focus on evaluating whether models can retrieve and utilize external textual knowledge for question answering. However, there are scenarios where retrieving visual information is either more beneficial or easier to access than textual data. In this paper, we introduce a multimodal retrieval-augmented generation benchmark, MRAG-Bench, in which we systematically identify and categorize scenarios where visually augmented knowledge is better than textual knowledge, for instance, more images from varying viewpoints. MRAG-Bench consists of 16,130 images and 1,353 human-annotated multiple-choice questions across 9 distinct scenarios. With MRAG-Bench, we conduct an evaluation of 10 open-source and 4 proprietary large vision-language models (LVLMs). Our results show that all LVLMs exhibit greater improvements when augmented with images compared to textual knowledge, confirming that MRAG-Bench is vision-centric. Additionally, we conduct extensive analysis with MRAG-Bench, which offers valuable insights into retrieval-augmented LVLMs. Notably, the top-performing model, GPT-4o, faces challenges in effectively leveraging retrieved knowledge, achieving only a 5.82\% improvement with ground-truth information, in contrast to a 33.16\% improvement observed in human participants. These findings highlight the importance of MRAG-Bench in encouraging the community to enhance LVLMs' ability to utilize retrieved visual knowledge more effectively.
Wenbo Hu 0006, Jia-Chen Gu, Zi-Yi Dou, Mohsen Fayyaz, Pan Lu, Kai-Wei Chang 0001, Nanyun Peng 0001
ICLR3
2024 Medical Vision-Language Pre-Training for Brain Abnormalities
abstract
Vision-language models have become increasingly powerful for tasks that require an understanding of both visual and linguistic elements, bridging the gap between these modalities. In the context of multimodal clinical AI, there is a growing need for models that possess domain-specific knowledge, as existing models often lack the expertise required for medical applications. In this paper, we take brain abnormalities as an example to demonstrate how to automatically collect medical image-text aligned data for pretraining from public resources such as PubMed. In particular, we present a pipeline that streamlines the pre-training process by initially collecting a large brain image-text dataset from case reports and published journals and subsequently constructing a high-performance vision-language model tailored to specific medical tasks. We also investigate the unique challenge of mapping subfigures to subcaptions in the medical domain. We evaluated the resulting model with quantitative and qualitative intrinsic evaluations. The resulting dataset will be released to the community.
Masoud Monajatipoor, Zi-Yi Dou, Aichi Chien, Nanyun Peng 0001, Kai-Wei Chang 0001
LREC/COLING2
2024 Re-ReST: Reflection-Reinforced Self-Training for Language Agents
abstract
Finetuning language agents with reasoningaction trajectories is effective, but obtaining these trajectories from human annotations or stronger models is costly and sometimes impractical.In this paper, we investigate the use of self-training in language agents, which can generate supervision from the agent itself, offering a promising alternative without relying on human or stronger model demonstrations.Self-training, however, requires high-quality model-generated samples, which are hard to obtain for challenging language agent tasks.To address this, we present Reflection-Reinforced Self-Training (Re-ReST), which uses a reflector to refine low-quality generated samples during self-training.The reflector takes the agent's output and feedback from an external environment (e.g., unit test results in code generation) to produce improved samples.This technique enhances the quality of inferior samples and efficiently enriches the self-training dataset with higher-quality samples.We conduct extensive experiments on open-source language agents across tasks, including multi-hop question answering, sequential decision-making, code generation, visual question answering, and text-toimage generation.The results demonstrate the effectiveness of self-training and Re-ReST in language agent tasks, with self-training improving baselines by 7.6% on HotpotQA and 28.4% on AlfWorld, and Re-ReST further boosting performance by 2.0% and 14.1%, respectively.Our studies also confirm the efficiency of using a reflector to generate high-quality samples for self-training.Moreover, we demonstrate a method to employ reflection during inference without ground-truth feedback, addressing the limitation of previous reflection work.
Zi-Yi Dou, Cheng-Fu Yang, Xueqing Wu 0001, Kai-Wei Chang 0001, Nanyun Peng 0001
EMNLP1
2024 Matryoshka Query Transformer for Large Vision-Language Models
abstract
Large Vision-Language Models (LVLMs) typically encode an image into a fixed number of visual tokens (e.g., 576) and process these tokens with a language model. Despite their strong performance, LVLMs face challenges in adapting to varying computational constraints. This raises the question: can we achieve flexibility in the number of visual tokens to suit different tasks and computational resources? We answer this with an emphatic yes. Inspired by Matryoshka Representation Learning, we introduce the Matryoshka Query Transformer (MQT), capable of encoding an image into $m$ visual tokens during inference, where $m$ can be any number up to a predefined maximum. This is achieved by employing a query transformer with $M$ latent query tokens to compress the visual embeddings. During each training step, we randomly select $m \leq M$ latent query tokens and train the model using only these first $m$ tokens, discarding the rest. Combining MQT with LLaVA, we train a single model once, and flexibly and drastically reduce the number of inference-time visual tokens while maintaining similar or better performance compared to training independent models for each number of tokens. Our model, MQT-LLaVA, matches LLaVA-1.5 performance across 11 benchmarks using a maximum of 256 tokens instead of LLaVA’s fixed 576. Reducing to 16 tokens (8x less TFLOPs) only sacrifices the performance by 2.4 points on MMBench. On certain tasks such as ScienceQA and MMMU, we can even go down to only 2 visual tokens with performance drops of just 3\% and 6\% each. Our exploration of the trade-off between the accuracy and computational cost brought about by the number of visual tokens facilitates future research to achieve the best of both worlds.
Wenbo Hu 0006, Zi-Yi Dou, Liunian Harold Li, Amita Kamath, Nanyun Peng 0001, Kai-Wei Chang 0001
NeurIPS2
2023 Generalized Decoding for Pixel, Image, and Language
abstract
We present X-Decoder, a generalized decoding model that can predict pixel-level segmentation and language tokens seamlessly. X-Decoder takes as input two types of queries: (i) generic non-semantic queries and (ii) semantic queries induced from text inputs, to decode different pixel-level and token-level outputs in the same semantic space. With such a novel design, X-Decoder is the first work that provides a unified way to support all types of image segmentation and a variety of vision-language (VL) tasks. Without any pseudo-labeling, our design enables seamless interactions across tasks at different granularities and brings mutual benefits by learning a common and rich pixel-level understanding. After pretraining on a mixed set of a limited amount of segmentation data and millions of image-text pairs, X-Decoder exhibits strong transferability to a wide range of downstream tasks in both zero-shot and finetuning settings. Notably, it achieves (1) state-of-the-art results on open-vocabulary segmentation and referring segmentation on seven datasets; (2) better or competitive finetuned performance to other generalist and specialist models on segmentation and VL tasks; and (3) flexibility for efficient fine-tuning and novel task composition (e.g., referring captioning and image editing shown in Fig. 1). Code, demo, video and visualization are available at: https://x-decoder-vl.github.io.
Xueyan Zou, Zi-Yi Dou, Zhe Gan, Chunyuan Li, Xiyang Dai, Harkirat Behl, Lu Yuan 0001, Nanyun Peng 0001, Yong Jae Lee, Jianfeng Gao 0001
CVPR2
2023 Gender Biases in Automatic Evaluation Metrics for Image Captioning
abstract
Model-based evaluation metrics (e.g., CLIP-Score and GPTScore) have demonstrated decent correlations with human judgments in various language generation tasks.However, their impact on fairness remains largely unexplored.It is widely recognized that pretrained models can inadvertently encode societal biases, thus employing these models for evaluation purposes may inadvertently perpetuate and amplify biases.For example, an evaluation metric may favor the caption "a woman is calculating an account book" over "a man is calculating an account book," even if the image only shows male accountants.In this paper, we conduct a systematic study of gender biases in modelbased automatic evaluation metrics for image captioning tasks.We start by curating a dataset comprising profession, activity, and object concepts associated with stereotypical gender associations.Then, we demonstrate the negative consequences of using these biased metrics, including the inability to differentiate between biased and unbiased generations, as well as the propagation of biases to generation models through reinforcement learning.Finally, we present a simple and effective way to mitigate the metric bias without hurting the correlations with human judgments.Our dataset and framework lay the foundation for understanding the potential harm of model-based evaluation metrics, and facilitate future works to develop more inclusive evaluation metrics. 1
Haoyi Qiu, Zi-Yi Dou, Asli Celikyilmaz, Nanyun Peng 0001
EMNLP2
2023 ACQUIRED: A Dataset for Answering Counterfactual Questions In Real-Life Videos
abstract
Te-Lin Wu, Zi-Yi Dou, Qingyuan Hu, Yu Hou, Nischal Chandra, Marjorie Freedman, Ralph Weischedel, Nanyun Peng. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 2023.
Te-Lin Wu, Zi-Yi Dou, Nischal Reddy Chandra, Marjorie Freedman, Ralph M. Weischedel, Nanyun Peng 0001
EMNLP2
2023 DesCo: Learning Object Recognition with Rich Language Descriptions
abstract
Recent development in vision-language approaches has instigated a paradigm shift in learning visual recognition models from language supervision. These approaches align objects with language queries (e.g. "a photo of a cat") and thus improve the models' adaptability to novel objects and domains. Recent studies have attempted to query these models with complex language expressions that include specifications of fine-grained details, such as colors, shapes, and relations. However, simply incorporating language descriptions into queries does not guarantee accurate interpretation by the models. In fact, our experiments show that GLIP, a state-of-the-art vision-language model for object detection, often disregards contextual information in the language descriptions and instead relies heavily on detecting objects solely by their names. To tackle the challenge, we propose a new description-conditioned (DesCo) paradigm of learning object recognition models with rich language descriptions consisting of two innovations: 1) we employ a large language model as a commonsense knowledge engine to generate rich language descriptions of objects; 2) we design context-sensitive queries to improve the model's ability in deciphering intricate nuances embedded within descriptions and enforce the model to focus on context rather than object names alone. On two novel object detection benchmarks, LVIS and OminiLabel, under the zero-shot detection setting, our approach achieves 34.8 APr minival (+9.1) and 29.3 AP (+3.6), respectively, surpassing the prior state-of-the-art models, GLIP and FIBER, by a large margin.
Liunian Harold Li, Zi-Yi Dou, Nanyun Peng 0001, Kai-Wei Chang 0001
NeurIPS2
2022 Zero-Shot Commonsense Question Answering with Cloze Translation and Consistency Optimization
abstract
Commonsense question answering (CQA) aims to test if models can answer questions regarding commonsense knowledge that everyone knows. Prior works that incorporate external knowledge bases have shown promising results, but knowledge bases are expensive to construct and are often limited to a fixed set of relations. In this paper, we instead focus on better utilizing the implicit knowledge stored in pre-trained language models. While researchers have found that the knowledge embedded in pre-trained language models can be extracted by having them fill in the blanks of carefully designed prompts for relation extraction and text classification, it remains unclear if we can adopt this paradigm in CQA where the inputs and outputs take much more flexible forms. To this end, we investigate four translation methods that can translate natural questions into cloze-style sentences to better solicit commonsense knowledge from language models, including a syntactic-based model, an unsupervised neural model, and two supervised neural models. In addition, to combine the different translation methods, we propose to encourage consistency among model predictions on different translated questions with unlabeled data. We demonstrate the effectiveness of our methods on three CQA datasets in zero-shot settings. We show that our methods are complementary to a knowledge base improved model, and combining them can lead to state-of-the-art zero-shot performance. Analyses also reveal distinct characteristics of the different cloze translation methods and provide insights on why combining them can lead to great improvements. Code/dataset is available at https://github.com/PlusLabNLP/zero_shot_cqa.
Zi-Yi Dou, Nanyun Peng 0001
AAAI1
2022 An Empirical Study of Training End-to-End Vision-and-Language Transformers
abstract
Vision-and-language (VL) pre-training has proven to be highly effective on various VL downstream tasks. While recent work has shown that fully transformer-based VL models can be more efficient than previous region-feature-based methods, their performance on downstream tasks often degrades significantly. In this paper, we present Meter, a Multimodal End-to-end TransformER framework, through which we investigate how to design and pre-train a fully transformer-based VL model in an end-to-end manner. Specifically, we dissect the model designs along multiple dimensions: vision encoders (e.g., CLIP-ViT, Swin transformer), text encoders (e.g., RoBERTa, De-BERTa), multimodal fusion module (e.g., merged attention vs. co-attention), architectural design (e.g., encoder-only vs. encoder-decoder), and pre-training objectives (e.g., masked image modeling). We conduct comprehensive experiments and provide insights on how to train a performant VL transformer. Meterachieves an accuracy of 77.64% on the VQAv2 test-std set using only 4M images for pre-training, surpassing the state-of-the-art region-feature-based model by 1.04%, and outperforming the previous best fully transformer-based model by 1.6%. Notably, when further scaled up, our best VQA model achieves an accuracy of 80.54%. Code and pre-trained models are released at https://github.com/zdou0830/METER.
Zi-Yi Dou, Yichong Xu, Zhe Gan, Shuohang Wang, Chenguang Zhu 0001, Pengchuan Zhang, Lu Yuan 0001, Nanyun Peng 0001, Zicheng Liu 0001, Michael Zeng 0001
CVPR1
2022 FOAM: A Follower-aware Speaker Model For Vision-and-Language Navigation
abstract
The speaker-follower models have proven to be effective in vision-and-language navigation, where a speaker model is used to synthesize new instructions to augment the training data for a follower navigation model.However, in many of the previous methods, the generated instructions are not directly trained to optimize the performance of the follower.In this paper, we present FOAM, a FOllower-Aware speaker Model that is constantly updated given the follower feedback, so that the generated instructions can be more suitable to the current learning state of the follower.Specifically, we optimize the speaker using a bi-level optimization framework and obtain its training signals by evaluating the follower on labeled data.Experimental results on the Room-to-Room and Room-across-Room datasets demonstrate that our methods can outperform strong baseline models across settings.Analyses also reveal that our generated instructions are of higher quality than the baselines.1
Zi-Yi Dou, Nanyun Peng 0001
NAACL-HLT1
2022 Coarse-to-Fine Vision-Language Pre-training with Fusion in the Backbone
abstract
Vision-language (VL) pre-training has recently received considerable attention. However, most existing end-to-end pre-training approaches either only aim to tackle VL tasks such as image-text retrieval, visual question answering (VQA) and image captioning that test high-level understanding of images, or only target region-level understanding for tasks such as phrase grounding and object detection. We present FIBER (Fusion-In-the-Backbone-based transformER), a new VL model architecture that can seamlessly handle both these types of tasks. Instead of having dedicated transformer layers for fusion after the uni-modal backbones, FIBER pushes multimodal fusion deep into the model by inserting cross-attention into the image and text backbones to better capture multimodal interactions. In addition, unlike previous work that is either only pre-trained on image-text data or on fine-grained data with box-level annotations, we present a two-stage pre-training strategy that uses both these kinds of data efficiently: (i) coarse-grained pre-training based on image-text data; followed by (ii) fine-grained pre-training based on image-text-box data. We conduct comprehensive experiments on a wide range of VL tasks, ranging from VQA, image captioning, and retrieval, to phrase grounding, referring expression comprehension, and object detection. Using deep multimodal fusion coupled with the two-stage pre-training, FIBER provides consistent performance improvements over strong baselines across all tasks, often outperforming methods using magnitudes more data. Code is released at https://github.com/microsoft/FIBER.
Zi-Yi Dou, Aishwarya Kamath, Zhe Gan, Pengchuan Zhang, Zicheng Liu 0001, Ce Liu 0001, Yann LeCun, Nanyun Peng 0001, Jianfeng Gao 0001
NeurIPS1
2021 Harnessing Social Media to Identify Homeless Youth At-Risk of Substance Use
abstract
Homeless youth are a highly vulnerable population and report highly elevated rates of substance use. Prior work on mitigating substance use among homeless youth has primarily relied on survey data to get information about substance use among homeless youth, which can then be used to inform the design of targeted intervention programs. However, such survey data is often onerous to collect, is limited by its reliance on self-reports and retrospective recall, and quickly becomes dated. The advent of social media has provided us with an important data source for understanding the health behaviors of homeless youth. In this paper, we target this specific population and demonstrate how to detect substance use based on texts from social media. We collect ~135K Facebook posts and comments together with survey responses from a group of homeless youth and use this data to build novel substance use detection systems with machine learning and natural language processing techniques. Experimental results show that our proposed methods achieve ROC-AUC scores of ~0.77 on identifying certain kinds of substance use among homeless youth using Facebook conversations only, and ROC-AUC scores of ~0.83 when combined with answers to four survey questions that are not about their demographic characteristics or substance use. Furthermore, we investigate connections between the characteristics of people's Facebook posts and substance use and provide insights about the problem.
Zi-Yi Dou, Anamika Barman-Adhikari, Fei Fang 0005, Amulya Yadav
AAAI1
2021 Word Alignment by Fine-tuning Embeddings on Parallel Corpora
abstract
Word alignment over parallel corpora has a wide variety of applications, including learning translation lexicons, cross-lingual transfer of language processing tools, and automatic evaluation or analysis of translation outputs.The great majority of past work on word alignment has worked by performing unsupervised learning on parallel text.Recently, however, other work has demonstrated that pre-trained contextualized word embeddings derived from multilingually trained language models (LMs) prove an attractive alternative, achieving competitive results on the word alignment task even in the absence of explicit training on parallel data.In this paper, we examine methods to marry the two approaches: leveraging pre-trained LMs but finetuning them on parallel text with objectives designed to improve alignment quality, and proposing methods to effectively extract alignments from these fine-tuned models.We perform experiments on five language pairs and demonstrate that our model can consistently outperform previous state-of-the-art models of all varieties.In addition, we demonstrate that we are able to train multilingual word aligners that can obtain robust performance on different language pairs.Our aligner, AWE-SOME (Aligning Word Embedding Spaces Of Multilingual Encoders), with pre-trained models is available at https://github. com/neulab/awesome-align.
Zi-Yi Dou, Graham Neubig
EACL1
2021 Improving Pre-trained Vision-and-Language Embeddings for Phrase Grounding
abstract
Phrase grounding aims to map textual phrases to their associated image regions, which can be a prerequisite for multimodal reasoning and can benefit tasks requiring identifying objects based on language.With pre-trained vision-and-language models achieving impressive performance across tasks, it remains unclear if we can directly utilize their learned embeddings for phrase grounding without finetuning.To this end, we propose a method to extract matched phrase-region pairs from pre-trained vision-and-language embeddings and propose four fine-tuning objectives to improve the model phrase grounding ability using image-caption data without any supervised grounding signals.Experiments on two representative datasets demonstrate the effectiveness of our objectives, outperforming baseline models in both weakly-supervised and supervised phrase grounding settings.In addition, we evaluate the aligned embeddings on several other downstream tasks and show that we can achieve better phrase grounding without sacrificing representation generality. 1
Zi-Yi Dou, Nanyun Peng 0001
EMNLP (1)1
2021 GSum: A General Framework for Guided Neural Abstractive Summarization
abstract
Zi-Yi Dou, Pengfei Liu, Hiroaki Hayashi, Zhengbao Jiang, Graham Neubig. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021.
Zi-Yi Dou, Pengfei Liu 0003, Hiroaki Hayashi, Zhengbao Jiang, Graham Neubig
NAACL-HLT1
2021 RefSum: Refactoring Neural Summarization
abstract
Although some recent works show potential complementarity among different state-of-theart systems, few works try to investigate this problem in text summarization.Researchers in other areas commonly refer to the techniques of reranking or stacking to approach this problem.In this work, we highlight several limitations of previous methods, which motivates us to present a new framework Refactor that provides a unified view of text summarization and summaries combination.Experimentally, we perform a comprehensive evaluation that involves twenty-two base systems, four datasets, and three different application scenarios.Besides new state-of-the-art results on CNN/DailyMail dataset (46.18 ROUGE-1), we also elaborate on how our proposed method addresses the limitations of the traditional methods and the effectiveness of the Refactor model sheds light on insight for performance improvement.Our system can be directly used by other researchers as an offthe-shelf tool to achieve further performance improvements.We open-source all the code and provide a convenient interface to use it:https://github.com/yixinL7/ Refactoring-Summarization.
Yixin Liu 0003, Zi-Yi Dou, Pengfei Liu 0003
NAACL-HLT2
2020 Dynamic Data Selection and Weighting for Iterative Back-Translation
abstract
Back-translation has proven to be an effective method to utilize monolingual data in neural machine translation (NMT), and iteratively conducting back-translation can further improve the model performance.Selecting which monolingual data to back-translate is crucial, as we require that the resulting synthetic data are of high quality and reflect the target domain.To achieve these two goals, data selection and weighting strategies have been proposed, with a common practice being to select samples close to the target domain but also dissimilar to the average general-domain text.In this paper, we provide insights into this commonly used approach and generalize it to a dynamic curriculum learning strategy, which is applied to iterative back-translation models.In addition, we propose weighting strategies based on both the current quality of the sentence and its improvement over the previous iteration.We evaluate our models on domain adaptation, low-resource, and high-resource MT settings and on two language pairs.Experimental results demonstrate that our methods achieve improvements of up to 1.8 BLEU points over competitive baselines.1
Zi-Yi Dou, Antonios Anastasopoulos, Graham Neubig
EMNLP (1)1
2020 Exploiting deep representations for natural language processing
Zi-Yi Dou, Xing Wang 0007, Shuming Shi 0001, Zhaopeng Tu
Neurocomputing1
2019 Dynamic Layer Aggregation for Neural Machine Translation with Routing-by-Agreement
abstract
With the promising progress of deep neural networks, layer aggregation has been used to fuse information across layers in various fields, such as computer vision and machine translation. However, most of the previous methods combine layers in a static fashion in that their aggregation strategy is independent of specific hidden states. Inspired by recent progress on capsule networks, in this paper we propose to use routing-by-agreement strategies to aggregate layers dynamically. Specifically, the algorithm learns the probability of a part (individual layer representations) assigned to a whole (aggregated representations) in an iterative way and combines parts accordingly. We implement our algorithm on top of the state-of-the-art neural machine translation model TRANSFORMER and conduct experiments on the widely-used WMT14 sh⇒German and WMT17 Chinese⇒English translation datasets. Experimental results across language pairs show that the proposed approach consistently outperforms the strong baseline model and a representative static aggregation model.
Zi-Yi Dou, Zhaopeng Tu, Xing Wang 0007, Longyue Wang, Shuming Shi 0001, Tong Zhang 0001
AAAI1
2019 Unsupervised Domain Adaptation for Neural Machine Translation with Domain-Aware Feature Embeddings
abstract
Zi-Yi Dou, Junjie Hu, Antonios Anastasopoulos, Graham Neubig. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Zi-Yi Dou, Junjie Hu 0001, Antonios Anastasopoulos, Graham Neubig
EMNLP/IJCNLP (1)1
2019 Investigating Meta-Learning Algorithms for Low-Resource Natural Language Understanding Tasks
abstract
Zi-Yi Dou, Keyi Yu, Antonios Anastasopoulos. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Zi-Yi Dou, Keyi Yu, Antonios Anastasopoulos
EMNLP/IJCNLP (1)1
2018 SkipNet: Learning Dynamic Routing in Convolutional Networks
Xin Wang 0066, Fisher Yu 0001, Zi-Yi Dou, Trevor Darrell, Joseph Gonzalez 0001
ECCV (13)3
2018 Exploiting Deep Representations for Neural Machine Translation
abstract
Advanced neural machine translation (NMT) models generally implement encoder and decoder as multiple layers, which allows systems to model complex functions and capture complicated linguistic structures.However, only the top layers of encoder and decoder are leveraged in the subsequent process, which misses the opportunity to exploit the useful information embedded in other layers.In this work, we propose to simultaneously expose all of these signals with layer aggregation and multi-layer attention mechanisms.In addition, we introduce an auxiliary regularization term to encourage different layers to capture diverse information.Experimental results on widely-used WMT14 English⇒German and WMT17 Chinese⇒English translation data demonstrate the effectiveness and universality of the proposed approach.
Zi-Yi Dou, Zhaopeng Tu, Xing Wang 0007, Shuming Shi 0001, Tong Zhang 0001
EMNLP1
2018 Unsupervised Bilingual Lexicon Induction via Latent Variable Models
abstract
Bilingual lexicon extraction has been studied for decades and most previous methods have relied on parallel corpora or bilingual dictionaries.Recent studies have shown that it is possible to build a bilingual dictionary by aligning monolingual word embedding spaces in an unsupervised way.With the recent advances in generative models, we propose a novel approach which builds cross-lingual dictionaries via latent variable models and adversarial training with no parallel corpora.To demonstrate the effectiveness of our approach, we evaluate our approach on several language pairs and the experimental results show that our model could achieve competitive and even superior performance compared with several state-of-the-art models.
Zi-Yi Dou, Zhi-Hao Zhou, Shujian Huang
EMNLP1
2018 Dynamic Oracle for Neural Machine Translation in Decoding Phase
Zi-Yi Dou, Hao Zhou 0012, Shujian Huang, Xinyu Dai, Jiajun Chen 0001
LREC1
2017 Capturing User and Product Information for Document Level Sentiment Analysis with Deep Memory Network
abstract
Document-level sentiment classification is a fundamental problem which aims to predict a user's overall sentiment about a product in a document.Several methods have been proposed to tackle the problem whereas most of them fail to consider the influence of users who express the sentiment and products which are evaluated.To address the issue, we propose a deep memory network for document-level sentiment classification which could capture the user and product information at the same time.To prove the effectiveness of our algorithm, we conduct experiments on IMDB and Yelp datasets and the results indicate that our model can achieve better performance than several existing methods.
Zi-Yi Dou
EMNLP1
2017 Compressing Neural Networks by Applying Frequent Item-Set Mining
Zi-Yi Dou, Shujian Huang, Yi-Fan Su
ICANN (2)1