VLDB 2026 Research / reviewers in the wild / expert
Nan Duan 0001
dblp:30/8160-1
· DBLP profile ↗
146ranked-venue papers
11as first author
79since 2021 · last 2026
0000-0002-3387-4674ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 135 · 11 first-author · 70 since 2021Graphics, computer vision, multimedia, augmented reality and games · 29 · 17 since 2021Databases, data management, data science and information retrieval · 7 · 4 since 2021Software engineering, systems software and programming languages · 3 · 3 since 2021Systems, architecture and hardware · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Dynamic Sparsity in Large-Scale Video DiT TrainingabstractDiffusion Transformers (DiTs) have shown remarkable performance in generating high-quality videos. However, the quadratic complexity of 3D full attention remains a bottleneck in scaling DiT training, especially with high-definition, lengthy videos, where it can consume up to 95% of processing time and demand specialized context parallelism. Xin Tan 0004, Yuetao Chen, Xing Chen 0009, Kun Yan 0004, Nan Duan 0001, Yibo Zhu 0001, Daxin Jiang, Hong Xu 0001 |
ASPLOS (1) | 6 |
| 2025 | Key-Point-Driven Data Synthesis with Its Enhancement on Mathematical ReasoningabstractLarge language models have shown great potential in complex reasoning tasks, yet their performance is often hampered by the scarcity of high-quality and reasoning-focused training datasets. Addressing this challenge, we propose Key-PointDriven Data Synthesis (KPDDS), a novel data synthesis framework that synthesizes question-answer pairs by leveraging key points and exemplar practices from authentic data sources. KPDDS ensures the generation of novel questions with rigorous quality control and substantial scalability. As a result, we present KPMath, an extensive synthetic dataset tailored for mathematical reasoning, comprising over 800K questionanswer pairs. Utilizing KPMath and augmenting it with additional reasoning-intensive corpora, we create the comprehensive KPMath-Plus dataset. Our experiments demonstrate that this dataset can enhance the mathematical reasoning performance of models across various architectures and sizes. The Qwen1.5-72B model, fine-tuned on KPMath-Plus, achieves 87.0% accuracy on GSM8K and 58.3% on MATH, surpassing competitors in the 7B to 72B range and best commercial models like GPT-4 across multiple math reasoning datasets. Xiao Liu 0029, Yeyun Gong, Zhibin Gou, Yelong Shen, Nan Duan 0001, Weizhu Chen |
AAAI | 6 |
| 2025 | Taming Teacher Forcing for Masked Autoregressive Video GenerationabstractWe introduce MAGI, a hybrid video generation framework that combines masked modeling for intra-frame generation with causal modeling for next-frame generation. Our key innovation, Complete Teacher Forcing (CTF), conditions masked frames on complete observation frames rather than masked ones (namely Masked Teacher Forcing, MTF), enabling a smooth transition from token-level (patch-level) to frame-level autoregressive generation. CTF significantly outperforms MTF, achieving a +23% improvement in FVD scores on first-frame conditioned video prediction. To address issues like exposure bias, we employ targeted training strategies, setting a new benchmark in autoregressive video generation. Experiments show that MAGI can generate long, coherent video sequences exceeding 100 frames, even when trained on as few as 16 frames, highlighting its potential for scalable, high-quality video generation. Yuang Peng, Kun Yan 0004, Runpei Dong, Duomin Wang, Zheng Ge, Nan Duan 0001, Xiangyu Zhang 0005 |
CVPR | 8 |
| 2025 | Automated Proof Generation for Rust Code via Self-EvolutionabstractEnsuring correctness is crucial for code generation. Formal verification offers a
definitive assurance of correctness, but demands substantial human effort in proof
construction and hence raises a pressing need for automation. The primary obsta-
cle lies in the severe lack of data—there is much fewer proofs than code snippets
for Large Language Models (LLMs) to train upon. In this paper, we introduce
SAFE, a framework that overcomes the lack of human-written proofs to enable
automated proof generation of Rust code. SAFE establishes a self-evolving cycle
where data synthesis and fine-tuning collaborate to enhance the model capability,
leveraging the definitive power of a symbolic verifier in telling correct proofs from
incorrect ones. SAFE also re-purposes the large number of synthesized incorrect
proofs to train the self-debugging capability of the fine-tuned models, empowering
them to fix incorrect proofs based on the verifier’s feedback. SAFE demonstrates
superior efficiency and precision compared to GPT-4o. Through tens of thousands
of synthesized proofs and the self-debugging mechanism, we improve the capa-
bility of open-source models, initially unacquainted with formal verification, to
automatically write proofs for Rust code. This advancement leads to a signifi-
cant improvement in performance, achieving a 52.52% accuracy rate in a bench-
mark crafted by human experts, a significant leap over GPT-4o’s performance of
14.39%. Tianyu Chen 0006, Shan Lu 0001, Yeyun Gong, Chenyuan Yang, Xuheng Li, Md Rakib Hossain Misu, Hao Yu 0016, Nan Duan 0001, Peng Cheng 0005, Fan Yang 0024, Shuvendu K. Lahiri, Tao Xie 0001, Lidong Zhou |
ICLR | 9 |
| 2025 | Alchemy: Amplifying Theorem-Proving Capability Through Symbolic MutationabstractFormal proofs are challenging to write even for experienced experts. Recent progress in Neural Theorem Proving (NTP) shows promise in expediting this process. However, the formal corpora available on the Internet are limited compared to the general text, posing a significant data scarcity challenge for NTP. To address this issue, this work proposes Alchemy, a general framework for data synthesis that constructs formal theorems through symbolic mutation. Specifically, for each candidate theorem in Mathlib, we identify all invocable theorems that can be used to rewrite or apply to it. Subsequently, we mutate the candidate theorem by replacing the corresponding term in the statement with its equivalent form or antecedent. As a result, our method increases the number of theorems in Mathlib by an order of magnitude, from 110k to 6M. Furthermore, we perform continual pretraining and supervised finetuning on this augmented corpus for large language models. Experimental results demonstrate the effectiveness of our approach, achieving a 4.70% absolute performance improvement on Leandojo benchmark. Additionally, our approach achieves a 2.47% absolute performance gain on the out-of-distribution miniF2F benchmark based on the synthetic data. To provide further insights, we conduct a comprehensive analysis of synthetic data composition and the training paradigm, offering valuable guidance for developing a strong theorem prover. Shaonan Wu, Yeyun Gong, Nan Duan 0001, Ping Wei 0001 |
ICLR | 4 |
| 2025 | Generative Pre-trained Autoregressive Diffusion TransformerabstractIn this work, we present GPDiT, a Generative Pre-trained Autoregressive Diffusion Transformer that unifies the strengths of diffusion and autoregressive modeling for long-range video synthesis, within a continuous latent space. Instead of predicting discrete tokens, GPDiT autoregressively predicts future latent frames using a diffusion loss, enabling natural modeling of motion dynamics and semantic consistency across frames. This continuous autoregressive framework not only enhances generation quality but also endows the model with representation capabilities. Additionally, we introduce a lightweight causal attention variant and a parameter-free rotation-based time-conditioning mechanism, improving both the training and inference efficiency. Extensive experiments demonstrate that GPDiT achieves strong performance in video generation quality, video representation ability, and few-shot learning tasks, highlighting its potential as an effective framework for video modeling in continuous space. Yuan Zhang 0022, Zhiying Lu, Haoyang Huang, Jianlong Yuan, Nan Duan 0001 |
NeurIPS | 8 |
| 2024 | ORES: Open-Vocabulary Responsible Visual SynthesisabstractAvoiding synthesizing specific visual concepts is an essential challenge in responsible visual synthesis. However, the visual concept that needs to be avoided for responsible visual synthesis tends to be diverse, depending on the region, context, and usage scenarios. In this work, we formalize a new task, Open-vocabulary Responsible Visual Synthesis (ORES), where the synthesis model is able to avoid forbidden visual concepts while allowing users to input any desired content. To address this problem, we present a Two-stage Intervention (TIN) framework. By introducing 1) rewriting with learnable instruction through a large-scale language model (LLM) and 2) synthesizing with prompt intervention on a diffusion synthesis model, it can effectively synthesize images avoiding any concepts but following the user's query as much as possible. To evaluate on ORES, we provide a publicly available dataset, baseline models, and benchmark. Experimental results demonstrate the effectiveness of our method in reducing risks of image generation. Our work highlights the potential of LLMs in responsible visual synthesis. Our code and dataset is public available in https://github.com/kodenii/ORES. Minheng Ni, Chenfei Wu, Xiaodong Wang 0023, Shengming Yin, Zicheng Liu 0001, Nan Duan 0001 |
AAAI | 7 |
| 2024 | HORIZON: High-Resolution Semantically Controlled Panorama SynthesisabstractPanorama synthesis endeavors to craft captivating 360-degree visual landscapes, immersing users in the heart of virtual worlds. Nevertheless, contemporary panoramic synthesis techniques grapple with the challenge of semantically guiding the content generation process. Although recent breakthroughs in visual synthesis have unlocked the potential for semantic control in 2D flat images, a direct application of these methods to panorama synthesis yields distorted content. In this study, we unveil an innovative framework for generating high-resolution panoramas, adeptly addressing the issues of spherical distortion and edge discontinuity through sophisticated spherical modeling. Our pioneering approach empowers users with semantic control, harnessing both image and text inputs, while concurrently streamlining the generation of high-resolution panoramas using parallel decoding. We rigorously evaluate our methodology on a diverse array of indoor and outdoor datasets, establishing its superiority over recent related work, in terms of both quantitative and qualitative performance metrics. Our research elevates the controllability, efficiency, and fidelity of panorama synthesis to new levels. Kun Yan 0004, Lei Ji 0001, Chenfei Wu, Ming Zhou 0001, Nan Duan 0001, Shuai Ma 0001 |
AAAI | 6 |
| 2024 | PROM: A Phrase-level Copying Mechanism with Pre-training for Abstractive SummarizationabstractBased on the remarkable achievements of pre-trained language models in abstractive summarization, the copying mechanism has proved helpful by improving the factuality, stability, and overall performance. This work proposes PROM, a new PhRase-level cOpying Mechanism that enhances attention on n-grams, which can be applied to zero-shot summarization with pre-training. PROM adds an indicator layer to explicitly pick up tokens in n-gram that can be copied from the source, and calculates an auxiliary loss for the copying prediction. Empirical studies show that PROM makes significant improvements in fine-tuning on benchmarks. In the zero-shot setting, PROM is utilized in the self-supervised pre-training on raw corpora and provides new general baselines on a wide range of summarization datasets. Further analysis shows that PROM performs more reasonable copying and contributes to faithfulness. Our code is publicly available at https://github.com/xbmxb/PROM. Xinbei Ma, Yeyun Gong, Hai Zhao 0001, Nan Duan 0001 |
LREC/COLING | 5 |
| 2024 | CRITIC: Large Language Models Can Self-Correct with Tool-Interactive CritiquingabstractRecent developments in large language models (LLMs) have been impressive. However, these models sometimes show inconsistencies and problematic behavior, such as hallucinating facts, generating flawed code, or creating offensive and toxic content. Unlike these models, humans typically utilize external tools to cross-check and refine their initial content, like using a search engine for fact-checking, or a code interpreter for debugging. Inspired by this observation, we introduce a framework called CRITIC that allows LLMs, which are essentially “black boxes” to validate and progressively amend their own outputs in a manner similar to human interaction with tools. More specifically, starting with an initial output, CRITIC interacts with appropriate tools to evaluate certain aspects of the text, and then revises the output based on the feedback obtained during this validation process. Comprehensive evaluations involving free-form question answering, mathematical program synthesis, and toxicity reduction demonstrate that CRITIC consistently enhances the performance of LLMs. Meanwhile, our research highlights the crucial importance of external feedback in promoting the ongoing self-improvement of LLMs. Zhibin Gou, Zhihong Shao, Yeyun Gong, Yelong Shen, Yujiu Yang 0001, Nan Duan 0001, Weizhu Chen |
ICLR | 6 |
| 2024 | ToRA: A Tool-Integrated Reasoning Agent for Mathematical Problem SolvingabstractLarge language models have made significant progress in various language tasks, yet they still struggle with complex mathematics. In this paper, we propose ToRA a series of Tool-integrated Reasoning Agents designed to solve challenging mathematical problems by seamlessly integrating natural language reasoning with the utilization of external tools (e.g., computation libraries and symbolic solvers), thereby amalgamating the analytical prowess of language and the computational efficiency of tools. To train ToRA, we curate interactive tool-use trajectories on mathematical datasets, apply imitation learning on the annotations, and propose output space shaping to further refine models' reasoning behavior. As a result, ToRA models significantly outperform open-source models on 10 mathematical reasoning datasets across all scales with 13%-19% absolute improvements on average. Notably, ToRA-7B reaches 44.6% on the competition-level dataset MATH, surpassing the best open-source model WizardMath-70B by 22% absolute. ToRA-34B is also the first open-source model that achieves an accuracy exceeding 50% on MATH, which significantly outperforms GPT-4's CoT result, and is competitive with GPT-4 solving problems with programs. Additionally, we conduct a comprehensive analysis of the benefits and remaining challenges of tool interaction for mathematical reasoning, providing valuable insights for future research. Zhibin Gou, Zhihong Shao, Yeyun Gong, Yelong Shen, Yujiu Yang 0001, Minlie Huang, Nan Duan 0001, Weizhu Chen |
ICLR | 7 |
| 2024 | LayoutNUWA: Revealing the Hidden Layout Expertise of Large Language ModelsabstractGraphic layout generation, a growing research field, plays a significant role in user engagement and information perception.
Existing methods primarily treat layout generation as a numerical optimization task, focusing on quantitative aspects while overlooking the semantic information of layout, such as the relationship between each layout element.
In this paper, we propose LayoutNUWA, the first model that treats layout generation as a code generation task to enhance semantic information and harness the hidden layout expertise of large language models~(LLMs).
Concretely, we develop a Code Instruct Tuning (CIT) approach comprising three interconnected modules: 1) the Code Initialization (CI) module quantifies the numerical conditions and initializes them as HTML code with strategically placed masks; 2) the Code Completion (CC) module employs the formatting knowledge of LLMs to fill in the masked portions within the HTML code; 3) the Code Rendering (CR) module transforms the completed code into the final layout output, ensuring a highly interpretable and transparent layout generation procedure that directly maps code to a visualized layout. We attain significant state-of-the-art performance (even over 50\% improvements compared to previous works) on multiple datasets, showcasing the strong capabilities of LayoutNUWA. Zecheng Tang, Chenfei Wu, Nan Duan 0001 |
ICLR | 4 |
| 2024 | Using Left and Right Brains Together: Towards Vision and Language PlanningabstractLarge Language Models (LLMs) and Large Multi-modality Models (LMMs) have demonstrated remarkable decision masking capabilities on a variety of tasks. However, they inherently operate planning within the language space, lacking the vision and spatial imagination ability. In contrast, humans utilize both left and right hemispheres of the brain for language and visual planning during the thinking process. Therefore, we introduce a novel vision-language planning framework in this work to perform concurrent visual and language planning for tasks with inputs of any form. Our framework incorporates visual planning to capture intricate environmental details, while language planning enhances the logical coherence of the overall system. We evaluate the effectiveness of our framework across vision-language tasks, vision-only tasks, and language-only tasks. The results demonstrate the superior performance of our approach, indicating that the integration of visual and language planning yields better contextually aware task execution. Jun Cen, Chenfei Wu, Xiao Liu 0029, Shengming Yin, Yixuan Pei, Jinglong Yang, Qifeng Chen 0001, Nan Duan 0001 |
ICML | 8 |
| 2024 | StrokeNUWA - Tokenizing Strokes for Vector Graphic SynthesisabstractTo leverage LLMs for visual synthesis, traditional methods convert raster image information into discrete grid tokens through specialized visual modules, while disrupting the model’s ability to capture the true semantic representation of visual scenes. This paper posits that an alternative representation of images, vector graphics, can effectively surmount this limitation by enabling a more natural and semantically coherent segmentation of the image information. Thus, we introduce StrokeNUWA, a pioneering work exploring a better visual representation "stroke" tokens on vector graphics, which is inherently visual semantics rich, naturally compatible with LLMs, and highly compressed. Equipped with stroke tokens, StrokeNUWA can significantly surpass traditional LLM-based and optimization-based methods across various metrics in the vector graphic generation task. Besides, StrokeNUWA achieves up to a $94\times$ speedup in inference over the speed of prior methods with an exceptional SVG code compression ratio of 6.9%. Zecheng Tang, Chenfei Wu, Minheng Ni, Shengming Yin, Zhengyuan Yang, Zicheng Liu 0001, Nan Duan 0001 |
ICML | 11 |
| 2024 | Contextualized Data-Wrangling Code Generation in Computational NotebooksabstractData wrangling, the process of preparing raw data for further analysis in computational notebooks, is a crucial yet time-consuming step in data science. Code generation has the potential to automate the data wrangling process to reduce analysts' overhead by translating user intents into executable code. Precisely generating data wrangling code necessitates a comprehensive consideration of the rich context present in notebooks, including textual context, code context and data context. However, notebooks often interleave multiple non-linear analysis tasks into linear sequence of code blocks, where the contextual dependencies are not clearly reflected. Directly training models with source code blocks fails to fully exploit the contexts for accurate wrangling code generation. Junjie Huang 0008, Daya Guo, Chenglong Wang 0005, Jiazhen Gu, Jeevana Priya Inala, Cong Yan, Jianfeng Gao 0001, Nan Duan 0001, Michael R. Lyu |
ASE | 9 |
| 2024 | Not All Tokens Are What You Need for PretrainingabstractPrevious language model pre-training methods have uniformly applied a next-token prediction loss to all training tokens. Challenging this norm, we posit that ''Not all tokens in a corpus are equally important for language model training''. Our initial analysis examines token-level training dynamics of language model, revealing distinct loss patterns for different tokens. Leveraging these insights, we introduce a new language model called Rho-1. Unlike traditional LMs that learn to predict every next token in a corpus, Rho-1 employs Selective Language Modeling (SLM), which selectively trains on useful tokens that aligned with the desired distribution. This approach involves scoring training tokens using a reference model, and then training the language model with a focused loss on tokens with higher scores. When continual continual pretraining on 15B OpenWebMath corpus, Rho-1 yields an absolute improvement in few-shot accuracy of up to 30% in 9 math tasks. After fine-tuning, Rho-1-1B and 7B achieved state-of-the-art results of 40.6% and 51.8% on MATH dataset, respectively - matching DeepSeekMath with only 3% of the pretraining tokens. Furthermore, when continual pretraining on 80B general tokens, Rho-1 achieves 6.8% average enhancement across 15 diverse tasks, increasing both data efficiency and performance of the language model pre-training. Zhenghao Lin, Zhibin Gou, Yeyun Gong, Xiao Liu 0029, Yelong Shen, Ruochen Xu, Chen Lin 0001, Yujiu Yang 0001, Jian Jiao 0007, Nan Duan 0001, Weizhu Chen |
NeurIPS | 10 |
| 2024 | Voila-A: Aligning Vision-Language Models with User's Gaze AttentionabstractIn recent years, the integration of vision and language understanding has led to significant advancements in artificial intelligence, particularly through Vision-Language Models (VLMs). However, existing VLMs face challenges in handling real-world applications with complex scenes and multiple objects, as well as aligning their focus with the diverse attention patterns of human users. In this paper, we introduce gaze information, feasibly collected by ubiquitous wearable devices such as MR glasses, as a proxy for human attention to guide VLMs. We propose a novel approach, Voila-A, for gaze alignment to enhance the effectiveness of these models in real-world applications. First, we collect hundreds of minutes of gaze data to demonstrate that we can mimic human gaze modalities using localized narratives. We then design an automatic data annotation pipeline utilizing GPT-4 to generate the VOILA-COCO dataset. Additionally, we introduce a new model VOILA-A that integrate gaze information into VLMs while maintain pretrained knowledge from webscale dataset. We evaluate Voila-A using a hold-out validation set and a newly collected VOILA-GAZE testset, which features real-life scenarios captured with a gaze-tracking device. Our experimental results demonstrate that Voila-A significantly outperforms several baseline models. By aligning model attention with human gaze patterns, Voila-A paves the way for more intuitive, user-centric VLMs and fosters engaging human-AI interaction across a wide range of applications. Kun Yan 0004, Lei Ji 0001, Yuntao Wang 0001, Nan Duan 0001, Shuai Ma 0001 |
NeurIPS | 5 |
| 2024 | LEAD: Liberal Feature-based Distillation for Dense RetrievalabstractKnowledge distillation is often used to transfer knowledge from a strong teacher model to a relatively weak student model. Traditional methods include response-based methods and feature-based methods. Response-based methods are widely used but suffer from lower upper limits of performance due to their ignorance of intermediate signals, while feature-based methods have constraints on vocabularies, tokenizers and model architectures. In this paper, we propose a liberal feature-based distillation method (LEAD). LEAD aligns the distribution between the intermediate layers of teacher model and student model, which is effective, extendable, portable and has no requirements on vocabularies, tokenizers, or model architectures. Extensive experiments show the effectiveness of LEAD on widely-used benchmarks, including MS MARCO Passage Ranking, TREC 2019 DL Track, MS MARCO Document Ranking and TREC 2020 DL Track. Our code is available in https://github.com/microsoft/SimXNS/tree/main/LEAD. Hao Sun 0015, Xiao Liu 0029, Yeyun Gong, Anlei Dong, Jingwen Lu, Yan Zhang 0117, Linjun Yang, Rangan Majumder, Nan Duan 0001 |
WSDM | 9 |
| 2023 | BridgeTower: Building Bridges between Encoders in Vision-Language Representation LearningabstractVision-Language (VL) models with the Two-Tower architecture have dominated visual-language representation learning in recent years. Current VL models either use lightweight uni-modal encoders and learn to extract, align and fuse both modalities simultaneously in a deep cross-modal encoder, or feed the last-layer uni-modal representations from the deep pre-trained uni-modal encoders into the top cross-modal encoder. Both approaches potentially restrict vision-language representation learning and limit model performance. In this paper, we propose BridgeTower, which introduces multiple bridge layers that build a connection between the top layers of uni-modal encoders and each layer of the cross-modal encoder. This enables effective bottom-up cross-modal alignment and fusion between visual and textual representations of different semantic levels of pre-trained uni-modal encoders in the cross-modal encoder. Pre-trained with only 4M images, BridgeTower achieves state-of-the-art performance on various downstream vision-language tasks. In particular, on the VQAv2 test-std set, BridgeTower achieves an accuracy of 78.73%, outperforming the previous state-of-the-art model METER by 1.09% with the same pre-training data and almost negligible additional parameters and computational costs. Notably, when further scaling the model, BridgeTower achieves an accuracy of 81.15%, surpassing models that are pre-trained on orders-of-magnitude larger datasets. Code and checkpoints are available at https://github.com/microsoft/BridgeTower. Xiao Xu 0005, Chenfei Wu, Shachar Rosenman, Vasudev Lal, Wanxiang Che, Nan Duan 0001 |
AAAI | 6 |
| 2023 | ManagerTower: Aggregating the Insights of Uni-Modal Experts for Vision-Language Representation LearningabstractXiao Xu, Bei Li, Chenfei Wu, Shao-Yen Tseng, Anahita Bhiwandiwalla, Shachar Rosenman, Vasudev Lal, Wanxiang Che, Nan Duan. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Xiao Xu 0005, Chenfei Wu, Shao-Yen Tseng, Anahita Bhiwandiwalla, Shachar Rosenman, Vasudev Lal, Wanxiang Che, Nan Duan 0001 |
ACL (1) | 9 |
| 2023 | Analyzing and Reducing the Performance Gap in Cross-Lingual Transfer with Fine-tuning Slow and FastabstractExisting research has shown that a multilingual pre-trained language model fine-tuned with one (source) language also performs well on downstream tasks for non-source languages, even though no fine-tuning is done on these languages.However, there is a clear gap between the performance of the source language and that of the non-source languages.This paper analyzes the fine-tuning process, discovers when the performance gap changes and identifies which network weights affect the overall performance most.Additionally, the paper seeks to answer to what extent the gap can be reduced by reducing forgetting.Based on the analysis results, a method named Fine-tuning slow and fast with four training policies is proposed to address these issues.Experimental results show the proposed method outperforms baselines by a clear margin. Yiduo Guo, Yaobo Liang, Dongyan Zhao 0001, Bing Liu 0001, Nan Duan 0001 |
ACL (1) | 5 |
| 2023 | CONE: An Efficient COarse-to-fiNE Alignment Framework for Long Video Temporal GroundingabstractZhijian Hou, Wanjun Zhong, Lei Ji, Difei Gao, Kun Yan, W.k. Chan, Chong-Wah Ngo, Mike Zheng Shou, Nan Duan. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Zhijian Hou, Wanjun Zhong, Lei Ji 0001, Difei Gao, Kun Yan 0004, Wing Kwong Chan, Chong-Wah Ngo, Zheng Shou 0001, Nan Duan 0001 |
ACL (1) | 9 |
| 2023 | NUWA-XL: Diffusion over Diffusion for eXtremely Long Video GenerationabstractShengming Yin, Chenfei Wu, Huan Yang, Jianfeng Wang, Xiaodong Wang, Minheng Ni, Zhengyuan Yang, Linjie Li, Shuguang Liu, Fan Yang, Jianlong Fu, Ming Gong, Lijuan Wang, Zicheng Liu, Houqiang Li, Nan Duan. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Shengming Yin, Chenfei Wu, Huan Yang 0005, Xiaodong Wang 0023, Minheng Ni, Zhengyuan Yang, Fan Yang 0024, Jianlong Fu, Ming Gong 0001, Zicheng Liu 0001, Houqiang Li, Nan Duan 0001 |
ACL (1) | 16 |
| 2023 | ReCo: Region-Controlled Text-to-Image GenerationabstractRecently, large-scale text-to-image (T2I) models have shown impressive performance in generating high-fidelity images, but with limited controllability, e.g., precisely specifying the content in a specific region with a free-form text description. In this paper, we propose an effective technique for such regional control in T2I generation. We augment T2I models' inputs with an extra set of position tokens, which represent the quantized spatial coordinates. Each region is specified by four position tokens to represent the top-left and bottom-right corners, followed by an open-ended natural language regional description. Then, we fine-tune a pre-trained T2I model with such new input interface. Our model, dubbed as ReCo (Region-Controlled T2I), enables the region control for arbitrary objects described by open-ended regional texts rather than by object labels from a constrained category set. Empirically, ReCo achieves better image quality than the T2I model strengthened by positional words (FID: 8.82 → 7.36, SceneFID: 15.54 → 6.51 on COCO), together with objects being more accurately placed, amounting to a 20.40% region classification accuracy improvement on COCO. Furthermore, we demonstrate that ReCo can better control the object count, spatial relationship, and region attributes such as color/size, with the free-form regional description. Human evaluation on PaintSkill shows that ReCo is +19.28% and +17.21% more accurate in generating images with correct object count and spatial relationship than the T2I model. Code is available at https://github.com/microsoft/Reeo. Zhengyuan Yang, Zhe Gan, Chenfei Wu, Nan Duan 0001, Zicheng Liu 0001, Ce Liu 0001, Michael Zeng 0001 |
CVPR | 7 |
| 2023 | CAPSTONE: Curriculum Sampling for Dense Retrieval with Document ExpansionabstractThe dual-encoder has become the de facto architecture for dense retrieval.Typically, it computes the latent representations of the query and document independently, thus failing to fully capture the interactions between the query and document.To alleviate this, recent research has focused on obtaining query-informed document representations.During training, it expands the document with a real query, but during inference, it replaces the real query with a generated one.This inconsistency between training and inference causes the dense retrieval model to prioritize query information while disregarding the document when computing the document representation.Consequently, it performs even worse than the vanilla dense retrieval model because its performance heavily relies on the relevance between the generated queries and the real query.In this paper, we propose a curriculum sampling strategy that utilizes pseudo queries during training and progressively enhances the relevance between the generated query and the real query.By doing so, the retrieval model learns to extend its attention from the document alone to both the document and query, resulting in high-quality queryinformed document representations.Experimental results on both in-domain and out-ofdomain datasets demonstrate that our approach outperforms previous dense retrieval models. Xingwei He 0003, Yeyun Gong, A-Long Jin, Hang Zhang 0029, Anlei Dong, Jian Jiao 0007, Siu-Ming Yiu, Nan Duan 0001 |
EMNLP | 8 |
| 2023 | Query Rewriting in Retrieval-Augmented Large Language ModelsabstractLarge Language Models (LLMs) play powerful, black-box readers in the retrieve-thenread pipeline, making remarkable progress in knowledge-intensive tasks.This work introduces a new framework, Rewrite-Retrieve-Read instead of the previous retrieve-then-read for the retrieval-augmented LLMs from the perspective of the query rewriting.Unlike prior studies focusing on adapting either the retriever or the reader, our approach pays attention to the adaptation of the search query itself, for there is inevitably a gap between the input text and the needed knowledge in retrieval.We first prompt an LLM to generate the query, then use a web search engine to retrieve contexts.Furthermore, to better align the query to the frozen modules, we propose a trainable scheme for our pipeline.A small language model is adopted as a trainable rewriter to cater to the black-box LLM reader.The rewriter is trained using the feedback of the LLM reader by reinforcement learning.Evaluation is conducted on downstream tasks, open-domain QA and multiple-choice QA.Experiments results show consistent performance improvement, indicating that our framework is proven effective and scalable, and brings a new framework for retrieval-augmented LLM 1 . Xinbei Ma, Yeyun Gong, Hai Zhao 0001, Nan Duan 0001 |
EMNLP | 5 |
| 2023 | Baize: An Open-Source Chat Model with Parameter-Efficient Tuning on Self-Chat DataabstractChat models, such as ChatGPT, have shown impressive capabilities and have been rapidly adopted across numerous domains.However, these models are only accessible through a restricted API, creating barriers for new research and progress in the field.We propose a pipeline that can automatically generate a highquality multi-turn chat corpus by leveraging ChatGPT to engage in a conversation with itself.Subsequently, we employ parameter-efficient tuning to enhance LLaMA, an open-source large language model.The resulting model, named Baize, demonstrates good performance in multi-turn dialogues with guardrails that minimize potential risks.Additionally, we propose a new technique called Self-Distill with Feedback, to further improve the performance of the Baize models with feedback from ChatGPT.The Baize models and data are released for research purposes only. 1 Canwen Xu, Daya Guo, Nan Duan 0001, Julian J. McAuley |
EMNLP | 3 |
| 2023 | Modeling Sequential Sentence Relation to Improve Cross-lingual Dense Retrieval
Shunyu Zhang, Yaobo Liang, Ming Gong 0001, Daxin Jiang, Nan Duan 0001 |
ICLR | 5 |
| 2023 | LongCoder: A Long-Range Pre-trained Language Model for Code CompletionabstractIn this paper, we introduce a new task for code completion that focuses on handling long code input and propose a sparse Transformer model, called LongCoder, to address this task. LongCoder employs a sliding window mechanism for self-attention and introduces two types of globally accessible tokens - bridge tokens and memory tokens - to improve performance and efficiency. Bridge tokens are inserted throughout the input sequence to aggregate local information and facilitate global interaction, while memory tokens are included to highlight important statements that may be invoked later and need to be memorized, such as package imports and definitions of classes, functions, or structures. We conduct experiments on a newly constructed dataset that contains longer code context and the publicly available CodeXGLUE benchmark. Experimental results demonstrate that LongCoder achieves superior performance on code completion tasks compared to previous models while maintaining comparable efficiency in terms of computational resources during inference. Daya Guo, Canwen Xu, Nan Duan 0001, Jian Yin 0001, Julian J. McAuley |
ICML | 3 |
| 2023 | Text Generation with Diffusion Language Models: A Pre-training Approach with Continuous Paragraph DenoiseabstractIn this paper, we introduce a novel dIffusion language modEl pre-training framework for text generation, which we call GENIE. GENIE is a large-scale pre-trained diffusion language model that consists of an encoder and a diffusion-based decoder, which can generate text by gradually transforming a random noise sequence into a coherent text sequence. To pre-train GENIE on a large-scale language corpus, we design a new continuous paragraph denoise objective, which encourages the diffusion-decoder to reconstruct a clean text paragraph from a corrupted version, while preserving the semantic and syntactic coherence. We evaluate GENIE on four downstream text generation benchmarks, namely XSum, CNN/DailyMail, Gigaword, and CommonGen. Our experimental results show that GENIE achieves comparable performance with the state-of-the-art autoregressive models on these benchmarks, and generates more diverse text samples. The code and models of GENIE are available at https://github.com/microsoft/ProphetNet/tree/master/GENIE. Zhenghao Lin, Yeyun Gong, Yelong Shen, Zhihao Fan, Chen Lin 0001, Nan Duan 0001, Weizhu Chen |
ICML | 7 |
| 2023 | Synthetic Prompting: Generating Chain-of-Thought Demonstrations for Large Language ModelsabstractLarge language models can perform various reasoning tasks by using chain-of-thought prompting, which guides them to find answers through step-by-step demonstrations. However, the quality of the prompts depends on the demonstrations given to the models, and creating many of them by hand is costly. We introduce Synthetic prompting, a method that leverages a few handcrafted examples to prompt the model to generate more examples by itself, and selects effective demonstrations to elicit better reasoning. Our method alternates between a backward and forward process to generate new examples. The backward process generates a question that match a sampled reasoning chain, so that the question is solvable and clear. The forward process produces a more detailed reasoning chain for the question, improving the quality of the example. We evaluate our method on numerical, symbolic, and algorithmic reasoning tasks, and show that it outperforms existing prompting techniques. Zhihong Shao, Yeyun Gong, Yelong Shen, Minlie Huang, Nan Duan 0001, Weizhu Chen |
ICML | 5 |
| 2023 | Learning 3D Photography Videos via Self-supervised Diffusion on Single Imagesabstract3D photography renders a static image into a video with appealing 3D visual effects. Existing approaches typically first conduct monocular depth estimation, then render the input frame to subsequent frames with various viewpoints, and finally use an inpainting model to fill those missing/occluded regions. The inpainting model plays a crucial role in rendering quality, but it is normally trained on out-of-domain data. To reduce the training and inference gap, we propose a novel self-supervised diffusion model as the inpainting module. Given a single input image, we automatically construct a training pair of the masked occluded image and the ground-truth image with random cycle rendering. The constructed training samples are closely aligned to the testing instances, without the need for data annotation. To make full use of the masked images, we designed a Masked Enhanced Block (MEB), which can be easily plugged into the UNet and enhance the semantic conditions. Towards real-world animation, we present a novel task: out-animation, which extends the space and time of input objects. Extensive experiments on real datasets show that our method achieves competitive results with existing SOTA methods. Xiaodong Wang 0023, Chenfei Wu, Shengming Yin, Minheng Ni, Zhengyuan Yang, Fan Yang 0024, Zicheng Liu 0001, Yuejian Fang, Nan Duan 0001 |
IJCAI | 12 |
| 2023 | AR-Diffusion: Auto-Regressive Diffusion Model for Text GenerationabstractDiffusion models have gained significant attention in the realm of image generation due to their exceptional performance. Their success has been recently expanded to text generation via generating all tokens within a sequence concurrently.
However, natural language exhibits a far more pronounced sequential dependency in comparison to images, and the majority of existing language models are trained with a left-to-right auto-regressive approach.
To account for the inherent sequential characteristic of natural language, we introduce Auto-Regressive Diffusion (AR-Diffusion). AR-Diffusion ensures that the generation of tokens on the right depends on the generated ones on the left, a mechanism achieved through employing a dynamic number of denoising steps that vary based on token position. This results in tokens on the left undergoing fewer denoising steps than those on the right, thereby enabling them to generate earlier and subsequently influence the generation of tokens on the right.
In a series of experiments on various text generation tasks, including text summarization, machine translation, and common sense generation, AR-Diffusion clearly demonstrated its superiority over existing diffusion language models and that it can be $100\times\sim600\times$ faster when achieving comparable results. Our code is available at https://github.com/microsoft/ProphetNet/tree/master/AR-diffusion. Zhihao Fan, Xiao Liu 0029, Hai-Tao Zheng 0002, Yeyun Gong, Yelong Shen, Jian Jiao 0007, Zhongyu Wei, Jian Guo 0016, Nan Duan 0001, Weizhu Chen |
NeurIPS | 11 |
| 2023 | MASTER: Multi-task Pre-trained Bottlenecked Masked Autoencoders Are Better Dense Retrievers
Kun Zhou 0002, Xiao Liu 0029, Yeyun Gong, Wayne Xin Zhao, Daxin Jiang, Nan Duan 0001, Ji-Rong Wen |
ECML/PKDD (2) | 6 |
| 2023 | PROD: Progressive Distillation for Dense RetrievalabstractKnowledge distillation is an effective way to transfer knowledge from a strong teacher to an efficient student model. Ideally, we expect the better the teacher is, the better the student performs. However, this expectation does not always come true. It is common that a strong teacher model results in a bad student via distillation due to the nonnegligible gap between teacher and student. To bridge the gap, we propose PROD, a PROgressive Distillation method, for dense retrieval. PROD consists of a teacher progressive distillation and a data progressive distillation to gradually improve the student. To alleviate catastrophic forgetting, we introduce a regularization term in each distillation process. We conduct extensive experiments on seven datasets including five widely-used publicly available benchmarks: MS MARCO Passage, TREC Passage 19, TREC Document 19, MS MARCO Document, and Natural Questions, as well as two industry datasets: Bing-Rel and Bing-Ads. PROD achieves the state-of-the-art in the distillation methods for dense retrieval. Our 6-layer student model even surpasses most of the existing 12-layer models on all five public benchmarks. The code and models are released in https://github.com/microsoft/SimXNS. Zhenghao Lin, Yeyun Gong, Xiao Liu 0029, Hang Zhang 0029, Chen Lin 0001, Anlei Dong, Jian Jiao 0007, Jingwen Lu, Daxin Jiang, Rangan Majumder, Nan Duan 0001 |
WWW | 11 |
| 2023 | Enhancing RDF Verbalization with Descriptive and Relational KnowledgeabstractRDF verbalization has received increasing interest, which aims to generate a natural language description of the knowledge base. Sequence-to-sequence models based on Transformer are able to obtain strong performance equipped with pre-trained language models such as BART and T5. However, in spite of the general performance gain introduced by the pre-trained models, the performance of the task is still limited by the small scale of the training dataset. To address the problem, we propose two orthogonal strategies to enhance the representation learning of RDF triples. Concretely, two types of knowledge are introduced, i.e., descriptive knowledge and relational knowledge, respectively. The descriptive knowledge indicates the semantic information of self definition, and the relational knowledge indicates the semantic information learned from the structural context. We further combine the descriptive and relational knowledge together to enhance the representation learning. Experimental results on the WebNLG and SemEval-2010 datasets show that the two types of knowledge can both enhance the model performance, and their combination is able to obtain further improvements in most cases, providing new state-of-the-art results. Meishan Zhang, Shuang Liu 0007, Yueheng Sun, Nan Duan 0001 |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 5 |
| 2023 | LogiQA 2.0 - An Improved Dataset for Logical Reasoning in Natural Language UnderstandingabstractNLP research on logical reasoning regains momentum with the recent releases of a handful of datasets, notably LogiQA and Reclor. Logical reasoning is exploited in many probing tasks over large Pre-trained Language Models (PLMs) and downstream tasks like question-answering and dialogue systems. In this paper, we release LogiQA 2.0. The dataset is an amendment and re-annotation of LogiQA in 2020, a large-scale logical reasoning reading comprehension dataset adapted from the Chinese Civil Service Examination. We increase the data size, refine the texts with manual translation by professionals, and improve the quality by removing items with distinctive cultural features like Chinese idioms. Furthermore, we conduct a fine-grained annotation on the dataset and turn it into a two-way natural language inference (NLI) task, resulting in 35k premise-hypothesis pairs with gold labels, making it the first large-scale NLI dataset for complex logical reasoning. Compared to Question Answering, Natural Language Inference excels in generalizability and helps downstream tasks better. We establish a baseline for logical reasoning in NLI and incite further research. Hanmeng Liu, Jian Liu 0030, Leyang Cui, Zhiyang Teng, Nan Duan 0001, Ming Zhou 0001, Yue Zhang 0004 |
IEEE ACM Trans. Audio Speech Lang. Process. | 5 |
| 2022 | DialogVED: A Pre-trained Latent Variable Encoder-Decoder Model for Dialog Response GenerationabstractWei Chen, Yeyun Gong, Song Wang, Bolun Yao, Weizhen Qi, Zhongyu Wei, Xiaowu Hu, Bartuer Zhou, Yi Mao, Weizhu Chen, Biao Cheng, Nan Duan. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022. Wei Chen 0088, Yeyun Gong, Song Wang 0012, Bolun Yao, Weizhen Qi, Zhongyu Wei, Xiaowu Hu, Bartuer Zhou, Weizhu Chen, Biao Cheng, Nan Duan 0001 |
ACL (1) | 12 |
| 2022 | Contextual Fine-to-Coarse Distillation for Coarse-grained Response Selection in Open-Domain ConversationsabstractWei Chen, Yeyun Gong, Can Xu, Huang Hu, Bolun Yao, Zhongyu Wei, Zhihao Fan, Xiaowu Hu, Bartuer Zhou, Biao Cheng, Daxin Jiang, Nan Duan. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022. Wei Chen 0088, Yeyun Gong, Can Xu 0002, Huang Hu, Bolun Yao, Zhongyu Wei, Zhihao Fan, Xiaowu Hu, Bartuer Zhou, Biao Cheng, Daxin Jiang, Nan Duan 0001 |
ACL (1) | 12 |
| 2022 | UniXcoder: Unified Cross-Modal Pre-training for Code RepresentationabstractPre-trained models for programming languages have recently demonstrated great success on code intelligence.To support both code-related understanding and generation tasks, recent works attempt to pre-train unified encoder-decoder models.However, such encoder-decoder framework is sub-optimal for auto-regressive tasks, especially code completion that requires a decoder-only manner for efficient inference.In this paper, we present UniXcoder, a unified cross-modal pre-trained model for programming language.The model utilizes mask attention matrices with prefix adapters to control the behavior of the model and leverages cross-modal contents like AST and code comment to enhance code representation.To encode AST that is represented as a tree in parallel, we propose a one-to-one mapping method to transform AST in a sequence structure that retains all structural information from the tree.Furthermore, we propose to utilize multi-modal contents to learn representation of code fragment with contrastive learning, and then align representations among programming languages using a cross-modal generation task.We evaluate UniXcoder on five code-related tasks over nine datasets.To further evaluate the performance of code fragment representation, we also construct a dataset for a new task, called zero-shot code-to-code search.Results show that our model achieves state-of-the-art performance on most tasks and analysis reveals that comment and AST can both enhance UniXcoder. Daya Guo, Nan Duan 0001, Yanlin Wang 0001, Ming Zhou 0001, Jian Yin 0001 |
ACL (1) | 3 |
| 2022 | ReACC: A Retrieval-Augmented Code Completion FrameworkabstractCode completion, which aims to predict the following code token(s) according to the code context, can improve the productivity of software development.Recent work has proved that statistical language modeling with transformers can greatly improve the performance in the code completion task via learning from large-scale source code datasets.However, current approaches focus only on code context within the file or project, i.e. internal context.Our distinction is utilizing "external" context, inspired by human behaviors of copying from the related code snippets when writing code.Specifically, we propose a retrieval-augmented code completion framework, leveraging both lexical copying and referring to code with similar semantics by retrieval.We adopt a stagewise training approach that combines a source code retriever and an auto-regressive language model for programming language.We evaluate our approach in the code completion task in Python and Java programming languages, achieving a state-of-the-art performance on CodeXGLUE benchmark. Nan Duan 0001, Hojae Han, Daya Guo, Seung-won Hwang, Alexey Svyatkovskiy |
ACL (1) | 2 |
| 2022 | Multi-View Document Representation Learning for Open-Domain Dense RetrievalabstractDense retrieval has achieved impressive advances in first-stage retrieval from a largescale document collection, which is built on bi-encoder architecture to produce single vector representation of query and document.However, a document can usually answer multiple potential queries from different views.So the single vector representation of a document is hard to match with multi-view queries, and faces a semantic mismatch problem.This paper proposes a multi-view document representation learning framework, aiming to produce multiview embeddings to represent documents and enforce them to align with different queries.First, we propose a simple yet effective method of generating multiple embeddings through viewers.Second, to prevent multi-view embeddings from collapsing to the same one, we further propose a global-local loss with annealed temperature to encourage the multiple viewers to better align with different potential queries.Experiments show our method outperforms recent works and achieves state-of-the-art results. * Work done during internship at Microsoft Research Asia.Q1: Where can people using iPods on planes view the device's interface?A1: Individual seat-back displays.Q2: What are two airlines that considered implementing iPod connections but did not join the 2007 agreement?A2: KLM and Air France. Shunyu Zhang, Yaobo Liang, Ming Gong 0001, Daxin Jiang, Nan Duan 0001 |
ACL (1) | 5 |
| 2022 | VL-InterpreT: An Interactive Visualization Tool for Interpreting Vision-Language TransformersabstractBreakthroughs in transformer-based models have revolutionized not only the NLP field, but also vision and multimodal systems. However, although visualization and interpretability tools have become available for NLP models, internal mechanisms of vision and multimodal transformers remain largely opaque. With the success of these transformers, it is increasingly critical to understand their inner workings, as unraveling these black-boxes will lead to more capable and trustworthy models. To contribute to this quest, we propose VL-InterpreT, which provides novel interactive visualizations for interpreting the attentions and hidden representations in multimodal transformers. VL-InterpreT is a task agnostic and integrated tool that (1) tracks a variety of statistics in attention heads throughout all layers for both vision and language components, (2) visualizes cross-modal and intra-modal attentions through easily readable heatmaps, and (3) plots the hidden representations of vision and language tokens as they pass through the transformer layers. In this paper, we demonstrate the functionalities of VL-InterpreT through the analysis of KD-VLP, an end-to-end pretraining vision-language multimodal transformer-based model, in the tasks of Visual Commonsense Reasoning (VCR) and WebQA, two visual question answering benchmarks. Furthermore, we also present a few interesting findings about multimodal transformer behaviors that were learned through our tool. Estelle Aflalo, Shao-Yen Tseng, Yongfei Liu, Chenfei Wu, Nan Duan 0001, Vasudev Lal |
CVPR | 6 |
| 2022 | NÜWA: Visual Synthesis Pre-training for Neural visUal World creAtion
Chenfei Wu, Lei Ji 0001, Fan Yang 0024, Yuejian Fang, Daxin Jiang, Nan Duan 0001 |
ECCV (16) | 7 |
| 2022 | Trace Controlled Text to Image Generation
Kun Yan 0004, Lei Ji 0001, Chenfei Wu, Jianmin Bao, Ming Zhou 0001, Nan Duan 0001, Shuai Ma 0001 |
ECCV (36) | 6 |
| 2022 | Metric-guided Distillation: Distilling Knowledge from the Metric to Ranker and Retriever for Generative Commonsense ReasoningabstractXingwei He, Yeyun Gong, A-Long Jin, Weizhen Qi, Hang Zhang, Jian Jiao, Bartuer Zhou, Biao Cheng, Sm Yiu, Nan Duan. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. 2022. Xingwei He 0003, Yeyun Gong, A-Long Jin, Weizhen Qi, Hang Zhang 0029, Jian Jiao 0007, Bartuer Zhou, Biao Cheng, Siu-Ming Yiu, Nan Duan 0001 |
EMNLP | 10 |
| 2022 | Sentiment-Aware Word and Sentence Level Pre-training for Sentiment AnalysisabstractMost existing pre-trained language representation models (PLMs) are sub-optimal in sentiment analysis tasks, as they capture the sentiment information from word-level while underconsidering sentence-level information.In this paper, we propose SentiWSP, a novel Sentiment-aware pre-trained language model with combined Word-level and Sentence-level Pre-training tasks.The word level pre-training task detects replaced sentiment words, via a generator-discriminator framework, to enhance the PLM's knowledge about sentiment words.The sentence level pre-training task further strengthens the discriminator via a contrastive learning framework, with similar sentences as negative samples, to encode sentiments in a sentence.Extensive experimental results show that SentiWSP achieves new state-of-the-art performance on various sentence-level and aspectlevel sentiment classification benchmarks.We have made our code and model publicly available at https://github.com/XMUDM/SentiWSP. Shuai Fan 0007, Chen Lin 0001, Haonan Li 0002, Zhenghao Lin, Jinsong Su, Hang Zhang 0029, Yeyun Gong, Jian Guo 0016, Nan Duan 0001 |
EMNLP | 9 |
| 2022 | Towards Compositional Generalization in Code SearchabstractWe study compositional generalization, which aims to generalize on unseen combinations of seen structural elements, for code search.Unlike existing approaches of partially pursuing this goal, we study how to extract structural elements, which we name a template that directly targets compositional generalization.Thus we propose CTBERT, or Code Template BERT, representing codes using automatically extracted templates as building blocks.We empirically validate CTBERT on two public code search benchmarks, AdvTest and CSN.Further, we show that templates are complementary to data flow graphs in GraphCodeBERT, by enhancing structural context around variables. Hojae Han, Seung-won Hwang, Nan Duan 0001, Seungtaek Choi |
EMNLP | 4 |
| 2022 | CodeRetriever: A Large Scale Contrastive Pre-Training Method for Code SearchabstractXiaonan Li, Yeyun Gong, Yelong Shen, Xipeng Qiu, Hang Zhang, Bolun Yao, Weizhen Qi, Daxin Jiang, Weizhu Chen, Nan Duan. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. 2022. Yeyun Gong, Yelong Shen, Xipeng Qiu, Hang Zhang 0029, Bolun Yao, Weizhen Qi, Daxin Jiang, Weizhu Chen, Nan Duan 0001 |
EMNLP | 10 |
| 2022 | Learning to Complete Code with Sketches
Daya Guo, Alexey Svyatkovskiy, Jian Yin 0001, Nan Duan 0001, Marc Brockschmidt, Miltiadis Allamanis |
ICLR | 4 |
| 2022 | Adversarial Retriever-Ranker for Dense Text Retrieval
Hang Zhang 0029, Yeyun Gong, Yelong Shen, Jiancheng Lv 0001, Nan Duan 0001, Weizhu Chen |
ICLR | 5 |
| 2022 | Unsupervised Context Aware Sentence Representation Pretraining for Multi-lingual Dense RetrievalabstractRecent research demonstrates the effectiveness of using pretrained language models (PLM) to improve dense retrieval and multilingual dense retrieval. In this work, we present a simple but effective monolingual pretraining task called contrastive context prediction (CCP) to learn sentence representation by modeling sentence level contextual relation. By pushing the embedding of sentences in a local context closer and pushing random negative samples away, different languages could form isomorphic structure, then sentence pairs in two different languages will be automatically aligned. Our experiments show that model collapse and information leakage are very easy to happen during contrastive training of language model, but language-specific memory bank and asymmetric batch normalization operation play an essential role in preventing collapsing and information leakage, respectively. Besides, a post-processing for sentence embedding is also very effective to achieve better retrieval performance. On the multilingual sentence retrieval task Tatoeba, our model achieves new SOTA results among methods without using bilingual data. Our model also shows larger gain on Tatoeba when transferring between non-English pairs. On two multi-lingual query-passage retrieval tasks, XOR Retrieve and Mr.TYDI, our model even achieves two SOTA results in both zero-shot and supervised setting among all pretraining models using bilingual data. Ning Wu 0013, Yaobo Liang, Houxing Ren, Linjun Shou, Nan Duan 0001, Ming Gong 0001, Daxin Jiang |
IJCAI | 5 |
| 2022 | Reasoning over Hybrid Chain for Table-and-Text Open Domain Question AnsweringabstractTabular and textual question answering requires systems to perform reasoning over heterogeneous information, considering table structure, and the connections among table and text. In this paper, we propose a ChAin-centric Reasoning and Pre-training framework (CARP). CARP utilizes hybrid chain to model the explicit intermediate reasoning process across table and text for question answering. We also propose a novel chain-centric pre-training method, to enhance the pre-trained model in identifying the cross-modality reasoning process and alleviating the data sparsity problem. This method constructs the large-scale reasoning corpus by synthesizing pseudo heterogeneous reasoning paths from Wikipedia and generating corresponding questions. We evaluate our system on OTT-QA, a large-scale table-and-text open-domain question answering benchmark, and our system achieves the state-of-the-art performance. Further analyses illustrate that the explicit hybrid chain offers substantial performance improvement and interpretablity of the intermediate reasoning process, and the chain-centric pre-training boosts the performance on the chain extraction. Wanjun Zhong, Junjie Huang 0008, Qian Liu 0033, Ming Zhou 0001, Jiahai Wang, Jian Yin 0001, Nan Duan 0001 |
IJCAI | 7 |
| 2022 | AdsCVLR: Commercial Visual-Linguistic Representation Modeling in Sponsored SearchabstractSponsored search advertisements (ads) appear next to search results when consumers look for products and services on search engines. As the fundamental basis of search ads, relevance modeling has attracted increasing attention due to the significant research challenges and tremendous practical value. In this paper, we address the problem of multi-modal modeling in sponsored search, which models the relevance between user query and commercial ads with multi-modal structured information. To solve this problem, we propose a transformer architecture with Ads data on Commercial Visual-Linguistic Representation (AdsCVLR) with contrastive learning that naturally extends the transformer encoder with the complementary multi-modal inputs, serving as a strong aggregator of image-text features. We also make a public advertising dataset, which includes 480K labeled query-ad pairwise data with structured information of image, title, seller, description, and so on. Empirically, we evaluate the AdsCVLR model over the large industry dataset, and the experimental results of online/offline tests show the superiority of our method. Yongjie Zhu, Chunhui Han, Yuefeng Zhan, Bochen Pang, Zhaoju Li, Hao Sun 0015, Si Li 0001, Boxin Shi, Nan Duan 0001, Ruofei Zhang, Liangjie Zhang, Qi Zhang 0066 |
ACM Multimedia | 9 |
| 2022 | ProQA: Structural Prompt-based Pre-training for Unified Question AnsweringabstractWanjun Zhong, Yifan Gao, Ning Ding, Yujia Qin, Zhiyuan Liu, Ming Zhou, Jiahai Wang, Jian Yin, Nan Duan. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Wanjun Zhong, Yifan Gao 0001, Ning Ding 0002, Yujia Qin, Zhiyuan Liu 0001, Ming Zhou 0001, Jiahai Wang, Jian Yin 0001, Nan Duan 0001 |
NAACL-HLT | 9 |
| 2022 | NUWA-Infinity: Autoregressive over Autoregressive Generation for Infinite Visual SynthesisabstractInfinite visual synthesis aims to generate high-resolution images, long-duration videos, and even visual generation of infinite size. Some recent work tried to solve this task by first dividing data into processable patches and then training the models on them without considering the dependencies between patches. However, since they fail to model global dependencies between patches, the quality and consistency of the generation can be limited. To address this issue, we propose NUWA-Infinity, a patch-level \emph{``render-and-optimize''} strategy for infinite visual synthesis. Given a large image or a long video, NUWA-Infinity first splits it into non-overlapping patches and uses the ordered patch chain as a complete training instance, a rendering model autoregressively predicts each patch based on its contexts. Once a patch is predicted, it is optimized immediately and its hidden states are saved as contexts for the next \emph{``render-and-optimize''} process. This brings two advantages: ($i$) The autoregressive rendering process with information transfer between contexts provides an implicit global probabilistic distribution modeling; ($ii$) The timely optimization process alleviates the optimization stress of the model and helps convergence. Based on the above designs, NUWA-Infinity shows a strong synthesis ability on high-resolution images and long-duration videos. The homepage link is \url{https://nuwa-infinity.microsoft.com}. Chenfei Wu, Xiaowei Hu 0006, Zhe Gan, Zicheng Liu 0001, Yuejian Fang, Nan Duan 0001 |
NeurIPS | 9 |
| 2022 | LogiGAN: Learning Logical Reasoning via Adversarial Pre-trainingabstractWe present LogiGAN, an unsupervised adversarial pre-training framework for improving logical reasoning abilities of language models. Upon automatic identification of logical reasoning phenomena in massive text corpus via detection heuristics, we train language models to predict the masked-out logical statements. Inspired by the facilitation effect of reflective thinking in human learning, we analogically simulate the learning-thinking process with an adversarial Generator-Verifier architecture to assist logic learning. LogiGAN implements a novel sequential GAN approach that (a) circumvents the non-differentiable challenge of the sequential GAN by leveraging the Generator as a sentence-level generative likelihood scorer with a learning objective of reaching scoring consensus with the Verifier; (b) is computationally feasible for large-scale pre-training with arbitrary target length. Both base and large size language models pre-trained with LogiGAN demonstrate obvious performance improvement on 12 datasets requiring general reasoning abilities, revealing the fundamental role of logic in broad reasoning, as well as the effectiveness of LogiGAN. Ablation studies on LogiGAN components reveal the relative orthogonality between linguistic and logic abilities and suggest that reflective thinking's facilitation effect might also generalize to machine learning. Xinyu Pi, Wanjun Zhong, Yan Gao 0002, Nan Duan 0001, Jian-Guang Lou |
NeurIPS | 4 |
| 2022 | Automating code review activities by large-scale pre-trainingabstractCode review is an essential part to software development lifecycle since it aims at guaranteeing the quality of codes. Modern code review activities necessitate developers viewing, understanding and even running the programs to assess logic, functionality, latency, style and other factors. It turns out that developers have to spend far too much time reviewing the code of their peers. Accordingly, it is in significant demand to automate the code review process. In this research, we focus on utilizing pre-training techniques for the tasks in the code review scenario. We collect a large-scale dataset of real-world code changes and code reviews from open-source projects in nine of the most popular programming languages. To better understand code diffs and reviews, we propose CodeReviewer, a pre-trained model that utilizes four pre-training tasks tailored specifically for the code review scenario. To evaluate our model, we focus on three key tasks related to code review activities, including code change quality estimation, review comment generation and code refinement. Furthermore, we establish a high-quality benchmark dataset based on our collected data for these three tasks and conduct comprehensive experiments on it. The experimental results demonstrate that our model outperforms the previous state-of-the-art pre-training approaches in all tasks. Further analysis show that our proposed pre-training tasks and the multilingual pre-training dataset benefit the model on the understanding of code changes and reviews. Daya Guo, Nan Duan 0001, Shailesh Jannu, Grant Jenks, Deep Majumder, Jared Green, Alexey Svyatkovskiy, Shengyu Fu, Neel Sundaresan |
ESEC/SIGSOFT FSE | 4 |
| 2022 | Learning Temporal Video Procedure Segmentation from an Automatically Collected Large DatasetabstractTemporal Video Segmentation (TVS) is a fundamental video understanding task and has been widely researched in recent years. There are two subtasks of TVS: Video Action Segmentation (VAS) and Video Procedure Segmentation (VPS): VAS aims to recognize what actions happen in-side the video while VPS aims to segment the video into a sequence of video clips as a procedure. The VAS task inevitably relies on pre-defined action labels and is thus hard to scale to various open-domain videos. To overcome this limitation, the VPS task tries to divide a video into several category-independent procedure segments. However, the existing dataset for the VPS task is small (2k videos) and lacks diversity (only cooking domain). To tackle these problems, we collect a large and diverse dataset called TIPS, specifically for the VPS task. TIPS contains 63k videos including more than 300k procedure segments from instructional videos on YouTube, which covers plenty of how-to areas such as cooking, health, beauty, parenting, gardening, etc. We then propose a multi-modal Transformer with Gaussian Boundary Detection (MT-GBD) model for VPS, with the backbone of the Transformer and Convolution. Furthermore, we propose a new EIOU metric for the VPS task, which helps better evaluate VPS quality in a more comprehensive way. Experimental results show the effectiveness of our proposed model and metric. Lei Ji 0001, Chenfei Wu, Daisy Zhou, Kun Yan 0004, Edward Dong Bo Cui, Xilin Chen 0001, Nan Duan 0001 |
WACV | 7 |
| 2022 | Multimodal graph neural network for video procedural captioning
Lei Ji 0001, Rongcheng Tu, Nan Duan 0001 |
Neurocomputing | 5 |
| 2022 | CLIP4Clip: An empirical study of CLIP for end to end video clip retrieval and captioning
Huaishao Luo, Lei Ji 0001, Ming Zhong 0015, Wen Lei, Nan Duan 0001, Tianrui Li 0001 |
Neurocomputing | 6 |
| 2022 | From LSAT: The Progress and Challenges of Complex ReasoningabstractComplex reasoning aims to draw a correct inference based on complex rules. As a hallmark of human intelligence, it involves a degree of explicit reading comprehension, interpretation of logical knowledge and complex rule application. In this paper, we take a step forward in complex reasoning by systematically studying the three challenging and domain-general tasks of the Law School Admission Test (LSAT), including analytical reasoning, logical reasoning and reading comprehension. We propose a hybrid reasoning system to integrate these three tasks and achieve impressive overall performance on the LSAT tests. The experimental results demonstrate that our system endows itself a certain complex reasoning ability, especially the fundamental reading comprehension and challenging logical reasoning capacities. Further analysis also shows the effectiveness of combining the pre-trained models with the task-specific reasoning module, and integrating symbolic knowledge into discrete interpretable reasoning steps in complex reasoning. We further shed a light on the potential future directions, like unsupervised symbolic knowledge extraction, model interpretability, few-shot learning and comprehensive benchmark for complex reasoning. Siyuan Wang 0025, Zhongkun Liu, Wanjun Zhong, Ming Zhou 0001, Zhongyu Wei, Zhumin Chen, Nan Duan 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 7 |
| 2022 | Conditional Sentence Generation and Cross-Modal Reranking for Sign Language TranslationabstractSign Language Translation (SLT) aims to generate spoken language translations from sign language videos. Currently, the available sign language datasets are relatively too small to learn the linguistic properties of spoken language. In this paper, towards effective SLT, we propose a novel framework which takes the advantage of the spoken language grammar learnt from a large corpus of text sentences. Our framework consists of three key modules: word existence verification, conditional sentence generation and cross-modal re-ranking. We first check the existence of words in the vocabulary by a series of binary classification in parallel. After that, the appearing words are assembled and guided by a pretrained spoken language generator to produce multiple candidate sentences in spoken language manner. Last but not least, we select the sentence most semantically similar to the input sign video as the translation result with a crossmodal re-ranking model. We evaluate our framework on two large scale continuous SLT benchmarks,i.e., CSL and RWTHPHOENIX-Weather 2014 T. Experimental results demonstrate that the proposed framework achieves promising performance on both datasets. Jian Zhao 0018, Weizhen Qi, Wengang Zhou 0001, Nan Duan 0001, Ming Zhou 0001, Houqiang Li |
IEEE Trans. Multim. | 4 |
| 2021 | Compare to The Knowledge: Graph Neural Fake News Detection with External KnowledgeabstractLinmei Hu, Tianchi Yang, Luhao Zhang, Wanjun Zhong, Duyu Tang, Chuan Shi, Nan Duan, Ming Zhou. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Linmei Hu, Tianchi Yang, Luhao Zhang, Wanjun Zhong, Duyu Tang, Chuan Shi 0001, Nan Duan 0001, Ming Zhou 0001 |
ACL/IJCNLP (1) | 7 |
| 2021 | CoSQA: 20, 000+ Web Queries for Code Search and Question AnsweringabstractJunjie Huang, Duyu Tang, Linjun Shou, Ming Gong, Ke Xu, Daxin Jiang, Ming Zhou, Nan Duan. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Junjie Huang 0008, Duyu Tang, Linjun Shou, Ming Gong 0001, Ke Xu 0001, Daxin Jiang, Ming Zhou 0001, Nan Duan 0001 |
ACL/IJCNLP (1) | 8 |
| 2021 | Syntax-Enhanced Pre-trained ModelabstractZenan Xu, Daya Guo, Duyu Tang, Qinliang Su, Linjun Shou, Ming Gong, Wanjun Zhong, Xiaojun Quan, Daxin Jiang, Nan Duan. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Zenan Xu, Daya Guo, Duyu Tang, Qinliang Su, Linjun Shou, Ming Gong 0001, Wanjun Zhong, Xiaojun Quan, Daxin Jiang, Nan Duan 0001 |
ACL/IJCNLP (1) | 10 |
| 2021 | Control Image Captioning Spatially and TemporallyabstractKun Yan, Lei Ji, Huaishao Luo, Ming Zhou, Nan Duan, Shuai Ma. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Kun Yan 0004, Lei Ji 0001, Huaishao Luo, Ming Zhou 0001, Nan Duan 0001, Shuai Ma 0001 |
ACL/IJCNLP (1) | 5 |
| 2021 | M3P: Learning Universal Representations via Multitask Multilingual Multimodal Pre-TrainingabstractWe present M3P, a Multitask Multilingual Multimodal Pre-trained model that combines multilingual pre-training and multimodal pre-training into a unified framework via multitask pre-training. Our goal is to learn universal representations that can map objects occurred in different modalities or texts expressed in different languages into a common semantic space. In addition, to explicitly encourage fine-grained alignment between images and non-English languages, we also propose Multimodal Code-switched Training (MCT) to combine monolingual pre-training and multimodal pre-training via a code-switch strategy. Experiments are performed on the multilingual image retrieval task across two benchmark datasets, including MSCOCO and Multi30K. M3P can achieve comparable results for English and new state-of-the-art results for non-English languages. Minheng Ni, Haoyang Huang, Edward Dong Bo Cui, Taroon Bharti, Dongdong Zhang 0001, Nan Duan 0001 |
CVPR | 8 |
| 2021 | Long-Range Modeling of Source Code Files with eWASH: Extended Window Access by Syntax HierarchyabstractColin Clement, Shuai Lu, Xiaoyu Liu, Michele Tufano, Dawn Drain, Nan Duan, Neel Sundaresan, Alexey Svyatkovskiy. Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing. 2021. Colin B. Clement, Michele Tufano, Dawn Drain, Nan Duan 0001, Neel Sundaresan, Alexey Svyatkovskiy |
EMNLP (1) | 6 |
| 2021 | GraphCodeBERT: Pre-training Code Representations with Data Flow
Daya Guo, Shuo Ren 0002, Zhangyin Feng, Duyu Tang, Shujie Liu 0001, Nan Duan 0001, Alexey Svyatkovskiy, Shengyu Fu, Michele Tufano, Shao Kun Deng, Colin B. Clement, Dawn Drain, Neel Sundaresan, Jian Yin 0001, Daxin Jiang, Ming Zhou 0001 |
ICLR | 8 |
| 2021 | BANG: Bridging Autoregressive and Non-autoregressive Generation with Large Scale PretrainingabstractIn this paper, we propose BANG, a new pretraining model to Bridge the gap between Autoregressive (AR) and Non-autoregressive (NAR) Generation. AR and NAR generation can be uniformly regarded as to what extent previous tokens can be attended, and BANG bridges AR and NAR generation through designing a novel model structure for large-scale pre-training. A pretrained BANG model can simultaneously support AR, NAR, and semi-NAR generation to meet different requirements. Experiments on question generation (SQuAD 1.1), summarization (XSum), and dialogue generation (PersonaChat) show that BANG improves NAR and semi-NAR performance significantly as well as attaining comparable performance with strong AR pretrained models. Compared with the semi-NAR strong baselines, BANG achieves absolute improvements of 14.01 and 5.24 in the overall scores of SQuAD 1.1 and XSum, respectively. In addition, BANG achieves absolute improvements of 10.73, 6.39, and 5.90 in the overall scores of SQuAD, XSUM, and PersonaChat compared with the NAR strong baselines, respectively. Our code will be made publicly available. Weizhen Qi, Yeyun Gong, Jian Jiao 0007, Weizhu Chen, Dayiheng Liu, Kewen Tang, Houqiang Li, Jiusheng Chen, Ruofei Zhang, Ming Zhou 0001, Nan Duan 0001 |
ICML | 12 |
| 2021 | EL-Attention: Memory Efficient Lossless Attention for GenerationabstractTransformer model with multi-head attention requires caching intermediate results for efficient inference in generation tasks. However, cache brings new memory-related costs and prevents leveraging larger batch size for faster speed. We propose memory-efficient lossless attention (called EL-attention) to address this issue. It avoids heavy operations for building multi-head keys and values, cache for them is not needed. EL-attention constructs an ensemble of attention results by expanding query while keeping key and value shared. It produces the same result as multi-head attention with less GPU memory and faster inference speed. We conduct extensive experiments on Transformer, BART, and GPT-2 for summarization and question generation tasks. The results show EL-attention speeds up existing models by 1.6x to 5.3x without accuracy loss. Jiusheng Chen, Weizhen Qi, Nikhil Bhendawade, Yeyun Gong, Nan Duan 0001, Ruofei Zhang |
ICML | 6 |
| 2021 | Poolingformer: Long Document Modeling with Pooling AttentionabstractIn this paper, we introduce a two-level attention schema, Poolingformer, for long document modeling. Its first level uses a smaller sliding window pattern to aggregate information from neighbors. Its second level employs a larger window to increase receptive fields with pooling attention to reduce both computational cost and memory consumption. We first evaluate Poolingformer on two long sequence QA tasks: the monolingual NQ and the multilingual TyDi QA. Experimental results show that Poolingformer sits atop three official leaderboards measured by F1, outperforming previous state-of-the-art models by 1.9 points (79.8 vs. 77.9) on NQ long answer, 1.9 points (79.5 vs. 77.6) on TyDi QA passage answer, and 1.6 points (67.6 vs. 66.0) on TyDi QA minimal answer. We further evaluate Poolingformer on a long sequence summarization task. Experimental results on the arXiv benchmark continue to demonstrate its superior performance. Hang Zhang 0029, Yeyun Gong, Yelong Shen, Weisheng Li 0001, Jiancheng Lv 0001, Nan Duan 0001, Weizhu Chen |
ICML | 6 |
| 2021 | Hybrid Reasoning Network for Video-based Commonsense CaptioningabstractThe task of video-based commonsense captioning aims to generate event-wise captions and meanwhile provide multiple commonsense descriptions (e.g., attribute, effect and intention) about the underlying event in the video. Prior works explore the commonsense captions by using separate networks for different commonsense types, which is time-consuming and lacks mining the interaction of different commonsense. In this paper, we propose a Hybrid Reasoning Network (HybridNet) to endow the neural networks with the capability of semantic-level reasoning and word-level reasoning. Firstly, we develop multi-commonsense learning for semantic-level reasoning by jointly training different commonsense types in a unified network, which encourages the interaction between the clues of multiple commonsense descriptions, event-wise captions and videos. Then, there are two steps to achieve the word-level reasoning: (1) a memory module records the history predicted sequence from the previous generation processes; (2) a memory-routed multi-head attention (MMHA) module updates the word-level attention maps by incorporating the history information from the memory module into the transformer decoder for word-level reasoning. Moreover, the multimodal features are used to make full use of diverse knowledge for commonsense reasoning. Experiments and abundant analysis on the large-scale Video-to-Commonsense benchmark show that our HybridNet achieves state-of-the-art performance compared with other methods. Weijiang Yu, Lei Ji 0001, Yuejian Fang, Nan Duan 0001 |
ACM Multimedia | 7 |
| 2021 | Mask Attention Networks: Rethinking and Strengthen TransformerabstractZhihao Fan, Yeyun Gong, Dayiheng Liu, Zhongyu Wei, Siyuan Wang, Jian Jiao, Nan Duan, Ruofei Zhang, Xuanjing Huang. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Zhihao Fan, Yeyun Gong, Dayiheng Liu, Zhongyu Wei, Siyuan Wang 0025, Jian Jiao 0007, Nan Duan 0001, Ruofei Zhang, Xuanjing Huang 0001 |
NAACL-HLT | 7 |
| 2021 | Learning from Inside: Self-driven Siamese Sampling and Reasoning for Video Question AnsweringabstractRecent advances in the video question answering (i.e., VideoQA) task have achieved strong success by following the paradigm of fine-tuning each clip-text pair independently on the pretrained transformer-based model via supervised learning. Intuitively, multiple samples (i.e., clips) should be interdependent to capture similar visual and key semantic information in the same video. To consider the interdependent knowledge between contextual clips into the network inference, we propose a Siamese Sampling and Reasoning (SiaSamRea) approach, which consists of a siamese sampling mechanism to generate sparse and similar clips (i.e., siamese clips) from the same video, and a novel reasoning strategy for integrating the interdependent knowledge between contextual clips into the network. The reasoning strategy contains two modules: (1) siamese knowledge generation to learn the inter-relationship among clips; (2) siamese knowledge reasoning to produce the refined soft label by propagating the weights of inter-relationship to the predicted candidates of all clips. Finally, our SiaSamRea can endow the current multimodal reasoning paradigm with the ability of learning from inside via the guidance of soft labels. Extensive experiments demonstrate our SiaSamRea achieves state-of-the-art performance on five VideoQA benchmarks, e.g., a significant +2.1% gain on MSRVTT-QA, +2.9% on MSVD-QA, +1.0% on ActivityNet-QA, +1.8% on How2QA and +4.3% (action) on TGIF-QA. Weijiang Yu, Haoteng Zheng, Lei Ji 0001, Nan Duan 0001 |
NeurIPS | 7 |
| 2021 | XGPT: Cross-modal Generative Pre-Training for Image Captioning
Qiaolin Xia, Haoyang Huang, Nan Duan 0001, Dongdong Zhang 0001, Lei Ji 0001, Zhifang Sui, Edward Dong Bo Cui, Taroon Bharti, Ming Zhou 0001 |
NLPCC (1) | 3 |
| 2021 | Question Generation from Code Snippets and Programming Error Messages
Bolun Yao, Wei Chen 0088, Yeyun Gong, Bartuer Zhou, Zhongyu Wei, Biao Cheng, Nan Duan 0001 |
NLPCC (1) | 8 |
| 2021 | Tree-Capsule: Tree-Structured Capsule Network for Improving Relation Extraction
Tianchi Yang, Linmei Hu, Luhao Zhang, Chuan Shi 0001, Cheng Yang 0002, Nan Duan 0001, Ming Zhou 0001 |
PAKDD (3) | 6 |
| 2020 | Segment-Then-Rank: Non-Factoid Question Answering on Instructional Videos
Kyungjae Lee 0002, Nan Duan 0001, Lei Ji 0001, Seung-won Hwang |
AAAI | 2 |
| 2020 | Unicoder-VL: A Universal Encoder for Vision and Language by Cross-Modal Pre-TrainingabstractWe propose Unicoder-VL, a universal encoder that aims to learn joint representations of vision and language in a pre-training manner. Borrow ideas from cross-lingual pre-trained models, such as XLM (Lample and Conneau 2019) and Unicoder (Huang et al. 2019), both visual and linguistic contents are fed into a multi-layer Transformer (Vaswani et al. 2017) for the cross-modal pre-training, where three pre-trained tasks are employed, including Masked Language Modeling(MLM), Masked Object Classification(MOC) and Visual-linguistic Matching(VLM). The first two tasks learn context-aware representations for input tokens based on linguistic and visual contents jointly. The last task tries to predict whether an image and a text describe each other. After pretraining on large-scale image-caption pairs, we transfer Unicoder-VL to caption-based image-text retrieval and visual commonsense reasoning, with just one additional output layer. We achieve state-of-the-art or comparable results on both two tasks and show the powerful ability of the cross-modal pre-training. Nan Duan 0001, Yuejian Fang, Ming Gong 0001, Daxin Jiang |
AAAI | 2 |
| 2020 | Graph-Based Reasoning over Heterogeneous External Knowledge for Commonsense Question AnsweringabstractCommonsense question answering aims to answer questions which require background knowledge that is not explicitly expressed in the question. The key challenge is how to obtain evidence from external knowledge and make predictions based on the evidence. Recent studies either learn to generate evidence from human-annotated evidence which is expensive to collect, or extract evidence from either structured or unstructured knowledge bases which fails to take advantages of both sources simultaneously. In this work, we propose to automatically extract evidence from heterogeneous knowledge sources, and answer questions based on the extracted evidence. Specifically, we extract evidence from both structured knowledge base (i.e. ConceptNet) and Wikipedia plain texts. We construct graphs for both sources to obtain the relational structures of evidence. Based on these graphs, we propose a graph-based approach consisting of a graph-based contextual word representation learning module and a graph-based inference module. The first module utilizes graph structural information to re-define the distance between words for learning better contextual word representations. The second module adopts graph convolutional network to encode neighbor information into the representations of nodes, and aggregates evidence with graph attention mechanism for predicting the final answer. Experimental results on CommonsenseQA dataset illustrate that our graph-based approach over both knowledge sources brings improvement over strong baselines. Our approach achieves the state-of-the-art accuracy (75.3%) on the CommonsenseQA dataset. Shangwen Lv, Daya Guo, Jingjing Xu 0001, Duyu Tang, Nan Duan 0001, Ming Gong 0001, Linjun Shou, Daxin Jiang, Guihong Cao, Songlin Hu 0001 |
AAAI | 5 |
| 2020 | Neural Semantic Parsing in Low-Resource Settings with Back-Translation and Meta-LearningabstractNeural semantic parsing has achieved impressive results in recent years, yet its success relies on the availability of large amounts of supervised data. Our goal is to learn a neural semantic parser when only prior knowledge about a limited number of simple rules is available, without access to either annotated programs or execution results. Our approach is initialized by rules, and improved in a back-translation paradigm using generated question-program pairs from the semantic parser and the question generator. A phrase table with frequent mapping patterns is automatically derived, also updated as training progresses, to measure the quality of generated instances. We train the model with model-agnostic meta-learning to guarantee the accuracy and stability on examples covered by rules, and meanwhile acquire the versatility to generalize well on examples uncovered by rules. Results on three benchmark datasets with different domains and programs show that our approach incrementally improves the accuracy. On WikiSQL, our best model is comparable to the state-of-the-art system learned from denotations. Duyu Tang, Nan Duan 0001, Yeyun Gong, Bing Qin 0001, Daxin Jiang |
AAAI | 3 |
| 2020 | Evidence-Aware Inferential Text Generation with Vector Quantised Variational AutoEncoderabstractGenerating inferential texts about an event in different perspectives requires reasoning over different contexts that the event occurs.Existing works usually ignore the context that is not explicitly provided, resulting in a context-independent semantic representation that struggles to support the generation.To address this, we propose an approach that automatically finds evidence for an event from a large text corpus, and leverages the evidence to guide the generation of inferential texts.Our approach works in an encoderdecoder manner and is equipped with a Vector Quantised-Variational Autoencoder, where the encoder outputs representations from a distribution over discrete variables.Such discrete representations enable automatically selecting relevant evidence, which not only facilitates evidence-aware generation, but also provides a natural way to uncover rationales behind the generation.Our approach provides state-ofthe-art performance on both Event2Mind and ATOMIC datasets.More importantly, we find that with discrete representations, our model selectively uses evidence to generate different inferential texts. Daya Guo, Duyu Tang, Nan Duan 0001, Jian Yin 0001, Daxin Jiang, Ming Zhou 0001 |
ACL | 3 |
| 2020 | Graph Neural News Recommendation with Unsupervised Preference DisentanglementabstractWith the explosion of news information, personalized news recommendation has become very important for users to quickly find their interested contents. Most existing methods usually learn the representations of users and news from news contents for recommendation. However, they seldom consider high-order connectivity underlying the user-news interactions. Moreover, existing methods failed to disentangle a user’s latent preference factors which cause her clicks on different news. In this paper, we model the user-news interactions as a bipartite graph and propose a novel Graph Neural News Recommendation model with Unsupervised Preference Disentanglement, named GNUD. Our model can encode high-order relationships into user and news representations by information propagation along the graph. Furthermore, the learned representations are disentangled with latent preference factors by a neighborhood routing algorithm, which can enhance expressiveness and interpretability. A preference regularizer is also designed to force each disentangled subspace to independently reflect an isolated preference, improving the quality of the disentangled representations. Experimental results on real-world news datasets demonstrate that our proposed model can effectively improve the performance of news recommendation and outperform state-of-the-art news recommendation methods. Linmei Hu, Siyong Xu, Cheng Yang 0002, Chuan Shi 0001, Nan Duan 0001, Xing Xie 0001, Ming Zhou 0001 |
ACL | 6 |
| 2020 | RikiNet: Reading Wikipedia Pages for Natural Question AnsweringabstractReading long documents to answer opendomain questions remains challenging in natural language understanding.In this paper, we introduce a new model, called RikiNet, which reads Wikipedia pages for natural question answering.RikiNet contains a dynamic paragraph dual-attention reader and a multi-level cascaded answer predictor.The reader dynamically represents the document and question by utilizing a set of complementary attention mechanisms.The representations are then fed into the predictor to obtain the span of the short answer, the paragraph of the long answer, and the answer type in a cascaded manner.On the Natural Questions (NQ) dataset, a single RikiNet achieves 74.3 F1 and 57.9 F1 on longanswer and short-answer tasks.To our best knowledge, it is the first single model that outperforms the single human performance.Furthermore, an ensemble RikiNet obtains 76.1 F1 and 61.3 F1 on long-answer and shortanswer tasks, achieving the best performance on the official NQ leaderboard 1 . Dayiheng Liu, Yeyun Gong, Jie Fu 0001, Jiusheng Chen, Daxin Jiang, Jiancheng Lv 0001, Nan Duan 0001 |
ACL | 8 |
| 2020 | Enhancing Answer Boundary Detection for Multilingual Machine Reading ComprehensionabstractMultilingual pre-trained models could leverage the training data from a rich source language (such as English) to improve the performance on low resource languages.However, the transfer effectiveness on the multilingual Machine Reading Comprehension (MRC) task is substantially poorer than that for sentence classification tasks, mainly due to the requirement of MRC to detect the word level answer boundary.In this paper, we propose two auxiliary tasks to introduce additional phrase boundary supervision in the fine-tuning stage:(1) a mixed MRC task, which translates the question or passage to other languages and builds cross-lingual question-passage pairs; and (2) a language-agnostic knowledge masking task by leveraging knowledge phrases mined from the Web.Extensive experiments on two cross-lingual MRC datasets show the effectiveness of our proposed approach.† Random N-gram Masking shows gains in English SQuAD. Fei Yuan 0010, Linjun Shou, Xuanyu Bai, Ming Gong 0001, Yaobo Liang, Nan Duan 0001, Daxin Jiang |
ACL | 6 |
| 2020 | Document Modeling with Graph Attention Networks for Multi-grained Machine Reading ComprehensionabstractNatural Questions is a new challenging machine reading comprehension benchmark with two-grained answers, which are a long answer (typically a paragraph) and a short answer (one or more entities inside the long answer).Despite the effectiveness of existing methods on this benchmark, they treat these two sub-tasks individually during training while ignoring their dependencies.To address this issue, we present a novel multi-grained machine reading comprehension framework that focuses on modeling documents at their hierarchical nature, which are different levels of granularity: documents, paragraphs, sentences, and tokens.We utilize graph attention networks to obtain different levels of representations so that they can be learned simultaneously.The long and short answers can be extracted from paragraphlevel representation and token-level representation, respectively.In this way, we can model the dependencies between the two-grained answers to provide evidence for each other.We jointly train the two sub-tasks, and our experiments show that our approach significantly outperforms previous systems at both long and short answer criteria. Bo Zheng 0010, Haoyang Wen, Yaobo Liang, Nan Duan 0001, Wanxiang Che, Daxin Jiang, Ming Zhou 0001, Ting Liu 0001 |
ACL | 4 |
| 2020 | LogicalFactChecker: Leveraging Logical Operations for Fact Checking with Graph Module NetworkabstractWanjun Zhong, Duyu Tang, Zhangyin Feng, Nan Duan, Ming Zhou, Ming Gong, Linjun Shou, Daxin Jiang, Jiahai Wang, Jian Yin. Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics. 2020. Wanjun Zhong, Duyu Tang, Zhangyin Feng, Nan Duan 0001, Ming Zhou 0001, Ming Gong 0001, Linjun Shou, Daxin Jiang, Jiahai Wang, Jian Yin 0001 |
ACL | 4 |
| 2020 | Reasoning Over Semantic-Level Graph for Fact CheckingabstractFact checking is a challenging task because verifying the truthfulness of a claim requires reasoning about multiple retrievable evidence.In this work, we present a method suitable for reasoning about the semantic-level structure of evidence.Unlike most previous works, which typically represent evidence sentences with either string concatenation or fusing the features of isolated evidence sentences, our approach operates on rich semantic structures of evidence obtained by semantic role labeling.We propose two mechanisms to exploit the structure of evidence while leveraging the advances of pre-trained models like BERT, GPT or XLNet.Specifically, using XLNet as the backbone, we first utilize the graph structure to re-define the relative distances of words, with the intuition that semantically related words should have short distances.Then, we adopt graph convolutional network and graph attention network to propagate and aggregate information from neighboring nodes on the graph.We evaluate our system on FEVER, a benchmark dataset for fact checking, and find that rich structural information is helpful and both our graph-based mechanisms improve the accuracy.Our model is the state-of-the-art system in terms of both official evaluation metrics, namely claim verification accuracy and FEVER score. Wanjun Zhong, Jingjing Xu 0001, Duyu Tang, Zenan Xu, Nan Duan 0001, Ming Zhou 0001, Jiahai Wang, Jian Yin 0001 |
ACL | 5 |
| 2020 | An Enhanced Knowledge Injection Model for Commonsense GenerationabstractCommonsense generation aims at generating plausible everyday scenario description based on a set of provided concepts. Digging the relationship of concepts from scratch is non-trivial, therefore, we retrieve prototypes from external knowledge to assist the understanding of the scenario for better description generation. We integrate two additional modules into the pretrained encoder-decoder model for prototype modeling to enhance the knowledge injection procedure. We conduct experiment on CommonGen benchmark, experimental results show that our method significantly improves the performance on all the metrics. Zhihao Fan, Yeyun Gong, Zhongyu Wei, Siyuan Wang 0025, Yameng Huang, Jian Jiao 0007, Xuanjing Huang 0001, Nan Duan 0001, Ruofei Zhang |
COLING | 8 |
| 2020 | Multi-level Alignment Pretraining for Multi-lingual Semantic ParsingabstractIn this paper, we present a multi-level alignment pretraining method in a unified architecture for multi-lingual semantic parsing.In this architecture, we use an adversarial training method to align the space of different languages and use sentence level and word level parallel corpus as supervision information to align the semantic of different languages.Finally, we jointly train the multi-level alignment and semantic parsing tasks.We conduct experiments on a publicly available multi-lingual semantic parsing dataset ATIS and a newly constructed dataset.Experimental results show that our model outperforms state-of-the-art methods on both datasets. Yeyun Gong, Weizhen Qi, Nan Duan 0001, Xiaola Lin |
COLING | 4 |
| 2020 | XGLUE: A New Benchmark Datasetfor Cross-lingual Pre-training, Understanding and GenerationabstractYaobo Liang, Nan Duan, Yeyun Gong, Ning Wu, Fenfei Guo, Weizhen Qi, Ming Gong, Linjun Shou, Daxin Jiang, Guihong Cao, Xiaodong Fan, Ruofei Zhang, Rahul Agrawal, Edward Cui, Sining Wei, Taroon Bharti, Ying Qiao, Jiun-Hung Chen, Winnie Wu, Shuguang Liu, Fan Yang, Daniel Campos, Rangan Majumder, Ming Zhou. Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP). 2020. Yaobo Liang, Nan Duan 0001, Yeyun Gong, Ning Wu 0013, Fenfei Guo, Weizhen Qi, Ming Gong 0001, Linjun Shou, Daxin Jiang, Guihong Cao, Xiaodong Fan, Ruofei Zhang, Rahul Agrawal, Edward Dong Bo Cui, Sining Wei, Taroon Bharti, Jiun-Hung Chen, Winnie Wu, Fan Yang 0024, Daniel Campos, Rangan Majumder, Ming Zhou 0001 |
EMNLP (1) | 2 |
| 2020 | Tell Me How to Ask Again: Question Data Augmentation with Controllable Rewriting in Continuous SpaceabstractIn this paper, we propose a novel data augmentation method, referred to as Controllable Rewriting based Question Data Augmentation (CRQDA), for machine reading comprehension (MRC), question generation, and question-answering natural language inference tasks.We treat the question data augmentation task as a constrained question rewriting problem to generate context-relevant, high-quality, and diverse question data samples.CRQDA utilizes a Transformer autoencoder to map the original discrete question into a continuous embedding space.It then uses a pre-trained MRC model to revise the question representation iteratively with gradientbased optimization.Finally, the revised question representations are mapped back into the discrete space, which serve as additional question data.Comprehensive experiments on SQuAD 2.0, SQuAD 1.1 question generation, and QNLI tasks demonstrate the effectiveness of CRQDA 1 . Dayiheng Liu, Yeyun Gong, Jie Fu 0001, Jiusheng Chen, Jiancheng Lv 0001, Nan Duan 0001, Ming Zhou 0001 |
EMNLP (1) | 7 |
| 2020 | Diverse, Controllable, and Keyphrase-Aware: A Corpus and Method for News Multi-Headline GenerationabstractNews headline generation aims to produce a short sentence to attract readers to read the news.One news article often contains multiple keyphrases that are of interest to different users, which can naturally have multiple reasonable headlines.However, most existing methods focus on the single headline generation.In this paper, we propose generating multiple headlines with keyphrases of user interests, whose main idea is to generate multiple keyphrases of interest to users for the news first, and then generate multiple keyphrase-relevant headlines.We propose a multi-source Transformer decoder, which takes three sources as inputs: (a) keyphrase, (b) keyphrase-filtered article, and (c) original article to generate keyphrase-relevant, highquality, and diverse headlines.Furthermore, we propose a simple and effective method to mine the keyphrases of interest in the news article and build a first large-scale keyphraseaware news headline corpus, which contains over 180K aligned triples of news article, headline, keyphrase .Extensive experimental comparisons on the real-world dataset show that the proposed method achieves state-of-theart results in terms of quality and diversity 1 . Dayiheng Liu, Yeyun Gong, Jie Fu 0001, Daxin Jiang, Jiancheng Lv 0001, Nan Duan 0001 |
EMNLP (1) | 8 |
| 2020 | Leveraging Declarative Knowledge in Text and First-Order Logic for Fine-Grained Propaganda DetectionabstractRuize Wang, Duyu Tang, Nan Duan, Wanjun Zhong, Zhongyu Wei, Xuanjing Huang, Daxin Jiang, Ming Zhou. Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP). 2020. Duyu Tang, Nan Duan 0001, Wanjun Zhong, Zhongyu Wei, Xuanjing Huang 0001, Daxin Jiang, Ming Zhou 0001 |
EMNLP (1) | 3 |
| 2020 | Neural Deepfake Detection with Factual Structure of TextabstractDeepfake detection, the task of automatically discriminating machine-generated text, is increasingly critical with recent advances in natural language generative models.Existing approaches to deepfake detection typically represent documents with coarse-grained representations.However, they struggle to capture factual structures of documents, which is a discriminative factor between machinegenerated and human-written text according to our statistical analysis.To address this, we propose a graph-based model that utilizes the factual structure of a document for deepfake detection of text.Our approach represents the factual structure of a given document as an entity graph, which is further utilized to learn sentence representations with a graph neural network.Sentence representations are then composed to a document representation for making predictions, where consistent relations between neighboring sentences are sequentially modeled.Results of experiments on two public deepfake datasets show that our approach significantly improves strong base models built with RoBERTa.Model analysis further indicates that our model can distinguish the difference in the factual structure between machine-generated text and humanwritten text. Wanjun Zhong, Duyu Tang, Zenan Xu, Nan Duan 0001, Ming Zhou 0001, Jiahai Wang, Jian Yin 0001 |
EMNLP (1) | 5 |
| 2020 | Learning Semantic Concepts and Temporal Alignment for Narrated Video Procedural CaptioningabstractVideo captioning is a fundamental task for visual understanding. Previous works employ end-to-end networks to learn from the low-level vision feature and generate descriptive captions, which are hard to recognize fine-grained objects and lacks the understanding of crucial semantic concepts. According to DPC [19], these concepts generally present in the narrative transcripts of the instructional videos. The incorporation of transcript and video can improve the captioning performance. However, DPC directly concatenates the embedding of transcript with video features, which is incapable of fusing language and vision features effectively and leads to the temporal mis-alignment between transcript and video. This motivates us to 1) learn the semantic concepts explicitly and 2) design a temporal alignment mechanism to better align the video and transcript for the captioning task. In this paper, we start with an encoder-decoder backbone using transformer models. Firstly, we design a semantic concept prediction module as a multi-task to train the encoder in a supervised way. Then, we develop an attention based cross-modality temporal alignment method that combines the sequential video frames and transcript sentences. Finally, we adopt a copy mechanism to enable the decoder(generation) module to copy important concepts from source transcript directly. The extensive experimental results demonstrate the effectiveness of our model, which achieves state-of-the-art results on YouCookII dataset. Botian Shi, Lei Ji 0001, Zhendong Niu, Nan Duan 0001, Ming Zhou 0001, Xilin Chen 0001 |
ACM Multimedia | 4 |
| 2020 | ProphetNet-Ads: A Looking Ahead Strategy for Generative Retrieval Models in Sponsored Search Engine
Weizhen Qi, Yeyun Gong, Jian Jiao 0007, Ruofei Zhang, Houqiang Li, Nan Duan 0001, Ming Zhou 0001 |
NLPCC (2) | 8 |
| 2020 | Joint Learning of Question Answering and Question GenerationabstractQuestion answering (QA) and question generation (QG) are closely related tasks that could improve each other; however, the connection of these two tasks is not well explored in the literature. In this paper, we present two training algorithms for learning better QA and QG models through leveraging one another. The first algorithm extends Generative Adversarial Network (GAN), which selectively incorporates artificially generated instances as additional QA training data. The second algorithm is an extension of dual learning, which incorporates the probabilistic correlation of QA and QG as additional regularization in training objectives. To test the scalability of our algorithms, we conduct experiments on both document based and table based question answering tasks. Results show that both algorithms improve a QA model in terms of accuracy and QG model in terms of BLEU score. Moreover, we find that the performance of a QG model could be easily improved by a QA model via policy gradient, however, directly applying GAN that regards all the generated questions as negative instances could not improve the accuracy of the QA model. Our algorithm that selectively assigns labels to generated questions would bring a performance boost. Duyu Tang, Nan Duan 0001, Tao Qin 0001, Shujie Liu 0001, Ming Zhou 0001, Yuanhua Lv, Wenpeng Yin 0001, Bing Qin 0001, Ting Liu 0001 |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2019 | Coupling Retrieval and Meta-Learning for Context-Dependent Semantic Parsingabstract5 9 469 9 65 67 1 016 9 93969696 5 5 67 21 22 67 !"626 ! 5 67#$12 !6%5 6&5 6 '()*+(,,-,./0100234,5671 )8*+9 )2+):*72;01 < = )2>29 ?)8 = 9 1 @A B702C3,2CD)@E0F,8 01 ,8 @,.G9 C/01 0H20-@= 9 =023I8 ,+)= = 9 2C:B702CJ(,7:IA KA 4(9 20 !L9 +8 ,= ,. 1K)= )08 +(H= 9 0:G)9 M 9 2C:4(9 20 NOPQRSTUVWXYZ[X\\]SX^UVWXY_`\S\P`aRP`b^ NRPcW^O[^W^RPW^[VX^OdeQP_UVXbfQ\Qgc`bQV hi j 21 (9 =606)8 :k)68 )= )21020668 ,0+(1 ,9 2+,8 < 6,8 01 )8 )1 8 9 )?)3301 06,9 21 =0== 766,8 1 9 2C)?9 < 3)2+).,8+,21 )l1 < 3)6)23)21= )5021 9 +608 = 9 2C: = 7+(0=C)2)8 01 9 2C= ,78 +)+,3)+,239 1 9 ,2)3,2 1 ()+-0= =)2?9 8 ,25)21 Am780668 ,0+(201 78 0--@ +,5F9 2)=08 )1 8 9 )?0-5,3)-02305)1 0< -)08 2)8 : k()8 )1 ().,8 5)8-)08 2=1 ,n23= 9 59 -08301 < 06,9 21 =. 8 ,51 ()1 8 09 29 2C301 0:0231 ()-01 1 )8 +,2= 9 3)8 =8 )1 8 9 )?)3301 06,9 21 =0=06= )73,1 0= o .,8. 0= 103061 01 9 ,2A*6)+9 n+0--@:,788 )1 8 9 )?< )89 =0+,21 )l1 < 0k08 ))2+,3)8 < 3)+,3)85,3)-k9 1 (0-01 )21?08 9 0F-)k(9 +(1 0o)=+,21 )l1)2< ?9 8 ,25)219 21 ,+,2= 9 3)8 01 9 ,2:023,785)1 0< -)08 2)8-)08 2=1 ,71 9 -9 J)8 )1 8 9 )?)3301 06,9 21 = 9 205,3)-< 0C2,= 1 9 +5)1 0< -)08 29 2C608 039 C5 .,8. 0= 103061 01 9 ,2Ap)+,237+1)l6)8 9 5)21 = ,24mq4m/r0234*sH 301 0= )1 = :k()8 ) 1 ()+,21 )l18 ).)8 =1 ,+-0= =)2?9 8 ,25)219 2t H< uH+,3)=023+,2?)8 = 01 9 ,20-(9 = 1 ,8 @:8 )= 6)+< 1 9 ?)-@Ap)7=)=)v7)2+)< 1 ,< 0+1 9 ,25,3)-0= 1 ()F0= )=)5021 9 +608 = )8 :k(9 +(6)8 .,8 5=1 () = 1 01 )< ,. < 1 ()< 08 10++78 0+@,2F,1 (301 0= )1 = AK)< = 7-1 == (,k1 (01F,1 (1 ()+,21 )l1 < 0k08 )8 )1 8 9 )?< )80231 ()5)1 0< -)08 29 2C=1 8 01 )C@9 568 ,?)0+< +78 0+@:023,780668 ,0+(6)8 .,8 5=F)1 1 )81 (02 8 )1 8 9 )?)< 023< )39 1F0= )-9 2)= A w x 6 12 5 16 4,21 )l1 < 3)6)23)21= )5021 9 +608 = 9 2C09 5=1 ,506 0201 78 0--02C70C)71 1 )8 02+)1 ,0=1 8 7+1 78 0--,C9 < +0-.,8 5y )A CA= ,78 +)+,3)z+,239 1 9 ,2)3,20C9 ?< )2+,21 )l1y )A CA+-0= =)2?9 8 ,25)21 zy E9 2C)10-A : {|}~E,2C)10-A :{|}~j @@)8)10-A :{|}j @< )8)10-A :{|}*7(8)10-A :{|}*7(8023H8 1 J9 : {|}z A*1 02308 30668 ,0+()=1 @69 +0--@-)08 20,2)< = 9 J)< n1 = < 0--5,3)-,21 ())21 9 8 )1 8 09 29 2C301 0= )1 : k(9 +(9 =. )3k9 1 ()0+()l056-)9 239 ?9 370--@9 21 () 1 8 09 29 2C6(0= )02350o)=68 )39 +1 9 ,2=.,8)0+(1 )= 1 )l056-)9 21 ()9 2. )8 )2+)6(0= )A,k)?)8 :1 0o9 2C p,8 o3,2)k(9 -)1 (9 =071 (,8k0=029 21 )8 201L9 +8 ,= ,. 1 K)= )08 +(A +,3)C)2)8 01 9 ,20=02)l056-):68 ,C8 055)8 =7= 7< 0--@3,2,1k8 9 1 )+,3)=.8 ,5= +8 01 +(9 21 ()8 )0k,8 -3Ap()21 ()@k8 9 1 )069 )+),.+,3)920608 < 1 9 +7-08)2?9 8 ,25)21 :1 ()@1 @69 +0--@-)?)8 0C)60= 1 )l6)8 9 )2+),2k8 9 1 9 2C,88 )039 2C+,3)=9 21 ()= 9 59 < -08= 9 1 701 9 ,20=0C79 302+)AL)02k(9 -):301 06,9 21 = .,801 0= o50@?08 @k9 3)-@y 702C)10-A :{|}0z : 1 (7=9 19 =3)= 9 8 0F-)1 ,-)08 206)8 = ,20-9 J)35,3< )-.,81 ()1 08 C)1301 06,9 21 Aj 21 (9 =k,8 o:k)= 1 73@ (,k1 ,071 ,501 9 +0--@8 )1 8 9 )?)= 9 59 -08301 06,9 21 =9 2 0+,21 )l1 < 3)6)23)21= +)208 9 ,0237= )1 ()50=1 () = 766,8 1 9 2C)?9 3)2+)1 ,. 0+9 -9 1 01 )= )5021 9 +608 = 9 2CA '()8 )08 )8 )+)2101 1 )561 =01)l6-,9 1 9 2C8 )< 1 8 9 )?)3)l056-)=1 ,9 568 ,?)1 ()C)2)8 01 9 ,2,.-,C< 9 +0-.,8 50231 )l1 AK)1 8 9 )?)< 023< )39 10668 ,0+()= y 0= (9 5,1 ,)10-A :{|}702C)10-A :{|}Fp7 )10-A :{|}B7)10-A :{|}z1 @69 +0--@n8 = 17= )0 +,21 )l1 < 9 23)6)23)218 )1 8 9 )?)81 ,n231 ()5,= 18 )-< )?021301 06,9 21 :0231 ()27= )9 10=020339 1 9 ,20-9 2671,.1 ())39 1 9 2C5,3)-A,k)?)8:0+,21 )l1 < 0k08 )8 )1 8 9 )?)89 =?)8 @9 56,8 1 021.,81 ()1 0= o,.+,21 )l1 < 3)6)23)21= )5021 9 +608 = 9 2CA,8)l05< 6-)= :0== (,k29 29 C78 )}:+-0= =)2?9 8 ,25)21+02 ()-61 ()8 )1 8 9 )?)83)+9 3)k()1 ()81 ()3)= 9 8 )3+,3) ,. 9 =C)2)8 01 )3F@39 8 )+1 -@ +0--9 2C ,89 1 )8 01 9 2C1 () ¡08 8 0@ 1 ,9 2+8 )5)21)0+()-)5)21 A78 1 ()8 5,8 ):8 )1 8 9 )?)< 023< )39 10668 ,0+()=1 @69 +0--@+,2= 9 3)8,2-@,2) = 9 59 -08)l056-)1 ,)39 1 Aj 2=)5021 9 +608 = 9 2C:1 () 601 1 )8 2,.0= 1 8 7+1 78 0-,71 67150@+,5).8 ,539 .< .)8 )218 )1 8 9 )?)3)l056-)= A'()8 )0-= ,)l9 = 1k,8 o= 1 ,71 9 -9 J)57-1 9 6-))l056-)=1 ,C79 3)1 ()= )5021 9 + 608 = )8y 0@01 9)10-A :{|}702C)10-A :{|}0z : (,k)?)8 :1 ()= )0668 ,0+()=)9 1 ()87= )0()78 9 = 1 9 + k0@1 ,)l6-,9 11 ()8 )1 8 9 )?)3-,C9 +0-.,8 5= 7+(0= 9 2+8 )0= 9 2C1 ()68 ,F0F9 -9 1 @,.0+1 9 ,2=y 0@01 9)10-A : {|}z,87= )08 )-)?02+).72+1 9 ,23)= 9 C2)3023 -)08 2)3F0= )3,2)l6)8 1 9 = )0F,711 ()1 08 C)1-,C9 < +0-.,8 5y 702C)10-A :{|}0z Ap()2k)+,2= 9 3)8 1 ()+,21 )l1)2?9 8 ,25)21 :9 1 ¢ =2,21 8 9 ?9 0-1 ,3)= 9 C2 Daya Guo, Duyu Tang, Nan Duan 0001, Ming Zhou 0001, Jian Yin 0001 |
ACL (1) | 3 |
| 2019 | Dense Procedure Captioning in Narrated Instructional VideosabstractUnderstanding narrated instructional videos is important for both research and real-world web applications.Motivated by video dense captioning, we propose a model to generate procedure captions from narrated instructional videos which are a sequence of stepwise clips with description.Previous works on video dense captioning learn video segments and generate captions without considering transcripts.We argue that transcripts in narrated instructional videos can enhance video representation by providing fine-grained complimentary and semantic textual information.In this paper, we introduce a framework to ( 1) extract procedures by a cross-modality module, which fuses video content with the entire transcript; and (2) generate captions by encoding video frames as well as a snippet of transcripts within each extracted procedure.Experiments show that our model can achieve state-of-the-art performance in procedure extraction and captioning, and the ablation studies demonstrate that both the video frames and the transcripts are important for the task. Botian Shi, Lei Ji 0001, Yaobo Liang, Nan Duan 0001, Peng Chen 0029, Zhendong Niu, Ming Zhou 0001 |
ACL (1) | 4 |
| 2019 | Joint Type Inference on Entities and Relations via Graph Convolutional NetworksabstractWe develop a new paradigm for the task of joint entity relation extraction.It first identifies entity spans, then performs a joint inference on entity types and relation types.To tackle the joint type inference task, we propose a novel graph convolutional network (GCN) running on an entity-relation bipartite graph.By introducing a binary relation classification task, we are able to utilize the structure of entity-relation bipartite graph in a more efficient and interpretable way.Experiments on ACE05 show that our model outperforms existing joint models in entity performance and is competitive with the state-of-the-art in relation performance. Changzhi Sun, Yeyun Gong, Yuanbin Wu, Ming Gong 0001, Daxin Jiang, Man Lan, Shiliang Sun, Nan Duan 0001 |
ACL (1) | 8 |
| 2019 | Unicoder: A Universal Language Encoder by Pre-training with Multiple Cross-lingual TasksabstractHaoyang Huang, Yaobo Liang, Nan Duan, Ming Gong, Linjun Shou, Daxin Jiang, Ming Zhou. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Haoyang Huang, Yaobo Liang, Nan Duan 0001, Ming Gong 0001, Linjun Shou, Daxin Jiang, Ming Zhou 0001 |
EMNLP/IJCNLP (1) | 3 |
| 2019 | Aggregating Bidirectional Encoder Representations Using MatchLSTM for Sequence MatchingabstractBo Shao, Yeyun Gong, Weizhen Qi, Nan Duan, Xiaola Lin. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Yeyun Gong, Weizhen Qi, Nan Duan 0001, Xiaola Lin |
EMNLP/IJCNLP (1) | 4 |
| 2019 | Multi-Task Learning for Conversational Question Answering over a Large-Scale Knowledge BaseabstractTao Shen, Xiubo Geng, Tao Qin, Daya Guo, Duyu Tang, Nan Duan, Guodong Long, Daxin Jiang. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Tao Shen 0001, Xiubo Geng, Tao Qin 0001, Daya Guo, Duyu Tang, Nan Duan 0001, Guodong Long, Daxin Jiang |
EMNLP/IJCNLP (1) | 6 |
| 2019 | Asking Clarification Questions in Knowledge-Based Question AnsweringabstractJingjing Xu, Yuechen Wang, Duyu Tang, Nan Duan, Pengcheng Yang, Qi Zeng, Ming Zhou, Xu Sun. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Jingjing Xu 0001, Yuechen Wang, Duyu Tang, Nan Duan 0001, Qi Zeng 0001, Ming Zhou 0001, Xu Sun 0001 |
EMNLP/IJCNLP (1) | 4 |
| 2019 | Weakly Supervised Multi-task Learning for Semantic ParsingabstractSemantic parsing is a challenging and important task which aims to convert a natural language sentence to a logical form. Existing neural semantic parsing methods mainly use (Q-L) pairs to train a sequence-to-sequence model. However, the amount of existing Q-L labeled data is limited and hard to obtain. We propose an effective method which substantially utilizes labeling information from other tasks to enhance the training of a semantic parser. We design a multi-task learning model to train question type classification, entity mention detection together with question semantic parsing using a shared encoder. We propose a weakly supervised learning method to enhance our multi-task learning model with paraphrase data, based on the idea that the paraphrased questions should have the same logical form and question type information. Finally, we integrate the weakly supervised multi-task learning method to an encoder-decoder framework. Experiments on a newly constructed dataset and ComplexWebQuestions show that our proposed method outperforms state-of-the-art methods which demonstrates the effectiveness and robustness of our method. Yeyun Gong, Junwei Bao 0001, Jianshu Ji, Guihong Cao, Xiaola Lin, Nan Duan 0001 |
IJCAI | 7 |
| 2019 | Knowledge Aware Semantic Concept Expansion for Image-Text MatchingabstractImage-text matching is a vital cross-modality task in artificial intelligence and has attracted increasing attention in recent years. Existing works have shown that learning semantic concepts is useful to enhance image representation and can significantly improve the performance of both image-to-text and text-to-image retrieval. However, existing models simply detect semantic concepts from a given image, which are less likely to deal with long-tail and occlusion concepts. Frequently co-occurred concepts in the same scene, e.g. bedroom and bed, can provide common-sense knowledge to discover other semantic-related concepts. In this paper, we develop a Scene Concept Graph (SCG) by aggregating image scene graphs and extracting frequently co-occurred concept pairs as scene common-sense knowledge. Moreover, we propose a novel model to incorporate this knowledge to improve image-text matching. Specifically, semantic concepts are detected from images and then expanded by the SCG. After learning to select relevant contextual concepts, we fuse their representations with the image embedding feature to feed into the matching module. Extensive experiments are conducted on Flickr30K and MSCOCO datasets, and prove that our model achieves state-of-the-art results due to the effectiveness of incorporating the external SCG. Botian Shi, Lei Ji 0001, Pan Lu, Zhendong Niu, Nan Duan 0001 |
IJCAI | 5 |
| 2019 | PasteGAN: A Semi-Parametric Method to Generate Image from Scene GraphabstractDespite some exciting progress on high-quality image generation from structured (scene graphs) or free-form (sentences) descriptions, most of them only guarantee the image-level semantical consistency, i.e. the generated image matching the semantic meaning of the description. They still lack the investigations on synthesizing the images in a more controllable way, like finely manipulating the visual appearance of every object. Therefore, to generate the images with preferred objects and rich interactions, we propose a semi-parametric method, PasteGAN, for generating the image from the scene graph and the image crops, where spatial arrangements of the objects and their pair-wise relationships are defined by the scene graph and the object appearances are determined by the given object crops. To enhance the interactions of the objects in the output, we design a Crop Refining Network and an Object-Image Fuser to embed the objects as well as their relationships into one map. Multiple losses work collaboratively to guarantee the generated images highly respecting the crops and complying with the scene graphs while maintaining excellent image quality. A crop selector is also proposed to pick the most-compatible crops from our external object tank by encoding the interactions around the objects in the scene graph if the crops are not provided. Evaluated on Visual Genome and COCO-Stuff dataset, our proposed method significantly outperforms the SOTA methods on Inception Score, Diversity Score and Fréchet Inception Distance. Extensive experiments also demonstrate our method’s ability to generate complex and diverse images with given objects. The code is available at https://github.com/yikang-li/PasteGAN. Yikang Li 0002, Tao Ma 0002, Yeqi Bai, Nan Duan 0001, Sining Wei, Xiaogang Wang 0001 |
NeurIPS | 4 |
| 2019 | A Tensorized Transformer for Language ModelingabstractLatest development of neural models has connected the encoder and decoder through a self-attention mechanism. In particular, Transformer, which is solely based on self-attention, has led to breakthroughs in Natural Language Processing (NLP) tasks. However, the multi-head attention mechanism, as a key component of Transformer, limits the effective deployment of the model to a resource-limited setting. In this paper, based on the ideas of tensor decomposition and parameters sharing, we propose a novel self-attention model (namely Multi-linear attention) with Block-Term Tensor Decomposition (BTD). We test and verify the proposed attention method on three language modeling tasks (i.e., PTB, WikiText-103 and One-billion) and a neural machine translation task (i.e., WMT-2016 English-German). Multi-linear attention can not only largely compress the model parameters but also obtain performance improvements, compared with a number of language modeling approaches, such as Transformer, Transformer-XL, and Transformer with tensor train decomposition. Xindian Ma, Peng Zhang 0002, Nan Duan 0001, Yuexian Hou, Ming Zhou 0001, Dawei Song 0001 |
NeurIPS | 4 |
| 2019 | Overview of the NLPCC 2019 Shared Task: Open Domain Semantic Parsing
Nan Duan 0001 |
NLPCC (2) | 1 |
| 2019 | Knowledge-Aware Conversational Semantic Parsing over Web Tables
Duyu Tang, Jingjing Xu 0001, Nan Duan 0001, Bing Qin 0001, Ting Liu 0001, Ming Zhou 0001 |
NLPCC (1) | 4 |
| 2019 | Improving Question Answering by Commonsense-Based Pre-training
Wanjun Zhong, Duyu Tang, Nan Duan 0001, Ming Zhou 0001, Jiahai Wang, Jian Yin 0001 |
NLPCC (1) | 3 |
| 2019 | Content-based table retrieval for web queries
Duyu Tang, Nan Duan 0001, Bing Qin 0001 |
Neurocomputing | 4 |
| 2019 | Text Generation From TablesabstractThis paper proposes a neural generative model, namely Table2Seq, to generate a natural language sentence based on a table. Specifically, the model maps a table to continuous vectors and then generates a natural language sentence by leveraging the semantics of a table. Since rare words, e.g., entities and values, usually appear in a table, we develop a flexible copying mechanism that selectively replicates contents from the table to the output sequence. We conduct extensive experiments to demonstrate the effectiveness of our Table2Seq model and the utility of the designed copying mechanism. On the WIKIBIO and SIMPLEQUESTIONS datasets, the Table2Seq model improves the state-of-the-art results from 34.70 to 40.26 and from 33.32 to 39.12 in terms of BLEU-4 scores, respectively. Moreover, we construct an open-domain dataset WIKITABLETEXT that includes 13 318 descriptive sentences for 4962 tables. Our Table2Seq model achieves a BLEU-4 score of 38.23 on WIKITABLETEXT outperforming template-based and language model based approaches. Furthermore, through experiments on 1 M table-query pairs from a search engine, our Table2Seq model considering the structured part of a table, i.e., table attributes and table cells, as additional information outperforms a sequence-to-sequence model considering only the sequential part of a table, i.e., table caption. Junwei Bao 0001, Duyu Tang, Nan Duan 0001, Ming Zhou 0001, Tiejun Zhao |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2018 | Table-to-Text: Describing Table Region With Natural LanguageabstractIn this paper, we present a generative model to generate a natural language sentence describing a table region, e.g., a row. The model maps a row from a table to a continuous vector and then generates a natural language sentence by leveraging the semantics of a table. To deal with rare words appearing in a table, we develop a flexible copying mechanism that selectively replicates contents from the table in the output sequence. Extensive experiments demonstrate the accuracy of the model and the power of the copying mechanism. On two synthetic datasets, WIKIBIO and SIMPLEQUESTIONS, our model improves the current state-of-the-art BLEU-4 score from 34.70 to 40.26 and from 33.32 to 39.12, respectively. Furthermore, we introduce an open-domain dataset WIKITABLETEXT including 13,318 explanatory sentences for 4,962 tables. Our model achieves a BLEU-4 score of 38.23, which outperforms template based and language model based approaches. Junwei Bao 0001, Duyu Tang, Nan Duan 0001, Yuanhua Lv, Ming Zhou 0001, Tiejun Zhao |
AAAI | 3 |
| 2018 | Assertion-Based QA With Question-Aware Open Information ExtractionabstractWe present assertion based question answering (ABQA), an open domain question answering task that takes a question and a passage as inputs, and outputs a semi-structured assertion consisting of a subject, a predicate and a list of arguments. An assertion conveys more evidences than a short answer span in reading comprehension, and it is more concise than a tedious passage in passage-based QA. These advantages make ABQA more suitable for human-computer interaction scenarios such as voice-controlled speakers. Further progress towards improving ABQA requires richer supervised dataset and powerful models of text understanding. To remedy this, we introduce a new dataset called WebAssertions, which includes hand-annotated QA labels for 358,427 assertions in 55,960 web passages. To address ABQA, we develop both generative and extractive approaches. The backbone of our generative approach is sequence to sequence learning. In order to capture the structure of the output assertion, we introduce a hierarchical decoder that first generates the structure of the assertion and then generates the words of each field. The extractive approach is based on learning to rank. Features at different levels of granularity are designed to measure the semantic relevance between a question and an assertion. Experimental results show that our approaches have the ability to infer question-aware assertions from a passage. We further evaluate our approaches by incorporating the ABQA results as additional features in passage-based QA. Results on two datasets show that ABQA features significantly improve the accuracy on passage-based QA. Duyu Tang, Nan Duan 0001, Shujie Liu 0001, Daxin Jiang, Ming Zhou 0001, Zhoujun Li 0001 |
AAAI | 3 |
| 2018 | Semantic Parsing with Syntax- and Table-Aware SQL GenerationabstractYibo Sun, Duyu Tang, Nan Duan, Jianshu Ji, Guihong Cao, Xiaocheng Feng, Bing Qin, Ting Liu, Ming Zhou. Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2018. Duyu Tang, Nan Duan 0001, Jianshu Ji, Guihong Cao, Bing Qin 0001, Ting Liu 0001, Ming Zhou 0001 |
ACL (1) | 3 |
| 2018 | Visual Question Generation as Dual Task of Visual Question AnsweringabstractVisual question answering (VQA) and visual question generation (VQG) are two trending topics in the computer vision, but they are usually explored separately despite their intrinsic complementary relationship. In this paper, we propose an end-to-end unified model, the Invertible Question Answering Network (iQAN), to introduce question generation as a dual task of question answering to improve the VQA performance. With our proposed invertible bilinear fusion module and parameter sharing scheme, our iQAN can accomplish VQA and its dual task VQG simultaneously. By jointly trained on two tasks with our proposed dual regularizes (termed as Dual Training), our model has a better understanding of the interactions among images, questions and answers. After training, iQAN can take either question or answer as input, and output the counterpart. Evaluated on the CLEVR and VQA2 datasets, our iQAN improves the top-1 accuracy of the prior art MUTAN VQA method by 1.33% and 0.88% (absolute increase) respectiely. We also show that our proposed dual training framework can consistently improve model performances of many popular VQA architectures1. Yikang Li 0002, Nan Duan 0001, Bolei Zhou, Xiao Chu, Wanli Ouyang, Xiaogang Wang 0001, Ming Zhou 0001 |
CVPR | 2 |
| 2018 | Question Generation from SQL Queries Improves Neural Semantic ParsingabstractWe study how to learn a semantic parser of state-of-the-art accuracy with less supervised training data.We conduct our study on WikiSQL, the largest hand-annotated semantic parsing dataset to date.First, we demonstrate that question generation is an effective method that empowers us to learn a state-ofthe-art neural network based semantic parser with thirty percent of the supervised training data.Second, we show that applying question generation to the full supervised training data further improves the state-of-the-art model.In addition, we observe that there is a logarithmic relationship between the accuracy of a semantic parser and the amount of training data. Daya Guo, Duyu Tang, Nan Duan 0001, Jian Yin 0001, Hong Chi, James Cao, Peng Chen 0029, Ming Zhou 0001 |
EMNLP | 4 |
| 2018 | R-VQA: Learning Visual Relation Facts with Semantic Attention for Visual Question AnsweringabstractRecently, Visual Question Answering (VQA) has emerged as one of the most significant tasks in multimodal learning as it requires understanding both visual and textual modalities. Existing methods mainly rely on extracting image and question features to learn their joint feature embedding via multimodal fusion or attention mechanism. Some recent studies utilize external VQA-independent models to detect candidate entities or attributes in images, which serve as semantic knowledge complementary to the VQA task. However, these candidate entities or attributes might be unrelated to the VQA task and have limited semantic capacities. To better utilize semantic knowledge in images, we propose a novel framework to learn visual relation facts for VQA. Specifically, we build up a Relation-VQA (R-VQA) dataset based on the Visual Genome dataset via a semantic similarity module, in which each data consists of an image, a corresponding question, a correct answer and a supporting relation fact. A well-defined relation detector is then adopted to predict visual question-related relation facts. We further propose a multi-step attention model composed of visual attention and semantic attention sequentially to extract related visual knowledge and semantic knowledge. We conduct comprehensive experiments on the two benchmark datasets, demonstrating that our model achieves state-of-the-art performance and verifying the benefit of considering visual relation facts. Pan Lu, Lei Ji 0001, Wei Zhang 0056, Nan Duan 0001, Ming Zhou 0001, Jianyong Wang 0001 |
KDD | 4 |
| 2018 | Learning to Collaborate for Question Answering and AskingabstractDuyu Tang, Nan Duan, Zhao Yan, Zhirui Zhang, Yibo Sun, Shujie Liu, Yuanhua Lv, Ming Zhou. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). 2018. Duyu Tang, Nan Duan 0001, Zhirui Zhang, Shujie Liu 0001, Yuanhua Lv, Ming Zhou 0001 |
NAACL-HLT | 2 |
| 2018 | Dialog-to-Action: Conversational Question Answering Over a Large-Scale Knowledge BaseabstractWe present an approach to map utterances in conversation to logical forms, which will be executed on a large-scale knowledge base. To handle enormous ellipsis phenomena in conversation, we introduce dialog memory management to manipulate historical entities, predicates, and logical forms when inferring the logical form of current utterances. Dialog memory management is embodied in a generative model, in which a logical form is interpreted in a top-down manner following a small and flexible grammar. We learn the model from denotations without explicit annotation of logical forms, and evaluate it on a large-scale dataset consisting of 200K dialogs over 12.8M entities. Results verify the benefits of modeling dialog memory, and show that our semantic parsing-based approach outperforms a memory network based encoder-decoder model by a huge margin. Daya Guo, Duyu Tang, Nan Duan 0001, Ming Zhou 0001, Jian Yin 0001 |
NeurIPS | 3 |
| 2018 | Overview of the NLPCC 2018 Shared Task: Open Domain QA
Nan Duan 0001 |
NLPCC (2) | 1 |
| 2018 | Response selection from unstructured documents for human-computer conversation systems
Nan Duan 0001, Junwei Bao 0001, Peng Chen 0029, Ming Zhou 0001, Zhoujun Li 0001 |
Knowl. Based Syst. | 2 |
| 2018 | Question Generation With Doubly Adversarial NetsabstractWe study the problem of question generation on a specific domain, where there are no labeled data. To address this problem, we propose a novel neural question generation approach called DoubAN, or doubly adversarial nets, which fully utilizes labeled data from other domains (source domains) and unlabeled data from the target domain. Learning a DoubAN involves two adversarial procedures between a question generator and two adversaries. One adversary is a domain-classification discriminator (DC-Dis), which is designed to help the generator learn domain-general representations of the input text. The other is a question-answering discriminator (QA-Dis), which provides more training data with estimated reward scores for generated text-question pairs. We conduct experiments on the SQuAD dataset as target-domain unlabeled data and the NewsQA dataset as source-domain labeled data. Experiment results show that our DoubAN achieves better results than baselines. Compared to model variants, which adopt only DC-Dis or QA-Dis, we find that the DC-Dis and QA-Dis indirectly interact with each other and jointly improve the quality of generated questions on the target domain. Moreover, extensive analysis and discussion prove the reasonableness and effectiveness of our proposed approach. Junwei Bao 0001, Yeyun Gong, Nan Duan 0001, Ming Zhou 0001, Tiejun Zhao |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2017 | Building Task-Oriented Dialogue Systems for Online ShoppingabstractWe present a general solution towards building task-oriented dialogue systems for online shopping, aiming to assist online customers in completing various purchase-related tasks, such as searching products and answering questions, in a natural language conversation manner. As a pioneering work, we show what & how existing NLP techniques, data resources, and crowdsourcing can be leveraged to build such task-oriented dialogue systems for E-commerce usage. To demonstrate its effectiveness, we integrate our system into a mobile online shopping app. To the best of our knowledge, this is the first time that an AI bot in Chinese is practically used in online shopping scenario with millions of real consumers. Interesting and insightful observations are shown in the experimental part, based on the analysis of human-bot conversation log. Several current challenges are also pointed out as our future directions. Nan Duan 0001, Peng Chen 0029, Ming Zhou 0001, Jianshe Zhou, Zhoujun Li 0001 |
AAAI | 2 |
| 2017 | Question Generation for Question AnsweringabstractThis paper presents how to generate questions from given passages using neural networks, where large scale QA pairs are automatically crawled and processed from Community-QA website, and used as training data.The contribution of the paper is 2-fold: First, two types of question generation approaches are proposed, one is a retrieval-based method using convolution neural network (CNN), the other is a generation-based method using recurrent neural network (RNN); Second, we show how to leverage the generated questions to improve existing question answering systems.We evaluate our question generation method for the answer sentence selection task on three benchmark datasets, including SQuAD, MS MARCO, and WikiQA.Experimental results show that, by using generated questions as an extra signal, significant QA improvement can be achieved. Nan Duan 0001, Duyu Tang, Peng Chen 0029, Ming Zhou 0001 |
EMNLP | 1 |
| 2017 | An Information Retrieval-Based Approach to Table-Based Question Answering
Junwei Bao 0001, Nan Duan 0001, Ming Zhou 0001, Tiejun Zhao |
NLPCC | 2 |
| 2017 | Overview of the NLPCC 2017 Shared Task: Open Domain Chinese Question Answering
Nan Duan 0001, Duyu Tang |
NLPCC | 1 |
| 2016 | DocChat: An Information Retrieval Approach for Chatbot Engines Using Unstructured DocumentsabstractMost current chatbot engines are designed to reply to user utterances based on existing utterance-response (or Q-R) 1 pairs.In this paper, we present DocChat, a novel information retrieval approach for chatbot engines that can leverage unstructured documents, instead of Q-R pairs, to respond to utterances.A learning to rank model with features designed at different levels of granularity is proposed to measure the relevance between utterances and responses directly.We evaluate our proposed approach in both English and Chinese: (i) For English, we evaluate Doc-Chat on WikiQA and QASent, two answer sentence selection tasks, and compare it with state-of-the-art methods.Reasonable improvements and good adaptability are observed.(ii) For Chinese, we compare DocChat with XiaoIce 2 , a famous chitchat engine in China, and side-by-side evaluation shows that DocChat is a perfect complement for chatbot engines using Q-R pairs as main source of responses. Nan Duan 0001, Junwei Bao 0001, Peng Chen 0029, Ming Zhou 0001, Zhoujun Li 0001, Jianshe Zhou |
ACL (1) | 2 |
| 2016 | Constraint-Based Question Answering with Knowledge GraphabstractWebQuestions and SimpleQuestions are two benchmark data-sets commonly used in recent knowledge-based question answering (KBQA) work. Most questions in them are ‘simple’ questions which can be answered based on a single relation in the knowledge base. Such data-sets lack the capability of evaluating KBQA systems on complicated questions. Motivated by this issue, we release a new data-set, namely ComplexQuestions, aiming to measure the quality of KBQA systems on ‘multi-constraint’ questions which require multiple knowledge base relations to get the answer. Beside, we propose a novel systematic KBQA approach to solve multi-constraint questions. Compared to state-of-the-art methods, our approach not only obtains comparable results on the two existing benchmark data-sets, but also achieves significant improvements on the ComplexQuestions. Junwei Bao 0001, Nan Duan 0001, Ming Zhou 0001, Tiejun Zhao |
COLING | 2 |
| 2015 | Answering Questions with Complex Semantic Constraints on Open Knowledge BasesabstractA knowledge-based question-answering system (KB-QA) is one that answers natural language questions with information stored in a large-scale knowledge base (KB). Existing KB-QA systems are either powered by curated KBs in which factual knowledge is encoded in entities and relations with well-structured schemas, or by open KBs, which contain assertions represented in the form of triples (e.g., subject; relation phrase; argument). We show that both approaches fall short in answering questions with complex prepositional or adverbial constraints. We propose using n-tuple assertions, which are assertions with an arbitrary number of arguments, and n-tuple open KB (nOKB), which is an open knowledge base of n-tuple assertions. We present TAQA, a novel KB-QA system that is based on an nOKB and illustrate via experiments how TAQA can effectively answer complex questions with rich semantic constraints. Our work also results in a new open KB containing 120M n-tuple assertions and a collection of 300 labeled complex questions, which is made publicly available for further research. Nan Duan 0001, Ben Kao, Junwei Bao 0001, Ming Zhou 0001 |
CIKM | 2 |
| 2015 | Overview of the NLPCC 2015 Shared Task: Open Domain QAabstractIn this paper, we give the overview of the open domain Question Answering (or open domain QA) shared task in NLPCC 2015. We first review the background of QA, and then describe open domain QA shared task in this year’s NLPCC, including the construction of the benchmark datasets, the auxiliary dataset, and the evaluation metrics. The evaluation results of submissions from participating teams are presented in the experimental part, together with a brief introduction to the techniques used in each participating team’s QA system. Nan Duan 0001 |
NLPCC | 1 |
| 2014 | Knowledge-Based Question Answering as Machine TranslationabstractA typical knowledge-based question answering (KB-QA) system faces two challenges: one is to transform natural language questions into their meaning representations (MRs); the other is to retrieve answers from knowledge bases (KBs) using generated MRs.Unlike previous methods which treat them in a cascaded manner, we present a translation-based approach to solve these two tasks in one unified framework.We translate questions to answers based on CYK parsing.Answers as translations of the span covered by each CYK cell are obtained by a question translation method, which first generates formal triple queries as MRs for the span based on question patterns and relation expressions, and then retrieves answers from a given KB based on triple queries generated.A linear model is defined over derivations, and minimum error rate training is used to tune feature weights based on a set of question-answer pairs.Compared to a KB-QA system using a state-of-the-art semantic parser, our method achieves better results. Junwei Bao 0001, Nan Duan 0001, Ming Zhou 0001, Tiejun Zhao |
ACL (1) | 2 |
| 2014 | Joint Relational Embeddings for Knowledge-based Question AnsweringabstractTransforming a natural language (NL) question into a corresponding logical form (LF) is central to the knowledge-based question answering (KB-QA) task.Unlike most previous methods that achieve this goal based on mappings between lexicalized phrases and logical predicates, this paper goes one step further and proposes a novel embedding-based approach that maps NL-questions into LFs for KB-QA by leveraging semantic associations between lexical representations and KBproperties in the latent space.Experimental results demonstrate that our proposed method outperforms three KB-QA baseline methods on two publicly released QA data sets. Nan Duan 0001, Ming Zhou 0001, Hae-Chang Rim |
EMNLP | 2 |
| 2013 | Answer Extraction from Passage Graph for Question Answering
Nan Duan 0001, Yajuan Duan, Ming Zhou 0001 |
IJCAI | 2 |
| 2012 | Forced Derivation Tree based Model Training to Statistical Machine Translation
Nan Duan 0001, Mu Li 0001, Ming Zhou 0001 |
EMNLP-CoNLL | 1 |
| 2011 | Hypothesis Mixture Decoding for Statistical Machine Translation
Nan Duan 0001, Mu Li 0001, Ming Zhou 0001 |
ACL | 1 |
| 2011 | Improving Phrase Extraction via MBR Phrase Scoring and Pruning
Nan Duan 0001, Mu Li 0001, Ming Zhou 0001, Lei Cui 0001 |
MTSummit | 1 |
| 2010 | Mixture Model-based Minimum Bayes Risk Decoding using Multiple Machine Translation Systems
Nan Duan 0001, Mu Li 0001, Dongdong Zhang 0001, Ming Zhou 0001 |
COLING | 1 |
| 2010 | Translation Model Generalization using Probability Averaging for Machine Translation
Nan Duan 0001, Ming Zhou 0001 |
COLING | 1 |
| 2009 | Collaborative Decoding: Partial Hypothesis Re-ranking Using Translation Consensus between Decoders
Mu Li 0001, Nan Duan 0001, Dongdong Zhang 0001, Chi-Ho Li, Ming Zhou 0001 |
ACL/IJCNLP | 2 |
| 2009 | The Feature Subspace Method for SMT System Combination
Nan Duan 0001, Mu Li 0001, Tong Xiao 0001, Ming Zhou 0001 |
EMNLP | 1 |
| 2008 | Measure Word Generation for English-Chinese SMT Systems
Dongdong Zhang 0001, Mu Li 0001, Nan Duan 0001, Chi-Ho Li, Ming Zhou 0001 |
ACL | 3 |