VLDB 2026 Research / reviewers in the wild / expert
Lei Cui 0001
dblp:47/5523-1
· DBLP profile ↗
33ranked-venue papers
5as first author
13since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 31 · 5 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 1 first-author · 5 since 2021Databases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | MMLU-CF: A Contamination-free Multi-task Language Understanding BenchmarkabstractMultiple-choice question (MCQ) datasets like Massive Multitask Language Understanding (MMLU) are widely used to evaluate the commonsense, understanding, and problem-solving abilities of large language models (LLMs). However, the open-source nature of these benchmarks and the broad sources of training data for LLMs have inevitably led to benchmark contamination, resulting in unreliable evaluation. To alleviate this issue, we propose the contamination-free MCQ benchmark called MMLU-CF, which reassesses LLMs’ understanding of world knowledge by averting both unintentional and malicious data contamination. To mitigate unintentional data contamination, we source questions from a broader domain of over 200 billion webpages and apply three specifically designed decontamination rules. To prevent malicious data contamination, we divide the benchmark into validation and test sets with similar difficulty and subject distributions. The test set remains closed-source to ensure reliable results, while the validation set is publicly available to promote transparency and facilitate independent evaluation. The performance gap between these two sets of LLMs will indicate the contamination degree on the validation set in the future. We evaluated over 40 mainstream LLMs on the MMLU-CF. Compared to the original MMLU, not only LLMs’ performances significantly dropped but also the performance rankings of them changed considerably. This indicates the effectiveness of our approach in establishing a contamination-free and fairer evaluation standard. Qihao Zhao, Yangyu Huang, Tengchao Lv, Lei Cui 0001, Qinzheng Sun, Shaoguang Mao, Qiufeng Yin, Scarlett Li, Furu Wei |
ACL (1) | 4 |
| 2025 | PEACE: Empowering Geologic Map Holistic Understanding with MLLMsabstractGeologic map, as a fundamental diagram in geology science, provides critical insights into the structure and composition of Earth’s subsurface and surface. These maps are indispensable in various fields, including disaster assessment, resource exploration, and civil engineering. Despite their significance, current Multimodal Large Language Models (MLLMs) often fall short in geologic map understanding. This gap is primarily due to the challenging nature of cartographic generalization, which involves handling high-resolution map, managing multiple associated components, and requiring domain-specific knowledge. To quantify this gap, we construct GeoMap-Bench, the first-ever benchmark for evaluating MLLMs in geologic map understanding, which assesses the full-scale abilities in extracting, referring, grounding, reasoning, and analyzing. To bridge this gap, we introduce GeoMap-Agent, the inaugural agent designed for geologic map understanding, which features three modules: Hierarchical Information Extraction (HIE), Domain Knowledge Injection (DKI), and Prompt-enhanced Question Answering (PEQA). Inspired by the interdisciplinary collaboration among human scientists, an AI expert group acts as consultants, utilizing a diverse tool pool to comprehensively analyze questions. Through comprehensive experiments, GeoMap-Agent achieves an overall score of 0.811 on GeoMap-Bench, significantly outperforming 0.369 of GPT-4o. Our work, emPowering gEologic mAp holistiC undErstanding (PEACE) with MLLMs, paves the way for advanced AI applications in geology, enhancing the efficiency and accuracy of geological investigations. The code and data are available at https://github.com/microsoft/PEACE. Yangyu Huang, Qihao Zhao, Zhipeng Gui, Tengchao Lv, Lei Cui 0001, Scarlett Li, Furu Wei |
CVPR | 9 |
| 2025 | Think Only When You Need with Large Hybrid-Reasoning ModelsabstractRecent Large Reasoning Models (LRMs) have shown substantially improved reasoning capabilities over traditional Large Language Models (LLMs) by incorporating extended thinking processes prior to producing final responses. However, excessively lengthy thinking introduces substantial overhead in terms of token consumption and latency, which is unnecessary for simple queries. In this work, we introduce Large Hybrid-Reasoning Models (LHRMs), the first kind of model capable of adaptively determining whether to perform reasoning based on the contextual information of user queries. To achieve this, we propose a two-stage training pipeline comprising Hybrid Fine-Tuning (HFT) as a cold start, followed by online reinforcement learning with the proposed Hybrid Group Policy Optimization (HGPO) to implicitly learn to select the appropriate reasoning mode. Furthermore, we introduce a metric called Hybrid Accuracy to quantitatively assess the model’s capability for hybrid reasoning. Extensive experimental results show that LHRMs can adaptively perform hybrid reasoning on queries of varying difficulty and type. It outperforms existing LRMs and LLMs in reasoning and general capabilities while significantly improving efficiency. Together, our work advocates for a reconsideration of the appropriate use of extended reasoning processes and provides a solid starting point for building hybrid reasoning systems. Lingjie Jiang, Shaohan Huang, Qingxiu Dong, Zewen Chi, Li Dong 0004, Xingxing Zhang 0002, Tengchao Lv, Lei Cui 0001, Furu Wei |
NeurIPS | 9 |
| 2024 | TextDiffuser-2: Unleashing the Power of Language Models for Text Rendering
Jingye Chen, Yupan Huang, Tengchao Lv, Lei Cui 0001, Qifeng Chen 0001, Furu Wei |
ECCV (5) | 4 |
| 2024 | Mind's Eye of LLMs: Visualization-of-Thought Elicits Spatial Reasoning in Large Language ModelsabstractLarge language models (LLMs) have exhibited impressive performance in language comprehension and various reasoning tasks. However, their abilities in spatial reasoning, a crucial aspect of human cognition, remain relatively unexplored. Human possess a remarkable ability to create mental images of unseen objects and actions through a process known as the Mind's Eye, enabling the imagination of the unseen world. Inspired by this cognitive capacity, we propose Visualization-of-Thought (VoT) prompting. VoT aims to elicit spatial reasoning of LLMs by visualizing their reasoning traces, thereby guiding subsequent reasoning steps. We employed VoT for multi-hop spatial reasoning tasks, including natural language navigation, visual navigation, and visual tiling in 2D grid worlds. Experimental results demonstrated that VoT significantly enhances the spatial reasoning abilities of LLMs. Notably, VoT outperformed existing multimodal large language models (MLLMs) in these tasks. While VoT works surprisingly well on LLMs, the ability to generate mental images to facilitate spatial reasoning resembles the mind's eye process, suggesting its potential viability in MLLMs. Please find the dataset and codes in our [project page](https://microsoft.github.io/visualization-of-thought). Wenshan Wu, Shaoguang Mao, Yan Xia 0005, Li Dong 0004, Lei Cui 0001, Furu Wei |
NeurIPS | 6 |
| 2023 | TrOCR: Transformer-Based Optical Character Recognition with Pre-trained ModelsabstractText recognition is a long-standing research problem for document digitalization. Existing approaches are usually built based on CNN for image understanding and RNN for char-level text generation. In addition, another language model is usually needed to improve the overall accuracy as a post-processing step. In this paper, we propose an end-to-end text recognition approach with pre-trained image Transformer and text Transformer models, namely TrOCR, which leverages the Transformer architecture for both image understanding and wordpiece-level text generation. The TrOCR model is simple but effective, and can be pre-trained with large-scale synthetic data and fine-tuned with human-labeled datasets. Experiments show that the TrOCR model outperforms the current state-of-the-art models on the printed, handwritten and scene text recognition tasks. The TrOCR models and code are publicly available at https://aka.ms/trocr. Minghao Li 0004, Tengchao Lv, Jingye Chen, Lei Cui 0001, Yijuan Lu, Dinei A. F. Florêncio, Cha Zhang, Zhoujun Li 0001, Furu Wei |
AAAI | 4 |
| 2023 | TextDiffuser: Diffusion Models as Text PaintersabstractDiffusion models have gained increasing attention for their impressive generation abilities but currently struggle with rendering accurate and coherent text. To address this issue, we introduce TextDiffuser, focusing on generating images with visually appealing text that is coherent with backgrounds. TextDiffuser consists of two stages: first, a Transformer model generates the layout of keywords extracted from text prompts, and then diffusion models generate images conditioned on the text prompt and the generated layout. Additionally, we contribute the first large-scale text images dataset with OCR annotations, MARIO-10M, containing 10 million image-text pairs with text recognition, detection, and character-level segmentation annotations. We further collect the MARIO-Eval benchmark to serve as a comprehensive tool for evaluating text rendering quality. Through experiments and user studies, we demonstrate that TextDiffuser is flexible and controllable to create high-quality text images using text prompts alone or together with text template images, and conduct text inpainting to reconstruct incomplete images with text. We will make the code, model and dataset publicly available. Jingye Chen, Yupan Huang, Tengchao Lv, Lei Cui 0001, Qifeng Chen 0001, Furu Wei |
NeurIPS | 4 |
| 2023 | Language Is Not All You Need: Aligning Perception with Language ModelsabstractA big convergence of language, multimodal perception, action, and world modeling is a key step toward artificial general intelligence. In this work, we introduce KOSMOS-1, a Multimodal Large Language Model (MLLM) that can perceive general modalities, learn in context (i.e., few-shot), and follow instructions (i.e., zero-shot). Specifically, we train KOSMOS-1 from scratch on web-scale multi-modal corpora, including arbitrarily interleaved text and images, image-caption pairs, and text data. We evaluate various settings, including zero-shot, few-shot, and multimodal chain-of-thought prompting, on a wide range of tasks without any gradient updates or finetuning. Experimental results show that KOSMOS-1 achieves impressive performance on (i) language understanding, generation, and even OCR-free NLP (directly fed with document images), (ii) perception-language tasks, including multimodal dialogue, image captioning, visual question answering, and (iii) vision tasks, such as image recognition with descriptions (specifying classification via text instructions). We also show that MLLMs can benefit from cross-modal transfer, i.e., transfer knowledge from language to multimodal, and from multimodal to language. In addition, we introduce a dataset of Raven IQ test, which diagnoses the nonverbal reasoning capability of MLLMs. Shaohan Huang, Li Dong 0004, Wenhui Wang 0003, Yaru Hao, Saksham Singhal, Shuming Ma, Tengchao Lv, Lei Cui 0001, Owais Khan Mohammed, Barun Patra, Kriti Aggarwal, Zewen Chi, Johan Bjorck, Vishrav Chaudhary, Subhojit Som, Furu Wei |
NeurIPS | 8 |
| 2022 | MarkupLM: Pre-training of Text and Markup Language for Visually Rich Document UnderstandingabstractMultimodal pre-training with text, layout, and image has made significant progress for Visually Rich Document Understanding (VRDU), especially the fixed-layout documents such as scanned document images.While, there are still a large number of digital documents where the layout information is not fixed and needs to be interactively and dynamically rendered for visualization, making existing layout-based pre-training approaches not easy to apply.In this paper, we propose MarkupLM for document understanding tasks with markup languages as the backbone, such as HTML/XMLbased documents, where text and markup information is jointly pre-trained.Experiment results show that the pre-trained MarkupLM significantly outperforms the existing strong baseline models on several document understanding tasks.The pre-trained model and code will be publicly available at https:// aka.ms/markuplm. Yiheng Xu, Lei Cui 0001, Furu Wei |
ACL (1) | 3 |
| 2022 | LayoutLMv3: Pre-training for Document AI with Unified Text and Image MaskingabstractSelf-supervised pre-training techniques have achieved remarkable progress in Document AI. Most multimodal pre-trained models use a masked language modeling objective to learn bidirectional representations on the text modality, but they differ in pre-training objectives for the image modality. This discrepancy adds difficulty to multimodal representation learning. In this paper, we propose LayoutLMv3 to pre-train multimodal Transformers for Document AI with unified text and image masking. Additionally, LayoutLMv3 is pre-trained with a word-patch alignment objective to learn cross-modal alignment by predicting whether the corresponding image patch of a text word is masked. The simple unified architecture and training objectives make LayoutLMv3 a general-purpose pre-trained model for both text-centric and image-centric Document AI tasks. Experimental results show that LayoutLMv3 achieves state-of-the-art performance not only in text-centric tasks, including form understanding, receipt understanding, and document visual question answering, but also in image-centric tasks such as document image classification and document layout analysis. The code and models are publicly available at https://aka.ms/layoutlmv3. Yupan Huang, Tengchao Lv, Lei Cui 0001, Yutong Lu, Furu Wei |
ACM Multimedia | 3 |
| 2022 | DiT: Self-supervised Pre-training for Document Image TransformerabstractImage Transformer has recently achieved significant progress for natural image understanding, either using supervised (ViT, DeiT, etc.) or self-supervised (BEiT, MAE, etc.) pre-training techniques. In this paper, we propose DiT, a self-supervised pre-trained Document Image Transformer model using large-scale unlabeled text images for Document AI tasks, which is essential since no supervised counterparts ever exist due to the lack of human-labeled document images. We leverage DiT as the backbone network in a variety of vision-based Document AI tasks, including document image classification, document layout analysis, table detection as well as text detection for OCR. Experiment results have illustrated that the self-supervised pre-trained DiT model achieves new state-of-the-art results on these downstream tasks, e.g. document image classification (91.11 - 92.69), document layout analysis (91.0 - 94.9), table detection (94.23 - 96.55) and text detection for OCR (93.07 - 94.29). The code and pre-trained models are publicly available at https://aka.ms/msdit. Yiheng Xu, Tengchao Lv, Lei Cui 0001, Cha Zhang, Furu Wei |
ACM Multimedia | 4 |
| 2021 | LayoutLMv2: Multi-modal Pre-training for Visually-rich Document UnderstandingabstractYang Xu, Yiheng Xu, Tengchao Lv, Lei Cui, Furu Wei, Guoxin Wang, Yijuan Lu, Dinei Florencio, Cha Zhang, Wanxiang Che, Min Zhang, Lidong Zhou. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Yang Xu 0049, Yiheng Xu, Tengchao Lv, Lei Cui 0001, Furu Wei, Yijuan Lu, Dinei A. F. Florêncio, Cha Zhang, Wanxiang Che, Min Zhang 0005, Lidong Zhou |
ACL/IJCNLP (1) | 4 |
| 2021 | LayoutReader: Pre-training of Text and Layout for Reading Order DetectionabstractReading order detection is the cornerstone to understanding visually-rich documents (e.g., receipts and forms).Unfortunately, no existing work took advantage of advanced deep learning models because it is too laborious to annotate a large enough dataset.We observe that the reading order of WORD documents is embedded in their XML metadata; meanwhile, it is easy to convert WORD documents to PDFs or images.Therefore, in an automated manner, we construct ReadingBank, a benchmark dataset that contains reading order, text, and layout information for 500,000 document images covering a wide spectrum of document types.This first-ever large-scale dataset unleashes the power of deep neural networks for reading order detection.Specifically, our proposed LayoutReader captures the text and layout information for reading order prediction using the seq2seq model.It performs almost perfectly in reading order detection and significantly improves both open-source and commercial OCR engines in ordering text lines in their results in our experiments.The dataset and models are publicly available at https: //aka.ms/layoutreader. Zilong Wang 0002, Yiheng Xu, Lei Cui 0001, Jingbo Shang, Furu Wei |
EMNLP (1) | 3 |
| 2020 | Unsupervised Fine-tuning for Text ClusteringabstractFine-tuning with pre-trained language models (e.g.BERT) has achieved great success in many language understanding tasks in supervised settings (e.g.text classification).However, relatively little work has been focused on applying pre-trained models in unsupervised settings, such as text clustering.In this paper, we propose a novel method to fine-tune pre-trained models unsupervisedly for text clustering, which simultaneously learns text representations and cluster assignments using a clustering oriented loss.Experiments on three text clustering datasets (namely TREC-6, Yelp, and DBpedia) show that our model outperforms the baseline methods and achieves stateof-the-art results. Shaohan Huang, Furu Wei, Lei Cui 0001, Xingxing Zhang 0002, Ming Zhou 0001 |
COLING | 3 |
| 2020 | DocBank: A Benchmark Dataset for Document Layout AnalysisabstractDocument layout analysis usually relies on computer vision models to understand documents while ignoring textual information that is vital to capture.Meanwhile, high quality labeled datasets with both visual and textual information are still insufficient.In this paper, we present DocBank, a benchmark dataset that contains 500K document pages with fine-grained tokenlevel annotations for document layout analysis.DocBank is constructed using a simple yet effective way with weak supervision from the L A T E X documents available on the arXiv.com.With DocBank, models from different modalities can be compared fairly and multi-modal approaches will be further investigated and boost the performance of document layout analysis.We build several strong baselines and manually split train/dev/test sets for evaluation.Experiment results show that models trained on DocBank accurately recognize the layout information for a variety of documents.The DocBank dataset is publicly available at https: //github.com/doc-analysis/DocBank. Minghao Li 0004, Yiheng Xu, Lei Cui 0001, Shaohan Huang, Furu Wei, Zhoujun Li 0001, Ming Zhou 0001 |
COLING | 3 |
| 2020 | Multimodal Matching Transformer for Live CommentingabstractAutomatic live commenting aims to provide real-time comments on videos for viewers. It encourages users engagement on online video sites, and is also a good benchmark for video-to-text generation. Recent work on this task adopts encoder-decoder models to generate comments. However, these methods do not model the interaction between videos and comments explicitly, so they tend to generate popular comments that are often irrelevant to the videos. In this work, we aim to improve the relevance between live comments and videos by modeling the cross-modal interactions among different modalities. To this end, we propose a multimodal matching transformer to capture the relationships among comments, vision, and audio. The proposed model is based on the transformer framework and can iteratively learn the attention-aware representations for each modality. We evaluate the model on a publicly available live commenting dataset. Experiments show that the multimodal matching transformer model outperforms the state-of-the-art methods. Chaoqun Duan, Lei Cui 0001, Shuming Ma, Furu Wei, Conghui Zhu, Tiejun Zhao |
ECAI | 2 |
| 2020 | LayoutLM: Pre-training of Text and Layout for Document Image UnderstandingabstractPre-training techniques have been verified successfully in a variety of NLP tasks in recent years. Despite the widespread use of pre-training models for NLP applications, they almost exclusively focus on text-level manipulation, while neglecting layout and style information that is vital for document image understanding. In this paper, we propose the LayoutLM to jointly model interactions between text and layout information across scanned document images, which is beneficial for a great number of real-world document image understanding tasks such as information extraction from scanned documents. Furthermore, we also leverage image features to incorporate words' visual information into LayoutLM. To the best of our knowledge, this is the first time that text and layout are jointly learned in a single framework for document-level pre-training. It achieves new state-of-the-art results in several downstream tasks, including form understanding (from 70.72 to 79.27), receipt understanding (from 94.02 to 95.24) and document image classification (from 93.07 to 94.42). The code and pre-trained LayoutLM models are publicly available at https://aka.ms/layoutlm. Yiheng Xu, Minghao Li 0004, Lei Cui 0001, Shaohan Huang, Furu Wei, Ming Zhou 0001 |
KDD | 3 |
| 2020 | TableBank: Table Benchmark for Image-based Table Detection and RecognitionabstractWe present TableBank, a new image-based table detection and recognition dataset built with novel weak supervision from Word and Latex documents on the internet. Existing research for image-based table detection and recognition usually fine-tunes pre-trained models on out-of-domain data with a few thousand human-labeled examples, which is difficult to generalize on real-world applications. With TableBank that contains 417K high quality labeled tables, we build several strong baselines using state-of-the-art models with deep neural networks. We make TableBank publicly available and hope it will empower more deep learning approaches in the table detection and recognition task. The dataset and models can be downloaded from https://github.com/doc-analysis/TableBank. Minghao Li 0004, Lei Cui 0001, Shaohan Huang, Furu Wei, Ming Zhou 0001, Zhoujun Li 0001 |
LREC | 2 |
| 2019 | LiveBot: Generating Live Video Comments Based on Visual and Textual ContextsabstractWe introduce the task of automatic live commenting. Live commenting, which is also called “video barrage”, is an emerging feature on online video sites that allows real-time comments from viewers to fly across the screen like bullets or roll at the right side of the screen. The live comments are a mixture of opinions for the video and the chit chats with other comments. Automatic live commenting requires AI agents to comprehend the videos and interact with human viewers who also make the comments, so it is a good testbed of an AI agent’s ability to deal with both dynamic vision and language. In this work, we construct a large-scale live comment dataset with 2,361 videos and 895,929 live comments. Then, we introduce two neural models to generate live comments based on the visual and textual contexts, which achieve better performance than previous neural baselines such as the sequence-to-sequence model. Finally, we provide a retrieval-based evaluation protocol for automatic live commenting where the model is asked to sort a set of candidate comments based on the log-likelihood score, and evaluated on metrics such as mean-reciprocal-rank. Putting it all together, we demonstrate the first “LiveBot”. The datasets and the codes can be found at https://github.com/lancopku/livebot. Shuming Ma, Lei Cui 0001, Damai Dai, Furu Wei, Xu Sun 0001 |
AAAI | 2 |
| 2019 | Retrieval-Enhanced Adversarial Training for Neural Response GenerationabstractDialogue systems are usually built on either generation-based or retrieval-based approaches, yet they do not benefit from the advantages of different models.In this paper, we propose a Retrieval-Enhanced Adversarial Training (REAT) method for neural response generation.Distinct from existing approaches, the REAT method leverages an encoder-decoder framework in terms of an adversarial training paradigm, while taking advantage of N-best response candidates from a retrieval-based system to construct the discriminator.An empirical study on a large scale public available benchmark dataset shows that the REAT method significantly outperforms the vanilla Seq2Seq model as well as the conventional adversarial training approach. Qingfu Zhu, Lei Cui 0001, Weinan Zhang 0003, Furu Wei, Ting Liu 0001 |
ACL (1) | 2 |
| 2019 | Neural Melody Composition from Lyrics
Hangbo Bao, Shaohan Huang, Furu Wei, Lei Cui 0001, Yu Wu 0012, Chuanqi Tan, Ming Zhou 0001 |
NLPCC (1) | 4 |
| 2018 | Fine-grained Coordinated Cross-lingual Text Stream Alignment for Endless Language Knowledge AcquisitionabstractThis paper proposes to study fine-grained coordinated cross-lingual text stream alignment through a novel information network decipherment paradigm.We use Burst Information Networks as media to represent text streams and present a simple yet effective network decipherment algorithm with diverse clues to decipher the networks for accurate text stream alignment.Experiments on Chinese-English news streams show our approach not only outperforms previous approaches on bilingual lexicon extraction from coordinated text streams but also can harvest high-quality alignments from large amounts of streaming data for endless language knowledge mining, which makes it promising to be a new paradigm for automatic language knowledge acquisition. Tao Ge 0001, Qing Dou, Heng Ji 0001, Lei Cui 0001, Baobao Chang, Zhifang Sui, Furu Wei, Ming Zhou 0001 |
EMNLP | 4 |
| 2018 | Attention-Fused Deep Matching Network for Natural Language InferenceabstractNatural language inference aims to predict whether a premise sentence can infer another hypothesis sentence. Recent progress on this task only relies on a shallow interaction between sentence pairs, which is insufficient for modeling complex relations. In this paper, we present an attention-fused deep matching network (AF-DMN) for natural language inference. Unlike existing models, AF-DMN takes two sentences as input and iteratively learns the attention-aware representations for each side by multi-level interactions. Moreover, we add a self-attention mechanism to fully exploit local context information within each sentence. Experiment results show that AF-DMN achieves state-of-the-art performance and outperforms strong baselines on Stanford natural language inference (SNLI), multi-genre natural language inference (MultiNLI), and Quora duplicate questions datasets. Chaoqun Duan, Lei Cui 0001, Xinchi Chen, Furu Wei, Conghui Zhu, Tiejun Zhao |
IJCAI | 2 |
| 2018 | EventWiki: A Knowledge Base of Major Events
Tao Ge 0001, Lei Cui 0001, Baobao Chang, Zhifang Sui, Furu Wei, Ming Zhou 0001 |
LREC | 2 |
| 2018 | SeRI: A Dataset for Sub-event Relation Inference from an Encyclopedia
Tao Ge 0001, Lei Cui 0001, Baobao Chang, Zhifang Sui, Furu Wei, Ming Zhou 0001 |
NLPCC (2) | 2 |
| 2016 | Event Detection with Burst Information NetworksabstractRetrospective event detection is an important task for discovering previously unidentified events in a text stream. In this paper, we propose two fast centroid-aware event detection models based on a novel text stream representation – Burst Information Networks (BINets) for addressing the challenge. The BINets are time-aware, efficient and can be easily analyzed for identifying key information (centroids). These advantages allow the BINet-based approaches to achieve the state-of-the-art performance on multiple datasets, demonstrating the efficacy of BINets for the task of event detection. Tao Ge 0001, Lei Cui 0001, Baobao Chang, Zhifang Sui, Ming Zhou 0001 |
COLING | 2 |
| 2016 | News Stream Summarization using Burst Information NetworksabstractThis paper studies summarizing key information from news streams. We propose simple yet effective models to solve the problem based on a novel and promising representation of text streams – Burst Information Networks (BINets). A BINet can be aware of redundant information, allows global analysis of a text stream, and can be efficiently built and dynamically updated, which perfectly fits the demands of text stream summarization. Extensive experiments show that the BINet-based approaches are not only efficient and can be used in a real-time online summarization setting, but also can generate high-quality summaries, outperforming the state-of-the-art approach. Tao Ge 0001, Lei Cui 0001, Baobao Chang, Sujian Li, Ming Zhou 0001, Zhifang Sui |
EMNLP | 2 |
| 2014 | Machine Translation with Real-Time Web Search
Lei Cui 0001, Ming Zhou 0001, Dongdong Zhang 0001, Mu Li 0001 |
AAAI | 1 |
| 2014 | Learning Topic Representation for SMT with Neural NetworksabstractStatistical Machine Translation (SMT) usually utilizes contextual information to disambiguate translation candidates. However, it is often limited to contexts within sentence boundaries, hence broader topical information cannot be leveraged. In this paper, we propose a novel approach to learning topic representation for paral-lel data using a neural network architec-ture, where abundant topical contexts are embedded via topic relevant monolingual data. By associating each translation rule with the topic representation, topic rele-vant rules are selected according to the dis-tributional similarity with the source text during SMT decoding. Experimental re-sults show that our method significantly improves translation accuracy in the NIST Chinese-to-English translation task com-pared to a state-of-the-art baseline. 1 Lei Cui 0001, Dongdong Zhang 0001, Shujie Liu 0001, Mu Li 0001, Ming Zhou 0001, Muyun Yang |
ACL (1) | 1 |
| 2013 | Multi-Domain Adaptation for SMT Using Multi-Task LearningabstractDomain adaptation for SMT usually adapts models to an individual specific domain.However, it often lacks some correlation among different domains where common knowledge could be shared to improve the overall translation quality.In this paper, we propose a novel multi-domain adaptation approach for SMT using Multi-Task Learning (MTL), with in-domain models tailored for each specific domain and a general-domain model shared by different domains.The parameters of these models are tuned jointly via MTL so that they can learn general knowledge more accurately and exploit domain knowledge better.Our experiments on a largescale English-to-Chinese translation task validate that the MTL-based adaptation approach significantly and consistently improves the translation quality compared to a non-adapted baseline.Furthermore, it also outperforms the individual adaptation of each specific domain. Lei Cui 0001, Xilun Chen 0002, Dongdong Zhang 0001, Shujie Liu 0001, Mu Li 0001, Ming Zhou 0001 |
EMNLP | 1 |
| 2013 | Collective Corpus Weighting and Phrase Scoring for SMT Using Graph-Based Random Walk
Lei Cui 0001, Dongdong Zhang 0001, Shujie Liu 0001, Mu Li 0001, Ming Zhou 0001 |
NLPCC | 1 |
| 2011 | Function Word Generation in Statistical Machine Translation Systems
Lei Cui 0001, Dongdong Zhang 0001, Mu Li 0001, Ming Zhou 0001 |
MTSummit | 1 |
| 2011 | Improving Phrase Extraction via MBR Phrase Scoring and Pruning
Nan Duan 0001, Mu Li 0001, Ming Zhou 0001, Lei Cui 0001 |
MTSummit | 4 |