VLDB 2026 Research / reviewers in the wild / expert
Lei Li 0005
dblp:13/7007-5
· DBLP profile ↗
151ranked-venue papers
10as first author
88since 2021 · last 2026
0000-0003-3095-9776ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 137 · 7 first-author · 86 since 2021Graphics, computer vision, multimedia, augmented reality and games · 30 · 16 since 2021Databases, data management, data science and information retrieval · 25 · 5 first-author · 5 since 2021Systems, architecture and hardware · 5 · 4 since 2021Security and privacy · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Hierarchical Policy Optimization for Simultaneous Translation of Unbounded SpeechabstractSiqi Ouyang, Shuoyang Ding, Oleksii Hrinchuk, Vitaly Lavrukhin, Brian Yan, Boris Ginsburg, Lei Li. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Siqi Ouyang, Shuoyang Ding, Oleksii Hrinchuk, Vitaly Lavrukhin, Brian Yan, Boris Ginsburg, Lei Li 0005 |
ACL (1) | 7 |
| 2026 | CodeEvo: Interaction-Driven Synthesis of Code-centric Data through Hybrid and Iterative FeedbackabstractAcquiring high-quality instruction-code pairs is essential for training Large Language Models for code generation.While automated synthesis has emerged as an alternative to expensive manual curation, current approaches often rely on rigid heuristics, yielding data that is ungrounded or lacks logical complexity.We propose CodeEvo, a dual-agent architecture comprising a Coder for iterative solution synthesis and a Reviewer to orchestrate the generation trajectory.To transcend the limitations of existing heuristics, the Reviewer formulates a Schema to systematically architect logic and complexity through an interleaved synthesis of instructions and code.This process is further reinforced by a hybrid verification protocol synergizing deterministic compiler feedback with semantic evaluation.Under this framework, we construct CodeEvo-100K, a large-scale dataset of instruction-code pairs with stepped difficulty levels.Extensive experiments demonstrate that models fine-tuned on CodeEvo data significantly outperform established baselines across code generation benchmarks.In-depth analyses further provide insights into effective code-centric data synthesis.Code and data are available at https://github.com/QiushiSun/CodeEvo. Qiushi Sun, Jingyang Gong, Lei Li 0005, Qipeng Guo, Fei Yuan 0006 |
ACL (1) | 3 |
| 2025 | Efficiently Identifying Watermarked Segments in Mixed-Source TextsabstractText watermarks in large language models (LLMs) are increasingly used to detect synthetic text, mitigating misuse cases like fake news and academic dishonesty. While existing watermarking detection techniques primarily focus on classifying entire documents as watermarked or not, they often neglect the common scenario of identifying individual watermark segments within longer, mixed-source documents. Drawing inspiration from plagiarism detection systems, we propose two novel methods for partial watermark detection. First, we develop a geometry cover detection framework aimed at determining whether there is a watermark segment in long text. Second, we introduce an adaptive online learning algorithm to pinpoint the precise location of watermark segments within the text. Evaluated on three popular watermarking techniques (KGW-Watermark, Unigram-Watermark, and Gumbel-Watermark), our approach achieves high accuracy, significantly outperforming baseline methods. Moreover, our framework is adaptable to other watermarking techniques, offering new insights for precise watermark detection. Our code is publicly available at https://github.com/XuandongZhao/llm-watermark-location. Xuandong Zhao, Chenwen Liao, Yu-Xiang Wang 0003, Lei Li 0005 |
ACL (1) | 4 |
| 2025 | Extrapolating to Unknown Opinions Using LLMsabstractFrom ice cream flavors to climate change, people exhibit a wide array of opinions on various topics, and understanding the rationale for these opinions can promote healthy discussion and consensus among them. As such, it can be valuable for a large language model (LLM), particularly as an AI assistant, to be able to empathize with or even explain these various standpoints. In this work, we hypothesize that different topic stances often manifest correlations that can be used to extrapolate to topics with unknown opinions. We explore various prompting and fine-tuning methods to improve an LLM’s ability to (a) extrapolate from opinions on known topics to unknown ones and (b) support their extrapolation with reasoning. Our findings suggest that LLMs possess inherent knowledge from training data about these opinion correlations, and with minimal data, the similarities between human opinions and model-extrapolated opinions can be improved by more than 50%. Furthermore, LLM can generate the reasoning process behind their extrapolation of opinions. Kexun Zhang, Jane Dwivedi-Yu, Zhaojiang Lin, Yuning Mao, William Yang Wang, Lei Li 0005, Yi-Chia Wang |
COLING | 6 |
| 2025 | TypedThinker: Diversify Large Language Model Reasoning with Typed ThinkingabstractLarge Language Models (LLMs) have demonstrated strong reasoning capabilities in solving complex problems. However, current approaches primarily enhance reasoning through the elaboration of thoughts while neglecting the diversity of reasoning types. LLMs typically employ deductive reasoning, proceeding step-by-step from given conditions, which limits their exploration during problem-solving. Our analysis reveals that certain problems are exclusively solvable through specific reasoning strategies like inductive, abductive, or analogical reasoning. However, incorporating diverse reasoning approaches presents two key challenges: identifying the appropriate reasoning type for each problem and exploiting this approach during problem-solving. Therefore, we propose the TypedThinker that predicts suitable reasoning types based on the problem and their previous effectiveness and provides relevant demonstrations to guide LLMs in applying these strategies. Experimental results show significant improvements across multiple benchmarks, with performance gains of 3.4\% for Mistral 7B, 6.5\% for LLaMA3 8B, and 7\% for Qwen 2 7B on logical and mathematical reasoning tasks. TypedThinker enhances LLM reasoning without requiring knowledge distillation from larger models. It can be integrated into more advanced systems like GPT-4o or specialized models like MetaMath to diversify their reasoning approaches and improve their problem-solving capabilities. Danqing Wang, Fei Fang 0001, Lei Li 0005 |
ICLR | 4 |
| 2025 | Speculative Knowledge Distillation: Bridging the Teacher-Student Gap Through Interleaved SamplingabstractRecent advances in knowledge distillation (KD) have enabled smaller student models to approach the performance of larger teacher models. However, popular methods such as supervised KD and on-policy KD, are adversely impacted by the knowledge gaps between teacher-student in practical scenarios. Supervised KD suffers from a distribution mismatch between training with a static dataset and inference over final student-generated outputs. Conversely, on-policy KD, which uses student-generated samples for training, can suffer from low-quality training examples with which teacher models are not familiar, resulting in inaccurate teacher feedback. To address these limitations, we introduce Speculative Knowledge Distillation (SKD), a novel approach that leverages cooperation between student and teacher models to generate high-quality training data on-the-fly while aligning with the student's inference-time distribution. In SKD, the student proposes tokens, and the teacher replaces poorly ranked ones based on its own distribution, transferring high-quality knowledge adaptively. We evaluate SKD on various text generation tasks, including translation, summarization, math, and instruction following, and show that SKD consistently outperforms existing KD methods across different domains, data sizes, and model initialization strategies. Wenda Xu, Rujun Han, Zifeng Wang 0002, Long T. Le, Dhruv Madeka, Lei Li 0005, William Yang Wang, Rishabh Agarwal, Chen-Yu Lee, Tomas Pfister |
ICLR | 6 |
| 2025 | Diversity Empowers Intelligence: Integrating Expertise of Software Engineering AgentsabstractLarge language model (LLM) agents have shown great potential in solving real-world software engineering (SWE) problems. The most advanced open-source SWE agent can resolve over 27% of real GitHub issues in SWE-Bench Lite. However, these sophisticated agent frameworks exhibit varying strengths, excelling in certain tasks while underperforming in others. To fully harness the diversity of these agents, we propose DEI (Diversity Empowered Intelligence), a framework that leverages their unique expertise. DEI functions as a meta-module atop existing SWE agent frameworks, managing agent collectives for enhanced problem-solving. Experimental results show that a DEI-guided committee of agents is able to surpass the best individual agent's performance by a large margin. For instance, a group of open-source SWE agents, with a maximum individual resolve rate of 27.3% on SWE-Bench Lite, can achieve a 34.3% resolve rate with DEI, making a 25% improvement and beating most closed-source solutions. Our best-performing group excels with a 55% resolve rate, securing the highest ranking on SWE-Bench Lite. Our findings contribute to the growing body of research on collaborative AI systems and their potential to solve complex software engineering challenges. Kexun Zhang, Weiran Yao, Zuxin Liu, Yihao Feng, Zhiwei Liu 0001, Rithesh R. N., Tian Lan 0006, Lei Li 0005, Renze Lou, Jiacheng Xu 0001, Bo Pang 0004, Yingbo Zhou 0002, Shelby Heinecke, Silvio Savarese, Huan Wang 0016, Caiming Xiong |
ICLR | 8 |
| 2025 | Permute-and-Flip: An optimally stable and watermarkable decoder for LLMsabstractIn this paper, we propose a new decoding method called Permute-and-Flip (PF) decoder. It enjoys stability properties similar to the standard sampling decoder, but is provably up to 2x better in its quality-stability tradeoff than sampling and never worse than any other decoder. We also design a cryptographic watermarking scheme analogous to Aaronson (2023)'s Gumbel watermark, but naturally tailored for PF decoder. The watermarking scheme does not change the distribution to sample, while allowing arbitrarily low false positive rate and high recall whenever the generated text has high entropy. Our experiments show that the PF decoder (and its watermarked counterpart) significantly outperform(s) naive sampling (and its Gumbel watermarked counterpart) in terms of perplexity, while retaining the same stability (and detectability), hence making it a promising new approach for LLM decoding. The code is available at https://github.com/XuandongZhao/pf-decoding Xuandong Zhao, Lei Li 0005, Yu-Xiang Wang 0003 |
ICLR | 2 |
| 2025 | DIS-CO: Discovering Copyrighted Content in VLMs Training Dataabstract*How can we verify whether copyrighted content was used to train a large vision-language model (VLM) without direct access to its training data?* Motivated by the hypothesis that a VLM is able to recognize images from its training corpus, we propose DIS-CO, a novel approach to infer the inclusion of copyrighted content during the model's development. By repeatedly querying a VLM with specific frames from targeted copyrighted material, DIS-CO extracts the content's identity through free-form text completions. To assess its effectiveness, we introduce MovieTection, a benchmark comprising 14,000 frames paired with detailed captions, drawn from films released both before and after a model’s training cutoff. Our results show that DIS-CO significantly improves detection performance, nearly doubling the average AUC of the best prior method on models with logits available. Our findings also highlight a broader concern: all tested models appear to have been exposed to some extent to copyrighted content. We provide the code in the supplementary materials. André V. Duarte, Xuandong Zhao, Arlindo L. Oliveira, Lei Li 0005 |
ICML | 4 |
| 2025 | PPDiff: Diffusing in Hybrid Sequence-Structure Space for Protein-Protein Complex DesignabstractDesigning protein-binding proteins with high affinity is critical in biomedical research and biotechnology. Despite recent advancements targeting specific proteins, the ability to create high-affinity binders for arbitrary protein targets on demand, without extensive rounds of wet-lab testing, remains a significant challenge. Here, we introduce PPDiff, a diffusion model to jointly design the sequence and structure of binders for arbitrary protein targets in a non-autoregressive manner. PPDiff builds upon our developed Sequence Structure Interleaving Network with Causal attention layers (SSINC), which integrates interleaved self-attention layers to capture global amino acid correlations, $k$-nearest neighbor ($k$NN) equivariant graph convolutional layers to model local interactions in three-dimensional (3D) space, and causal attention layers to simplify the intricate interdependencies within the protein sequence. To assess PPDiff, we curate PPBench, a general protein-protein complex dataset comprising 706,360 complexes from the Protein Data Bank (PDB). The model is pretrained on PPBench and finetuned on two real-world applications: target-protein mini-binder complex design and antigen-antibody complex design. PPDiff consistently surpasses baseline methods, achieving success rates of 50.00\%, 23.16\%, and 16.89\% for the pretraining task and the two downstream applications, respectively. Zhenqiao Song, Tianxiao Li 0001, Lei Li 0005, Martin Renqiang Min |
ICML | 3 |
| 2025 | Weak-to-Strong Jailbreaking on Large Language ModelsabstractLarge language models (LLMs) are vulnerable to jailbreak attacks -- resulting in harmful, unethical, or biased text generations. However, existing jailbreaking methods are computationally costly. In this paper, we propose the **weak-to-strong** jailbreaking attack, an efficient inference time attack for aligned LLMs to produce harmful text. Our key intuition is based on the observation that jailbroken and aligned models only differ in their initial decoding distributions. The weak-to-strong attack's key technical insight is using two smaller models (a safe and an unsafe one) to adversarially modify a significantly larger safe model's decoding probabilities. We evaluate the weak-to-strong attack on 5 diverse open-source LLMs from 3 organizations. The results show our method can increase the misalignment rate to over 99\% on two datasets with just one forward pass per example. Our study exposes an urgent safety issue that needs to be addressed when aligning LLMs. As an initial attempt, we propose a defense strategy to protect against such attacks, but creating more advanced defenses remains challenging. The code for replicating the method is available at https://github.com/XuandongZhao/weak-to-strong. Xuandong Zhao, Xianjun Yang, Tianyu Pang, Lei Li 0005, Yu-Xiang Wang 0003, William Yang Wang |
ICML | 5 |
| 2025 | Anticipating Future with Large Language Model for Simultaneous Machine TranslationabstractSiqi Ouyang, Oleksii Hrinchuk, Zhehuai Chen, Vitaly Lavrukhin, Jagadeesh Balam, Lei Li, Boris Ginsburg. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Siqi Ouyang, Oleksii Hrinchuk, Zhehuai Chen, Vitaly Lavrukhin, Jagadeesh Balam, Lei Li 0005, Boris Ginsburg |
NAACL (Long Papers) | 6 |
| 2025 | Revealing the Barriers of Language Agents in PlanningabstractJian Xie, Kexun Zhang, Jiangjie Chen, Siyu Yuan, Kai Zhang, Yikai Zhang, Lei Li, Yanghua Xiao. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Kexun Zhang, Jiangjie Chen, Kai Zhang 0033, Yikai Zhang 0004, Lei Li 0005, Yanghua Xiao |
NAACL (Long Papers) | 7 |
| 2025 | KS-Lottery: Finding Certified Lottery Tickets for Multilingual Transfer in Large Language ModelsabstractFei Yuan, Chang Ma, Shuai Yuan, Qiushi Sun, Lei Li. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Fei Yuan 0006, Shuai Yuan 0018, Qiushi Sun, Lei Li 0005 |
NAACL (Long Papers) | 5 |
| 2025 | Scaling LLM Inference Efficiently with Optimized Sample Compute AllocationabstractKexun Zhang, Shang Zhou, Danqing Wang, William Yang Wang, Lei Li. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Kexun Zhang, Shang Zhou, Danqing Wang, William Yang Wang, Lei Li 0005 |
NAACL (Long Papers) | 5 |
| 2025 | A Technical Report on "Erasing the Invisible": The 2024 NeurIPS Competition on Stress Testing Image WatermarksabstractAI-generated images have become pervasive, raising critical concerns around content authenticity, intellectual property, and the spread of misinformation. Invisible watermarks offer a promising solution for identifying AI-generated images, preserving content provenance without degrading visual quality. However, their real-world robustness remains uncertain due to the lack of standardized evaluation protocols and large-scale stress testing. To bridge this gap, we organized “Erasing the Invisible,” a NeurIPS 2024 competition and newly established benchmark designed to systematically stress testing the resilience of watermarking techniques. The competition introduced two attack tracks—Black-box and Beige-box—that simulate practical scenarios with varying levels of attacker knowledge on watermarks, providing a comprehensive assessment of watermark robustness. The competition attracted significant global participation, with 2,722 submissions from 298 teams. Through a rigorous evaluation pipeline featuring real-time feedback and human-verified final rankings, participants developed and demonstrated new attack strategies that revealed critical vulnerabilities in state-of-the-art watermarking methods. On average, the top-5 teams in both tracks could remove watermarks from $\geq$ 89% of the images while preserving high visual quality, setting strong baselines for future research on watermark attacks and defenses. To support continued progress in this field, we summarize the insights and lessons learned from this competition in this paper, and release the benchmark dataset, evaluation toolkit, and competition results. “Erasing the Invisible” establishes a valuable open resource for advancing more robust watermarking techniques and strengthening content provenance in the era of generative AI. Mucong Ding, Bang An 0001, Tahseen Rabbani, Chenghao Deng, Anirudh Satheesh, Souradip Chakraborty, Mehrdad Saberi, Yuxin Wen, Kyle Sang, Aakriti Agrawal, Xuandong Zhao, Mary-Anne Hartley, Lei Li 0005, Yu-Xiang Wang 0003, Vishal M. Patel, Soheil Feizi, Tom Goldstein, Furong Huang |
NeurIPS | 14 |
| 2025 | SoK: Watermarking for AI-Generated ContentabstractAs the outputs of generative AI (GenAl) techniques improve in quality, it becomes increasingly challenging to distinguish them from human-created content. Watermarking schemes are a promising approach to address the problem of distinguishing between AI and human-generated content. These schemes embed hidden signals within AI -generated content to enable reliable detection. While watermarking is not a silver bullet for addressing all risks associated with GenAl, it can play a crucial role in enhancing AI safety and trustworthiness by combating misinformation and deception. This paper presents a comprehensive overview of water-marking techniques for GenAl, beginning with the need for watermarking from historical and regulatory perspectives. We formalize the definitions and desired properties of watermarking schemes and examine the key objectives and threat models for existing approaches. Practical evaluation strategies are also explored, providing insights into the development of robust watermarking techniques capable of resisting various attacks. Additionally, we review recent representative works, highlight open challenges, and discuss potential directions for this emerging field. By offering a thorough understanding of watermarking in GenAl, this work aims to guide researchers in advancing watermarking methods and applications, and support policymakers in addressing the broader implications of GenAl. Xuandong Zhao, Sam Gunn, Miranda Christ, Jaiden Fairoze, Andrés Fábrega, Nicholas Carlini, Sanjam Garg, Sanghyun Hong 0001, Milad Nasr, Florian Tramèr, Somesh Jha, Lei Li 0005, Yu-Xiang Wang 0003, Dawn Song |
SP | 12 |
| 2024 | Where It Really Matters: Few-Shot Environmental Conservation Media Monitoring for Low-Resource LanguagesabstractEnvironmental conservation organizations routinely monitor news content on conservation in protected areas to maintain situational awareness of developments that can have an environmental impact. Existing automated media monitoring systems require large amounts of data labeled by domain experts, which is only feasible at scale for high-resource languages like English. However, such tools are most needed in the global south where the news of interest is mainly in local low-resource languages, and far fewer experts are available to annotate datasets on a sustainable basis. In this paper, we propose NewsSerow, a method to automatically recognize environmental conservation content in low-resource languages. NewsSerow is a pipeline of summarization, in-context few-shot classification, and self-reflection using large language models (LLMs). Using at most 10 demonstration example news articles in Nepali, NewsSerow significantly outperforms other few-shot methods and can achieve comparable performance with models fully fine-tuned using thousands of examples. With NewsSerow, Organization X has been able to deploy the media monitoring tool in Nepal, significantly reducing their operational burden, and ensuring that AI tools for conservation actually reach the communities that need them the most. NewsSerow has also been deployed for countries with other languages like Colombia. Sedrick Keh, Shova Chettri, Karun Dewan, Pablo Izquierdo, Johanna Prussman, Pooja Shrestha, César Suárez, Zheyuan Shi, Lei Li 0005, Fei Fang 0001 |
AAAI | 10 |
| 2024 | Pride and Prejudice: LLM Amplifies Self-Bias in Self-RefinementabstractRecent studies show that large language models (LLMs) improve their performance through self-feedback on certain tasks while degrade on others.We discovered that such a contrary is due to LLM's bias in evaluating their own output.In this paper, we formally define LLM's self-bias -the tendency to favor its own generation -using two statistics.We analyze six LLMs (GPT-4, GPT-3.5, Gemini, LLaMA2, Mixtral and DeepSeek) on translation, constrained text generation, and mathematical reasoning tasks.We find that self-bias is prevalent in all examined LLMs across multiple languages and tasks.Our analysis reveals that while the self-refine pipeline improves the fluency and understandability of model outputs, it further amplifies self-bias.To mitigate such biases, we discover that larger model size and external feedback with accurate assessment can significantly reduce bias in the self-refine pipeline, leading to actual performance improvement in downstream tasks.The code and data are released at https://github. com/xu1998hz/llm_self_bias. Wenda Xu, Guanglei Zhu, Xuandong Zhao, Liangming Pan, Lei Li 0005, William Yang Wang |
ACL (1) | 5 |
| 2024 | A Survey on In-context LearningabstractQingxiu Dong, Lei Li, Damai Dai, Ce Zheng, Jingyuan Ma, Rui Li, Heming Xia, Jingjing Xu, Zhiyong Wu, Baobao Chang, Xu Sun, Lei Li, Zhifang Sui. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024. Qingxiu Dong, Lei Li 0039, Damai Dai, Jingyuan Ma, Rui Li 0094, Heming Xia, Jingjing Xu 0001, Zhiyong Wu 0011, Baobao Chang, Xu Sun 0001, Lei Li 0005, Zhifang Sui |
EMNLP | 12 |
| 2024 | Learning Personalized Alignment for Evaluating Open-ended Text GenerationabstractRecent research has increasingly focused on evaluating large language models' (LLMs) alignment with diverse human values and preferences, particularly for open-ended tasks like story generation.Traditional evaluation metrics rely heavily on lexical similarity with humanwritten references, often showing poor correlation with human judgments and failing to account for alignment with the diversity of human preferences.To address these challenges, we introduce PERSE, an interpretable evaluation framework designed to assess alignment with specific human preferences.It is tuned to infer specific preferences from an in-context personal profile and evaluate the alignment between the generated content and personal preferences.PERSE enhances interpretability by providing detailed comments and fine-grained scoring, facilitating more personalized content generation.Our 13B LLaMA-2-based PERSE shows a 15.8% increase in Kendall correlation and a 13.7% rise in accuracy with zero-shot reviewers compared to GPT-4.It also outperforms GPT-4 by 46.01% in Kendall correlation on new domains, indicating its transferability 1 . Danqing Wang, Kevin Yang, Hanlin Zhu, Andrew Cohen, Lei Li 0005, Yuandong Tian |
EMNLP | 6 |
| 2024 | BPO: Staying Close to the Behavior LLM Creates Better Online LLM AlignmentabstractDirect alignment from preferences (DAP) has emerged as a promising paradigm for aligning large language models (LLMs) to human desiderata from pre-collected, offline preference datasets. While recent studies indicate that existing offline DAP methods can directly benefit from online training samples, we highlight the need to develop specific online DAP algorithms to fully harness the power of online training. Specifically, we identify that the learned LLM should adhere to the proximity of the behavior LLM, which collects the training samples. To this end, we propose online Preference Optimization in proximity to the Behavior LLM (BPO), emphasizing the importance of constructing a proper trust region for LLM alignment.We conduct extensive experiments to validate the effectiveness and applicability of our approach by integrating it with various DAP methods, resulting in significant performance improvements across a wide range of tasks when training with the same amount of preference data. Even when only introducing one additional data collection phase, our online BPO improves its offline DAP baseline from 72.0% to 80.2% on TL;DR and from 82.2% to 89.1% on Anthropic Helpfulness in terms of win rate against human reference text. Wenda Xu, William Yang Wang, Lei Li 0005 |
EMNLP | 4 |
| 2024 | Provable Robust Watermarking for AI-Generated TextabstractWe study the problem of watermarking large language models (LLMs) generated text — one of the most promising approaches for addressing the safety challenges of LLM usage. In this paper, we propose a rigorous theoretical framework to quantify the effectiveness and robustness of LLM watermarks. We propose a robust and high-quality watermark method, Unigram-Watermark, by extending an existing approach with a simplified fixed grouping strategy. We prove that our watermark method enjoys guaranteed generation quality, correctness in watermark detection, and is robust against text editing and paraphrasing. Experiments on three varying LLMs and two datasets verify that our Unigram-Watermark achieves superior detection accuracy and comparable generation quality in perplexity, thus promoting the responsible use of LLMs. Xuandong Zhao, Prabhanjan Vijendra Ananth, Lei Li 0005, Yu-Xiang Wang 0003 |
ICLR | 3 |
| 2024 | DE-COP: Detecting Copyrighted Content in Language Models Training Dataabstract*How can we detect if copyrighted content was used in the training process of a language model, considering that the training data is typically undisclosed?* We are motivated by the premise that a language model is likely to identify verbatim excerpts from its training text. We propose DE-COP, a method to determine whether a piece of copyrighted content is included in training. DE-COP's core approach is to probe an LLM with multiple-choice questions, whose options include both verbatim text and their paraphrases. We construct BookTection, a benchmark with excerpts from 165 books published prior and subsequent to a model's training cutoff, along with their paraphrases. Our experiments show that DE-COP outperforms the prior best method by 8.6% in detection accuracy (AUC) on models with logits available. Moreover, DE-COP also achieves an average accuracy of 72% for detecting suspect books on fully black-box models where prior methods give approximately 0% accuracy. The code and datasets are available at https://github.com/LeiLiLab/DE-COP. André V. Duarte, Xuandong Zhao, Arlindo L. Oliveira, Lei Li 0005 |
ICML | 4 |
| 2024 | SurfPro: Functional Protein Design Based on Continuous SurfaceabstractHow can we design proteins with desired functions? We are motivated by a chemical intuition that both geometric structure and biochemical properties are critical to a protein’s function. In this paper, we propose SurfPro, a new method to generate functional proteins given a desired surface and its associated biochemical properties. SurfPro comprises a hierarchical encoder that progressively models the geometric shape and biochemical features of a protein surface, and an autoregressive decoder to produce an amino acid sequence. We evaluate SurfPro on a standard inverse folding benchmark CATH 4.2 and two functional protein design tasks: protein binder design and enzyme design. Our SurfPro consistently surpasses previous state-of-the-art inverse folding methods, achieving a recovery rate of 57.78% on CATH 4.2 and higher success rates in terms of protein-protein binding and enzyme-substrate interaction scores Zhenqiao Song, Lei Li 0005, Wengong Jin |
ICML | 3 |
| 2024 | Generative Enzyme Design Guided by Functionally Important Sites and Small-Molecule SubstratesabstractEnzymes are genetically encoded biocatalysts capable of accelerating chemical reactions. How can we automatically design functional enzymes? In this paper, we propose EnzyGen, an approach to learn a unified model to design enzymes across all functional families. Our key idea is to generate an enzyme's amino acid sequence and their three-dimensional (3D) coordinates based on functionally important sites and substrates corresponding to a desired catalytic function. These sites are automatically mined from enzyme databases. EnzyGen consists of a novel interleaving network of attention and neighborhood equivariant layers, which captures both long-range correlation in an entire protein sequence and local influence from nearest amino acids in 3D space. To learn the generative model, we devise a joint training objective, including a sequence generation loss, a position prediction loss and an enzyme-substrate interaction loss. We further construct EnzyBench, a dataset with 3157 enzyme families, covering all available enzymes within the protein data bank (PDB). Experimental results show that our EnzyGen consistently achieves the best performance across all 323 testing families, surpassing the best baseline by 10.79% in terms of substrate binding affinity. These findings demonstrate EnzyGen's superior capability in designing well-folded and effective enzymes binding to specific substrates with high affinities. Our code, model and dataset are provided at https://github.com/LeiLiLab/EnzyGen. Zhenqiao Song, Yunlong Zhao 0002, Wenxian Shi, Wengong Jin, Yang Yang 0059, Lei Li 0005 |
ICML | 6 |
| 2024 | Global Human-guided Counterfactual Explanations for Molecular Properties via Reinforcement LearningabstractCounterfactual explanations of Graph Neural Networks (GNNs) offer a powerful way to understand data that can naturally be represented by a graph structure. Furthermore, in many domains, it is highly desirable to derive data-driven global explanations or rules that can better explain the high-level properties of the models and data in question. However, evaluating global counterfactual explanations is hard in real-world datasets due to a lack of human-annotated ground truth, which limits their use in areas like molecular sciences. Additionally, the increasing scale of these datasets provides a challenge for random search-based methods. In this paper, we develop a novel global explanation model RLHEX for molecular property prediction. It aligns the counterfactual explanations with human-defined principles, making the explanations more interpretable and easy for experts to evaluate. RLHEX includes a VAE-based graph generator to generate global explanations and an adapter to adjust the latent representation space to human-defined principles. Optimized by Proximal Policy Optimization (PPO), the global explanations produced by RLHEX cover 4.12% more input graphs and reduce the distance between the counterfactual explanation set and the input set by 0.47% on average across three molecular datasets. RLHEX provides a flexible framework to incorporate different human-designed principles into the counterfactual explanation generation process, aligning these explanations with domain expertise. The code and data are released at https://github.com/dqwang122/RLHEX. Danqing Wang, Antonis Antoniades, Kha-Dinh Luong, Edwin Zhang, Mert Kosan, Ambuj K. Singh, William Yang Wang, Lei Li 0005 |
KDD | 9 |
| 2024 | MindMerger: Efficiently Boosting LLM Reasoning in non-English LanguagesabstractReasoning capabilities are crucial for Large Language Models~(LLMs), yet a notable gap exists between English and non-English languages. To bridge this disparity, some works fine-tune LLMs to relearn reasoning capabilities in non-English languages, while others replace non-English inputs with an external model's outputs such as English translation text to circumvent the challenge of LLM understanding non-English. Unfortunately, these methods often underutilize the built-in skilled reasoning and useful language understanding capabilities of LLMs. In order to better utilize the minds of reasoning and language understanding in LLMs, we propose a new method, namely MergeMinds, which merges LLMs with the external language understanding capabilities from multilingual models to boost the multilingual reasoning performance. Furthermore, a two-step training scheme is introduced to first train to embeded the external capabilities into LLMs and then train the collaborative utilization of the external capabilities and the built-in capabilities in LLMs. Experiments on three multilingual reasoning datasets and a language understanding dataset demonstrate that MergeMinds consistently outperforms all baselines, especially in low-resource languages. Without updating the parameters of LLMs, the average accuracy improved by 6.7 and 8.0 across all languages and low-resource languages on the MGSM dataset, respectively. Zixian Huang, Gong Cheng 0001, Lei Li 0005, Fei Yuan 0006 |
NeurIPS | 4 |
| 2024 | Invisible Image Watermarks Are Provably Removable Using Generative AIabstractInvisible watermarks safeguard images' copyrights by embedding hidden messages only detectable by owners. They also prevent people from misusing images, especially those generated by AI models.
We propose a family of regeneration attacks to remove these invisible watermarks.
The proposed attack method first adds random noise to an image to destroy the watermark and then reconstructs the image.
This approach is flexible and can be instantiated with many existing image-denoising algorithms and pre-trained generative models such as diffusion models. Through formal proofs and extensive empirical evaluations, we demonstrate that pixel-level invisible watermarks are vulnerable to this regeneration attack.
Our results reveal that, across four different pixel-level watermarking schemes, the proposed method consistently achieves superior performance compared to existing attack techniques, with lower detection rates and higher image quality.
However, watermarks that keep the image semantically similar can be an alternative defense against our attacks.
Our finding underscores the need for a shift in research/industry emphasis from invisible watermarks to semantic-preserving watermarks. Code is available at https://github.com/XuandongZhao/WatermarkAttacker Xuandong Zhao, Kexun Zhang, Zihao Su, Saastha Vasan, Ilya Grishchenko, Christopher Krügel, Giovanni Vigna, Yu-Xiang Wang 0003, Lei Li 0005 |
NeurIPS | 9 |
| 2023 | Converge to the Truth: Factual Error Correction via Iterative Constrained EditingabstractGiven a possibly false claim sentence, how can we automatically correct it with minimal editing? Existing methods either require a large number of pairs of false and corrected claims for supervised training or do not handle well errors spanning over multiple tokens within an utterance. In this paper, we propose VENCE, a novel method for factual error correction (FEC) with minimal edits. VENCE formulates the FEC problem as iterative sampling editing actions with respect to a target density function. We carefully design the target function with predicted truthfulness scores from an offline trained fact verification model. VENCE samples the most probable editing positions based on back-calculated gradients of the truthfulness score concerning input tokens and the editing actions using a distantly-supervised language model (T5). Experiments on a public dataset show that VENCE improves the well-adopted SARI metric by 5.3 (or a relative improvement of 11.8%) over the previous best distantly-supervised methods. Jiangjie Chen, Rui Xu 0026, Wenxuan Zeng, Changzhi Sun, Lei Li 0005, Yanghua Xiao |
AAAI | 5 |
| 2023 | Say What You Mean! Large Language Models Speak Too Positively about Negative Commonsense KnowledgeabstractLarge language models (LLMs) have been widely studied for their ability to store and utilize positive knowledge.However, negative knowledge, such as "lions don't live in the ocean", is also ubiquitous in the world but rarely mentioned explicitly in the text.What do LLMs know about negative knowledge?This work examines the ability of LLMs to negative commonsense knowledge.We design a constrained keywords-to-sentence generation task (CG) and a Boolean question-answering task (QA) to probe LLMs.Our experiments reveal that LLMs frequently fail to generate valid sentences grounded in negative commonsense knowledge, yet they can correctly answer polar yes-or-no questions.We term this phenomenon the belief conflict of LLMs.Our further analysis shows that statistical shortcuts and negation reporting bias from language modeling pre-training cause this conflict.1 * Work done while at Brain Technologies, Inc. Jiangjie Chen, Ziquan Fu, Sijie Cheng, Lei Li 0005, Yanghua Xiao |
ACL (1) | 5 |
| 2023 | WACO: Word-Aligned Contrastive Learning for Speech TranslationabstractEnd-to-end Speech Translation (E2E ST) aims to directly translate source speech into target text.Existing ST methods perform poorly when only extremely small speech-text data are available for training.We observe that an ST model's performance closely correlates with its embedding similarity between speech and source transcript.In this paper, we propose Word-Aligned COntrastive learning (WACO), a simple and effective method for extremely low-resource speech-to-text translation.Our key idea is bridging word-level representations for both speech and text modalities via contrastive learning.We evaluate WACO and other methods on the MuST-C dataset, a widely used ST benchmark, and on a low-resource direction Maltese-English from IWSLT 2023.Our experiments demonstrate that WACO outperforms the best baseline by 9+ BLEU points with only 1-hour parallel ST data. Siqi Ouyang, Rong Ye, Lei Li 0005 |
ACL (1) | 3 |
| 2023 | SESCORE2: Learning Text Generation Evaluation via Synthesizing Realistic MistakesabstractIs it possible to train a general metric for evaluating text generation quality without humanannotated ratings?Existing learned metrics either perform unsatisfactorily across text generation tasks or require human ratings for training on specific tasks.In this paper, we propose SESCORE2, a self-supervised approach for training a model-based metric for text generation evaluation.The key concept is to synthesize realistic model mistakes by perturbing sentences retrieved from a corpus.The primary advantage of the SESCORE2 is its ease of extension to many other languages while providing reliable severity estimation.We evaluate SESCORE2 and previous methods on four text generation tasks across three languages.SESCORE2 outperforms unsupervised metric PRISM on four text generation evaluation benchmarks, with a Kendall improvement of 0.078.Surprisingly, SESCORE2 even outperforms the supervised BLEURT and COMET on multiple text generation tasks.The code and data are available at https://github.com/ xu1998hz/SEScore2 1 . Wenda Xu, Xian Qian, Mingxuan Wang, Lei Li 0005, William Yang Wang |
ACL (1) | 4 |
| 2023 | Pre-trained Language Models Can be Fully Zero-Shot LearnersabstractHow can we extend a pre-trained model to many language understanding tasks, without labeled or additional unlabeled data?Pre-trained language models (PLMs) have been effective for a wide range of NLP tasks.However, existing approaches either require fine-tuning on downstream labeled datasets or manually constructing proper prompts.In this paper, we propose nonparametric prompting PLM (NPPrompt) for fully zero-shot language understanding.Unlike previous methods, NPPrompt uses only pre-trained language models and does not require any labeled data or additional raw corpus for further fine-tuning, nor does it rely on humans to construct a comprehensive set of prompt label words.We evaluate NPPrompt against previous major fewshot and zero-shot learning methods on diverse NLP tasks: text classification, text entailment, similar text retrieval, paraphrasing, and multiple-choice question answering.Experimental results demonstrate that our NPPrompt outperforms the previous best fully zero-shot method by big margins, with absolute gains of 12.8% in accuracy on text classification and 15.6% on the GLUE benchmark. Xuandong Zhao, Siqi Ouyang, Lei Li 0005 |
ACL (1) | 5 |
| 2023 | Learning from Mistakes via Cooperative Study Assistant for Large Language ModelsabstractLarge language models (LLMs) have demonstrated their potential to refine their generation based on their own feedback.However, the feedback from LLM itself is often inaccurate, thereby limiting its benefits.In this paper, we propose Study Assistant for Large LAnguage Model (SALAM), a novel framework with an auxiliary agent to assist the main LLM in learning from mistakes through interactive cooperation.In the gathering phase, the student assistant agent probes the main LLM, analyzes its errors, and collects the interaction in a mistake memory.During the examination phase, the study assistant provides guidelines by retrieving relevant cases to help the main LLM anticipate and avoid similar errors.We first investigate the effectiveness of a general study assistant and then customize it to provide LLMspecific guidance through imitation learning from successful guidance experiences.Our experiments on three LLMs using two challenging frameworks demonstrate that SALAM can significantly boost LLMs by an accuracy margin of up to 6.6 on BBH and 12.6 on BBQ 1 . Danqing Wang, Lei Li 0005 |
EMNLP | 2 |
| 2023 | INSTRUCTSCORE: Towards Explainable Text Generation Evaluation with Automatic FeedbackabstractAutomatically evaluating the quality of language generation is critical.Although recent learned metrics show high correlation with human judgement, these metrics do not provide explicit explanation of their verdict, nor associate the scores with defects in the generated text.To address this limitation, we present IN-STRUCTSCORE, a fine-grained explainable evaluation metric for text generation.By harnessing both explicit human instruction and the implicit knowledge of GPT-4, we fine-tune a text evaluation metric based on LLaMA, producing both a score for generated text and a human readable diagnostic report.We evaluate INSTRUCTSCORE on a variety of generation tasks, including translation, captioning, data-to-text, and commonsense generation.Experiments show that our 7B model surpasses all other unsupervised metrics, including those based on 175B GPT-3 and GPT-4.Surprisingly, our INSTRUCTSCORE, even without direct supervision from human-rated data, achieves performance levels on par with state-of-the-art metrics like COMET22, which were fine-tuned on human ratings.Prompt: You are evaluating a model output based on a reference.Reference: Normally the administration office downstairs would call me when there's a delivery.Output: Usually when there is takeaway, the management office downstairs will call. Wenda Xu, Danqing Wang, Liangming Pan, Zhenqiao Song, Markus Freitag, William Yang Wang, Lei Li 0005 |
EMNLP | 7 |
| 2023 | Importance Weighted Expectation-Maximization for Protein Sequence DesignabstractDesigning protein sequences with desired biological function is crucial in biology and chemistry. Recent machine learning methods use a surrogate sequence-function model to replace the expensive wet-lab validation. How can we efficiently generate diverse and novel protein sequences with high fitness? In this paper, we propose IsEM-Pro, an approach to generate protein sequences towards a given fitness criterion. At its core, IsEM-Pro is a latent generative model, augmented by combinatorial structure features from a separately learned Markov random fields (MRFs). We develop an Monte Carlo Expectation-Maximization method (MCEM) to learn the model. During inference, sampling from its latent space enhances diversity while its MRFs features guide the exploration in high fitness regions. Experiments on eight protein sequence design tasks show that our IsEM-Pro outperforms the previous best methods by at least 55% on average fitness score and generates more diverse and novel protein sequences. Zhenqiao Song, Lei Li 0005 |
ICML | 2 |
| 2023 | ReDi: Efficient Learning-Free Diffusion Inference via Trajectory RetrievalabstractDiffusion models show promising generation capability for a variety of data. Despite their high generation quality, the inference for diffusion models is still time-consuming due to the numerous sampling iterations required. To accelerate the inference, we propose ReDi, a simple yet learning-free Retrieval-based Diffusion sampling framework. From a precomputed knowledge base, ReDi retrieves a trajectory similar to the partially generated trajectory at an early stage of generation, skips a large portion of intermediate steps, and continues sampling from a later step in the retrieved trajectory. We theoretically prove that the generation performance of ReDi is guaranteed. Our experiments demonstrate that ReDi improves the model inference efficiency by 2$\times$ speedup. Furthermore, ReDi is able to generalize well in zero-shot cross-domain image generation such as image stylization. The code and demo for ReDi is available at https://github.com/zkx06111/ReDiffusion. Kexun Zhang, Xianjun Yang, William Yang Wang, Lei Li 0005 |
ICML | 4 |
| 2023 | Protecting Language Generation Models via Invisible WatermarkingabstractLanguage generation models have been an increasingly powerful enabler to many applications. Many such models offer free or affordable API access which makes them potentially vulnerable to model extraction attacks through distillation. To protect intellectual property (IP) and make fair use of these models, various techniques such as lexical watermarking and synonym replacement have been proposed. However, these methods can be nullified by obvious countermeasures such as ``synonym randomization''. To address this issue, we propose GINSW, a novel method to protect text generation models from being stolen through distillation. The key idea of our method is to inject secret signals into the probability vector of the decoding steps for each target token. We can then detect the secret message by probing a suspect model to tell if it is distilled from the protected one. Experimental results show that GINSW can effectively identify instances of IP infringement with minimal impact on the generation quality of protected APIs. Our method demonstrates an absolute improvement of 19 to 29 points on mean average precision (mAP) in detecting suspects compared to previous methods against watermark removal attacks. Xuandong Zhao, Yu-Xiang Wang 0003, Lei Li 0005 |
ICML | 3 |
| 2023 | Accelerating Antimicrobial Peptide Discovery with Latent StructureabstractAntimicrobial peptides (AMPs) are promising therapeutic approaches against drug-resistant pathogens. Recently, deep generative models are used to discover new AMPs. However, previous studies mainly focus on peptide sequence attributes and do not consider crucial structure information. In this paper, we propose a latent sequence-structure model for designing AMPs (LSSAMP). LSSAMP exploits multi-scale vector quantization in the latent space to represent secondary structures (e.g. alpha helix and beta sheet). By sampling in the latent space, LSSAMP can simultaneously generate peptides with ideal sequence attributes and secondary structures. Experimental results show that the peptides generated by LSSAMP have a high probability of antimicrobial activity. Our wet laboratory experiments verified that two of the 21 candidates exhibit strong antimicrobial activity. The code is released at https://github.com/dqwang122/LSSAMP. Danqing Wang, Zeyu Wen, Lei Li 0005, Hao Zhou 0012 |
KDD | 4 |
| 2023 | Statistical Knowledge Assessment for Large Language ModelsabstractGiven varying prompts regarding a factoid question, can a large language model (LLM) reliably generate factually correct answers? Existing LLMs may generate distinct responses for different prompts. In this paper, we study the problem of quantifying knowledge contained in an LLM regarding a given set of facts. We propose KaRR, a statistical approach to assess factual knowledge for LLMs. The main idea is to estimate the ratio of LLM generating text corresponding to the answer entity given diverse prompts of the subject and the querying relation, versus it generating by random chances. Our assessment suite contains a comprehensive set of 994,123 entities and 600 relations, with 1,395,905 text aliases. We use our method to evaluate 20 LLMs of various sizes, including LLaMA, Alpaca, OPT, etc. Experiments show that our results have a strong correlation (0.43 Kendall's $\tau$) with the results of human assessment on LLMs. Our results reveal that the knowledge in LLMs with the same backbone architecture adheres to the scaling law, while tuning on instruction-following data sometimes compromises the model's capability to generate factually correct text reliably. Qingxiu Dong, Jingjing Xu 0001, Lingpeng Kong, Zhifang Sui, Lei Li 0005 |
NeurIPS | 5 |
| 2023 | ALGO: Synthesizing Algorithmic Programs with Generated Oracle VerifiersabstractLarge language models (LLMs) excel at implementing code from functionality descriptions but struggle with algorithmic problems that require not only implementation but also identification of the suitable algorithm. Moreover, LLM-generated programs lack guaranteed correctness and require human verification. To address these challenges, we propose ALGO, a framework that synthesizes Algorithmic programs with LLM-Generated Oracles to guide the generation and verify their correctness. ALGO first generates a reference oracle by prompting an LLM to exhaustively enumerate all the combinations of relevant variables. This oracle is then utilized to guide an arbitrary search strategy in exploring the algorithm space and to verify the synthesized algorithms. Our study shows that the LLM-generated
oracles are correct for 88% of the cases. With the oracles as verifiers, ALGO can be integrated with any existing code generation model in a model-agnostic manner to enhance its performance. Experiments show that when equipped with ALGO, we achieve an 8× better one-submission pass rate over the Codex model and a 2.6× better one-submission pass rate over CodeT, the current state-of-the-art model on CodeContests. We can also get 1.3× better pass rate over the ChatGPT Code Interpreter on unseen problems. The problem set we used for testing, the prompts we used, the verifier and solution programs, and the test cases generated by ALGO
are available at https://github.com/zkx06111/ALGO. Kexun Zhang, Danqing Wang, Jingtao Xia, William Yang Wang, Lei Li 0005 |
NeurIPS | 5 |
| 2022 | Non-autoregressive Translation with Layer-Wise Prediction and Deep SupervisionabstractHow do we perform efficient inference while retaining high translation quality? Existing neural machine translation models, such as Transformer, achieve high performance, but they decode words one by one, which is inefficient. Recent non-autoregressive translation models speed up the inference, but their quality is still inferior. In this work, we propose DSLP, a highly efficient and high-performance model for machine translation. The key insight is to train a non-autoregressive Transformer with Deep Supervision and feed additional Layer-wise Predictions. We conducted extensive experiments on four translation tasks (both directions of WMT'14 EN-DE and WMT'16 EN-RO). Results show that our approach consistently improves the BLEU scores compared with respective base models. Specifically, our best variant outperforms the autoregressive model on three translation tasks, while being 14.8 times more efficient in inference. Chenyang Huang 0001, Hao Zhou 0012, Osmar R. Zaïane, Lili Mou, Lei Li 0005 |
AAAI | 5 |
| 2022 | LOREN: Logic-Regularized Reasoning for Interpretable Fact VerificationabstractGiven a natural language statement, how to verify its veracity against a large-scale textual knowledge source like Wikipedia? Most existing neural models make predictions without giving clues about which part of a false claim goes wrong. In this paper, we propose LOREN, an approach for interpretable fact verification. We decompose the verification of the whole claim at phrase-level, where the veracity of the phrases serves as explanations and can be aggregated into the final verdict according to logical rules. The key insight of LOREN is to represent claim phrase veracity as three-valued latent variables, which are regularized by aggregation logical rules. The final claim verification is based on all latent variables. Thus, LOREN enjoys the additional benefit of interpretability --- it is easy to explain how it reaches certain results with claim phrase veracity. Experiments on a public fact verification benchmark show that LOREN is competitive against previous approaches while enjoying the merit of faithful and accurate interpretability. The resources of LOREN are available at: https://github.com/jiangjiechen/LOREN. Jiangjie Chen, Qiaoben Bao, Changzhi Sun, Xinbo Zhang, Jiaze Chen, Hao Zhou 0012, Yanghua Xiao, Lei Li 0005 |
AAAI | 8 |
| 2022 | Unsupervised Editing for Counterfactual StoriesabstractCreating what-if stories requires reasoning about prior statements and possible outcomes of the changed conditions. One can easily generate coherent endings under new conditions, but it would be challenging for current systems to do it with minimal changes to the original story. Therefore, one major challenge is the trade-off between generating a logical story and rewriting with minimal-edits. In this paper, we propose EDUCAT, an editing-based unsupervised approach for counterfactual story rewriting. EDUCAT includes a target position detection strategy based on estimating causal effects of the what-if conditions, which keeps the causal invariant parts of the story. EDUCAT then generates the stories under fluency, coherence and minimal-edits constraints. We also propose a new metric to alleviate the shortcomings of current automatic metrics and better evaluate the trade-off. We evaluate EDUCAT on a public counterfactual story rewriting benchmark. Experiments show that EDUCAT achieves the best trade-off over unsupervised SOTA methods according to both automatic and human evaluation. The resources of EDUCAT are available at: https://github.com/jiangjiechen/EDUCAT. Jiangjie Chen, Chun Gan, Sijie Cheng, Hao Zhou 0012, Yanghua Xiao, Lei Li 0005 |
AAAI | 6 |
| 2022 | latent-GLAT: Glancing at Latent Variables for Parallel Text GenerationabstractYu Bao, Hao Zhou, Shujian Huang, Dongqi Wang, Lihua Qian, Xinyu Dai, Jiajun Chen, Lei Li. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022. Hao Zhou 0012, Shujian Huang, Dongqi Wang 0005, Lihua Qian, Xinyu Dai, Jiajun Chen 0001, Lei Li 0005 |
ACL (1) | 8 |
| 2022 | Learning When to Translate for Streaming SpeechabstractHow to find proper moments to generate partial sentence translation given a streaming speech input?Existing approaches waitingand-translating for a fixed duration often break the acoustic units in speech, since the boundaries between acoustic units in speech are not even.In this paper, we propose MoSST, a simple yet effective method for translating streaming speech content.Given a usually long speech sequence, we develop an efficient monotonic segmentation module inside an encoder-decoder model to accumulate acoustic information incrementally and detect proper speech unit boundaries for the input in speech translation task.Experiments on multiple translation directions of the MuST-C dataset show that MoSST outperforms existing methods and achieves the best trade-off between translation quality (BLEU) and latency.Our code is available at https://github. com/dqqcasia/mosst. Yaoming Zhu, Mingxuan Wang, Lei Li 0005 |
ACL (1) | 4 |
| 2022 | STEMM: Self-learning with Speech-text Manifold Mixup for Speech TranslationabstractHow to learn a better speech representation for end-to-end speech-to-text translation (ST) with limited labeled data?Existing techniques often attempt to transfer powerful machine translation (MT) capabilities to ST, but neglect the representation discrepancy across modalities.In this paper, we propose the Speech-TExt Manifold Mixup (STEMM) method to calibrate such discrepancy.Specifically, we mix up the representation sequences of different modalities, and take both unimodal speech sequences and multimodal mixed sequences as input to the translation model in parallel, and regularize their output predictions with a selflearning framework.Experiments on MuST-C speech translation benchmark and further analysis show that our method effectively alleviates the cross-modal representation discrepancy, and achieves significant improvements over a strong baseline on eight translation directions.* indicates corresponding authors. Qingkai Fang, Rong Ye, Lei Li 0005, Yang Feng 0004, Mingxuan Wang |
ACL (1) | 3 |
| 2022 | Contextual Representation Learning beyond Masked Language ModelingabstractHow do masked language models (MLMs) such as BERT learn contextual representations?In this work, we analyze the learning dynamics of MLMs.We find that MLMs adopt sampled embeddings as anchors to estimate and inject contextual semantics to representations, which limits the efficiency and effectiveness of MLMs.To address these issues, we propose TACO, a simple yet effective representation learning approach to directly model global semantics.TACO extracts and aligns contextual semantics hidden in contextualized representations to encourage models to attend global semantics when generating contextualized representations.Experiments on the GLUE benchmark show that TACO achieves up to 5x speedup and up to 1.2 points average improvement over existing MLMs.The code is available at https:// github.com/FUZHIYI/TACO. Zhiyi Fu, Wangchunshu Zhou, Jingjing Xu 0001, Hao Zhou 0012, Lei Li 0005 |
ACL (1) | 5 |
| 2022 | switch-GLAT: Multilingual Parallel Machine Translation Via Code-Switch Decoder
Zhenqiao Song, Hao Zhou 0012, Lihua Qian, Jingjing Xu 0001, Shanbo Cheng, Mingxuan Wang, Lei Li 0005 |
ICLR | 7 |
| 2022 | Enhancing Cross-lingual Transfer by Manifold Mixup
Huiyun Yang, Huadong Chen, Hao Zhou 0012, Lei Li 0005 |
ICLR | 4 |
| 2022 | On the Learning of Non-Autoregressive TransformersabstractNon-autoregressive Transformer (NAT) is a family of text generation models, which aims to reduce the decoding latency by predicting the whole sentences in parallel. However, such latency reduction sacrifices the ability to capture left-to-right dependencies, thereby making NAT learning very challenging. In this paper, we present theoretical and empirical analyses to reveal the challenges of NAT learning and propose a unified perspective to understand existing successes. First, we show that simply training NAT by maximizing the likelihood can lead to an approximation of marginal distributions but drops all dependencies between tokens, where the dropped information can be measured by the dataset’s conditional total correlation. Second, we formalize many previous objectives in a unified framework and show that their success can be concluded as maximizing the likelihood on a proxy distribution, leading to a reduced information loss. Empirical studies show that our perspective can explain the phenomena in NAT learning and guide the design of new training methods. Fei Huang 0005, Tianhua Tao, Hao Zhou 0012, Lei Li 0005, Minlie Huang |
ICML | 4 |
| 2022 | Learning Design and Construction with Varying-Sized Materials via Prioritized Memory ResetsabstractCan a robot autonomously learn to design and construct a bridge from varying-sized blocks without a blueprint? It is a challenging task with long horizon and sparse reward - the robot has to figure out physically stable design schemes and feasible actions to manipulate and transport blocks. Due to diverse block sizes, the state space and action trajectories are vast to explore. In this paper, we propose a hierarchical approach for this problem. It consists of a reinforcement-learning designer to propose high-level building instructions and a motion-planning-based action generator to manipulate blocks at the low level. For high-level learning, we develop a novel technique, prioritized memory resetting (PMR) to improve exploration. PMR adaptively resets the state to those most critical configurations from a replay buffer so that the robot can resume training on partial architectures instead of from scratch. Furthermore, we augment PMR with auxiliary training objectives and fine-tune the designer with the locomotion generator. Our experiments in simulation and on a real deployed robotic system demonstrate that it is able to effectively construct bridges with blocks of varying sizes at a high success rate. Demos can be found at https://sites.google.com/view/bridge-pmr. Yunfei Li 0005, Tao Kong, Lei Li 0005, Yi Wu 0013 |
ICRA | 3 |
| 2022 | Uncovering the Heterogeneous Effects of Preference Diversity on User Activeness: A Dynamic Mixture ModelabstractPreference diversity arouses much research attention in recent years, as it is believed to be closely related to many profound problems such as user activeness in social media or recommendation systems. However, due to the lack of large-scale data with comprehensive user behavior log and accurate content labels, the real quantitative effect of preference diversity on user activeness is still largely unknown. This paper studies the heterogeneous effect of preference diversity on user activeness in social media. We examine large-scale real-world datasets collected from two of the most popular video-sharing social platforms in China, including the behavior logs of more than 787 thousand users and 1.95 million videos, with accurate content category information. We investigate the distribution and evolution of preference diversity, and find rich heterogeneity in the effect of preference diversity on the dynamic activeness. Furthermore, we discover the divergence of preference diversity mechanisms for the same user under different usage scenarios, such as active (where users actively seek information) and passive (where users passively receive information) modes. Unlike existing qualitative studies, we propose a universal mixture model with the capability of accurately fitting dynamic activeness curves while reflecting the heterogeneous patterns of preference diversity. To our best knowledge, this is the first quantitative model that incorporates the effect of preference diversity on user activeness. With the modeling parameters, we are able to make accurate churn and activeness predictions and provide decision support for increasing user activity through the intervention of diversity. Our findings and model comprehensively reveal the significance of preference diversity and provide potential implications for the design of future recommendation systems and social media. Yunfei Lu, Peng Cui 0001, Linyun Yu, Lei Li 0005, Wenwu Zhu 0001 |
KDD | 4 |
| 2022 | Cross-modal Contrastive Learning for Speech TranslationabstractHow can we learn unified representations for spoken utterances and their written text?Learning similar representations for semantically similar speech and text is important for speech translation.To this end, we propose ConST, a cross-modal contrastive learning method for end-to-end speech-to-text translation.We evaluate ConST and a variety of previous baselines on a popular benchmark MuST-C.Experiments show that the proposed ConST consistently outperforms the previous methods, and achieves an average BLEU of 29.4.The analysis further verifies that ConST indeed closes the representation gap of different modalities -its learned representation improves the accuracy of cross-modal speechtext retrieval from 4% to 88%.Code and models are available at https://github. com/ReneeYe/ConST. Rong Ye, Mingxuan Wang, Lei Li 0005 |
NAACL-HLT | 3 |
| 2022 | Provably Confidential Language ModellingabstractLarge language models are shown to memorize privacy information such as social security numbers in training data.Given the sheer scale of the training corpus, it is challenging to screen and filter all privacy data, either manually or automatically.In this paper, we propose Confidentially Redacted Training (CRT), a method to train language generation models while protecting the confidential segments.We borrow ideas from differential privacy (which solves a related but distinct problem) and show that our method is able to provably prevent unintended memorization by randomizing parts of the training process.Moreover, we show that redaction with an approximately correct screening policy amplifies the confidentiality guarantee.We implement the method for both LSTM and GPT language models.Our experimental results show that the models trained by CRT obtain almost the same perplexity while preserving strong confidentiality 1 . Xuandong Zhao, Lei Li 0005, Yu-Xiang Wang 0003 |
NAACL-HLT | 2 |
| 2022 | LightSeq2: Accelerated Training for Transformer-Based Models on GPUsabstractTransformer-based neural models are used in many AI applications. Training these models is expensive, as it takes huge GPU resources and long duration. It is challenging because typical data like sentences have variable lengths, and Transformer's computation patterns are more complex than convolutional neural networks. Existing systems either only focus on model inference or optimization for only BERT-like encoder models. In this paper, we present LightSeq2, a system to accelerate training for a general family of Transformer models on GPUs. We propose a series of GPU optimization techniques tailored to the specific computation flow and memory access patterns of Transformer models. LightSeq2 supports many model architectures, including BERT (encoder-only), GPT (decoder-only), Transformer (encoder-decoder), and vision Transformer. Our experiments for a variety of models and benchmarks show that LightSeq2 is consistently faster (1.4-3.5 x) than previous systems on different GPUs. In particular, it gains 308 % training speedup compared with existing systems on a large public machine translation benchmark (WMTI4 English-German). Guyue Huang, Xian Qian, Yufei Ding 0001, Mingxuan Wang, Lei Li 0005 |
SC | 8 |
| 2022 | SOLO: A Simple Framework for Instance SegmentationabstractCompared to many other dense prediction tasks, e.g., semantic segmentation, it is the arbitrary number of instances that has made instance segmentation much more challenging. In order to predict a mask for each instance, mainstream approaches either follow the "detect-then-segment" strategy (e.g., Mask R-CNN), or predict embedding vectors first then cluster pixels into individual instances. In this paper, we view the task of instance segmentation from a completely new perspective by introducing the notion of "instance categories", which assigns categories to each pixel within an instance according to the instance's location. With this notion, we propose segmenting objects by locations (SOLO), a simple, direct, and fast framework for instance segmentation with strong performance. We derive a few SOLO variants (e.g., Vanilla SOLO, Decoupled SOLO, Dynamic SOLO) following the basic principle. Our method directly maps a raw input image to the desired object categories and instance masks, eliminating the need for the grouping post-processing or the bounding box detection. Our approach achieves state-of-the-art results for instance segmentation in terms of both speed and accuracy, while being considerably simpler than the existing methods. Besides instance segmentation, our method yields state-of-the-art results in object detection (from our mask byproduct) and panoptic segmentation. We further demonstrate the flexibility and high-quality segmentation of SOLO by extending it to perform one-stage instance-level image matting. Code is available at: https://git.io/AdelaiDet. Rufeng Zhang, Chunhua Shen, Tao Kong, Lei Li 0005 |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2021 | Consecutive Decoding for Speech-to-text TranslationabstractSpeech-to-text translation (ST), which directly translates the source language speech to the target language text, has attracted intensive attention recently. However, the combination of speech recognition and machine translation in a single model poses a heavy burden on the direct cross-modal cross-lingual mapping. To reduce the learning difficulty, we propose COnSecutive Transcription and Translation (COSTT), an integral approach for speech-to-text translation. The key idea is to generate source transcript and target translation text with a single decoder. It benefits the model training so that additional large parallel text corpus can be fully exploited to enhance the speech translation training. Our method is verified on three mainstream datasets, including Augmented LibriSpeech English-French dataset, TED English-German dataset, and TED English-Chinese dataset. Experiments show that our proposed COSTT outperforms the previous state-of-the-art methods. The code is available at https://github.com/dqqcasia/st. Qianqian Dong, Mingxuan Wang, Hao Zhou 0012, Bo Xu 0002, Lei Li 0005 |
AAAI | 6 |
| 2021 | Listen, Understand and Translate: Triple Supervision Decouples End-to-end Speech-to-text TranslationabstractAn end-to-end speech-to-text translation (ST) takes audio in a source language and outputs the text in a target language. Existing methods are limited by the amount of parallel corpus. Can we build a system to fully utilize signals in a parallel ST corpus? We are inspired by human understanding system which is composed of auditory perception and cognitive processing. In this paper, we propose Listen-Understand-Translate, (LUT), a unified framework with triple supervision signals to decouple the end-to-end speech-to-text translation task. LUT is able to guide the acoustic encoder to extract as much information from the auditory input. In addition, LUT utilizes a pre-trained BERT model to enforce the upper encoder to produce as much semantic information as possible, without extra data. We perform experiments on a diverse set of speech translation benchmarks, including Librispeech English-French, IWSLT English-German and TED English-Chinese. Our results demonstrate LUT achieves the state-of-the-art performance, outperforming previous methods. The code is available at https://github.com/dqqcasia/st. Qianqian Dong, Rong Ye, Mingxuan Wang, Hao Zhou 0012, Bo Xu 0002, Lei Li 0005 |
AAAI | 7 |
| 2021 | ACMo: Angle-Calibrated Moment Methods for Stochastic Optimization
Xunpeng Huang, Runxin Xu, Hao Zhou 0012, Zhengyang Liu 0002, Lei Li 0005 |
AAAI | 6 |
| 2021 | Finding Sparse Structures for Domain Specific Neural Machine TranslationabstractNeural machine translation often adopts the fine-tuning approach to adapt to specific domains. However, nonrestricted fine-tuning can easily degrade on the general domain and over-fit to the target domain. To mitigate the issue, we propose Prune-Tune, a novel domain adaptation method via gradual pruning. It learns tiny domain-specific sub-networks during fine-tuning on new domains. Prune-Tune alleviates the over-fitting and the degradation problem without model modification. Furthermore, Prune-Tune is able to sequentially learn a single network with multiple disjoint domain-specific sub-networks for multiple domains. Empirical experiment results show that Prune-Tune outperforms several strong competitors in the target domain test set without sacrificing the quality on the general domain in both single and multi-domain settings. The source code and data are available at https://github.com/ohlionel/Prune-Tune. Jianze Liang, Chengqi Zhao, Mingxuan Wang, Xipeng Qiu, Lei Li 0005 |
AAAI | 5 |
| 2021 | TextGAIL: Generative Adversarial Imitation Learning for Text GenerationabstractGenerative Adversarial Networks (GANs) for text generation have recently received many criticisms, as they perform worse than their MLE counterparts. We suspect previous text GANs' inferior performance is due to the lack of a reliable guiding signal in their discriminators. To address this problem, we propose a generative adversarial imitation learning framework for text generation that uses large pre-trained language models to provide more reliable reward guidance. As previous text GANs suffer from high variance of gradients, we apply contrastive discriminator, and proximal policy optimization (PPO) to stabilize and improve text generation performance. For evaluation, we conduct experiments on a diverse set of unconditional and conditional text generation tasks. Experimental results show that TextGAIL achieves better performance in terms of both quality and diversity than the MLE baseline. We also validate our intuition that TextGAIL's discriminator demonstrates the capability of providing reasonable rewards with an additional task. Qingyang Wu, Lei Li 0005 |
AAAI | 2 |
| 2021 | Taxonomy Completion via Triplet Matching NetworkabstractAutomatically constructing taxonomy finds many applications in e-commerce and web search. One critical challenge is as data and business scope grow in real applications, new concepts are emerging and needed to be added to the existing taxonomy. Previous approaches focus on the taxonomy expansion, i.e. finding an appropriate hypernym concept from the taxonomy for a new query concept. In this paper, we formulate a new task, “taxonomy completion”, by discovering both the hypernym and hyponym concepts for a query. We propose Triplet Matching Network (TMN), to find the appropriate pairs for a given query concept. TMN consists of one primal scorer and multiple auxiliary scorers. These auxiliary scorers capture various fine-grained signals (e.g., query to hypernym or query to hyponym semantics), and the primal scorer makes a holistic prediction on triplet based on the internal feature representations of all auxiliary scorers. Also, an innovative channel-wise gating mechanism that retains task-specific information in concept representations is introduced to further boost model performance. Experiments on four real-world large-scale datasets show that TMN achieves the best performance on both taxonomy completion task and the previous taxonomy expansion task, outperforming existing methods. Jieyu Zhang 0001, Xiangchen Song, Jiaze Chen, Yuning Mao, Lei Li 0005 |
AAAI | 7 |
| 2021 | Learning Language Specific Sub-network for Multilingual Machine TranslationabstractZehui Lin, Liwei Wu, Mingxuan Wang, Lei Li. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Mingxuan Wang, Lei Li 0005 |
ACL/IJCNLP (1) | 4 |
| 2021 | Contrastive Learning for Many-to-many Multilingual Neural Machine TranslationabstractXiao Pan, Mingxuan Wang, Liwei Wu, Lei Li. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Mingxuan Wang, Lei Li 0005 |
ACL/IJCNLP (1) | 4 |
| 2021 | Glancing Transformer for Non-Autoregressive Neural Machine TranslationabstractLihua Qian, Hao Zhou, Yu Bao, Mingxuan Wang, Lin Qiu, Weinan Zhang, Yong Yu, Lei Li. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Lihua Qian, Hao Zhou 0012, Mingxuan Wang, Weinan Zhang 0001, Yong Yu 0001, Lei Li 0005 |
ACL/IJCNLP (1) | 8 |
| 2021 | UniRE: A Unified Label Space for Entity Relation ExtractionabstractYijun Wang, Changzhi Sun, Yuanbin Wu, Hao Zhou, Lei Li, Junchi Yan. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Changzhi Sun, Yuanbin Wu, Hao Zhou 0012, Lei Li 0005, Junchi Yan |
ACL/IJCNLP (1) | 5 |
| 2021 | Document-level Event Extraction via Heterogeneous Graph-based Interaction Model with a TrackerabstractRunxin Xu, Tianyu Liu, Lei Li, Baobao Chang. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Runxin Xu, Tianyu Liu 0001, Lei Li 0005, Baobao Chang |
ACL/IJCNLP (1) | 3 |
| 2021 | Vocabulary Learning via Optimal Transport for Neural Machine TranslationabstractJingjing Xu, Hao Zhou, Chun Gan, Zaixiang Zheng, Lei Li. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Jingjing Xu 0001, Hao Zhou 0012, Chun Gan, Zaixiang Zheng, Lei Li 0005 |
ACL/IJCNLP (1) | 5 |
| 2021 | Scale-Aware Automatic Augmentation for Object DetectionabstractWe propose Scale-aware AutoAug to learn data augmentation policies for object detection. We define a new scaleaware search space, where both image- and box-level augmentations are designed for maintaining scale invariance. Upon this search space, we propose a new search metric, termed Pareto Scale Balance, to facilitate search with high efficiency. In experiments, Scale-aware AutoAug yields significant and consistent improvement on various object detectors (e.g., RetinaNet, Faster R-CNN, Mask R-CNN, and FCOS), even compared with strong multi-scale training baselines. Our searched augmentation policies are transferable to other datasets and box-level tasks beyond object detection (e.g., instance segmentation and keypoint estimation) to improve performance. The search cost is much less than previous automated augmentation approaches for object detection. It is notable that our searched policies have meaningful patterns, which intuitively provide valuable insight for human data augmentation design. Code and models are available at https://github.com/Jia-ResearchLab/SA-AutoAug. Yukang Chen, Tao Kong, Lu Qi 0001, Ruihang Chu, Lei Li 0005, Jiaya Jia |
CVPR | 6 |
| 2021 | Locate Then Segment: A Strong Pipeline for Referring Image SegmentationabstractReferring image segmentation aims to segment the objects referred by a natural language expression. Previous methods usually focus on designing an implicit and recurrent feature interaction mechanism to fuse the visual-linguistic features to directly generate the final segmentation mask without explicitly modeling the localization information of the referent instances. To tackle these problems, we view this task from another perspective by decoupling it into a "Locate-Then-Segment" (LTS) scheme. Given a language expression, people generally first perform attention to the corresponding target image regions, then generate a fine segmentation mask about the object based on its context. The LTS first extracts and fuses both visual and textual features to get a cross-modal representation, then applies a cross-model interaction on the visual-textual features to locate the referred object with position prior, and finally generates the segmentation result with a light-weight segmentation network. Our LTS is simple but surprisingly effective. On three popular benchmark datasets, the LTS outperforms all the previous state-of-the-arts methods by a large margin (e.g., +3.2% on RefCOCO+ and +3.4% on RefCOCOg). In addition, our model is more interpretable with explicitly locating the object, which is also proved by visualization experiments. We believe this framework is promising to serve as a strong baseline for referring image segmentation. Ya Jing, Tao Kong, Wei Wang 0115, Liang Wang 0001, Lei Li 0005, Tieniu Tan |
CVPR | 5 |
| 2021 | Sparse R-CNN: End-to-End Object Detection With Learnable ProposalsabstractWe present Sparse R-CNN, a purely sparse method for object detection in images. Existing works on object detection heavily rely on dense object candidates, such as k anchor boxes pre-defined on all grids of image feature map of size H × W. In our method, however, a fixed sparse set of learned object proposals, total length of N, are provided to object recognition head to perform classification and location. By eliminating HWk (up to hundreds of thousands) hand-designed object candidates to N (e.g. 100) learnable proposals, Sparse R-CNN completely avoids all efforts related to object candidates design and many-to-one label assignment. More importantly, final predictions are directly output without non-maximum suppression post-procedure. Sparse R-CNN demonstrates accuracy, run-time and training convergence performance on par with the well-established detector baselines on the challenging COCO dataset, e.g., achieving 45.0 AP in standard 3× training schedule and running at 22 fps using ResNet-50 FPN model. We hope our work could inspire re-thinking the convention of dense prior in object detectors. The code is available at: https://github.com/PeizeSun/SparseR-CNN. Peize Sun, Rufeng Zhang, Yi Jiang 0009, Tao Kong, Chenfeng Xu, Masayoshi Tomizuka, Lei Li 0005, Zehuan Yuan, Changhu Wang, Ping Luo 0002 |
CVPR | 8 |
| 2021 | Dense Contrastive Learning for Self-Supervised Visual Pre-TrainingabstractTo date, most existing self-supervised learning methods are designed and optimized for image classification. These pre-trained models can be sub-optimal for dense prediction tasks due to the discrepancy between image-level prediction and pixel-level prediction. To fill this gap, we aim to design an effective, dense self-supervised learning method that directly works at the level of pixels (or local features) by taking into account the correspondence between local features. We present dense contrastive learning (DenseCL), which implements self-supervised learning by optimizing a pairwise contrastive (dis)similarity loss at the pixel level between two views of input images.Compared to the baseline method MoCo-v2, our method introduces negligible computation overhead (only <1% slower), but demonstrates consistently superior performance when transferring to downstream dense prediction tasks including object detection, semantic segmentation and instance segmentation; and outperforms the state-of-the-art methods by a large margin. Specifically, over the strong MoCo-v2 baseline, our method achieves significant improvements of 2.0% AP on PASCAL VOC object detection, 1.1% AP on COCO object detection, 0.9% AP on COCO instance segmentation, 3.0% mIoU on PASCAL VOC semantic segmentation and 1.8% mIoU on Cityscapes semantic segmentation.Code and models are available at: https://git.io/DenseCL Rufeng Zhang, Chunhua Shen, Tao Kong, Lei Li 0005 |
CVPR | 5 |
| 2021 | ENPAR: Enhancing Entity and Entity Pair Representations for Joint Entity Relation ExtractionabstractCurrent state-of-the-art systems for joint entity relation extraction (Luan et al., 2019;Wadden et al., 2019) usually adopt the multi-task learning framework.However, annotations for these additional tasks such as coreference resolution and event extraction are always equally hard (or even harder) to obtain.In this work, we propose a pre-training method ENPAR to improve the joint extraction performance.EN-PAR requires only the additional entity annotations that are much easier to collect.Unlike most existing works that only consider incorporating entity information into the sentence encoder, we further utilize the entity pair information.Specifically, we devise four novel objectives, i.e., masked entity typing, masked entity prediction, adversarial context discrimination, and permutation prediction, to pretrain an entity encoder and an entity pair encoder.Comprehensive experiments show that the proposed pre-training method achieves significant improvement over BERT on ACE05, SciERC, and NYT, and outperforms current state-of-the-art on ACE05. Changzhi Sun, Yuanbin Wu, Hao Zhou 0012, Lei Li 0005, Junchi Yan |
EACL | 5 |
| 2021 | Learning Kernel-Smoothed Machine Translation with Retrieved ExamplesabstractHow to effectively adapt neural machine translation (NMT) models according to emerging cases without retraining?Despite the great success of neural machine translation, updating the deployed models online remains a challenge.Existing non-parametric approaches that retrieve similar examples from a database to guide the translation process are promising but are prone to overfit the retrieved examples.However, non-parametric methods are prone to overfit the retrieved examples.In this work, we propose to learn Kernel-Smoothed Translation with Example Retrieval (KSTER), an effective approach to adapt neural machine translation models online.Experiments on domain adaptation and multi-domain machine translation datasets show that even without expensive retraining, KSTER is able to achieve improvement of 1.1 to 1.5 BLEU scores over the best existing online adaptation methods.The code and trained models are released at https://github.com/jiangqn/KSTER. Qingnan Jiang, Mingxuan Wang, Shanbo Cheng, Shujian Huang, Lei Li 0005 |
EMNLP (1) | 6 |
| 2021 | Learning Logic Rules for Document-Level Relation ExtractionabstractDocument-level relation extraction aims to identify relations between entities in a whole document.Prior efforts to capture long-range dependencies have relied heavily on implicitly powerful representations learned through (graph) neural networks, which makes the model less transparent.To tackle this challenge, in this paper, we propose LogiRE, a novel probabilistic model for document-level relation extraction by learning logic rules.Lo-giRE treats logic rules as latent variables and consists of two modules: a rule generator and a relation extractor.The rule generator is to generate logic rules potentially contributing to final predictions, and the relation extractor outputs final predictions based on the generated logic rules.Those two modules can be efficiently optimized with the expectationmaximization (EM) algorithm.By introducing logic rules into neural networks, LogiRE can explicitly capture long-range dependencies as well as enjoy better interpretation.Empirical results show that LogiRE significantly outperforms several strong baselines in terms of relation performance (∼1.8 F1 score) and logical consistency (over 3.3 logic score).Our code is available at https://github.com/rudongyu/LogiRE. Dongyu Ru, Changzhi Sun, Jiangtao Feng, Hao Zhou 0012, Weinan Zhang 0001, Yong Yu 0001, Lei Li 0005 |
EMNLP (1) | 8 |
| 2021 | Gradient-Based Adversarial Factual Consistency Evaluation for Abstractive SummarizationabstractNeural abstractive summarization systems have gained significant progress in recent years.However, abstractive summarization often produce inconsisitent statements or false facts.How to automatically generate highly abstract yet factually correct summaries?In this paper, we proposed an efficient weaksupervised adversarial data augmentation approach to form the factual consistency dataset.Based on the artificial dataset, we train an evaluation model that can not only make accurate and robust factual consistency discrimination but is also capable of making interpretable factual errors tracing by backpropagated gradient distribution on token embeddings.Experiments and analysis conduct on public annotated summarization and factual consistency datasets demonstrate our approach effective and reasonable.Our codes can be found at https://github.com/ parZival27/GrAdualCC Zhiyuan Zeng 0002, Jiaze Chen, Weiran Xu, Lei Li 0005 |
EMNLP (1) | 4 |
| 2021 | MARS: Markov Molecular Sampling for Multi-objective Drug Discovery
Yutong Xie 0007, Chence Shi, Hao Zhou 0012, Weinan Zhang 0001, Yong Yu 0001, Lei Li 0005 |
ICLR | 7 |
| 2021 | Adversarial Option-Aware Hierarchical Imitation LearningabstractIt has been a challenge to learning skills for an agent from long-horizon unannotated demonstrations. Existing approaches like Hierarchical Imitation Learning(HIL) are prone to compounding errors or suboptimal solutions. In this paper, we propose Option-GAIL, a novel method to learn skills at long horizon. The key idea of Option-GAIL is modeling the task hierarchy by options and train the policy via generative adversarial optimization. In particular, we propose an Expectation-Maximization(EM)-style algorithm: an E-step that samples the options of expert conditioned on the current learned policy, and an M-step that updates the low- and high-level policies of agent simultaneously to minimize the newly proposed option-occupancy measurement between the expert and the agent. We theoretically prove the convergence of the proposed algorithm. Experiments show that Option-GAIL outperforms other counterparts consistently across a variety of tasks. Mingxuan Jing, Wenbing Huang 0001, Fuchun Sun 0001, Xiaojian Ma 0001, Tao Kong, Chuang Gan 0001, Lei Li 0005 |
ICML | 7 |
| 2021 | End-to-End Speech Translation via Cross-Modal Progressive TrainingabstractEnd-to-end speech translation models have become a new trend in research due to their potential of reducing error propagation.However, these models still suffer from the challenge of data scarcity.How to effectively use unlabeled or other parallel corpora from machine translation is promising but still an open problem.In this paper, we propose Cross Speech-Text Network (XSTNet), an end-to-end model for speech-to-text translation.XSTNet takes both speech and text as input and outputs both transcription and translation text.The model benefits from its three key design aspects: a self-supervised pretrained sub-network as the audio encoder, a multi-task training objective to exploit additional parallel bilingual text, and a progressive training procedure.We evaluate the performance of XSTNet and baselines on the MuST-C En-X and LibriSpeech En-Fr datasets.In particular, XSTNet achieves state-of-the-art results on all language directions with an average BLEU of 28.8, outperforming the previous best method by 3.2 BLEU.Code, models, cases, and more detailed analysis are available at https://github.com/ReneeYe/XSTNet. Rong Ye, Mingxuan Wang, Lei Li 0005 |
Interspeech | 3 |
| 2021 | Simultaneous Semantic and Collision Learning for 6-DoF Grasp Pose EstimationabstractGrasping in cluttered scenes has always been a great challenge for robots, due to the requirement of the ability to well understand the scene and object information. Previous works usually assume that the geometry information of the objects is available, or utilize a step-wise, multi-stage strategy to predict the feasible 6-DoF grasp poses. In this work, we propose to formalize the 6-DoF grasp pose estimation as a simultaneous multi-task learning problem. In a unified framework, we jointly predict the feasible 6-DoF grasp poses, instance semantic segmentation, and collision information. The whole framework is jointly optimized and end-to-end differentiable. Our model is evaluated on large-scale benchmarks as well as the real robot system. On the public dataset, our method outperforms prior state-of-the-art methods by a large margin (+4.08 AP). We also demonstrate the implementation of our model on a real robotic platform and show that the robot can accurately grasp target objects in cluttered scenarios with a high success rate. Project link: https://openbyterobotics.github.io/sscl. Tao Kong, Ruihang Chu, Peng Wang 0024, Lei Li 0005 |
IROS | 6 |
| 2021 | Learning to Design and Construct Bridge without BlueprintabstractAutonomous assembly has been a desired functionality of many intelligent robot systems. We study a new challenging assembly task, designing and constructing a bridge without a blueprint. In this task, the robot needs to first design a feasible bridge architecture for arbitrarily wide cliffs and then manipulate the blocks reliably to construct a stable bridge according to the proposed design. In this paper, we propose a bi-level approach to tackle this task. At the high level, the system learns a bridge blueprint policy in a physical simulator using deep reinforcement learning and curriculum learning. A policy is represented as an attention-based neural network with object-centric input, which enables generalization to different number of blocks and cliff widths. For low-level control, we implement a motion-planning-based policy for real-robot motion control, which can be directly combined with a trained blueprint policy for real-world bridge construction without tuning. In our field study, our bi-level robot system demonstrates the capability of manipulating blocks to construct a diverse set of bridges with different architectures. Yunfei Li 0005, Tao Kong, Lei Li 0005, Yi Wu 0013 |
IROS | 3 |
| 2021 | Generative Imagination Elevates Machine TranslationabstractThere are common semantics shared across text and images.Given a sentence in a source language, whether depicting the visual scene helps translation into a target language?Existing multimodal neural machine translation methods (MNMT) require triplets of bilingual sentence -image for training and tuples of source sentence -image for inference.In this paper, we propose ImagiT, a novel machine translation method via visual imagination.ImagiT first learns to generate visual representation from the source sentence, and then utilizes both source sentence and the "imagined representation" to produce a target translation.Unlike previous methods, it only needs the source sentence at the inference time.Experiments demonstrate that ImagiT benefits from visual imagination and significantly outperforms the text-only neural machine translation baselines.Further analysis reveals that the imagination process in ImagiT helps fill in missing information when performing the degradation strategy. Quanyu Long, Mingxuan Wang, Lei Li 0005 |
NAACL-HLT | 3 |
| 2021 | Duplex Sequence-to-Sequence Learning for Reversible Machine TranslationabstractSequence-to-sequence learning naturally has two directions. How to effectively utilize supervision signals from both directions? Existing approaches either require two separate models, or a multitask-learned model but with inferior performance. In this paper, we propose REDER (Reversible Duplex Transformer), a parameter-efficient model and apply it to machine translation. Either end of REDER can simultaneously input and output a distinct language. Thus REDER enables {\em reversible machine translation} by simply flipping the input and output ends. Experiments verify that REDER achieves the first success of reversible machine translation, which helps outperform its multitask-trained baselines by up to 1.3 BLEU. Zaixiang Zheng, Hao Zhou 0012, Shujian Huang, Jiajun Chen 0001, Jingjing Xu 0001, Lei Li 0005 |
NeurIPS | 6 |
| 2021 | CNewSum: A Large-Scale Summarization Dataset with Human-Annotated Adequacy and Deducibility Level
Danqing Wang, Jiaze Chen, Xianze Wu, Hao Zhou 0012, Lei Li 0005 |
NLPCC (1) | 5 |
| 2021 | Follow Your Path: A Progressive Method for Knowledge Distillation
Wenxian Shi, Yuxuan Song 0002, Hao Zhou 0012, Lei Li 0005 |
ECML/PKDD (3) | 5 |
| 2021 | Triangular Bidword Generation for Sponsored Search AuctionabstractSponsored search auction is a crucial component of modern search engines. It requires a set of candidate bidwords that advertisers can place bids on. Existing methods generate bidwords from search queries or advertisement content. However, they suffer from the data noise in (query, bidword) and (advertisement, bidword) pairs. In this paper, we propose a triangular bidword generation model (TRIDENT), which takes the high-quality data of paired (query, advertisement) as a supervision signal to indirectly guide the bidword generation process. Our proposed model is simple yet effective: by using bidword as the bridge between search query and advertisement, the generation of search query, advertisement and bidword can be jointly learned in the triangular training framework. This alleviates the problem that the training data of bidword may be noisy. Experimental results, including automatic and human evaluations, show that our proposed TRIDENT can generate relevant and diverse bidwords for both search queries and advertisements. Our evaluation on online real data validates the effectiveness of the TRIDENT's generated bidwords for product search. Zhenqiao Song, Jiaze Chen, Hao Zhou 0012, Lei Li 0005 |
WSDM | 4 |
| 2020 | SPAN: A Stochastic Projected Approximate Newton MethodabstractSecond-order optimization methods have desirable convergence properties. However, the exact Newton method requires expensive computation for the Hessian and its inverse. In this paper, we propose SPAN, a novel approximate and fast Newton method. SPAN computes the inverse of the Hessian matrix via low-rank approximation and stochastic Hessian-vector products. Our experiments on multiple benchmark datasets demonstrate that SPAN outperforms existing first-order and second-order optimization methods in terms of the convergence wall-clock time. Furthermore, we provide a theoretical analysis of the per-iteration complexity, the approximation error, and the convergence rate. Both the theoretical analysis and experimental results show that our proposed method achieves a better trade-off between the convergence rate and the per-iteration efficiency. Xunpeng Huang, Xianfeng Liang, Zhengyang Liu 0002, Lei Li 0005, Yitan Li |
AAAI | 4 |
| 2020 | Task-Aware Monocular Depth Estimation for 3D Object DetectionabstractMonocular depth estimation enables 3D perception from a single 2D image, thus attracting much research attention for years. Almost all methods treat foreground and background regions (“things and stuff”) in an image equally. However, not all pixels are equal. Depth of foreground objects plays a crucial role in 3D object recognition and localization. To date how to boost the depth prediction accuracy of foreground objects is rarely discussed. In this paper, we first analyze the data distributions and interaction of foreground and background, then propose the foreground-background separated monocular depth estimation (ForeSeE) method, to estimate the foreground and background depth using separate optimization objectives and decoders. Our method significantly improves the depth estimation performance on foreground objects. Applying ForeSeE to 3D object detection, we achieve 7.5 AP gains and set new state-of-the-art results among other monocular methods. Code will be available at: https://github.com/WXinlong/ForeSeE. Wei Yin 0006, Tao Kong, Yuning Jiang 0001, Lei Li 0005, Chunhua Shen |
AAAI | 5 |
| 2020 | Importance-Aware Learning for Neural Headline EditingabstractMany social media news writers are not professionally trained. Therefore, social media platforms have to hire professional editors to adjust amateur headlines to attract more readers. We propose to automate this headline editing process through neural network models to provide more immediate writing support for these social media news writers. To train such a neural headline editing model, we collected a dataset which contains articles with original headlines and professionally edited headlines. However, it is expensive to collect a large number of professionally edited headlines. To solve this low-resource problem, we design an encoder-decoder model which leverages large scale pre-trained language models. We further improve the pre-trained model's quality by introducing a headline generation task as an intermediate task before the headline editing task. Also, we propose Self Importance-Aware (SIA) loss to address the different levels of editing in the dataset by down-weighting the importance of easily classified tokens and sentences. With the help of Pre-training, Adaptation, and SIA, the model learns to generate headlines in the professional editor's style. Experimental results show that our method significantly improves the quality of headline editing comparing against previous methods. Qingyang Wu, Lei Li 0005, Hao Zhou 0012 |
AAAI | 2 |
| 2020 | Towards Making the Most of BERT in Neural Machine TranslationabstractGPT-2 and BERT demonstrate the effectiveness of using pre-trained language models (LMs) on various natural language processing tasks. However, LM fine-tuning often suffers from catastrophic forgetting when applied to resource-rich tasks. In this work, we introduce a concerted training framework (CTnmt) that is the key to integrate the pre-trained LMs to neural machine translation (NMT). Our proposed CTnmt} consists of three techniques: a) asymptotic distillation to ensure that the NMT model can retain the previous pre-trained knowledge; b) a dynamic switching gate to avoid catastrophic forgetting of pre-trained knowledge; and c) a strategy to adjust the learning paces according to a scheduled policy. Our experiments in machine translation show CTnmt gains of up to 3 BLEU score on the WMT14 English-German language pair which even surpasses the previous state-of-the-art pre-training aided NMT by 1.4 BLEU score. While for the large WMT14 English-French task with 40 millions of sentence-pairs, our base model still significantly improves upon the state-of-the-art Transformer big model by more than 1 BLEU score. Mingxuan Wang, Hao Zhou 0012, Chengqi Zhao, Weinan Zhang 0001, Yong Yu 0001, Lei Li 0005 |
AAAI | 7 |
| 2020 | Do you have the right scissors? Tailoring Pre-trained Language Models via Monte-Carlo MethodsabstractIt has been a common approach to pre-train a language model on a large corpus and finetune it on task-specific data.In practice, we observe that fine-tuning a pre-trained model on a small dataset may lead to over-and/or under-estimation problem.In this paper, we propose MC-Tailor, a novel method to alleviate the above issue in text generation tasks by truncating and transferring the probability mass from over-estimated regions to underestimated ones.Experiments on a variety of text generation datasets show that MC-Tailor consistently and significantly outperforms the fine-tuning approach.Our code is available at https://github.com/NingMiao/ MC-tailor. Ning Miao, Yuxuan Song 0002, Hao Zhou 0012, Lei Li 0005 |
ACL | 4 |
| 2020 | Improving Maximum Likelihood Training for Text Generation with Density Ratio EstimationabstractAutoregressive neural sequence generative models trained by Maximum Likelihood Estimation suffer the exposure bias problem in practical finite sample scenarios. The crux is that the number of training samples for Maximum Likelihood Estimation is usually limited and the input data distributions are different at training and inference stages. Many methods have been proposed to solve the above problem, which relies on sampling from the non-stationary model distribution and suffers from high variance or biased estimations. In this paper, we propose $\psi$-MLE, a new training scheme for autoregressive sequence generative models, which is effective and stable when operating at large sample space encountered in text generation. We derive our algorithm from a new perspective of self-augmentation and introduce bias correction with density ratio estimation. Extensive experimental results on synthetic data and real-world text generation tasks demonstrate that our method stably outperforms Maximum Likelihood Estimation and other state-of-the-art sequence generative models in terms of both quality and diversity. Yuxuan Song 0002, Ning Miao, Hao Zhou 0012, Lantao Yu, Mingxuan Wang, Lei Li 0005 |
AISTATS | 6 |
| 2020 | SOLO: Segmenting Objects by Locations
Tao Kong, Chunhua Shen, Yuning Jiang 0001, Lei Li 0005 |
ECCV (18) | 5 |
| 2020 | On the Sentence Embeddings from Pre-trained Language ModelsabstractPre-trained contextual representations like BERT have achieved great success in natural language processing.However, the sentence embeddings from the pre-trained language models without fine-tuning have been found to poorly capture semantic meaning of sentences.In this paper, we argue that the semantic information in the BERT embeddings is not fully exploited.We first reveal the theoretical connection between the masked language model pre-training objective and the semantic similarity task theoretically, and then analyze the BERT sentence embeddings empirically.We find that BERT always induces a non-smooth anisotropic semantic space of sentences, which harms its performance of semantic similarity.To address this issue, we propose to transform the anisotropic sentence embedding distribution to a smooth and isotropic Gaussian distribution through normalizing flows that are learned with an unsupervised objective.Experimental results show that our proposed BERT-flow method obtains significant performance gains over the state-of-the-art sentence embeddings on a variety of semantic textual similarity tasks.The code is available at https://github.com/ bohanli/BERT-flow. Hao Zhou 0012, Junxian He, Mingxuan Wang, Yiming Yang 0002, Lei Li 0005 |
EMNLP (1) | 6 |
| 2020 | Pre-training Multilingual Neural Machine Translation by Leveraging Alignment InformationabstractWe investigate the following question for machine translation (MT): can we develop a single universal MT model to serve as the common seed and obtain derivative and improved models on arbitrary language pairs?We propose mRASP, an approach to pre-train a universal multilingual neural machine translation model.Our key idea in mRASP is its novel technique of random aligned substitution, which brings words and phrases with similar meanings across multiple languages closer in the representation space.We pre-train a mRASP model on 32 language pairs jointly with only public datasets.The model is then fine-tuned on downstream language pairs to obtain specialized MT models.We carry out extensive experiments on 42 translation directions across a diverse settings, including low, medium, rich resource, and as well as transferring to exotic language pairs.Experimental results demonstrate that mRASP achieves significant performance improvement compared to directly training on those target pairs.It is the first time to verify that multiple lowresource language pairs can be utilized to improve rich resource MT.Surprisingly, mRASP is even able to improve the translation quality on exotic languages that never occur in the pretraining corpus.Code, data, and pre-trained models are available at https://github. com/linzehui/mRASP. Mingxuan Wang, Xipeng Qiu, Jiangtao Feng, Hao Zhou 0012, Lei Li 0005 |
EMNLP (1) | 7 |
| 2020 | Double Graph Based Reasoning for Document-level Relation ExtractionabstractDocument-level relation extraction aims to extract relations among entities within a document.Different from sentence-level relation extraction, it requires reasoning over multiple sentences across paragraphs.In this paper, we propose Graph Aggregation-and-Inference Network (GAIN), a method to recognize such relations for long paragraphs.GAIN constructs two graphs, a heterogeneous mentionlevel graph (MG) and an entity-level graph (EG).The former captures complex interaction among different mentions and the latter aggregates mentions underlying for the same entities.Based on the graphs we propose a novel path reasoning mechanism to infer relations between entities.Experiments on the public dataset, DocRED, show GAIN achieves a significant performance improvement (2.85 on F1) over the previous state-of-the-art.Our code is available at https://github.com/ PKUnlp-icler/GAIN. Shuang Zeng, Runxin Xu, Baobao Chang, Lei Li 0005 |
EMNLP (1) | 4 |
| 2020 | Variational Template Machine for Data-to-Text Generation
Rong Ye, Wenxian Shi, Hao Zhou 0012, Zhongyu Wei, Lei Li 0005 |
ICLR | 5 |
| 2020 | Mirror-Generative Neural Machine Translation
Zaixiang Zheng, Hao Zhou 0012, Shujian Huang, Lei Li 0005, Xinyu Dai, Jiajun Chen 0001 |
ICLR | 4 |
| 2020 | Dispersed Exponential Family Mixture VAEs for Interpretable Text GenerationabstractDeep generative models are commonly used for generating images and text. Interpretability of these models is one important pursuit, other than the generation quality. Variational auto-encoder (VAE) with Gaussian distribution as prior has been successfully applied in text generation, but it is hard to interpret the meaning of the latent variable. To enhance the controllability and interpretability, one can replace the Gaussian prior with a mixture of Gaussian distributions (GM-VAE), whose mixture components could be related to hidden semantic aspects of data. In this paper, we generalize the practice and introduce DEM-VAE, a class of models for text generation using VAEs with a mixture distribution of exponential family. Unfortunately, a standard variational training algorithm fails due to the \emph{mode-collapse} problem. We theoretically identify the root cause of the problem and propose an effective algorithm to train DEM-VAE. Our method penalizes the training with an extra \emph{dispersion term} to induce a well-structured latent space. Experimental results show that our approach does obtain a meaningful space, and it outperforms strong baselines in text generation benchmarks. The code is available at \url{https://github.com/wenxianxian/demvae}. Wenxian Shi, Hao Zhou 0012, Ning Miao, Lei Li 0005 |
ICML | 4 |
| 2020 | SOLOv2: Dynamic and Fast Instance SegmentationabstractIn this work, we design a simple, direct, and fast framework for instance segmentation with strong performance. To this end, we propose a novel and effective approach, termed SOLOv2, following the principle of the SOLO method [32]. First, our new framework is empowered by an efficient and holistic instance mask representation scheme, which dynamically segments each instance in the image, without resorting to bounding box detection. Specifically, the object mask generation is decoupled into a mask kernel prediction and mask feature learning, which are responsible for generating convolution kernels and the feature maps to be convolved with, respectively. Second, SOLOv2 significantly reduces inference overhead with our novel matrix non-maximum suppression (NMS) technique. Our Matrix NMS performs NMS with parallel matrix operations in one shot, and yields better results. We demonstrate that the proposed SOLOv2 achieves the state-of-the- art performance with high efficiency, making it suitable for both mobile and cloud applications. A light-weight version of SOLOv2 executes at 31.3 FPS and yields 37.1% AP on COCO test-dev. Moreover, our state-of-the-art results in object detection (from our mask byproduct) and panoptic segmentation show the potential of SOLOv2 to serve as a new strong baseline for many instance-level recognition tasks. Code is available at https://git.io/AdelaiDet Rufeng Zhang, Tao Kong, Lei Li 0005, Chunhua Shen |
NeurIPS | 4 |
| 2020 | QuAChIE: Question Answering based Chinese Information Extraction SystemabstractIn this paper, we present the design of QuAChIE, a Question Answering based Chinese Information Extraction system. QuAChIE mainly depends on a well-trained question answering model to extract high-quality triples. The group of head entity and relation are regarded as a question given the input text as the context. For the training and evaluation of each model in the system, we build a large-scale information extraction dataset using Wikidata and Wikipedia pages by distant supervision. The advanced models implemented on top of the pre-trained language model and the enormous distant supervision data enable QuAChIE to extract relation triples from documents with cross-sentence correlations. The experimental results on the test set and the case study based on the interactive demonstration show its satisfactory Information Extraction quality on Chinese document-level texts. Dongyu Ru, Zhenghui Wang, Hao Zhou 0012, Lei Li 0005, Weinan Zhang 0001, Yong Yu 0001 |
SIGIR | 5 |
| 2020 | FoveaBox: Beyound Anchor-Based Object DetectionabstractWe present FoveaBox, an accurate, flexible, and completely anchor-free framework for object detection. While almost all state-of-the-art object detectors utilize predefined anchors to enumerate possible locations, scales and aspect ratios for the search of the objects, their performance and generalization ability are also limited to the design of anchors. Instead, FoveaBox directly learns the object existing possibility and the bounding box coordinates without anchor reference. This is achieved by: (a) predicting category-sensitive semantic maps for the object existing possibility, and (b) producing category-agnostic bounding box for each position that potentially contains an object. The scales of target boxes are naturally associated with feature pyramid representations. In FoveaBox, an instance is assigned to adjacent feature levels to make the model more accurate.We demonstrate its effectiveness on standard benchmarks and report extensive experimental analysis. Without bells and whistles, FoveaBox achieves state-of-the-art single model performance on the standard COCO and Pascal VOC object detection benchmark. More importantly, FoveaBox avoids all computation and hyper-parameters related to anchor boxes, which are often sensitive to the final detection performance. We believe the simple and effective approach will serve as a solid baseline and help ease future research for object detection. The code has been made publicly available athttps://github.com/taokong/FoveaBox. Tao Kong, Fuchun Sun 0001, Huaping Liu 0001, Yuning Jiang 0001, Lei Li 0005, Jianbo Shi |
IEEE Trans. Image Process. | 5 |
| 2019 | CGMH: Constrained Sentence Generation by Metropolis-Hastings SamplingabstractIn real-world applications of natural language generation, there are often constraints on the target sentences in addition to fluency and naturalness requirements. Existing language generation techniques are usually based on recurrent neural networks (RNNs). However, it is non-trivial to impose constraints on RNNs while maintaining generation quality, since RNNs generate sentences sequentially (or with beam search) from the first word to the last. In this paper, we propose CGMH, a novel approach using Metropolis-Hastings sampling for constrained sentence generation. CGMH allows complicated constraints such as the occurrence of multiple keywords in the target sentences, which cannot be handled in traditional RNN-based approaches. Moreover, CGMH works in the inference stage, and does not require parallel corpora for training. We evaluate our method on a variety of tasks, including keywords-to-sentence generation, unsupervised sentence paraphrasing, and unsupervised sentence error correction. CGMH achieves high performance compared with previous supervised methods for sentence generation. Our code is released at https://github.com/NingMiao/CGMH Ning Miao, Hao Zhou 0012, Lili Mou, Rui Yan 0001, Lei Li 0005 |
AAAI | 5 |
| 2019 | Generating Sentences from Disentangled Syntactic and Semantic SpacesabstractVariational auto-encoders (VAEs) are widely used in natural language generation due to the regularization of the latent space.However, generating sentences from the continuous latent space does not explicitly model the syntactic information.In this paper, we propose to generate sentences from disentangled syntactic and semantic spaces.Our proposed method explicitly models syntactic information in the VAE's latent space by using the linearized tree sequence, leading to better performance of language generation.Additionally, the advantage of sampling in the disentangled syntactic and semantic latent spaces enables us to perform novel applications, such as the unsupervised paraphrase generation and syntaxtransfer generation.Experimental results show that our proposed model achieves similar or better performance in various tasks, compared with state-of-the-art related work. Hao Zhou 0012, Shujian Huang, Lei Li 0005, Lili Mou, Olga Vechtomova, Xinyu Dai, Jiajun Chen 0001 |
ACL (1) | 4 |
| 2019 | Dynamically Fused Graph Network for Multi-hop ReasoningabstractText-based question answering (TBQA) has been studied extensively in recent years.Most existing approaches focus on finding the answer to a question within a single paragraph.However, many difficult questions require multiple supporting evidence from scattered text across two or more documents.In this paper, we propose the Dynamically Fused Graph Network (DFGN), a novel method to answer those questions requiring multiple scattered evidence and reasoning over them.Inspired by human's step-by-step reasoning behavior, DFGN includes a dynamic fusion layer that starts from the entities mentioned in the given query, explores along the entity graph dynamically built from the text, and gradually finds relevant supporting entities from the given documents.We evaluate DFGN on HotpotQA, a public TBQA dataset requiring multi-hop reasoning.DFGN achieves competitive results on the public board.Furthermore, our analysis shows DFGN could produce interpretable reasoning chains. Yunxuan Xiao, Yanru Qu, Hao Zhou 0012, Lei Li 0005, Weinan Zhang 0001, Yong Yu 0001 |
ACL (1) | 5 |
| 2019 | Generating Fluent Adversarial Examples for Natural LanguagesabstractEfficiently building an adversarial attacker for natural language processing (NLP) tasks is a real challenge.Firstly, as the sentence space is discrete, it is difficult to make small perturbations along the direction of gradients.Secondly, the fluency of the generated examples cannot be guaranteed.In this paper, we propose MHA, which addresses both problems by performing Metropolis-Hastings sampling, whose proposal is designed with the guidance of gradients.Experiments on IMDB and SNLI show that our proposed MHA outperforms the baseline model on attacking capability.Adversarial training with MHA also leads to better robustness and performance. Huangzhao Zhang, Hao Zhou 0012, Ning Miao, Lei Li 0005 |
ACL (1) | 4 |
| 2019 | What You Look Matters?: Offline Evaluation of Advertising Creatives for Cold-start ProblemabstractModern online auction-based advertising systems combine item and user features to promote ad creatives with the most revenue.However, new ad creatives have to display for certain initial users before enough click statistics could collected and utilized in later ads ranking and bidding processes. This leads to a well-known challenging cold start problem.In this paper, we argue that the content of the creatives intrinsically determines their performance (e.g. ctr, cvr), and we add a pre-ranking stage based on the content. The stage prunes inferior creatives and thus makes online impressions more effective. Since the pre-ranking stage can be executed offline, we can use deep features and take their well generalization to navigate the cold start problem.Specifically, we propose Pre Evaluation Ad Creation Model (PEAC), a novel method to evaluate creatives even before they were shown in the online ads system. Our proposed PEAC only utilizes ads information such as verbal and visual content, but requires no user data as features. During the online A/B testing, PEAC shows significant improvement in revenue. The method has been implemented and deployed in the large scale online advertising system at ByteDance. Furthermore, we provide detailed analysis on what the model learns, which also gives suggestions for ad creative design. Zhichen Zhao, Lei Li 0005, Bowen Zhang 0007, Yuning Jiang 0001, Fengkun Wang, Wei-Ying Ma |
CIKM | 2 |
| 2019 | Unified Visual-Semantic Embeddings: Bridging Vision and Language With Structured Meaning RepresentationsabstractWe propose the Unified Visual-Semantic Embeddings (Unified VSE) for learning a joint space of visual representation and textual semantics. The model unifies the embeddings of concepts at different levels: objects, attributes, relations, and full scenes. We view the sentential semantics as a combination of different semantic components such as objects and relations; their embeddings are aligned with different image regions. A contrastive learning approach is proposed for the effective learning of this fine-grained alignment from only image-caption pairs. We also present a simple yet effective approach that enforces the coverage of caption embeddings on the semantic components that appear in the sentence. We demonstrate that the Unified VSE outperforms baselines on cross-modal retrieval tasks; the enforcement of the semantic coverage improves the model's robustness in defending text-domain adversarial attacks. Moreover, our model empowers the use of visual cues to accurately resolve word dependencies in novel sentences. Hao Wu 0011, Jiayuan Mao, Yuning Jiang 0001, Lei Li 0005, Weiwei Sun 0008, Wei-Ying Ma |
CVPR | 5 |
| 2019 | Towards Linear Time Neural Machine Translation with Capsule NetworksabstractMingxuan Wang, Jun Xie, Zhixing Tan, Jinsong Su, Deyi Xiong, Lei Li. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Mingxuan Wang, Zhixing Tan, Jinsong Su, Deyi Xiong, Lei Li 0005 |
EMNLP/IJCNLP (1) | 6 |
| 2019 | SVD: A Large-Scale Short Video Dataset for Near-Duplicate Video RetrievalabstractWith the explosive growth of video data in real applications, near-duplicate video retrieval (NDVR) has become indispensable and challenging, especially for short videos. However, all existing NDVR datasets are introduced for long videos. Furthermore, most of them are small-scale and lack of diversity due to the high cost of collecting and labeling near-duplicate videos. In this paper, we introduce a large-scale short video dataset, called SVD, for the NDVR task. SVD contains over 500,000 short videos and over 30,000 labeled videos of near-duplicates. We use multiple video mining techniques to construct positive/negative pairs. Furthermore, we design temporal and spatial transformations to mimic user-attack behavior in real applications for constructing more difficult variants of SVD. Experiments show that existing state-of-the-art NDVR methods, including real-value based and hashing based methods, fail to achieve satisfactory performance on this challenging dataset. The release of SVD dataset will foster research and system engineering in the NDVR area. The SVD dataset is available at https://svdbase.github.io. Qing-Yuan Jiang, Lei Li 0005, Wu-Jun Li |
ICCV | 5 |
| 2019 | VaTeX: A Large-Scale, High-Quality Multilingual Dataset for Video-and-Language ResearchabstractWe present a new large-scale multilingual video description dataset, VATEX1, which contains over 41,250 videos and 825, 000 captions in both English and Chinese. Among the captions, there are over 206,000 English-Chinese parallel translation pairs. Compared to the widely-used MSRVTT dataset [64], VATEX is multilingual, larger, linguistically complex, and more diverse in terms of both video and natural language descriptions. We also introduce two tasks for video-and-language research based on VATEX: (1) Multilingual Video Captioning, aimed at describing a video in various languages with a compact unified captioning model, and (2) Video-guided Machine Translation, to translate a source language description into the target language using the video information as additional spatiotemporal context. Extensive experiments on the VATEX dataset show that, first, the unified multilingual model can not only produce both English and Chinese descriptions for a video more efficiently, but also offer improved performance over the monolingual models. Furthermore, we demonstrate that the spatiotemporal video context can be effectively utilized to align source and target languages and thus assist machine translation. In the end, we discuss the potentials of using VATEXfor other video-and-language research. Xin Wang 0061, Jiawei Wu 0003, Jun-Kun Chen, Lei Li 0005, Yuan-Fang Wang, William Yang Wang |
ICCV | 4 |
| 2019 | GraspSnooker: Automatic Chinese Commentary Generation for Snooker VideosabstractWe demonstrate a web-based software system, GraspSnooker, which is able to automatically generate Chinese text commentaries for snooker game videos. It consists of a video analyzer, a strategy predictor and a commentary generator. As far as we know, it is the first attempt on snooker commentary generation, which might be helpful for snooker learners to understand the game. Zhaoyue Sun, Jiaze Chen, Hao Zhou 0012, Lei Li 0005, Mingmin Jiang |
IJCAI | 5 |
| 2019 | Correct-and-Memorize: Learning to Translate from Interactive RevisionsabstractState-of-the-art machine translation models are still not on a par with human translators. Previous work takes human interactions into the neural machine translation process to obtain improved results in target languages. However, not all model--translation errors are equal -- some are critical while others are minor. In the meanwhile, same translation mistakes occur repeatedly in similar context. To solve both issues, we propose CAMIT, a novel method for translating in an interactive environment. Our proposed method works with critical revision instructions, therefore allows human to correct arbitrary words in model-translated sentences. In addition, CAMIT learns from and softly memorizes revision actions based on the context, alleviating the issue of repeating mistakes. Experiments in both ideal and real interactive translation settings demonstrate that our proposed CAMIT enhances machine translation results significantly while requires fewer revision instructions from human compared to previous methods. Rongxiang Weng, Hao Zhou 0012, Shujian Huang, Lei Li 0005, Jiajun Chen 0001 |
IJCAI | 4 |
| 2019 | Rethinking Text Attribute Transfer: A Lexical AnalysisabstractText attribute transfer is modifying certain linguistic attributes (e.g.sentiment, style, authorship, etc.) of a sentence and transforming them from one type to another.In this paper, we aim to analyze and interpret what is changed during the transfer process.We start from the observation that in many existing models and datasets, certain words within a sentence play important roles in determining the sentence attribute class.These words are referred to as the Pivot Words.Based on these pivot words, we propose a lexical analysis framework, the Pivot Analysis, to quantitatively analyze the effects of these words in text attribute classification and transfer.We apply this framework to existing datasets and models, and show that: (1) the pivot words are strong features for the classification of sentence attributes; (2) to change the attribute of a sentence, many datasets only requires to change certain pivot words; (3) consequently, many transfer models only perform the lexical-level modification, while leaving higher-level sentence structures unchanged.Our work provides an in-depth understanding of linguistic attribute transfer and further identifies the future requirements and challenges of this task 1 . Hao Zhou 0012, Jiaze Chen, Lei Li 0005 |
INLG | 4 |
| 2019 | Uncovering the Co-driven Mechanism of Social and Content Links in User Churn PhenomenaabstractRecent years witness the merge of social networks and user-generated content (UGC) platforms. In these new platforms, users establish links to others not only driven by their social relationships in the physical world but also driven by the contents published by others. During this merging process, social networks gradually integrate both social and content links and become unprecedentedly complicated, with the motivation to exploit both the advantages of social viscosity and content attractiveness to reach the best customer retention situation. However, due to the lack of fine-grained data recording such merging phenomena, the co-driven mechanism of social and content links in churn remains unexplored. How do social and content factors jointly influence customers' churn? What is the best ratio of social and content links for retention? Is there a model to capture this co-driven mechanism in churn phenomena? In this paper, we collect a real-world dataset with more than 5.77 million users and 1.15 billion links, with each link being tagged as a social one or a content one. We find that both social and content links have a significant impact on users' churn and they work jointly as a complicated mixture effect. As a result, we propose a novel survival model, which incorporates both social and content factors, to predict churn probability over time. Our model successfully fits the churn distribution in reality and accurately predicts the churn rate of different subpopulations in the future. By analyzing the modeling parameters, we try to strike a balance between social-driven and content-driven links in a user's social network to reach the lowest churn rate. Our model and findings may have potential implications for the design of future social media. Yunfei Lu, Linyun Yu, Peng Cui 0001, Chengxi Zang, Renzhe Xu, Lei Li 0005, Wenwu Zhu 0001 |
KDD | 7 |
| 2019 | Kernelized Bayesian Softmax for Text GenerationabstractNeural models for text generation require a softmax layer with proper token embeddings during the decoding phase. Most existing approaches adopt single point embedding for each token. However, a word may have multiple senses according to different context, some of which might be distinct. In this paper, we propose KerBS, a novel approach for learning better embeddings for text generation. KerBS embodies two advantages: (a) it employs a Bayesian composition of embeddings for words with multiple senses; (b) it is adaptive to semantic variances of words and robust to rare sentence context by imposing learned kernels to capture the closeness of words (senses) in the embedding space. Empirical studies show that KerBS significantly boosts the performance of several text generation tasks. Ning Miao, Hao Zhou 0012, Chengqi Zhao, Wenxian Shi, Lei Li 0005 |
NeurIPS | 5 |
| 2018 | On Tree-Based Neural Sentence ModelingabstractNeural networks with tree-based sentence encoders have shown better results on many downstream tasks.Most of existing tree-based encoders adopt syntactic parsing trees as the explicit structure prior.To study the effectiveness of different tree structures, we replace the parsing trees with trivial trees (i.e., binary balanced tree, left-branching tree and right-branching tree) in the encoders.Though trivial trees contain no syntactic information, those encoders get competitive or even better results on all of the ten downstream tasks we investigated.This surprising result indicates that explicit syntax guidance may not be the main contributor to the superior performances of tree-based neural sentence modeling.Further analysis show that tree modeling gives better results when crucial words are closer to the final representation.Additional experiments give more clues on how to design an effective tree-based encoder.Our code is opensource and available at https://github.com/ExplorerFreda/TreeEnc. Freda Shi, Hao Zhou 0012, Jiaze Chen, Lei Li 0005 |
EMNLP | 4 |
| 2018 | Reinforced Co-TrainingabstractJiawei Wu, Lei Li, William Yang Wang. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). 2018. Jiawei Wu 0003, Lei Li 0005, William Yang Wang |
NAACL-HLT | 2 |
| 2018 | BRITS: Bidirectional Recurrent Imputation for Time SeriesabstractTime series are widely used as signals in many classification/regression tasks. It is ubiquitous that time series contains many missing values. Given multiple correlated time series data, how to fill in missing values and to predict their class labels? Existing imputation methods often impose strong assumptions of the underlying data generating process, such as linear dynamics in the state space. In this paper, we propose BRITS, a novel method based on recurrent neural networks for missing value imputation in time series data. Our proposed method directly learns the missing values in a bidirectional recurrent dynamical system, without any specific assumption. The imputed values are treated as variables of RNN graph and can be effectively updated during the backpropagation. BRITS has three advantages: (a) it can handle multiple correlated missing values in time series; (b) it generalizes to time series with nonlinear dynamics underlying; (c) it provides a data-driven imputation procedure and applies to general settings with missing data. We evaluate our model on three real-world datasets, including an air quality dataset, a health-care data, and a localization data for human activity. Experiments show that our model outperforms the state-of-the-art methods in both imputation and classification/regression accuracies. Wei Cao 0007, Dong Wang 0037, Jian Li 0015, Hao Zhou 0012, Lei Li 0005, Yitan Li |
NeurIPS | 5 |
| 2018 | Overview of the NLPCC 2018 Shared Task: Single Document Summarization
Lei Li 0005, Xiaojun Wan 0001 |
NLPCC (2) | 1 |
| 2017 | A Nearly-Black-Box Online Algorithm for Joint Parameter and State Estimation in Temporal ModelsabstractOnline joint parameter and state estimation is a core problem for temporal models.Most existing methods are either restricted to a particular class of models (e.g., the Storvik filter) or computationally expensive (e.g., particle MCMC). We propose a novel nearly-black-box algorithm, the Assumed Parameter Filter (APF), a hybrid of particle filtering for state variables and assumed density filtering for parameter variables.It has the following advantages:(a) it is online and computationally efficient;(b) it is applicable to both discrete and continuous parameter spaces with arbitrary transition dynamics.On a variety of toy and real models, APF generates more accurate results within a fixed computation budget compared to several standard algorithms from the literature. Yusuf Erol, Yi Wu 0013, Lei Li 0005, Stuart Russell 0001 |
AAAI | 3 |
| 2017 | Overview of the NLPCC 2017 Shared Task: Single Document Summarization
Lifeng Hua, Xiaojun Wan 0001, Lei Li 0005 |
NLPCC | 3 |
| 2017 | Nonlinear Dynamics of Information Diffusion in Social NetworksabstractThe recent explosion in the adoption of search engines and new media such as blogs and Twitter have facilitated the faster propagation of news and rumors. How quickly does a piece of news spread over these media? How does its popularity diminish over time? Does the rising and falling pattern follow a simple universal law? In this article, we propose SpikeM, a concise yet flexible analytical model of the rise and fall patterns of information diffusion. Our model has the following advantages. First, unification power: it explains earlier empirical observations and generalizes theoretical models including the SI and SIR models. We provide the threshold of the take-off versus die-out conditions for SpikeM and discuss the generality of our model by applying it to an arbitrary graph topology. Second, practicality: it matches the observed behavior of diverse sets of real data. Third, parsimony: it requires only a handful of parameters. Fourth, usefulness: it makes it possible to perform analytic tasks such as forecasting, spotting anomalies, and interpretation by reverse engineering the system parameters of interest (quality of news, number of interested bloggers, etc.). We also introduce an efficient and effective algorithm for the real-time monitoring of information diffusion, namely SpikeStream, which identifies multiple diffusion patterns in a large collection of online event streams. Extensive experiments on real datasets demonstrate that SpikeM accurately and succinctly describes all patterns of the rise and fall spikes in social networks. Yasuko Matsubara, Yasushi Sakurai, B. Aditya Prakash, Lei Li 0005, Christos Faloutsos |
ACM Trans. Web | 4 |
| 2016 | CFO: Conditional Focused Neural Question Answering with Large-scale Knowledge BasesabstractHow can we enable computers to automatically answer questions like "Who created the character Harry Potter"? Carefully built knowledge bases provide rich sources of facts.However, it remains a challenge to answer factoid questions raised in natural language due to numerous expressions of one question.In particular, we focus on the most common questions -ones that can be answered with a single fact in the knowledge base.We propose CFO, a Conditional Focused neuralnetwork-based approach to answering factoid questions with knowledge bases.Our approach first zooms in a question to find more probable candidate subject mentions, and infers the final answers with a unified conditional probabilistic framework.Powered by deep recurrent neural networks and neural embeddings, our proposed CFO achieves an accuracy of 75.7% on a dataset of 108k questions -the largest public one to date.It outperforms the current state of the art by an absolute margin of 11.8%. Zihang Dai, Lei Li 0005, Wei Xu 0017 |
ACL (1) | 2 |
| 2016 | Swift: Compiled Inference for Probabilistic Programming Languages
Yi Wu 0013, Lei Li 0005, Stuart Russell 0001, Rastislav Bodík |
IJCAI | 2 |
| 2016 | Concept over time: the combination of probabilistic topic model with wikipedia knowledge
Yin Zhang 0006, Baogang Wei, Lei Li 0005, Fei Wu 0001, Peng Zhang 0075, Yali Bian |
Expert Syst. Appl. | 4 |
| 2014 | Beyond Poisson: Modeling Inter-Arrival Time of Requests in a Datacenter
Da-Cheng Juan, Lei Li 0005, Huan-Kai Peng, Diana Marculescu, Christos Faloutsos |
PAKDD (2) | 2 |
| 2013 | Dynamic Scaled Sampling for Deterministic ConstraintsabstractDeterministic and near-deterministic relationships among subsets of random variables in multivariate systems are known to cause serious problems for Monte Carlo algorithms. We examine the case in which the relationship Z = f(X_1,...,X_k) holds, where each X_i has a continuous prior pdf and we wish to obtain samples from the conditional distribution P(X_1,...,X_k | Z= s). When f is addition, the problem is NP-hard even when the X_i are independent. In more restricted cases — for example, i.i.d. Boolean or categorical X_i — efficient exact samplers have been obtained previously. For the general continuous case, we propose a dynamic scaling algorithm (DYSC), and prove that it has O(k) expected running time and finite variance. We discuss generalizations of DYSC to functions f described by binary operation trees. We evaluate the algorithm on several examples. Lei Li 0005, Bharath Ramsundar, Stuart Russell 0001 |
AISTATS | 1 |
| 2013 | Hibernating Process: Modelling Mobile Calls at Multiple ScalesabstractDo mobile phone calls at larger granularities behave in the same pattern as in smaller ones? How can we forecast the distribution of a whole month's phone calls with only one day's observation? There are many models developed to interpret large scale social graphs. However, all of the existing models focus on graph at one time scale. Many dynamical behaviors were either ignored, or handled at one scale. In particular new users might join or current users quit social networks at any time. In this paper, we propose HiP, a novel model to capture longitudinal behaviors in modeling degree distribution of evolving social graphs. We analyze a large scale phone call dataset using HiP, and compare with several previous models in literature. Our model is able to fit phone call distribution at multiple scales with 30% to 75% improvement over the best existing method on each scale. Siyuan Liu 0001, Lei Li 0005, Rammaya Krishnan |
ICDM | 2 |
| 2013 | The Extended Parameter FilterabstractThe parameters of temporal models, such as dynamic Bayesian networks, may be modelled in a Bayesian context as static or atemporal variables that influence transition probabilities at every time step. Particle filters fail for models that include such variables, while methods that use Gibbs sampling of parameter variables may incur a per-sample cost that grows linearly with the length of the observation sequence. Storvik devised a method for incremental computation of exact sufficient statistics that, for some cases, reduces the per-sample cost to a constant. In this paper, we demonstrate a connection between Storvik’s filter and a Kalman filter in parameter space and establish more general conditions under which Storvik’s filter works. Drawing on an analogy to the extended Kalman filter, we develop and analyze, both theoretically and experimentally, a Taylor approximation to the parameter posterior that allows Storvik’s method to be applied to a broader class of models. Our experiments on both synthetic examples and real applications show improvement over existing methods. Yusuf Erol, Lei Li 0005, Bharath Ramsundar, Stuart Russell 0001 |
ICML (3) | 2 |
| 2013 | Why people hate your app: making sense of user feedback in a mobile app storeabstractUser review is a crucial component of open mobile app markets such as the Google Play Store. How do we automatically summarize millions of user reviews and make sense out of them? Unfortunately, beyond simple summaries such as histograms of user ratings, there are few analytic tools that can provide insights into user reviews. In this paper, we propose Wiscom, a system that can analyze tens of millions user ratings and comments in mobile app markets at three different levels of detail. Our system is able to (a) discover inconsistencies in reviews; (b) identify reasons why users like or dislike a given app, and provide an interactive, zoomable view of how users' reviews evolve over time; and (c) provide valuable insights into the entire app market, identifying users' major concerns and preferences of different types of apps. Results using our techniques are reported on a 32GB dataset consisting of over 13 million user reviews of 171,493 Android apps in the Google Play Store. We discuss how the techniques presented herein can be deployed to help a mobile app market operator such as Google as well as individual app developers and end-users. Jialiu Lin, Lei Li 0005, Christos Faloutsos, Jason I. Hong, Norman M. Sadeh |
KDD | 3 |
| 2013 | Multilinear Dynamical Systems for Tensor Time SeriesabstractMany scientific data occur as sequences of multidimensional arrays called tensors. How can hidden, evolving trends in such data be extracted while preserving the tensor structure? The model that is traditionally used is the linear dynamical system (LDS), which treats the observation at each time slice as a vector. In this paper, we propose the multilinear dynamical system (MLDS) for modeling tensor time series and an expectation-maximization (EM) algorithm to estimate the parameters. The MLDS models each time slice of the tensor time series as the multilinear projection of a corresponding member of a sequence of latent, low-dimensional tensors. Compared to the LDS with an equal number of parameters, the MLDS achieves higher prediction accuracy and marginal likelihood for both simulated and real datasets. Mark Rogers, Lei Li 0005, Stuart Russell 0001 |
NIPS | 2 |
| 2013 | F-Trail: Finding Patterns in Taxi Trajectories
Yasuko Matsubara, Lei Li 0005, Evangelos E. Papalexakis, David Lo 0001, Yasushi Sakurai, Christos Faloutsos |
PAKDD (1) | 2 |
| 2012 | RolX: structural role extraction & mining in large graphsabstractGiven a network, intuitively two nodes belong to the same role if they have similar structural behavior. Roles should be automatically determined from the data, and could be, for example, "clique-members," "periphery-nodes," etc. Roles enable numerous novel and useful network-mining tasks, such as sense-making, searching for similar nodes, and node classification. This paper addresses the question: Given a graph, how can we automatically discover roles for nodes? We propose RolX (Role eXtraction), a scalable (linear in the number of edges), unsupervised learning approach for automatically extracting structural roles from general network data. We demonstrate the effectiveness of RolX on several network-mining tasks: from exploratory data analysis to network transfer learning. Moreover, we compare network role discovery with network community discovery. We highlight fundamental differences between the two (e.g., roles generalize across disconnected networks, communities do not); and show that the two approaches are complimentary in nature. Keith Henderson, Brian Gallagher, Tina Eliassi-Rad, Hanghang Tong, Sugato Basu, Leman Akoglu, Danai Koutra, Christos Faloutsos, Lei Li 0005 |
KDD | 9 |
| 2012 | Rise and fall patterns of information diffusion: model and implicationsabstractThe recent explosion in the adoption of search engines and new media such as blogs and Twitter have facilitated faster propagation of news and rumors. How quickly does a piece of news spread over these media? How does its popularity diminish over time? Does the rising and falling pattern follow a simple universal law? Yasuko Matsubara, Yasushi Sakurai, B. Aditya Prakash, Lei Li 0005, Christos Faloutsos |
KDD | 4 |
| 2012 | A Novel Violent Videos Classification Scheme Based on the Bag of Audio Words FeaturesabstractThis paper has been removed because it was brought to our attention that Lei Li from CMU is not the author of this paper. Lei Li from CMU informed that the e-mail used in this paper is not the e-mail of Lei Li from CMU. Please contact [email protected] Lei Li 0005 |
Int. J. Comput. Intell. Appl. | 1 |
| 2011 | Time Series Clustering: Complex is Simpler!
Lei Li 0005, B. Aditya Prakash |
ICML | 1 |
| 2011 | It's who you know: graph mining using recursive structural featuresabstractGiven a graph, how can we extract good features for the nodes? For example, given two large graphs from the same domain, how can we use information in one to do classification in the other (i.e., perform across-network classification or transfer learning on graphs)? Also, if one of the graphs is anonymized, how can we use information in one to de-anonymize the other? The key step in all such graph mining tasks is to find effective node features. We propose ReFeX (Recursive Feature eXtraction), a novel algorithm, that recursively combines local (node-based) features with neighborhood (egonet-based) features; and outputs regional features -- capturing "behavioral" information. We demonstrate how these powerful regional features can be used in within-network and across-network classification and de-anonymization tasks -- without relying on homophily, or the availability of class labels. The contributions of our work are as follows: (a) ReFeX is scalable and (b) it is effective, capturing regional ("behavioral") information in large graphs. We report experiments on real graphs from various domains with over 1M edges, where ReFeX outperforms its competitors on typical graph mining tasks like network classification and de-anonymization. Keith Henderson, Brian Gallagher, Lei Li 0005, Leman Akoglu, Tina Eliassi-Rad, Hanghang Tong, Christos Faloutsos |
KDD | 3 |
| 2011 | ThermoCast: a cyber-physical forecasting model for datacentersabstractEfficient thermal management is important in modern data centers as cooling consumes up to 50% of the total energy. Unlike previous work, we consider proactive thermal management, whereby servers can predict potential overheating events due to dynamics in data center configuration and workload, giving operators enough time to react. However, such forecasting is very challenging due to data center scales and complexity. Moreover, such a physical system is influenced by cyber effects, including workload scheduling in servers. We propose ThermoCast, a novel thermal forecasting model to predict the temperatures surrounding the servers in a data center, based on continuous streams of temperature and airflow measurements. Our approach is (a) capable of capturing cyberphysical interactions and automatically learning them from data; (b) computationally and physically scalable to data center scales; (c) able to provide online prediction with real-time sensor measurements. The paper's main contributions are: (i) We provide a systematic approach to integrate physical laws and sensor observations in a data center; (ii) We provide an algorithm that uses sensor data to learn the parameters of a data center's cyber-physical system. In turn, this ability enables us to reduce model complexity compared to full-fledged fluid dynamics models, while maintaining forecast accuracy; (iii) Unlike previous simulation-based studies, we perform experiments in a production data center. Using real data traces, we show that ThermoCast forecasts temperature better than a machine learning approach solely driven by data, and can successfully predict thermal alarms 4.2 minutes ahead of time. Lei Li 0005, Chieh-Jan Mike Liang, Jie Liu 0001, Suman Nath, Andreas Terzis, Christos Faloutsos |
KDD | 1 |
| 2011 | WindMine: Fast and Effective Mining of Web-click SequencesabstractGiven a large stream of users clicking on web sites, how can we find trends, patterns and anomalies? We have developed a novel method, WindMine, and its fine-tuning sibling, WindMine-part, to find patterns and anomalies in such datasets. Our approach has the following advantages: (a) it is effective in discovering meaningful “building blocks” and patterns such as the lunch-break trend and anomalies, (b) it automatically determines suitable window sizes, and (c) it is fast, with its wall clock time linear on the duration of sequences. Moreover, it can be made sub-quadratic on the number of sequences (WindMine-part), with little loss of accuracy. We examine the effectiveness and scalability by performing experiments on 67 GB of real data (one billion clicks for 30 days). Our proposed WindMine does produce concise, informative and interesting patterns. We also show that WindMine-part can be easily implemented in a parallel or distributed setting, and that, even in a single-machine setting, it can be an order of magnitude faster (up to 70 times) than the plain version. Yasushi Sakurai, Lei Li 0005, Yasuko Matsubara, Christos Faloutsos |
SDM | 2 |
| 2010 | Metric forensics: a multi-level approach for mining volatile graphsabstractAdvances in data collection and storage capacity have made it increasingly possible to collect highly volatile graph data for analysis. Existing graph analysis techniques are not appropriate for such data, especially in cases where streaming or near-real-time results are required. An example that has drawn significant research interest is the cyber-security domain, where internet communication traces are collected and real-time discovery of events, behaviors, patterns, and anomalies is desired. We propose MetricForensics, a scalable framework for analysis of volatile graphs. MetricForensics combines a multi-level "drill down" approach, a collection of user-selected graph metrics, and a collection of analysis techniques. At each successive level, more sophisticated metrics are computed and the graph is viewed at finer temporal resolutions. In this way, MetricForensics scales to highly volatile graphs by only allocating resources for computationally expensive analysis when an interesting event is discovered at a coarser resolution first. We test MetricForensics on three real-world graphs: an enterprise IP trace, a trace of legitimate and malicious network traffic from a research institution, and the MIT Reality Mining proximity sensor data. Our largest graph has 3M vertices and 32M edges, spanning 4.5 days. The results demonstrate the scalability and capability of MetricForensics in analyzing volatile graphs; and highlight four novel phenomena in such graphs: elbows, broken correlations, prolonged spikes, and lightweight stars. Keith Henderson, Tina Eliassi-Rad, Christos Faloutsos, Leman Akoglu, Lei Li 0005, Koji Maruhashi, B. Aditya Prakash, Hanghang Tong |
KDD | 5 |
| 2010 | Parsimonious Linear Fingerprinting for Time SeriesabstractWe study the problem of mining and summarizing multiple time series effectively and efficiently. We propose PLiF, a novel method to discover essential characteristics ("fingerprints"), by exploiting the joint dynamics in numerical sequences. Our fingerprinting method has the following benefits: (a) it leads to interpretable features; (b) it is versatile: PLiF enables numerous mining tasks, including clustering, compression, visualization, forecasting, and segmentation, matching top competitors in each task; and (c) it is fast and scalable , with linear complexity on the length of the sequences. We did experiments on both synthetic and real datasets, including human motion capture data (17MB of human motions), sensor data (166 sensors), and network router traffic data (18 million raw updates over 2 years). Despite its generality, PLiF outperforms the top clustering methods on clustering; the top compression methods on compression (3 times better reconstruction error, for the same compression ratio); it gives meaningful visualization and at the same time, enjoys a linear scale-up. Lei Li 0005, B. Aditya Prakash, Christos Faloutsos |
Proc. VLDB Endow. | 1 |
| 2009 | DynaMMo: mining and summarization of coevolving sequences with missing valuesabstractGiven multiple time sequences with missing values, we propose DynaMMo which summarizes, compresses, and finds latent variables. The idea is to discover hidden variables and learn their dynamics, making our algorithm able to function even when there are missing values.We performed experiments on both real and synthetic datasets spanning several megabytes, including motion capture sequences and chlorine levels in drinking water. We show that our proposed DynaMMo method (a) can successfully learn the latent variables and their evolution; (b) can provide high compression for little loss of reconstruction accuracy; (c) can extract compact but powerful features for segmentation, interpretation, and forecasting; (d) has complexity linear on the duration of sequences. Lei Li 0005, James McCann, Nancy S. Pollard, Christos Faloutsos |
KDD | 1 |
| 2008 | GMIP: A Novel Optical Interconnect Gridded Memory Service ProtocolabstractDynamic self-organized computer architecture (DSAG) based on Grid-components departs computer components to grid components, and dynamically aggregates and organizes these components to realize architecture-on-demand. We take the important feature, that CPU centered design principle should possibly become memory centered design and optimization. We propose a novel computer architecture Gridded Memory Service (GMS), which is based on DSAG. We design and implement a serial, light weight and packet switching optical interconnect protocol, Gridded Memory Interconnect Protocol (GMIP), which is featured as high bandwidth and low latency protocol. It optimizes the link and physical layer to take advantage of very short reach optical interconnect technology. At last we study interconnect effects, and propose the main evaluation principles of bandwidth latency compensation. Siyuan Liu 0001, Lei Li 0005, Jianping Fan 0002 |
ICPADS | 2 |
| 2008 | Cut-and-stitch: efficient parallel learning of linear dynamical systems on smpsabstractMulti-core processors with ever increasing number of cores per chip are becoming prevalent in modern parallel computing. Our goal is to make use of the multi-core as well as multi-processor architectures to speed up data mining algorithms. Specifically, we present a parallel algorithm for approximate learning of Linear Dynamical Systems (LDS), also known as Kalman Filters (KF). LDSs are widely used in time series analysis such as motion capture modeling, visual tracking etc. We propose Cut-And-Stitch (CAS), a novel method to handle the data dependencies from the chain structure of hidden variables in LDS, so as to parallelize the EM-based parameter learning algorithm. We implement the algorithm using OpenMP on both a supercomputer and a quad-core commercial desktop. The experimental results show that parallel algorithms using Cut-And-Stitch achieve comparable accuracy and almost linear speedups over the serial version. In addition, Cut-And-Stitch can be generalized to other models with similar linear structures such as Hidden Markov Models (HMM) and Switching Kalman Filters (SKF). Lei Li 0005, Wenjie Fu 0002, Fan Guo 0006, Todd C. Mowry, Christos Faloutsos |
KDD | 1 |
| 2008 | Efficient Distribution Mining and ClassificationabstractWe define and solve the problem of “distribution classification”, and, in general, “distribution mining”. Given n distributions (i.e., clouds) of multi-dimensional points, we want to classify them into k classes, to find patterns, rules and out-lier clouds. For example, consider the 2-d case of sales of items, where, for each item sold, we record the unit price and quantity; then, each customer is represented as a distribution/cloud of 2-d points (one for each item he bought). We want to group similar users together, e.g., for market segmentation, anomaly/fraud detection. We propose D-Mine to achieve this goal. Our main contribution is Theorem 3.1, which shows how to use wavelets to speed up the cloud-similarity computations. Extensive experiments on both synthetic and real multi-dimensional data sets show that our method achieves up to 400 faster wall-clock time over the naive implementation, with comparable (and occasionally better) classification quality. Yasushi Sakurai, Rosalynn Chong, Lei Li 0005, Christos Faloutsos |
SDM | 3 |
| 2008 | C-DEM: a multi-modal query system for Drosophila Embryo databasesabstractThe amount of biological data publicly available has experienced an exponential growth as the technology advances. Online databases are now playing an important role as information repositories as well as easily accessible platforms for researchers to communicate and contribute. Recent research projects in image bioinformatics produce a number of databases of images, which visualize the spatial expression pattern of a gene (eg. "fj"), and most of which also have one or several annotation keywords (eg., "embryonic hindgut"). C-DEM is an online system for Drosophila (= fruit-fly) Embryo images Mining. It supports queries from all three modalities to all three, namely, (a) genes, (b) images of gene expression, and (c) annotation keywords of the images. Thus, it can find images that are similar to a given image, and/or related to the desirable annotation keywords, and/or related to specific genes. Typical queries are what are most suitable keywords to assign to image insitu28465.jpg or find images that are related to gene "fj", and to the keyword "embryonic hindgut" . C-DEM uses state-of-the-art feature extraction methods for images (wavelets and principal component analysis). It envisions the whole database as a tri-partite graph (one type for each modality), and it uses fast and flexible proximity measures, namely, random walk with restarts (RWR). In addition to flexible querying, C-DEM allows for navigation: the user can click on the results of an earlier query (image thumbnails and/or keywords and/or genes), and the system will report the most related images (and keywords, and genes). The demo is on a real Drosophila Embryo database, with 10,204 images, 2,969 distinct genes, and 113 annotation keywords. The query response time is below one second on a commodity desktop. Fan Guo 0006, Lei Li 0005, Christos Faloutsos, Eric P. Xing |
Proc. VLDB Endow. | 2 |
| 2006 | Providing an Uncertainty Reasoning Service for Semantic Web Application
Lei Li 0005, Qiaoling Liu, Yunfeng Tao, Lei Zhang 0007, Yong Yu 0001 |
APWeb | 1 |
| 2005 | A Reconfigurable Optical Interconnect System for DSAGabstractHigh performance computing research is facing challenges and innovation on architecture is urgent. DSAG architecture is proposed and delivers "Architecture on Demand" feature. In DSAG, components in different catalogs are parted, while the ones in same catalog are congregated. This architecture can be enabled by optical interconnect and reconfigurable computing technology. Using advanced optical devices and enhanced reconfigurable computing devices (FPGA), we build a prototype system for DSAG. Optical interconnect can reach 16Gbps bandwidth; DDRAM interface is selected as host communication interface to match the bandwidth of optical channel; Reconfigurable logic and embedded processors are employed for flexible reconfiguration. The system is featured by high bandwidth, owerful, flexible. Lei Li 0005, Zheng Cao 0003, Mingyu Chen 0001, Jianping Fan 0002 |
PDCAT | 1 |