VLDB 2026 Research / reviewers in the wild / expert
Yuan Yao 0013
dblp:25/4120-13
· DBLP profile ↗
39ranked-venue papers
7as first author
25since 2021 · last 2026
0000-0002-8276-3620ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 33 · 6 first-author · 22 since 2021Graphics, computer vision, multimedia, augmented reality and games · 16 · 3 first-author · 14 since 2021Databases, data management, data science and information retrieval · 3 · 1 since 2021Computer networks · 2Systems, architecture and hardware · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | LLaVA-UHD v2: Exploiting Hierarchical Vision Granularity in MLLMs via Inverse Semantic PyramidabstractVision transformers (ViTs) are widely employed in multimodal large language models (MLLMs) for visual encoding. However, they exhibit inferior performance on tasks regarding fine-grained visual perception. We attribute this to the inner limitations of ViTs in capturing diverse visual semantic levels. To address this, we present Hierarchical window (Hiwin) transformer as a plug-and-play solution for MLLMs, centered around our inverse semantic pyramid (ISP). Hiwin transformer comprises two key modules: (i) a visual detail injection module, which progressively injects low-level visual details into high-level language-aligned semantics features, thereby constructing an ISP, and (ii) a hierarchical window attention module, which leverages cross-scale windows to condense multi-level semantics from the ISP. Notably, our design achieves an average boost of 3.7% across 14 benchmarks compared with the baseline method, 9.3% on DocVQA for instance. Zonghao Guo, Xuesong Yang, Chi Chen 0005, Yuan Yao 0013, Tat-Seng Chua, Maosong Sun 0001 |
AAAI | 9 |
| 2025 | GUICourse: From General Vision Language Model to Versatile GUI AgentabstractUtilizing Graphic User Interfaces (GUIs) for human-computer interaction is essential for accessing various digital tools. Recent advancements in Vision Language Models (VLMs) reveal significant potential for developing versatile agents that assist humans in navigating GUIs. However, current VLMs face challenges related to fundamental abilities, such as OCR and grounding, as well as a lack of knowledge about GUI elements functionalities and control methods. These limitations hinder their effectiveness as practical GUI agents. To address these challenges, we introduce GUICourse, a series of datasets for training visual-based GUI agents using general VLMs. First, we enhance the OCR and grounding capabilities of VLMs using the GUIEnv dataset. Next, we enrich the GUI knowledge of VLMs using the GUIAct and GUIChat datasets. Our experiments demonstrate that even a small-sized GUI agent (with 3.1 billion parameters) performs effectively on both single-step and multi-step GUI tasks. We further finetune our GUI agents on other GUI tasks with different action spaces (AITW and Mind2Web), and the results show that our agents are better than their baseline VLMs. Additionally, we analyze the impact of OCR and grounding capabilities through an ablation study, revealing a positive correlation with GUI navigation ability. Wentong Chen, Junbo Cui, Jinyi Hu, Yujia Qin, Junjie Fang, Chongyi Wang, Guirong Chen, Yupeng Huo, Yuan Yao 0013, Yankai Lin 0001, Zhiyuan Liu 0001, Maosong Sun 0001 |
ACL (1) | 11 |
| 2025 | RLAIF-V: Open-Source AI Feedback Leads to Super GPT-4V TrustworthinessabstractTraditional feedback learning for hallucination reduction relies on labor-intensive manual labeling or expensive proprietary models. This leaves the community without foundational knowledge about how to build high-quality feedback with open-source MLLMs. In this work, we introduce RLAIF-V, a novel framework that aligns MLLMs in a fully open-source paradigm. RLAIF-V maximally explores open-source MLLMs from two perspectives, including high-quality feedback data generation for preference learning and self-feedback guidance for inference-time scaling. Extensive experiments on six benchmarks in both automatic and human evaluation show that RLAIF-V substantially enhances the trustworthiness of models at both preference learning and inference time. RLAIF-V 7B reduces object hallucination by 80.7% and overall hallucination by 33.7%. Remarkably, RLAIF-V 12B further reveals the self-alignment potential of open-source MLLMs, where the model can learn from feedback of itself to achieve super GPT-4V trustworthiness. Tianyu Yu 0002, Haoye Zhang, Qixin Xu, Yuan Yao 0013, Xiaoman Lu, Ganqu Cui, Yunkai Dang, Taiwen He, Bo Zheng 0007, Zhiyuan Liu 0001, Tat-Seng Chua, Maosong Sun 0001 |
CVPR | 5 |
| 2025 | Thoroughly Modeling Multi-domain Pre-trained Recommendation as LanguageabstractWith the thriving of the pre-trained language model (PLM) widely verified in various NLP tasks, pioneer efforts attempt to explore the possible cooperation of the general textual information in PLM with the personalized behavioral information in user historical behavior sequences to enhance sequential recommendation (SR). However, despite the commonalities of input format and task goal, there are huge gaps between the behavioral and textual information, which obstruct thoroughly modeling SR as language modeling via PLM. To bridge the gap, we propose a novel unified pre-trained language model enhanced sequential recommendation (UPSR) that thoroughly transfers the next item prediction task to a text generation task, aiming to build a unified pre-trained recommendation model for multi-domain recommendation tasks. We formally design five key indicators, namely naturalness, domain consistency, informativeness, noise and ambiguity, and text length, to guide the text \(\rightarrow\) item adaptation (selecting appropriate text to form the item textual representation) and behavior sequence \(\rightarrow\) text sequence adaptation (transferring the sequence of item textual representations into a text sequence) differently for pre-training and fine-tuning stages, which are essential but under-explored by previous works. In experiments, we conduct extensive evaluations on seven datasets with both supervised and zero-shot settings and achieve the overall best performance. Comprehensive model analyses also provide valuable insights for behavior modeling via PLM, shedding light on large pre-trained recommendation models. The source codes will be released in the future. Zekai Qu, Ruobing Xie, Chaojun Xiao, Yuan Yao 0013, Zhiyuan Liu 0001, Fengzong Lian, Zhanhui Kang, Jie Zhou 0016 |
ACM Trans. Inf. Syst. | 4 |
| 2024 | Fine-Grained Legal Argument-Pair Extraction via Coarse-Grained Pre-trainingabstractLegal Argument-Pair Extraction (LAE) is dedicated to the identification of interactive arguments targeting the same subject matter within legal complaints and corresponding defenses. This process serves as a foundation for automatically recognizing the focal points of disputes. Current methodologies predominantly conceptualize LAE as a supervised sentence-pair classification problem and usually necessitate extensive manual annotations, thereby constraining their scalability and general applicability. To this end, we present an innovative approach to LAE that focuses on fine-grained alignment of argument pairs, building upon coarse-grained complaint-defense pairs. This strategy stems from two key observations: 1) In general, every argument presented in a legal complaint is likely to be addressed by at least one corresponding argument in the defense. 2) It’s rare for multiple complaint arguments to be addressed by a single defense argument; rather, each complaint argument usually corresponds to a unique defense argument. Motivated by these insights, we develop a specialized pre-training framework. Our model employs pre-training objectives designed to exploit the coarse-grained supervision signals. This enables expressive representations of legal arguments for LAE, even when working with a limited amount of labeled data. To verify the effectiveness of our model, we construct the largest LAE datasets from two representative causes, private lending, and contract dispute. The experimental results demonstrate that our model can effectively capture informative argument knowledge from unlabeled complaint-defense pairs and outperform the unsupervised and supervised baselines by 3.7 and 2.4 points on average respectively. Besides, our model can reach superior accuracy with only half manually annotated data. The datasets and code can be found in https://github.com/thunlp/LAE. Chaojun Xiao, Yutao Sun, Yuan Yao 0013, Zhiyuan Liu 0001, Maosong Sun 0001 |
LREC/COLING | 3 |
| 2024 | En3D: An Enhanced Generative Model for Sculpting 3D Humans from 2D Synthetic DataabstractWe present En3D, an enhanced generative scheme for sculpting high-quality 3D human avatars. Unlike previous works that rely on scarce 3D datasets or limited 2D collections with imbalanced viewing angles and imprecise pose priors, our approach aims to develop a zero-shot 3D generative scheme capable of producing visually realistic, ge-ometrically accurate and content-wise diverse 3D humans without directly relying on pre-existing 3D or 2D assets. To address this challenge, we introduce a meticulously crafted workflow that implements accurate physical modeling to learn the enhanced 3D generative model from synthetic 2D data. During inference, we integrate optimization modules to bridge the gap between realistic appearances and coarse 3D shapes. Specifically, En3D comprises three modules: a 3D generator that accurately models generalizable 3D humans with realistic appearance from synthesized balanced, diverse, and structured human images; a geometry sculptor that enhances shape quality using multi-view normal constraints for intricate human structure; and a texturing module that disentangles explicit texture maps with fidelity and editability, leveraging semantical UV partitioning and a differentiable rasterizer: Experimental results show that our approach significantly outperforms prior works in terms of image quality, geometry accuracy and content diversity. We also showcase the applicability of our generated avatars for animation and editing, as well as the scalability of our approach for content-style free adaptation. Yifang Men, Biwen Lei, Yuan Yao 0013, Miaomiao Cui, Zhouhui Lian, Xuansong Xie |
CVPR | 3 |
| 2024 | Revisiting Non-Autoregressive Transformers for Efficient Image SynthesisabstractThe field of image synthesis is currently flourishing due to the advancements in diffusion models. While diffusion models have been successful, their computational inten-sity has prompted the pursuit of more efficient alternatives. As a representative work, non-autoregressive Transformers (NATs) have been recognized for their rapid generation. However, a major drawback of these models is their in-ferior performance compared to diffusion models. In this paper, we aim to re-evaluate the full potential of NATs by revisiting the design of their training and inference strategies. Specifically, we identify the complexities in properly configuring these strategies and indicate the possible sub-optimality in existing heuristic-driven designs. Recognizing this, we propose to go beyond existing methods by directly solving the optimal strategies in an automatic framework. The resulting method, named AutoNAT, advances the performance boundaries of NATs notably, and is able to perform comparably with the latest diffusion models with a significantly reduced inference cost. The effectiveness of AutoNAT is comprehensively validated on four benchmark datasets, i.e., ImageNet-256 & 512, MS-COCO, and CC3M. Code and pretrained models will be available at htt P s: / /gi thub. com/LeapLabTHU/ImprovedNAT. Zanlin Ni, Yulin Wang 0002, Renping Zhou, Jinyi Hu, Zhiyuan Liu 0001, Shiji Song, Yuan Yao 0013, Gao Huang 0001 |
CVPR | 8 |
| 2024 | RLHF-V: Towards Trustworthy MLLMs via Behavior Alignment from Fine-Grained Correctional Human FeedbackabstractMultimodal Large Language Models (MLLMs) have recently demonstrated impressive capabilities in multimodal understanding, reasoning, and interaction. However, existing MLLMs prevalently suffer from serious hallucination problems, generating text that is not factually grounded in associated images. The problem makes existing MLLMs untrustworthy and thus impractical in real-world (especially high-stakes) applications. To address the challenge, we present RLHF-V, which enhances MLLM trustworthiness via behavior alignment from fine-grained correctional human feedback. Specifically, RLHF-V collects human preference in the form of segment-level corrections on hallucinations, and performs dense direct preference optimization over the human feedback. Comprehensive experiments on five benchmarks in both automatic and human evaluation show that, RLHF-V can enable substantially more trustworthy MLLM behaviors with promising data and computation efficiency. Remarkably, using 1.4k annotated data samples, RLHF-V significantly reduces the hallucination rate of the base MLLM by 34.8%, outperforming the concurrent LLaVA-RLHF trained on 10k annotated data. The final model achieves state-of-the-art performance in trustwor-thiness among open-source MLLMs, and shows better ro-bustness than GPT-4V in preventing hallucinations aroused from over-generalization. Tianyu Yu 0002, Yuan Yao 0013, Haoye Zhang, Taiwen He, Yifeng Han, Ganqu Cui, Jinyi Hu, Zhiyuan Liu 0001, Hai-Tao Zheng 0002, Maosong Sun 0001 |
CVPR | 2 |
| 2024 | LLaVA-UHD: An LMM Perceiving Any Aspect Ratio and High-Resolution Images
Zonghao Guo, Ruyi Xu, Yuan Yao 0013, Junbo Cui, Zanlin Ni, Chunjiang Ge, Tat-Seng Chua, Zhiyuan Liu 0001, Gao Huang 0001 |
ECCV (83) | 3 |
| 2024 | AdaNAT: Exploring Adaptive Policy for Token-Based Image Generation
Zanlin Ni, Yulin Wang 0002, Renping Zhou, Rui Lu 0001, Jinyi Hu, Zhiyuan Liu 0001, Yuan Yao 0013, Gao Huang 0001 |
ECCV (16) | 8 |
| 2024 | Large Multilingual Models Pivot Zero-Shot Multimodal Learning across LanguagesabstractRecently there has been a significant surge in multimodal learning in terms of both image-to-text and text-to-image generation. However, the success is typically limited to English, leaving other languages largely behind. Building a competitive counterpart in other languages is highly challenging due to the low-resource nature of non-English multimodal data (i.e., lack of large-scale, high-quality image-text data). In this work, we propose MPM, an effective training paradigm for training large multimodal models in low-resource languages. MPM demonstrates that Multilingual language models can Pivot zero-shot Multimodal learning across languages. Specifically, based on a strong multilingual large language model, multimodal models pretrained on English-only image-text data can well generalize to other languages in a (quasi)-zero-shot manner, even surpassing models trained on image-text data in native languages. Taking Chinese as a practice of MPM, we build large multimodal models VisCPM in image-to-text and text-to-image generation, which achieve state-of-the-art (open-source) performance in Chinese. To facilitate future research, we open-source codes and model weights at https://github.com/OpenBMB/VisCPM. Jinyi Hu, Yuan Yao 0013, Chongyi Wang, Shan Wang 0015, Yinxu Pan, Tianyu Yu 0002, Hanghao Wu, Haoye Zhang, Xu Han 0007, Yankai Lin 0001, Jiao Xue, Dahai Li, Zhiyuan Liu 0001, Maosong Sun 0001 |
ICLR | 2 |
| 2024 | NExT-Chat: An LMM for Chat, Detection and SegmentationabstractThe development of large language models (LLMs) has greatly advanced the field of multimodal understanding, leading to the emergence of large multimodal models (LMMs). In order to enhance visual comprehension, recent studies have equipped LMMs with region-level understanding capabilities by representing object bounding box coordinates as a series of text sequences (pix2seq). In this paper, we introduce a novel paradigm for object location modeling called the pix2emb method, where we ask the LMM to output the location embeddings and then decode them with different decoders. This paradigm allows us to use different location formats (such as bounding boxes and masks) in multimodal conversations. Leveraging the proposed pix2emb method, we train an LMM named NExT-Chat and demonstrate its capability of handling multiple tasks like visual grounding, region captioning, and grounded reasoning. Comprehensive experiments show the effectiveness of our NExT-Chat on various tasks, e.g., NExT-Chat (87.7) vs. Shikra (86.9) on POPE-Random, NExT-Chat (71.3) vs. LISA (67.9) on referring expression segmentation task, and NExT-Chat (79.6) vs. Kosmos-2 (62.3) on region caption task. Yuan Yao 0013, Wei Ji 0008, Zhiyuan Liu 0001, Tat-Seng Chua |
ICML | 2 |
| 2024 | Fact : Teaching MLLMs with Faithful, Concise and Transferable RationalesabstractThe remarkable performance of Multimodal Large Language Models (MLLMs) has demonstrated their proficient understanding capabilities in handling various visual tasks. Nevertheless, the opaque nature of black-box reasoning processes persists as an enigma, rendering them uninterpretable and struggling with hallucination. Their ability to execute intricate reasoning tasks is also constrained, culminating in stagnation of progression. In this work, we introduce Fact, a novel paradigm designed to generate multimodal rationales that are faithful, concise, and transferable for teaching MLLMs. This paradigm utilizes verifiable visual programming to generate executable code guaranteeing faithfulness. Through a series of operations including pruning, merging, and bridging, the rationale enhances its conciseness. Furthermore, we filter rationales that can be transferred to end-to-end paradigms from programming paradigms to guarantee transferability. Empirical evidence from experiments demonstrates the superiority of Fact across models of varying parameter sizes, significantly enhancing their compositional reasoning and generalization ability and reducing hallucinations owing to its high correlation between images and text. Minghe Gao, Liang Pang 0001, Yuan Yao 0013, Jisheng Dang, Wenqiao Zhang, Juncheng Li 0006, Siliang Tang, Yueting Zhuang, Tat-Seng Chua |
ACM Multimedia | 4 |
| 2024 | ENAT: Rethinking Spatial-temporal Interactions in Token-based Image SynthesisabstractRecently, token-based generation approaches have demonstrated their effectiveness in synthesizing visual content. As a representative example, non-autoregressive Transformers (NATs) can generate decent-quality images in just a few steps. NATs perform generation in a progressive manner, where the latent tokens of a resulting image are incrementally revealed step-by-step. At each step, the unrevealed image regions are padded with [MASK] tokens and inferred by NAT, with the most reliable predictions preserved as newly revealed, visible tokens. In this paper, we delve into understanding the mechanisms behind the effectiveness of NATs and uncover two important interaction patterns that naturally emerge from NAT’s paradigm: Spatially (within a step), although [MASK] and visible tokens are processed uniformly by NATs, the interactions between them are highly asymmetric. In specific, [MASK] tokens mainly gather information for decoding. On the contrary, visible tokens tend to primarily provide information, and their deep representations can be built only upon themselves. Temporally (across steps), the interactions between adjacent generation steps mostly concentrate on updating the representations of a few critical tokens, while the computation for the majority of tokens is generally repetitive. Driven by these findings, we propose EfficientNAT (ENAT), a NAT model that explicitly encourages these critical interactions inherent in NATs. At the spatial level, we disentangle the computations of visible and [MASK] tokens by encoding visible tokens independently, while decoding [MASK] tokens conditioned on the fully encoded visible tokens. At the temporal level, we prioritize the computation of the critical tokens at each step, while maximally reusing previously computed token representations to supplement necessary information. ENAT improves the performance of NATs notably with significantly reduced computational cost. Experiments on ImageNet-256 2 & 512 2 and MS-COCO validate the effectiveness of ENAT. Code and pre-trained models will be released at https://github.com/LeapLabTHU/ENAT. Zanlin Ni, Yulin Wang 0002, Renping Zhou, Yizeng Han, Zhiyuan Liu 0001, Yuan Yao 0013, Gao Huang 0001 |
NeurIPS | 7 |
| 2023 | Visually Grounded Commonsense Knowledge AcquisitionabstractLarge-scale commonsense knowledge bases empower a broad range of AI applications, where the automatic extraction of commonsense knowledge (CKE) is a fundamental and challenging problem. CKE from text is known for suffering from the inherent sparsity and reporting bias of commonsense in text. Visual perception, on the other hand, contains rich commonsense knowledge about real-world entities, e.g., (person, can_hold, bottle), which can serve as promising sources for acquiring grounded commonsense knowledge. In this work, we present CLEVER, which formulates CKE as a distantly supervised multi-instance learning problem, where models learn to summarize commonsense relations from a bag of images about an entity pair without any human annotation on image instances. To address the problem, CLEVER leverages vision-language pre-training models for deep understanding of each image in the bag, and selects informative instances from the bag to summarize commonsense entity relations via a novel contrastive attention mechanism. Comprehensive experimental results in held-out and human evaluation show that CLEVER can extract commonsense knowledge in promising quality, outperforming pre-trained language model-based methods by 3.9 AUC and 6.4 mAUC points. The predicted commonsense scores show strong correlation with human judgment with a 0.78 Spearman coefficient. Moreover, the extracted commonsense can also be grounded into images with reasonable interpretability. The data and codes can be obtained at https://github.com/thunlp/CLEVER. Yuan Yao 0013, Tianyu Yu 0002, Mengdi Li 0006, Ruobing Xie, Cornelius Weber, Zhiyuan Liu 0001, Hai-Tao Zheng 0002, Stefan Wermter, Tat-Seng Chua, Maosong Sun 0001 |
AAAI | 1 |
| 2022 | Unpaired Cartoon Image Synthesis via Gated Cycle MappingabstractIn this paper, we present a general-purpose solution to cartoon image synthesis with unpaired training data. In contrast to previous works learning pre-defined cartoon styles for specified usage scenarios (portrait or scene), we aim to train a common cartoon translator which can not only simultaneously render exaggerated anime faces and realistic cartoon scenes, but also provide flexible user controls for desired cartoon styles. It is challenging due to the complexity of the task and the absence of paired data. The core idea of the proposed method is to introduce gated cycle mapping, that utilizes a novel gated mapping unit to produce the category-specific style code and embeds this code into cycle networks to control the translation process. For the concept of category, we classify images into different categories (e.g., 4 types: photo/cartoon portrait/scene) and learn finer-grained category translations rather than overall mappings between two domains (e.g., photo and cartoon). Furthermore, the proposed method can be easily extended to cartoon video generation with an auxiliary dataset and a new adaptive style loss. Experimental results demonstrate the superiority of the proposed method over the state of the art and validate its effectiveness in the brand-new task of general cartoon image synthesis. Yifang Men, Yuan Yao 0013, Miaomiao Cui, Zhouhui Lian, Xuansong Xie, Xian-Sheng Hua 0001 |
CVPR | 2 |
| 2022 | Structure-Aware Flow Generation for Human Body ReshapingabstractBody reshaping is an important procedure in portrait photo retouching. Due to the complicated structure and multifarious appearance of human bodies, existing methods either fall back on the 3D domain via body morphable model or resort to keypoint-based image deformation, leading to inefficiency and unsatisfied visual quality. In this paper, we address these limitations by formulating an end-to-end flow generation architecture under the guidance of body structural priors, including skeletons and Part Affinity Fields, and achieve unprecedentedly controllable performance under arbitrary poses and garments. A compositional attention mechanism is introduced for capturing both visual perceptual correlations and structural associations of the human body to reinforce the manipulation consistency among related parts. For a comprehensive evaluation, we construct the first large-scale body reshaping dataset, namely BR-5K, which contains 5,000 portrait photos as well as professionally retouched targets. Extensive experiments demonstrate that our approach significantly outperforms existing state-of-the-art methods in terms of visual performance, controllability, and efficiency. The dataset is available at our website: https://github.com/JianqiangRen/FlowBasedBodyReshaping. Jianqiang Ren, Yuan Yao 0013, Biwen Lei, Miaomiao Cui, Xuansong Xie |
CVPR | 2 |
| 2022 | Fine-Grained Scene Graph Generation with Data Transfer
Yuan Yao 0013, Wei Ji 0008, Zhiyuan Liu 0001, Maosong Sun 0001, Tat-Seng Chua |
ECCV (27) | 2 |
| 2022 | PEVL: Position-enhanced Pre-training and Prompt Tuning for Vision-language ModelsabstractVision-language pre-training (VLP) has shown impressive performance on a wide range of cross-modal tasks, where VLP models without reliance on object detectors are becoming the mainstream due to their superior computation efficiency and competitive performance.However, the removal of object detectors also deprives the capability of VLP models in explicit object modeling, which is essential to various position-sensitive vision-language (VL) tasks, such as referring expression comprehension and visual commonsense reasoning.To address the challenge, we introduce PEVL that enhances the pre-training and prompt tuning of VLP models with explicit object position modeling.Specifically, PEVL reformulates discretized object positions and language in a unified language modeling framework, which facilitates explicit VL alignment during pretraining, and also enables flexible prompt tuning for various downstream tasks.We show that PEVL enables state-of-the-art performance of detector-free VLP models on position-sensitive tasks such as referring expression comprehension and phrase grounding, and also improves the performance on position-insensitive tasks with grounded inputs.We make the data and code for this paper publicly available at https://github.com/thunlp/PEVL. Yuan Yao 0013, Wei Ji 0008, Zhiyuan Liu 0001, Tat-Seng Chua, Maosong Sun 0001 |
EMNLP | 1 |
| 2022 | DCT-net: domain-calibrated translation for portrait stylizationabstractThis paper introduces DCT-Net, a novel image translation architecture for few-shot portrait stylization. Given limited style exemplars (~100), the new architecture can produce high-quality style transfer results with advanced ability to synthesize high-fidelity contents and strong generality to handle complicated scenes (e.g., occlusions and accessories). Moreover, it enables full-body image translation via one elegant evaluation network trained by partial observations (i.e., stylized heads). Few-shot learning based style transfer is challenging since the learned model can easily become overfitted in the target domain, due to the biased distribution formed by only a few training examples. This paper aims to handle the challenge by adopting the key idea of "calibration first, translation later" and exploring the augmented global structure with locally-focused translation. Specifically, the proposed DCT-Net consists of three modules: a content adapter borrowing the powerful prior from source photos to calibrate the content distribution of target samples; a geometry expansion module using affine transformations to release spatially semantic constraints; and a texture translation module leveraging samples produced by the calibrated distribution to learn a fine-grained conversion. Experimental results demonstrate the proposed method's superiority over the state of the art in head stylization and its effectiveness on full image translation with adaptive deformations. Our code is publicly available at https://github.com/menyifang/DCT-Net. Yifang Men, Yuan Yao 0013, Miaomiao Cui, Zhouhui Lian, Xuansong Xie |
ACM Trans. Graph. | 2 |
| 2021 | Adversarial Language Games for Advanced Natural Language IntelligenceabstractWe study the problem of adversarial language games, in which multiple agents with conflicting goals compete with each other via natural language interactions. While adversarial language games are ubiquitous in human activities, little attention has been devoted to this field in natural language processing. In this work, we propose a challenging adversarial language game called Adversarial Taboo as an example, in which an attacker and a defender compete around a target word. The attacker is tasked with inducing the defender to utter the target word invisible to the defender, while the defender is tasked with detecting the target word before being induced by the attacker. In Adversarial Taboo, a successful attacker and defender need to hide or infer the intention, and induce or defend during conversations. This requires several advanced language abilities, such as adversarial pragmatic reasoning and goal-oriented language interactions in open domain, which will facilitate many downstream NLP tasks. To instantiate the game, we create a game environment and a competition platform. Comprehensive experiments on several baseline attack and defense strategies show promising and interesting results, based on which we discuss some directions for future research. Yuan Yao 0013, Haoxi Zhong, Zhengyan Zhang, Xu Han 0007, Xiaozhi Wang, Kai Zhang 0033, Chaojun Xiao, Guoyang Zeng, Zhiyuan Liu 0001, Maosong Sun 0001 |
AAAI | 1 |
| 2021 | Turn the Combination Lock: Learnable Textual Backdoor Attacks via Word SubstitutionabstractFanchao Qi, Yuan Yao, Sophia Xu, Zhiyuan Liu, Maosong Sun. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Fanchao Qi, Yuan Yao 0013, Sophia Xu, Zhiyuan Liu 0001, Maosong Sun 0001 |
ACL/IJCNLP (1) | 2 |
| 2021 | ONION: A Simple and Effective Defense Against Textual Backdoor AttacksabstractBackdoor attacks are a kind of emergent training-time threat to deep neural networks (DNNs).They can manipulate the output of DNNs and possess high insidiousness.In the field of natural language processing, some attack methods have been proposed and achieve very high attack success rates on multiple popular models.Nevertheless, there are few studies on defending against textual backdoor attacks.In this paper, we propose a simple and effective textual backdoor defense named ONION, which is based on outlier word detection and, to the best of our knowledge, is the first method that can handle all the textual backdoor attack situations.Experiments demonstrate the effectiveness of our model in defending BiLSTM and BERT against five different backdoor attacks.All the code and data of this paper can be obtained at https: //github.com/thunlp/ONION. Fanchao Qi, Yangyi Chen, Mukai Li, Yuan Yao 0013, Zhiyuan Liu 0001, Maosong Sun 0001 |
EMNLP (1) | 4 |
| 2021 | CodRED: A Cross-Document Relation Extraction Dataset for Acquiring Knowledge in the WildabstractExisting relation extraction (RE) methods typically focus on extracting relational facts between entity pairs within single sentences or documents.However, a large quantity of relational facts in knowledge bases can only be inferred across documents in practice.In this work, we present the problem of crossdocument RE, making an initial step towards knowledge acquisition in the wild.To facilitate the research, we construct the first human-annotated cross-document RE dataset CodRED.Compared to existing RE datasets, CodRED presents two key challenges: Given two entities, (1) it requires finding the relevant documents that can provide clues for identifying their relations; (2) it requires reasoning over multiple documents to extract the relational facts.We conduct comprehensive experiments to show that CodRED is challenging to existing RE methods including strong BERT-based models.We make CodRED and the code for our baselines publicly available at https://github.com/thunlp/CodRED. Yuan Yao 0013, Jiaju Du, Yankai Lin 0001, Peng Li 0030, Zhiyuan Liu 0001, Jie Zhou 0016, Maosong Sun 0001 |
EMNLP (1) | 1 |
| 2021 | Open Hierarchical Relation ExtractionabstractKai Zhang, Yuan Yao, Ruobing Xie, Xu Han, Zhiyuan Liu, Fen Lin, Leyu Lin, Maosong Sun. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Kai Zhang 0033, Yuan Yao 0013, Ruobing Xie, Xu Han 0007, Zhiyuan Liu 0001, Fen Lin 0002, Leyu Lin, Maosong Sun 0001 |
NAACL-HLT | 2 |
| 2020 | Meta-Information Guided Meta-Learning for Few-Shot Relation ClassificationabstractFew-shot classification requires classifiers to adapt to new classes with only a few training instances.State-of-the-art meta-learning approaches such as MAML learn how to initialize and fast adapt parameters from limited instances, which have shown promising results in few-shot classification.However, existing meta-learning models solely rely on implicit instance-based statistics, and thus suffer from instance unreliability and weak interpretability.To solve this problem, we propose a novel meta-information guided meta-learning (MIML) framework, where semantic concepts of classes provide strong guidance for meta-learning in both initialization and adaptation.In effect, our model can establish connections between instance-based information and semantic-based information, which enables more effective initialization and faster adaptation.Comprehensive experimental results on few-shot relation classification demonstrate the effectiveness of the proposed framework.Notably, MIML achieves comparable or superior performance to humans with only one shot on FewRel evaluation.The source code and experiment details of this paper can be obtained from https://github.com/thunlp/MIML. Bowen Dong 0005, Yuan Yao 0013, Ruobing Xie, Tianyu Gao 0001, Xu Han 0007, Zhiyuan Liu 0001, Fen Lin 0002, Leyu Lin, Maosong Sun 0001 |
COLING | 2 |
| 2020 | Boosting Semantic Human Matting With Coarse AnnotationsabstractSemantic human matting aims to estimate the per-pixel opacity of the foreground human regions. It is quite challenging that usually requires user interactive trimaps and plenty of high quality annotated data. Annotating such kind of data is labor intensive and requires great skills beyond normal users, especially considering the very detailed hair part of humans. In contrast, coarse annotated human dataset is much easier to acquire and collect from the public dataset. In this paper, we propose to leverage coarse annotated data coupled with fine annotated data to boost end-to-end semantic human matting without trimaps as extra input. Specifically, We train a mask prediction network to estimate the coarse semantic mask using the hybrid data, and then propose a quality unification network to unify the quality of the previous coarse mask outputs. A matting refinement network takes the unified mask and the input image to predict the final alpha matte. The collected coarse annotated dataset enriches our dataset significantly, allows generating high quality alpha matte for real images. Experimental results show that the proposed method performs comparably against state-of-the-art methods. Moreover, the proposed method can be used for refining coarse annotated public dataset, as well as semantic segmentation methods, which reduces the cost of annotating high quality human data to a great extent. Jinlin Liu, Yuan Yao 0013, Wendi Hou, Miaomiao Cui, Xuansong Xie, Changshui Zhang, Xian-Sheng Hua 0001 |
CVPR | 2 |
| 2020 | Denoising Relation Extraction from Document-level Distant SupervisionabstractDistant supervision (DS) has been widely used to generate auto-labeled data for sentencelevel relation extraction (RE), which improves RE performance.However, the existing success of DS cannot be directly transferred to the more challenging document-level relation extraction (DocRE), since the inherent noise in DS may be even multiplied in document level and significantly harm the performance of RE.To address this challenge, we propose a novel pre-trained model for DocRE, which denoises the document-level DS data via multiple pre-training tasks.Experimental results on the large-scale DocRE benchmark show that our model can capture useful information from noisy DS data and achieve promising results.The source code of this paper can be found in https://github.com/thunlp/DSDocRE. Chaojun Xiao, Yuan Yao 0013, Ruobing Xie, Xu Han 0007, Zhiyuan Liu 0001, Maosong Sun 0001, Fen Lin 0002, Leyu Lin |
EMNLP (1) | 2 |
| 2019 | DocRED: A Large-Scale Document-Level Relation Extraction DatasetabstractYuan Yao, Deming Ye, Peng Li, Xu Han, Yankai Lin, Zhenghao Liu, Zhiyuan Liu, Lixin Huang, Jie Zhou, Maosong Sun. Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics. 2019. Yuan Yao 0013, Deming Ye, Peng Li 0030, Xu Han 0007, Yankai Lin 0001, Zhenghao Liu 0001, Zhiyuan Liu 0001, Lixin Huang, Jie Zhou 0016, Maosong Sun 0001 |
ACL (1) | 1 |
| 2019 | Attention-Aware Multi-Stroke Style TransferabstractNeural style transfer has drawn considerable attention from both academic and industrial field. Although visual effect and efficiency have been significantly improved, existing methods are unable to coordinate spatial distribution of visual attention between the content image and stylized image, or render diverse level of detail via different brush strokes. In this paper, we tackle these limitations by developing an attention-aware multi-stroke style transfer model. We first propose to assemble self-attention mechanism into a style-agnostic reconstruction autoencoder framework, from which the attention map of a content image can be derived. By performing multi-scale style swap on content features and style features, we produce multiple feature maps reflecting different stroke patterns. A flexible fusion strategy is further presented to incorporate the salient characteristics from the attention map, which allows integrating multiple stroke patterns into different spatial regions of the output image harmoniously. We demonstrate the effectiveness of our method, as well as generate comparable stylized images with multiple stroke patterns against the state-of-the-art methods. Yuan Yao 0013, Jianqiang Ren, Xuansong Xie, Weidong Liu 0001 |
CVPR | 1 |
| 2019 | Open Relation Extraction: Relational Knowledge Transfer from Supervised Data to Unsupervised DataabstractRuidong Wu, Yuan Yao, Xu Han, Ruobing Xie, Zhiyuan Liu, Fen Lin, Leyu Lin, Maosong Sun. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Ruidong Wu, Yuan Yao 0013, Xu Han 0007, Ruobing Xie, Zhiyuan Liu 0001, Fen Lin 0002, Leyu Lin, Maosong Sun 0001 |
EMNLP/IJCNLP (1) | 2 |
| 2018 | FewRel: A Large-Scale Supervised Few-shot Relation Classification Dataset with State-of-the-Art EvaluationabstractWe present a Few-Shot Relation Classification Dataset (FewRel), consisting of 70, 000 sentences on 100 relations derived from Wikipedia and annotated by crowdworkers.The relation of each sentence is first recognized by distant supervision methods, and then filtered by crowdworkers.We adapt the most recent state-of-the-art few-shot learning methods for relation classification and conduct thorough evaluation of these methods.Empirical results show that even the most competitive few-shot learning models struggle on this task, especially as compared with humans.We also show that a range of different reasoning skills are needed to solve our task.These results indicate that few-shot relation classification remains an open problem and still requires further research.Our detailed analysis points multiple directions for future research.All details and resources about the dataset and baselines are released on http://zhuhao.me/ fewrel. Xu Han 0007, Hao Zhu 0006, Pengfei Yu 0001, Yuan Yao 0013, Zhiyuan Liu 0001, Maosong Sun 0001 |
EMNLP | 5 |
| 2018 | CAMF: Context Aware Matrix Factorization for Social RecommendationabstractSocial Networks have experienced increased popularity and rapid growth in recent years. Recommendation is significant for users due to the extremely large amount of information in Social Networks. Most existing recommender systems rely on collaborative filtering techniques which focus on recommending the most relevant items to users based on past rating information of users or items. In Social Networks, the cold-start and data sparsity problems are very serious because new users and items are growing rapidly. Taking the Event Recommendation problem in Event-Based Social Networks as a scenario, many events are newly created and have few feedbacks. Existed collaborative filtering based methods will fail for Social Recommendation due to these problems. Therefore, a more sophisticated recommendation mechanism that can efficiently combine various contextual information to further improve recommendation quality is desired. In this paper, we propose a Context Aware Matrix Factorization model called CAMF which models implicit feedbacks and various contextual information simultaneously for Social Recommendation. Specifically, CAMF is a unified model that combines the Matrix Factorization model which models implicit feedbacks with the Linear Contextual Features model which models explicit contextual features. Extensive experiments on a large real-world dataset demonstrate that the CAMF model significantly outperforms state-of-the-art methods by 12.7% in terms of accuracy for the Event Recommendation problem. Yulong Gu, Weidong Liu 0001, Lixin Zou, Yuan Yao 0013 |
Web Intell. | 5 |
| 2016 | We Know What You Are Doing or Going to Do: Towards Accurate Human Activities SensingabstractUnderstanding Activities of Human Daily Life is a fundamental and essential AI problem for Pervasive Computing and Human-Computer Interaction. Activity Sensing has attracted enormous research on activity recognition from mobile sensor data. However, there are two challenging problems: There is no standard taxonomy of activities and there is a lack of research on sensing high level activities. To this end, firstly, we built AHDL, the first knowledge base of Activities of Human Daily Life in this planet leveraging a large time use surveys. AHDL not only has a taxonomy of activities but also has common sense knowledge of these activities. Secondly, we designed ActivitySensor, a Conditional Random Fields based Sensor for sensing high level activities in AHDL. To be specific, ActivitySensor performs activity sensing using Conditional Random Fields model by combining contextual signals (time, location, previous activity and related person) and demographical signals. Extensive experiments demonstrated that ActivitySensor can improve the accuracy of activity recognition about 15% comparing to state-of-the-art methods on the same dataset. What's more, we revealed that ActivitySensor can predict what will you do next with high accuracy. Yulong Gu, Mengjia Feng, Yuan Yao 0013, Weidong Liu 0001 |
ICCCN | 3 |
| 2016 | We Know Where You Are: Home Location Identification in Location-Based Social NetworksabstractThe rapid spread of smartphones has led to the increasing popularity of Location-Based Social Networks(LBSNs) like Foursquare, Gowalla, Facebook Places and so on where users can publish information about their current location. In LBSNs, identifying home locations of users is very important for various applications like effective location-based advertisement and recommendation. However, this problem is rather challenging because the location information in LBSNs is sparse and noisy: Only a small percentage of users share their home location information due to privacy concerns; users may check in at diverse places far from their home and make friends far away; many users even do not have any check-in information. In this paper, we propose a trust-based influence model, named as TSU to solve the problem. To be specific, TSU is a Trust-based unified probabilistic model that models edges in LBSNs based on signals from Social relationship data(social friendship, social trust) and User-centric data(check-in data) in LBSN. We proposed a Home Location Identification method based on TSU model and evaluate it on a large real-world LBSNs dataset. Extensive experiments demonstrate that our method significantly outperforms state-of-the-art methods. Yulong Gu, Yuan Yao 0013, Weidong Liu 0001 |
ICCCN | 2 |
| 2016 | Grouped Text Clustering Using Non-Parametric Gaussian Mixture Experts
Yu Rong 0003, Yuan Yao 0013, Weidong Liu 0001 |
PRICAI | 3 |
| 2016 | Towards Accurate Relation Extraction from WikipediaabstractEnormous efforts of human volunteers have made Wikipedia become a treasure of textual knowledge. Relation extraction that aims at extracting structured knowledge in the unstructured texts in Wikipedia is an appealing but quite challenging problem because it's hard for machines to understand plain texts. Existing methods are not effective enough because they understand relation types in textual level without exploiting knowledge behind plain texts. In this paper, we propose a novel framework called Athena 2.0 leveraging Semantic Patterns which are patterns that can understand relation types in semantic level to solve this problem. Extensive experiments show that Athena 2.0 significantly outperforms existing methods. Yulong Gu, Weidong Liu 0001, Yuan Yao 0013, Lixin Zou |
WI | 4 |
| 2016 | Context Aware Matrix Factorization for Event Recommendation in Event-Based Social NetworksabstractEvent-based Social Networks(EBSNs) which combine online interactions and offline events among users have experienced increased popularity and rapid growth recently. In EBSNs, event recommendation is significant for users due to the extremely large amount of events. However, the event recommendation problem is rather challenging because it faces a serious cold-start problem: Events have short life time and new events are registered by only a few users. What's more, there are only implicit feedback information. Existing approaches like collaborative filtering methods are not suitable for this scenario. In this paper, we propose a Context Aware Matrix Factorization model called AlphaMF to tackle with the problem. Specifically, AlphaMF is a unified model that combines the Matrix Factorization model which models implicit feedbacks with the Linear contextual features model which models explicit contextual features. Extensive experiments on a large real-world EBSN dataset demonstrate that the AlphaMF model significantly outperforms state-of-the-art methods by 11%. Yulong Gu, Weidong Liu 0001, Lixin Zou, Yuan Yao 0013 |
WI | 5 |
| 2015 | A Mixture Distribution Based System in BitTorrent-Like P2P NetworksabstractIn this paper, we develop a novel file sharing system based on a mixture distribution model working in the BitTorrentlike p2p networks. The BitTorrent's built-in “tit-for-tat” unchoking mechanism delays the initial file sharing process for newly joined peers as well as brings the problem of free-riding that peers only download from others but never contribute to the network. We demonstrate a file sharing mechanism which allows peers to share pieces according to different mixture distributions. The mechanism utilizes the historical contributions of peers in the network to inspire cooperation among peers, Along with the mixture distribution model, the peers can only obtain the whole file by contributing to the network continuously which deters the free-riding behaviors. We theoretically prove that the peers take the truthful revealing as their dominant strategy and our system can speed up the initial process of file sharing. The experiments show that the proposed system performs well and has good scalability, as well as prevents the free-riding problem elegantly. Yuan Yao 0013, Weidong Liu 0001 |
ICPADS | 1 |