EDBT 2026 Demo / reviewers in the wild / expert
Yangguang Li 0001
dblp:132/4829-1
· DBLP profile ↗
26ranked-venue papers
3as first author
25since 2021 · last 2026
0000-0002-6090-3899ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 22 · 2 first-author · 22 since 2021Graphics, computer vision, multimedia, augmented reality and games · 16 · 1 first-author · 15 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | DetailGen3D: Generative 3D Geometry Enhancement via Data-Dependent FlowabstractModern 3D generation methods can rapidly create shapes from sparse or single views, but their outputs often lack geometric detail due to computational constraints. We present DetailGen3D, a generative approach specifically designed to enhance these generated 3D shapes. Our key insight is to model the coarse-to-fine transformation directly through data-dependent flows in latent space, avoiding the computational overhead of large-scale 3D generative models. We introduce a token matching strategy that ensures accurate spatial correspondence during refinement, enabling local detail synthesis while preserving global structure. By carefully designing our training data to match the characteristics of synthesized coarse shapes, our method can effectively enhance shapes produced by various 3D generation and reconstruction approaches, from single-view to sparse multi-view inputs. Extensive experiments demonstrate that DetailGen3D achieves high-fidelity geometric detail synthesis while maintaining efficiency in training. Our project page is https://detailgen3d.github.io/DetailGen3D/ Ken Deng, Jingxiang Sun, Zixin Zou, Yangguang Li 0001, Yan-Pei Cao 0001, Yebin Liu, Ding Liang |
3DV | 5 |
| 2025 | MIDI: Multi-Instance Diffusion for Single Image to 3D Scene GenerationabstractThis paper introduces MIDI, a novel paradigm for compositional 3D scene generation from a single image. Unlike existing methods that rely on reconstruction or retrieval techniques or recent approaches that employ multi-stage object-by-object generation, MIDI extends pre-trained image-to-3D object generation models to multi-instance diffusion models, enabling the simultaneous generation of multiple 3D instances with accurate spatial relationships and high generalizability. At its core, MIDI incorporates a novel multi-instance attention mechanism, that effectively captures inter-object interactions and spatial coherence directly within the generation process, without the need for complex multi-step processes. The method utilizes partial object images and global scene context as inputs, directly modeling object completion during 3D generation. During training, we effectively supervise the interactions between 3D instances using a limited amount of scene-level data, while incorporating single-object data for regularization, thereby maintaining the pre-trained generalization ability. MIDI demonstrates state-of-the-art performance in image-to-scene generation, validated through evaluations on synthetic data, real-world scene data, and stylized scene images generated by text-to-image diffusion models. Zehuan Huang, Xingqiao An, Yunhan Yang, Yangguang Li 0001, Zixin Zou, Ding Liang, Xihui Liu, Yan-Pei Cao 0001, Lu Sheng |
CVPR | 5 |
| 2025 | PSHuman: Photorealistic Single-image 3D Human Reconstruction using Cross-Scale Multiview Diffusion and Explicit RemeshingabstractPhotorealistic 3D human modeling is essential for various applications and has seen tremendous progress. However, existing methods for monocular full-body reconstruction, typically relying on front and/or predicted back view, still struggle with satisfactory performance due to the ill-posed nature of the problem and sophisticated self-occlusions. In this paper, we propose PSHuman, a novel framework that explicitly reconstructs human meshes utilizing priors from the multiview diffusion model. It is found that directly applying multiview diffusion on single-view human images leads to severe geometric distortions, especially on generated faces. To address it, we propose a cross-scale diffusion that models the joint probability distribution of global full-body shape and local facial characteristics, enabling identity-preserved novel-view generation without geometric distortion. Moreover, to enhance cross-view body shape consistency of varied human poses, we condition the generative model on parametric models (SMPL-X), which provide body priors and prevent unnatural views inconsistent with human anatomy. Leveraging the generated multiview normal and color images, we present SMPLX-initialized explicit human carving to recover realistic textured human meshes efficiently. Extensive experiments on CAPE and THuman2.1 demonstrate PSHuman’s superiority in geometry details, texture fidelity, and generalization capability. Wangguandong Zheng, Yuan Liu 0025, Tao Yu 0007, Yangguang Li 0001, Xingqun Qi, Xiaowei Chi, Si-Yu Xia, Yan-Pei Cao 0001, Wei Xue 0002, Wenhan Luo, Yike Guo |
CVPR | 5 |
| 2025 | SparseFlex: High-Resolution and Arbitrary-Topology 3D Shape ModelingabstractCreating high-fidelity 3D meshes with arbitrary topology, including open surfaces and complex interiors, remains a significant challenge. Existing implicit field methods often require costly and detail-degrading watertight conversion, while other approaches struggle with high resolutions. This paper introduces SparseFlex, a novel sparse-structured isosurface representation that enables differentiable mesh reconstruction at resolutions up to $1024^3$ directly from rendering losses. SparseFlex combines the accuracy of Flexicubes with a sparse voxel structure, focusing computation on surface-adjacent regions and efficiently handling open surfaces. Crucially, we introduce a frustum-aware sectional voxel training strategy that activates only relevant voxels during rendering, dramatically reducing memory consumption and enabling high-resolution training. This also allows, for the first time, the reconstruction of mesh interiors using only rendering supervision. Building upon this, we demonstrate a complete shape modeling pipeline by training a variational autoencoder (VAE) and a rectified flow transformer for high-quality 3D shape generation. Our experiments show state-of-the-art reconstruction accuracy, with a ~82% reduction in Chamfer Distance and a ~88% increase in F-score compared to previous methods, and demonstrate the generation of high-resolution, detailed 3D shapes with arbitrary topology. By enabling high-resolution, differentiable mesh reconstruction and generation with rendering losses, SparseFlex significantly advances the state-of-the-art in 3D shape representation and modeling. Xianglong He, Zixin Zou, Chia-Hao Chen, Ding Liang, Chun Yuan 0003, Wanli Ouyang, Yan-Pei Cao 0001, Yangguang Li 0001 |
ICCV | 9 |
| 2025 | TAR3D: Creating High-Quality 3D Assets Via Next-Part PredictionabstractWe present TAR3D, a novel framework that consists of a 3D-aware Vector Quantized-Variational AutoEncoder (VQ-VAE) and a Generative Pre-trained Transformer (GPT) to generate high-quality 3D assets. The core insight of this work is to migrate the multimodal unification and promising learning capabilities of the next-token prediction paradigm to conditional 3D object generation. To achieve this, the 3D VQ-VAE first encodes a wide range of 3D shapes into a compact triplane latent space and utilizes a set of discrete representations from a trainable codebook to reconstruct fine-grained geometries under the supervision of query point occupancy. Then, the 3D GPT, equipped with a custom triplane position embedding called TriPE, predicts the codebook index sequence with prefilling prompt tokens in an autoregressive manner so that the composition of 3D geometries can be modeled part by part. Extensive experiments on ShapeNet and Objaverse demonstrate that TAR3D can achieve superior generation quality over existing methods in text-to-3D and image-to-3D tasks Xuying Zhang, Yangguang Li 0001, Renrui Zhang, Kai Wang 0001, Wanli Ouyang, Zhiwei Xiong, Peng Gao 0007, Qibin Hou, Ming-Ming Cheng |
ICCV | 3 |
| 2025 | DREAM: Disentangling Risks to Enhance Safety Alignment in Multimodal Large Language ModelsabstractJianyu Liu, Hangyu Guo, Ranjie Duan, Xingyuan Bu, Yancheng He, Shilong Li, Hui Huang, Jiaheng Liu, Yucheng Wang, Chenchen Jing, Xingwei Qu, Xiao Zhang, Pei Wang, Yanan Wu, Jihao Gu, Yangguang Li, Jianke Zhu. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Jianyu Liu, Hangyu Guo, Ranjie Duan, Xingyuan Bu, Yancheng He, Hui Huang 0021, Chenchen Jing, Xingwei Qu, Jihao Gu, Yangguang Li 0001, Jianke Zhu |
NAACL (Long Papers) | 16 |
| 2025 | Flow-GRPO: Training Flow Matching Models via Online RLabstractWe propose Flow-GRPO, the first method to integrate online policy gradient reinforcement learning (RL) into flow matching models. Our approach uses two key strategies: (1) an ODE-to-SDE conversion that transforms a deterministic Ordinary Differential Equation (ODE) into an equivalent Stochastic Differential Equation (SDE) that matches the original model's marginal distribution at all timesteps, enabling statistical sampling for RL exploration; and (2) a Denoising Reduction strategy that reduces training denoising steps while retaining the original number of inference steps, significantly improving sampling efficiency without sacrificing performance. Empirically, Flow-GRPO is effective across multiple text-to-image tasks. For compositional generation, RL-tuned SD3.5-M generates nearly perfect object counts, spatial relations, and fine-grained attributes, increasing GenEval accuracy from $63\%$ to $95\%$. In visual text rendering, accuracy improves from $59\%$ to $92\%$, greatly enhancing text generation. Flow-GRPO also achieves substantial gains in human preference alignment. Notably, very little reward hacking occurred, meaning rewards did not increase at the cost of appreciable image quality or diversity degradation. Jie Liu 0047, Gongye Liu, Jiajun Liang, Yangguang Li 0001, Xintao Wang 0002, Pengfei Wan 0001, Di Zhang 0026, Wanli Ouyang |
NeurIPS | 4 |
| 2025 | ShapeGen: Towards High-Quality 3D Shape SynthesisabstractInspired by generative paradigms in image and video, 3D shape generation has made notable progress, enabling the rapid synthesis of high-fidelity 3D assets from a single image. However, current methods still face challenges, including the lack of intricate details, overly smoothed surfaces, and fragmented thin-shell structures. These limitations leave the generated 3D assets still one step short of meeting the standards favored by artists. In this paper, we present ShapeGen, which achieves high-quality image-to-3D shape generation through 3D representation and supervision improvements, resolution scaling up, and the advantages of linear transformers. These advancements allow the generated assets to be seamlessly integrated into 3D pipelines, facilitating their widespread adoption across various applications. Specifically, in contrast to existing methods: 1) We investigate how different representations and VAE supervision strategies affect the generation process, and address issues like aliasing artifacts and fragmented thin-shell structures by using an TSDF-based representation supervised with BCE loss. 2) We scale up the resolution of 3D data, image conditioning inputs, and the number of latent tokens to enhance generation fidelity. 3) We adopt mixed conditioning using raw RGB images and normal maps during training, effectively resolving ambiguities caused by inconsistencies between ControlNet-generated RGB images and the underlying geometry from untextured assets. 4) We replace the original softmax attention with linear attention to improve training and inference efficiency when handling a large number of latent tokens. 5) We introduce an inference-time scaling strategy that enhances generation quality at test time. Through extensive experiments, we validate the impact of these improvements on overall performance. Ultimately, thanks to the synergistic effects of these enhancements, ShapeGen achieves a significant leap in image-to-3D generation, establishing a new state-of-the-art performance. Yangguang Li 0001, Xianglong He, Zixin Zou, Zexiang Liu, Wanli Ouyang, Ding Liang, Yan-Pei Cao 0001 |
SIGGRAPH Asia | 1 |
| 2024 | Exploring Temporal Feature Correlation for Efficient and Stable Video Semantic SegmentationabstractThis paper tackles the problem of efficient and stable video semantic segmentation. While stability has been under-explored, prevalent work in efficient video semantic segmentation uses the keyframe paradigm. They efficiently process videos by only recomputing the low-level features and reusing high-level features computed at selected keyframes. In addition, the reused features stabilize the predictions across frames, thereby improving video consistency. However, dynamic scenes in the video can easily lead to misalignments between reused and recomputed features, which hampers performance. Moreover, relying on feature reuse to improve prediction consistency is brittle; an erroneous alignment of the features can easily lead to unstable predictions. Therefore, the keyframe paradigm exhibits a dilemma between stability and performance. We address this efficiency and stability challenge using a novel yet simple Temporal Feature Correlation (TFC) module. It uses the cosine similarity between two frames’ low-level features to inform the semantic label’s consistency across frames. Specifically, we selectively reuse label-consistent features across frames through linear interpolation and update others through sparse multi-scale deformable attention. As a result, we no longer directly reuse features to improve stability and thus effectively solve feature misalignment. This work provides a significant step towards efficient and stable video semantic segmentation. On the VSPW dataset, our method significantly improves the prediction consistency of image-based methods while being as fast and accurate. Matthieu Lin, Jenny Sheng, Yubin Hu 0001, Yangguang Li 0001, Andrew Zhao, Gao Huang 0001, Yong-Jin Liu 0001 |
AAAI | 4 |
| 2024 | EpiDiff: Enhancing Multi-View Synthesis via Localized Epipolar-Constrained DiffusionabstractGenerating multiview images from a single view facilitates the rapid generation of a 3D mesh conditioned on a single image. Recent methods [31] that introduce 3D global representation into diffusion models have shown the potential to generate consistent multiviews, but they have reduced generation speed and face challenges in maintaining generalizability and quality. To address this issue, we propose EpiDiff, a localized interactive multiview diffusion model. At the core of the proposed approach is to insert a lightweight epipolar attention block into the frozen diffusion model, leveraging epipolar constraints to enable cross-view interaction among feature maps of neighboring views. The newly initialized 3D modeling module preserves the original feature distribution of the diffusion model, exhibiting compatibility with a variety of base diffusion models. Experiments show that EpiDiff generates 16 multiview images in just 12 seconds, and it surpasses previous methods in quality evaluation metrics, including PSNR, SSIM and LPIPS. Additionally, EpiDiff can generate a more diverse distribution of views, improving the reconstruction quality from generated multiviews. Please see the project page at huanngzh.github.io/EpiDiff/. Zehuan Huang, Junting Dong, Yaohui Wang 0001, Yangguang Li 0001, Yan-Pei Cao 0001, Ding Liang, Yu Qiao 0001, Bo Dai 0002, Lu Sheng |
CVPR | 5 |
| 2024 | Triplane Meets Gaussian Splatting: Fast and Generalizable Single-View 3D Reconstruction with TransformersabstractRecent advancements in 3D reconstruction from single images have been driven by the evolution of generative models. Prominent among these are methods based on Score Distillation Sampling (SDS) and the adaptation ofdiffusion models in the 3D domain. Despite their progress, these techniques often face limitations due to slow optimization or rendering processes, leading to extensive training and optimization times. In this paper, we introduce a novel approach for single-view reconstruction that efficiently generates a 3D model from a single image via feed-forward inference. Our method utilizes two transformer-based networks, namely a point decoder and a triplane decoder, to reconstruct 3D objects using a hybrid Triplane-Gaussian intermediate representation. This hybrid representation strikes a balance, achieving a faster rendering speed compared to implicit representations while simultaneously delivering superior rendering quality than explicit representations. The point decoder is designed for generating point clouds from single images, offering an explicit representation which is then utilized by the triplane decoder to query Gaussian features for each point. This design choice addresses the challenges associated with directly regressing explicit 3D Gaussian attributes characterized by their non-structural nature. Subsequently, the 3D Gaussians are decoded by an MLP to enable rapid rendering through splatting. Both decoders are built upon a scalable, transformer-based architecture and have been efficiently trained on large-scale 3D datasets. The evaluations conducted on both synthetic datasets and real-world images demonstrate that our method not only achieves higher quality but also ensures a faster runtime in comparison to previous state-of-the-art techniques. Please see our project page at https://zouzx.github.io/TriplaneGaussian/ Zixin Zou, Yangguang Li 0001, Ding Liang, Yan-Pei Cao 0001, Song-Hai Zhang |
CVPR | 4 |
| 2024 | GVGEN: Text-to-3D Generation with Volumetric Representation
Xianglong He, Sida Peng, Yangguang Li 0001, Xiaoshui Huang, Chun Yuan 0003, Wanli Ouyang, Tong He 0001 |
ECCV (8) | 5 |
| 2024 | UniDream: Unifying Diffusion Priors for Relightable Text-to-3D Generation
Zexiang Liu, Yangguang Li 0001, Youtian Lin, Xin Yu 0004, Sida Peng, Yan-Pei Cao 0001, Xiaojuan Qi 0001, Xiaoshui Huang, Ding Liang, Wanli Ouyang |
ECCV (5) | 2 |
| 2024 | Text-to-3D with Classifier Score DistillationabstractText-to-3D generation has made remarkable progress recently, particularly with methods based on Score Distillation Sampling (SDS) that leverages pre-trained 2D diffusion models. While the usage of classifier-free guidance is well acknowledged to be crucial for successful optimization, it is considered an auxiliary trick rather than the most essential component. In this paper, we re-evaluate the role of classifier-free guidance in score distillation and discover a surprising finding: the guidance alone is enough for effective text-to-3D generation tasks.
We name this method Classifier Score Distillation (CSD), which can be interpreted as using an implicit classification model for generation. This new perspective reveals new insights for understanding existing techniques. We validate the effectiveness of CSD across a variety of text-to-3D tasks including shape generation, texture synthesis, and shape editing, achieving results superior to those of state-of-the-art methods. Our project page is https://xinyu-andy.github.io/Classifier-Score-Distillation Xin Yu 0004, Yangguang Li 0001, Ding Liang, Song-Hai Zhang, Xiaojuan Qi 0001 |
ICLR | 3 |
| 2024 | Lumina-Next : Making Lumina-T2X Stronger and Faster with Next-DiTabstractLumina-T2X is a nascent family of Flow-based Large Diffusion Transformers (Flag-DiT) that establishes a unified framework for transforming noise into various modalities, such as images and videos, conditioned on text instructions. Despite its promising capabilities, Lumina-T2X still encounters challenges including training instability, slow inference, and extrapolation artifacts. In this paper, we present Lumina-Next, an improved version of Lumina-T2X, showcasing stronger generation performance with increased training and inference efficiency. We begin with a comprehensive analysis of the Flag-DiT architecture and identify several suboptimal components, which we address by introducing the Next-DiT architecture with 3D RoPE and sandwich normalizations. To enable better resolution extrapolation, we thoroughly compare different context extrapolation methods applied to text-to-image generation with 3D RoPE, and propose Frequency- and Time-Aware Scaled RoPE tailored for diffusion transformers. Additionally, we introduce a sigmoid time discretization schedule for diffusion sampling, which achieves high-quality generation in 5-10 steps combined with higher-order ODE solvers. Thanks to these improvements, Lumina-Next not only improves the basic text-to-image generation but also demonstrates superior resolution extrapolation capabilities as well as multilingual generation using decoder-based LLMs as the text encoder, all in a zero-shot manner. To further validate Lumina-Next as a versatile generative framework, we instantiate it on diverse tasks including visual recognition, multi-views, audio, music, and point cloud generation, showcasing strong performance across these domains. By releasing all codes and model weights at https://github.com/Alpha-VLLM/Lumina-T2X, we aim to advance the development of next-generation generative AI capable of universal modeling. Le Zhuo, Ruoyi Du, Han Xiao 0010, Yangguang Li 0001, Rongjie Huang 0001, Wenze Liu, Fu-Yun Wang, Zhanyu Ma, Zehan Wang 0001, Kaipeng Zhang, Lirui Zhao, Si Liu 0001, Xiangyu Yue 0001, Wanli Ouyang, Yu Qiao 0001, Hongsheng Li 0001, Peng Gao 0007 |
NeurIPS | 4 |
| 2024 | Fast-BEV: A Fast and Strong Bird's-Eye View Perception BaselineabstractRecently, perception task based on Bird's-Eye View (BEV) representation has drawn more and more attention, and BEV representation is promising as the foundation for next-generation Autonomous Vehicle (AV) perception. However, most existing BEV solutions either require considerable resources to execute on-vehicle inference or suffer from modest performance. This paper proposes a simple yet effective framework, termed Fast-BEV, which is capable of performing faster BEV perception on the on-vehicle chips. Towards this goal, we first empirically find that the BEV representation can be sufficiently powerful without expensive transformer based transformation or depth representation. Our Fast-BEV consists of five parts, we innovatively propose (1) a lightweight deployment-friendly view transformation which fast transfers 2D image features to 3D voxel space, (2) a multi-scale image encoder which leverages multi-scale information for better performance, (3) an efficient BEV encoder which is particularly designed to speed up on-vehicle inference. We further introduce (4) a strong data augmentation strategy for both image and BEV space to avoid over-fitting, (5) a multi-frame feature fusion mechanism to leverage the temporal information. Among them, (1) and (3) enable Fast-BEV to be fast inference and deployment friendly on the on-vehicle chips, (2), (4) and (5) ensure that Fast-BEV has competitive performance. All these make Fast-BEV a solution with high performance, fast inference speed, and deployment-friendly on the on-vehicle chips of autonomous driving. Through experiments, on 2080Ti platform, our R50 model can run 52.6 FPS with 47.3% NDS on the nuScenes validation set, exceeding the 41.3 FPS and 47.5% NDS of the BEVDepth-R50 model (Li et al. 2022) and 30.2 FPS and 45.7% NDS of the BEVDet4D-R50 model (J. Huang and G. Huang, 2022). Our largest model (R101@900×1600) establishes a competitive 53.5% NDS on the nuScenes validation set. We further develop a benchmark with considerable accuracy and efficiency on current popular on-vehicle chips. Yangguang Li 0001, Bin Huang 0001, Zeren Chen, Yufeng Cui, Mingzhu Shen, Fenggang Liu, Enze Xie, Lu Sheng, Wanli Ouyang |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2024 | TEXGen: a Generative Diffusion Model for Mesh TexturesabstractWhile high-quality texture maps are essential for realistic 3D asset rendering, few studies have explored learning directly in the texture space, especially on large-scale datasets. In this work, we depart from the conventional approach of relying on pre-trained 2D diffusion models for testtime optimization of 3D textures. Instead, we focus on the fundamental problem of learning in the UV texture space itself. For the first time, we train a large diffusion model capable of directly generating high-resolution texture maps in a feed-forward manner. To facilitate efficient learning in high-resolution UV spaces, we propose a scalable network architecture that interleaves convolutions on UV maps with attention layers on point clouds. Leveraging this architectural design, we train a 700 million parameter diffusion model that can generate UV texture maps guided by text prompts and single-view images. Once trained, our model naturally supports various extended applications, including text-guided texture inpainting, sparse-view texture completion, and text-driven texture synthesis. The code is available at https://github.com/CVMI-Lab/TEXGen. Xin Yu 0004, Ze Yuan, Ying-Tian Liu, Yangguang Li 0001, Yan-Pei Cao 0001, Ding Liang, Xiaojuan Qi 0001 |
ACM Trans. Graph. | 6 |
| 2023 | Task-balanced distillation for object detection
Ruining Tang, Zhenyu Liu 0005, Yangguang Li 0001, Yiguo Song, Hui Liu 0037, Qide Wang, Guifang Duan, Jianrong Tan |
Pattern Recognit. | 3 |
| 2022 | Neighbor Regularized Bayesian Optimization for Hyperparameter Optimization
Yangguang Li 0001, Dong An 0002, Fenggang Liu |
BMVC | 2 |
| 2022 | IMCI: Integrate Multi-view Contextual Information for Fact Extraction and VerificationabstractWith the rapid development of automatic fake news detection technology, fact extraction and verification (FEVER) has been attracting more attention. The task aims to extract the most related fact evidences from millions of open-domain Wikipedia documents and then verify the credibility of corresponding claims. Although several strong models have been proposed for the task and they have made great process, we argue that they fail to utilize multi-view contextual information and thus cannot obtain better performance. In this paper, we propose to integrate multi-view contextual information (IMCI) for fact extraction and verification. For each evidence sentence, we define two kinds of context, i.e. intra-document context and inter-document context. Intra-document context consists of the document title and all the other sentences from the same document. Inter-document context consists of all other evidences which may come from different documents. Then we integrate the multi-view contextual information to encode the evidence sentences to handle the task. Our experimental results on FEVER 1.0 shared task show that our IMCI framework makes great progress on both fact extraction and verification, and achieves state-of-the-art performance with a winning FEVER score of 73.96% and label accuracy of 77.25% on the online blind test set. We also conduct ablation study to detect the impact of multi-view contextual information. Yangguang Li 0001, Zhen Huang 0006, Yong Dou |
COLING | 2 |
| 2022 | Towards Accurate Binary Neural Networks via Modeling Contextual Dependencies
Xingrun Xing, Yangguang Li 0001, Wei Li 0022, Wenrui Ding, Yalong Jiang, Yufeng Wang 0004, Chunlei Liu 0001, Xianglong Liu 0001 |
ECCV (11) | 2 |
| 2022 | R2F: A General Retrieval, Reading and Fusion Framework for Document-level Natural Language InferenceabstractDocument-level natural language inference (DOCNLI) is a new challenging task in natural language processing, aiming at judging the entailment relationship between a pair of hypothesis and premise documents.Current datasets and baselines largely follow sentence-level settings, but fail to address the issues raised by longer documents.In this paper, we establish a general solution, named Retrieval, Reading and Fusion (R 2 F) framework, and a new setting, by analyzing the main challenges of DOCNLI: interpretability, long-range dependency, and cross-sentence inference.The basic idea of the framework is to simplify documentlevel task into a set of sentence-level tasks, and improve both performance and interpretability with the power of evidence.For each hypothesis sentence, the framework retrieves evidence sentences from the premise, and reads to estimate its credibility.Then the sentencelevel results are fused to judge the relationship between the documents.For the setting, we contribute complementary evidence and entailment label annotation on hypothesis sentences, for interpretability study.Our experimental results show that R 2 F framework can obtain state-of-the-art performance and is robust for diverse evidence retrieval methods.Moreover, it can give more interpretable prediction results.Our model and code are released at https://github.com/phoenixsecularbird/R2F. Yixin Cao 0002, Yangguang Li 0001, Zhen Huang 0006, Kun Wang 0056 |
EMNLP | 3 |
| 2022 | Supervision Exists Everywhere: A Data Efficient Contrastive Language-Image Pre-training Paradigm
Yangguang Li 0001, Lichen Zhao, Yufeng Cui, Wanli Ouyang, Fengwei Yu |
ICLR | 1 |
| 2022 | RePre: Improving Self-Supervised Vision Transformer with Reconstructive Pre-trainingabstractRecently, self-supervised vision transformers have attracted unprecedented attention for their impressive representation learning ability. However, the dominant method, contrastive learning, mainly relies on an instance discrimination pretext task, which learns a global understanding of the image. This paper incorporates local feature learning into self-supervised vision transformers via Reconstructive Pre-training (RePre). Our RePre extends contrastive frameworks by adding a branch for reconstructing raw image pixels in parallel with the existing contrastive objective. RePre equips with a lightweight convolution-based decoder that fuses the multi-hierarchy features from the transformer encoder. The multi-hierarchy features provide rich supervisions from low to high semantic information, crucial for our RePre. Our RePre brings decent improvements on various contrastive frameworks with different vision transformer architectures. Transfer performance in downstream tasks outperforms supervised pre-training and state-of-the-art (SOTA) self-supervised counterparts. Luya Wang, Yangguang Li 0001, Wanli Ouyang |
IJCAI | 3 |
| 2022 | A Mixture Of Surprises for Unsupervised Reinforcement LearningabstractUnsupervised reinforcement learning aims at learning a generalist policy in a reward-free manner for fast adaptation to downstream tasks. Most of the existing methods propose to provide an intrinsic reward based on surprise. Maximizing or minimizing surprise drives the agent to either explore or gain control over its environment. However, both strategies rely on a strong assumption: the entropy of the environment's dynamics is either high or low. This assumption may not always hold in real-world scenarios, where the entropy of the environment's dynamics may be unknown. Hence, choosing between the two objectives is a dilemma. We propose a novel yet simple mixture of policies to address this concern, allowing us to optimize an objective that simultaneously maximizes and minimizes the surprise. Concretely, we train one mixture component whose objective is to maximize the surprise and another whose objective is to minimize the surprise. Hence, our method does not make assumptions about the entropy of the environment's dynamics. We call our method a $\textbf{M}\text{ixture }\textbf{O}\text{f }\textbf{S}\text{urprise}\textbf{S}$ (MOSS) for unsupervised reinforcement learning. Experimental results show that our simple method achieves state-of-the-art performance on the URLB benchmark, outperforming previous pure surprise maximization-based objectives. Our code is available at: https://github.com/LeapLabTHU/MOSS. Andrew Zhao, Matthieu Lin, Yangguang Li 0001, Yong-Jin Liu 0001, Gao Huang 0001 |
NeurIPS | 3 |
| 2017 | Depth map super-resolution via low-resolution depth guided joint trilateral up-sampling
Xin Jin 0002, Yangguang Li 0001, Chun Yuan 0003 |
J. Vis. Commun. Image Represent. | 3 |