VLDB 2026 Research / reviewers in the wild / expert
Yao Yao 0008
dblp:07/4410-8
· DBLP profile ↗
45ranked-venue papers
7as first author
33since 2021 · last 2026
0000-0001-9866-4291ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 42 · 6 first-author · 31 since 2021Graphics, computer vision, multimedia, augmented reality and games · 29 · 5 first-author · 17 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | BoYaEval: Evaluating Multimodal Large Language Models on Understanding Ancient Chinese Musical ScoresabstractJiajia Li, Weizhi Xue, Yao Yao, Qiwei Li, Chenchong, Zuchao Li, Ping Wang, Hai Zhao. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Jiajia Li 0005, Weizhi Xue, Yao Yao 0008, Qiwei Li 0002, Zuchao Li, Ping Wang 0028, Hai Zhao 0001 |
ACL (1) | 3 |
| 2026 | PAR: Training-Free Positional Perturbation and Attention Recycling for Faithful OCRabstractIn high-precision scenarios, vision language models suffer from Linguistic Priors Hallucination.When processing familiar text, models tend to over-rely on internal parametric knowledge, effectively "reciting" the content rather than "reading" the image.In this paper, we first systematically investigate this phenomenon by constructing the GlitchText Probing Dataset.We discover that the model's reliance on visual grounding diminishes significantly as the generation length increases.To mitigate this, we propose PAR (Positional Perturbation and Attention Recycling), a training-free, inferencetime intervention framework.PAR consists of two parts: (1) Positional Perturbation (PP) injects structured phase noise into the rotary positional embeddings; (2) Foveal Attention Recycling (FAR) detects over-confident linguistic priors and dynamically redistributes attention mass back to important visual regions.Extensive experiments across state-of-the-art models, demonstrate that PAR significantly reduces hallucination rates (reducing CER by 12%), particularly in long-context scenarios, while maintaining robust generalization on standard benchmarks.Our code is publicly available at https://github.com/Zoeyyao27/PAR-for- Faithful-OCR. Yao Yao 0008, Manwen Liao, Weitian Zhang, Zuchao Li, Hai Zhao 0001 |
ACL (1) | 1 |
| 2025 | 4D Diffusion for Dynamic Protein Structure Prediction with Reference and Motion GuidanceabstractProtein structure prediction is pivotal for understanding the structure-function relationship of proteins, advancing biological research, and facilitating pharmaceutical development and experimental design. While deep learning methods and the expanded availability of experimental 3D protein structures have accelerated structure prediction, the dynamic nature of protein structures has received limited attention. This study introduces an innovative 4D diffusion model incorporating molecular dynamics (MD) simulation data to learn dynamic protein structures. Our approach is distinguished by the following components: (1) a unified diffusion model capable of generating dynamic protein structures, including both the backbone and side chains, utilizing atomic grouping and side-chain dihedral angle predictions; (2) a reference network that enhances structural consistency by integrating the latent embeddings of the initial 3D protein structures; and (3) a motion alignment module aimed at improving temporal structural coherence across multiple time steps. To our knowledge, this is the first diffusion-based model aimed at predicting protein trajectories across multiple time steps simultaneously. Validation on benchmark datasets demonstrates that our model exhibits high accuracy in predicting dynamic 3D structures of proteins containing up to 256 amino acids over 32 time steps, effectively capturing both local flexibility in stable states and significant conformational changes. Kaihui Cheng, Ce Liu 0004, Qingkun Su, Yining Tang, Yao Yao 0008, Siyu Zhu 0001, Yuan Qi 0001 |
AAAI | 7 |
| 2025 | Caution for the Environment: Multimodal LLM Agents are Susceptible to Environmental DistractionsabstractXinbei Ma, Yiting Wang, Yao Yao, Tongxin Yuan, Aston Zhang, Zhuosheng Zhang, Hai Zhao. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Xinbei Ma, Yao Yao 0008, Tongxin Yuan, Aston Zhang, Zhuosheng Zhang 0001, Hai Zhao 0001 |
ACL (1) | 3 |
| 2025 | LESA: Learnable LLM Layer Scaling-UpabstractTraining Large Language Models (LLMs) from scratch requires immense computational resources, making it prohibitively expensive.Model scaling-up offers a promising solution by leveraging the parameters of smaller models to create larger ones.However, existing depth scaling-up methods rely on empirical heuristic rules for layer duplication, which result in poorer initialization and slower convergence during continual pre-training.We propose LESA, a novel learnable method for depth scaling-up.By concatenating parameters from each layer and applying Singular Value Decomposition, we uncover latent patterns between layers, suggesting that inter-layer parameters can be learned.LESA uses a neural network to predict the parameters inserted between adjacent layers, enabling better initialization and faster training.Experiments show that LESA outperforms existing baselines, achieving superior performance with less than half the computational cost during continual pre-training.Extensive analyses demonstrate its effectiveness across different model sizes and tasks. 1 Zouying Cao, Xinbei Ma, Yao Yao 0008, Zhi Chen 0006, Libo Qin 0001, Hai Zhao 0001 |
ACL (1) | 4 |
| 2025 | Mani-GS: Gaussian Splatting Manipulation with Triangular MeshabstractNeural 3D representations, such as Neural Radiation Fields (NeRF), excel at producing photorealistic rendering results but lack the flexibility for manipulation and editing which is crucial for content creation. However, manipulating NeRF is not highly controllable and requires a long training and inference time. With the emergence of 3D Gaussian Splatting (3DGS), extremely high-fidelity novel view synthesis can be achieved using an explicit point-based 3D representation with much faster training and rendering speed. However, there is still a lack of effective means to manipulate 3DGS freely while maintaining rendering quality. In this work, we aim to tackle the challenge of achieving manipulable photo-realistic rendering. We propose to utilize a triangular mesh to manipulate 3DGS directly with self-adaptation. This approach reduces the need to design various algorithms for different types of 3DGS manipulation. By utilizing a triangle shape-aware Gaussian binding and adapting method, we can achieve 3DGS manipulation and preserve high-fidelity rendering. In addition, our method is also effective with inaccurate meshes extracted from 3DGS. Experiments demonstrate our method’s effectiveness and superiority over baseline approaches. Xiangjun Gao, Xiaoyu Li 0002, Yiyu Zhuang, Qi Zhang 0029, Wenbo Hu 0002, Chaopeng Zhang, Yao Yao 0008, Ying Shan, Long Quan |
CVPR | 7 |
| 2025 | Matrix3D: Large Photogrammetry Model All-in-OneabstractWe present Matrix3D, a unified model that performs several photogrammetry subtasks, including pose estimation, depth prediction, and novel view synthesis using just the same model. Matrix3D utilizes a multi-modal diffusion transformer (DiT) to integrate transformations across several modalities, such as images, camera parameters, and depth maps. The key to Matrix3D’s large-scale multi-modal training lies in the incorporation of a mask learning strategy. This enables full-modality model training even with partially complete data, such as bi-modality data of image-pose and image-depth pairs, thus significantly increases the pool of available training data. Matrix3D demonstrates state-of-the-art performance in pose estimation and novel view synthesis tasks. Additionally, it offers fine-grained control through multi-round interactions, making it an innovative tool for 3D content creation. Project page: https://nju-3dv.github.io/projects/matrix3d. Yuanxun Lu, Jingyang Zhang, Tian Fang, Jean-Daniel Nahmias, Yanghai Tsin, Long Quan, Xun Cao, Yao Yao 0008, Shiwei Li 0001 |
CVPR | 8 |
| 2025 | FATE: Full-head Gaussian Avatar with Textural Editing from Monocular VideoabstractReconstructing high-fidelity, animatable 3D head avatars from effortlessly captured monocular videos is a pivotal yet formidable challenge. Although significant progress has been made in rendering performance and manipulation capabilities, notable challenges remain, including incomplete reconstruction and inefficient Gaussian representation. To address these challenges, we introduce FATE — a novel method for reconstructing an editable full-head avatar from a single monocular video. FATE integrates a sampling-based densification strategy to ensure optimal positional distribution of points, improving rendering efficiency. A neural baking technique is introduced to convert discrete Gaussian representations into continuous attribute maps, facilitating intuitive appearance editing. Furthermore, we propose a universal completion framework to recover non-frontal appearance, culminating in a 360° -renderable 3D head avatar. FATE outperforms previous approaches in both qualitative and quantitative evaluations, achieving state-of-the-art performance. To the best of our knowledge, FATE is the first animatable and 360° full-head monocular reconstruction method for a 3D head avatar. Project page and code are available at this link. Zhiyang Liang 0002, Dongfang Hu, Yao Yao 0008, Xun Cao, Hao Zhu 0004 |
CVPR | 6 |
| 2025 | XQuant: Achieving Ultra-Low Bit KV Cache Quantization with Cross-Layer CompressionabstractLarge Language Models (LLMs) have demonstrated remarkable capabilities across diverse natural language processing tasks.However, their extensive memory requirements, particularly due to KV cache growth during long-text understanding and generation, present significant challenges for deployment in resourceconstrained environments.Quantization has emerged as a promising solution to reduce memory consumption while preserving historical information.We propose XQuant, a training-free and plug-and-play framework that achieves ultra-low equivalent bit-width KV cache quantization.XQuant introduces two key innovations: a computationally negligible data-free calibration method and cross-layer KV cache compression, enabling quantization to sub-1.4 bits.Extensive experiments on TruthfulQA and LongBench demonstrate that XQuant outperforms state-of-the-art methods (e.g., KIVI-2bit and AsymKV-1.5bit)by achieving lower bit-width while maintaining superior performance, establishing a better trade-off between memory efficiency and model accuracy.The source code is available at https: //github.com/brinenick511/XQuant.KeyCache[l][0], 14 KeyCache[l][1], 15 KeyCache[l][2] 16 else 17 DequantizedKey ← Dequantize 18 KeyCache[l -1][0], 19 KeyCache[l -1][1], 20 KeyCache[l][2] 21 if l < vm or l mod 2 == 0 then 22 DequantizedValue ← Dequantize 23 ValueCache[l][0], 24 ValueCache[l][1], 25 ValueCache[l][2] 26 else 27 DequantizedValue ← Dequantize 28 ValueCache[l -1][0], 29 ValueCache[l -1][1], 30 ValueCache[l][2] Haoqi Yang 0001, Yao Yao 0008, Zuchao Li, Baoyuan Qi, Guoming Liu, Hai Zhao 0001 |
EMNLP | 2 |
| 2025 | Hallo2: Long-Duration and High-Resolution Audio-Driven Portrait Image AnimationabstractRecent advances in latent diffusion-based generative models for portrait image animation, such as Hallo, have achieved impressive results in short-duration video synthesis. In this paper, we present updates to Hallo, introducing several design enhancements to extend its capabilities.First, we extend the method to produce long-duration videos. To address substantial challenges such as appearance drift and temporal artifacts, we investigate augmentation strategies within the image space of conditional motion frames. Specifically, we introduce a patch-drop technique augmented with Gaussian noise to enhance visual consistency and temporal coherence over long duration.Second, we achieve 4K resolution portrait video generation. To accomplish this, we implement vector quantization of latent codes and apply temporal alignment techniques to maintain coherence across the temporal dimension. By integrating a high-quality decoder, we realize visual synthesis at 4K resolution.Third, we incorporate adjustable semantic textual labels for portrait expressions as conditional inputs. This extends beyond traditional audio cues to improve controllability and increase the diversity of the generated content. To the best of our knowledge, Hallo2, proposed in this paper, is the first method to achieve 4K resolution and generate hour-long, audio-driven portrait image animations enhanced with textual prompts. We have conducted extensive experiments to evaluate our method on publicly available datasets, including HDTF, CelebV, and our introduced ''Wild'' dataset. The experimental results demonstrate that our approach achieves state-of-the-art performance in long-duration portrait video animation, successfully generating rich and controllable content at 4K resolution for duration extending up to tens of minutes. Jiahao Cui 0003, Yao Yao 0008, Hao Zhu 0004, Hanlin Shang, Kaihui Cheng, Hang Zhou 0009, Siyu Zhu 0001, Jingdong Wang 0001 |
ICLR | 3 |
| 2025 | Flow Distillation Sampling: Regularizing 3D Gaussians with Pre-trained Matching Priorsabstract3D Gaussian Splatting (3DGS) has achieved excellent rendering quality with fast training and rendering speed. However, its optimization process lacks explicit geometric constraints, leading to suboptimal geometric reconstruction in regions with sparse or no observational input views. In this work, we try to mitigate the issue by incorporating a pre-trained matching prior to the 3DGS optimization process. We introduce Flow Distillation Sampling (FDS), a technique that leverages pre-trained geometric knowledge to bolster the accuracy of the Gaussian radiance field. Our method employs a strategic sampling technique to target unobserved views adjacent to the input views, utilizing the optical flow calculated from the matching model (Prior Flow) to guide the flow analytically calculated from the 3DGS geometry (Radiance Flow). Comprehensive experiments in depth rendering, mesh reconstruction, and novel view synthesis showcase the significant advantages of FDS over state-of-the-art methods. Additionally, our interpretive experiments and analysis aim to shed light on the effects of FDS on geometric accuracy and rendering quality, potentially providing readers with insights into its performance. Lin-Zhuo Chen, Kangjie Liu, Youtian Lin, Zhihao Li 0002, Siyu Zhu 0001, Xun Cao, Yao Yao 0008 |
ICLR | 7 |
| 2025 | Direct3D-S2: Gigascale 3D Generation Made Easy with Spatial Sparse AttentionabstractGenerating high-resolution 3D shapes using volumetric representations such as Signed Distance Functions (SDFs) presents substantial computational and memory challenges. We introduce Direct3D-S2, a scalable 3D generation framework based on sparse volumes that achieves superior output quality with dramatically reduced training costs.
Our key innovation is the Spatial Sparse Attention (SSA) mechanism, which greatly enhances the efficiency of Diffusion Transformer (DiT) computations on sparse volumetric data. SSA allows the model to effectively process large token sets within sparse volumes, significantly reducing computational overhead and achieving a 3.9$\times$ speedup in the forward pass and a 9.6$\times$ speedup in the backward pass.
Our framework also includes a variational autoencoder (VAE) that maintains a consistent sparse volumetric format across input, latent, and output stages. Compared to previous methods with heterogeneous representations in 3D VAE, this unified design significantly improves training efficiency and stability.
Our model is trained on public datasets, and experiments demonstrate that Direct3D-S2 not only surpasses state-of-the-art methods in generation quality and efficiency, but also enables training at 1024³ resolution using only 8 GPUs—a task typically requiring at least 32 GPUs for volumetric representations at $256^3$ resolution, thus making gigascale 3D generation both practical and accessible. Project page: https://www.neural4d.com/research-page/direct3d-s2. Youtian Lin, Feihu Zhang, Yifei Zeng, Yajie Bao, Jiachen Qian, Siyu Zhu 0001, Xun Cao, Philip Torr 0001, Yao Yao 0008 |
NeurIPS | 11 |
| 2025 | High-Fidelity Dynamic Portrait Animation via Direct Preference Optimization and Temporal Motion ModulationabstractGenerating highly dynamic and photorealistic portrait animations driven by audio and skeletal motion remains challenging due to the need for precise lip synchronization, natural facial expressions, and high-fidelity body motion dynamics. We propose a human-preference-aligned diffusion framework that addresses these challenges through two key innovations. First, we introduce direct preference optimization tailored for human-centric animation, leveraging a curated dataset of human preferences to align generated outputs with perceptual metrics for portrait motion-video alignment and naturalness of expression. Second, the proposed temporal motion modulation resolves spatiotemporal resolution mismatches by reshaping motion conditions into dimensionally aligned latent features through temporal channel redistribution and proportional feature expansion, preserving the fidelity of high-frequency motion details in diffusion-based synthesis. The proposed mechanism is complementary to existing UNet and DiT-based portrait diffusion approaches, and experiments demonstrate obvious improvements in lip-audio synchronization, expression vividness, body motion coherence over baseline methods, alongside notable gains in human preference metrics. Code and data for this paper are at https://github.com/fudan-generative-vision/hallo4. Jiahao Cui 0003, Baoyou Chen, Mingwang Xu, Hanlin Shang, Qinkun Su, Zilong Dong, Yao Yao 0008, Jingdong Wang 0001, Siyu Zhu 0001 |
SIGGRAPH Asia | 8 |
| 2024 | SirLLM: Streaming Infinite Retentive LLMabstractAs Large Language Models (LLMs) become increasingly prevalent in various domains, their ability to process inputs of any length and maintain a degree of memory becomes essential.However, the one-off input of overly long texts is limited, as studies have shown that when input lengths exceed the LLMs' pre-trained text length, there is a dramatic decline in text generation capabilities.Moreover, simply extending the length of pre-training texts is impractical due to the difficulty in obtaining long text data and the substantial memory consumption costs this would entail for LLMs.Recent efforts have employed streaming inputs to alleviate the pressure of excessively long text inputs, but this approach can significantly impair the model's long-term memory capabilities.Motivated by this challenge, we introduce Streaming Infinite Retentive LLM (SirLLM), which allows LLMs to maintain longer memory during infinite-length dialogues without the need for fine-tuning.SirLLM utilizes the Token Entropy metric and a memory decay mechanism to filter key phrases, endowing LLMs with both long-lasting and flexible memory.We designed three distinct tasks and constructed three datasets to measure the effectiveness of SirLLM from various angles: (1) DailyDialog; (2) Grocery Shopping; (3) Rock-Paper-Scissors.Our experimental results robustly demonstrate that SirLLM can achieve stable and significant improvements across different LLMs and tasks, compellingly proving its effectiveness.When having a coversation, "A sir could forget himself," but SirLLM never does!Our Yao Yao 0008, Zuchao Li, Hai Zhao 0001 |
ACL (1) | 1 |
| 2024 | Gaussian-Flow: 4D Reconstruction with Dynamic 3D Gaussian ParticleabstractWe introduce Gaussian-Flow, a novel point-based approach for fast dynamic scene reconstruction and real-time rendering from both multi-view and monocular videos. In contrast to the prevalent NeRF-based approaches hampered by slow training and rendering speeds, our approach harnesses recent advancements in point-based 3D Gaussian Splatting (3DGS). Specifically, a novel Dual-Domain Deformation Model (DDDM) is proposed to explicitly model attribute deformations of each Gaussian point, where the time-dependent residual of each attribute is captured by a polynomial fitting in the time domain, and a Fourier series fitting in the frequency domain. The proposed DDDM is capable of modeling complex scene deformations across long video footage, eliminating the need for training separate 3DGS for each frame or introducing an additional implicit neural field to model 3D dynamics. Moreover, the explicit deformation modeling for discretized Gaussian points ensures ultra-fast training and rendering of a 4D scene, which is comparable to the original 3DGS designed for static 3D reconstruction. Our proposed approach showcases a substantial efficiency improvement, achieving a 5 x faster training speed compared to the per-frame 3DGS modeling. In addition, quantitative results demonstrate that the proposed Gaussian-Flow significantly outperforms previous leading methods in novel view rendering quality. Project page: https://nju-3dv.github.io/projects/Gaussian-Flow. Youtian Lin, Zuozhuo Dai, Siyu Zhu 0001, Yao Yao 0008 |
CVPR | 4 |
| 2024 | Direct2.5: Diverse Text-to-3D Generation via Multi-view 2.5D DiffusionabstractRecent advances in generative AI have unveiled significant potential for the creation of 3D content. However, current methods either apply a pre-trained 2D diffusion model with the time-consuming score distillation sampling (SDS), or a direct 3D diffusion model trained on limited 3D data losing generation diversity. In this work, we approach the problem by employing a multi-view 2.5D diffusion fine-tuned from a pre-trained 2D diffusion model. The multi-view 2.5D diffusion directly models the structural distribution of 3D data, while still maintaining the strong generalization ability of the original 2D diffusion model, filling the gap between 2D diffusion-based and direct 3D diffusion-based methods for 3D content generation. During inference, multi-view normal maps are generated using the 2.5D diffusion, and a novel differentiable rasterization scheme is introduced to fuse the almost consistent multi-view normal maps into a consistent 3D model. We further design a normal-conditioned multi-view image generation module for fast appearance generation given the 3D geometry. Our method is a one-pass diffusion process and does not require any SDS optimization as post-processing. We demonstrate through extensive experiments that, our direct 2.5D generation with the specially-designed fusion scheme can achieve diverse, mode-seeking-free, and high-fidelity 3D content generation in only 10 seconds. Project page: https://nju-3dv.github.io/projects/direct25. Yuanxun Lu, Jingyang Zhang, Shiwei Li 0001, Tian Fang, David McKinnon, Yanghai Tsin, Long Quan, Xun Cao, Yao Yao 0008 |
CVPR | 9 |
| 2024 | Relightable 3D Gaussians: Realistic Point Cloud Relighting with BRDF Decomposition and Ray Tracing
Jian Gao 0009, Chun Gu, Youtian Lin, Zhihao Li 0002, Hao Zhu 0004, Xun Cao, Li Zhang 0040, Yao Yao 0008 |
ECCV (45) | 8 |
| 2024 | EmoTalk3D: High-Fidelity Free-View Synthesis of Emotional 3D Talking Head
Qianyun He, Xinya Ji, Yuanxun Lu, Zhengyu Diao, Linjia Huang, Yao Yao 0008, Siyu Zhu 0001, Zhan Ma 0001, Songcen Xu, Zixiao Zhang, Xun Cao, Hao Zhu 0004 |
ECCV (57) | 7 |
| 2024 | Head360: Learning a Parametric 3D Full-Head for Free-View Synthesis in 360$^\circ $
Yuxiao He, Yiyu Zhuang, Yao Yao 0008, Siyu Zhu 0001, Xiaoyu Li 0002, Qi Zhang 0029, Xun Cao, Hao Zhu 0004 |
ECCV (56) | 4 |
| 2024 | STAG4D: Spatial-Temporal Anchored Generative 4D Gaussians
Yifei Zeng, Yanqin Jiang, Siyu Zhu 0001, Yuanxun Lu, Youtian Lin, Hao Zhu 0004, Weiming Hu 0004, Xun Cao, Yao Yao 0008 |
ECCV (36) | 9 |
| 2024 | Champ: Controllable and Consistent Human Image Animation with 3D Parametric Guidance
Shenhao Zhu, Junming Leo Chen, Zuozhuo Dai, Zilong Dong, Xun Cao, Yao Yao 0008, Hao Zhu 0004, Siyu Zhu 0001 |
ECCV (55) | 7 |
| 2024 | Consistent4D: Consistent 360° Dynamic Object Generation from Monocular VideoabstractIn this paper, we present Consistent4D, a novel approach for generating 4D dynamic objects from uncalibrated monocular videos. Uniquely, we cast the 360-degree dynamic object reconstruction as a 4D generation problem, eliminating the need for tedious multi-view data collection and camera calibration. This is achieved by leveraging the object-level 3D-aware image diffusion model as the primary supervision signal for training dynamic Neural Radiance Fields (DyNeRF). Specifically, we propose a cascade DyNeRF to facilitate stable convergence and temporal continuity under the time-discrete supervision signal. To achieve spatial and temporal consistency of the 4D generation, an interpolation-driven consistency loss is further introduced, which aligns the rendered frames with the interpolated frames from a pre-trained video interpolation model. Extensive experiments show that the proposed Consistent4D significantly outperforms previous 4D reconstruction approaches as well as per-frame 3D generation approaches, opening up new possibilities for 4D dynamic object generation from a single-view uncalibrated video. Project page: https://consistent4d.github.io Yanqin Jiang, Li Zhang 0040, Weiming Hu 0004, Yao Yao 0008 |
ICLR | 5 |
| 2024 | JointNet: Extending Text-to-Image Diffusion for Dense Distribution ModelingabstractWe introduce JointNet, a novel neural network architecture for modeling the joint distribution of images and an additional dense modality (e.g., depth maps).
JointNet is extended from a pre-trained text-to-image diffusion model, where a copy of the original network is created for the new dense modality branch and is densely connected with the RGB branch.
The RGB branch is locked during network fine-tuning, which enables efficient learning of the new modality distribution while maintaining the strong generalization ability of the large-scale pre-trained diffusion model.
We demonstrate the effectiveness of JointNet by using the RGB-D diffusion as an example and through extensive experiments, showcasing its applicability in a variety of applications, including joint RGB-D generation, dense depth prediction, depth-conditioned image generation, and high-resolution 3D panorama generation. Jingyang Zhang, Shiwei Li 0001, Yuanxun Lu, Tian Fang, David McKinnon, Yanghai Tsin, Long Quan, Yao Yao 0008 |
ICLR | 8 |
| 2024 | Stereo Risk: A Continuous Modeling Approach to Stereo MatchingabstractWe introduce Stereo Risk, a new deep-learning approach to solve the classical stereo-matching problem in computer vision. As it is well-known that stereo matching boils down to a per-pixel disparity estimation problem, the popular state-of-the-art stereo-matching approaches widely rely on regressing the scene disparity values, yet via discretization of scene disparity values. Such discretization often fails to capture the nuanced, continuous nature of scene depth. Stereo Risk departs from the conventional discretization approach by formulating the scene disparity as an optimal solution to a continuous risk minimization problem, hence the name "stereo risk". We demonstrate that $L^1$ minimization of the proposed continuous risk function enhances stereo-matching performance for deep networks, particularly for disparities with multi-modal probability distributions. Furthermore, to enable the end-to-end network training of the non-differentiable $L^1$ risk optimization, we exploited the implicit function theorem, ensuring a fully differentiable network. A comprehensive analysis demonstrates our method's theoretical soundness and superior performance over the state-of-the-art methods across various benchmark datasets, including KITTI 2012, KITTI 2015, ETH3D, SceneFlow, and Middlebury 2014. Ce Liu 0004, Suryansh Kumar 0001, Shuhang Gu, Radu Timofte, Yao Yao 0008, Luc Van Gool |
ICML | 5 |
| 2024 | GaussianPro: 3D Gaussian Splatting with Progressive Propagationabstract3D Gaussian Splatting (3DGS) has recently revolutionized the field of neural rendering with its high fidelity and efficiency. However, 3DGS heavily depends on the initialized point cloud produced by Structure-from-Motion (SfM) techniques. When tackling large-scale scenes that unavoidably contain texture-less surfaces, SfM techniques fail to produce enough points in these surfaces and cannot provide good initialization for 3DGS. As a result, 3DGS suffers from difficult optimization and low-quality renderings. In this paper, inspired by classic multi-view stereo (MVS) techniques, we propose GaussianPro, a novel method that applies a progressive propagation strategy to guide the densification of the 3D Gaussians. Compared to the simple split and clone strategies used in 3DGS, our method leverages the priors of the existing reconstructed geometries of the scene and utilizes patch matching to produce new Gaussians with accurate positions and orientations. Experiments on both large-scale and small-scale scenes validate the effectiveness of our method. Our method significantly surpasses 3DGS on the Waymo dataset, exhibiting an improvement of 1.15dB in terms of PSNR. Codes and data are available at https://github.com/kcheng1021/GaussianPro. Xiaoxiao Long, Kaizhi Yang, Yao Yao 0008, Wei Yin 0006, Yuexin Ma, Wenping Wang 0001, Xuejin Chen |
ICML | 4 |
| 2024 | Reference Trustable Decoding: A Training-Free Augmentation Paradigm for Large Language ModelsabstractLarge language models (LLMs) have rapidly advanced and demonstrated impressive capabilities. In-Context Learning (ICL) and Parameter-Efficient Fine-Tuning (PEFT) are currently two mainstream methods for augmenting LLMs to downstream tasks. ICL typically constructs a few-shot learning scenario, either manually or by setting up a Retrieval-Augmented Generation (RAG) system, helping models quickly grasp domain knowledge or question-answering patterns without changing model parameters. However, this approach involves trade-offs, such as slower inference speed and increased space occupancy. PEFT assists the model in adapting to tasks through minimal parameter modifications, but the training process still demands high hardware requirements, even with a small number of parameters involved. To address these challenges, we propose Reference Trustable Decoding (RTD), a paradigm that allows models to quickly adapt to new tasks without fine-tuning, maintaining low inference costs. RTD constructs a reference datastore from the provided training examples and optimizes the LLM's final vocabulary distribution by flexibly selecting suitable references based on the input, resulting in more trustable responses and enabling the model to adapt to downstream tasks at a low cost. Experimental evaluations on various LLMs using different benchmarks demonstrate that RTD establishes a new paradigm for augmenting models to downstream tasks. Furthermore, our method exhibits strong orthogonality with traditional methods, allowing for concurrent usage. Our code can be found at https://github.com/ShiLuohe/ReferenceTrustableDecoding. Luohe Shi, Yao Yao 0008, Zuchao Li, Lefei Zhang, Hai Zhao 0001 |
NeurIPS | 2 |
| 2024 | Direct3D: Scalable Image-to-3D Generation via 3D Latent Diffusion TransformerabstractGenerating high-quality 3D assets from text and images has long been challenging, primarily due to the absence of scalable 3D representations capable of capturing intricate geometry distributions. In this work, we introduce Direct3D, a native 3D generative model scalable to in-the-wild input images, without requiring a multi-view diffusion model or SDS optimization. Our approach comprises two primary components: a Direct 3D Variational Auto-Encoder (D3D-VAE) and a Direct 3D Diffusion Transformer (D3D-DiT). D3D-VAE efficiently encodes high-resolution 3D shapes into a compact and continuous latent triplane space. Notably, our method directly supervises the decoded geometry using a semi-continuous surface sampling strategy, diverging from previous methods relying on rendered images as supervision signals. D3D-DiT models the distribution of encoded 3D latents and is specifically designed to fuse positional information from the three feature maps of the triplane latent, enabling a native 3D generative model scalable to large-scale 3D datasets. Additionally, we introduce an innovative image-to-3D generation pipeline incorporating semantic and pixel-level image conditions, allowing the model to produce 3D shapes consistent with the provided conditional image input. Extensive experiments demonstrate the superiority of our large-scale pre-trained Direct3D over previous image-to-3D approaches, achieving significantly better generation quality and generalization ability, thus establishing a new state-of-the-art for 3D content creation. Project page: https://www.neural4d.com/research/direct3d. Youtian Lin, Yifei Zeng, Feihu Zhang, Jingxi Xu 0001, Philip Torr 0001, Xun Cao, Yao Yao 0008 |
NeurIPS | 8 |
| 2023 | NeILF++: Inter-Reflectable Light Fields for Geometry and Material EstimationabstractWe present a novel differentiable rendering framework for joint geometry, material, and lighting estimation from multi-view images. In contrast to previous methods which assume a simplified environment map or co-located flashlights, in this work, we formulate the lighting of a static scene as one neural incident light field (NeILF) and one outgoing neural radiance field (NeRF). The key insight of the proposed method is the union of the incident and outgoing light fields through physically-based rendering and inter-reflections between surfaces, making it possible to disentangle the scene geometry, material, and lighting from image observations in a physically-based manner. The proposed incident light and inter-reflection framework can be easily applied to other NeRF systems. We show that our method can not only decompose the outgoing radiance into incident lights and surface materials, but also serve as a surface refinement module that further improves the reconstruction detail of the neural surface. We demonstrate on several datasets that the proposed method is able to achieve state-of-the-art results in terms of geometry reconstruction quality, material estimation accuracy, and the fidelity of novel view rendering. Jingyang Zhang, Yao Yao 0008, Shiwei Li 0001, Tian Fang, David McKinnon, Yanghai Tsin, Long Quan |
ICCV | 2 |
| 2023 | Anti-Aliased Neural Implicit Surfaces with Encoding Level of DetailabstractWe present LoD-NeuS, an efficient neural representation for high-frequency geometry detail recovery and anti-aliased novel view rendering. Drawing inspiration from voxel-based representations with the level of detail (LoD), we introduce a multi-scale tri-plane-based scene representation that is capable of capturing the LoD of the signed distance function (SDF) and the space radiance. Our representation aggregates space features from a multi-convolved featurization within a conical frustum along a ray and optimizes the LoD feature volume through differentiable rendering. Additionally, we propose an error-guided sampling strategy to guide the growth of the SDF during the optimization. Both qualitative and quantitative evaluations demonstrate that our method achieves superior surface reconstruction and photorealistic view synthesis compared to state-of-the-art approaches. Yiyu Zhuang, Qi Zhang 0029, Hao Zhu 0004, Yao Yao 0008, Xiaoyu Li 0002, Yan-Pei Cao 0001, Ying Shan, Xun Cao |
SIGGRAPH Asia | 5 |
| 2023 | Vis-MVSNet: Visibility-Aware Multi-view Stereo Network
Jingyang Zhang, Shiwei Li 0001, Zixin Luo, Tian Fang, Yao Yao 0008 |
Int. J. Comput. Vis. | 5 |
| 2022 | Critical Regularizations for Neural Surface Reconstruction in the WildabstractNeural implicit functions have recently shown promising results on surface reconstructions from multiple views. However, current methods still suffer from excessive time complexity and poor robustness when reconstructing unbounded or complex scenes. In this paper, we present RegSDF, which shows that proper point cloud supervisions and geometry regularizations are sufficient to produce high-quality and robust reconstruction results. Specifically, RegSDF takes an additional oriented point cloud as input, and optimizes a signed distance field and a surface light field within a differentiable rendering framework. We also introduce the two critical regularizations for this optimization. The first one is the Hessian regularization that smoothly diffuses the signed distance values to the entire distance field given noisy and incomplete input. And the second one is the minimal surface regularization that compactly interpolates and extrapolates the missing geometry. Extensive experiments are conducted on DTU, Blended-MVS, and Tanks and Temples datasets. Compared with recent neural surface reconstruction approaches, RegSDF is able to reconstruct surfaces with fine details even for open scenes with complex topologies and unstructured camera trajectories. Jingyang Zhang, Yao Yao 0008, Shiwei Li 0001, Tian Fang, David McKinnon, Yanghai Tsin, Long Quan |
CVPR | 2 |
| 2022 | NeILF: Neural Incident Light Field for Physically-based Material Estimation
Yao Yao 0008, Jingyang Zhang, Yihang Qu, Tian Fang, David McKinnon, Yanghai Tsin, Long Quan |
ECCV (31) | 1 |
| 2021 | Learning Signed Distance Field for Multi-view Surface ReconstructionabstractRecent works on implicit neural representations have shown promising results for multi-view surface reconstruction. However, most approaches are limited to relatively simple geometries and usually require clean object masks for reconstructing complex and concave objects. In this work, we introduce a novel neural surface reconstruction framework that leverages the knowledge of stereo matching and feature consistency to optimize the implicit surface representation. More specifically, we apply a signed distance field (SDF) and a surface light field to represent the scene geometry and appearance respectively. The SDF is directly supervised by geometry from stereo matching, and is refined by optimizing the multi-view feature consistency and the fidelity of rendered images. Our method is able to improve the robustness of geometry estimation and support reconstruction of complex scene topologies. Extensive experiments have been conducted on DTU, EPFL and Tanks and Temples datasets. Compared to previous state-of-the-art methods, our method achieves better mesh reconstruction in wide open scenes without masks as input. Jingyang Zhang, Yao Yao 0008, Long Quan |
ICCV | 2 |
| 2020 | Visibility-aware Multi-view Stereo Network
Jingyang Zhang, Yao Yao 0008, Shiwei Li 0001, Zixin Luo, Tian Fang |
BMVC | 2 |
| 2020 | BlendedMVS: A Large-Scale Dataset for Generalized Multi-View Stereo NetworksabstractWhile deep learning has recently achieved great success on multi-view stereo (MVS), limited training data makes the trained model hard to be generalized to unseen scenarios. Compared with other computer vision tasks, it is rather difficult to collect a large-scale MVS dataset as it requires expensive active scanners and labor-intensive process to obtain ground truth 3D structures. In this paper, we introduce BlendedMVS, a novel large-scale dataset, to provide sufficient training ground truth for learning-based MVS. To create the dataset, we apply a 3D reconstruction pipeline to recover high-quality textured meshes from images of well-selected scenes. Then, we render these mesh models to color images and depth maps. To introduce the ambient lighting information during training, the rendered color images are further blended with the input images to generate the training input. Our dataset contains over 17k high-resolution images covering a variety of scenes, including cities, architectures, sculptures and small objects. Extensive experiments demonstrate that BlendedMVS endows the trained model with significantly better generalization ability compared with other MVS datasets. The dataset and pretrained models are available at https://github.com/YoYo000/BlendedMVS. Yao Yao 0008, Zixin Luo, Shiwei Li 0001, Jingyang Zhang, Yufan Ren, Lei Zhou 0011, Tian Fang, Long Quan |
CVPR | 1 |
| 2020 | ASLFeat: Learning Local Features of Accurate Shape and LocalizationabstractThis work focuses on mitigating two limitations in the joint learning of local feature detectors and descriptors. First, the ability to estimate the local shape (scale, orientation, etc.) of feature points is often neglected during dense feature extraction, while the shape-awareness is crucial to acquire stronger geometric invariance. Second, the localization accuracy of detected keypoints is not sufficient to reliably recover camera geometry, which has become the bottleneck in tasks such as 3D reconstruction. In this paper, we present ASLFeat, with three light-weight yet effective modifications to mitigate above issues. First, we resort to deformable convolutional networks to densely estimate and apply local transformation. Second, we take advantage of the inherent feature hierarchy to restore spatial resolution and low-level details for accurate keypoint localization. Finally, we use a peakiness measurement to relate feature responses and derive more indicative detection scores. The effect of each modification is thoroughly studied, and the evaluation is extensively conducted across a variety of practical scenarios. State-of-the-art results are reported that demonstrate the superiority of our methods. Zixin Luo, Lei Zhou 0011, Xuyang Bai, Yao Yao 0008, Shiwei Li 0001, Tian Fang, Long Quan |
CVPR | 6 |
| 2020 | KFNet: Learning Temporal Camera Relocalization Using Kalman FilteringabstractTemporal camera relocalization estimates the pose with respect to each video frame in sequence, as opposed to one-shot relocalization which focuses on a still image. Even though the time dependency has been taken into account, current temporal relocalization methods still generally underperform the state-of-the-art one-shot approaches in terms of accuracy. In this work, we improve the temporal relocalization method by using a network architecture that incorporates Kalman filtering (KFNet) for online camera relocalization. In particular, KFNet extends the scene coordinate regression problem to the time domain in order to recursively establish 2D and 3D correspondences for the pose determination. The network architecture design and the loss formulation are based on Kalman filtering in the context of Bayesian learning. Extensive experiments on multiple relocalization benchmarks demonstrate the high accuracy of KFNet at the top of both one-shot and temporal relocalization approaches. Lei Zhou 0011, Zixin Luo, Tianwei Shen, Mingmin Zhen, Yao Yao 0008, Tian Fang, Long Quan |
CVPR | 6 |
| 2020 | Learning Stereo Matchability in Disparity Regression NetworksabstractLearning-based stereo matching has recently achieved promising results, yet still suffers difficulties in establishing reliable matches in weakly matchable regions that are textureless, non-Lambertian, or occluded. In this paper, we address this challenge by proposing a stereo matching network that considers pixel-wise matchability. Specifically, the network jointly regresses disparity and matchability maps from 3D probability volume through expectation and entropy operations. Next, a learned attenuation is applied as the robust loss function to alleviate the influence of weakly matchable pixels in the training. Finally, a matchability-aware disparity refinement is introduced to improve the depth inference in weakly matchable regions. The proposed deep stereo matchability (DSM) framework can improve the matching result or accelerate the computation while still guaranteeing the quality. Moreover, the DSM framework is portable to many recent stereo networks. Extensive experiments are conducted on Scene Flow and KITTI stereo datasets to demonstrate the effectiveness of the proposed framework over the state-of-the-art learning-based stereo methods. Jingyang Zhang, Yao Yao 0008, Zixin Luo, Shiwei Li 0001, Tianwei Shen, Tian Fang, Long Quan |
ICPR | 2 |
| 2019 | Recurrent MVSNet for High-Resolution Multi-View Stereo Depth InferenceabstractDeep learning has recently demonstrated its excellent performance for multi-view stereo (MVS). However, one major limitation of current learned MVS approaches is the scalability: the memory-consuming cost volume regularization makes the learned MVS hard to be applied to high-resolution scenes. In this paper, we introduce a scalable multi-view stereo framework based on the recurrent neural network. Instead of regularizing the entire 3D cost volume in one go, the proposed Recurrent Multi-view Stereo Network (R-MVSNet) sequentially regularizes the 2D cost maps along the depth direction via the gated recurrent unit (GRU). This reduces dramatically the memory consumption and makes high-resolution reconstruction feasible. We first show the state-of-the-art performance achieved by the proposed R-MVSNet on the recent MVS benchmarks. Then, we further demonstrate the scalability of the proposed method on several large-scale scenarios, where previous learned approaches often fail due to the memory constraint. Code is available at https://github.com/YoYo000/MVSNet. Yao Yao 0008, Zixin Luo, Shiwei Li 0001, Tianwei Shen, Tian Fang, Long Quan |
CVPR | 1 |
| 2019 | Cross-Atlas Convolution for Parameterization Invariant Learning on Textured Mesh SurfaceabstractWe present a convolutional network architecture for direct feature learning on mesh surfaces through their atlases of texture maps. The texture map encodes the parameterization from 3D to 2D domain, rendering not only RGB values but also rasterized geometric features if necessary. Since the parameterization of texture map is not pre-determined, and depends on the surface topologies, we therefore introduce a novel cross-atlas convolution to recover the original mesh geodesic neighborhood, so as to achieve the invariance property to arbitrary parameterization. The proposed module is integrated into classification and segmentation architectures, which takes the input texture map of a mesh, and infers the output predictions. Our method not only shows competitive performances on classification and segmentation public benchmarks, but also paves the way for the broad mesh surfaces learning. Shiwei Li 0001, Zixin Luo, Mingmin Zhen, Yao Yao 0008, Tianwei Shen, Tian Fang, Long Quan |
CVPR | 4 |
| 2019 | ContextDesc: Local Descriptor Augmentation With Cross-Modality ContextabstractMost existing studies on learning local features focus on the patch-based descriptions of individual keypoints, whereas neglecting the spatial relations established from their keypoint locations. In this paper, we go beyond the local detail representation by introducing context awareness to augment off-the-shelf local feature descriptors. Specifically, we propose a unified learning framework that leverages and aggregates the cross-modality contextual information, including (i) visual context from high-level image representation, and (ii) geometric context from 2D keypoint distribution. Moreover, we propose an effective N-pair loss that eschews the empirical hyper-parameter search and improves the convergence. The proposed augmentation scheme is lightweight compared with the raw local feature description, meanwhile improves remarkably on several large-scale benchmarks with diversified scenes, which demonstrates both strong practicality and generalization ability in geometric matching applications. Zixin Luo, Tianwei Shen, Lei Zhou 0011, Yao Yao 0008, Shiwei Li 0001, Tian Fang, Long Quan |
CVPR | 5 |
| 2018 | Reconstructing Thin Structures of Manifold Surfaces by Integrating Spatial CurvesabstractThe manifold surface reconstruction in multi-view stereo often fails in retaining thin structures due to incomplete and noisy reconstructed point clouds. In this paper, we address this problem by leveraging spatial curves. The curve representation in nature is advantageous in modeling thin and elongated structures, implying topology and connectivity information of the underlying geometry, which exactly compensates the weakness of scattered point clouds. We present a novel surface reconstruction method using both curves and point clouds. First, we propose a 3D curve reconstruction algorithm based on the initialize-optimize-extend strategy. Then, tetrahedra are constructed from points and curves, where the volumes of thin structures are robustly preserved by the Curve-conformed Delaunay Refinement. Finally, the mesh surface is extracted from tetrahedra by a graph optimization. The method has been intensively evaluated on both synthetic and real-world datasets, showing significant improvements over state-of-the-art methods. Shiwei Li 0001, Yao Yao 0008, Tian Fang, Long Quan |
CVPR | 2 |
| 2018 | GeoDesc: Learning Local Descriptors by Integrating Geometry Constraints
Zixin Luo, Tianwei Shen, Lei Zhou 0011, Siyu Zhu 0001, Yao Yao 0008, Tian Fang, Long Quan |
ECCV (9) | 6 |
| 2018 | MVSNet: Depth Inference for Unstructured Multi-view Stereo
Yao Yao 0008, Zixin Luo, Shiwei Li 0001, Tian Fang, Long Quan |
ECCV (8) | 1 |
| 2017 | Relative Camera Refinement for Accurate Dense ReconstructionabstractMulti-view stereo (MVS) depends on the pre-determined camera geometry, often from structure from motion (SfM) or simultaneous localization and mapping (SLAM). However, cameras may not be locally optimal for dense stereo matching, especially when it comes from the large scale SfM or the SLAM with multiple sensor fusion. In this paper, we propose a local camera refinement approach for accurate dense reconstruction. Firstly, we refines the relative geometry of independent camera pair using a tailored bundle adjustment. The refinement is also extended to a multi-view version for general MVS reconstructions. Then, the non-rigid dense alignment is formulated as an inverse-distortion problem to transfer point clouds from each local coordinate system to a global coordinate system. The proposed framework has been intensively validated in both SfM and SLAM based dense reconstructions. Results on different datasets show that our method can significantly improve the dense reconstruction quality. Yao Yao 0008, Shiwei Li 0001, Siyu Zhu 0001, Hanyu Deng, Tian Fang, Long Quan |
3DV | 1 |