EDBT 2026 Demo / reviewers in the wild / expert
Jiawei Yang 0002
dblp:96/2976-2
· DBLP profile ↗
20ranked-venue papers
8as first author
18since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 13 · 6 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 10 · 5 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 2 first-author · 5 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | InfiniCube: Unbounded and Controllable Dynamic 3D Driving Scene Generation with World-Guided Video ModelsabstractWe present InfiniCube, a scalable method for generating unbounded dynamic 3D driving scenes with high fidelity and controllability. Previous methods for scene generation either suffer from limited scales or lack geometric and appearance consistency along generated sequences. In contrast, we leverage the recent advancements in scalable 3D representation and video models to achieve large dynamic scene generation that allows flexible controls through HD maps, vehicle bounding boxes, and text descriptions. First, we construct a map-conditioned sparse-voxel-based 3D generative model to unleash its power for unbounded voxel world generation. Then, we re-purpose a video model and ground it on the voxel world through a set of carefully designed pixel-aligned guidance buffers, synthesizing a consistent appearance. Finally, we propose a fast feed-forward approach that employs both voxel and pixel branches to lift the dynamic videos to dynamic 3D Gaussians with controllable objects. Our method can generate controllable and realistic 3D driving scenes, and extensive experiments validate the effectiveness and superiority of our model. Xuanchi Ren, Jiawei Yang 0002, Tianchang Shen, Jay Zhangjie Wu, Jun Gao 0004, Yue Wang 0041, Siheng Chen, Sanja Fidler |
ICCV | 3 |
| 2025 | STORM: Spatio-TempOral Reconstruction Model For Large-Scale Outdoor ScenesabstractWe present STORM, a spatio-temporal reconstruction model designed for reconstructing dynamic outdoor scenes from sparse observations. Existing dynamic reconstruction methods often rely on per-scene optimization, dense observations across space and time, and strong motion supervision, resulting in lengthy optimization times, limited generalization to novel views or scenes, and degenerated quality caused by noisy pseudo-labels for dynamics. To address these challenges, STORM leverages a data-driven Transformer architecture that directly infers dynamic 3D scene representations—parameterized by 3D Gaussians and their velocities—in a single forward pass. Our key design is to aggregate 3D Gaussians from all frames using self-supervised scene flows, transforming them to the target timestep to enable complete (i.e., "amodal") reconstructions from arbitrary viewpoints at any moment in time. As an emergent property, STORM automatically captures dynamic instances and generates high-quality masks using only reconstruction losses. Extensive experiments on public datasets show that STORM achieves precise dynamic scene reconstruction, surpassing state-of-the-art per-scene optimization methods (+4.3 to 6.6 PSNR) and existing feed-forward approaches (+2.1 to 4.7 PSNR) in dynamic regions. STORM reconstructs large-scale outdoor scenes in 200ms, supports real-time rendering, and outperforms competitors in scene flow estimation, improving 3D EPE by 0.422m and Acc5 by 28.02%. Beyond reconstruction, we showcase four additional applications of our model, illustrating the potential of self-supervised learning for broader dynamic scene understanding. For more details, please visit our project at https://jiawei-yang.github.io/STORM/. Jiawei Yang 0002, Boris Ivanovic, Yuxiao Chen 0008, Yan Wang 0051, Boyi Li 0001, Yurong You, Apoorva Sharma, Maximilian Igl, Péter Karkus, Danfei Xu, Yue Wang 0041, Marco Pavone 0001 |
ICLR | 1 |
| 2025 | OmniRe: Omni Urban Scene ReconstructionabstractWe introduce OmniRe, a comprehensive system for efficiently creating high-fidelity digital twins of dynamic real-world scenes from on-device logs. Recent methods using neural fields or Gaussian Splatting primarily focus on vehicles, hindering a holistic framework for all dynamic foregrounds demanded by downstream applications, e.g., the simulation of human behavior. OmniRe extends beyond vehicle modeling to enable accurate, full-length reconstruction of diverse dynamic objects in urban scenes. Our approach builds scene graphs on 3DGS and constructs multiple Gaussian representations in canonical spaces that model various dynamic actors, including vehicles, pedestrians, cyclists, and others. OmniRe allows holistically reconstructing any dynamic object in the scene, enabling advanced simulations (~60 Hz) that include human-participated scenarios, such as pedestrian behavior simulation and human-vehicle interaction. This comprehensive simulation capability is unmatched by existing methods. Extensive evaluations on the Waymo dataset show that our approach outperforms prior state-of-the-art methods quantitatively and qualitatively by a large margin. We further extend our results to 5 additional popular driving datasets to demonstrate its generalizability on common urban scenes. Code and results are available at [omnire](https://ziyc.github.io/omnire/). Jiawei Yang 0002, Riccardo de Lutio, Janick Martinez Esturo, Boris Ivanovic, Or Litany, Zan Gojcic, Sanja Fidler, Marco Pavone 0001, Yue Wang 0041 |
ICLR | 2 |
| 2024 | Denoising Vision Transformers
Jiawei Yang 0002, Katie Luo, Congyue Deng, Leonidas J. Guibas, Dilip Krishnan, Kilian Q. Weinberger, Yonglong Tian, Yue Wang 0041 |
ECCV (85) | 1 |
| 2024 | EmerNeRF: Emergent Spatial-Temporal Scene Decomposition via Self-SupervisionabstractWe present EmerNeRF, a simple yet powerful approach for learning spatial-temporal representations of dynamic driving scenes. Grounded in neural fields, EmerNeRF simultaneously captures scene geometry, appearance, motion, and semantics via self-bootstrapping. EmerNeRF hinges upon two core components: First, it stratifies scenes into static and dynamic fields. This decomposition emerges purely from self-supervision, enabling our model to learn from general, in-the-wild data sources. Second, EmerNeRF parameterizes an induced flow field from the dynamic field and uses this flow field to further aggregate multi-frame features, amplifying the rendering precision of dynamic objects. Coupling these three fields (static, dynamic, and flow) enables EmerNeRF to represent highly-dynamic scenes self-sufficiently, without relying on ground truth object annotations or pre-trained models for dynamic object segmentation or optical flow estimation. Our method achieves state-of-the-art performance in sensor simulation, significantly outperforming previous methods when reconstructing static (+2.93 PSNR) and dynamic (+3.70 PSNR) scenes. In addition, to bolster EmerNeRF's semantic generalization, we lift 2D visual foundation model features into 4D space-time and address a general positional bias in modern Transformers, significantly boosting 3D perception performance (e.g., 37.50% relative improvement in occupancy prediction accuracy on average). Finally, we construct a diverse and challenging 120-sequence dataset to benchmark neural fields under extreme and highly-dynamic settings. See the project page for code, data, and request pre-trained models: https://emernerf.github.io Jiawei Yang 0002, Boris Ivanovic, Or Litany, Xinshuo Weng, Seung Wook Kim 0001, Boyi Li 0001, Tong Che, Danfei Xu, Sanja Fidler, Marco Pavone 0001, Yue Wang 0041 |
ICLR | 1 |
| 2024 | Parallelized Spatiotemporal Slot Binding for VideosabstractWhile modern best practices advocate for scalable architectures that support long-range interactions, object-centric models are yet to fully embrace these architectures. In particular, existing object-centric models for handling sequential inputs, due to their reliance on RNN-based implementation, show poor stability and capacity and are slow to train on long sequences. We introduce Parallelizable Spatiotemporal Binder or PSB, the first temporally-parallelizable slot learning architecture for sequential inputs. Unlike conventional RNN-based approaches, PSB produces object-centric representations, known as slots, for all time-steps in parallel. This is achieved by refining the initial slots across all time-steps through a fixed number of layers equipped with causal attention. By capitalizing on the parallelism induced by our architecture, the proposed model exhibits a significant boost in efficiency. In experiments, we test PSB extensively as an encoder within an auto-encoding framework paired with a wide variety of decoder options. Compared to the state-of-the-art, our architecture demonstrates stable training on longer sequences, achieves parallelization that results in a 60% increase in training speed, and yields performance that is on par with or better on unsupervised 2D and 3D object-centric scene decomposition and understanding. Gautam Singh, Yue Wang 0041, Jiawei Yang 0002, Boris Ivanovic, Sungjin Ahn, Marco Pavone 0001, Tong Che |
ICML | 3 |
| 2024 | DistillNeRF: Perceiving 3D Scenes from Single-Glance Images by Distilling Neural Fields and Foundation Model FeaturesabstractWe propose DistillNeRF, a self-supervised learning framework addressing the challenge of understanding 3D environments from limited 2D observations in outdoor autonomous driving scenes. Our method is a generalizable feedforward model that predicts a rich neural scene representation from sparse, single-frame multi-view camera inputs with limited view overlap, and is trained self-supervised with differentiable rendering to reconstruct RGB, depth, or feature images. Our first insight is to exploit per-scene optimized Neural Radiance Fields (NeRFs) by generating dense depth and virtual camera targets from them, which helps our model to learn enhanced 3D geometry from sparse non-overlapping image inputs. Second, to learn a semantically rich 3D representation, we propose distilling features from pre-trained 2D foundation models, such as CLIP or DINOv2, thereby enabling various downstream tasks without the need for costly 3D human annotations. To leverage these two insights, we introduce a novel model architecture with a two-stage lift-splat-shoot encoder and a parameterized sparse hierarchical voxel representation. Experimental results on the NuScenes and Waymo NOTR datasets demonstrate that DistillNeRF significantly outperforms existing comparable state-of-the-art self-supervised methods for scene reconstruction, novel view synthesis, and depth estimation; and it allows for competitive zero-shot 3D semantic occupancy prediction, as well as open-world scene understanding through distilled foundation model features. Demos and code will be available at https://distillnerf.github.io/. Seung Wook Kim 0001, Jiawei Yang 0002, Cunjun Yu, Boris Ivanovic, Steven Lake Waslander, Yue Wang 0041, Sanja Fidler, Marco Pavone 0001, Péter Karkus |
NeurIPS | 3 |
| 2023 | FreeNeRF: Improving Few-Shot Neural Rendering with Free Frequency RegularizationabstractNovel view synthesis with sparse inputs is a challenging problem for neural radiance fields (NeRF). Recent efforts alleviate this challenge by introducing external supervision, such as pre-trained models and extra depth signals, or by using non-trivial patch-based rendering. In this paper, we present Frequency regularized NeRF (FreeNeRF), a surprisingly simple baseline that outperforms previous methods with minimal modifications to plain NeRF. We analyze the key challenges in few-shot neural rendering and find that frequency plays an important role in NeRF's training. Based on this analysis, we propose two regularization terms: one to regularize the frequency range of NeRF's inputs, and the other to penalize the near-camera density fields. Both techniques are “free lunches” that come at no additional computational cost. We demonstrate that even with just one line of code change, the original NeRF can achieve similar performance to other complicated methods in the few-shot setting. FreeNeRF achieves state-of-the-art performance across diverse datasets, including Blender, DTU, and LLFF. We hope that this simple baseline will motivate a rethinking of the fundamental role of frequency in NeRF's training, under both the low-data regime and beyond. This project is released at FreeNeRF. Jiawei Yang 0002, Marco Pavone 0001, Yue Wang 0041 |
CVPR | 1 |
| 2022 | ConCL: Concept Contrastive Learning for Dense Prediction Pre-training in Pathology Images
Jiawei Yang 0002, Hanbo Chen, Yuan Liang 0001, Junzhou Huang, Lei He 0001, Jianhua Yao 0001 |
ECCV (21) | 1 |
| 2022 | Towards Better Understanding and Better Generalization of Low-shot Classification in Histology Images with Contrastive Learning
Jiawei Yang 0002, Hanbo Chen, Jiangpeng Yan, Jianhua Yao 0001 |
ICLR | 1 |
| 2022 | ReMix: A General and Efficient Framework for Multiple Instance Learning Based Whole Slide Image Classification
Jiawei Yang 0002, Hanbo Chen, Yu Zhao 0009, Fan Yang 0081, Yao Zhang 0010, Lei He 0001, Jianhua Yao 0001 |
MICCAI (2) | 1 |
| 2022 | mmFormer: Multimodal Medical Transformer for Incomplete Multimodal Learning of Brain Tumor Segmentation
Yao Zhang 0010, Nanjun He, Jiawei Yang 0002, Yuexiang Li, Dong Wei 0004, Yawen Huang, Yang Zhang 0002, Zhiqiang He 0002, Yefeng Zheng 0001 |
MICCAI (5) | 3 |
| 2022 | TreeMoCo: Contrastive Neuron Morphology Representation LearningabstractMorphology of neuron trees is a key indicator to delineate neuronal cell-types, analyze brain development process, and evaluate pathological changes in neurological diseases. Traditional analysis mostly relies on heuristic features and visual inspections. A quantitative, informative, and comprehensive representation of neuron morphology is largely absent but desired. To fill this gap, in this work, we adopt a Tree-LSTM network to encode neuron morphology and introduce a self-supervised learning framework named TreeMoCo to learn features without the need for labels. We test TreeMoCo on 2403 high-quality 3D neuron reconstructions of mouse brains from three different public resources. Our results show that TreeMoCo is effective in both classifying major brain cell-types and identifying sub-types. To our best knowledge, TreeMoCo is the very first to explore learning the representation of neuron tree morphology with contrastive learning. It has a great potential to shed new light on quantitative neuron morphology analysis. Code is available at https://github.com/TencentAILabHealthcare/NeuronRepresentation. Hanbo Chen, Jiawei Yang 0002, Daniel Maxim Iascone, Lei He 0001, Hanchuan Peng, Jianhua Yao 0001 |
NeurIPS | 2 |
| 2021 | Oral-3D: Reconstructing the 3D Structure of Oral Cavity from Panoramic X-rayabstractPanoramic X-ray (PX) provides a 2D picture of the patient's mouth in a panoramic view to help dentists observe the invisible disease inside the gum. However, it provides limited 2D information compared with cone-beam computed tomography (CBCT), another dental imaging method that generates a 3D picture of the oral cavity but with more radiation dose and a higher price. Consequently, it is of great interest to reconstruct the 3D structure from a 2D X-ray image, which can greatly explore the application of X-ray imaging in dental surgeries. In this paper, we propose a framework, named Oral-3D, to reconstruct the 3D oral cavity from a single PX image and prior information of the dental arch. Specifically, we first train a generative model to learn the cross-dimension transformation from 2D to 3D. Then we restore the shape of the oral cavity with a deformation module with the dental arch curve, which can be obtained simply by taking a photo of the patient's mouth. To be noted, Oral-3D can restore both the density of bony tissues and the curved mandible surface. Experimental results show that Oral-3D can efficiently and effectively reconstruct the 3D oral structure and show critical information in clinical applications, e.g., tooth pulling and dental implants. To the best of our knowledge, we are the first to explore this domain transformation problem between these two imaging methods. Weinan Song, Yuan Liang 0001, Jiawei Yang 0002, Kun Wang 0005, Lei He 0001 |
AAAI | 3 |
| 2021 | OralViewer: 3D Demonstration of Dental Surgeries for Patient Education with Oral Cavity Reconstruction from a 2D Panoramic X-rayabstractPatient’s understanding on forthcoming dental surgeries is required by patient-centered care and helps reduce anxiety. Due to the complexity of dental surgeries and the patient-dentist expertise gap, conventional techniques of patient education are usually not effective for explaining surgical steps. In this paper, we present OralViewer—the first interactive application that enables dentist’s demonstration of dental surgeries in 3D to promote patients’ understanding. OralViewer takes a single 2D panoramic dental X-ray to reconstruct patient-specific 3D teeth structures, which are then assembled with registered gum and jaw bone models for complete oral cavity modeling. During the demonstration, OralViewer enables dentists to show surgery steps with virtual dental instruments that can animate effects on a 3D model in real-time. A technical evaluation shows that our deep learning model achieves a mean Intersection over Union (IoU) of 0.771 for 3D teeth reconstruction. A patient study with 12 participants shows OralViewer can improve patients’ understanding of surgeries. A preliminary expert study with 3 board-certified dentists further verifies the clinical validity of our system. Yuan Liang 0001, Liang Qiu 0001, Tiancheng Lu, Zhujun Fang, Dezhan Tu, Jiawei Yang 0002, Yiting Shao, Kun Wang 0005, Xiang 'Anthony' Chen, Lei He 0001 |
IUI | 6 |
| 2021 | TumorCP: A Simple but Effective Object-Level Data Augmentation for Tumor Segmentation
Jiawei Yang 0002, Yao Zhang 0010, Yuan Liang 0001, Yang Zhang 0002, Lei He 0001, Zhiqiang He 0002 |
MICCAI (1) | 1 |
| 2021 | Modality-Aware Mutual Learning for Multi-modal Medical Image Segmentation
Yao Zhang 0010, Jiawei Yang 0002, Jiang Tian, Zhongchao Shi, Yang Zhang 0002, Zhiqiang He 0002 |
MICCAI (1) | 2 |
| 2021 | The state of the art in kidney and kidney tumor segmentation in contrast-enhanced CT imaging: Results of the KiTS19 challenge
Nicholas Heller, Fabian Isensee, Klaus H. Maier-Hein, Xiaoshuai Hou, Chunmei Xie, Fengyi Li, Yang Nan 0002, Guangrui Mu, Miofei Han, Guang Yao, Yaozong Gao, Yao Zhang 0010, Yixin Wang 0003, Feng Hou, Jiawei Yang 0002, Guangwei Xiong, Jiang Tian, Christopher J. Weight |
Medical Image Anal. | 16 |
| 2020 | Atlas-aware ConvNet for Accurate yet Robust Anatomical SegmentationabstractConvolutional networks (ConvNets) have achieved promising accuracy for various anatomical segmentation tasks. Despite the success, these methods can be sensitive to appearance variations that unforeseen from the training distributions. Considering the large variability of scans caused by artifacts, pathologies, and scanning setups, the robustness of ConvNets poses as a major challenge for their clinical applications, yet has not been much explored. In this paper, we propose to mitigate the challenge by enabling ConvNets’ awareness of the underlying anatomical invariances among imaging scans. Specifically, we introduce a fully convolutional Constraint Adoption Module (CAM) that incorporates probabilistic atlas priors as explicit constraints for predictions over a locally connected Conditional Random Field (CFR), which effectively reinforces the anatomical consistency of the labeling outputs. We design the CAM to be flexible for boosting various ConvNet, and compact for co-optimizing with ConvNets for fusion parameters that leads to the optimal performance. We show the advantage of such atlas priors fusion is two-fold with two brain parcellation tasks. First, our models achieve state-of-the-art accuracy among ConvNet-based methods on both datasets, by significantly reducing structural abnormalities of predictions. Second, we can largely boost the robustness of existing ConvNets, proved by: (i) testing on scans with synthetic pathologies, and (ii) training and evaluation on scans of different scanning setups across datasets. Our method is proposing to be easily adopted to existing ConvNets by fine-tuning with CAM plugged in for accuracy and robustness boosts. Yuan Liang 0001, Weinan Song, Jiawei Yang 0002, Liang Qiu 0001, Kun Wang 0005, Lei He 0001 |
ACML | 3 |
| 2020 | X2Teeth: 3D Teeth Reconstruction from a Single Panoramic Radiograph
Yuan Liang 0001, Weinan Song, Jiawei Yang 0002, Liang Qiu 0001, Kun Wang 0005, Lei He 0001 |
MICCAI (2) | 3 |