EDBT 2026 Demo / reviewers in the wild / expert
Shenlong Wang
dblp:117/4842
· DBLP profile ↗
83ranked-venue papers
11as first author
49since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 67 · 10 first-author · 36 since 2021Graphics, computer vision, multimedia, augmented reality and games · 61 · 9 first-author · 36 since 2021Systems, architecture and hardware · 7 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | CropCraft: Complete Structural Characterization of Crop Plants from ImagesabstractThe ability to automatically build 3D digital twins of plants from images has countless applications in agriculture, environmental science, robotics, and other fields. However, current 3D reconstruction methods fail to recover complete shapes of plants due to heavy occlusion and complex geometries. In this work, we present a novel method for 3D modeling of agricultural crops based on optimizing a parametric model of plant morphology via inverse procedural modeling. Our method first estimates depth maps by fitting a neural radiance field and then optimizes a specialized loss to estimate morphological parameters that result in consistent depth renderings. The resulting 3D model is complete and biologically plausible. We validate our method on a dataset of real images of agricultural fields, and demonstrate that the reconstructed canopies can be used for a variety of monitoring and simulation applications. Project page: https://ajzhai.github.io/CropCraft Albert J. Zhai, Zhao Jiang, Sheng Wang 0020, Zhenong Jin, Kaiyu Guan, Shenlong Wang |
3DV | 9 |
| 2026 | Design, Modeling, and Experiment of a Rigid-Flexible Coupling Robotic Manta Ray With Fast Motion PerformanceabstractUnderwater bioinspired robots’ pursuit of high speed locomotion is fundamentally limited by a well-recognized physical trade-off. Increasing the flapping frequency to achieve higher speeds induces a quadratic increase in hydrodynamic resistance, which severely attenuates the oscillation amplitude of the fin and consequently limits the maximum thrust that can be generated. To address this core limitation, this paper presents the design, modeling, and experimental validation of a magnetically actuated rigid-flexible coupling robotic manta ray (MARRM). The robot’s design strategy is to address both aspects of this tradeoff simultaneously: coil-magnet (CM) pairs enable high-frequency actuation, while a biomimetic pectoral fin structure, coupling a rigid leading edge with a flexible membrane, is designed to passively maintain large undulatory amplitudes. We develop a comprehensive dynamic model to analyze and optimize this system. The model integrates an analytical magnetic torque calculation, an underactuated joint dynamics formulation, and a boundary-constrained mechanical wave function that captures the compound deformation of the fin for hydrodynamic force prediction. Simulations and comparative experiments validate the model and the design’s effectiveness. Fabricated based on this design, an untethered prototype achieves a linear swimming speed of 322.4 mm/s (3.27 body lengths per second), a turning speed of 290 °/s, and a minimal turning radius of 23.9 mm. This work demonstrates an integrated approach to mitigating the frequency-amplitude trade-off in oscillatory propulsion and provides a modeling framework applicable to the design of rigid-flexible underwater robots. Kaiyi Xu, Shenlong Wang, Shizhuo Zhang |
IEEE Trans Autom. Sci. Eng. | 2 |
| 2025 | AutoVFX: Physically Realistic Video Editing from Natural Language InstructionsabstractModern visual effects (VFX) software has made it possible for skilled artists to create imagery of virtually anything. However, the creation process remains laborious, complex, and largely inaccessible to everyday users. In this work, we present AutoVFX, a framework that automatically creates realistic and dynamic VFX videos from a single video and natural language instructions. By carefully integrating neural scene modeling, LLM-based code generation, and physical simulation, AutoVFX is able to provide physically-grounded, photorealistic editing effects that can be controlled directly using natural language instructions. We conduct extensive experiments to validate AutoVFX's efficacy across a diverse spectrum of videos and instructions. Quantitative and qualitative results suggest that AutoVFX outperforms all competing methods by a large margin in generative quality, instruction alignment, editing versatility, and physical plausibility. Hao-Yu Hsu, Chih-Hao Lin, Albert J. Zhai, Hongchi Xia, Shenlong Wang |
3DV | 5 |
| 2025 | Plenoptic PNG: Real-Time Neural Radiance Fields in 150 KBabstractThe goal of this paper is to encode a 3D scene into an extremely compact representation from$2 D$images and to enable its transmittance, decoding and rendering in real-time across various platforms. Despite the progress in NeRFs and Gaussian Splats, their large model size and specialized renderers make it challenging to distribute free-viewpoint 3D content as easily as images. To address this, we have designed a novel 3D representation that encodes the plenoptic function into sinusoidal function indexed dense volumes. This approach facilitates feature sharing across different locations, improving compactness over traditional spatial voxels. The memory footprint of the dense 3D feature grid can be further reduced using spatial decomposition techniques. This design combines the strengths of spatial hashing functions and voxel decomposition, resulting in a model size as small as 150 KB for each 3D scene. Moreover, PPNG features a lightweight rendering pipeline with only 300 lines of code that decodes its representation into standard GL textures and fragment shaders. This enables realtime rendering using the traditional GL pipeline, ensuring universal compatibility and efficiency across various platforms without additional dependencies. Our results are available at: https://jyl.kr/ppng Jae Yong Lee 0006, Yuqun Wu, Chuhang Zou, Derek Hoiem, Shenlong Wang |
3DV | 5 |
| 2025 | UrbanIR: Large-Scale Urban Scene Inverse Rendering from a Single VideoabstractWe present UrbanIR (Urban Scene Inverse Rendering), a new inverse graphics model that enables realistic, free-viewpoint renderings of scenes under various lighting conditions with a single video. It accurately infers shape, albedo, visibility, and sun and sky illumination from wide-baseline videos, such as those from car-mounted cameras, differing from NeRF's dense view settings. In this context, standard methods often yield subpar geometry and material estimates, such as inaccurate roof representations and numerous ‘floaters’. UrbanIR addresses these issues with novel losses that reduce errors in inverse graphics inference and rendering artifacts. Its techniques allow for precise shadow volume estimation in the original scene. The model's outputs support controllable editing, enabling photorealistic free-viewpoint renderings of night simulations, relit scenes, and inserted objects, marking a significant improvement over existing state-of-the-art methods. Our code and data will be made publicly available upon acceptance. Chih-Hao Lin, Kuan-Sheng Chen, David A. Forsyth, Jia-Bin Huang 0001, Anand Bhattad, Shenlong Wang |
3DV | 8 |
| 2025 | MonoPatchNeRF: Improving Neural Radiance Fields with Patch-Based Monocular GuidanceabstractThe latest regularized Neural Radiance Field (NeRF) approaches produce poor geometry and view extrapolation for large scale sparse view scenes, such as ETH3D. Density-based approaches tend to be under-constrained, while surface-based approaches tend to miss details. In this paper, we take a density-based approach, sampling patches instead of individual rays to better incorporate monocular depth and normal estimates and patch-based photometric consistency constraints between training views and sampled virtual views. Loosely constraining densities based on estimated depth aligned to sparse points further improves geometric accuracy. While maintaining similar view synthesis quality, our approach significantly improves geometric accuracy on the ETH3D benchmark, e.g. increasing the F1@2cm score by 4x-8x compared to other regularized density-based approaches, with much lower training and inference time than other approaches. Yuqun Wu, Jae Yong Lee 0006, Chuhang Zou, Shenlong Wang, Derek Hoiem |
3DV | 4 |
| 2025 | PhysGen3D: Crafting a Miniature Interactive World from a Single ImageabstractEnvisioning physically plausible outcomes from a single image requires a deep understanding of the world’s dynamics. To address this, we introduce PhysGen3D, a novel framework that transforms a single image into an amodal, camera-centric, interactive 3D scene. By combining advanced image-based geometric and semantic understanding with physics-based simulation, PhysGen3D creates an interactive 3D world from a static image, enabling us to "imagine" and simulate future scenarios based on user input. At its core, PhysGen3D estimates 3D shapes, poses, physical and lighting properties of objects, thereby capturing essential physical attributes that drive realistic object interactions. This framework allows users to specify precise initial conditions, such as object speed or material properties, for enhanced control over generated video outcomes. We evaluate PhysGen3D’s performance against closed-source state-of-the-art (SOTA) image-to-video models, including Pika, Kling, and Gen-3, showing PhysGen3D’s capacity to generate videos with realistic physics while offering greater flexibility and fine-grained control. Our results show that PhysGen3D achieves a unique balance of photorealism, physical plausibility, and user-driven interactivity, opening new possibilities for generating dynamic, physics-grounded video from an image. Project page: https://by-luckk.github.io/PhysGen3D. Hanxiao Jiang 0001, Saurabh Gupta 0001, Yunzhu Li, Shenlong Wang |
CVPR | 7 |
| 2025 | IRIS: Inverse Rendering of Indoor Scenes from Low Dynamic Range ImagesabstractInverse rendering seeks to recover 3D geometry, surface material, and lighting from captured images, enabling advanced applications such as novel-view synthesis, relighting, and virtual object insertion. However, most existing techniques rely on high dynamic range (HDR) images as input, limiting accessibility for general users. In response, we introduce IRIS, an inverse rendering framework that recovers the physically based material, spatially-varying HDR lighting, and camera response functions from multi-view, low-dynamic-range (LDR) images. By eliminating the dependence on HDR input, we make inverse rendering technology more accessible. We evaluate our approach on real-world and synthetic scenes and compare it with state-of-the-art methods. Our results show that IRIS effectively recovers HDR lighting, accurate material, and plausible camera response functions, supporting photorealistic relighting and object insertion. Chih-Hao Lin, Jia-Bin Huang 0001, Zhengqin Li, Zhao Dong 0001, Christian Richardt, Tuotuo Li, Michael Zollhöfer, Johannes Kopf 0001, Shenlong Wang, Changil Kim 0001 |
CVPR | 9 |
| 2025 | DRAWER: Digital Reconstruction and Articulation With Environment RealismabstractCreating virtual digital replicas from real-world data unlocks significant potential across domains like gaming and robotics. In this paper, we present DRAWER, a novel framework that converts a video of a static indoor scene into a photorealistic and interactive digital environment. Our approach centers on two main contributions: (i) a reconstruction module based on a dual scene representation that reconstructs the scene with fine-grained geometric details, and (ii) an articulation module that identifies articulation types and hinge positions, reconstructs simulatable shapes and appearances and integrates them into the scene. The resulting virtual environment is photorealistic, interactive, and runs in real time, with compatibility for game engines and robotic simulation platforms. We demonstrate the potential of DRAWER by using it to automatically create an interactive game in Unreal Engine and to enable real-to-sim-to-real transfer for robotics applications. Project page: here. Hongchi Xia, Entong Su, Marius Memmel, Arhan Jain, Raymond Yu, Numfor Mbiziwo-Tiapo, Ali Farhadi, Abhishek Gupta 0004, Shenlong Wang, Wei-Chiu Ma |
CVPR | 9 |
| 2025 | Uni4D: Unifying Visual Foundation Models for 4D Modeling from a Single VideoabstractThis paper presents a unified approach to understanding dynamic scenes from casual videos. Large pretrained vision foundation models, such as vision-language, video depth prediction, motion tracking, and segmentation models, offer promising capabilities. However, training a single model for comprehensive 4D understanding remains challenging. We introduce Uni4D, a multi-stage optimization framework that harnesses multiple pretrained models to advance dynamic 3D modeling, including static/dynamic reconstruction, camera pose estimation, and dense 3D motion tracking. Our results show state-of-the-art performance in dynamic 4D modeling with superior visual quality. Notably, Uni4D requires no retraining or fine- tuning, highlighting the effectiveness of repurposing visual foundation models for 4D understanding. Code and more results are available at: https://davidyao99.github.io/uni4d. David Yifan Yao, Albert J. Zhai, Shenlong Wang |
CVPR | 3 |
| 2025 | VoxelSplat: Dynamic Gaussian Splatting as an Effective Loss for Occupancy and Flow PredictionabstractRecent advancements in camera-based occupancy prediction have focused on the simultaneous prediction of 3D semantics and scene flow, a task that presents significant challenges due to specific difficulties, e.g., occlusions and unbalanced dynamic environments. In this paper, we analyze these challenges and their underlying causes. To address them, we propose a novel regularization framework called VoxelSplat. This framework leverages recent developments in 3D Gaussian Splatting to enhance model performance in two key ways: (i) Enhanced Semantics Supervision through 2D Projection: During training, our method decodes sparse semantic 3D Gaussians from 3D representations and projects them onto the 2D camera view. This provides additional supervision signals in the camera-visible space, allowing 2D labels to improve the learning of 3D semantics. (ii) Scene Flow Learning: Our framework uses the predicted scene flow to model the motion of Gaussians, and is thus able to learn the scene flow of moving objects in a self-supervised manner using the labels of adjacent frames. Our method can be seamlessly integrated into various existing occupancy models, enhancing performance without increasing inference time. Extensive experiments on benchmark datasets demonstrate the effectiveness of Voxel-Splat in improving the accuracy of both semantic occupancy and scene flow estimation. The project page and codes are available at https://zzy816.github.io/VoxelSplat-Demo/. Ziyue Zhu, Shenlong Wang, Jin Xie 0001, Jiang-jiang Liu, Jingdong Wang 0001, Jian Yang 0003 |
CVPR | 2 |
| 2025 | InvRGB+L: Inverse Rendering of Complex Scenes with Unified Color and LiDAR Reflectance ModelingabstractWe present InvRGB+L, a novel inverse rendering model that reconstructs large, relightable, and dynamic scenes from a single RGB+LiDAR sequence. Conventional inverse graphics methods rely primarily on RGB observations and use LiDAR mainly for geometric information, often resulting in suboptimal material estimates due to visible light interference. We find that LiDAR's intensity values-captured with active illumination in a different spectral range-offer complementary cues for robust material estimation under variable lighting. Inspired by this, InvRGB+L leverages LiDAR intensity cues to overcome challenges inherent in RGB-centric inverse graphics through two key innovations: (1) a novel physics-based LiDAR shading model and (2) RGB-LiDAR material consistency losses. The model produces novel-view RGB and LiDAR renderings of urban and indoor scenes and supports relighting, night simulations, and dynamic object insertions, achieving results that surpass current state-of-the-art methods in both scene-level urban inverse rendering and LiDAR simulation. Xiaoxue Chen, Bhargav Chandaka, Chih-Hao Lin, Ya-Qin Zhang, David A. Forsyth, Shenlong Wang |
ICCV | 7 |
| 2025 | Demeter: A Parametric Model of Crop Plant Morphology from the Real WorldabstractLearning 3D parametric shape models of objects has gained popularity in vision and graphics and has showed broad utility in 3D reconstruction, generation, understanding, and simulation. While powerful models exist for humans and animals, equally expressive approaches for modeling plants are lacking. In this work, we present Demeter, a data-driven parametric model that encodes key factors of a plant morphology, including topology, shape, articulation, and deformation into a compact learned representation. Unlike previous parametric models, Demeter handles varying shape topology across various species and models three sources of shape variation: articulation, subcomponent shape variation, and non-rigid deformation. To advance crop plant modeling, we collected a large-scale, ground-truthed dataset from a soybean farm as a testbed. Experiments show that Demeter effectively synthesizes shapes, reconstructs structures, and simulates biophysical processes. Code and data is available at https://tianhang-cheng.github.io/Demeter/. Tianhang Cheng, Akbert J. Zhai, Evan Z. Chen, Kejie Zhao, Janice Shiu, Qianyu Zhao, Yide Xu, Sheng Wang 0020, Lisa Ainsworth, Kaiyu Guan, Shenlong Wang |
ICCV | 16 |
| 2025 | PhysTwin: Physics-Informed Reconstruction and Simulation of Deformable Objects from Videos
Hanxiao Jiang 0001, Hao-Yu Hsu, Hsin-Ni Yu, Shenlong Wang, Yunzhu Li |
ICCV | 5 |
| 2025 | Controllable Weather Synthesis and Removal with Video Diffusion Models
Chih-Hao Lin, Ruofan Liang, Yuxuan Zhang 0001, Sanja Fidler, Shenlong Wang, Zan Gojcic |
ICCV | 6 |
| 2025 | AD-GS: Object-Aware B-Spline Gaussian Splatting for Self-Supervised Autonomous DrivingabstractModeling and rendering dynamic urban driving scenes is crucial for self-driving simulation. Current high-quality methods typically rely on costly manual object tracklet annotations, while self-supervised approaches fail to capture dynamic object motions accurately and decompose scenes properly, resulting in rendering artifacts. We introduce AD-GS, a novel self-supervised framework for high-quality free-viewpoint rendering of driving scenes from a single log. At its core is a novel learnable motion model that integrates locality-aware B-spline curves with global-aware trigonometric functions, enabling flexible yet precise dynamic object modeling. Rather than requiring comprehensive semantic labeling, AD-GS automatically segments scenes into objects and background with the simplified pseudo 2D segmentation, representing objects using dynamic Gaussians and bidirectional temporal visibility masks. Further, our model incorporates visibility reasoning and physically rigid regularization to enhance robustness. Extensive evaluations demonstrate that our annotation-free model significantly outperforms current state-of-the-art annotation-free methods and is competitive with annotation-dependent approaches. Zexin Fan, Shenlong Wang, Jin Xie 0001, Jian Yang 0003 |
ICCV | 4 |
| 2025 | LIFe-GoM: Generalizable Human Rendering with Learned Iterative Feedback Over Multi-Resolution Gaussians-on-MeshabstractGeneralizable rendering of an animatable human avatar from sparse inputs relies on data priors and inductive biases extracted from training on large data to avoid scene-specific optimization and to enable fast reconstruction. This raises two main challenges: First, unlike iterative gradient-based adjustment in scene-specific optimization, generalizable methods must reconstruct the human shape representation in a single pass at inference time.
Second, rendering is preferably computationally efficient yet of high resolution.
To address both challenges we augment the recently proposed
dual shape representation, which combines the benefits of a mesh and Gaussian points, in two ways.
To improve reconstruction, we propose an iterative feedback update framework, which successively improves the canonical human shape representation during reconstruction.
To achieve computationally efficient yet high-resolution rendering, we study a coupled-multi-resolution Gaussians-on-Mesh representation.
We evaluate the proposed approach on the challenging THuman2.0, XHuman and AIST++ data. Our approach reconstructs an animatable representation from sparse inputs in less than 1s, renders views with 95.1FPS at $1024 \times 1024$, and achieves PSNR/LPIPS*/FID of 24.65/110.82/51.27 on THuman2.0, outperforming the state-of-the-art in rendering quality. Alexander G. Schwing, Shenlong Wang |
ICLR | 3 |
| 2025 | LidarDM: Generative LiDAR Simulation in a Generated WorldabstractWe present LidarDM, a novel LiDAR generative model capable of producing realistic, layout-aware, physically plausible, and temporally coherent LiDAR videos. LidarDM stands out with two unprecedented capabilities in LiDAR generative modeling: (i) LiDAR generation guided by driving scenarios, offering significant potential for autonomous driving simulations, and (ii) 4D LiDAR point cloud generation, enabling the creation of realistic and temporally coherent sequences. At the heart of our model is a novel integrated 4D world generation framework. Specifically, we employ latent diffusion models to generate the 3D scene, combine it with dynamic actors to form the underlying 4D world, and subsequently produce realistic sensory observations within this virtual environment. Our experiments indicate that our approach outperforms competing algorithms in realism, temporal coherency, and layout consistency. We additionally show that LidarDM can be used as a generative world model simulator for training and testing perception models. We release our source code and checkpoints at https://github.com/vzyrianov/LidarDM Vlas Zyrianov, Henry Che, Shenlong Wang |
ICRA | 4 |
| 2025 | Visual Sync: Multi-Camera Synchronization via Cross-View Object MotionabstractToday, people can easily record memorable moments, ranging from concerts, sports events, lectures, family gatherings, and birthday parties with multiple consumer cameras. However, synchronizing these cross‑camera streams remains challenging. Existing methods assume controlled settings, specific targets, manual correction, or costly hardware.
We present VisualSync, an optimization framework based on multi‑view dynamics that aligns unposed, unsynchronized videos at millisecond accuracy. Our key insight is that any moving 3D point, when co‑visible in two cameras, obeys epipolar constraints once properly synchronized. To exploit this, VisualSync leverages off‑the‑shelf 3D reconstruction, feature matching, and dense tracking to extract tracklets, relative poses, and cross‑view correspondences. It then jointly minimizes the epipolar error to estimate each camera’s time offset. Experiments on four diverse, challenging datasets show that VisualSync outperforms baseline methods, achieving an average synchronization error below 130 ms. David Yifan Yao, Saurabh Gupta 0001, Shenlong Wang |
NeurIPS | 4 |
| 2025 | NoPo-Avatar: Generalizable and Animatable Avatars from Sparse Inputs without Human PosesabstractWe tackle the task of recovering an animatable 3D human avatar from a single or a sparse set of images. For this task, beyond a set of images, many prior state-of-the-art methods use accurate “ground-truth” camera poses and human poses as input to guide reconstruction at test-time. We show that pose‑dependent reconstruction degrades results significantly if pose estimates are noisy.
To overcome this, we introduce NoPo-Avatar, which reconstructs avatars solely from images, without any pose input.
By removing the dependence of test-time reconstruction on human poses, NoPo-Avatar is not affected by noisy human pose estimates, making it more widely applicable. Experiments on challenging THuman2.0, XHuman, and HuGe100K data show that NoPo-Avatar outperforms existing baselines in practical settings (without ground‑truth poses) and delivers comparable results in lab settings (with ground‑truth poses). Alexander G. Schwing, Shenlong Wang |
NeurIPS | 3 |
| 2025 | HoloScene: Simulation-Ready Interactive 3D Worlds from a Single VideoabstractDigitizing the physical world into accurate simulation‑ready virtual environments offers significant opportunities in a variety of fields such as augmented and virtual reality, gaming, and robotics. However, current 3D reconstruction and scene-understanding methods commonly fall short in one or more critical aspects, such as geometry completeness, object interactivity, physical plausibility, photorealistic rendering, or realistic physical properties for reliable dynamic simulation. To address these limitations, we introduce HoloScene, a novel interactive 3D reconstruction framework that simultaneously achieves these requirements. HoloScene leverages a comprehensive interactive scene-graph representation, encoding object geometry, appearance, and physical properties alongside hierarchical and inter-object relationships. Reconstruction is formulated as an energy-based optimization problem, integrating observational data, physical constraints, and generative priors into a unified, coherent objective. Optimization is efficiently performed via a hybrid approach combining sampling-based exploration with gradient-based refinement. The resulting digital twins exhibit complete and precise geometry, physical stability, and realistic rendering from novel viewpoints. Evaluations conducted on multiple benchmark datasets demonstrate superior performance, while practical use-cases in interactive gaming and real-time digital-twin manipulation illustrate HoloScene's broad applicability and effectiveness. Hongchi Xia, Chih-Hao Lin, Hao-Yu Hsu, Quentin Leboutet, Katelyn Gao, Michael Paulitsch, Benjamin Ummenhofer, Shenlong Wang |
NeurIPS | 8 |
| 2025 | Ada: A Distributed, Power-Aware, Real-Time Scene Provider for XRabstractReal-time scene provisioning-reconstructing and delivering scene data to requesting XR applications during runtime-is central to enabling spatial computing in modern XR systems. However, existing solutions struggle to balance latency, power and scene fidelity under XR device constraints, and often rely on designs that are either closed, application-specific designs, or both. We present Ada, the first open distributed, power-aware, application-agnostic real-time scene provisioning system. Through computation offloading along with algorithmic and system innovations, Ada provides high-fidelity scenes with stable performance across all evaluated scene sizes and with low power consumption. To isolate the benefits of Ada's algorithmic and design innovations over the closest prior work [82], which is on-device and CPU-based, we configure a comparable on-device, CPU-based variant of Ada (AdaLocal-CPU). We show this variant achieves up to 6.8× lower scene request latency and higher scene fidelity compared to the prior work. Furthermore, Ada's final distributed GPU-accelerated implementation reduces latency by an additional 2×, highlighting the benefits of GPU acceleration and distributed computing. Additionally, Ada also lowers the incremental power cost of scene provisioning by 24% compared to the best on-device variant (AdaLocal-GPU). Finally, Ada flexibly adapts to diverse latency, power, scene fidelity, and network bandwidth requirements. Yihan Pang, Sushant Kondguli, Shenlong Wang, Sarita V. Adve |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2024 | Objects With Lighting: A Real-World Dataset for Evaluating Reconstruction and Rendering for Object RelightingabstractReconstructing an object from photos and placing it virtually in a new environment goes beyond the standard novel view synthesis task as the appearance of the object has to not only adapt to the novel viewpoint but also to the new lighting conditions and yet evaluations of inverse rendering methods rely on novel view synthesis data or simplistic synthetic datasets for quantitative analysis. This work presents a real-world dataset for measuring the reconstruction and rendering of objects for relighting. To this end, we capture the environment lighting and ground truth images of the same objects in multiple environments allowing to reconstruct the objects from images taken in one environment and quantify the quality of the rendered views for the unseen lighting environments. Further, we introduce a simple baseline composed of off-the-shelf methods and test several state-of-the-art methods on the relighting task and show that novel view synthesis is not a reliable proxy to measure performance. Code and dataset are available at https://github.com/isl-org/objects-with-lighting. Benjamin Ummenhofer, Sanskar Agrawal, Rene Sepúlveda, Yixing Lao, Tianhang Cheng, Stephan R. Richter, Shenlong Wang, Germán Ros 0001 |
3DV | 8 |
| 2024 | GoMAvatar: Efficient Animatable Human Modeling from Monocular Video Using Gaussians-on-MeshabstractWe introduce GoMAvatar, a novel approach for real-time, memory-efficient, high-quality animatable human modeling. GoMAvatar takes as input a single monocular video to create a digital avatar capable of re-articulation in new poses and real-time rendering from novel view-points, while seamlessly integrating with rasterization-based graphics pipelines. Central to our method is the Gaussians-on-Mesh (GoM) representation, a hybrid 3D model combining rendering quality and speed of Gaussian splatting with geometry modeling and compatibility of deformable meshes. We assess GoMAvatar on ZJU-MoCap, PeopleSnapshot, and various YouTube videos. GoMAvatar matches or surpasses current monocular human modeling algorithms in rendering quality and significantly outperforms them in computational efficiency (43 FPS) while being memory-efficient (3.63 MB per subject). Xiaoming Zhao 0001, Zhongzheng Ren, Alexander G. Schwing, Shenlong Wang |
CVPR | 5 |
| 2024 | Video2Game: Real-time, Interactive, Realistic and Browser-Compatible Environment from a Single VideoabstractCreating high-quality and interactive virtual environments, such as games and simulators, often involves complex and costly manual modeling processes. In this paper, we present Video2Game, a novel approach that automatically converts videos of real-world scenes into realistic and interactive game environments. At the heart of our system are three core components: (i) a neural radiance fields (NeRF) module that effectively captures the geometry and visual appearance of the scene; (ii) a mesh module that distills the knowledge from NeRF for faster rendering; and (iii) a physics module that models the interactions and physical dynamics among the objects. By following the carefully designed pipeline, one can construct an interactable and actionable digital replica of the real world. We benchmark our system on both indoor and large-scale outdoor scenes. We show that we can not only produce highly-realistic renderings in real-time, but also build interactive games on top. Hongchi Xia, Zhi-Hao Lin, Wei-Chiu Ma, Shenlong Wang |
CVPR | 4 |
| 2024 | Physical Property Understanding from Language-Embedded Feature FieldsabstractCan computers perceive the physical properties of objects solely through vision? Research in cognitive science and vision science has shown that humans excel at identifying materials and estimating their physical properties based purely on visual appearance. In this paper, we present a novel approach for dense prediction of the physical properties of objects using a collection of images. Inspired by how humans reason about physics through vision, we leverage large language models to propose candidate materials for each object. We then construct a language-embedded point cloud and estimate the physical properties of each 3D point using a zero-shot kernel regression approach. Our method is accurate, annotation-free, and applicable to any object in the open world. Experiments demonstrate the effectiveness of the proposed approach in various physical property reasoning tasks, such as estimating the mass of common objects, as well as other properties like friction and hardness. Code is available at https://ajzhai.github.io/NeRF2Physics. Albert J. Zhai, Emily Y. Chen, Gloria X. Wang, Sheng Wang 0020, Kaiyu Guan, Shenlong Wang |
CVPR | 8 |
| 2024 | PhysGen: Rigid-Body Physics-Grounded Image-to-Video Generation
Zhongzheng Ren, Saurabh Gupta 0001, Shenlong Wang |
ECCV (82) | 4 |
| 2024 | SUPERGAUSSIAN: Repurposing Video Models for 3D Super Resolution
Duygu Ceylan, Paul Guerrero 0001, Zexiang Xu, Niloy J. Mitra, Shenlong Wang, Anna Frühstück |
ECCV (29) | 6 |
| 2024 | On the Overconfidence Problem in Semantic 3D MappingabstractSemantic 3D mapping, the process of fusing depth and image segmentation information between multiple views to build 3D maps annotated with object classes in real-time, is a recent topic of interest. This paper highlights the fusion overconfidence problem, in which conventional mapping methods assign high confidence to the entire map even when they are incorrect, leading to miscalibrated outputs. Several methods to improve uncertainty calibration at different stages in the fusion pipeline are presented and compared on the ScanNet dataset. We show that the most widely used Bayesian fusion strategy is among the worst calibrated, and propose a learned pipeline that combines fusion and calibration, GLFS, which achieves simultaneously higher accuracy and 3D map calibration while retaining real-time capability and adding only 525 learned parameters to the pipeline. We further illustrate the importance of map calibration on a downstream task by showing that incorporating proper semantic fusion to an indoor object search agent improves its success rates. João Marcos Correia Marques, Albert J. Zhai, Shenlong Wang, Kris Hauser |
ICRA | 3 |
| 2024 | Towards Energy-Efficiency by Navigating the Trilemma of Energy, Latency, and AccuracyabstractExtended Reality (XR) enables immersive experiences through untethered headsets but suffers from stringent battery and resource constraints. Energy-efficient design is crucial to ensure both longevity and high performance in XR devices. However, latency and accuracy are often prioritized over energy, leading to a gap in achieving energy efficiency. This paper examines scene reconstruction, a key building block for immersive XR experiences, and demonstrates how energy efficiency can be achieved by navigating the trilemma of energy, latency, and accuracy. We explore three classes of energy-oriented optimizations, covering the algorithm, execution, and data, that reveal a broad de-sign space through configurable parameters. Our resulting 72 designs expose a wide range of latency and energy trade-offs, with a smaller range of accuracy loss. We identify a Pareto-optimal curve and show that the designs on the curve are achievable only through synergistic co-optimization of all three optimization classes and by considering the latency and accuracy needs of downstream scene reconstruction consumers. Our analysis covering various use cases and measurements on an embedded class system shows that, relative to the baseline, our designs offer energy benefits of up to $60 \times$ with potential latency range of $4 \times$ slowdown to $2 \times$ speedup. Detailed exploration of a use case across representative data sequences from ScanNet showed about $25 \times$ energy savings with $1.5 \times$ latency reduction and negligible reconstruction quality loss. Boyuan Tian, Yihan Pang, Muhammad Huzaifa, Shenlong Wang, Sarita V. Adve |
ISMAR | 4 |
| 2024 | DiffTune: Autotuning Through AutodifferentiationabstractThe performance of robots in high-level tasks depends on the quality of their lower level controller, which requires fine-tuning. However, the intrinsically nonlinear dynamics and controllers make tuning a challenging task when it is done by hand. In this article, we present DiffTune, a novel, gradient-based automatic tuning framework. We formulate the controller tuning as a parameter optimization problem. Our method unrolls the dynamical system and controller as a computational graph and updates the controller parameters through gradient-based optimization. The gradient is obtained using sensitivity propagation, which is the only method for gradient computation when tuning for a physical system instead of its simulated counterpart. Furthermore, we use$\mathcal {L}_{1}$adaptive control to compensate for the uncertainties (that unavoidably exist in a physical system) such that the gradient is not biased by the unmodeled uncertainties. We validate the DiffTune on a Dubin's car and a quadrotor in challenging simulation environments. In comparison with state-of-the-art autotuning methods, DiffTune achieves the best performance in a more efficient manner owing to its effective usage of the first-order information of the system. Experiments on tuning a nonlinear controller for quadrotor show promising results, where DiffTune achieves 3.5× tracking error reduction on an aggressive trajectory in only ten trials over a 12-D controller parameter space. Sheng Cheng 0001, Minkyung Kim 0011, Lin Song 0001, Yiquan Jin, Shenlong Wang, Naira Hovakimyan |
IEEE Trans. Robotics | 6 |
| 2023 | Building Rearticulable Models for Arbitrary 3D Objects from 4D Point CloudsabstractWe build rearticulable models for arbitrary everyday man-made objects containing an arbitrary number of parts that are connected together in arbitrary ways via 1 degree-of-freedom joints. Given point cloud videos of such everyday objects, our method identifies the distinct object parts, what parts are connected to what other parts, and the properties of the joints connecting each part pair. We do this by jointly optimizing the part segmentation, transformation, and kinematics using a novel energy minimization frame-work. Our inferred animatable models, enables retargeting to novel poses with sparse point correspondences guidance. We test our method on a new articulating robot dataset, and the Sapiens dataset with common daily objects. Experiments show that our method outperforms two leading prior works on various metrics. Saurabh Gupta 0001, Shenlong Wang |
CVPR | 3 |
| 2023 | ClimateNeRF: Extreme Weather Synthesis in Neural Radiance FieldabstractPhysical simulations produce excellent predictions of weather effects. Neural radiance fields produce SOTA scene models. We describe a novel NeRF-editing procedure that can fuse physical simulations with NeRF models of scenes, producing realistic movies of physical phenomena in those scenes. Our application – Climate NeRF – allows people to visualize what climate change outcomes will do to them.ClimateNeRF allows us to render realistic weather effects, including smog, snow, and flood. Results can be controlled with physically meaningful variables like water level. Qualitative and quantitative studies show that our simulated results are significantly more realistic than those from SOTA 2D image editing and SOTA 3D NeRF stylization. Zhi-Hao Lin, David A. Forsyth, Jia-Bin Huang 0001, Shenlong Wang |
ICCV | 5 |
| 2023 | ContactGen: Generative Contact Modeling for Grasp GenerationabstractThis paper presents a novel object-centric contact representation ContactGen for hand-object interaction. The ContactGen comprises 3 components: a contact map indicates the contact location, a part map represents the contact hand part, and a direction map tells the contact direction within each part. Given an input object, we propose a conditional generative model to predict ContactGen and adopt model-based optimization to predict diverse and geometrically feasible grasps. Experimental results demonstrate our method can generate high-fidelity and diverse human grasps for various objects. Jimei Yang, Saurabh Gupta 0001, Shenlong Wang |
ICCV | 5 |
| 2023 | PEANUT: Predicting and Navigating to Unseen TargetsabstractEfficient ObjectGoal navigation (ObjectNav) in novel environments requires an understanding of the spatial and semantic regularities in environment layouts. In this work, we present a straightforward method for learning these regularities by predicting the locations of unobserved objects from incomplete semantic maps. Our method differs from previous prediction-based navigation methods, such as frontier potential prediction or egocentric map completion, by directly predicting unseen targets while leveraging the global context from all previously explored areas. Our prediction model is lightweight and can be trained in a supervised manner using a relatively small amount of passively collected data. Once trained, the model can be incorporated into a modular pipeline for ObjectNav without the need for any reinforcement learning. We validate the effectiveness of our method on the HM3D and MP3D ObjectNav datasets. We find that it achieves the state-of-the-art on both datasets, despite not using any additional data for training. Code is available at https://ajzhai.github.io/PEANUT. Albert J. Zhai, Shenlong Wang |
ICCV | 2 |
| 2023 | MapPrior: Bird's-Eye View Map Layout Estimation with Generative ModelsabstractDespite tremendous advancements in bird’s-eye view (BEV) perception, existing models fall short in generating realistic and coherent semantic map layouts, and they fail to account for uncertainties arising from partial sensor information (such as occlusion or limited coverage). In this work, we introduce MapPrior, a novel BEV perception framework that combines a traditional discriminative BEV perception model with a learned generative model for semantic map layouts. Our MapPrior delivers predictions with better accuracy, realism and uncertainty awareness. We evaluate our model on the large-scale nuScenes benchmark. At the time of submission, MapPrior outperforms the strongest competing method, with significantly improved MMD and ECE scores in camera- and LiDAR-based BEV perception. Furthermore, our method can be used to perpetually generate layouts with unconditional sampling. Xiyue Zhu, Vlas Zyrianov, Shenlong Wang |
ICCV | 4 |
| 2023 | Structure from Duplicates: Neural Inverse Graphics from a Pile of ObjectsabstractAbstract Our world is full of identical objects (\emph{e.g.}, cans of coke, cars of same model). These duplicates, when seen together, provide additional and strong cues for us to effectively reason about 3D. Inspired by this observation, we introduce Structure from Duplicates (SfD), a novel inverse graphics framework that reconstructs geometry, material, and illumination from a single image containing multiple identical objects. SfD begins by identifying multiple instances of an object within an image, and then jointly estimates the 6DoF pose for all instances. An inverse graphics pipeline is subsequently employed to jointly reason about the shape, material of the object, and the environment light, while adhering to the shared geometry and material constraint across instances.
Our primary contributions involve utilizing object duplicates as a robust prior for single-image inverse graphics and proposing an in-plane rotation-robust Structure from Motion (SfM) formulation for joint 6-DoF object pose estimation. By leveraging multi-view cues from a single image, SfD generates more realistic and detailed 3D reconstructions, significantly outperforming existing single image reconstruction models and multi-view reconstruction approaches with a similar or greater number of observations. Tianhang Cheng, Wei-Chiu Ma, Kaiyu Guan, Antonio Torralba 0001, Shenlong Wang |
NeurIPS | 5 |
| 2022 | NeurMiPs: Neural Mixture of Planar Experts for View SynthesisabstractWe present Neural Mixtures of Planar Experts (Neur-MiPs), a novel planar-based scene representation for modeling geometry and appearance. NeurMiPs leverages a collection of local planar experts in 3D space as the scene representation. Each planar expert consists of the parameters of the local rectangular shape representing geometry and a neural radiance field modeling the color and opacity. We render novel views by calculating ray-plane intersections and composite output colors and densities at intersected points to the image. NeurMiPs blends the efficiency of explicit mesh rendering and flexibility of the neural radiance field. Experiments demonstrate superior performance and speed of our proposed method, compared to other 3D representations in novel view synthesis. Zhi-Hao Lin, Wei-Chiu Ma, Hao-Yu Hsu, Yu-Chiang Frank Wang, Shenlong Wang |
CVPR | 5 |
| 2022 | Virtual Correspondence: Humans as a Cue for Extreme-View GeometryabstractRecovering the spatial layout of the cameras and the geometry of the scene from extreme-view images is a longstanding challenge in computer vision. Prevailing 3D reconstruction algorithms often adopt the image matching paradigm and presume that a portion of the scene is covisible across images, yielding poor performance when there is little overlap among inputs. In contrast, humans can associate visible parts in one image to the corresponding invisible components in another image via prior knowledge of the shapes. Inspired by this fact, we present a novel concept called virtual correspondences (VCs). VCs are a pair of pixels from two images whose camera rays intersect in 3D. Similar to classic correspondences, VCs conform with epipolar geometry; unlike classic correspondences, VCs do not need to be co-visible across views. Therefore VCs can be established and exploited even if images do not overlap. We introduce a method to find virtual correspondences based on humans in the scene. We showcase how VCs can be seamlessly integrated with classic bundle adjustment to recover camera poses across extreme views. Experiments show that our method significantly outperforms state-of-the-art camera pose estimation methods in challenging scenarios and is comparable in the traditional densely captured setup. Our approach also unleashes the potential of multiple down-stream tasks such as scene reconstruction from multi-view stereo and novel view synthesis in extreme-view scenarios11Project page: https://people.csail.mit.edu/weichium/virtual-correspondence/. Wei-Chiu Ma, Anqi Joyce Yang, Shenlong Wang, Raquel Urtasun, Antonio Torralba 0001 |
CVPR | 3 |
| 2022 | Learning to Generate Realistic LiDAR Point Clouds
Vlas Zyrianov, Xiyue Zhu, Shenlong Wang |
ECCV (23) | 3 |
| 2022 | SGAM: Building a Virtual 3D World through Simultaneous Generation and MappingabstractWe present simultaneous generation and mapping (SGAM), a novel 3D scene generation algorithm. Our goal is to produce a realistic, globally consistent 3D world on a large scale. Achieving this goal is challenging and goes beyond the capacities of existing 3D generation or video generation approaches, which fail to scale up to create large, globally consistent 3D scene structures. Towards tackling the challenges, we take a hybrid approach that integrates generative sensor model- ing with 3D reconstruction. Our proposed approach is an autoregressive generative framework that simultaneously generates sensor data at novel viewpoints and builds a 3D map at each timestamp. Given an arbitrary camera trajectory, our method repeatedly applies this generation-and-mapping process for thousands of steps, allowing us to create a gigantic virtual world. Our model can be trained from RGB-D sequences without having access to the complete 3D scene structure. The generated scenes are readily compatible with various interactive environments and rendering engines. Experiments on CLEVER and GoogleEarth datasets demon- strates ours can generate consistent, realistic, and geometrically-plausible scenes that compare favorably to existing view synthesis methods. Our project page is available at https://yshen47.github.io/sgam. Wei-Chiu Ma, Shenlong Wang |
NeurIPS | 3 |
| 2022 | CASA: Category-agnostic Skeletal Animal ReconstructionabstractRecovering a skeletal shape from a monocular video is a longstanding challenge. Prevailing nonrigid animal reconstruction methods often adopt a control-point driven animation model and optimize bone transforms individually without considering skeletal topology, yielding unsatisfactory shape and articulation. In contrast, humans can easily infer the articulation structure of an unknown character by associating it with a seen articulated object in their memory. Inspired by this fact, we present CASA, a novel category-agnostic articulated animal reconstruction method. Our method consists of two components, a video-to-shape retrieval process and a neural inverse graphics framework. During inference, CASA first finds a matched articulated shape from a 3D character assets bank so that the input video scores highly with the rendered image, according to a pretrained image-language model. It then integrates the retrieved character into an inverse graphics framework and jointly infers the shape deformation, skeleton structure, and skinning weights through optimization. Experiments validate the efficacy of our method in shape reconstruction and articulation. We further show that we can use the resulting skeletal-animated character for re-animation. Yuefan Wu, Zhongzheng Ren, Shenlong Wang |
NeurIPS | 5 |
| 2022 | Mending Neural Implicit Modeling for 3D Vehicle Reconstruction in the WildabstractReconstructing high-quality 3D objects from sparse, partial observations from a single view is of crucial importance for various applications in computer vision, robotics, and graphics. While recent neural implicit modeling methods show promising results on synthetic or dense data, they perform poorly on sparse and noisy real-world data. We discover that the limitations of a popular neural implicit model are due to lack of robust shape priors and lack of proper regularization. In this work, we demonstrate high-quality in-the-wild shape reconstruction using: (i) a deep encoder as a robust-initializer of the shape latent-code; (ii) regularized test-time optimization of the latent-code; (iii) a deep discriminator as a learned high-dimensional shape prior; (iv) a novel curriculum learning strategy that allows the model to learn shape priors on synthetic data and smoothly transfer them to sparse real world data. Our approach better captures the global structure, performs well on occluded and sparse observations, and registers well with the ground-truth shape. We demonstrate superior performance over state-of-the-art 3D object reconstruction methods on two real-world datasets. Shivam Duggal, Wei-Chiu Ma, Sivabalan Manivasagam, Justin Liang 0001, Shenlong Wang, Raquel Urtasun |
WACV | 6 |
| 2022 | Intracity Temperature Estimation by Physics Informed Neural Network Using Modeled Forcing Meteorology and Multispectral Satellite ImageryabstractEstimating urban surface temperature at high resolution is crucial for effective urban planning for climate-driven risks. This high-resolution surface temperature over broader scales can usually be obtained via satellite remote sensing for historical period. However, it can be hard for future predictions. This paper presents a Physics Informed Hierarchical Perception (PIHP) network, a novel approach for accurate, high-resolution and generalizable urban surface temperature estimation. The key to our approach is leveraging the implied temperature-related physics information of the land surface structure from high-resolution multi-spectral satellite images, thus achieving precise estimation or prediction for high spatial resolution urban surface temperature. Specifically, a semantic category histogram is first designed to describe the land surface structures. Based on this, a hierarchical urban surface perception network is proposed to capture the complex relationship between the underlying land surface features, upper atmosphere conditions and the intracity temperature. The proposed PIHP-Net makes it possible to generate models that can generalize across different cities, thus to estimating or predicting high-resolution urban surface temperature when the satellite land surface temperature (LST) observation is not available. Experiments over various cities in different climate regions in China show, for the first time, errors less than 2 Kelvin (for most of the cases) at the high resolution (60-by-60 meters grids), thus making it possible to predict futureintracity temperaturefrom forcing meteorology and multi-spectral satellite imagery. Donghang Wu, Weiquan Liu, Lei Zhao 0023, Shenlong Wang, Cheng Wang 0003, José Marcato Junior, Jonathan Li 0001 |
IEEE Trans. Geosci. Remote. Sens. | 7 |
| 2021 | GeoSim: Realistic Video Simulation via Geometry-Aware Composition for Self-DrivingabstractScalable sensor simulation is an important yet challenging open problem for safety-critical domains such as self-driving. Current works in image simulation either fail to be photorealistic or do not model the 3D environment and the dynamic objects within, losing high-level control and physical realism. In this paper, we present GeoSim, a geometry-aware image composition process which synthesizes novel urban driving scenarios by augmenting existing images with dynamic objects extracted from other scenes and rendered at novel poses. Towards this goal, we first build a diverse bank of 3D objects with both realistic geometry and appearance from sensor data. During simulation, we perform a novel geometry-aware simulation-by-composition procedure which 1) proposes plausible and realistic object placements into a given scene, 2) renders novel views of dynamic objects from the asset bank, and 3) composes and blends the rendered image segments. The resulting synthetic images are realistic, traffic-aware, and geometrically consistent, allowing our approach to scale to complex use cases. We demonstrate two such important applications: long-range realistic video simulation across multiple camera sensors, and synthetic data generation for data augmentation on downstream segmentation tasks. Please check https://tmux.top/publication/geosim/ for high-resolution video results. Yun Chen 0014, Frieda Rong, Shivam Duggal, Shenlong Wang, Xinchen Yan, Sivabalan Manivasagam, Shangjie Xue, Ersin Yumer, Raquel Urtasun |
CVPR | 4 |
| 2021 | SceneGen: Learning To Generate Realistic Traffic ScenesabstractWe consider the problem of generating realistic traffic scenes automatically. Existing methods typically insert actors into the scene according to a set of hand-crafted heuristics and are limited in their ability to model the true complexity and diversity of real traffic scenes, thus inducing a content gap between synthesized traffic scenes versus real ones. As a result, existing simulators lack the fidelity necessary to train and test self-driving vehicles. To address this limitation, we present SceneGen—a neural autoregressive model of traffic scenes that eschews the need for rules and heuristics. In particular, given the ego-vehicle state and a high definition map of surrounding area, SceneGen inserts actors of various classes into the scene and synthesizes their sizes, orientations, and velocities. We demonstrate on two large-scale datasets SceneGen’s ability to faithfully model distributions of real traffic scenes. Moreover, we show that SceneGen coupled with sensor simulation can be used to train perception models that generalize to the real world. Shuhan Tan, Shenlong Wang, Sivabalan Manivasagam, Mengye Ren, Raquel Urtasun |
CVPR | 3 |
| 2021 | S3: Neural Shape, Skeleton, and Skinning Fields for 3D Human ModelingabstractConstructing and animating humans is an important component for building virtual worlds in a wide variety of applications such as virtual reality or robotics testing in simulation. As there are exponentially many variations of humans with different shape, pose and clothing, it is critical to develop methods that can automatically reconstruct and animate humans at scale from real world data. Towards this goal, we represent the pedestrian’s shape, pose and skinning weights as neural implicit functions that are directly learned from data. This representation enables us to handle a wide variety of different pedestrian shapes and poses without explicitly fitting a human parametric body model, allowing us to handle a wider range of human geometries and topologies. We demonstrate the effectiveness of our approach on various datasets and show that our reconstructions outperform existing state-of-the-art methods. Furthermore, our re-animation experiments show that we can generate 3D human animations at scale from a single RGB image (and/or an optional LiDAR sweep) as input. Ze Yang 0003, Shenlong Wang, Sivabalan Manivasagam, Zeng Huang, Wei-Chiu Ma, Xinchen Yan, Ersin Yumer, Raquel Urtasun |
CVPR | 2 |
| 2021 | Asynchronous Multi-View SLAMabstractExisting multi-camera SLAM systems assume synchronized shutters for all cameras, which is often not the case in practice. In this work, we propose a generalized multi-camera SLAM formulation which accounts for asynchronous sensor observations. Our framework integrates a continuous-time motion model to relate information across asynchronous multi-frames during tracking, local mapping, and loop closing. For evaluation, we collected AMV-Bench, a challenging new SLAM dataset covering 482 km of driving recorded using our asynchronous multi-camera robotic platform. AMV-Bench is over an order of magnitude larger than previous multi-view HD outdoor SLAM datasets, and covers diverse and challenging motions and environments. Our experiments emphasize the necessity of asynchronous sensor modeling, and show that the use of multiple cameras is critical towards robust and accurate SLAM in challenging outdoor scenes. The supplementary material is located at: https://www.cs.toronto.edu/~ajyang/amv-slam Anqi Joyce Yang, Ioan Andrei Barsan, Raquel Urtasun, Shenlong Wang |
ICRA | 5 |
| 2021 | Deep Learning for Visual Data CompressionabstractIn this paper, we will introduce the recent progress in deep learning based visual data compression, including image compression, video compression and point cloud compression. In the past few years, deep learning techniques have been successfully applied to various computer vision and image processing applications. However, for the data compression task, the traditional approaches (i.e., block based motion estimation and motion compensation, etc.) are still widely employed in the mainstream codecs. Considering the powerful representation capability of neural networks, it is feasible to improve the data compression performance by employing the advanced deep learning technologies. To this end, the deep leaning based compression approaches have recently received increasing attention from both academia and industry in the field of computer vision and signal processing. Guo Lu, Shenlong Wang, Shan Liu 0001, Radu Timofte |
ACM Multimedia | 3 |
| 2020 | OctSqueeze: Octree-Structured Entropy Model for LiDAR CompressionabstractWe present a novel deep compression algorithm to reduce the memory footprint of LiDAR point clouds. Our method exploits the sparsity and structural redundancy between points to reduce the bitrate. Towards this goal, we first encode the point cloud into an octree, a data-efficient structure suitable for sparse point clouds. We then design a tree-structured conditional entropy model that can be directly applied to octree structures to predict the probability of a symbol's occurrence. We validate the effectiveness of our method over two large-scale datasets. The results demonstrate that our approach reduces the bitrate by 10- 20% at the same reconstruction quality, compared to the previous state-of-the-art. Importantly, we also show that for the same bitrate, our approach outperforms other compression algorithms when performing downstream 3D segmentation and detection tasks using compressed representations. This helps advance the feasibility of using point cloud compression to reduce the onboard and offboard storage for safety-critical applications such as self-driving cars, where a single vehicle captures 84 billion points per day. Lila Huang, Shenlong Wang, Jerry Liu, Raquel Urtasun |
CVPR | 2 |
| 2020 | LiDARsim: Realistic LiDAR Simulation by Leveraging the Real WorldabstractWe tackle the problem of producing realistic simulations of LiDAR point clouds, the sensor of preference for most self-driving vehicles. We argue that, by leveraging real data, we can simulate the complex world more realistically compared to employing virtual worlds built from CAD/procedural models. Towards this goal, we first build a large catalog of 3D static maps and 3D dynamic objects by driving around several cities with our self-driving fleet. We can then generate scenarios by selecting a scene from our catalog and "virtually" placing the self-driving vehicle (SDV) and a set of dynamic objects from the catalog in plausible locations in the scene. To produce realistic simulations, we develop a novel simulator that captures both the power of physics-based and learning-based simulation. We first utilize raycasting over the 3D scene and then use a deep neural network to produce deviations from the physics-based simulation, producing realistic LiDAR point clouds. We showcase LiDARsim's usefulness for perception algorithms-testing on long-tail events and end-to-end closed-loop evaluation on safety-critical scenarios. Sivabalan Manivasagam, Shenlong Wang, Wenyuan Zeng, Mikita Sazanovich, Shuhan Tan, Bin Yang 0021, Wei-Chiu Ma, Raquel Urtasun |
CVPR | 2 |
| 2020 | Conditional Entropy Coding for Efficient Video Compression
Jerry Liu, Shenlong Wang, Wei-Chiu Ma, Meet Shah 0001, Rui Hu 0001, Pranaab Dhawan, Raquel Urtasun |
ECCV (17) | 2 |
| 2020 | Deep Feedback Inverse Problem Solver
Wei-Chiu Ma, Shenlong Wang, Jiayuan Gu, Sivabalan Manivasagam, Antonio Torralba 0001, Raquel Urtasun |
ECCV (5) | 2 |
| 2020 | DSDNet: Deep Structured Self-driving Network
Wenyuan Zeng, Shenlong Wang, Renjie Liao 0001, Yun Chen 0014, Bin Yang 0021, Raquel Urtasun |
ECCV (21) | 2 |
| 2020 | Pit30M: A Benchmark for Global Localization in the Age of Self-Driving CarsabstractWe are interested in understanding whether retrieval-based localization approaches are good enough in the context of self-driving vehicles. Towards this goal, we introduce Pit30M, a new image and LiDAR dataset with over 30 million frames, which is 10 to 100 times larger than those used in previous work. Pit30M is captured under diverse conditions (i.e., season, weather, time of the day, traffic), and provides accurate localization ground truth. We also automatically annotate our dataset with historical weather and astronomical data, as well as with image and LiDAR semantic segmentation as a proxy measure for occlusion. We benchmark multiple existing methods for image and LiDAR retrieval and, in the process, introduce a simple, yet effective convolutional network-based LiDAR retrieval method that is competitive with the state of the art. Our work provides, for the first time, a benchmark for sub-metre retrieval-based localization at city scale.The dataset, additional experimental results, as well as more information about the sensors, calibration, and metadata, are available on the project website: https://uber.com/atg/datasets/pit30m. Julieta Martinez 0001, Sasha Doubov, Jack Fan, Ioan Andrei Barsan, Shenlong Wang, Gellért Máttyus, Raquel Urtasun |
IROS | 5 |
| 2020 | MuSCLE: Multi Sweep Compression of LiDAR using Deep Entropy ModelsabstractWe present a novel compression algorithm for reducing the storage of LiDAR sensory data streams. Our model exploits spatio-temporal relationships across multiple LIDAR sweeps to reduce the bitrate of both geometry and intensity values. Towards this goal, we propose a novel conditional entropy model that models the probabilities of the octree symbols, by considering both coarse level geometry and previous sweeps’ geometric and intensity information. We then exploit the learned probability to encode the full data-stream into a compact one. Our experiments demonstrate that our method significantly reduces the joint geometry and intensity bitrate over prior state-of-the-art LiDAR compression methods, with a reduction of 7–17% and 15–35% on the UrbanCity and SemanticKITTI datasets respectively. Sourav Biswas 0001, Jerry Liu, Shenlong Wang, Raquel Urtasun |
NeurIPS | 4 |
| 2019 | Convolutional Recurrent Network for Road Boundary ExtractionabstractCreating high definition maps that contain precise information of static elements of the scene is of utmost importance for enabling self driving cars to drive safely. In this paper, we tackle the problem of drivable road boundary extraction from LiDAR and camera imagery. Towards this goal, we design a structured model where a fully convolutional network obtains deep features encoding the location and direction of road boundaries and then, a convolutional recurrent network outputs a polyline representation for each one of them. Importantly, our method is fully automatic and does not require a user in the loop. We showcase the effectiveness of our method on a large North American city where we obtain perfect topology of road boundaries 99.3% of the time at a high precision and recall. Justin Liang 0001, Namdar Homayounfar, Wei-Chiu Ma, Shenlong Wang, Raquel Urtasun |
CVPR | 4 |
| 2019 | Deep Rigid Instance Scene FlowabstractIn this paper we tackle the problem of scene flow estimation in the context of self-driving. We leverage deep learning techniques as well as strong priors as in our application domain the motion of the scene can be composed by the motion of the robot and the 3D motion of the actors in the scene. We formulate the problem as energy minimization in a deep structured model, which can be solved efficiently in the GPU by unrolling a Gaussian-Newton solver. Our experiments in the challenging KITTI scene flow dataset show that we outperform the state-of-the-art by a very large margin, while being 800 times faster. Wei-Chiu Ma, Shenlong Wang, Rui Hu 0001, Yuwen Xiong, Raquel Urtasun |
CVPR | 2 |
| 2019 | Learning to Localize Through Compressed Binary MapsabstractOne of the main difficulties of scaling current localization systems to large environments is the on-board storage required for the maps. In this paper we propose to learn to compress the map representation such that it is optimal for the localization task. As a consequence, higher compression rates can be achieved without loss of localization accuracy when compared to standard coding schemes that optimize for reconstruction, thus ignoring the end task. Our experiments show that it is possible to learn a task-specific compression which reduces storage requirements by two orders of magnitude over general-purpose codecs such as WebP without sacrificing performance. Xinkai Wei, Ioan Andrei Barsan, Shenlong Wang, Julieta Martinez 0001, Raquel Urtasun |
CVPR | 3 |
| 2019 | DeepPruner: Learning Efficient Stereo Matching via Differentiable PatchMatchabstractOur goal is to significantly speed up the runtime of current state-of-the-art stereo algorithms to enable real-time inference. Towards this goal, we developed a differentiable PatchMatch module that allows us to discard most disparities without requiring full cost volume evaluation. We then exploit this representation to learn which range to prune for each pixel. By progressively reducing the search space and effectively propagating such information, we are able to efficiently compute the cost volume for high likelihood hypotheses and achieve savings in both memory and computation.Finally, an image guided refinement module is exploited to further improve the performance. Since all our components are differentiable, the full network can be trained end-to-end. Our experiments show that our method achieves competitive results on KITTI and SceneFlow datasets while running in real-time at 62ms. Shivam Duggal, Shenlong Wang, Wei-Chiu Ma, Rui Hu 0001, Raquel Urtasun |
ICCV | 2 |
| 2019 | DSIC: Deep Stereo Image CompressionabstractIn this paper we tackle the problem of stereo image compression, and leverage the fact that the two images have overlapping fields of view to further compress the representations. Our approach leverages state-of-the-art single-image compression autoencoders and enhances the compression with novel parametric skip functions to feed fully differentiable, disparity-warped features at all levels to the encoder/decoder of the second image. Moreover, we model the probabilistic dependence between the image codes using a conditional entropy model. Our experiments show an impressive 30 - 50% reduction in the second image bitrate at low bitrates compared to deep single-image compression, and a 10 - 20% reduction at higher bitrates. Jerry Liu, Shenlong Wang, Raquel Urtasun |
ICCV | 2 |
| 2019 | Exploiting Sparse Semantic HD Maps for Self-Driving Vehicle LocalizationabstractIn this paper we propose a novel semantic localization algorithm that exploits multiple sensors and has precision on the order of a few centimeters. Our approach does not require detailed knowledge about the appearance of the world, and our maps require orders of magnitude less storage than maps utilized by traditional geometry- and LiDAR intensity-based localizers. This is important as self-driving cars need to operate in large environments. Towards this goal, we formulate the problem in a Bayesian filtering framework, and exploit lanes, traffic signs, as well as vehicle dynamics to localize robustly with respect to a sparse semantic map. We validate the effectiveness of our method on a new highway dataset consisting of 312km of roads. Our experiments show that the proposed approach is able to achieve 0.05m lateral accuracy and 1.12m longitudinal accuracy on average while taking up only 0.3% of the storage required by previous LiDAR intensity-based approaches. Wei-Chiu Ma, Raquel Urtasun, Ignacio Tartavull, Ioan Andrei Barsan, Shenlong Wang, Min Bai, Gellért Máttyus, Namdar Homayounfar, Shrinidhi Kowshika Lakshmikanth, Andrei Pokrovsky |
IROS | 5 |
| 2019 | Efficient Graph Generation with Graph Recurrent Attention NetworksabstractWe propose a new family of efficient and expressive deep generative models of graphs, called Graph Recurrent Attention Networks (GRANs). Our model generates graphs one block of nodes and associated edges at a time. The block size and sampling stride allow us to trade off sample quality for efficiency. Compared to previous RNN-based graph generative models, our framework better captures the auto-regressive conditioning between the already-generated and to-be-generated parts of the graph using Graph Neural Networks (GNNs) with attention. This not only reduces the dependency on node ordering but also bypasses the long-term bottleneck caused by the sequential nature of RNNs. Moreover, we parameterize the output distribution per block using a mixture of Bernoulli, which captures the correlations among generated edges within the block. Finally, we propose to handle node orderings in generation by marginalizing over a family of canonical orderings. On standard benchmarks, we achieve state-of-the-art time efficiency and sample quality compared to previous models. Additionally, we show our model is capable of generating large graphs of up to 5K nodes with good quality. Our code is released at: \url{https://github.com/lrjconan/GRAN}. Renjie Liao 0001, Yujia Li 0001, Yang Song 0011, Shenlong Wang, William L. Hamilton, David Duvenaud, Raquel Urtasun, Richard S. Zemel |
NeurIPS | 4 |
| 2018 | Monocular Depth Estimation via Deep Structured Models with Ordinal ConstraintsabstractUser interaction provides useful information for solving challenging computer vision problems in practice. In this paper, we show that a very limited number of user clicks could greatly boost monocular depth estimation performance and overcome monocular ambiguities. We formulate this task as a deep structured model, in which the structured pixel-wise depth estimation has ordinal constraints introduced by user clicks. We show that the inference of the proposed model could be efficiently solved through a feed-forward network. We demonstrate the effectiveness of the proposed model on NYU Depth V2 and Stanford 2D-3D datasets. On both datasets, we achieve state-of-the-art performance when encoding user interaction into our deep models. Daniel Ron, Kun Duan, Chongyang Ma, Shenlong Wang, Sumant Hanumante, Dhritiman Sagar |
3DV | 5 |
| 2018 | Deep Parametric Continuous Convolutional Neural NetworksabstractStandard convolutional neural networks assume a grid structured input is available and exploit discrete convolutions as their fundamental building blocks. This limits their applicability to many real-world applications. In this paper we propose Parametric Continuous Convolution, a new learnable operator that operates over non-grid structured data. The key idea is to exploit parameterized kernel functions that span the full continuous vector space. This generalization allows us to learn over arbitrary data structures as long as their support relationship is computable. Our experiments show significant improvement over the state-of-the-art in point cloud segmentation of indoor and outdoor scenes, and lidar motion estimation of driving scenes. Shenlong Wang, Simon Suo, Wei-Chiu Ma, Andrei Pokrovsky, Raquel Urtasun |
CVPR | 1 |
| 2018 | Deep Continuous Fusion for Multi-sensor 3D Object Detection
Bin Yang 0021, Shenlong Wang, Raquel Urtasun |
ECCV (16) | 3 |
| 2018 | Deep Multi-Sensor Lane DetectionabstractReliable and accurate lane detection has been a long-standing problem in the field of autonomous driving. In recent years, many approaches have been developed that use images (or videos) as input and reason in image space. In this paper we argue that accurate image estimates do not translate to precise 3D lane boundaries, which are the input required by modern motion planning algorithms. To address this issue, we propose a novel deep neural network that takes advantage of both LiDAR and camera sensors and produces very accurate estimates directly in 3D space. We demonstrate the performance of our approach on both highways and in cities, and show very accurate estimates in complex scenarios such as heavy traffic (which produces occlusion), fork, merges and intersections. Min Bai, Gellért Máttyus, Namdar Homayounfar, Shenlong Wang, Shrinidhi Kowshika Lakshmikanth, Raquel Urtasun |
IROS | 4 |
| 2017 | AutoScaler: Scale-Attention Networks for Visual Correspondence
Shenlong Wang, Linjie Luo |
BMVC | 1 |
| 2017 | TorontoCity: Seeing the World with a Million EyesabstractIn this paper we introduce the TorontoCity benchmark, which covers the full greater Toronto area (GTA) with 712.5km2 of land, 8439km of road and around 400, 000 buildings. Our benchmark provides different perspectives of the world captured from airplanes, drones and cars driving around the city. Manually labeling such a large scale dataset is infeasible. Instead, we propose to utilize different sources of high-precision maps to create our ground truth. Towards this goal, we develop algorithms that allow us to align all data sources with the maps while requiring minimal human supervision. We have designed a wide variety of tasks including building height estimation (reconstruction), road centerline and curb extraction, building instance segmentation, building contour extraction (reorganization), semantic labeling and scene type classification (recognition). Our pilot study shows that most of these tasks are still difficult for modern convolutional neural networks. Shenlong Wang, Min Bai, Gellért Máttyus, Hang Chu, Wenjie Luo 0002, Bin Yang 0021, Justin Liang 0001, Joel Cheverie, Sanja Fidler, Raquel Urtasun |
ICCV | 1 |
| 2017 | Find your way by observing the sun and other semantic cuesabstractIn this paper we present a robust, efficient and affordable approach to self-localization which requires neither GPS nor knowledge about the appearance of the world. Towards this goal, we utilize freely available cartographic maps and derive a probabilistic model that exploits semantic cues in the form of sun direction, presence of an intersection, road type, speed limit and ego-car trajectory to produce very reliable localization results. Our experimental evaluation shows that our approach can localize much faster (in terms of driving time) with less computation and more robustly than competing approaches, which ignore semantic information. Wei-Chiu Ma, Shenlong Wang, Marcus A. Brubaker, Sanja Fidler, Raquel Urtasun |
ICRA | 2 |
| 2016 | HD Maps: Fine-Grained Road Segmentation by Parsing Ground and Aerial ImagesabstractIn this paper we present an approach to enhance existing maps with fine grained segmentation categories such as parking spots and sidewalk, as well as the number and location of road lanes. Towards this goal, we propose an efficient approach that is able to estimate these fine grained categories by doing joint inference over both, monocular aerial imagery, as well as ground images taken from a stereo camera pair mounted on top of a car. Important to this is reasoning about the alignment between the two types of imagery, as even when the measurements are taken with sophisticated GPS+IMU systems, this alignment is not sufficiently accurate. We demonstrate the effectiveness of our approach on a new dataset which enhances KITTI [8] with aerial images taken with a camera mounted on an airplane and flying around the city of Karlsruhe, Germany. Gellért Máttyus, Shenlong Wang, Sanja Fidler, Raquel Urtasun |
CVPR | 2 |
| 2016 | The Global Patch ColliderabstractThis paper proposes a novel extremely efficient, fully-parallelizable, task-specific algorithm for the computation of global point-wise correspondences in images and videos. Our algorithm, the Global Patch Collider, is based on detecting unique collisions between image points using a collection of learned tree structures that act as conditional hash functions. In contrast to conventional approaches that rely on pairwise distance computation, our algorithm isolates distinctive pixel pairs that hit the same leaf during traversal through multiple learned tree structures. The split functions stored at the intermediate nodes of the trees are trained to ensure that only visually similar patches or their geometric or photometric transformed versions fall into the same leaf node. The matching process involves passing all pixel positions in the images under analysis through the tree structures. We then compute matches by isolating points that uniquely collide with each other ie. fell in the same empty leaf in multiple trees. Our algorithm is linear in the number of pixels but can be made constant time on a parallel computation architecture as the tree traversal for individual image points is decoupled. We demonstrate the efficacy of our method by using it to perform optical flow matching and stereo matching on some challenging benchmarks. Experimental results show that not only is our method extremely computationally efficient, but it is also able to match or outperform state of the art methods that are much more complex. Shenlong Wang, Sean Ryan Fanello, Christoph Rhemann, Shahram Izadi, Pushmeet Kohli |
CVPR | 1 |
| 2016 | HouseCraft: Building Houses from Rental Ads and Street Views
Hang Chu, Shenlong Wang, Raquel Urtasun, Sanja Fidler |
ECCV (6) | 2 |
| 2016 | Proximal Deep Structured ModelsabstractMany problems in real-world applications involve predicting continuous-valued random variables that are statistically related. In this paper, we propose a powerful deep structured model that is able to learn complex non-linear functions which encode the dependencies between continuous output variables. We show that inference in our model using proximal methods can be efficiently solved as a feed-foward pass of a special type of deep recurrent neural network. We demonstrate the effectiveness of our approach in the tasks of image denoising, depth refinement and optical flow estimation. Shenlong Wang, Sanja Fidler, Raquel Urtasun |
NIPS | 1 |
| 2016 | Holoportation: Virtual 3D Teleportation in Real-timeabstractWe present an end-to-end system for augmented and virtual reality telepresence, called Holoportation. Our system demonstrates high-quality, real-time 3D reconstructions of an entire space, including people, furniture and objects, using a set of new depth cameras. These 3D models can also be transmitted in real-time to remote users. This allows users wearing virtual or augmented reality displays to see, hear and interact with remote participants in 3D, almost as if they were present in the same physical space. From an audio-visual perspective, communicating and interacting with remote users edges closer to face-to-face communication. This paper describes the Holoportation technical system in full, its key interactive capabilities, the application scenarios it enables, and an initial qualitative study of using this new communication medium. Sergio Orts, Christoph Rhemann, Sean Ryan Fanello, Wayne Chang, Adarsh Kowdle, Yury Degtyarev, David Kim 0002, Philip Davidson, Sameh Khamis, Mingsong Dou, Vladimir Tankovich, Charles T. Loop, Qin Cai, Philip A. Chou, Sarah Mennicken, Julien P. C. Valentin, Vivek Pradeep, Shenlong Wang, Sing Bing Kang, Pushmeet Kohli, Yuliya Lutchyn, Cem Keskin, Shahram Izadi |
UIST | 18 |
| 2015 | Holistic 3D scene understanding from a single geo-tagged imageabstractIn this paper we are interested in exploiting geographic priors to help outdoor scene understanding. Towards this goal we propose a holistic approach that reasons jointly about 3D object detection, pose estimation, semantic segmentation as well as depth reconstruction from a single image. Our approach takes advantage of large-scale crowd-sourced maps to generate dense geographic, geometric and semantic priors by rendering the 3D world. We demonstrate the effectiveness of our holistic model on the challenging KITTI dataset [13], and show significant improvements over the baselines in all metrics and tasks. Shenlong Wang, Sanja Fidler, Raquel Urtasun |
CVPR | 1 |
| 2015 | Enhancing Road Maps by Parsing Aerial Images Around the WorldabstractIn recent years, contextual models that exploit maps have been shown to be very effective for many recognition and localization tasks. In this paper we propose to exploit aerial images in order to enhance freely available world maps. Towards this goal, we make use of OpenStreetMap and formulate the problem as the one of inference in a Markov random field parameterized in terms of the location of the road-segment centerlines as well as their width. This parameterization enables very efficient inference and returns only topologically correct roads. In particular, we can segment all OSM roads in the whole world in a single day using a small cluster of 10 computers. Importantly, our approach generalizes very well, it can be trained using only 1.5 km2aerial imagery and produce very accurate results in any location across the globe. We demonstrate the effectiveness of our approach outperforming the state-of-the-art in two new benchmarks that we collect. We then show how our enhanced maps are beneficial for semantic segmentation of ground images. Gellért Máttyus, Shenlong Wang, Sanja Fidler, Raquel Urtasun |
ICCV | 2 |
| 2015 | Lost Shopping! Monocular Localization in Large Indoor SpacesabstractIn this paper we propose a novel approach to localization in very large indoor spaces (i.e., 200+ store shopping malls) that takes a single image and a floor plan of the environment as input. We formulate the localization problem as inference in a Markov random field, which jointly reasons about text detection (localizing shop's names in the image with precise bounding boxes), shop facade segmentation, as well as camera's rotation and translation within the entire shopping mall. The power of our approach is that it does not use any prior information about appearance and instead exploits text detections corresponding to the shop names. This makes our method applicable to a variety of domains and robust to store appearance variation across countries, seasons, and illumination conditions. We demonstrate the performance of our approach in a new dataset we collected of two very large shopping malls, and show the power of holistic reasoning. Shenlong Wang, Sanja Fidler, Raquel Urtasun |
ICCV | 1 |
| 2014 | Transductive Gaussian processes for image denoisingabstractIn this paper we are interested in exploiting self-similarity information for discriminative image denoising. Towards this goal, we propose a simple yet powerful denoising method based on transductive Gaussian processes, which introduces self-similarity in the prediction stage. Our approach allows to build a rich similarity measure by learning hyper parameters defining multi-kernel combinations. We introduce perceptual-driven kernels to capture pixel-wise, gradient-based and local-structure similarities. In addition, our algorithm can integrate several initial estimates as input features to boost performance even further. We demonstrate the effectiveness of our approach on several benchmarks. The experiments show that our proposed denoising algorithm has better performance than competing discriminative denoising methods, and achieves competitive result with respect to the state-of-the-art. Shenlong Wang, Lei Zhang 0006, Raquel Urtasun |
ICCP | 1 |
| 2014 | Efficient Inference of Continuous Markov Random Fields with Polynomial Potentials
Shenlong Wang, Alexander G. Schwing, Raquel Urtasun |
NIPS | 1 |
| 2012 | Nonlocal Spectral Prior Model for Low-Level Vision
Shenlong Wang, Lei Zhang 0006, Yan Liang 0001 |
ACCV (3) | 1 |
| 2012 | Semi-coupled dictionary learning with applications to image super-resolution and photo-sketch synthesisabstractIn various computer vision applications, often we need to convert an image in one style into another style for better visualization, interpretation and recognition; for examples, up-convert a low resolution image to a high resolution one, and convert a face sketch into a photo for matching, etc. A semi-coupled dictionary learning (SCDL) model is proposed in this paper to solve such cross-style image synthesis problems. Under SCDL, a pair of dictionaries and a mapping function will be simultaneously learned. The dictionary pair can well characterize the structural domains of the two styles of images, while the mapping function can reveal the intrinsic relationship between the two styles' domains. In SCDL, the two dictionaries will not be fully coupled, and hence much flexibility can be given to the mapping function for an accurate conversion across styles. Moreover, clustering and image nonlocal redundancy are introduced to enhance the robustness of SCDL. The proposed SCDL model is applied to image super-resolution and photo-sketch synthesis, and the experimental results validated its generality and effectiveness in cross-style image synthesis. Shenlong Wang, Lei Zhang 0006, Yan Liang 0001, Quan Pan 0001 |
CVPR | 1 |
| 2012 | Relaxed collaborative representation for pattern classificationabstractRegularized linear representation learning has led to interesting results in image classification, while how the object should be represented is a critical issue to be investigated. Considering the fact that the different features in a sample should contribute differently to the pattern representation and classification, in this paper we present a novel relaxed collaborative representation (RCR) model to effectively exploit the similarity and distinctiveness of features. In RCR, each feature vector is coded on its associated dictionary to allow flexibility of feature coding, while the variance of coding vectors is minimized to address the similarity among features. In addition, the distinctiveness of different features is exploited by weighting its distance to other features in the coding domain. The proposed RCR is simple, while our extensive experimental results on benchmark image databases (e.g., various face and flower databases) show that it is very competitive with state-of-the-art image classification methods. Meng Yang 0001, Lei Zhang 0006, David Zhang 0001, Shenlong Wang |
CVPR | 4 |