VLDB 2026 Research / reviewers in the wild / expert
Peter Wonka
dblp:98/5522
· DBLP profile ↗
219ranked-venue papers
4as first author
80since 2021 · last 2026
0000-0003-0627-9746ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 180 · 4 first-author · 61 since 2021Artificial intelligence and machine learning · 83 · 55 since 2021Applied, interdisciplinary, general and emerging computing · 6Human-computer interaction and ubiquitous computing · 4 · 2 since 2021Databases, data management, data science and information retrieval · 3
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | DissolveStereo: Coarse Depth Injection for Zero-Shot Stereo Video GenerationabstractGenerating high-quality stereo videos requires consistent depth perception and temporal coherence across frames. Despite advances in image and video synthesis using diffusion models, producing high-quality stereo videos remains a challenging task due to the difficulty of maintaining consistent temporal and spatial coherence between left and right views. We introduce DissolveStereo , a novel framework for zero-shot stereo video generation that leverages video diffusion priors without requiring paired training data. Our key innovations include a noisy restart strategy to initialize stereo-aware latent representations and an iterative refinement process that progressively harmonizes the latent space, addressing issues like temporal flickering and view inconsistencies. Importantly, we propose the use of dissolved depth maps to streamline latent space operations by reducing high-frequency depth information. Our comprehensive evaluations, including quantitative metrics and user studies, demonstrate that DissolveStereo produces high-quality stereo videos with enhanced depth consistency and temporal smoothness. In terms of epipolar consistency, our method achieves an 11.7% improvement in MEt3R score over the current state-of-the-art. Furthermore, user studies indicate strong perceptual gains over the previous arts, with an 8.0% higher perceived frame quality and 10.9% higher perceived temporal coherence. Our code is in https://github.com/shijianjian/DissolveStereo. Zhenyu Li 0007, Wenqing Cui, Ramzi Idoughi, Peter Wonka |
ACM Trans. Graph. | 6 |
| 2026 | HierRelTriple: Guiding Indoor Layout Generation With Hierarchical Relationship Triplet LossesabstractWe present a hierarchical triplet-based indoor relationship learning method, coined HierRelTriple, with a focus on spatial relationship learning. Existing approaches often depend on attention-based network designs and manually defined training objectives using handcrafted spatial rules or simplified pairwise relationships. However, these methods fail to capture complex, multi-object relationships found in real scenarios, leading to overcrowded or physically implausible arrangements. We introduce HierRelTriple, a hierarchical framework for modeling relational triplets, which first partitions functional regions and then automatically extracts three levels of spatial relationships: object-to-region (O2R), object-to-object (O2O), and corner-to-corner (C2C). By representing these relationships as geometric triplets and employing approaches based on Delaunay Triangulation to establish spatial priors, we derive IoU-based losses between denoised and ground-truth triplets and integrate them seamlessly into the diffusion denoising process. The joint formulation of inter-object distances, angular orientations, and spatial relationships enhances the physical realism of the generated scenes. Extensive experiments on unconditional layout synthesis, floorplan-conditioned layout generation, and scene rearrangement demonstrate that HierRelTriple improves spatial-relation metrics by over 15% and substantially reduces collisions and boundary violations compared to state-of-the-art methods. Kaifan Sun, Bingchen Yang, Peter Wonka, Jun Xiao 0005, Haiyong Jiang |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2026 | E$^{3}$3-Net: Efficient E(3)-Equivariant Normal Estimation NetworkabstractPoint cloud normal estimation is a fundamental task in 3D geometry processing, playing a crucial role in applications such as 3D reconstruction, object recognition, and surface analysis. While recent learning-based methods achieve notable advancements in normal prediction, they often overlook the critical aspect of equivariance. This oversight leads to inefficient learning of symmetric patterns inherent in geometric data. To address this issue, we propose E$^{3}$3-Net, an innovative neural network architecture designed to inherently achieve equivariance for normal estimation. We introduce an efficient random frame method, which significantly reduces the training resources required for this task to just 1/8 of previous work, while simultaneously enhancing prediction accuracy. Furthermore, we design a Gaussian-weighted loss function and a receptive-aware inference strategy that effectively leverage the local properties of point clouds, ensuring more precise and reliable normal estimation. Our method demonstrates superior performance across both synthetic and real-world datasets, consistently outperforming current state-of-the-art techniques by a substantial margin. Specifically, we achieve a 4% improvement in RMSE on the PCPNet dataset, 2.67% on the SceneNN dataset, and 2.44% on the FamousShape dataset, highlighting the robustness and scalability of E$^{3}$3-Net in diverse environments. Mingyang Zhao 0001, Weize Quan, Zhen Chen 0013, Dong-Ming Yan 0001, Peter Wonka |
IEEE Trans. Vis. Comput. Graph. | 6 |
| 2025 | 4Real-Video: Learning Generalizable Photo-Realistic 4D Video DiffusionabstractWe propose 4Real-Video, a novel framework for generating 4D videos, organized as a grid of video frames with both time and viewpoint axes. In this grid, each row contains frames sharing the same timestep, while each column contains frames from the same viewpoint. We propose a novel two-stream architecture. One stream performs viewpoint updates on columns, and the other stream performs temporal updates on rows. After each diffusion transformer layer, a synchronization layer exchanges information between the two token streams. We propose two implementations of the synchronization layer, using either hard or soft synchronization. This feedforward architecture improves upon previous work in three ways: higher inference speed, enhanced visual quality (measured by FVD, CLIP, and VideoScore), and improved temporal and viewpoint consistency (measured by VideoScore and Dust3R-Confidence). Chaoyang Wang 0001, Peiye Zhuang, Tuan Duc Ngo, Willi Menapace, Aliaksandr Siarohin, Michael Vasilkovsky, Ivan Skorokhodov, Sergey Tulyakov, Peter Wonka, Hsin-Ying Lee 0001 |
CVPR | 9 |
| 2025 | PrEditor3D: Fast and Precise 3D Shape EditingabstractWe propose a training-free approach to 3D editing that enables the editing of a single shape within a few minutes. The edited 3D mesh aligns well with the prompts, and remains identical for regions that are not intended to be altered. To this end, we first project the 3D object onto 4-view images and perform synchronized multi-view image editing along with user-guided text prompts and user-provided rough masks. However, the targeted regions to be edited are ambiguous due to projection from 3D to 2D. To ensure precise editing only in intended regions, we develop a 3D segmentation pipeline that detects edited areas in 3D space, followed by a merging algorithm to seamlessly integrate edited 3D regions with the original input. Extensive experiments demonstrate the superiority of our method over previous approaches, enabling fast, high-quality editing while preserving unintended regions. Ziya Erkoç, Can Gümeli, Chaoyang Wang 0001, Matthias Nießner, Angela Dai, Peter Wonka, Hsin-Ying Lee 0001, Peiye Zhuang |
CVPR | 6 |
| 2025 | Factored-NeuS: Reconstructing Surfaces, Illumination, and Materials of Possibly Glossy ObjectsabstractWe develop a method that recovers the surface, materials, and illumination of a scene from its posed multi-view images. In contrast to prior work, it does not require any additional data and can handle glossy objects or bright lighting. It is a progressive inverse rendering approach, which consists of three stages. In the first stage, we reconstruct the scene radiance and signed distance function (SDF) with a novel regularization strategy for specular reflections. We propose to explain a pixel color using both surface and volume rendering jointly, which allows for handling complex view-dependent lighting effects for surface reconstruction. In the second stage, we distill light visibility and indirect illumination from the learned SDF and radiance field using learnable mapping functions. Finally, we design a method for estimating the ratio of incoming direct light reflected in a specular manner and use it to reconstruct the materials and direct illumination. Experimental results demonstrate that the proposed method outperforms the current state-of-the-art in recovering surfaces, materials, and lighting without relying on any additional data. Ningjing Fan, Ivan Skorokhodov, Oleg Voynov, Savva Ignatyev, Evgeny Burnaev, Peter Wonka, Yiqun Wang 0001 |
CVPR | 7 |
| 2025 | VidSeg: Training-free Video Semantic Segmentation based on Diffusion ModelsabstractWe introduce the first training-free approach for Video Semantic Segmentation (VSS) based on pre-trained diffusion models, termed VidSeg. A growing research direction attempts to employ diffusion models to perform downstream vision tasks by exploiting their deep understanding of image semantics. Yet, the majority of these approaches have focused on image-related tasks like semantic segmentation, with less emphasis on video tasks such as VSS. Ideally, diffusion-based image semantic segmentation approaches can be applied to videos in a frame-by-frame manner. However, we find their performance on videos to be subpar due to the absence of any modeling of temporal information inherent in the video data. To this end, we tackle this problem and introduce a framework tailored for VSS based on pre-trained image and video diffusion models. We propose building a scene context model based on the diffusion features, where the model is autoregressively updated to adapt to scene changes. This context model predicts per-frame coarse segmentation maps that are temporally consistent. To refine these maps further, we propose a correspondence-based refinement strategy that aggregates predictions temporally, resulting in more confi- Abdelrahman Eldesokey, Mohit Mendiratta, Fangneng Zhan, Adam Kortylewski, Christian Theobalt, Peter Wonka |
CVPR | 7 |
| 2025 | Placeit3d: Language-Guided Object Placement in Real 3D ScenesabstractWe introduce the novel task of Language-Guided Object Placement in Real 3D Scenes. Our model is given a 3D scene's point cloud, a 3D asset, and a textual prompt broadly describing where the 3D asset should be placed. The task here is to find a valid placement for the 3D asset that respects the prompt. Compared with other language-guided localization tasks in 3D scenes such as grounding, this task has specific challenges: it is ambiguous because it has multiple valid solutions, and it requires reasoning about 3D geometric relationships and free space. We inaugurate this task by proposing a new benchmark and evaluation protocol. We also introduce a new dataset for training 3D LLMs on this task, as well as the first method to serve as a non-trivial baseline. We believe that this challenging task and our new benchmark could become part of the suite of benchmarks used to evaluate and compare generalist 3D LLM models. Ahmed Abdelreheem 0002, Filippo Aleotti, Jamie Watson, Zawar Qureshi, Abdelrahman Eldesokey, Peter Wonka, Gabriel J. Brostow, Sara Vicente, Guillermo Garcia-Hernando |
ICCV | 6 |
| 2025 | V2M4: 4D Mesh Animation Reconstruction from a Single Monocular VideoabstractWe present V2M4, a novel 4D reconstruction method that directly generates a usable 4D mesh animation asset from a single monocular video. Unlike existing approaches that rely on priors from multi-view image and video generation models, our method is based on native 3D mesh generation models. Naively applying 3D mesh generation models to generate a mesh for each frame in a 4D task can lead to issues such as incorrect mesh poses, misalignment of mesh appearance, and inconsistencies in mesh geometry and texture maps. To address these problems, we propose a structured workflow that includes camera search and mesh reposing, condition embedding optimization for mesh appearance refinement, pairwise mesh registration for topology consistency, and global texture map optimization for texture consistency. Our method outputs high-quality 4D animated assets that are compatible with mainstream graphics and game software. Experimental results across a variety of animation types and motion amplitudes demonstrate the generalization and effectiveness of our method. Project page: https://windvchen.github.io/V2M4/. Jianqi Chen, Biao Zhang 0005, Xiangjun Tang, Peter Wonka |
ICCV | 4 |
| 2025 | ZeroKey: Point-Level Reasoning and Zero-Shot 3D Keypoint Detection from Large Language ModelsabstractWe propose a novel zero-shot approach for keypoint detection on 3D shapes. Point-level reasoning on visual data is challenging as it requires precise localization capability, posing problems even for powerful models like DINO or CLIP. Traditional methods for 3D keypoint detection rely heavily on annotated 3D datasets and extensive supervised training, limiting their scalability and applicability to new categories or domains. In contrast, our method utilizes the rich knowledge embedded within Multi-Modal Large Language Models (MLLMs). Specifically, we demonstrate, for the first time, that pixel-level annotations used to train recent MLLMs can be exploited for both extracting and naming salient keypoints on 3D models without any ground truth labels or supervision. Experimental evaluations demonstrate that our approach achieves competitive performance on standard benchmarks compared to supervised methods, despite not requiring any 3D keypoint annotations during training. Our results highlight the potential of integrating language models for localized 3D shape understanding. This work opens new avenues for cross-modal learning and underscores the effectiveness of MLLMs in contributing to 3D computer vision challenges. Bingchen Gong, Diego Gomez 0003, Abdullah Hamdi, Abdelrahman Eldesokey, Ahmed Abdelreheem 0002, Peter Wonka, Maks Ovsjanikov |
ICCV | 6 |
| 2025 | Amodal Depth Anything: Amodal Depth Estimation in the WildabstractAmodal depth estimation aims to predict the depth of occluded (invisible) parts of objects in a scene. This task addresses the question of whether models can effectively perceive the geometry of occluded regions based on visible cues. Prior methods primarily rely on synthetic datasets and focus on metric depth estimation, limiting their generalization to real-world settings due to domain shifts and scalability challenges. In this paper, we propose a novel formulation of amodal depth estimation in the wild, focusing on relative depth prediction to improve model generalization across diverse natural images. We introduce a new large-scale dataset, Amodal Depth In the Wild (ADIW), created using a scalable pipeline that leverages segmentation datasets and compositing techniques. Depth maps are generated using large pre-trained depth models, and a scale-and-shift alignment strategy is employed to refine and blend depth predictions, ensuring consistency in ground-truth annotations. To tackle the amodal depth task, we present two complementary frameworks: Amodal-DAV2, a deterministic model based on Depth Anything V2, and Amodal-DepthFM, a generative model that integrates conditional flow matching principles. Our proposed frameworks effectively leverage the capabilities of large pre-trained models with minimal modifications to achieve high-quality amodal depth predictions. Experiments validate our design choices, demonstrating the flexibility of our models in generating diverse, plausible depth structures for occluded regions. Our method achieves a 69.5% improvement in accuracy over the previous SoTA on the ADIW dataset. Zhenyu Li 0007, Mykola Lavrenyuk, Shariq Farooq Bhat, Peter Wonka |
ICCV | 5 |
| 2025 | T2Bs: Text-to-Character Blendshapes via Video GenerationabstractWe present T2Bs, a framework for generating high-quality, animatable character head morphable models from text by combining static text-to-3D generation with video diffusion. Text-to-3D models produce detailed static geometry but lack motion synthesis, while video diffusion models generate motion with temporal and multi-view geometric inconsistencies. T2Bs bridges this gap by leveraging deformable 3D Gaussian splatting to align static 3D assets with video outputs. By constraining motion with static geometry and employing a view-dependent deformation MLP, T2Bs (i) outperforms existing 4D generation methods in accuracy and expressiveness while reducing video artifacts and view inconsistencies, and (ii) reconstructs smooth, coherent, fully registered 3D geometries designed to scale for building morphable models with diverse, realistic facial motions. This enables synthesizing expressive, animatable character heads that surpass current 4D generation techniques. Jiahao Luo, Chaoyang Wang 0001, Michael Vasilkovsky, Vladislav Shakhrai, Di Liu 0003, Peiye Zhuang, Sergey Tulyakov, Peter Wonka, Hsin-Ying Lee 0001, Jian Wang 0100 |
ICCV | 8 |
| 2025 | VoxelKP: A Voxel-Based Network Architecture for Human Keypoint Estimation in LiDAR DataabstractWe present \textit{VoxelKP}, a novel fully sparse network architecture tailored for human keypoint estimation in LiDAR data. The key challenge is that objects are distributed sparsely in 3D space, while human keypoint detection requires detailed local information wherever humans are present. We propose four novel ideas in this paper. First, we propose sparse selective kernels to capture multi-scale context. Second, we introduce sparse box-attention to focus on learning spatial correlations between keypoints within each human instance. Third, we incorporate a spatial encoding to leverage absolute 3D coordinates when projecting 3D voxels to a 2D grid encoding a bird's eye view. Finally, we propose hybrid feature learning to combine the processing of per-voxel features with sparse convolution. We evaluate our method on the Waymo dataset and achieve an improvement of $27\%$ on the MPJPE metric compared to the state-of-the-art, \textit{HUM3DIL}, trained on the same data, and $12\%$ against the state-of-the-art, \textit{GC-KPL}, pretrained on a $25\times$ larger dataset. To the best of our knowledge, \textit{VoxelKP} is the first single-staged, fully sparse network that is specifically designed for addressing the challenging task of 3D keypoint estimation from LiDAR data, achieving state-of-the-art performances. Our code is available at \url{https://github.com/shijianjian/VoxelKP}. Peter Wonka |
ICCV | 2 |
| 2025 | EditClip: Representation Learning for Image EditingabstractWe introduce EditCLIP, a novel representation-learning approach for image editing. Our method learns a unified representation of edits by jointly encoding an input image and its edited counterpart, effectively capturing their transformation. To evaluate its effectiveness, we employ EditCLIP to solve two tasks: exemplar-based image editing and automated edit evaluation. In exemplar-based image editing, we replace text-based instructions in InstructPix2Pix with EditCLIP embeddings computed from a reference exemplar image pair. Experiments demonstrate that our approach outperforms state-of-the-art methods while being more efficient and versatile. For automated evaluation, EditCLIP assesses image edits by measuring the similarity between the EditCLIP embedding of a given image pair and either a textual editing instruction or the EditCLIP embedding of another reference image pair. Experiments show that EditCLIP aligns more closely with human judgments than existing CLIP-based metrics, providing a reliable measure of edit quality and structural preservation. Aleksandar Cvejic, Abdelrahman Eldesokey, Peter Wonka |
ICCV | 4 |
| 2025 | Geometry DistributionsabstractNeural representations of 3D data have been widely adopted across various applications, particularly in recent work leveraging coordinate-based networks to model scalar or vector fields. However, these approaches face inherent challenges, such as handling thin structures and non-watertight geometries, which limit their flexibility and accuracy. In contrast, we propose a novel geometric data representation that models geometry as distributions-a powerful representation that makes no assumptions about surface genus, connectivity, or boundary conditions. Our approach uses diffusion models with a novel network architecture to learn surface point distributions, capturing fine-grained geometric details. We evaluate our representation qualitatively and quantitatively across various object types, demonstrating its effectiveness in achieving high geometric fidelity. Additionally, we explore applications using our representation, such as textured mesh representation, neural surface compression, dynamic object modeling, and rendering, highlighting its potential to advance 3D geometric learning. Biao Zhang 0005, Jing Ren 0004, Peter Wonka |
ICCV | 3 |
| 2025 | LaGeM: A Large Geometry Model for 3D Representation Learning and DiffusionabstractThis paper introduces a novel hierarchical autoencoder that maps 3D models into a highly compressed latent space. The hierarchical autoencoder is specifically designed to tackle the challenges arising from large-scale datasets and generative modeling using diffusion. Different from previous approaches that only work on a regular image or volume grid, our hierarchical autoencoder operates on unordered sets of vectors. Each level of the autoencoder controls different geometric levels of detail. We show that the model can be used to represent a wide range of 3D models while faithfully representing high-resolution geometry details. The training of the new architecture takes 0.70x time and 0.58x memory compared to the baseline.
We also explore how the new representation can be used for generative modeling. Specifically, we propose a cascaded diffusion framework where each stage is conditioned on the previous stage. Our design extends existing cascaded designs for image and volume grids to vector sets. Biao Zhang 0005, Peter Wonka |
ICLR | 2 |
| 2025 | Build-A-Scene: Interactive 3D Layout Control for Diffusion-Based Image GenerationabstractWe propose a diffusion-based approach for Text-to-Image (T2I) generation with interactive 3D layout control.
Layout control has been widely studied to alleviate the shortcomings of T2I diffusion models in understanding objects' placement and relationships from text descriptions.
Nevertheless, existing approaches for layout control are limited to 2D layouts, require the user to provide a static layout beforehand, and fail to preserve generated images under layout changes.
This makes these approaches unsuitable for applications that require 3D object-wise control and iterative refinements, e.g., interior design and complex scene generation.
To this end, we leverage the recent advancements in depth-conditioned T2I models and propose a novel approach for interactive 3D layout control.
We replace the traditional 2D boxes used in layout control with 3D boxes.
Furthermore, we revamp the T2I task as a multi-stage generation process, where at each stage, the user can insert, change, and move an object in 3D while preserving objects from earlier stages.
We achieve this through a novel Dynamic Self-Attention (DSA) module and a consistent 3D object translation strategy.
To evaluate our approach, we establish a benchmark and an evaluation protocol for interactive 3D layout control.
Experiments show that our approach can generate complicated scenes based on 3D layouts, outperforming the standard depth-conditioned T2I methods by two-folds on object generation success rate.
Moreover, it outperforms all methods in comparison on preserving objects under layout changes.
Project Page: https://abdo-eldesokey.github.io/build-a-scene/ Abdelrahman Eldesokey, Peter Wonka |
ICLR | 2 |
| 2025 | A3D: Does Diffusion Dream about 3D Alignment?abstractWe tackle the problem of text-driven 3D generation from a geometry alignment perspective. Given a set of text prompts, we aim to generate a collection of objects with semantically corresponding parts aligned across them. Recent methods based on Score Distillation have succeeded in distilling the knowledge from 2D diffusion models to high-quality representations of the 3D objects. These methods handle multiple text queries separately, and therefore the resulting objects have a high variability in object pose and structure. However, in some applications, such as 3D asset design, it may be desirable to obtain a set of objects aligned with each other. In order to achieve the alignment of the corresponding parts of the generated objects, we propose to embed these objects into a common latent space and optimize the continuous transitions between these objects. We enforce two kinds of properties of these transitions: smoothness of the transition and plausibility of the intermediate objects along the transition. We demonstrate that both of these properties are essential for good alignment. We provide several practical scenarios that benefit from alignment between the objects, including 3D editing and object hybridization, and experimentally demonstrate the effectiveness of our method. Savva Ignatyev, Nina Konovalova, Daniil Selikhanovych, Oleg Voynov, Nikolay Patakin, Ilya Olkov, Dmitry Senushkin, Alexey Artemov, Anton Konushin, Alexander Filippov, Peter Wonka, Evgeny Burnaev |
ICLR | 11 |
| 2025 | Mind-the-Glitch: Visual Correspondence for Detecting Inconsistencies in Subject-Driven GenerationabstractWe propose a novel approach for disentangling visual and semantic features from the backbones of pre-trained diffusion models, enabling visual correspondence in a manner analogous to the well-established semantic correspondence.
While diffusion model backbones are known to encode semantically rich features, they must also contain visual features to support their image synthesis capabilities.
However, isolating these visual features is challenging due to the absence of annotated datasets.
To address this, we introduce an automated pipeline that constructs image pairs with annotated semantic and visual correspondences based on existing subject-driven image generation datasets, and design a contrastive architecture to separate the two feature types.
Leveraging the disentangled representations, we propose a new metric, Visual Semantic Matching (VSM), that quantifies visual inconsistencies in subject-driven image generation.
Empirical results show that our approach outperforms global feature-based metrics such as CLIP, DINO, and vision--language models in quantifying visual inconsistencies while also enabling spatial localization of inconsistent regions.
To our knowledge, this is the first method that supports both quantification and localization of inconsistencies in subject-driven generation, offering a valuable tool for advancing this task. Abdelrahman Eldesokey, Aleksandar Cvejic, Bernard Ghanem, Peter Wonka |
NeurIPS | 4 |
| 2025 | Fused View-Time Attention and Feedforward Reconstruction for 4D Scene GenerationabstractWe propose the first framework capable of computing a 4D spatio-temporal grid of video frames and 3D Gaussian particles for each time step using a feed-forward architecture. Our architecture has two main components, a 4D video model and a 4D reconstruction model. In the first part, we analyze current 4D video diffusion architectures that perform spatial and temporal attention either sequentially or in parallel within a two-stream design. We highlight the limitations of existing approaches and introduce a novel fused architecture that performs spatial and temporal attention within a single layer. The key to our method is a sparse attention pattern, where tokens attend to others in the same frame, at the same timestamp, or from the same viewpoint.
In the second part, we extend existing 3D reconstruction algorithms by introducing a Gaussian head, a camera token replacement algorithm, and additional dynamic layers and training. Overall, we establish a new state of the art for 4D generation, improving both visual quality and reconstruction capability. Chaoyang Wang 0001, Ashkan Mirzaei, Vidit Goel, Willi Menapace, Aliaksandr Siarohin, Michael Vasilkovsky, Ivan Skorokhodov, Vladislav Shakhrai, Sergei Korolev, Sergey Tulyakov, Peter Wonka |
NeurIPS | 11 |
| 2025 | Autoregressive Generation of Static and Growing TreesabstractWe propose a transformer architecture and training strategy for tree generation. The architecture processes data at multiple resolutions and has an hourglass shape, with middle layers processing fewer tokens than outer layers. Similar to convolutional networks, we introduce longer-range skip connections to complement this multi-resolution approach. The key advantages of this architecture are the faster processing speed and lower memory consumption. We are, therefore, able to process more complex trees than would be possible with a vanilla transformer architecture. Furthermore, we extend this approach to perform image-to-tree and point-cloud-to-tree conditional generation and to simulate the tree growth processes, generating 4D trees. Empirical results validate our approach in terms of speed, memory consumption, and generation quality. Biao Zhang 0005, Jonathan Klein, Dominik L. Michels, Dong-Ming Yan 0001, Peter Wonka |
SIGGRAPH Asia | 6 |
| 2025 | 3DCoMPaT++: An Improved Large-Scale 3D Vision Dataset for Compositional RecognitionabstractIn this work, we present 3DCOMPAT++, a multimodal 2D/3D dataset with 160 million rendered views of more than 10 million stylized 3D shapes carefully annotated at the partinstance level, alongside matching RGB point clouds, 3D textured meshes, depth maps, and segmentation masks. 3DCOMPAT ++ covers 42 shape categories, 275 fine-grained part categories, and 293 fine-grained material classes that can be compositionally applied to parts of 3D objects. We render a subset of one million stylized shapes from four equally spaced views as well as four randomized views, leading to a total of 160 million renderings. Parts are segmented at the instance level, with coarse-grained and fine-grained semantic levels. We introduce a new task, called Grounded CoMPaT Recognition (GCR), to collectively recognize and ground compositions of materials on parts of 3D objects. Additionally, we report the outcomes of a data challenge organized at the CVPR conference, showcasing the winning method's utilization of a modified PointNet++model trained on 6D inputs, and exploring alternative techniques for GCR enhancement. We hope our work will help ease future research on compositional 3D Vision. The dataset and code have been made publicly available at https://3dcompat-dataset.org/v2/.3D vision, dataset, 3D modeling, multimodal learning, compositional learning. Habib Slim, Xiang Li 0046, Yuchen Li 0010, Mohamed Ayman, Ujjwal Upadhyay, Ahmed Abdelreheem 0002, Arpit Prajapati, Suhail Pothigara, Peter Wonka, Mohamed Elhoseiny 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 10 |
| 2025 | BrepGPT: Autoregressive B-rep Generation with Voronoi Half-PatchabstractBoundary representation (B-rep) is the de facto standard for CAD model representation in modern industrial design. The intricate coupling between geometric and topological elements in B-rep structures has forced existing generative methods to rely on cascaded multi-stage networks, resulting in error accumulation and computational inefficiency. We present BrepGPT, a single-stage autoregressive framework for B-rep generation. Our key innovation lies in the Voronoi Half-Patch (VHP) representation, which decomposes B-reps into unified local units by assigning geometry to nearest half-edges and sampling their next pointers. Unlike hierarchical representations that require multiple distinct encodings for different structural levels, our VHP representation facilitates unifying geometric attributes and topological relations in a single, coherent format. We further leverage dual VQ-VAEs to encode both vertex topology and Voronoi Half-Patches into vertex-based tokens, achieving a more compact sequential encoding. A decoder-only Transformer is then trained to autoregressively predict these tokens, which are subsequently mapped to vertex-based features and decoded into complete B-rep models. Experiments demonstrate that BrepGPT achieves state-of-the-art performance in unconditional B-rep generation. The framework also exhibits versatility in various applications, including conditional generation from category labels, point clouds, text descriptions, and images, as well as B-rep autocompletion and interpolation. Weize Quan, Biao Zhang 0005, Peter Wonka, Dong-Ming Yan 0001 |
ACM Trans. Graph. | 5 |
| 2025 | PS-CAD: Local Geometry Guidance via Prompting and Selection for CAD ReconstructionabstractReverse engineering CAD models from raw geometry is a classic but challenging research problem. In particular, reconstructing the CAD modeling sequence from point clouds provides great interpretability and convenience for editing. Analyzing previous work, we observed that a CAD modeling sequence represented by tokens and processed by a generative model does not have an immediate geometric interpretation. To improve upon this problem, we introduce geometric guidance into the reconstruction network. Our proposed model, PS-CAD, reconstructs the CAD modeling sequence one step at a time as illustrated in Figure 1 . At each step, we provide three forms of geometric guidance. First, we provide the geometry of surfaces where the current reconstruction differs from the complete model as a point cloud. This helps the framework to focus on regions that still need work. Second, we use geometric analysis to extract a set of planar prompts, that correspond to candidate surfaces where a CAD extrusion step could be started. Third, we present a step-wise sampling to generate multiple complete candidate CAD modeling steps instead of single-tokens without direct geometric interpretation. Our framework has three major components. Geometric guidance computation extracts the first two types of geometric guidance. Single-step reconstruction computes a single candidate CAD modeling step for each provided prompt. Single-step selection selects among the candidate CAD modeling steps. The process continues until the reconstruction is completed. Our quantitative results show a significant improvement across all metrics. For example, on the dataset DeepCAD, PS-CAD improves upon the best published SOTA method by reducing the geometry errors (CD and HD) by 10%, and the structural error (ECD metric) by about 13%. Bingchen Yang, Haiyong Jiang, Hao Pan 0001, Guosheng Lin, Jun Xiao 0005, Peter Wonka |
ACM Trans. Graph. | 6 |
| 2024 | Back to 3D: Few-Shot 3D Keypoint Detection with Back-Projected 2D FeaturesabstractWith the immense growth of dataset sizes and computing resources in recent years, so-called foundation models have become popular in NLP and vision tasks. In this work, we propose to explore foundation models for the task of key-point detection on 3D shapes. A unique characteristic of keypoint detection is that it requires semantic and geomet-ric awareness while demanding high localization accuracy. To address this problem, we propose, first, to back-project features from large pre-trained 2D vision models onto 3D shapes and employ them for this task. We show that we ob-tain robust 3D features that contain rich semantic information and analyze multiple candidate features stemming from different 2D foundation models. Second, we employ a key-point candidate optimization module which aims to match the average observed distribution of keypoints on the shape and is guided by the back-projected features. The resulting approach achieves a new state of the art for few-shot key-point detection on the KeyPointNet dataset, almost doubling the performance of the previous best methods. Thomas Wimmer 0001, Peter Wonka, Maks Ovsjanikov |
CVPR | 2 |
| 2024 | Functional DiffusionabstractWe propose functional diffusion, a generative diffusion model focused on infinite-dimensional function data samples. In contrast to previous work, functional diffusion works on samples that are represented by functions with a continuous domain. Functional diffusion can be seen as an extension of classical diffusion models to an infinite-dimensional domain. Functional diffusion is very versatile as images, videos, audio, 3D shapes, deformations, etc., can be handled by the same framework with minimal changes. In addition, functional diffusion is especially suited for irregular data or data defined in non-standard domains. In our work, we derive the necessary foundations for functional diffusion and propose a first implementation based on the transformer architecture. We show generative results on complicated signed distance functions and deformation functions defined on 3D surfaces. Biao Zhang 0005, Peter Wonka |
CVPR | 2 |
| 2024 | 4D-fy: Text-to-4D Generation Using Hybrid Score Distillation SamplingabstractRecent breakthroughs in text-to-4D generation rely on pre-trained text-to-image and text-to-video models to generate dynamic 3D scenes. However, current text-to-4D methods face a three-way tradeoff between the quality of scene appearance, 3D structure, and motion. For example, text-to-image models and their 3D-aware variants are trained on internet-scale image datasets and can be used to produce scenes with realistic appearance and 3D structure—but no motion. Text-to-video models are trained on relatively smaller video datasets and can produce scenes with motion, but poorer appearance and 3D structure. While these models have complementary strengths, they also have opposing weaknesses, making it difficult to combine them in a way that alleviates this three-way tradeoff. Here, we introduce hybrid score distillation sampling, an alternating optimization procedure that blends supervision signals from multiple pre-trained diffusion models and incorporates benefits of each for high-fidelity text-to-4D generation. Using hybrid SDS, we demonstrate synthesis of 4D scenes with compelling appearance, 3D structure, and motion. Sherwin Bahmani, Ivan Skorokhodov, Victor Rong, Gordon Wetzstein, Leonidas J. Guibas, Peter Wonka, Sergey Tulyakov, Jeong Joon Park, Andrea Tagliasacchi, David B. Lindell |
CVPR | 6 |
| 2024 | WinSyn: A High Resolution Testbed for Synthetic DataabstractWe present WinSyn, a unique dataset and testbed for cre-ating high-quality synthetic data with procedural modeling techniques. The dataset contains high-resolution pho-tographs of windows, selected from locations around the world, with 89,318 individual window crops showcasing diverse geometric and material characteristics. We evaluate a procedural model by training semantic segmentation networks on both synthetic and real images and then comparing their performances on a shared test set of real images. Specifically, we measure the difference in mean Intersection over Union (mIoU) and determine the effective number of real images to match synthetic data's training performance. We design a baseline procedural model as a benchmark and provide 21,290 synthetically generated images. By tuning the procedural model, key factors are identified which significantly influence the model's fidelity in replicating real-world scenarios. Importantly, we highlight the challenge of procedural modeling using current techniques, especially in their ability to replicate the spatial semantics of real-world scenarios. This insight is critical because of the potential of procedural models to bridge to hidden scene aspects such as depth, reflectivity, material properties, and lighting conditions. John Femiani 0001, Peter Wonka |
CVPR | 3 |
| 2024 | PatchFusion: An End-to-End Tile-Based Framework for High-Resolution Monocular Metric Depth EstimationabstractSingle image depth estimation is a foundational task in computer vision and generative modeling. However, prevailing depth estimation models grapple with accom-modating the increasing resolutions commonplace in to-day's consumer cameras and devices. Existing high-resolution strategies show promise, but they often face limi-tations, ranging from error propagation to the loss of high-frequency details. We present PatchFusion, a novel tile-based framework with three key components to improve the current state of the art: (1) A patch-wise fusion network that fuses a globally-consistent coarse prediction with finer, in-consistent tiled predictions via high-level feature guidance, (2) A Global-to-Local (G2L) module that adds vital con-text to the fusion network, discarding the need for patch selection heuristics, and (3) A Consistency-Aware Training (CAT) and Inference (CAI) approach, emphasizing patch overlap consistency and thereby eradicating the neces-sity for post-processing. Experiments on Unreal. Stereo 4K, MVS-Synth, and Middleburry 2014 demonstrate that our framework can generate high-resolution depth maps with intricate details. PatchFusion is independent of the base model for depth estimation. Notably, our framework built on top of SOTA ZoeDepth brings improvements for a total of 17.3% and 29.4% in terms of the root mean squared error (RMSE) on UnrealStereo4K and MVS-Synth, respectively. Zhenyu Li 0007, Shariq Farooq Bhat, Peter Wonka |
CVPR | 3 |
| 2024 | PatchRefiner: Leveraging Synthetic Data for Real-Domain High-Resolution Monocular Metric Depth Estimation
Zhenyu Li 0007, Shariq Farooq Bhat, Peter Wonka |
ECCV (67) | 3 |
| 2024 | Dissolving is Amplifying: Towards Fine-Grained Anomaly Detection
Hakim Ghazzai, Peter Wonka |
ECCV (59) | 5 |
| 2024 | LLM Blueprint: Enabling Text-to-Image Generation with Complex and Detailed PromptsabstractDiffusion-based generative models have significantly advanced text-to-image generation but encounter challenges when processing lengthy and intricate text prompts describing complex scenes with multiple objects. While excelling in generating images from short, single-object descriptions, these models often struggle to faithfully capture all the nuanced details within longer and more elaborate textual inputs. In response, we present a novel approach leveraging Large Language Models (LLMs) to extract critical components from text prompts, including bounding box coordinates for foreground objects, detailed textual descriptions for individual objects, and a succinct background context. These components form the foundation of our layout-to-image generation model, which operates in two phases. The initial Global Scene Generation utilizes object layouts and background context to create an initial scene but often falls short in faithfully representing object characteristics as specified in the prompts. To address this limitation, we introduce an Iterative Refinement Scheme that iteratively evaluates and refines box-level content to align them with their textual descriptions, recomposing objects as needed to ensure consistency. Our evaluation on complex prompts featuring multiple objects demonstrates a substantial improvement in recall compared to baseline diffusion models. This is further validated by a user study, underscoring the efficacy of our approach in generating coherent and detailed scenes from intricate textual inputs. Our iterative framework offers a promising solution for enhancing text-to-image generation models' fidelity with lengthy, multifaceted descriptions, opening new possibilities for accurate and diverse image synthesis from textual inputs. Hanan Gani, Shariq Farooq Bhat, Muzammal Naseer, Salman Khan 0001, Peter Wonka |
ICLR | 5 |
| 2024 | Magic123: One Image to High-Quality 3D Object Generation Using Both 2D and 3D Diffusion PriorsabstractWe present ``Magic123'', a two-stage coarse-to-fine approach for high-quality, textured 3D mesh generation from a single image in the wild using *both 2D and 3D priors*. In the first stage, we optimize a neural radiance field to produce a coarse geometry. In the second stage, we adopt a memory-efficient differentiable mesh representation to yield a high-resolution mesh with a visually appealing texture. In both stages, the 3D content is learned through reference-view supervision and novel-view guidance by a joint 2D and 3D diffusion prior. We introduce a trade-off parameter between the 2D and 3D priors to control the details and 3D consistencies of the generation. Magic123 demonstrates a significant improvement over previous image-to-3D techniques, as validated through extensive experiments on diverse synthetic and real-world images. Guocheng Qian, Jinjie Mai, Abdullah Hamdi, Jian Ren 0005, Aliaksandr Siarohin, Bing Li 0024, Hsin-Ying Lee 0001, Ivan Skorokhodov, Peter Wonka, Sergey Tulyakov, Bernard Ghanem |
ICLR | 9 |
| 2024 | Vivid-ZOO: Multi-View Video Generation with Diffusion ModelabstractWhile diffusion models have shown impressive performance in 2D image/video generation, diffusion-based Text-to-Multi-view-Video (T2MVid) generation remains underexplored. The new challenges posed by T2MVid generation lie in the lack of massive captioned multi-view videos and the complexity of modeling such multi-dimensional distribution. To this end, we propose a novel diffusion-based pipeline that generates high-quality multi-view videos centered around a dynamic 3D object from text. Specifically, we factor the T2MVid problem into viewpoint-space and time components. Such factorization allows us to combine and reuse layers of advanced pre-trained multi-view image and 2D video diffusion models to ensure multi-view consistency as well as temporal coherence for the generated multi-view videos, largely reducing the training cost. We further introduce alignment modules to align the latent spaces of layers from the pre-trained multi-view and the 2D video diffusion models, addressing the reused layers' incompatibility that arises from the domain gap between 2D and multi-view data. In support of this and future research, we further contribute a captioned multi-view video dataset. Experimental results demonstrate that our method generates high-quality multi-view videos, exhibiting vivid motions, temporal coherence, and multi-view consistency, given a variety of text prompts. Bing Li 0024, Cheng Zheng 0002, Wenxuan Zhu, Jinjie Mai, Biao Zhang 0005, Peter Wonka, Bernard Ghanem |
NeurIPS | 6 |
| 2024 | ScanEnts3D: Exploiting Phrase-to-3D-Object Correspondences for Improved Visio-Linguistic Models in 3D ScenesabstractThe two popular datasets ScanRefer [20] and ReferIt3D [5] connect natural language to real-world 3D scenes. In this paper, we curate a complementary dataset extending both the aforementioned ones. We associate all objects mentioned in a referential sentence with their underlying instances inside a 3D scene. In contrast, previous work did this only for a single object per sentence. Our Scan Entities in 3D (ScanEnts3D) dataset provides explicit correspondences between 369k objects across 84k referential sentences, covering 705 real-world scenes. We propose novel architecture modifications and losses that enable learning from this new type of data and improve the performance for both neural listening and language generation. For neural listening, we improve the SoTA in both the Nr3D and ScanRefer benchmarks by 4.3% and 5.0%, respectively. For language generation, we improve the SoTA by 13.2 CIDEr points on the Nr3D benchmark. For both of these tasks, the new type of data is only used to improve training, but no additional annotations are required at inference time. Our introduced dataset is available on the project’s webpage at https://scanents3d.github.io/. Ahmed Abdelreheem 0002, Kyle Olszewski, Hsin-Ying Lee 0001, Peter Wonka, Panos Achlioptas |
WACV | 4 |
| 2024 | State of the Art on Diffusion Models for Visual ComputingabstractAbstract The field of visual computing is rapidly advancing due to the emergence of generative artificial intelligence (AI), which unlocks unprecedented capabilities for the generation, editing, and reconstruction of images, videos, and 3D scenes. In these domains, diffusion models are the generative AI architecture of choice. Within the last year alone, the literature on diffusion‐based tools and applications has seen exponential growth and relevant papers are published across the computer graphics, computer vision, and AI communities with new works appearing daily on arXiv. This rapid growth of the field makes it difficult to keep up with all recent developments. The goal of this state‐of‐the‐art report (STAR) is to introduce the basic mathematical concepts of diffusion models, implementation details and design choices of the popular Stable Diffusion model, as well as overview important aspects of these generative AI tools, including personalization, conditioning, inversion, among others. Moreover, we give a comprehensive overview of the rapidly growing literature on diffusion‐based generation and editing, categorized by the type of generated medium, including 2D images, videos, 3D objects, locomotion, and 4D scenes. Finally, we discuss available datasets, metrics, open challenges, and social implications. This STAR provides an intuitive starting point to explore this exciting topic for researchers, artists, and practitioners alike. Ryan Po, Wang Yifan 0001, Vladislav Golyanik, Kfir Aberman, Jonathan T. Barron, Amit Bermano, Eric R. Chan, Tali Dekel, Aleksander Holynski, Angjoo Kanazawa, C. Karen Liu, Lingjie Liu, Ben Mildenhall, Matthias Nießner, Björn Ommer, Christian Theobalt, Peter Wonka, Gordon Wetzstein |
Comput. Graph. Forum | 17 |
| 2024 | Deep Learning-Based Image and Video Inpainting: A Survey
Weize Quan, Jiaxi Chen, Dong-Ming Yan 0001, Peter Wonka |
Int. J. Comput. Vis. | 5 |
| 2023 | PET-NeuS: Positional Encoding Tri-Planes for Neural SurfacesabstractA signed distance function (SDF) parametrized by an MLP is a common ingredient of neural surface reconstruction. We build on the successful recent method NeuS to extend it by three new components. The first component is to borrow the tri-plane representation from EG3D and represent signed distance fields as a mixture of tri-planes and MLPs instead of representing it with MLPs only. Using tri-planes leads to a more expressive data structure but will also introduce noise in the reconstructed surface. The second component is to use a new type of positional encoding with learnable weights to combat noise in the reconstruction process. We divide the features in the tri-plane into multiple frequency scales and modulate them with sin and cos functions of different frequencies. The third component is to use learnable convolution operations on the tri-plane features using self-attention convolution to produce features with different frequency bands. The experiments show that PET-NeuS achieves high-fidelity surface reconstruction on standard datasets. Following previous work and using the Chamfer metric as the most important way to measure surface reconstruction quality, we are able to improve upon the NeuS baseline by 57% on Nerf-synthetic (0.84 compared to 1.97) and by 15.5% on DTU (0.71 compared to 0.84). The qualitative evaluation reveals how our method can better control the interference of high-frequency noise. Yiqun Wang 0001, Ivan Skorokhodov, Peter Wonka |
CVPR | 3 |
| 2023 | 3DAvatarGAN: Bridging Domains for Personalized Editable AvatarsabstractModern 3D-GANs synthesize geometry and texture by training on large-scale datasets with a consistent structure. Training such models on stylized, artistic data, with often unknown, highly variable geometry, and camera information has not yet been shown possible. Can we train a 3D GAN on such artistic data, while maintaining multi-view consistency and texture quality? To this end, we propose an adaptation framework, where the source domain is a pre-trained 3D-GAN, while the target domain is a 2D-GAN trained on artistic datasets. We, then, distill the knowledge from a 2D generator to the source 3D generator. To do that, we first propose an optimization-based method to align the distributions of camera parameters across domains. Second, we propose regularizations necessary to learn high-quality texture, while avoiding degenerate geometric solutions, such as flat shapes. Third, we show a deformation-based technique for modeling exaggerated geometry of artistic domains, enabling-as a byproduct- personalized geometric editing. Finally, we propose a novel inversion method for 3D-GANs linking the latent spaces of the source and the target domains. Our contributions-for the first time-allow for the generation, editing, and animation of personalized artistic 3D avatars on artistic datasets. Project Page: https:/rameenabdal.github.io/3DAvatarGAN Rameen Abdal, Hsin-Ying Lee 0001, Peihao Zhu 0001, Menglei Chai, Aliaksandr Siarohin, Peter Wonka, Sergey Tulyakov |
CVPR | 6 |
| 2023 | VIVE3D: Viewpoint-Independent Video Editing using 3D-Aware GANsabstractWe introduce VIVE3D, a novel approach that extends the capabilities of image-based 3D GANs to video editing and is able to represent the input video in an identity-preserving and temporally consistent way. We propose two new building blocks. First, we introduce a novel GAN inversion technique specifically tailored to 3D GANs by jointly embedding multiple frames and optimizing for the camera parameters. Second, besides traditional semantic face edits (e.g. for age and expression), we are the first to demonstrate edits that show novel views of the head enabled by the inherent prop-erties of 3D GANs and our optical flow-guided compositing technique to combine the head with the background video. Our experiments demonstrate that VIVE3D generates high-fidelity face edits at consistent quality from a range of camera viewpoints which are composited with the original video in a temporally and spatially consistent manner. Anna Frühstück, Nikolaos Sarafianos, Yuanlu Xu, Peter Wonka, Tony Tung |
CVPR | 4 |
| 2023 | SATR: Zero-Shot Semantic Segmentation of 3D ShapesabstractWe explore the task of zero-shot semantic segmentation of 3D shapes by using large-scale off-the-shelf 2D image recognition models. Surprisingly, we find that modern zero-shot 2D object detectors are better suited for this task than contemporary text/image similarity predictors or even zero-shot 2D segmentation networks. Our key finding is that it is possible to extract accurate 3D segmentation maps from multi-view bounding box predictions by using the topological properties of the underlying surface. For this, we develop the Segmentation Assignment with Topological Reweighting (SATR) algorithm and evaluate it on ShapeNetPart and our proposed FAUST benchmarks. SATR achieves state-of-the-art performance and outperforms a baseline algorithm by 1.3% and 4% average mIoU on the FAUST coarse and fine-grained benchmarks, respectively, and by 5.2% average mIoU on the ShapeNetPart benchmark. Our source code and data will be publicly released. Project webpage: https://samir55.github.io/SATR/. Ahmed Abdelreheem 0002, Ivan Skorokhodov, Maks Ovsjanikov, Peter Wonka |
ICCV | 4 |
| 2023 | 3D generation on ImageNet
Ivan Skorokhodov, Aliaksandr Siarohin, Yinghao Xu 0001, Jian Ren 0005, Hsin-Ying Lee 0001, Peter Wonka, Sergey Tulyakov |
ICLR | 6 |
| 2023 | SLIBO-Net: Floorplan Reconstruction via Slicing Box Representation with Local Geometry RegularizationabstractThis paper focuses on improving the reconstruction of 2D floorplans from unstructured 3D point clouds. We identify opportunities for enhancement over the existing methods in three main areas: semantic quality, efficient representation, and local geometric details. To address these, we presents SLIBO-Net, an innovative approach to reconstructing 2D floorplans from unstructured 3D point clouds. We propose a novel transformer-based architecture that employs an efficient floorplan representation, providing improved room shape supervision and allowing for manageable token numbers. By incorporating geometric priors as a regularization mechanism and post-processing step, we enhance the capture of local geometric details. We also propose a scale-independent evaluation metric, correcting the discrepancy in error treatment between varying floorplan sizes. Our approach notably achieves a new state-of-the-art on the Structure3D dataset. The resultant floorplans exhibit enhanced semantic plausibility, substantially improving the overall quality and realism of the reconstructions. Our code and dataset are available online. Jheng-Wei Su, Kuei-Yu Tung, Chihan Peng, Peter Wonka, Hung-Kuo Chu |
NeurIPS | 4 |
| 2023 | Zero-Shot 3D Shape CorrespondenceabstractWe propose a novel zero-shot approach to computing correspondences between 3D shapes. Existing approaches mainly focus on isometric and near-isometric shape pairs (e.g., human vs. human), but less attention has been given to strongly non-isometric and inter-class shape matching (e.g., human vs. cow). To this end, we introduce a fully automatic method that exploits the exceptional reasoning capabilities of recent foundation models in language and vision to tackle difficult shape correspondence problems. Our approach comprises multiple stages. First, we classify the 3D shapes in a zero-shot manner by feeding rendered shape views to a language-vision model (e.g., BLIP2) to generate a list of class proposals per shape. These proposals are unified into a single class per shape by employing the reasoning capabilities of ChatGPT. Second, we attempt to segment the two shapes in a zero-shot manner, but in contrast to the co-segmentation problem, we do not require a mutual set of semantic regions. Instead, we propose to exploit the in-context learning capabilities of ChatGPT to generate two different sets of semantic regions for each shape and a semantic mapping between them. This enables our approach to match strongly non-isometric shapes with significant differences in geometric structure. Finally, we employ the generated semantic mapping to produce coarse correspondences that can further be refined by the functional maps framework to produce dense point-to-point maps. Our approach, despite its simplicity, produces highly plausible results in a zero-shot manner, especially between strongly non-isometric shapes. Ahmed Abdelreheem 0002, Abdelrahman Eldesokey, Maks Ovsjanikov, Peter Wonka |
SIGGRAPH Asia | 4 |
| 2023 | ART-Owen ScramblingabstractWe present a novel algorithm for implementing Owen-scrambling, combining the generation and distribution of the scrambling bits in a single self-contained compact process. We employ a context-free grammar to build a binary tree of symbols, and equip each symbol with a scrambling code that affects all descendant nodes. We nominate the grammar of adaptive regular tiles (ART) derived from the repetition-avoiding Thue-Morse word, and we discuss its potential advantages and shortcomings. Our algorithm has many advantages, including random access to samples, fixed time complexity, GPU friendliness, and scalability to any memory budget. Further, it provides two unique features over known methods: it admits optimization, and it is in-vertible, enabling screen-space scrambling of the high-dimensional Sobol sampler. Abdalla G. M. Ahmed, Matt Pharr, Peter Wonka |
ACM Trans. Graph. | 3 |
| 2023 | Analysis and Synthesis of Digital Dyadic SequencesabstractWe explore the space of matrix-generated (0, m , 2)-nets and (0, 2)-sequences in base 2, also known as digital dyadic nets and sequences. In computer graphics, they are arguably leading the competition for use in rendering. We provide a complete characterization of the design space and count the possible number of constructions with and without considering possible reorderings of the point set. Based on this analysis, we then show that every digital dyadic net can be reordered into a sequence, together with a corresponding algorithm. Finally, we present a novel family of self-similar digital dyadic sequences, to be named ξ -sequences, that spans a subspace with fewer degrees of freedom. Those ξ -sequences are extremely efficient to sample and compute, and we demonstrate their advantages over the classic Sobol (0, 2)-sequence. Abdalla G. M. Ahmed, Mikhail Skopenkov, Markus Hadwiger, Peter Wonka |
ACM Trans. Graph. | 4 |
| 2023 | 3DShape2VecSet: A 3D Shape Representation for Neural Fields and Generative Diffusion ModelsabstractWe introduce 3DShape2VecSet, a novel shape representation for neural fields designed for generative diffusion models. Our shape representation can encode 3D shapes given as surface models or point clouds, and represents them as neural fields. The concept of neural fields has previously been combined with a global latent vector, a regular grid of latent vectors, or an irregular grid of latent vectors. Our new representation encodes neural fields on top of a set of vectors. We draw from multiple concepts, such as the radial basis function representation, and the cross attention and self-attention function, to design a learnable representation that is especially suitable for processing with transformers. Our results show improved performance in 3D shape encoding and 3D shape generative modeling tasks. We demonstrate a wide variety of generative applications: unconditioned generation, category-conditioned generation, text-conditioned generation, point-cloud completion, and image-conditioned generation. Code: https://1zb.github.io/3DShape2VecSet/. Biao Zhang 0005, Jiapeng Tang, Matthias Nießner, Peter Wonka |
ACM Trans. Graph. | 4 |
| 2023 | Finding Nano-Ötzi: Cryo-Electron Tomography Visualization Guided by Learned SegmentationabstractCryo-electron tomography (cryo-ET) is a new 3D imaging technique with unprecedented potential for resolving submicron structural details. Existing volume visualization methods, however, are not able to reveal details of interest due to low signal-to-noise ratio. In order to design more powerful transfer functions, we propose leveraging soft segmentation as an explicit component of visualization for noisy volumes. Our technical realization is based on semi-supervised learning, where we combine the advantages of two segmentation algorithms. First, the weak segmentation algorithm provides good results for propagating sparse user-provided labels to other voxels in the same volume and is used to generate dense pseudo-labels. Second, the powerful deep-learning-based segmentation algorithm learns from these pseudo-labels to generalize the segmentation to other unseen volumes, a task that the weak segmentation algorithm fails at completely. The proposed volume visualization uses deep-learning-based segmentation as a component for segmentation-aware transfer function design. Appropriate ramp parameters can be suggested automatically through frequency distribution analysis. Furthermore, our visualization uses gradient-free ambient occlusion shading to further suppress the visual presence of noise, and to give structural detail the desired prominence. The cryo-ET data studied in our technical experiments are based on the highest-quality tilted series of intact SARS-CoV-2 virions. Our technique shows the high impact in target sciences for visual data analysis of very noisy volumes that cannot be visualized with existing techniques. Ngan V. T. Nguyen, Ciril Bohak, Dominik Engel 0001, Peter Mindek, Ondrej Strnad, Peter Wonka, Timo Ropinski, Ivan Viola |
IEEE Trans. Vis. Comput. Graph. | 6 |
| 2022 | Learning to Construct 3D Building Wireframes from 3D Line Clouds
Yicheng Luo, Jing Ren 0004, Xuefei Zhe, Peter Wonka, Linchao Bao |
BMVC | 6 |
| 2022 | InsetGAN for Full-Body Image GenerationabstractWhile GANs can produce photo-realistic images in ideal conditions for certain domains, the generation of full-body human images remains difficult due to the diversity of identities, hairstyles, clothing, and the variance in pose. In-stead of modeling this complex domain with a single GAN, we propose a novel method to combine multiple pretrained GANs, where one GAN generates a global canvas (e.g., human body) and a set of specialized GANs, or insets, focus on different parts (e.g., faces, shoes) that can be seamlessly inserted onto the global canvas. We model the problem as jointly exploring the respective latent spaces such that the generated images can be combined, by inserting the parts from the specialized generators onto the global canvas, without introducing seams. We demonstrate the setup by combining a full body GAN with a dedicated high-quality face GAN to produce plausible-looking humans. We evalu-ate our results with quantitative metrics and user studies. Anna Frühstück, Krishna Kumar Singh, Eli Shechtman, Niloy J. Mitra, Peter Wonka, Jingwan Lu |
CVPR | 5 |
| 2022 | On the Robustness of Quality Measures for GANs
Motasem Alfarra, Juan C. Pérez, Anna Frühstück, Philip Torr 0001, Peter Wonka, Bernard Ghanem |
ECCV (17) | 5 |
| 2022 | LocalBins: Improving Depth Estimation by Learning Local Distributions
Shariq Farooq Bhat, Ibraheem Alhashim, Peter Wonka |
ECCV (1) | 3 |
| 2022 | 3D CoMPaT: Composition of Materials on Parts of 3D Things
Yuchen Li 0010, Ujjwal Upadhyay, Habib Slim, Ahmed Abdelreheem 0002, Arpit Prajapati, Suhail Pothigara, Peter Wonka, Mohamed Elhoseiny 0001 |
ECCV (8) | 7 |
| 2022 | HairNet: Hairstyle Transfer with Pose Changes
Peihao Zhu 0001, Rameen Abdal, John Femiani 0001, Peter Wonka |
ECCV (16) | 4 |
| 2022 | Training Data Generating Networks: Shape Reconstruction via Bi-level Optimization
Biao Zhang 0005, Peter Wonka |
ICLR | 2 |
| 2022 | Mind the Gap: Domain Gap Control for Single Shot Domain Adaptation for Generative Adversarial Networks
Peihao Zhu 0001, Rameen Abdal, John Femiani 0001, Peter Wonka |
ICLR | 4 |
| 2022 | HF-NeuS: Improved Surface Reconstruction Using High-Frequency DetailsabstractNeural rendering can be used to reconstruct implicit representations of shapes without 3D supervision. However, current neural surface reconstruction methods have difficulty learning high-frequency geometry details, so the reconstructed shapes are often over-smoothed. We develop HF-NeuS, a novel method to improve the quality of surface reconstruction in neural rendering. We follow recent work to model surfaces as signed distance functions (SDFs). First, we offer a derivation to analyze the relationship between the SDF, the volume density, the transparency function, and the weighting function used in the volume rendering equation and propose to model transparency as a transformed SDF. Second, we observe that attempting to jointly encode high-frequency and low-frequency components in a single SDF leads to unstable optimization. We propose to decompose the SDF into base and displacement functions with a coarse-to-fine strategy to increase the high-frequency details gradually. Finally, we design an adaptive optimization strategy that makes the training process focus on improving those regions near the surface where the SDFs have artifacts. Our qualitative and quantitative results show that our method can reconstruct fine-grained surface details and obtain better surface reconstruction quality than the current state of the art. Code available at https://github.com/yiqun-wang/HFS. Yiqun Wang 0001, Ivan Skorokhodov, Peter Wonka |
NeurIPS | 3 |
| 2022 | 3DILG: Irregular Latent Grids for 3D Generative ModelingabstractWe propose a new representation for encoding 3D shapes as neural fields. The representation is designed to be compatible with the transformer architecture and to benefit both shape reconstruction and shape generation. Existing works on neural fields are grid-based representations with latents being defined on a regular grid. In contrast, we define latents on irregular grids which facilitates our representation to be sparse and adaptive. In the context of shape reconstruction from point clouds, our shape representation built on irregular grids improves upon grid-based methods in terms of reconstruction accuracy. For shape generation, our representation promotes high-quality shape generation using auto-regressive probabilistic models. We show different applications that improve over the current state of the art. First, we show results of probabilistic shape reconstruction from a single higher resolution image. Second, we train a probabilistic model conditioned on very low resolution images. Third, we apply our model to category-conditioned generation. All probabilistic experiments confirm that we are able to generate detailed and high quality shapes to yield the new state of the art in generative 3D shape modeling. Biao Zhang 0005, Matthias Nießner, Peter Wonka |
NeurIPS | 3 |
| 2022 | EpiGRAF: Rethinking training of 3D GANsabstractA recent trend in generative modeling is building 3D-aware generators from 2D image collections. To induce the 3D bias, such models typically rely on volumetric rendering, which is expensive to employ at high resolutions. Over the past months, more than ten works have addressed this scaling issue by training a separate 2D decoder to upsample a low-resolution image (or a feature tensor) produced from a pure 3D generator. But this solution comes at a cost: not only does it break multi-view consistency (i.e., shape and texture change when the camera moves), but it also learns geometry in low fidelity. In this work, we show that obtaining a high-resolution 3D generator with SotA image quality is possible by following a completely different route of simply training the model patch-wise. We revisit and improve this optimization scheme in two ways. First, we design a location- and scale-aware discriminator to work on patches of different proportions and spatial positions. Second, we modify the patch sampling strategy based on an annealed beta distribution to stabilize training and accelerate the convergence. The resulting model, named EpiGRAF, is an efficient, high-resolution, pure 3D generator, and we test it on four datasets (two introduced in this work) at (256^2) and (512^2) resolutions. It obtains state-of-the-art image quality, high-fidelity geometry and trains ({\approx})2.5 faster than the upsampler-based counterparts. Code/data/visualizations: https://universome.github.io/epigraf. Ivan Skorokhodov, Sergey Tulyakov, Yiqun Wang 0001, Peter Wonka |
NeurIPS | 4 |
| 2022 | Self-Supervised Learning of Domain Invariant Features for Depth EstimationabstractWe tackle the problem of unsupervised synthetic-to-real domain adaptation for single image depth estimation. An essential building block of single image depth estimation is an encoder-decoder task network that takes RGB images as input and produces depth maps as output. In this paper, we propose a novel training strategy to force the task network to learn domain invariant representations in a self-supervised manner. Specifically, we extend self-supervised learning from traditional representation learning, which works on images from a single domain, to domain invariant representation learning, which works on images from two different domains by utilizing an image-to-image translation network. Firstly, we use an image-to-image translation network to transfer domain-specific styles between synthetic and real domains. This style transfer operation allows us to obtain similar images from the different domains. Secondly, we jointly train our task network and Siamese network with the same images from the different domains to obtain domain invariance for the task network. Finally, we fine-tune the task network using labeled synthetic and unlabeled real-world data. Our training strategy yields improved generalization capability in the real-world domain. We carry out an extensive evaluation on two popular datasets for depth estimation, KITTI and Make3D. The results demonstrate that our proposed method outperforms the state-of-the-art on all metrics, e.g. by 14.7% on Sq Rel on KITTI. The source code and model weights will be made available. Hiroyasu Akada, Shariq Farooq Bhat, Ibraheem Alhashim, Peter Wonka |
WACV | 4 |
| 2022 | RLSS: A Deep Reinforcement Learning Algorithm for Sequential Scene GenerationabstractWe present RLSS: a reinforcement learning algorithm for sequential scene generation. This is based on employing the proximal policy optimization (PPO) algorithm for generative problems. In particular, we consider how to effectively reduce the action space by including a greedy search algorithm in the learning process. Our experiments demonstrate that our method converges for a relatively large number of actions and learns to generate scenes with predefined design objectives. This approach is placing objects iteratively in the virtual scene. In each step, the network chooses which objects to place and selects positions which result in maximal reward. A high reward is assigned if the last action resulted in desired properties whereas the violation of constraints is penalized. We demonstrate the capability of our method to generate plausible and diverse scenes efficiently by solving indoor planning problems and generating Angry Birds levels. Azimkhon Ostonov, Peter Wonka, Dominik L. Michels |
WACV | 2 |
| 2022 | Gaussian Blue NoiseabstractAmong the various approaches for producing point distributions with blue noise spectrum, we argue for an optimization framework using Gaussian kernels. We show that with a wise selection of optimization parameters, this approach attains unprecedented quality, provably surpassing the current state of the art attained by the optimal transport (BNOT) approach. Further, we show that our algorithm scales smoothly and feasibly to high dimensions while maintaining the same quality, realizing unprecedented high-quality high-dimensional blue noise sets. Finally, we show an extension to adaptive sampling. Abdalla G. M. Ahmed, Jing Ren 0004, Peter Wonka |
ACM Trans. Graph. | 3 |
| 2022 | Large-Scale Architectural Asset Extraction from Panoramic ImageryabstractWe present a system to extract architectural assets from large-scale collections of panoramic imagery. We automatically rectify and crop parts of the panoramic image that contain dominant planes, and then use object detection to extract assets such as façades and windows. We also provide various tools to identify attributes of the assets to determine the asset quality and index the assets for search. In addition, we propose a User Interface (UI) to visualize and query assets. Finally, we present applications for urban modeling and texture synthesis. Peihao Zhu 0001, Wamiq Para, Anna Frühstück, John Femiani 0001, Peter Wonka |
IEEE Trans. Vis. Comput. Graph. | 5 |
| 2021 | AdaBins: Depth Estimation Using Adaptive BinsabstractWe address the problem of estimating a high quality dense depth map from a single RGB input image. We start out with a baseline encoder-decoder convolutional neural network architecture and pose the question of how the global processing of information can help improve overall depth estimation. To this end, we propose a transformer-based architecture block that divides the depth range into bins whose center value is estimated adaptively per image. The final depth values are estimated as linear combinations of the bin centers. We call our new building block AdaBins. Our results show a decisive improvement over the state-of-the-art on several popular depth datasets across all metrics. We also validate the effectiveness of the proposed block with an ablation study and provide the code and corresponding pre-trained weights of the new state-of-the-art model. Shariq Farooq Bhat, Ibraheem Alhashim, Peter Wonka |
CVPR | 3 |
| 2021 | Fast Sinkhorn Filters: Using Matrix Scaling for Non-Rigid Shape Correspondence With Functional MapsabstractIn this paper, we provide a theoretical foundation for pointwise map recovery from functional maps and highlight its relation to a range of shape correspondence methods based on spectral alignment. With this analysis in hand, we develop a novel spectral registration technique: Fast Sinkhorn Filters, which allows for the recovery of accurate and bijective pointwise correspondences with a superior time and memory complexity in comparison to existing approaches. Our method combines the simple and concise representation of correspondence using functional maps with the matrix scaling schemes from computational optimal transport. By exploiting the sparse structure of the kernel matrices involved in the transport map computation, we provide an efficient trade-off between acceptable accuracy and complexity for the problem of dense shape correspondence, while promoting bijectivity.1 Gautam Pai 0001, Jing Ren 0004, Simone Melzi, Peter Wonka, Maks Ovsjanikov |
CVPR | 4 |
| 2021 | Point Cloud Instance Segmentation Using Probabilistic EmbeddingsabstractIn this paper we propose a new framework for point cloud instance segmentation. Our framework has two steps: an embedding step and a clustering step. In the embedding step, our main contribution is to propose a probabilistic embedding space for point cloud embedding. Specifically, each point is represented as a tri-variate normal distribution. In the clustering step, we propose a novel loss function, which benefits both the semantic segmentation and the clustering. Our experimental results show important improvements to the SOTA, i.e., 3.1% increased average per-category mAP on the PartNet dataset. Biao Zhang 0005, Peter Wonka |
CVPR | 2 |
| 2021 | Labels4Free: Unsupervised Segmentation using StyleGANabstractWe propose an unsupervised segmentation framework for StyleGAN generated objects. We build on two main observations. First, the features generated by StyleGAN hold valuable information that can be utilized towards training segmentation networks. Second, the foreground and background can often be treated to be largely independent and be swapped across images to produce plausible composited images. For our solution, we propose to augment the StyleGAN2 generator architecture with a segmentation branch and to split the generator into a foreground and background network. This enables us to generate soft segmentation masks for the foreground object in an unsupervised fashion. On multiple object classes, we report comparable results against state-of-the-art supervised segmentation networks, while against the best unsupervised segmentation approach we demonstrate a clear improvement, both in qualitative and quantitative metricsProject Page : https:/rameenabdal.github.io/Labels4Free Rameen Abdal, Peihao Zhu 0001, Niloy J. Mitra, Peter Wonka |
ICCV | 4 |
| 2021 | Flow-Guided Video Inpainting with Scene TemplatesabstractWe consider the problem of filling in missing spatiotemporal regions of a video. We provide a novel flow-based solution by introducing a generative model of images in relation to the scene (without missing regions) and mappings from the scene to images. We use the model to jointly infer the scene template, a 2D representation of the scene, and the mappings. This ensures consistency of the frame-to-frame flows generated to the underlying scene, reducing geometric distortions in flow based inpainting. The template is mapped to the missing regions in the video by a new (L2-L1) interpolation scheme, creating crisp inpaintings and reducing common blur and distortion artifacts. We show on two benchmark datasets that our approach out-performs state-of-the-art quantitatively and in user studies.1 Dong Lao, Peihao Zhu 0001, Peter Wonka, Ganesh Sundaramoorthi |
ICCV | 3 |
| 2021 | Generative Layout Modeling using Constraint GraphsabstractWe propose a new generative model for layout generation. We generate layouts in three steps. First, we generate the layout elements as nodes in a layout graph. Second, we compute constraints between layout elements as edges in the layout graph. Third, we solve for the final layout using constrained optimization. For the first two steps, we build on recent transformer architectures. The layout optimization implements the constraints efficiently. We show three practical contributions compared to the state of the art: our work requires no user input, produces higher quality layouts, and enables many novel capabilities for conditional layout generation. Wamiq Para, Paul Guerrero 0001, Leonidas J. Guibas, Peter Wonka |
ICCV | 5 |
| 2021 | IntraTomo: Self-supervised Learning-based Tomography via Sinogram Synthesis and PredictionabstractWe propose IntraTomo, a powerful framework that combines the benefits of learning-based and model-based approaches for solving highly ill-posed inverse problems in the Computed Tomography (CT) context. IntraTomo is composed of two core modules: a novel sinogram prediction module, and a geometry refinement module, which are applied iteratively. In the first module, the unknown density field is represented as a continuous and differentiable function, parameterized by a deep neural network. This network is learned, in a self-supervised fashion, from the incomplete or/and degraded input sinogram. After getting estimated through the sinogram prediction module, the density field is consistently refined in the second module using local and non-local geometrical priors. With these two core modules, we show that IntraTomo significantly outperforms existing approaches on several ill-posed inverse problems, such as limited angle tomography with a range of 45 degrees, sparse view tomographic reconstruction with as few as eight views, or super-resolution tomography with eight times increased resolution. The experiments on simulated and real data show that our approach can achieve results of unprecedented quality. Guangming Zang, Ramzi Idoughi, Rui Li 0054, Peter Wonka, Wolfgang Heidrich |
ICCV | 4 |
| 2021 | SketchGen: Generating Constrained CAD SketchesabstractComputer-aided design (CAD) is the most widely used modeling approach for technical design. The typical starting point in these designs is 2D sketches which can later be extruded and combined to obtain complex three-dimensional assemblies. Such sketches are typically composed of parametric primitives, such as points, lines, and circular arcs, augmented with geometric constraints linking the primitives, such as coincidence, parallelism, or orthogonality. Sketches can be represented as graphs, with the primitives as nodes and the constraints as edges. Training a model to automatically generate CAD sketches can enable several novel workflows, but is challenging due to the complexity of the graphs and the heterogeneity of the primitives and constraints. In particular, each type of primitive and constraint may require a record of different size and parameter types.We propose SketchGen as a generative model based on a transformer architecture to address the heterogeneity problem by carefully designing a sequential language for the primitives and constraints that allows distinguishing between different primitive or constraint types and their parameters, while encouraging our model to re-use information across related parameters, encoding shared structure. A particular highlight of our work is the ability to produce primitives linked via constraints that enables the final output to be further regularized via a constraint solver. We evaluate our model by demonstrating constraint prediction for given sets of primitives and full sketch generation from scratch, showing that our approach significantly out performs the state-of-the-art in CAD sketch generation. Wamiq Para, Shariq Farooq Bhat, Paul Guerrero 0001, Niloy J. Mitra, Leonidas J. Guibas, Peter Wonka |
NeurIPS | 7 |
| 2021 | Computational Design of Lightweight Trusses
Caigui Jiang, Chengcheng Tang, Hans-Peter Seidel, Renjie Chen 0001, Peter Wonka |
Comput. Aided Des. | 5 |
| 2021 | Discrete Optimization for Shape MatchingabstractAbstract We propose a novel discrete solver for optimizing functional map‐based energies, including descriptor preservation and promoting structural properties such as area‐preservation, bijectivity and Laplacian commutativity among others. Unlike the commonly‐used continuous optimization methods, our approach enforces the functional map to be associated with a pointwise correspondence as a hard constraint, which provides a stronger link between optimized properties of functional and point‐to‐point maps. Under this hard constraint, our solver obtains functional maps with lower energy values compared to the standard continuous strategies. Perhaps more importantly, the recovered pointwise maps from our discrete solver preserve the optimized for functional properties and are thus of higher overall quality. We demonstrate the advantages of our discrete solver on a range of energies and shape categories, compared to existing techniques for promoting pointwise maps within the functional map framework. Finally, with this solver in hand, we introduce a novel Effective Functional Map Refinement (EFMR) method which achieves the state‐of‐the‐art accuracy on the SHREC'19 benchmark. Jing Ren 0004, Simone Melzi, Peter Wonka, Maks Ovsjanikov |
Comput. Graph. Forum | 3 |
| 2021 | Customized Summarizations of Visual Data CollectionsabstractAbstract We propose a framework to generate customized summarizations of visual data collections, such as collections of images, materials, 3D shapes, and 3D scenes. We assume that the elements in the visual data collections can be mapped to a set of vectors in a feature space, in which a fitness score for each element can be defined, and we pose the problem of customized summarizations as selecting a subset of these elements. We first describe the design choices a user should be able to specify for modeling customized summarizations and propose a corresponding user interface. We then formulate the problem as a constrained optimization problem with binary variables and propose a practical and fast algorithm based on the alternating direction method of multipliers (ADMM). Our results show that our problem formulation enables a wide variety of customized summarizations, and that our solver is both significantly faster than state‐of‐the‐art commercial integer programming solvers and produces better solutions than fast relaxation‐based solvers. Mengke Yuan, Bernard Ghanem, Dong-Ming Yan 0001, Baoyuan Wu, Xiaopeng Zhang 0001, Peter Wonka |
Comput. Graph. Forum | 6 |
| 2021 | Manhattan Room Layout Reconstruction from a Single $360^{\circ }$ Image: A Comparative Study of State-of-the-Art Methods
Chuhang Zou, Jheng-Wei Su, Chihan Peng, Alex Colburn, Qi Shan, Peter Wonka, Hung-Kuo Chu, Derek Hoiem |
Int. J. Comput. Vis. | 6 |
| 2021 | StyleFlow: Attribute-conditioned Exploration of StyleGAN-Generated Images using Conditional Continuous Normalizing FlowsabstractHigh-quality, diverse, and photorealistic images can now be generated by unconditional GANs (e.g., StyleGAN). However, limited options exist to control the generation process using (semantic) attributes while still preserving the quality of the output. Further, due to the entangled nature of the GAN latent space, performing edits along one attribute can easily result in unwanted changes along other attributes. In this article, in the context of conditional exploration of entangled latent spaces, we investigate the two sub-problems of attribute-conditioned sampling and attribute-controlled editing. We present StyleFlow as a simple, effective, and robust solution to both the sub-problems by formulating conditional exploration as an instance of conditional continuous normalizing flows in the GAN latent space conditioned by attribute features. We evaluate our method using the face and the car latent space of StyleGAN, and demonstrate fine-grained disentangled edits along various attributes on both real photographs and StyleGAN generated images. For example, for faces, we vary camera pose, illumination variation, expression, facial hair, gender, and age. Finally, via extensive qualitative and quantitative comparisons, we demonstrate the superiority of StyleFlow over prior and several concurrent works. Project Page and Video: https://rameenabdal.github.io/StyleFlow . Rameen Abdal, Peihao Zhu 0001, Niloy J. Mitra, Peter Wonka |
ACM Trans. Graph. | 4 |
| 2021 | Optimizing dyadic netsabstractWe explore the space of (0, m , 2)-nets in base 2 commonly used for sampling. We present a novel constructive algorithm that can exhaustively generate all nets --- up to m -bit resolution --- and thereby compute the exact number of distinct nets. We observe that the construction algorithm holds the key to defining a transformation operation that lets us transform one valid net into another one. This enables the optimization of digital nets using arbitrary objective functions. For example, we define an analytic energy function for blue noise, and use it to generate nets with high-quality blue-noise frequency power spectra. We also show that the space of (0, 2)-sequences is significantly smaller than nets with the same number of points, which drastically limits the optimizability of sequences. Abdalla G. M. Ahmed, Peter Wonka |
ACM Trans. Graph. | 2 |
| 2021 | Intuitive and efficient roof modeling for reconstruction and synthesisabstractWe propose a novel and flexible roof modeling approach that can be used for constructing planar 3D polygon roof meshes. Our method uses a graph structure to encode roof topology and enforces the roof validity by optimizing a simple but effective planarity metric we propose. This approach is significantly more efficient than using general purpose 3D modeling tools such as 3ds Max or SketchUp, and more powerful and expressive than specialized tools such as the straight skeleton. Our optimization-based formulation is also flexible and can accommodate different styles and user preferences for roof modeling. We showcase two applications. The first application is an interactive roof editing framework that can be used for roof design or roof reconstruction from aerial images. We highlight the efficiency and generality of our approach by constructing a mesh-image paired dataset consisting of 2539 roofs. Our second application is a generative model to synthesize new roof meshes from scratch. We use our novel dataset to combine machine learning and our roof optimization techniques, by using transformers and graph convolutional networks to model roof topology, and our roof optimization methods to enforce the planarity constraint. Jing Ren 0004, Biao Zhang 0005, Bojian Wu, Jianqiang Huang 0001, Lubin Fan, Maks Ovsjanikov, Peter Wonka |
ACM Trans. Graph. | 7 |
| 2021 | Barbershop: GAN-based image compositing using segmentation masksabstractSeamlessly blending features from multiple images is extremely challenging because of complex relationships in lighting, geometry, and partial occlusion which cause coupling between different parts of the image. Even though recent work on GANs enables synthesis of realistic hair or faces, it remains difficult to combine them into a single, coherent, and plausible image rather than a disjointed set of image patches. We present a novel solution to image blending, particularly for the problem of hairstyle transfer, based on GAN-inversion. We propose a novel latent space for image blending which is better at preserving detail and encoding spatial information, and propose a new GAN-embedding algorithm which is able to slightly modify images to conform to a common segmentation mask. Our novel representation enables the transfer of the visual properties from multiple reference images including specific details such as moles and wrinkles, and because we do image blending in a latent-space we are able to synthesize images that are coherent. Our approach avoids blending artifacts present in other approaches and finds a globally consistent image. Our results demonstrate a significant improvement over the current state of the art in a user study, with users preferring our blending solution over 95 percent of the time. Source code for the new approach is available at https://zpdesu.github.io/Barbershop. Peihao Zhu 0001, Rameen Abdal, John Femiani 0001, Peter Wonka |
ACM Trans. Graph. | 4 |
| 2021 | Modeling in the Time of COVID-19: Statistical and Rule-based Mesoscale ModelsabstractWe present a new technique for the rapid modeling and construction of scientifically accurate mesoscale biological models. The resulting 3D models are based on a few 2D microscopy scans and the latest knowledge available about the biological entity, represented as a set of geometric relationships. Our new visual-programming technique is based on statistical and rule-based modeling approaches that are rapid to author, fast to construct, and easy to revise. From a few 2D microscopy scans, we determine the statistical properties of various structural aspects, such as the outer membrane shape, the spatial properties, and the distribution characteristics of the macromolecular elements on the membrane. This information is utilized in the construction of the 3D model. Once all the imaging evidence is incorporated into the model, additional information can be incorporated by interactively defining the rules that spatially characterize the rest of the biological entity, such as mutual interactions among macromolecules, and their distances and orientations relative to other structures. These rules are defined through an intuitive 3D interactive visualization as a visual-programming feedback loop. We demonstrate the applicability of our approach on a use case of the modeling procedure of the SARS-CoV-2 virion ultrastructure. This atomistic model, which we present here, can steer biological research to new promising directions in our efforts to fight the spread of the virus. Ngan V. T. Nguyen, Ondrej Strnad, Tobias Klein, Deng Luo, Ruwayda Alharbi, Peter Wonka, Martina Maritan, Peter Mindek, Ludovic Autin, David S. Goodsell, Ivan Viola |
IEEE Trans. Vis. Comput. Graph. | 6 |
| 2020 | Image2StyleGAN++: How to Edit the Embedded Images?abstractWe propose Image2StyleGAN++, a flexible image editing framework with many applications. Our framework extends the recent Image2StyleGAN in three ways. First, we introduce noise optimization as a complement to the W+ latent space embedding. Our noise optimization can restore high frequency features in images and thus significantly improves the quality of reconstructed images, e.g. a big increase of PSNR from 20 dB to 45 dB. Second, we extend the global W+ latent space embedding to enable local embeddings. Third, we combine embedding with activation tensor manipulation to perform high quality local edits along with global semantic edits on images. Such edits motivate various high quality image editing applications, e.g. image reconstruction, image inpainting, image crossover, local style transfer, image editing using scribbles, and attribute level feature transfer. Examples of the edited images are shown across the paper for visual inspection. Rameen Abdal, Yipeng Qin, Peter Wonka |
CVPR | 3 |
| 2020 | Disentangled Image Generation Through Structured Noise InjectionabstractWe explore different design choices for injecting noise into generative adversarial networks (GANs) with the goal of disentangling the latent space. Instead of traditional approaches, we propose feeding multiple noise codes through separate fully-connected layers respectively. The aim is restricting the influence of each noise code to specific parts of the generated image. We show that disentanglement in the first layer of the generator network leads to disentanglement in the generated image. Through a grid-based structure, we achieve several aspects of disentanglement without complicating the network architecture and without requiring labels. We achieve spatial disentanglement, scale-space disentanglement, and disentanglement of the foreground object from the background style allowing fine-grained control over the generated images. Examples include changing facial expressions in face images, changing beak length in bird images, and changing car dimensions in car images. This empirically leads to better disentanglement scores than state-of-the-art methods on the FFHQ dataset. Yazeed Alharbi, Peter Wonka |
CVPR | 2 |
| 2020 | StructEdit: Learning Structural Shape VariationsabstractLearning to encode differences in the geometry and (topological) structure of the shapes of ordinary objects is key to generating semantically plausible variations of a given shape, transferring edits from one shape to another, and for many other applications in 3D content creation. The common approach of encoding shapes as points in a high-dimensional latent feature space suggests treating shape differences as vectors in that space. Instead, we treat shape differences as primary objects in their own right and propose to encode them in their own latent space. In a setting where the shapes themselves are encoded in terms of fine-grained part hierarchies, we demonstrate that a separate encoding of shape deltas or differences provides a principled way to deal with inhomogeneities in the shape space due to different combinatorial part structures, while also allowing for compactness in the representation, as well as edit abstraction and transfer. Our approach is based on a conditional variational autoencoder for encoding and decoding shape deltas, conditioned on a source shape. We demonstrate the effectiveness and robustness of our approach in multiple shape modification and generation tasks, and provide comparison and ablation studies on the PartNet dataset, one of the largest publicly available 3D datasets. Kaichun Mo, Paul Guerrero 0001, Li Yi 0001, Hao Su 0001, Peter Wonka, Niloy J. Mitra, Leonidas J. Guibas |
CVPR | 5 |
| 2020 | TomoFluid: Reconstructing Dynamic Fluid From Sparse View VideosabstractVisible light tomography is a promising and increasingly popular technique for fluid imaging. However, the use of a sparse number of viewpoints in the capturing setups makes the reconstruction of fluid flows very challenging. In this paper, we present a state-of-the-art 4D tomographic reconstruction framework that integrates several regularizers into a multi-scale matrix free optimization algorithm. In addition to existing regularizers, we propose two new regularizers for improved results: a regularizer based on view interpolation of projected images and a regularizer to encourage reprojection consistency. We demonstrate our method with extensive experiments on both simulated and real data. Guangming Zang, Ramzi Idoughi, Congli Wang, Anthony Bennett, Jianguo Du 0003, Scott Skeen, William L. Roberts, Peter Wonka, Wolfgang Heidrich |
CVPR | 8 |
| 2020 | SEAN: Image Synthesis With Semantic Region-Adaptive NormalizationabstractWe propose semantic region-adaptive normalization (SEAN), a simple but effective building block for Generative Adversarial Networks conditioned on segmentation masks that describe the semantic regions in the desired output image. Using SEAN normalization, we can build a network architecture that can control the style of each semantic region individually, e.g., we can specify one style reference image per region. SEAN is better suited to encode, transfer, and synthesize style than the best previous method in terms of reconstruction quality, variability, and visual quality. We evaluate SEAN on multiple datasets and report better quantitative metrics (e.g. FID, PSNR) than the current state of the art. SEAN also pushes the frontier of interactive image editing. We can interactively edit images by changing segmentation masks or the style for any given region. We can also interpolate styles from two reference images per region. Peihao Zhu 0001, Rameen Abdal, Yipeng Qin, Peter Wonka |
CVPR | 4 |
| 2020 | How Does Lipschitz Regularization Influence GAN Training?
Yipeng Qin, Niloy J. Mitra, Peter Wonka |
ECCV (16) | 3 |
| 2020 | Consistent ZoomOut: Efficient Spectral Map SynchronizationabstractAbstract In this paper, we propose a novel method, which we call C onsistent Z oom O ut , for efficiently refining correspondences among deformable 3D shape collections, while promoting the resulting map consistency. Our formulation is closely related to a recent unidirectional spectral refinement framework, but naturally integrates map consistency constraints into the refinement. Beyond that, we show further that our formulation can be adapted to recover the underlying isometry among near‐isometric shape collections with a theoretical guarantee, which is absent in the other spectral map synchronization frameworks. We demonstrate that our method improves the accuracy compared to the competing methods when synchronizing correspondences in both near‐isometric and heterogeneous shape collections, but also significantly outperforms the baselines in terms of map consistency. Ruqi Huang, Jing Ren 0004, Peter Wonka, Maks Ovsjanikov |
Comput. Graph. Forum | 3 |
| 2020 | Photorealistic Material Editing Through Direct Image ManipulationabstractAbstract Creating photorealistic materials for light transport algorithms requires carefully fine‐tuning a set of material properties to achieve a desired artistic effect. This is typically a lengthy process that involves a trained artist with specialized knowledge. In this work, we present a technique that aims to empower novice and intermediate‐level users to synthesize high‐quality photorealistic materials by only requiring basic image processing knowledge. In the proposed workflow, the user starts with an input image and applies a few intuitive transforms (e.g., colorization, image inpainting) within a 2D image editor of their choice, and in the next step, our technique produces a photorealistic result that approximates this target image. Our method combines the advantages of a neural network‐augmented optimizer and an encoder neural network to produce high‐quality output results within 30 seconds. We also demonstrate that it is resilient against poorly‐edited target images and propose a simple extension to predict image sequences with a strict time budget of 1–2 seconds per image. Károly Zsolnai-Fehér, Peter Wonka, Michael Wimmer 0001 |
Comput. Graph. Forum | 2 |
| 2020 | PLADE: A Plane-Based Descriptor for Point Cloud Registration With Small OverlapabstractTraditional point cloud registration methods require large overlap between scans, which imposes strict constraints on data acquisition. To facilitate registration, users have to carefully position scanners to ensure sufficient overlap. In this article, we propose to use high-level structural information (i.e., plane/line features and their interrelationship) for registration, which is capable of registering point clouds with small overlap, allowing more freedom in data acquisition. We design a novel plane-/line-based descriptor dedicated to establishing structure-level correspondences between point clouds. Based on this descriptor, we propose a simple but effective registration algorithm. We also provide a data set of real-world scenes containing a larger number of scans with a wide range of overlap. Experiments and comparisons with state-of-the-art methods on various data sets reveal that our method is superior to existing techniques. Though the proposed algorithm outperforms state-of-the-art methods on the most challenging data set, the point cloud registration problem is still far from being solved, leaving significant room for improvement and future work. Liangliang Nan, Renbo Xia, Peter Wonka |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2020 | Screen-space blue-noise diffusion of monte carlo sampling error via hierarchical ordering of pixelsabstractWe present a novel technique for diffusing Monte Carlo sampling error as a blue noise in screen space. We show that automatic diffusion of sampling error can be achieved by ordering the pixels in a way that preserves locality, such as Morton's Z-ordering, and assigning the samples to the pixels from successive sub-sequences of a single low-discrepancy sequence, thus securing well-distributed samples for each pixel, local neighborhoods, and the whole image. We further show that a blue-noise distribution of the error is attainable by scrambling the Z-ordering to induce isotropy. We present an efficient technique to implement this hierarchical scrambling by defining a context-free grammar that describes infinite self-similar lookup trees. Our concept is scalable to arbitrary image resolutions, sample dimensions, and sample count, and supports progressive and adaptive sampling. Abdalla G. M. Ahmed, Peter Wonka |
ACM Trans. Graph. | 2 |
| 2020 | MapTree: recovering multiple solutions in the space of mapsabstractIn this paper we propose an approach for computing multiple high-quality near-isometric dense correspondences between a pair of 3D shapes. Our method is fully automatic and does not rely on user-provided landmarks or descriptors. This allows us to analyze the full space of maps and extract multiple diverse and accurate solutions, rather than optimizing for a single optimal correspondence as done in most previous approaches. To achieve this, we propose a compact tree structure based on the spectral map representation for encoding and enumerating possible rough initializations, and a novel efficient approach for refining them to dense pointwise maps. This leads to a new method capable of both producing multiple high-quality correspondences across shapes and revealing the symmetry structure of a shape without a priori information. In addition, we demonstrate through extensive experiments that our method is robust and results in more accurate correspondences than state-of-the-art for shape matching and symmetry detection. Jing Ren 0004, Simone Melzi, Maks Ovsjanikov, Peter Wonka |
ACM Trans. Graph. | 4 |
| 2020 | MGCN: descriptor learning using multiscale GCNsabstractWe propose a novel framework for computing descriptors for characterizing points on three-dimensional surfaces. First, we present a new non-learned feature that uses graph wavelets to decompose the Dirichlet energy on a surface. We call this new feature Wavelet Energy Decomposition Signature (WEDS). Second, we propose a new Multiscale Graph Convolutional Network (MGCN) to transform a non-learned feature to a more discriminative descriptor. Our results show that the new descriptor WEDS is more discriminative than the current state-of-the-art non-learned descriptors and that the combination of WEDS and MGCN is better than the state-of-the-art learned descriptors. An important design criterion for our descriptor is the robustness to different surface discretizations including triangulations with varying numbers of vertices. Our results demonstrate that previous graph convolutional networks significantly overfit to a particular resolution or even a particular triangulation, but MGCN generalizes well to different surface discretizations. In addition, MGCN is compatible with previous descriptors and it can also be used to improve the performance of other descriptors, such as the heat kernel signature, the wave kernel signature, or the local point signature. Yiqun Wang 0001, Jing Ren 0004, Dong-Ming Yan 0001, Jianwei Guo 0003, Xiaopeng Zhang 0001, Peter Wonka |
ACM Trans. Graph. | 6 |
| 2020 | Selection Expressions for Procedural ModelingabstractWe introduce a new approach for procedural modeling. Our main idea is to select shapes using selection-expressions instead of simple string matching used in current state-of-the-art grammars like CGA shape and CGA++. A selection-expression specifies how to select a potentially complex subset of shapes from a shape hierarchy, e.g., "select all tall windows in the second floor of the main building facade". This new way of modeling enables us to express modeling ideas in their global context rather than traditional rules that operate only locally. To facilitate selection-based procedural modeling we introduce the procedural modeling language SelEx. An important implication of our work is that enforcing important constraints, such as alignment and same size constraints can be done by construction. Therefore, our procedural descriptions can generate facade and building variations without violating alignment and sizing constraints that plague the current state of the art. While the procedural modeling of architecture is our main application domain, we also demonstrate that our approach nicely extends to other man-made objects. Haiyong Jiang, Dong-Ming Yan 0001, Xiaopeng Zhang 0001, Peter Wonka |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2019 | Latent Filter Scaling for Multimodal Unsupervised Image-To-Image TranslationabstractIn multimodal unsupervised image-to-image translation tasks, the goal is to translate an image from the source domain to many images in the target domain. We present a simple method that produces higher quality images than current state-of-the-art while maintaining the same amount of multimodal diversity. Previous methods follow the unconditional approach of trying to map the latent code directly to a full-size image. This leads to complicated network architectures with several introduced hyperparameters to tune. By treating the latent code as a modifier of the convolutional filters, we produce multimodal output while maintaining the traditional Generative Adversarial Network (GAN) loss and without additional hyperparameters. The only tuning required by our method controls the tradeoff between variability and quality of generated images. Furthermore, we achieve disentanglement between source domain content and target domain style for free as a by-product of our formulation. We perform qualitative and quantitative experiments showing the advantages of our method compared with the state-of-the art on multiple benchmark image-to-image translation datasets. Yazeed Alharbi, Neil Smith, Peter Wonka |
CVPR | 3 |
| 2019 | DuLa-Net: A Dual-Projection Network for Estimating Room Layouts From a Single RGB PanoramaabstractWe present a deep learning framework, called DuLa-Net, to predict Manhattan-world 3D room layouts from a single RGB panorama. To achieve better prediction accuracy, our method leverages two projections of the panorama at once, namely the equirectangular panorama-view and the perspective ceiling-view, that each contains different clues about the room layouts. Our network architecture consists of two encoder-decoder branches for analyzing each of the two views. In addition, a novel feature fusion structure is proposed to connect the two branches, which are then jointly trained to predict the 2D floor plans and layout heights. To learn more complex room layouts, we introduce the Realtor360 dataset that contains panoramas of Manhattan-world room layouts with different numbers of corners. Experimental results show that our work outperforms recent state-of-the-art in prediction accuracy and performance, especially in the rooms with non-cuboid layouts. Shang-Ta Yang, Fu-En Wang, Chihan Peng, Peter Wonka, Min Sun 0001, Hung-Kuo Chu |
CVPR | 4 |
| 2019 | Image2StyleGAN: How to Embed Images Into the StyleGAN Latent Space?abstractWe propose an efficient algorithm to embed a given image into the latent space of StyleGAN. This embedding enables semantic image editing operations that can be applied to existing photographs. Taking the StyleGAN trained on the FFHD dataset as an example, we show results for image morphing, style transfer, and expression transfer. Studying the results of the embedding algorithm provides valuable insights into the structure of the StyleGAN latent space. We propose a set of experiments to test what class of images can be embedded, how they are embedded, what latent space is suitable for embedding, and if the embedding is semantically meaningful. Rameen Abdal, Yipeng Qin, Peter Wonka |
ICCV | 3 |
| 2019 | Local Editing of Procedural ModelsabstractAbstract Procedural modeling is used across many industries for rapid 3D content creation. However, professional procedural tools often lack artistic control, requiring manual edits on baked results, diminishing the advantages of a procedural modeling pipeline. Previous approaches to enable local artistic control require special annotations of the procedural system and manual exploration of potential edit locations. Therefore, we propose a novel approach to discover meaningful and non‐redundant good edit locations (GELs). We introduce a bottom‐up algorithm for finding GELs directly from the attributes in procedural models, without special annotations. To make attribute edits at GELs persistent, we analyze their local spatial context and construct a meta‐locator to uniquely specify their structure. Meta‐locators are calculated independently per attribute, making them robust against changes in the procedural system. Functions on meta‐locators enable intuitive and robust multi‐selections. Finally, we introduce an algorithm to transfer meta‐locators to a different procedural model. We show that our approach greatly simplifies the exploration of the local edit space, and we demonstrate its usefulness in a user study and multiple real‐world examples. Markus Lipp, Matthias Specht, Cheryl Lau, Peter Wonka, Pascal Müller |
Comput. Graph. Forum | 4 |
| 2019 | Structured Regularization of Functional Map ComputationsabstractAbstract We consider the problem of non‐rigid shape matching using the functional map framework. Specifically, we analyze a commonly used approach for regularizing functional maps, which consists in penalizing the failure of the unknown map to commute with the Laplace‐Beltrami operators on the source and target shapes. We show that this approach has certain undesirable fundamental theoretical limitations, and can be undefined even for trivial maps in the smooth setting. Instead we propose a novel, theoretically well‐justified approach for regularizing functional maps, by using the notion of the resolvent of the Laplacian operator. In addition, we provide a natural one‐parameter family of regularizers, that can be easily tuned depending on the expected approximate isometry of the input shape pair. We show on a wide range of shape correspondence scenarios that our novel regularization leads to an improvement in the quality of the estimated functional, and ultimately pointwise correspondences before and after commonly‐used refinement techniques. Jing Ren 0004, Mikhail Panine, Peter Wonka, Maks Ovsjanikov |
Comput. Graph. Forum | 3 |
| 2019 | TileGAN: synthesis of large-scale non-homogeneous texturesabstractWe tackle the problem of texture synthesis in the setting where many input images are given and a large-scale output is required. We build on recent generative adversarial networks and propose two extensions in this paper. First, we propose an algorithm to combine outputs of GANs trained on a smaller resolution to produce a large-scale plausible texture map with virtually no boundary artifacts. Second, we propose a user interface to enable artistic control. Our quantitative and qualitative results showcase the generation of synthesized high-resolution maps consisting of up to hundreds of megapixels as a case in point. Anna Frühstück, Ibraheem Alhashim, Peter Wonka |
ACM Trans. Graph. | 3 |
| 2019 | ZoomOut: spectral upsampling for efficient shape correspondenceabstractWe present a simple and efficient method for refining maps or correspondences by iterative upsampling in the spectral domain that can be implemented in a few lines of code. Our main observation is that high quality maps can be obtained even if the input correspondences are noisy or are encoded by a small number of coefficients in a spectral basis. We show how this approach can be used in conjunction with existing initialization techniques across a range of application scenarios, including symmetry detection, map refinement across complete shapes, non-rigid partial shape matching and function transfer. In each application we demonstrate an improvement with respect to both the quality of the results and the computational speed compared to the best competing methods, with up to two orders of magnitude speed-up in some applications. We also demonstrate that our method is both robust to noisy input and is scalable with respect to shape complexity. Finally, we present a theoretical justification for our approach, shedding light on structural properties of functional maps. Simone Melzi, Jing Ren 0004, Emanuele Rodolà, Abhishek Sharma 0013, Peter Wonka, Maks Ovsjanikov |
ACM Trans. Graph. | 5 |
| 2019 | StructureNet: hierarchical graph networks for 3D shape generationabstractThe ability to generate novel, diverse, and realistic 3D shapes along with associated part semantics and structure is central to many applications requiring high-quality 3D assets or large volumes of realistic training data. A key challenge towards this goal is how to accommodate diverse shape variations, including both continuous deformations of parts as well as structural or discrete alterations which add to, remove from, or modify the shape constituents and compositional structure. Such object structure can typically be organized into a hierarchy of constituent object parts and relationships, represented as a hierarchy of n -ary graphs. We introduce StructureNet, a hierarchical graph network which (i) can directly encode shapes represented as such n -ary graphs, (ii) can be robustly trained on large and complex shape families, and (iii) be used to generate a great diversity of realistic structured shape geometries. Technically, we accomplish this by drawing inspiration from recent advances in graph neural networks to propose an order-invariant encoding of n -ary graphs, considering jointly both part geometry and inter-part relations during network training. We extensively evaluate the quality of the learned latent spaces for various shape families and show significant advantages over baseline and competing methods. The learned latent spaces enable several structure-aware geometry processing applications, including shape generation and interpolation, shape editing, or shape structure discovery directly from un-annotated images, point clouds, or partial scans. Kaichun Mo, Paul Guerrero 0001, Li Yi 0001, Hao Su 0001, Peter Wonka, Niloy J. Mitra, Leonidas J. Guibas |
ACM Trans. Graph. | 5 |
| 2019 | Checkerboard patterns with black rectanglesabstractCheckerboard patterns with black rectangles can be derived from quad meshes with orthogonal diagonals. First, we present an initial theoretical analysis of these quad meshes. The analysis reveals many possible applications in geometry processing and also motivates the numerical optimization for aesthetic and functional checkerboard pattern design. Second, we describe an optimization algorithm that transforms initial 2D and 3D quad meshes into quad meshes with orthogonal diagonals. Third, we present a 2D checkerboard pattern design framework based on integer programming inspired by the logo design of the 2020 Olympic games. Our results show a variety of 2D and 3D checkerboard patterns that can be derived from 2D or 3D quad meshes with orthogonal diagonals. Chihan Peng, Caigui Jiang, Peter Wonka, Helmut Pottmann |
ACM Trans. Graph. | 3 |
| 2019 | Warp-and-project tomography for rapidly deforming objectsabstractComputed tomography has emerged as the method of choice for scanning complex shapes as well as interior structures of stationary objects. Recent progress has also allowed the use of CT for analyzing deforming objects and dynamic phenomena, although the deformations have been constrained to be either slow or periodic motions. In this work we improve the tomographic reconstruction of time-varying geometries undergoing faster, non-periodic deformations. Our method uses a warp-and-project approach that allows us to introduce an essentially continuous time axis where consistency of the reconstructed shape with the projection images is enforced for the specific time and deformation state at which the image was captured. The method uses an efficient, time-adaptive solver that yields both the moving geometry as well as the deformation field. We validate our method with extensive experiments using both synthetic and real data from a range of different application scenarios. Guangming Zang, Ramzi Idoughi, Ran Tao 0008, Gilles Lubineau, Peter Wonka, Wolfgang Heidrich |
ACM Trans. Graph. | 5 |
| 2019 | Isotropic Surface Remeshing without Large and Small AnglesabstractWe introduce a novel algorithm for isotropic surface remeshing which progressively eliminates obtuse triangles and improves small angles. The main novelty of the proposed approach is a simple vertex insertion scheme that facilitates the removal of large angles, and a vertex removal operation that improves the distribution of small angles. In combination with other standard local mesh operators, e.g., connectivity optimization and local tangential smoothing, our algorithm is able to remesh efficiently a low-quality mesh surface. Our approach can be applied directly or used as a post-processing step following other remeshing approaches. Our method has a similar computational efficiency to the fastest approach available, i.e., real-time adaptive remeshing [1]. In comparison with state-of-the-art approaches, our method consistently generates better results based on evaluations using different metrics. Yiqun Wang 0001, Dong-Ming Yan 0001, Chengcheng Tang, Jianwei Guo 0003, Xiaopeng Zhang 0001, Peter Wonka |
IEEE Trans. Vis. Comput. Graph. | 7 |
| 2018 | Super-Resolution and Sparse View CT Reconstruction
Guangming Zang, Mohamed Aly 0001, Ramzi Idoughi, Peter Wonka, Wolfgang Heidrich |
ECCV (16) | 4 |
| 2018 | String Art: Towards Computational Fabrication of String ImagesabstractAbstract In this paper we propose a novel method for the automatic computation and digital fabrication of artistic string images. String art is a technique used by artists for the creation of abstracted images which are composed of straight lines of strings tensioned between pins distributed on a frame. Together the strings fuse to a perceptible image. Traditionally, artists craft such images manually in a highly sophisticated and tedious design process. To achieve this goal fully automatically we propose a computational setup driven by a discrete optimization algorithm which takes an ordinary picture as input and converts it into a connected graph of strings that tries to reassemble the input image best possibly. Furthermore, we propose a hardware setup for automatic digital fabrication of these images using an industrial robot that spans the strings. Finally, we demonstrate the applicability of our approach by generating and fabricating a set of real string art images. Michael Birsak, Florian Rist 0001, Peter Wonka, Przemyslaw Musialski |
Comput. Graph. Forum | 3 |
| 2018 | MIQP-based Layout Design for Building InteriorsabstractAbstract We propose a hierarchical framework for the generation of building interiors. Our solution is based on a mixed integer quadratic programming (MIQP) formulation. We parametrize a layout by polygons that are further decomposed into small rectangles. We identify important high‐level constraints, such as room size, room position, room adjacency, and the outline of the building, and formulate them in a way that is compatible with MIQP and the problem parametrization. We also propose a hierarchical framework to improve the scalability of the approach. We demonstrate that our algorithm can be used for residential building layouts and can be scaled up to large layouts such as office buildings, shopping malls, and supermarkets. We show that our method is faster by multiple orders of magnitude than previous methods. Wenming Wu 0001, Lubin Fan, Ligang Liu 0001, Peter Wonka |
Comput. Graph. Forum | 4 |
| 2018 | FrankenGAN: guided detail synthesis for building mass models using style-synchonized GANsabstractCoarse building mass models are now routinely generated at scales ranging from individual buildings to whole cities. Such models can be abstracted from raw measurements, generated procedurally, or created manually. However, these models typically lack any meaningful geometric or texture details, making them unsuitable for direct display. We introduce the problem of automatically and realistically decorating such models by adding semantically consistent geometric details and textures. Building on the recent success of generative adversarial networks (GANs), we propose F ranken GAN, a cascade of GANs that creates plausible details across multiple scales over large neighborhoods. The various GANs are synchronized to produce consistent style distributions over buildings and neighborhoods. We provide the user with direct control over the variability of the output. We allow him/her to interactively specify the style via images and manipulate style-adapted sliders to control style variability. We test our system on several large-scale examples. The generated outputs are qualitatively evaluated via a set of perceptual studies and are found to be realistic, semantically plausible, and consistent in style. Paul Guerrero 0001, Anthony Steed, Peter Wonka, Niloy J. Mitra |
ACM Trans. Graph. | 4 |
| 2018 | Designing patterns using triangle-quad hybrid meshesabstractWe present a framework to generate mesh patterns that consist of a hybrid of both triangles and quads. Given a 3D surface, the generated patterns fit the surface boundaries and curvatures. Such regular and near regular triangle-quad hybrid meshes provide two key advantages: first, novel-looking polygonal patterns achieved by mixing different arrangements of triangles and quads together; second, a finer discretization of angle deficits than utilizing triangles or quads alone. Users have controls over the generated patterns in global and local levels. We demonstrate applications of our approach in architectural geometry and pattern design on surfaces. Chihan Peng, Helmut Pottmann, Peter Wonka |
ACM Trans. Graph. | 3 |
| 2018 | Continuous and orientation-preserving correspondences via functional mapsabstractWe propose a method for efficiently computing orientation-preserving and approximately continuous correspondences between non-rigid shapes, using the functional maps framework. We first show how orientation preservation can be formulated directly in the functional (spectral) domain without using landmark or region correspondences and without relying on external symmetry information. This allows us to obtain functional maps that promote orientation preservation, even when using descriptors, that are invariant to orientation changes. We then show how higher quality, approximately continuous and bijective pointwise correspondences can be obtained from initial functional maps by introducing a novel refinement technique that aims to simultaneously improve the maps both in the spectral and spatial domains. This leads to a general pipeline for computing correspondences between shapes that results in high-quality maps, while admitting an efficient optimization scheme. We show through extensive evaluation that our approach improves upon state-of-the-art results on challenging isometric and non-isometric correspondence benchmarks according to both measures of continuity and coverage as well as producing semantically meaningful correspondences as measured by the distance to ground truth maps. Jing Ren 0004, Adrien Poulenard, Peter Wonka, Maks Ovsjanikov |
ACM Trans. Graph. | 3 |
| 2018 | Space-time tomography for continuously deforming objectsabstractX-ray computed tomography (CT) is a valuable tool for analyzing objects with interesting internal structure or complex geometries that are not accessible with optical means. Unfortunately, tomographic reconstruction of complex shapes requires a multitude (often hundreds or thousands) of projections from different viewpoints. Such a large number of projections can only be acquired in a time-sequential fashion. This significantly limits the ability to use x-ray tomography for either objects that undergo uncontrolled shape change at the time scale of a scan, or else for analyzing dynamic phenomena, where the motion itself is under investigation. In this work, we present a non-parametric space-time tomographic method for tackling such dynamic settings. Through a combination of a new CT image acquisition strategy, a space-time tomographic image formation model, and an alternating, multi-scale solver, we achieve a general approach that can be used to analyze a wide range of dynamic phenomena. We demonstrate our method with extensive experiments on both real and simulated data. Guangming Zang, Ramzi Idoughi, Ran Tao 0008, Gilles Lubineau, Peter Wonka, Wolfgang Heidrich |
ACM Trans. Graph. | 5 |
| 2018 | Gaussian material synthesisabstractWe present a learning-based system for rapid mass-scale material synthesis that is useful for novice and expert users alike. The user preferences are learned via Gaussian Process Regression and can be easily sampled for new recommendations. Typically, each recommendation takes 40-60 seconds to render with global illumination, which makes this process impracticable for real-world workflows. Our neural network eliminates this bottleneck by providing high-quality image predictions in real time, after which it is possible to pick the desired materials from a gallery and assign them to a scene in an intuitive manner. Workflow timings against Disney's "principled" shader reveal that our system scales well with the number of sought materials, thus empowering even novice users to generate hundreds of high-quality material models without any expertise in material modeling. Similarly, expert users experience a significant decrease in the total modeling time when populating a scene with materials. Furthermore, our proposed solution also offers controllable recommendations and a novel latent space variant generation step to enable the real-time fine-tuning of materials without requiring any domain expertise. Károly Zsolnai-Fehér, Peter Wonka, Michael Wimmer 0001 |
ACM Trans. Graph. | 2 |
| 2018 | Dynamic Path Exploration on Mobile DevicesabstractWe present a novel framework for visualizing routes on mobile devices. Our framework is suitable for helping users explore their environment. First, given a starting point and a maximum route length, the system retrieves nearby points of interest (POIs). Second, we automatically compute an attractive walking path through the environment trying to pass by as many highly ranked POIs as possible. Third, we automatically compute a route visualization that shows the current user position, POI locations via pins, and detail lenses for more information about the POIs. The visualization is an animation of an orthographic map view that follows the current user position. We propose an optimization based on a binary integer program (BIP) that models multiple requirements for an effective placement of detail lenses. We show that our path computation method outperforms recently proposed methods and we evaluate the overall impact of our framework in two user studies. Michael Birsak, Przemyslaw Musialski, Peter Wonka, Michael Wimmer 0001 |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2018 | How Do Users Map Points Between Dissimilar Shapes?abstractFinding similar points in globally or locally similar shapes has been studied extensively through the use of various point descriptors or shape-matching methods. However, little work exists on finding similar points in dissimilar shapes. In this paper, we present the results of a study where users were given two dissimilar two-dimensional shapes and asked to map a given point in the first shape to the point in the second shape they consider most similar. We find that user mappings in this study correlate strongly with simple geometric relationships between points and shapes. To predict the probability distribution of user mappings between any pair of simple two-dimensional shapes, two distinct statistical models are defined using these relationships. We perform a thorough validation of the accuracy of these predictions and compare our models qualitatively and quantitatively to well-known shape-matching methods. Using our predictive models, we propose an approach to map objects or procedural content between different shapes in different design scenarios. Michael Hecher, Paul Guerrero 0001, Peter Wonka, Michael Wimmer 0001 |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2018 | Joint Graph Layouts for Visualizing Collections of Segmented MeshesabstractWe present a novel and efficient approach for computing joint graph layouts and then use it to visualize collections of segmented meshes. Our joint graph layout algorithm takes as input the adjacency matrices for a set of graphs along with partial, possibly soft, correspondences between nodes of different graphs. We then use a two stage procedure, where in the first step, we extend spectral graph drawing to include a consistency term so that a collection of graphs can be handled jointly. Our second step extends metric multi-dimensional scaling with stress majorization to the joint layout setting, while using the output of the spectral approach as initialization. Further, we discuss a user interface for exploring a collection of graphs. Finally, we show multiple example visualizations of graphs stemming from collections of segmented meshes and we present qualitative and quantitative comparisons with previous work. Jing Ren 0004, Jens Schneider 0002, Maks Ovsjanikov, Peter Wonka |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2017 | PolyFit: Polygonal Surface Reconstruction from Point CloudsabstractWe propose a novel framework for reconstructing lightweight polygonal surfaces from point clouds. Unlike traditional methods that focus on either extracting good geometric primitives or obtaining proper arrangements of primitives, the emphasis of this work lies in intersecting the primitives (planes only) and seeking for an appropriate combination of them to obtain a manifold polygonal surface model without boundary. We show that reconstruction from point clouds can be cast as a binary labeling problem. Our method is based on a hypothesizing and selection strategy. We first generate a reasonably large set of face candidates by intersecting the extracted planar primitives. Then an optimal subset of the candidate faces is selected through optimization. Our optimization is based on a binary linear programming formulation under hard constraints that enforce the final polygonal surface model to be manifold and watertight. Experiments on point clouds from various sources demonstrate that our method can generate lightweight polygonal surface models of arbitrary piecewise planar objects. Besides, our method is capable of recovering sharp features and is robust to noise, outliers, and missing data. Liangliang Nan, Peter Wonka |
ICCV | 2 |
| 2017 | Design Transformations for Rule-based Procedural ModelingabstractWe introduce design transformations for rule-based procedural models, e.g., for buildings and plants. Given two or more procedural designs, each specified by a grammar, a design transformation combines elements of the existing designs to generate new designs. We introduce two technical components to enable design transformations. First, we extend the concept of discrete rule switching to rule merging, leading to a very large shape space for combining procedural models. Second, we propose an algorithm to jointly derive two or more grammars, called grammar co-derivation. We demonstrate two applications of our work: we show that our framework leads to a larger variety of models than previous work, and we show fine-grained transformation sequences between two procedural models. Stefan Lienhard, Cheryl Lau, Pascal Müller, Peter Wonka, Mark Pauly |
Comput. Graph. Forum | 4 |
| 2017 | Design and volume optimization of space structuresabstractWe study the design and optimization of statically sound and materially efficient space structures constructed by connected beams. We propose a systematic computational framework for the design of space structures that incorporates static soundness, approximation of reference surfaces, boundary alignment, and geometric regularity. To tackle this challenging problem, we first jointly optimize node positions and connectivity through a nonlinear continuous optimization algorithm. Next, with fixed nodes and connectivity, we formulate the assignment of beam cross sections as a mixed-integer programming problem with a bilinear objective function and quadratic constraints. We solve this problem with a novel and practical alternating direction method based on linear programming relaxation. The capability and efficiency of the algorithms and the computational framework are validated by a variety of examples and comparisons. Caigui Jiang, Chengcheng Tang, Hans-Peter Seidel, Peter Wonka |
ACM Trans. Graph. | 4 |
| 2017 | BigSUR: large-scale structured urban reconstructionabstractThe creation of high-quality semantically parsed 3D models for dense metropolitan areas is a fundamental urban modeling problem. Although recent advances in acquisition techniques and processing algorithms have resulted in large-scale imagery or 3D polygonal reconstructions, such data-sources are typically noisy, and incomplete, with no semantic structure. In this paper, we present an automatic data fusion technique that produces high-quality structured models of city blocks. From coarse polygonal meshes, street-level imagery, and GIS footprints, we formulate a binary integer program that globally balances sources of error to produce semantically parsed mass models with associated facade elements. We demonstrate our system on four city regions of varying complexity; our examples typically contain densely built urban blocks spanning hundreds of buildings. In our largest example, we produce a structured model of 37 city blocks spanning a total of 1, 011 buildings at a scale and quality previously impossible to achieve automatically. John Femiani 0001, Peter Wonka, Niloy J. Mitra |
ACM Trans. Graph. | 3 |
| 2016 | Large Scale Asset Extraction for Urban Images
Lama Affara, Liangliang Nan, Bernard Ghanem, Peter Wonka |
ECCV (3) | 4 |
| 2016 | Manhattan-World Urban Reconstruction from Point Clouds
Minglei Li 0003, Peter Wonka, Liangliang Nan |
ECCV (4) | 2 |
| 2016 | Tetrahedral meshing via maximal Poisson-disk sampling
Jianwei Guo 0003, Dong-Ming Yan 0001, Li Chen 0031, Xiaopeng Zhang 0001, Oliver Deussen, Peter Wonka |
Comput. Aided Geom. Des. | 6 |
| 2016 | Reconstructing building mass models from UAV images
Minglei Li 0003, Liangliang Nan, Neil Smith, Peter Wonka |
Comput. Graph. | 4 |
| 2016 | Capacity constrained blue-noise sampling on surfaces
Sen Zhang 0005, Jianwei Guo 0003, Hui Zhang 0013, Xiaohong Jia 0001, Dong-Ming Yan 0001, Jun-Hai Yong, Peter Wonka |
Comput. Graph. | 7 |
| 2016 | A Probabilistic Model for Exteriors of Residential BuildingsabstractWe propose a new framework to model the exterior of residential buildings. The main goal of our work is to design a model that can be learned from data that is observable from the outside of a building and that can be trained with widely available data such as aerial images and street-view images. First, we propose a parametric model to describe the exterior of a building (with a varying number of parameters) and propose a set of attributes as a building representation with fixed dimensionality. Second, we propose a hierarchical graphical model with hidden variables to encode the relationships between building attributes and learn both the structure and parameters of the model from the database. Third, we propose optimization algorithms to generate three-dimensional models based on building attributes sampled from the graphical model. Finally, we demonstrate our framework by synthesizing new building models and completing partially observed building models from photographs. Lubin Fan, Peter Wonka |
ACM Trans. Graph. | 2 |
| 2016 | RAID: a relation-augmented image descriptorabstractAs humans, we regularly interpret scenes based on how objects arerelated, rather than based on the objects themselves. For example, we see a personridingan object X or a plankbridgingtwo objects. Current methods provide limited support to search for content based on such relations. We presentraid, a relation-augmented image descriptor that supports queries based on inter-region relations. The key idea of our descriptor is to encode region-to-region relations as the spatial distribution of point-to-region relationships between two image regions.raidallows sketch-based retrieval and requires minimal training data, thus making it suited even for querying uncommon relations. We evaluate the proposed descriptor by querying into large image databases and successfully extract non-trivial images demonstrating complex inter-region relations, which are easily missed or erroneously classified by existing methods. We assess the robustness ofraidon multiple datasets even when the region segmentation is computed automatically or very noisy. Paul Guerrero 0001, Niloy J. Mitra, Peter Wonka |
ACM Trans. Graph. | 3 |
| 2016 | Block assembly for global registration of building scansabstractWe propose a framework for global registration of building scans. The first contribution of our work is to detect and use portals (e.g., doors and windows) to improve the local registration between two scans. Our second contribution is an optimization based on a linear integer programming formulation. We abstract each scan as a block and model the blocks registration as an optimization problem that aims at maximizing the overall matching score of the entire scene. We propose an efficient solution to this optimization problem by iteratively detecting and adding local constraints. We demonstrate the effectiveness of the proposed method on buildings of various styles and that our approach is superior to the current state of the art. Feilong Yan, Liangliang Nan, Peter Wonka |
ACM Trans. Graph. | 3 |
| 2016 | Computational network design from functional specificationsabstractConnectivity and layout of underlying networks largely determine agent behavior and usage in many environments. For example, transportation networks determine the flow of traffic in a neighborhood, whereas building floorplans determine the flow of people in a workspace. Designing such networks from scratch is challenging as even local network changes can have large global effects. We investigate how to computationally create networks starting from only high-level functional specifications. Such specifications can be in the form of network density, travel time versus network length, traffic type, destination location, etc. We propose an integer programming-based approach that guarantees that the resultant networks are valid by fulfilling all the specified hard constraints and that they score favorably in terms of the objective function. We evaluate our algorithm in two different design settings, street layout and floorplans to demonstrate that diverse networks can emerge purely from high-level functional specifications. Chihan Peng, Fan Bao, Dong-Ming Yan 0001, Peter Wonka, Niloy J. Mitra |
ACM Trans. Graph. | 6 |
| 2016 | Automatic Constraint Detection for 2D Layout RegularizationabstractIn this paper, we address the problem of constraint detection for layout regularization. The layout we consider is a set of two-dimensional elements where each element is represented by its bounding box. Layout regularization is important in digitizing plans or images, such as floor plans and facade images, and in the improvement of user-created contents, such as architectural drawings and slide layouts. To regularize a layout, we aim to improve the input by detecting and subsequently enforcing alignment, size, and distance constraints between layout elements. Similar to previous work, we formulate layout regularization as a quadratic programming problem. In addition, we propose a novel optimization algorithm that automatically detects constraints. We evaluate the proposed framework using a variety of input layouts from different applications. Our results demonstrate that our method has superior performance to the state of the art. Haiyong Jiang, Liangliang Nan, Dong-Ming Yan 0001, Weiming Dong, Xiaopeng Zhang 0001, Peter Wonka |
IEEE Trans. Vis. Comput. Graph. | 6 |
| 2016 | Non-Obtuse Remeshing with Centroidal Voronoi TessellationabstractWe present a novel remeshing algorithm that avoids triangles with small (acute) angles and those with large (obtuse) angles. Our solution is based on an extension of Centroidal Voronoi Tesselation (CVT). We augment the original CVT formulation with a penalty term that penalizes short Voronoi edges, while the CVT term helps to avoid small angles. Our results show significant improvements in remeshing quality over the state of the art. Dong-Ming Yan 0001, Peter Wonka |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2015 | Intrinsic Scene Decomposition from RGB-D ImagesabstractIn this paper, we address the problem of computing an intrinsic decomposition of the colors of a surface into an albedo and a shading term. The surface is reconstructed from a single or multiple RGB-D images of a static scene obtained from different views. We thereby extend and improve existing works in the area of intrinsic image decomposition. In a variational framework, we formulate the problem as a minimization of an energy composed of two terms: a data term and a regularity term. The first term is related to the image formation process and expresses the relation between the albedo, the surface normals, and the incident illumination. We use an affine shading model, a combination of a Lambertian model, and an ambient lighting term. This model is relevant for Lambertian surfaces. When available, multiple views can be used to handle view-dependent non-Lambertian reflections. The second term contains an efficient combination of l2and l1-regularizers on the illumination vector field and albedo respectively. Unlike most previous approaches, especially Retinex-like techniques, these terms do not depend on the image gradient or texture, thus reducing the mixing shading/reflectance artifacts and leading to better results. The obtained non-linear optimization problem is efficiently solved using a cyclic block coordinate descent algorithm. Our method outperforms a range of state-of-the-art algorithms on a popular benchmark dataset. Mohammed Hachama, Bernard Ghanem, Peter Wonka |
ICCV | 3 |
| 2015 | Structural Graphical Lasso for Learning Mouse Brain ConnectivityabstractInvestigations into brain connectivity aim to recover networks of brain regions connected by anatomical tracts or by functional associations. The inference of brain networks has recently attracted much interest due to the increasing availability of high-resolution brain imaging data. Sparse inverse covariance estimation with lasso and group lasso penalty has been demonstrated to be a powerful approach to discover brain networks. Motivated by the hierarchical structure of the brain networks, we consider the problem of estimating a graphical model with tree-structural regularization in this paper. The regularization encourages the graphical model to exhibit a brain-like structure. Specifically, in this hierarchical structure, hundreds of thousands of voxels serve as the leaf nodes of the tree. A node in the intermediate layer represents a region formed by voxels in the subtree rooted at that node. The whole brain is considered as the root of the tree. We propose to apply the tree-structural regularized graphical model to estimate the mouse brain network. However, the dimensionality of whole-brain data, usually on the order of hundreds of thousands, poses significant computational challenges. Efficient algorithms that are capable of estimating networks from high-dimensional data are highly desired. To address the computational challenge, we develop a screening rule which can quickly identify many zero blocks in the estimated graphical model, thereby dramatically reducing the computational cost of solving the proposed model. It is based on a novel insight on the relationship between screening and the so-called proximal operator that we first establish in this paper. We perform experiments on both synthetic data and real data from the Allen Developing Mouse Brain Atlas; results demonstrate the effectiveness and efficiency of the proposed approach. Sen Yang 0004, Qian Sun 0002, Shuiwang Ji, Peter Wonka, Ian Davidson, Jieping Ye |
KDD | 4 |
| 2015 | Patch layout generation by detecting feature networks
Yuanhao Cao, Dong-Ming Yan 0001, Peter Wonka |
Comput. Graph. | 3 |
| 2015 | Designing Camera Networks by Convex Quadratic ProgrammingabstractAbstract In this paper, we study the problem of automatic camera placement for computer graphics and computer vision applications. We extend the problem formulations of previous work by proposing a novel way to incorporate visibility constraints and camera‐to‐camera relationships. For example, the placement solution can be encouraged to have cameras that image the same important locations from different viewing directions, which can enable reconstruction and surveillance tasks to perform better. We show that the general camera placement problem can be formulated mathematically as a convex binary quadratic program (BQP) under linear constraints. Moreover, we propose an optimization strategy with a favorable trade‐off between speed and solution quality. Our solution is almost as fast as a greedy treatment of the problem, but the quality is significantly higher, so much so that it is comparable to exact solutions that take orders of magnitude more computation time. Because it is computationally attractive, our method also allows users to explore the space of solutions for variations in input parameters. To evaluate its effectiveness, we show a range of 3D results on real‐world floorplans (garage, hotel, mall, and airport). Bernard Ghanem, Yuanhao Cao, Peter Wonka |
Comput. Graph. Forum | 3 |
| 2015 | Interactive Dimensioning of Parametric ModelsabstractAbstract We propose a solution for the dimensioning of parametric and procedural models. Dimensioning has long been a staple of technical drawings, and we present the first solution for interactive dimensioning: a dimension line positioning system that adapts to the view direction, given behavioral properties. After proposing a set of design principles for interactive dimensioning, we describe our solution consisting of the following major components. First, we describe how an author can specify the desired interactive behavior of a dimension line. Second, we propose a novel algorithm to place dimension lines at interactive speeds. Third, we introduce multiple extensions, including chained dimension lines, controls for different parameter types (e.g. discrete choices, angles), and the use of dimension lines for interactive editing. Our results show the use of dimension lines in an interactive parametric modeling environment for architectural, botanical, and mechanical models. Peter Wonka, Pascal Müller |
Comput. Graph. Forum | 2 |
| 2015 | Template Assembly for Detailed Urban ReconstructionabstractAbstract We propose a new framework to reconstruct building details by automatically assembling 3D templates on coarse textured building models. In a preprocessing step, we generate an initial coarse model to approximate a point cloud computed using Structure from Motion and Multi View Stereo, and we model a set of 3D templates of facade details. Next, we optimize the initial coarse model to enforce consistency between geometry and appearance (texture images). Then, building details are reconstructed by assembling templates on the textured faces of the coarse model. The 3D templates are automatically chosen and located by our optimization‐based template assembly algorithm that balances image matching and structural regularity. In the results, we demonstrate how our framework can enrich the details of coarse models using various data sets. Liangliang Nan, Caigui Jiang, Bernard Ghanem, Peter Wonka |
Comput. Graph. Forum | 4 |
| 2015 | A Survey of Blue-Noise Sampling and Its Applications
Dong-Ming Yan 0001, Jianwei Guo 0003, Bin Wang 0021, Xiaopeng Zhang 0001, Peter Wonka |
J. Comput. Sci. Technol. | 5 |
| 2015 | Lasso screening rules via dual polytope projection
Jie Wang 0005, Peter Wonka, Jieping Ye |
J. Mach. Learn. Res. | 2 |
| 2015 | Robust Rooftop Extraction From Visible Band Images Using Higher Order CRFabstractIn this paper, we propose a robust framework for building extraction in visible band images. We first get an initial classification of the pixels based on an unsupervised presegmentation. Then, we develop a novel conditional random field (CRF) formulation to achieve accurate rooftops extraction, which incorporates pixel-level information and segment-level information for the identification of rooftops. Comparing with the commonly used CRF model, a higher order potential defined on segment is added in our model, by exploiting region consistency and shape feature at segment level. Our experiments show that the proposed higher order CRF model outperforms the state-of-the-art methods both at pixel and object levels on rooftops with complex structures and sizes in challenging environments. Er Li, John Femiani 0001, Shibiao Xu, Xiaopeng Zhang 0001, Peter Wonka |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2015 | Interactive design of probability density functions for shape grammarsabstractA shape grammar defines a procedural shape space containing a variety of models of the same class, e.g. buildings, trees, furniture, airplanes, bikes, etc. We present a framework that enables a user to interactively design a probability density function (pdf) over such a shape space and to sample models according to the designed pdf. First, we propose a user interface that enables a user to quickly provide preference scores for selected shapes and suggest sampling strategies to decide which models to present to the user to evaluate. Second, we propose a novel kernel function to encode the similarity between two procedural models. Third, we propose a framework to interpolate user preference scores by combining multiple techniques: function factorization, Gaussian process regression, autorelevance detection, and l 1 regularization. Fourth, we modify the original grammars to generate models with a pdf proportional to the user preference scores. Finally, we provide evaluations of our user interface and framework parameters and a comparison to other exploratory modeling techniques using modeling tasks in five example shape spaces: furniture, low-rise buildings, skyscrapers, airplanes, and vegetation. Minh Dang, Stefan Lienhard, Duygu Ceylan, Boris Neubert, Peter Wonka, Mark Pauly |
ACM Trans. Graph. | 5 |
| 2015 | Polyhedral patternsabstractWe study the design and optimization of polyhedral patterns, which are patterns of planar polygonal faces on freeform surfaces. Working with polyhedral patterns is desirable in architectural geometry and industrial design. However, the classical tiling patterns on the plane must take on various shapes in order to faithfully and feasibly approximate curved surfaces. We define and analyze the deformations these tiles must undertake to account for curvature, and discover the symmetries that remain invariant under such deformations. We propose a novel method to regularize polyhedral patterns while maintaining these symmetries into a plethora of aesthetic and feasible patterns. Caigui Jiang, Chengcheng Tang, Amir Vaxman, Peter Wonka, Helmut Pottmann |
ACM Trans. Graph. | 4 |
| 2015 | Learning shape placements by exampleabstractWe present a method to learn and propagate shape placements in 2D polygonal scenes from a few examples provided by a user. The placement of a shape is modeled as an oriented bounding box. Simple geometric relationships between this bounding box and nearby scene polygons define a feature set for the placement. The feature sets of all example placements are then used to learn a probabilistic model over all possible placements and scenes. With this model, we can generate a new set of placements with similar geometric relationships in any given scene. We introduce extensions that enable propagation and generation of shapes in 3D scenes, as well as the application of a learned modeling session to large scenes without additional user interaction. These concepts allow us to generate complex scenes with thousands of objects with relatively little user interaction. Paul Guerrero 0001, Stefan Jeschke, Michael Wimmer 0001, Peter Wonka |
ACM Trans. Graph. | 4 |
| 2014 | A Highly Scalable Parallel Algorithm for Isotropic Total Variation ModelsabstractTotal variation (TV) models are among the most popular and successful tools in signal processing. However, due to the complex nature of the TV term, it is challenging to efficiently compute a solution for large-scale problems. State-of-the-art algorithms that are based on the alternating direction method of multipliers (ADMM) often involve solving large-size linear systems. In this paper, we propose a highly scalable parallel algorithm for TV models that is based on a novel decomposition strategy of the problem domain. As a result, the TV models can be decoupled into a set of small and independent subproblems, which admit closed form solutions. This makes our approach particularly suitable for parallel implementation. Our algorithm is guaranteed to converge to its global minimum. With N variables and n_p processes, the time complexity is O(N/(εn_p)) to reach an epsilon-optimal solution. Extensive experiments demonstrate that our approach outperforms existing state-of-the-art algorithms, especially in dealing with high-resolution, mega-size images. Jie Wang 0005, Qingyang Li 0001, Sen Yang 0004, Wei Fan 0001, Peter Wonka, Jieping Ye |
ICML | 5 |
| 2014 | Scaling SVM and Least Absolute Deviations via Exact Data ReductionabstractThe support vector machine (SVM) is a widely used method for classification. Although many efforts have been devoted to develop efficient solvers, it remains challenging to apply SVM to large-scale problems. A nice property of SVM is that the non-support vectors have no effect on the resulting classifier. Motivated by this observation, we present fast and efficient screening rules to discard non-support vectors by analyzing the dual problem of SVM via variational inequalities (DVI). As a result, the number of data instances to be entered into the optimization can be substantially reduced. Some appealing features of our screening method are: (1) DVI is safe in the sense that the vectors discarded by DVI are guaranteed to be non-support vectors; (2) the data set needs to be scanned only once to run the screening, and its computational cost is negligible compared to that of solving the SVM problem; (3) DVI is independent of the solvers and can be integrated with any existing efficient solver. We also show that the DVI technique can be extended to detect non-support vectors in the least absolute deviations regression (LAD). To the best of our knowledge, there are currently no screening methods for LAD. We have evaluated DVI on both synthetic and real data sets. Experiments indicate that DVI significantly outperforms the existing state-of-the-art screening rules for SVM, and it is very effective in discarding non-support vectors for LAD. The speedup gained by DVI rules can be up to two orders of magnitude. Jie Wang 0005, Peter Wonka, Jieping Ye |
ICML | 2 |
| 2014 | A Safe Screening Rule for Sparse Logistic Regression
Jie Wang 0005, Jun Liu 0003, Peter Wonka, Jieping Ye |
NIPS | 4 |
| 2014 | What Makes London Work Like London?abstractAbstract Urban data ranging from images and laser scans to traffic flows are regularly analyzed and modeled leading to better scene understanding. Commonly used computational approaches focus on geometric descriptors, both for images and for laser scans. In contrast, in urban planning, a large body of work has qualitatively evaluated street networks to understand their effects on the functionality of cities, both for pedestrians and for cars. In this work, we analyze street networks, both their topology (i.e., connectivity) and their geometry (i.e., layout), in an attempt to understand which factors play dominant roles in determining the characteristic of cities. We propose a set of street network descriptors to capture the essence of city layouts and use them, in a supervised setting, to classify and categorize various cities across the world. We evaluate our method on a range of cities, of various styles, and demonstrate that while standard image‐level descriptors perform poorly, the proposed network‐level descriptors can distinguish between different cities reliably and with high accuracy. Sawsan AlHalawani, Peter Wonka, Niloy J. Mitra |
Comput. Graph. Forum | 3 |
| 2014 | Automatic generation of tourist brochuresabstractAbstract We present a novel framework for the automatic generation of tourist brochures that include routing instructions and additional information presented in the form of so‐called detail lenses. The first contribution of this paper is the automatic creation of layouts for the brochures. Our approach is based on the minimization of an energy function that combines multiple goals: positioning of the lenses as close as possible to the corresponding region shown in an overview map, keeping the number of lenses low, and an efficient numbering of the lenses. The second contribution is a route‐aware simplification of the graph of streets used for traveling between the points of interest (POIs). This is done by reducing the graph consisting of all shortest paths through the minimization of an energy function. The output is a subset of street segments that enable traveling between all the POIs without considerable detours, while at the same time guaranteeing a clutter‐free visualization. Michael Birsak, Przemyslaw Musialski, Peter Wonka, Michael Wimmer 0001 |
Comput. Graph. Forum | 3 |
| 2014 | Parallel generation of architecture on the GPUabstractAbstract In this paper, we present a novel approach for the parallel evaluation of procedural shape grammars on the graphics processing unit (GPU). Unlike previous approaches that are either limited in the kind of shapes they allow, the amount of parallelism they can take advantage of, or both, our method supports state of the art procedural modeling including stochasticity and context‐sensitivity. To increase parallelism, we explicitly express independence in the grammar, reduce inter‐rule dependencies required for context‐sensitive evaluation, and introduce intra‐rule parallelism. Our rule scheduling scheme avoids unnecessary back and forth between CPU and GPU and reduces round trips to slow global memory by dynamically grouping rules in on‐chip shared memory. Our GPU shape grammar implementation is multiple orders of magnitude faster than the standard in CPU‐based rule evaluation, while offering equal expressive power. In comparison to the state of the art in GPU shape grammar derivation, our approach is nearly 50 times faster, while adding support for geometric context‐sensitivity. Markus Steinberger, Michael Kenzel, Bernhard Kainz, Joerg H. Mueller, Peter Wonka, Dieter Schmalstieg |
Comput. Graph. Forum | 5 |
| 2014 | On-the-fly generation and rendering of infinite cities on the GPUabstractAbstract In this paper, we present a new approach for shape‐grammar‐based generation and rendering of huge cities in real‐time on the graphics processing unit (GPU). Traditional approaches rely on evaluating a shape grammar and storing the geometry produced as a preprocessing step. During rendering, the pregenerated data is then streamed to the GPU. By interweaving generation and rendering, we overcome the problems and limitations of streaming pregenerated data. Using our methods ofvisibility pruningand adaptive level of detail, we are able to dynamically generate only the geometry needed to render the current view in real‐time directly on the GPU. We also present a robust and efficient way to dynamically update a scene's derivation tree and geometry, enabling us to exploit frame‐to‐frame coherence. Our combined generation and rendering is significantly faster than all previous work. For detailed scenes, we are capable of generating geometry more rapidly than even just copying pregenerated data from main memory, enabling us to render cities with thousands of buildings at up to 100 frames per second, even with the camera moving at supersonic speed. Markus Steinberger, Michael Kenzel, Bernhard Kainz, Peter Wonka, Dieter Schmalstieg |
Comput. Graph. Forum | 4 |
| 2014 | Blue-Noise Remeshing with Farthest Point OptimizationabstractAbstract In this paper, we present a novel method for surface sampling and remeshing with good blue‐noise properties. Our approach is based on the farthest point optimization (FPO), a relaxation technique that generates high quality blue‐noise point sets in 2D. We propose two important generalizations of the original FPO framework: adaptive sampling and sampling on surfaces. A simple and efficient algorithm for accelerating the FPO framework is also proposed. Experimental results show that the generalized FPO generates point sets with excellent blue‐noise properties for adaptive and surface sampling. Furthermore, we demonstrate that our remeshing quality is superior to the current state‐of‐theߚart approaches. Dong-Ming Yan 0001, Jianwei Guo 0003, Xiaohong Jia 0001, Xiaopeng Zhang 0001, Peter Wonka |
Comput. Graph. Forum | 5 |
| 2014 | Edit propagation using geometric relationship functionsabstractWe propose a method for propagating edit operations in 2D vector graphics, based on geometric relationship functions. These functions quantify the geometric relationship of a point to a polygon, such as the distance to the boundary or the direction to the closest corner vertex. The level sets of the relationship functions describe points with the same relationship to a polygon. For a given query point, we first determine a set of relationships to local features, construct all level sets for these relationships, and accumulate them. The maxima of the resulting distribution are points with similar geometric relationships. We show extensions to handle mirror symmetries, and discuss the use of relationship functions as local coordinate systems. Our method can be applied, for example, to interactive floorplan editing, and it is especially useful for large layouts, where individual edits would be cumbersome. We demonstrate populating 2D layouts with tens to hundreds of objects by propagating relatively few edit operations. Paul Guerrero 0001, Stefan Jeschke, Michael Wimmer 0001, Peter Wonka |
ACM Trans. Graph. | 4 |
| 2014 | Exploring quadrangulationsabstractWe present a framework for exploring topologically unique quadrangulations of an input shape. First, the input shape is segmented into surface patches. Second, different topologies are enumerated and explored in each patch. This is realized by an efficient subdivision-based quadrangulation algorithm that can exhaustively enumerate all mesh topologies within a patch. To help users navigate the potentially huge collection of variations, we propose tools to preview and arrange the results. Furthermore, the requirement that all patches need to be jointly quadrangulatable is formulated as a linear integer program. Finally, we apply the framework to shape-space exploration, remeshing, and design to underline the importance of topology exploration. Chihan Peng, Michael Barton 0002, Caigui Jiang, Peter Wonka |
ACM Trans. Graph. | 4 |
| 2014 | Procedural Design of Exterior Lighting for Buildings with Complex ConstraintsabstractWe present a system for the lighting design of procedurally modeled buildings. The design is procedurally specified as part of the ordinary modeling workflow by defining goals for the illumination that should be attained and locations where luminaires may be installed to realize these goals. Additionally, constraints can be modeled that make the arrangement of the installed luminaires respect certain aesthetic and structural considerations. From this specification, the system automatically generates a lighting solution for any concrete model instance. The underlying, intricate joint optimization and constraint satisfaction problem is approached with a stochastic scheme that operates directly in the complex subspace where all constraints are observed. To navigate this subspace efficaciously, the actual lighting situation is taken into account. We demonstrate our system on multiple examples spanning a variety of architectural structures and lighting designs. Michael Schwarz 0003, Peter Wonka |
ACM Trans. Graph. | 2 |
| 2014 | PushPull++abstractPushPull tools are implemented in most commercial 3D modeling suites. Their purpose is to intuitively transform a face, edge, or vertex, and then to adapt the polygonal mesh locally. However, previous approaches have limitations: Some allow adjustments only when adjacent faces are orthogonal; others support slanted surfaces but never create new details. Moreover, self-intersections and edge-collapses during editing are either ignored or work only partially for solid geometry. To overcome these limitations, we introduce the PushPull++ tool for rapid polygonal modeling. In our solution, we contribute novel methods for adaptive face insertion, adjacent face updates, edge collapse handling, and an intuitive user interface that automatically proposes useful drag directions. We show that PushPull++ reduces the complexity of common modeling tasks by up to an order of magnitude when compared with existing tools. Markus Lipp, Peter Wonka, Pascal Müller |
ACM Trans. Graph. | 2 |
| 2014 | Structure completion for facade layoutsabstractWe present a method to complete missing structures in facade layouts. Starting from an abstraction of the partially observed layout as a set of shapes, we can propose one or multiple possible completed layouts. Structure completion with large missing parts is an ill-posed problem. Therefore, we combine two sources of information to derive our solution: the observed shapes and a database of complete layouts. The problem is also very difficult, because shape positions and attributes have to be estimated jointly. Our proposed solution is to break the problem into two components: a statistical model to evaluate layouts and a planning algorithm to generate candidate layouts. This ensures that the completed result is consistent with the observation and the layouts in the database. Lubin Fan, Przemyslaw Musialski, Ligang Liu 0001, Peter Wonka |
ACM Trans. Graph. | 4 |
| 2014 | Inverse procedural modeling of facade layoutsabstractIn this paper, we address the following research problem: How can we generate a meaningful split grammar that explains a given facade layout? To evaluate if a grammar is meaningful, we propose a cost function based on the description length and minimize this cost using an approximate dynamic programming framework. Our evaluation indicates that our framework extracts meaningful split grammars that are competitive with those of expert users, while some users and all competing automatic solutions are less successful. Fuzhang Wu, Dong-Ming Yan 0001, Weiming Dong, Xiaopeng Zhang 0001, Peter Wonka |
ACM Trans. Graph. | 5 |
| 2014 | Computing layouts with deformable templatesabstractIn this paper, we tackle the problem of tiling a domain with a set of deformable templates. A valid solution to this problem completely covers the domain with templates such that the templates do not overlap. We generalize existing specialized solutions and formulate a general layout problem by modeling important constraints and admissible template deformations. Our main idea is to break the layout algorithm into two steps: a discrete step to lay out the approximate template positions and a continuous step to refine the template shapes. Our approach is suitable for a large class of applications, including floorplans, urban layouts, and arts and design. Chihan Peng, Peter Wonka |
ACM Trans. Graph. | 3 |
| 2014 | Low-Resolution Remeshing Using the Localized Restricted Voronoi DiagramabstractA big problem in triangular remeshing is to generate meshes when the triangle size approaches the feature size in the mesh. The main obstacle for Centroidal Voronoi Tessellation (CVT)-based remeshing is to compute a suitable Voronoi diagram. In this paper, we introduce the localized restricted Voronoi diagram (LRVD) on mesh surfaces. The LRVD is an extension of the restricted Voronoi diagram (RVD), but it addresses the problem that the RVD can contain Voronoi regions that consist of multiple disjoint surface patches. Our definition ensures that each Voronoi cell in the LRVD is a single connected region. We show that the LRVD is a useful extension to improve several existing mesh-processing techniques, most importantly surface remeshing with a low number of vertices. While the LRVD and RVD are identical in most simple configurations, the LRVD is essential when sampling a mesh with a small number of points and for sampling surface areas that are in close proximity to other surface areas, e.g., nearby sheets. To compute the LRVD, we combine local discrete clustering with a global exact computation. Dong-Ming Yan 0001, Guanbo Bao, Xiaopeng Zhang 0001, Peter Wonka |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2014 | Unbiased Sampling and Meshing of IsosurfacesabstractIn this paper, we present a new technique to generate unbiased samples on isosurfaces. An isosurface, F(x; y; z) = c, of a function, F, is implicitly defined by trilinear interpolation of background grid points. The key idea of our approach is that of treating the isosurface within a grid cell as a graph (height) function in one of the three coordinate axis directions, restricted to where the slope is not too high, and integrating / sampling from each of these three. We use this unbiased sampling algorithm for applications in Monte Carlo integration, Poisson-disk sampling, and isosurface meshing. Dong-Ming Yan 0001, Johannes Wallner 0001, Peter Wonka |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2014 | Efficient triangulation of Poisson-disk sampled point sets
Jianwei Guo 0003, Dong-Ming Yan 0001, Guanbo Bao, Weiming Dong, Xiaopeng Zhang 0001, Peter Wonka |
Vis. Comput. | 6 |
| 2013 | An efficient ADMM algorithm for multidimensional anisotropic total variation regularization problemsabstractTotal variation (TV) regularization has important applications in signal processing including image denoising, image deblurring, and image reconstruction. A significant challenge in the practical use of TV regularization lies in the nondifferentiable convex optimization, which is difficult to solve especially for large-scale problems. In this paper, we propose an efficient alternating augmented Lagrangian method (ADMM) to solve total variation regularization problems. The proposed algorithm is applicable for tensors, thus it can solve multidimensional total variation regularization problems. One appealing feature of the proposed algorithm is that it does not need to solve a linear system of equations, which is often the most expensive part in previous ADMM-based methods. In addition, each step of the proposed algorithm involves a set of independent and smaller problems, which can be solved in parallel. Thus, the proposed algorithm scales to large size problems. Furthermore, the global convergence of the proposed algorithm is guaranteed, and the time complexity of the proposed algorithm is O(dN/ε) on a d-mode tensor with N entries for achieving an ε-optimal solution. Extensive experimental results demonstrate the superior performance of the proposed algorithm in comparison with current state-of-the-art methods. Sen Yang 0004, Jie Wang 0005, Wei Fan 0001, Peter Wonka, Jieping Ye |
KDD | 5 |
| 2013 | Lasso Screening Rules via Dual Polytope ProjectionabstractLasso is a widely used regression technique to find sparse representations. When the dimension of the feature space and the number of samples are extremely large, solving the Lasso problem remains challenging. To improve the efficiency of solving large-scale Lasso problems, El Ghaoui and his colleagues have proposed the SAFE rules which are able to quickly identify the inactive predictors, i.e., predictors that have $0$ components in the solution vector. Then, the inactive predictors or features can be removed from the optimization problem to reduce its scale. By transforming the standard Lasso to its dual form, it can be shown that the inactive predictors include the set of inactive constraints on the optimal dual solution. In this paper, we propose an efficient and effective screening rule via Dual Polytope Projections (DPP), which is mainly based on the uniqueness and nonexpansiveness of the optimal dual solution due to the fact that the feasible set in the dual space is a convex and closed polytope. Moreover, we show that our screening rule can be extended to identify inactive groups in group Lasso. To the best of our knowledge, there is currently no exact" screening rule for group Lasso. We have evaluated our screening rule using many real data sets. Results show that our rule is more effective to identify inactive predictors than existing state-of-the-art screening rules for Lasso." Jie Wang 0005, Peter Wonka, Jieping Ye |
NIPS | 3 |
| 2013 | Illustrating the disassembly of 3D models
Jianwei Guo 0003, Dong-Ming Yan 0001, Er Li, Weiming Dong, Peter Wonka, Xiaopeng Zhang 0001 |
Comput. Graph. | 5 |
| 2013 | A Survey of Urban ReconstructionabstractAbstract This paper provides a comprehensive overview of urban reconstruction. While there exists a considerable body of literature, this topic is still under active research. The work reviewed in this survey stems from the following three research communities: computer graphics, computer vision and photogrammetry and remote sensing. Our goal is to provide a survey that will help researchers to better position their own work in the context of existing solutions, and to help newcomers and practitioners in computer graphics to quickly gain an overview of this vast field. Further, we would like to bring the mentioned research communities to even more interdisciplinary work, since the reconstruction problem itself is by far not solved. Przemyslaw Musialski, Peter Wonka, Daniel G. Aliaga, Michael Wimmer 0001, Luc Van Gool, Werner Purgathofer |
Comput. Graph. Forum | 2 |
| 2013 | Connectivity Editing for Quad-Dominant MeshesabstractAbstract We propose a connectivity editing framework for quad‐dominant meshes. In our framework, the user can edit the mesh connectivity to control the location, type, and number of irregular vertices (with more or fewer than four neighbors) and irregular faces (non‐quads). We provide a theoretical analysis of the problem, discuss what edits are possible and impossible, and describe how to implement an editing framework that realizes all possible editing operations. In the results, we show example edits and illustrate the advantages and disadvantages of different strategies for quad‐dominant mesh design. Chihan Peng, Peter Wonka |
Comput. Graph. Forum | 2 |
| 2013 | Tensor Completion for Estimating Missing Values in Visual DataabstractIn this paper, we propose an algorithm to estimate missing values in tensors of visual data. The values can be missing due to problems in the acquisition process or because the user manually identified unwanted outliers. Our algorithm works even with a small amount of samples and it can propagate structure to fill larger missing regions. Our methodology is built on recent studies about matrix completion using the matrix trace norm. The contribution of our paper is to extend the matrix case to the tensor case by proposing the first definition of the trace norm for tensors and then by building a working algorithm. First, we propose a definition for the tensor trace norm that generalizes the established definition of the matrix trace norm. Second, similarly to matrix completion, the tensor completion is formulated as a convex optimization problem. Unfortunately, the straightforward problem extension is significantly harder to solve than the matrix case because of the dependency among multiple constraints. To tackle this problem, we developed three algorithms: simple low rank tensor completion (SiLRTC), fast low rank tensor completion (FaLRTC), and high accuracy low rank tensor completion (HaLRTC). The SiLRTC algorithm is simple to implement and employs a relaxation technique to separate the dependent relationships and uses the block coordinate descent (BCD) method to achieve a globally optimal solution; the FaLRTC algorithm utilizes a smoothing scheme to transform the original nonsmooth problem into a smooth one and can be used to solve a general tensor trace norm minimization problem; the HaLRTC algorithm applies the alternating direction method of multipliers (ADMMs) to our problem. Our experiments show potential applications of our algorithms and the quantitative evaluation indicates that our methods are more accurate and robust than heuristic approaches. The efficiency comparison indicates that FaLTRC and HaLRTC are more efficient than SiLRTC and between FaLRTC an- HaLRTC the former is more efficient to obtain a low accuracy solution and the latter is preferred if a high-accuracy solution is desired. Ji Liu 0002, Przemyslaw Musialski, Peter Wonka, Jieping Ye |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2013 | Procedural facade variations from a single layoutabstractWe introduce a framework to generate many variations of a facade design that look similar to a given facade layout. Starting from an input image, the facade is hierarchically segmented and labeled with a collection of manual and automatic tools. The user can then model constraints that should be maintained in any variation of the input facade design. Subsequently, facade variations are generated for different facade sizes, where multiple variations can be produced for a certain size. Computing such new facade variations has many unique challenges, and we propose a new algorithm based on interleaving heuristic search and quadratic programming. In contrast to most previous work, we focus on the generation of new design variations and not on the automatic analysis of the input's structure. Adding a modeling step with the user in the loop ensures that our results routinely are of high quality. Fan Bao, Michael Schwarz 0003, Peter Wonka |
ACM Trans. Graph. | 3 |
| 2013 | Generating and exploring good building layoutsabstractGood building layouts are required to conform to regulatory guidelines, while meeting certain quality measures. While different methods can sample the space of such good layouts, there exists little support for a user to understand and systematically explore the samples. Starting from a discrete set of good layouts, we analytically characterize the local shape space of good layouts around each initial layout, compactly encode these spaces, and link them to support transitions across the different local spaces. We represent such transitions in the form of a portal graph. The user can then use the portal graph, along with the family of local shape spaces, to globally and locally explore the space of good building layouts. We use our framework on a variety of different test scenarios to showcase an intuitive design, navigation, and exploration interface. Fan Bao, Dong-Ming Yan 0001, Niloy J. Mitra, Peter Wonka |
ACM Trans. Graph. | 4 |
| 2013 | Gap processing for adaptive maximal poisson-disk samplingabstractIn this article, we study the generation of maximal Poisson-disk sets with varying radii. First, we present a geometric analysis of gaps in such disk sets. This analysis is the basis for maximal and adaptive sampling in Euclidean space and on manifolds. Second, we propose efficient algorithms and data structures to detect gaps and update gaps when disks are inserted, deleted, moved, or when their radii are changed. We build on the concepts of regular triangulations and the power diagram. Third, we show how our analysis contributes to the state-of-the-art in surface remeshing. Dong-Ming Yan 0001, Peter Wonka |
ACM Trans. Graph. | 2 |
| 2013 | Urban pattern: layout design by hierarchical domain splittingabstractWe present a framework for generating street networks and parcel layouts. Our goal is the generation of high-quality layouts that can be used for urban planning and virtual environments. We propose a solution based on hierarchical domain splitting using two splitting types: streamline-based splitting, which splits a region along one or multiple streamlines of a cross field, and template-based splitting, which warps pre-designed templates to a region and uses the interior geometry of the template as the splitting lines. We combine these two splitting approaches into a hierarchical framework, providing automatic and interactive tools to explore the design space. Etienne Vouga, Peter Wonka |
ACM Trans. Graph. | 4 |
| 2013 | A framework for interactive image color editing
Przemyslaw Musialski, Ming Cui, Jieping Ye, Anshuman Razdan, Peter Wonka |
Vis. Comput. | 5 |
| 2012 | Feature grouping and selection over an undirected graphabstractHigh-dimensional regression/classification continues to be an important and challenging problem, especially when features are highly correlated. Feature selection, combined with additional structure information on the features has been considered to be promising in promoting regression/classification performance. Graph-guided fused lasso (GFlasso) has recently been proposed to facilitate feature selection and graph structure exploitation, when features exhibit certain graph structures. However, the formulation in GFlasso relies on pairwise sample correlations to perform feature grouping, which could introduce additional estimation bias. In this paper, we propose three new feature grouping and selection methods to resolve this issue. The first method employs a convex function to penalize the pairwise l∞ norm of connected regression/classification coefficients, achieving simultaneous feature grouping and selection. The second method improves the first one by utilizing a non-convex function to reduce the estimation bias. The third one is the extension of the second method using a truncated l1 regularization to further reduce the estimation bias. The proposed methods combine feature grouping and feature selection to enhance estimation accuracy. We employ the alternating direction method of multipliers (ADMM) and difference of convex functions (DC) programming to solve the proposed formulations. Our experimental results on synthetic data and two real datasets demonstrate the effectiveness of the proposed methods. Sen Yang 0004, Lei Yuan 0001, Ying-Cheng Lai, Xiaotong Shen, Peter Wonka, Jieping Ye |
KDD | 5 |
| 2012 | Interactive Coherence-Based Façade ModelingabstractAbstract We propose a novel interactive framework for modeling building façades from images. Our method is based on the notion of coherence‐based editing which allows exploiting partial symmetries across the façade at any level of detail. The proposed workflow mixes manual interaction with automatic splitting and grouping operations based on unsupervised cluster analysis. In contrast to previous work, our approach leads to detailed 3d geometric models with up to several thousand regions per façade. We compare our modeling scheme to others and evaluate our approach in a user study with an experienced user and several novice users. Przemyslaw Musialski, Michael Wimmer 0001, Peter Wonka |
Comput. Graph. Forum | 3 |
| 2012 | A Multi-Stage Framework for Dantzig Selector and LASSO
Ji Liu 0002, Peter Wonka, Jieping Ye |
J. Mach. Learn. Res. | 2 |
| 2012 | Sparse non-negative tensor factorization using columnwise coordinate descent
Ji Liu 0002, Jun Liu 0003, Peter Wonka, Jieping Ye |
Pattern Recognit. | 3 |
| 2011 | Estimating Color and Texture Parameters for Vector GraphicsabstractAbstract Diffusion curves are a powerful vector graphic representation that stores an image as a set of 2D Bezier curves with colors defined on either side. These colors are diffused over the image plane, resulting in smooth color regions as well as sharp boundaries. In this paper, we introduce a new automatic diffusion curve coloring algorithm. We start by defining a geometric heuristic for the maximum density of color control points along the image curves. Following this, we present a new algorithm to set the colors of these points so that the resulting diffused image is as close as possible to a source image in a least squares sense. We compare our coloring solution to the existing one which fails for textured regions, small features, and inaccurately placed curves. The second contribution of the paper is to extend the diffusion curve representation to include texture details based on Gabor noise. Like the curves themselves, the defined texture is resolution independent, and represented compactly. We define methods to automatically make an initial guess for the noise texure, and we provide intuitive manual controls to edit the parameters of the Gabor noise. Finally, we show that the diffusion curve representation itself extends to storing any number of attributes in an image, and we demonstrate this functionality with image stippling an hatching applications. Stefan Jeschke, David Cline, Peter Wonka |
Comput. Graph. Forum | 3 |
| 2011 | Interactive Modeling of City Layouts using Layers of Procedural ContentabstractAbstract In this paper, we present new solutions for the interactive modeling of city layouts that combine the power of procedural modeling with the flexibility of manual modeling. Procedural modeling enables us to quickly generate large city layouts, while manual modeling allows us to hand‐craft every aspect of a city. We introduce transformation and merging operators for both topology preserving and topology changing transformations based on graph cuts. In combination with a layering system, this allows intuitive manipulation of urban layouts using operations such as drag and drop, translation, rotation etc. In contrast to previous work, these operations always generate valid, i.e., intersection‐free layouts. Furthermore, we introduce anchored assignments to make sure that modifications are persistent even if the whole urban layout is regenerated. Markus Lipp, Daniel Scherzer, Peter Wonka, Michael Wimmer 0001 |
Comput. Graph. Forum | 3 |
| 2011 | A New QEM for Parametrization of Raster ImagesabstractAbstract We present an image processing method that converts a raster image to a simplical two‐complex which has only a small number of vertices (base mesh) plus a parametrization that maps each pixel in the original image to a combination of the barycentric coordinates of the triangle it is finally mapped into. Such a conversion of a raster image into a base mesh plus parametrization can be useful for many applications such as segmentation, image retargeting, multi‐resolution editing with arbitrary topologies, edge preserving smoothing, compression, etc. The goal of the algorithm is to produce a base mesh such that it has a small colour distortion as well as high shape fairness, and a parametrization that is globally continuous visually and numerically. Inspired by multi‐resolution adaptive parametrization of surfaces and quadric error metric, the algorithm converts pixels in the image to a dense triangle mesh and performs error‐bounded simplification jointly considering geometry and colour. The eliminated vertices are projected to an existing face. The implementation is iterative and stops when it reaches a prescribed error threshold. The algorithm is feature‐sensitive, i.e. salient feature edges in the images are preserved where possible and it takes colour into account thereby producing a better quality triangulation. Xuetao Yin, John Femiani 0001, Peter Wonka, Anshuman Razdan |
Comput. Graph. Forum | 3 |
| 2011 | Interactive architectural modeling with procedural extrusionsabstractWe present an interactive procedural modeling system for the exterior of architectural models. Our modeling system is based on procedural extrusions of building footprints. The main novelty of our work is that we can model difficult architectural surfaces in a procedural framework, for example, curved roofs, overhanging roofs, dormer windows, interior dormer windows, roof constructions with vertical walls, buttresses, chimneys, bay windows, columns, pilasters, and alcoves. We present a user interface to interactively specify procedural extrusions, a sweep plane algorithm to compute a two-manifold architectural surface, and applications to architectural modeling. Peter Wonka |
ACM Trans. Graph. | 2 |
| 2011 | Connectivity editing for quadrilateral meshesabstractWe propose new connectivity editing operations for quadrilateral meshes with the unique ability to explicitly control the location, orientation, type, and number of the irregular vertices (valence not equal to four) in the mesh while preserving sharp edges. We provide theoretical analysis on what editing operations are possible and impossible and introduce threefundamentaloperations to move and re-orient a pair of irregular vertices. We argue that our editing operations are fundamental, because they only change the quad mesh in the smallest possible region and involve the fewest irregular vertices (i.e., two). The irregular vertex movement operations are supplemented by operations for the splitting, merging, canceling, and aligning of irregular vertices. We explain how the proposed high-level operations are realized through graph-level editing operations such as quad collapses, edge flips, and edge splits. The utility of these mesh editing operations are demonstrated by improving the connectivity of quad meshes generated from state-of-art quadrangulation techniques. Chihan Peng, Eugene Zhang, Yoshihiro Kobayashi, Peter Wonka |
ACM Trans. Graph. | 4 |
| 2011 | Geometry Synthesis on Surfaces Using Field-Guided Shape GrammarsabstractWe show how to model geometric patterns on surfaces. We build on the concept of shape grammars to allow the grammars to be guided by a vector or tensor field. Our approach affords greater artistic freedom in design and enables the use of grammars to create patterns on manifold surfaces. We show several application examples in visualization, anisotropic tiling of mosaics, and geometry synthesis on surfaces. In contrast to previous work, we can create patterns that adapt to the underlying surface rather than distorting the geometry with a texture parameterization. Additionally, we are the first to model patterns with a global structure thanks to the ability to derive field-guided shape grammars on surfaces. Fan Bao, Eugene Zhang, Yoshihiro Kobayashi, Peter Wonka |
IEEE Trans. Vis. Comput. Graph. | 5 |
| 2010 | Multi-Stage Dantzig SelectorabstractWe consider the following sparse signal recovery (or feature selection) problem: given a design matrix $X\in \mathbb{R}^{n\times m}$ $(m\gg n)$ and a noisy observation vector $y\in \mathbb{R}^{n}$ satisfying $y=X\beta^*+\epsilon$ where $\epsilon$ is the noise vector following a Gaussian distribution $N(0,\sigma^2I)$, how to recover the signal (or parameter vector) $\beta^*$ when the signal is sparse? The Dantzig selector has been proposed for sparse signal recovery with strong theoretical guarantees. In this paper, we propose a multi-stage Dantzig selector method, which iteratively refines the target signal $\beta^*$. We show that if $X$ obeys a certain condition, then with a large probability the difference between the solution $\hat\beta$ estimated by the proposed method and the true solution $\beta^*$ measured in terms of the $l_p$ norm ($p\geq 1$) is bounded as \begin{equation*} \|\hat\beta-\beta^*\|_p\leq \left(C(s-N)^{1/p}\sqrt{\log m}+\Delta\right)\sigma, \end{equation*} $C$ is a constant, $s$ is the number of nonzero entries in $\beta^*$, $\Delta$ is independent of $m$ and is much smaller than the first term, and $N$ is the number of entries of $\beta^*$ larger than a certain value in the order of $\mathcal{O}(\sigma\sqrt{\log m})$. The proposed method improves the estimation bound of the standard Dantzig selector approximately from $Cs^{1/p}\sqrt{\log m}\sigma$ to $C(s-N)^{1/p}\sqrt{\log m}\sigma$ where the value $N$ depends on the number of large entries in $\beta^*$. When $N=s$, the proposed algorithm achieves the oracle solution with a high probability. In addition, with a large probability, the proposed method can select the same number of correct features under a milder condition than the Dantzig selector. Ji Liu 0002, Peter Wonka, Jieping Ye |
NIPS | 2 |
| 2010 | Parallel generation of multiple L-systems
Markus Lipp, Peter Wonka, Michael Wimmer 0001 |
Comput. Graph. | 2 |
| 2010 | Editorial
Michael Wimmer 0001, Peter Wonka |
Comput. Graph. | 2 |
| 2010 | Grammar-based Encoding of FacadesabstractAbstract In this paper we propose a real‐time rendering approach for procedural cities. Our first contribution is a new lightweight grammar representation that compactly encodes facade structures and allows fast per‐pixel access. We call this grammarF‐shade. Our second contribution is a prototype rendering system that renders an urban model from the compact representation directly on the GPU. Our suggested approach explores an interesting connection from procedural modeling to real‐time rendering. Evaluating procedural descriptions at render time uses less memory than the generation of intermediate geometry. This enables us to render large urban models directly from GPU memory. Simon Haegler, Peter Wonka, Stefan Müller Arisona, Luc Van Gool, Pascal Müller |
Comput. Graph. Forum | 2 |
| 2010 | Modelling the Appearance and Behaviour of Urban SpacesabstractAbstract Urban spaces consist of a complex collection of buildings, parcels, blocks and neighbourhoods interconnected by streets. Accurately modelling both the appearance and the behaviour of dense urban spaces is a significant challenge. The recent surge in urban data and its availability via the Internet has fomented a significant amount of research in computer graphics and in a number of applications in urban planning, emergency management and visualization. In this paper, we seek to provide an overview of methods spanning computer graphics and related fields involved in this goal. Our paper reports the most prominent methods in urban modelling and rendering, urban visualization and urban simulation models. A reader will be well versed in the key problems and current solution methods. Carlos A. Vanegas, Daniel G. Aliaga, Peter Wonka, Pascal Müller, Paul Waddell, Benjamin Watson 0001 |
Comput. Graph. Forum | 3 |
| 2010 | Editing operations for irregular vertices in triangle meshesabstractWe describe an interactive editing framework that provides control over the type, location, and number of irregular vertices in a triangle mesh. We first provide a theoretical analysis to identify the simplest possible operations for editing irregular vertices and then introduce a hierarchy of editing operations to control the type, location, and number of irregular vertices. We demonstrate the power of our editing framework with an example application in pattern design on surfaces. Eugene Zhang, Yoshihiro Kobayashi, Peter Wonka |
ACM Trans. Graph. | 4 |
| 2010 | Route Visualization Using Detail LensesabstractWe present a method designed to address some limitations of typical route map displays of driving directions. The main goal of our system is to generate a printable version of a route map that shows the overview and detail views of the route within a single, consistent visual frame. Our proposed visualization provides a more intuitive spatial context than a simple list of turns. We present a novel multifocus technique to achieve this goal, where the foci are defined by points of interest (POI) along the route. A detail lens that encapsulates the POI at a finer geospatial scale is created for each focus. The lenses are laid out on the map to avoid occlusion with the route and each other, and to optimally utilize the free space around the route. We define a set of layout metrics to evaluate the quality of a lens layout for a given route map visualization. We compare standard lens layout methods to our proposed method and demonstrate the effectiveness of our method in generating aesthetically pleasing layouts. Finally, we perform a user study to evaluate the effectiveness of our layout choices. Pushpak Karnick, David Cline, Stefan Jeschke, Anshuman Razdan, Peter Wonka |
IEEE Trans. Vis. Comput. Graph. | 5 |
| 2010 | Color-to-gray conversion using ISOMAP
Ming Cui, Jiuxiang Hu, Anshuman Razdan, Peter Wonka |
Vis. Comput. | 4 |
| 2009 | Tensor completion for estimating missing values in visual dataabstractIn this paper we propose an algorithm to estimate missing values in tensors of visual data. The values can be missing due to problems in the acquisition process, or because the user manually identified unwanted outliers. Our algorithm works even with a small amount of samples and it can propagate structure to fill larger missing regions. Our methodology is built on recent studies about matrix completion using the matrix trace norm. The contribution of our paper is to extend the matrix case to the tensor case by laying out the theoretical foundations and then by building a working algorithm. First, we propose a definition for the tensor trace norm, that generalizes the established definition of the matrix trace norm. Second, similar to matrix completion, the tensor completion is formulated as a convex optimization problem. Unfortunately, the straightforward problem extension is significantly harder to solve than the matrix case because of the dependency among multiple constraints. To tackle this problem, we employ a relaxation technique to separate the dependant relationships and use the block coordinate descent (BCD) method to achieve a globally optimal solution. Our experiments show potential applications of our algorithm and the quantitative evaluation indicates that our method is more accurate and robust than heuristic approaches. Ji Liu 0002, Przemyslaw Musialski, Peter Wonka, Jieping Ye |
ICCV | 3 |
| 2009 | GPU Rendering of Relief Mapped Conical FrustaabstractAbstract This paper proposes to use relief‐mapped conical frusta (cones cut by planes) to skin skeletal objects. Based on this representation, current programmable graphics hardware can perform the rendering with only minimal communication between the CPU and GPU. A consistent definition of conical frusta including texture parametrization and a continuous surface normal is provided. Rendering is performed by analytical ray casting of the relief‐mapped frusta directly on the GPU. We demonstrate both static and animated objects rendered using our technique and compare to polygonal renderings of similar quality. D. Bhagvat, Stefan Jeschke, David Cline, Peter Wonka |
Comput. Graph. Forum | 4 |
| 2009 | Dart Throwing on SurfacesabstractAbstract In this paper we present dart throwing algorithms to generate maximal Poisson disk point sets directly on 3D surfaces. We optimize dart throwing by efficiently excluding areas of the domain that are already covered by existing darts. In the case of triangle meshes, our algorithm shows dramatic speed improvement over comparable sampling methods. The simplicity of our basic algorithm naturally extends to the sampling of other surface types, including spheres, NURBS, subdivision surfaces, and implicits. We further extend the method to handle variable density points, and the placement of arbitrary ellipsoids without overlap. Finally, we demonstrate how to adapt our algorithm to work with geodesic instead of Euclidean distance. Applications for our method include fur modeling, the placement of mosaic tiles and polygon remeshing. David Cline, Stefan Jeschke, K. White, Anshuman Razdan, Peter Wonka |
Comput. Graph. Forum | 5 |
| 2009 | A Comparison of Tabular PDF Inversion MethodsabstractAbstract The most common form of tabular inversion used in computer graphics is to compute the cumulative distribution table of a probability distribution (PDF) and then search within it to transform points, using an O(log n) binary search. Besides the standard inversion method, however, several other discrete inversion algorithms exist that can perform the same transformation inO(1) time per point. In this paper, we examine the performance of three of these alternate methods, two of which are new. David Cline, Anshuman Razdan, Peter Wonka |
Comput. Graph. Forum | 3 |
| 2009 | A Shape Grammar for Developing Glyph-based VisualizationsabstractAbstract In this paper we address the question of how to quickly model glyph‐based Geographic Information System visualizations. Our solution is based on using shape grammars to set up the different aspects of a visualization, including the geometric content of the visualization, methods for resolving layout conflicts and interaction methods. Our approach significantly increases modelling efficiency over similarly flexible systems currently in use. Pushpak Karnick, Stefan Jeschke, David Cline, Anshuman Razdan, E. Wentz, Peter Wonka |
Comput. Graph. Forum | 6 |
| 2009 | Interactive Geometric Simulation of 4D CitiesabstractAbstract We present a simulation system that can simulate a three‐dimensional urban model over time. The main novelty of our approach is that we do not rely on land‐use simulation on a regular grid, but instead build a complete and inherently geometric simulation that includes exact parcel boundaries, streets of arbitrary orientation, street widths, 3D street geometry, building footprints, and 3D building envelopes. The second novelty is the fast simulation time and user interaction at interactive speed of about 1 second per time step. Basil Weber, Pascal Müller, Peter Wonka, Markus Gross 0001 |
Comput. Graph. Forum | 3 |
| 2009 | Curve matching for open 2D curves
Ming Cui, John Femiani 0001, Jiuxiang Hu, Peter Wonka, Anshuman Razdan |
Pattern Recognit. Lett. | 4 |
| 2009 | Interactive Hyperspectral Image Visualization Using Convex OptimizationabstractIn this paper, we propose a new framework to visualize hyperspectral images. We present three goals for such a visualization: 1) preservation of spectral distances; 2) discriminability of pixels with different spectral signatures; 3) and interactive visualization for analysis. The introduced method considers all three goals at the same time and produces higher quality output than existing methods. The technical contribution of our mapping is to derive a simplified convex optimization from a complex nonlinear optimization problem. During interactive visualization, we can map the spectral signature of pixels to red, green, and blue colors using a combination of principal component analysis and linear programming. In the results, we present a quantitative analysis to demonstrate the favorable attributes of our algorithm. Ming Cui, Anshuman Razdan, Jiuxiang Hu, Peter Wonka |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2009 | Adaptive global visibility samplingabstractIn this paper we propose a global visibility algorithm which computes from-region visibility for all view cells simultaneously in a progressive manner. We cast rays to sample visibility interactions and use the information carried by a ray for all view cells it intersects. The main contribution of the paper is a set of adaptive sampling strategies based on ray mutations that exploit the spatial coherence of visibility. Our method achieves more than an order of magnitude speedup compared to per-view cell sampling. This provides a practical solution to visibility preprocessing and also enables a new type of interactive visibility analysis application, where it is possible to quickly inspect and modify a coarse global visibility solution that is constantly refined. Jirí Bittner, Oliver Mattausch, Peter Wonka, Vlastimil Havran, Michael Wimmer 0001 |
ACM Trans. Graph. | 3 |
| 2009 | A GPU Laplacian solver for diffusion curves and Poisson image editingabstractWe present a new Laplacian solver for minimal surfaces---surfaces having a mean curvature of zero everywhere except at some fixed (Dirichlet) boundary conditions. Our solution has two main contributions: First, we provide a robust rasterization technique to transform continuous boundary values (diffusion curves) to a discrete domain. Second, we define a variable stencil size diffusion solver that solves the minimal surface problem. We prove that the solver converges to the right solution, and demonstrate that it is at least as fast as commonly proposed multigrid solvers, but much simpler to implement. It also works for arbitrary image resolutions, as well as 8 bit data. We show examples of robust diffusion curve rendering where our curve rasterization and diffusion solver eliminate the strobing artifacts present in previous methods. We also show results for real-time seamless cloning and stitching of large image panoramas. Stefan Jeschke, David Cline, Peter Wonka |
ACM Trans. Graph. | 3 |
| 2009 | Rendering surface details with diffusion curvesabstractDiffusion curve images (DCI) provide a powerful tool for efficient 2D image generation, storage and manipulation. A DCI consist of curves with colors defined on either side. By diffusing these colors over the image, the final result includes sharp boundaries along the curves with smoothly shaded regions between them. This paper extends the application of diffusion curves to render high quality surface details on 3D objects. The first extension is a view dependent warping technique that dynamically reallocates texture space so that object parts that appear large on screen get more texture for increased detail. The second extension is a dynamic feature embedding technique that retains crisp, anti-aliased curve details even in extreme closeups. The third extension is the application of dynamic feature embedding to displacement mapping and geometry images. Our results show high quality renderings of diffusion curve textures, displacements, and geometry images, all rendered interactively. Stefan Jeschke, David Cline, Peter Wonka |
ACM Trans. Graph. | 3 |
| 2009 | Compressed Facade Displacement MapsabstractWe describe an approach to render massive urban models. To prevent a memory transfer bottleneck we propose to render the models from a compressed representation directly. Our solution is based on rendering crude building outlines as polygons and generating details by ray-tracing displacement maps in the fragment shader. We demonstrate how to compress a displacement map so that a decompression algorithm can selectively and quickly access individual entries in a fragment shader. Our prototype implementation shows how a massive urban model can be compressed by a factor of 85 and outperform a basic geometry-based renderer by a factor of 40 to 80 in rendering speed. Saif Ali, Jieping Ye, Anshuman Razdan, Peter Wonka |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2008 | Tiamat: A Three-Dimensional Editing Tool for Complex DNA Structures
Sean Williams, Kyle Lund, Chenxiang Lin, Peter Wonka, Stuart Lindsay |
DNA | 4 |
| 2008 | Interactive procedural street modelingabstractThis paper addresses the problem of interactively modeling large street networks. We introduce an intuitive and flexible modeling framework in which a user can create a street network from scratch or modify an existing street network. This is achieved through designing an underlying tensor field and editing the graph representing the street network. The framework is intuitive because it uses tensor fields to guide the generation of a street network. The framework is flexible because it allows the user to combine various global and local modeling operations such as brush strokes, smoothing, constraints, noise and rotation fields. Our results will show street networks and three-dimensional urban geometry of high visual quality. Guoning Chen, Gregory Esch, Peter Wonka, Pascal Müller, Eugene Zhang |
ACM Trans. Graph. | 3 |
| 2008 | Interactive visual editing of grammars for procedural architectureabstractWe introduce a real-time interactive visual editing paradigm for shape grammars, allowing the creation of rulebases from scratch without text file editing. In previous work, shape-grammar based procedural techniques were successfully applied to the creation of architectural models. However, those methods are text based, and may therefore be difficult to use for artists with little computer science background. Therefore the goal was to enable a visual work-flow combining the power of shape grammars with traditional modeling techniques. We extend previous shape grammar approaches by providing direct and persistent local control over the generated instances, avoiding the combinatorial explosion of grammar rules for modifications that should not affect all instances. The resulting visual editor is flexible: All elements of a complex state-of-the-art grammar can be created and modified visually. Markus Lipp, Peter Wonka, Michael Wimmer 0001 |
ACM Trans. Graph. | 2 |
| 2008 | Visibility-driven Mesh Analysis and Visualization through Graph CutsabstractIn this paper we present an algorithm that operates on a triangular mesh and classifies each face of a triangle as either inside or outside. We present three example applications of this core algorithm: normal orientation, inside removal, and layer-based visualization. The distinguishing feature of our algorithm is its robustness even if a difficult input model that includes holes, coplanar triangles, intersecting triangles, and lost connectivity is given. Our algorithm works with the original triangles of the input model and uses sampling to construct a visibility graph that is then segmented using graph cut. Kaichi Zhou, Eugene Zhang, Jirí Bittner, Peter Wonka |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2007 | Optimized subdivisions for preprocessed visibilityabstractThis paper describes a new tool for preprocessed visibility. It puts together view space and object space partitioning in order to control the render cost and memory cost of the visibility description generated by a visibility solver. The presented method progressively refines view space and object space subdivisions while minimizing the associated render and memory costs. Contrary to previous techniques, both subdivisions are driven by actual visibility information. We show that treating view space and object space together provides a powerful method for controlling the efficiency of the resulting visibility data structures. Oliver Mattausch, Jirí Bittner, Peter Wonka, Michael Wimmer 0001 |
Graphics Interface | 3 |
| 2007 | Fourier Shape Descriptors of Pixel Footprints for Road Extraction from Satellite ImagesabstractIn this paper, an automatic road tracking method is presented for detecting roads from satellite images. This method is based on shape classification of a local homogeneous region around a pixel. The local homogeneous region is enclosed by a polygon, called the pixel footprint. We introduce a spoke wheel operator to obtain the pixel footprint and propose a Fourier-based approach to classify footprints for automatic seeding and growing of the road tracker. We experimentally demonstrate that our proposed road tracker can extract the centerlines of roads with sharp turns and intersections effectively, and has relatively small amount of leakage. Jiuxiang Hu, Anshuman Razdan, John Femiani 0001, Peter Wonka, Ming Cui |
ICIP (1) | 4 |
| 2007 | Road Network Extraction and Intersection Detection From Aerial Images by Tracking Road FootprintsabstractIn this paper, a new two-step approach (detecting and pruning) for automatic extraction of road networks from aerial images is presented. The road detection step is based on shape classification of a local homogeneous region around a pixel. The local homogeneous region is enclosed by a polygon, called the footprint of the pixel. This step involves detecting road footprints, tracking roads, and growing a road tree. We use a spoke wheel operator to obtain the road footprint. We propose an automatic road seeding method based on rectangular approximations to road footprints and a toe-finding algorithm to classify footprints for growing a road tree. The road tree pruning step makes use of a Bayes decision model based on the area-to-perimeter ratio (the A/P ratio) of the footprint to prune the paths that leak into the surroundings. We introduce a lognormal distribution to characterize the conditional probability of A/P ratios of the footprints in the road tree and present an automatic method to estimate the parameters that are related to the Bayes decision model. Results are presented for various aerial images. Evaluation of the extracted road networks using representative aerial images shows that the completeness of our road tracker ranges from 84% to 94%, correctness is above 81%, and quality is from 82% to 92%. Jiuxiang Hu, Anshuman Razdan, John Femiani 0001, Ming Cui, Peter Wonka |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2007 | Image-based procedural modeling of facadesabstractThis paper describes algorithms to automatically derive 3D models of high visual quality from single facade images of arbitrary resolutions. We combine the procedural modeling pipeline of shape grammars with image analysis to derive a meaningful hierarchical facade subdivision. Our system gives rise to three exciting applications: urban reconstruction based on low resolution oblique aerial imagery, reconstruction of facades based on higher resolution ground-based imagery, and the automatic derivation of shape grammar rules from facade images to build a rule base for procedural modeling technology. Pascal Müller, Peter Wonka, Luc Van Gool |
ACM Trans. Graph. | 3 |
| 2007 | A new image registration scheme based on curvature scale space curve matching
Ming Cui, Peter Wonka, Anshuman Razdan, Jiuxiang Hu |
Vis. Comput. | 2 |
| 2006 | Procedural modeling of buildingsabstractCGA shape , a novel shape grammar for the procedural modeling of CG architecture, produces building shells with high visual quality and geometric detail. It produces extensive architectural models for computer games and movies, at low cost. Context sensitive shape rules allow the user to specify interactions between the entities of the hierarchical shape descriptions. Selected examples demonstrate solutions to previously unsolved modeling problems, especially to consistent mass modeling with volumetric shapes of arbitrary orientation. CGA shape is shown to efficiently generate massive urban models with unprecedented level of detail, with the virtual rebuilding of the archaeological site of Pompeii as a case in point. Pascal Müller, Peter Wonka, Simon Haegler, Andreas Ulmer, Luc Van Gool |
ACM Trans. Graph. | 2 |
| 2006 | Guided visibility samplingabstractThis paper addresses the problem of computing the triangles visible from a region in space. The proposed aggressive visibility solution is based on stochastic ray shooting and can take any triangular model as input. We do not rely on connectivity information, volumetric occluders, or the availability of large occluders, and can therefore process any given input scene. The proposed algorithm is practically memoryless, thereby alleviating the large memory consumption problems prevalent in several previous algorithms. The strategy of our algorithm is to use ray mutations in ray space to cast rays that are likely to sample new triangles. Our algorithm improves the sampling efficiency of previous work by over two orders of magnitude. Peter Wonka, Michael Wimmer 0001, Kaichi Zhou, Stefan Maierhofer, Gerd Hesina, Alexander Reshetov |
ACM Trans. Graph. | 1 |
| 2006 | Punctuated simplification of man-made objects
Justin Jang, Peter Wonka, William Ribarsky, Chris Shaw 0002 |
Vis. Comput. | 2 |
| 2005 | Fast Exact From-Region Visibility in Urban Scenes
Jirí Bittner, Peter Wonka, Michael Wimmer 0001 |
Rendering Techniques | 2 |
| 2003 | Appearance-Preserving View-Dependent VisualizationabstractIn this paper a new quadric-based view-dependent simplification scheme is presented. The scheme provides a method to connect mesh simplification controlled by a quadric error metric with a level-of-detail hierarchy that is accessed continuously and efficiently based on current view parameters. A variety of methods for determining the screen-space metric for the view calculation are implemented and evaluated, including an appearance-preserving method that has both geometry- and texture-preserving aspects. Results are presented and compared for a variety of models. Justin Jang, William Ribarsky, Chris Shaw 0002, Peter Wonka |
IEEE Visualization | 4 |
| 2003 | Instant architectureabstractThis paper presents a new method for the automatic modeling of architecture. Building designs are derived using split grammars, a new type of parametric set grammar based on the concept of shape. The paper also introduces an attribute matching system and a separate control grammar, which offer the flexibility required to model buildings using a large variety of different styles and design ideas. Through the adaptive nature of the design grammar used, the created building designs can either be generic or adhere closely to a specified goal, depending on the amount of data available. Peter Wonka, Michael Wimmer 0001, François X. Sillion, William Ribarsky |
ACM Trans. Graph. | 1 |
| 2001 | Visibility Preprocessing for Urban Scenes using Line Space SubdivisionabstractWe present an algorithm for visibility preprocessing of urban environments. The algorithm uses a subdivision of line space to analytically calculate a conservative potentially visible set for a given region in the scene. We present a detailed evaluation of our method, including a comparison to another recently published visibility preprocessing algorithm. To the best of our knowledge, the proposed method is the first algorithm that scales to large scenes and efficiently handles large view cells. Jirí Bittner, Peter Wonka, Michael Wimmer 0001 |
PG | 2 |
| 2001 | Instant VisibilityabstractWe present an online occlusion culling system which computes visibility in parallel to the rendering pipeline. We show how to use point visibility algorithms to quickly calculate a tight potentially visible set (PVS) which is valid for several frames, by shrinking the occluders used in visibility calculations by an adequate amount. These visibility calculations can be performed on a visibility server, possibly a distinct computer communicating with the display host over a local network. The resulting system essentially combines the advantages of online visibility processing and region-based visibility calculations, allowing asynchronous processing of visibility and display operations. We analyze two different types of hardware-based point visibility algorithms and address the problem of bounded calculation time which is the basis for true real-time behavior. Our results show reliable, sustained 60 Hz performance in a walkthrough with an urban environment of nearly 2 million polygons, and a terrain flyover. Peter Wonka, Michael Wimmer 0001, François X. Sillion |
Comput. Graph. Forum | 1 |
| 1999 | Occluder Shadows for Fast Walkthroughs of Urban EnvironmentsabstractThis paper describes a new algorithm that employs image‐based rendering for fast occlusion culling in complex urban environments. It exploits graphics hardware to render and automatically combine a relatively large set of occluders. The algorithm is fast to calculate and therefore also useful for scenes of moderate complexity and walkthroughs with over 20 frames per second. Occlusion is calculated dynamically and does not rely on any visibility precalculation or occluder preselection. Speed‐ups of one order of magnitude can be obtained. Peter Wonka, Dieter Schmalstieg |
Comput. Graph. Forum | 1 |