VLDB 2026 Research / reviewers in the wild / expert
Baoquan Chen
dblp:23/4197
· DBLP profile ↗
197ranked-venue papers
7as first author
71since 2021 · last 2026
0000-0003-4702-036XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 173 · 3 first-author · 61 since 2021Artificial intelligence and machine learning · 40 · 24 since 2021Human-computer interaction and ubiquitous computing · 10 · 3 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 2Systems, architecture and hardware · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Spatial-Spectral Homogeneous Attacks on Physical-World Large Vision-Language ModelsabstractAlthough large vision-language models (LVLMs) have demonstrated promising versatile capabilities on various downstream tasks, they are shown to be susceptible to adversarial examples. Existing LVLM attackers simply implement adversarial patterns in an impracticable setting: i) add digital global perturbations to entire input image; ii) access prior knowledge of LVLMs for optimization; iii) do not consider realistic transformations. These make them difficult to deploy in the physical-world attack scenarios. Motivated by the research gap and counter-practice phenomenon, this paper proposes the first practical LVLM attack method based on a novel adversarial patch design, which can achieve physical and digital attack settings without using any LVLM details. In particular, we introduce adversarial homogeneous constraints in both spatial and spectral domains to improve the patch stealthy for resisting potential real-world defenses. Besides, we also develop a new technique for synthesizing reasonably realistic transformations that capture the expected patch appearance variations in daily life. Extensive experiments are conducted to verify the strong adversarial capabilities of our proposed attack against prevalent LVLMs spanning a spectrum of tasks. Daizong Liu, Baoquan Chen, Wei Hu 0003 |
AAAI | 2 |
| 2026 | Visual permeability-driven generation of support-free stochastic porous structures
Kaifeng Tian, Lingxin Cao, Longdu Liu, Lihao Tian, Bingteng Sun, Changhe Tu, Lin Lu 0001, Baoquan Chen |
Comput. Aided Des. | 9 |
| 2026 | STAGED: Stress-Tensor Assisted Global-local-global solver for interactive Elastic shape DesignabstractAbstract We present an efficient and scalable method for the inverse shape design problem of elastic objects, with broad applicability to diverse materials and interactive editing. The core idea is to decouple material nonlinearity from geometry optimization by introducing the Cauchy stress tensor as an auxiliary variable. We design a three‐stage scheme that iteratively optimizes the stress tensors and the rest shape, with each stage being well‐posed and efficiently‐solvable. To address the lack of a theoretical convergence guarantee arising from the decoupled energy formulation, we incorporate a relaxation method that ensures robust stability in practice. As a result, our method achieves a 3 × speedup over the state‐of‐the‐art asymptotic method [Jia21] on a model with 40k vertices and 112k elements (Fig. 2), and exhibits near‐linear scalability to large systems (Fig. 8). We demonstrate applications including rest shape design for various materials (ranging from standard models to complex spline‐based materials [XSZB15]), interactive material and force editing, and elastic object reconstruction from images. Liangwang Ruan, Bin Wang 0069, Tiantian Liu 0002, Baoquan Chen |
Comput. Graph. Forum | 4 |
| 2026 | Floating-Point Robustness in Parametric Surface Continuous Collision Detection: From Algorithm to BenchmarkingabstractContinuous Collision Detection is essential in simulation and modeling for accurately identifying object collisions. While robust CCD techniques have matured for triangle meshes, ensuring floating-point robustness for parametric surfaces remains an open challenge due to their representational complexity and heightened algorithmic sensitivity. In this paper, we present the first floating-point-robust CCD framework for parametric surfaces. Built on the Time-Dependent Inclusion-Based Method (TDIBM), our approach introduces a novel error decomposition strategy that separates coefficient and arithmetic errors, enabling structured analysis and safety guarantees. To rigorously benchmark robustness, we develop a rational-arithmetic-based dataset by inverting the CCD process: we generate exact ground-truth datasets from prescribed collision outcomes. Our construction captures both typical scenarios and near-degenerate cases. We evaluate several CCD algorithms using this benchmark to provide an in-depth analysis. Together, our method and dataset establish a comprehensive foundation for analyzing, benchmarking, and improving floating-point robustness in parametric surface CCD. Code and dataset will be published upon acceptance. Xingyu Ni, Meng Zhang 0043, Bin Wang 0069, Mengyu Chu, Baoquan Chen |
ACM Trans. Graph. | 8 |
| 2025 | RainyGS: Efficient Rain Synthesis with Physically-Based Gaussian SplattingabstractWe consider the problem of adding dynamic rain effects to in-the-wild scenes in a physically-correct manner. Recent advances in scene modeling have made significant progress, with NeRF and Gaussian Splatting techniques emerging as powerful tools for reconstructing complex scenes. However, while effective for novel view synthesis, these methods typically struggle with challenging scene editing tasks, such as physics-based rain simulation. In contrast, traditional physics-based simulations can generate realistic rain effects, such as raindrops and splashes, but they often rely on skilled artists to carefully set up high-fidelity scenes. This process lacks flexibility and scalability, limiting its applicability to broader, open-world environments. In this work, we introduce RainyGS, a novel approach that leverages the strengths of both physics-based modeling and Gaussian Splatting to generate photorealistic, dynamic rain effects in open-world scenes with physical accuracy. At the core of our method is the integration of physically-based raindrop and shallow water simulation techniques within the fast Gaussian Splatting rendering framework, enabling realistic and efficient simulations of raindrop behavior, splashes, and reflections. Our method supports synthesizing rain effects at over 30 fps, offering users flexible control over rain intensity—from light drizzles to heavy downpours. We demonstrate that RainyGS performs effectively for both real-world outdoor scenes and large-scale driving scenarios, delivering more photorealistic and physically-accurate rain effects compared to state-of-the-art methods. Project page can be found at https://pku-vcl-geometry.github.io/RainyGS/. Qiyu Dai, Xingyu Ni, Qianfan Shen, Wenzheng Chen, Baoquan Chen, Mengyu Chu |
CVPR | 5 |
| 2025 | One-shot 3D Object Canonicalization based on Geometric and Semantic Consistencyabstract3D object canonicalization is a fundamental task, essential for various downstream tasks. Existing methods rely on either cumbersome manual processes or priors learned from extensive, per- category training samples. Real- world datasets, however, often exhibit long- tail distributions, challenging existing learning- based methods, especially in categories with limited samples. We address this by introducing the first one- shot category- level object canonicalization framework that operates under arbitrary poses, requiring only a single canonical model as a reference (the "prior model") for each category. To canonicalize any object, our framework first extracts semantic cues with large language models (LLMs) and vision- language models (VLMs) to establish correspondences with the prior model. We introduce a novel joint energy function to enforce geometric and semantic consistency, aligning object orientations precisely despite significant shape variations. Moreover, we adopt a support- plane strategy to reduce search space for initial poses and utilize a semantic relationship map to select the canonical pose from multiple hypotheses. Extensive experiments on multiple datasets demonstrate that our framework achieves state- of- the- art performance and validates key design choices. Using our framework, we create the Canonical Objaverse Dataset (COD), canonicalizing 32K samples in the Objaverse-LVIS dataset, underscoring the effectiveness of our framework on handling large- scale datasets. Project page at https://github.com/JinLi998/CanonObjaverseDataset Wenzheng Chen, Qiyu Dai, Qingzhe Gao, Xueying Qin, Baoquan Chen |
CVPR | 7 |
| 2025 | SLAM3R: Real-Time Dense Scene Reconstruction from Monocular RGB VideosabstractIn this paper, we introduce SLAM3R, a novel and effective system for real-time, high-quality, dense 3D reconstruction using RGB videos. SLAM3R provides an end-to-end solution by seamlessly integrating local 3D reconstruction and global coordinate registration through feed-forward neural networks. Given an input video, the system first converts it into overlapping clips using a sliding window mechanism. Unlike traditional pose optimization-based methods, SLAM3R directly regresses 3D pointmaps from RGB images in each window and progressively aligns and deforms these local pointmaps to create a globally consistent scene reconstruction-all without explicitly solving any camera parameters. Experiments across datasets consistently show that SLAM3R achieves state-of-the-art reconstruction accuracy and completeness while maintaining real-time performance at 20+ FPS. Code available at: https://github.com/PKU-VCL-3DV/SLAM3R. Yuzheng Liu, Siyan Dong, Shuzhe Wang, Yingda Yin, Yanchao Yang 0001, Qingnan Fan, Baoquan Chen |
CVPR | 7 |
| 2025 | DOF-GS: Adjustable Depth-of-Field 3D Gaussian Splatting for Post-Capture Refocusing, Defocus Rendering and Blur Removalabstract3D Gaussian Splatting (3DGS) techniques have recently enabled high-quality 3D scene reconstruction and real-time novel view synthesis. These approaches, however, are limited by the pinhole camera model and lack effective modeling of defocus effects. Departing from this, we introduce DOF-GS — a new 3DGS-based framework with a finite-aperture camera model and explicit, differentiable defocus rendering, enabling it to function as a post-capture control tool. By training with multi-view images with moderate defocus blur, DOF-GS learns inherent camera characteristics and reconstructs sharp details of the underlying scene, particularly, enabling rendering of varying DOF effects through on-demand aperture and focal distance control, post-capture and optimization. Additionally, our framework extracts circle-of-confusion cues during optimization to identify in-focus regions in input views, enhancing the reconstructed 3D scene details. Experimental results demonstrate that DOF-GS supports post-capture refocusing, adjustable defocus and high-quality all-in-focus rendering, from multi-view images with uncalibrated defocus blur. Praneeth Chakravarthula, Baoquan Chen |
CVPR | 3 |
| 2025 | GFPack++: Attention-Driven Gradient Fields for Optimizing 2D Irregular Packing
Tianyang Xue, Lin Lu 0001, Yang Liu 0014, Mingdong Wu, Hao Dong 0003, Renmin Han, Baoquan Chen |
ICCV | 8 |
| 2025 | GeoSplatting: Towards Geometry Guided Gaussian Splatting for Physically-Based Inverse RenderingabstractRecent 3D Gaussian Splatting (3DGS) representations have demonstrated remarkable performance in novel view synthesis; further, material-lighting disentanglement on 3DGS warrants relighting capabilities and its adaptability to broader applications. While the general approach to the latter operation lies in integrating differentiable physically-based rendering (PBR) techniques to jointly recover BRDF materials and environment lighting, achieving a precise disentanglement remains an inherently difficult task due to the challenge of accurately modeling light transport. Existing approaches typically approximate Gaussian points' normals, which constitute an implicit geometric constraint. However, they usually suffer from inaccuracies in normal estimation that subsequently degrade light transport, resulting in noisy material decomposition and flawed relighting results. To address this, we propose GeoSplatting, a novel approach that augments 3DGS with explicit geometry guidance for precise light transport modeling. By differentiably constructing a surface-grounded 3DGS from an optimizable mesh, our approach leverages well-defined mesh normals and the opaque mesh surface, and additionally facilitates the use of mesh-based ray tracing techniques for efficient, occlusion-aware light transport calculations. This enhancement ensures precise material decomposition while preserving the efficiency and high-quality rendering capabilities of 3DGS. Comprehensive evaluations across diverse datasets demonstrate the effectiveness of GeoSplatting, highlighting its superior efficiency and state-of-the-art inverse rendering performance. The project page can be found at https://pku-vcl-geometry.github.io/GeoSplatting/. Kai Ye 0007, Guanbin Li, Wenzheng Chen, Baoquan Chen |
ICCV | 5 |
| 2025 | Robust Single-shot Structured Light 3D Imaging via Neural Feature DecodingabstractWe consider the problem of active 3D imaging using single-shot structured light systems, which are widely employed in commercial 3D sensing devices such as Apple Face ID and Intel RealSense. Traditional structured light methods typically decode depth correspondences through pixel-domain matching algorithms, resulting in limited robustness under challenging scenarios like occlusions, fine-structured details, and non-Lambertian surfaces. Inspired by recent advances in neural feature matching, we propose a learning-based structured light decoding framework that performs robust correspondence matching within feature space rather than the fragile pixel domain. Our method extracts neural features from the projected patterns and captured infrared (IR) images, explicitly incorporating their geometric priors by building cost volumes in feature space, achieving substantial performance improvements over pixel-domain decoding approaches. To further enhance depth quality, we introduce a depth refinement module that leverages strong priors from large-scale monocular depth estimation models, improving fine detail recovery and global structural coherence. To facilitate effective learning, we develop a physically-based structured light rendering pipeline, generating nearly one million synthetic pattern-image pairs with diverse objects and materials for indoor settings. Experiments demonstrate that our method, trained exclusively on synthetic data with multiple structured light patterns, generalizes well to real-world indoor environments, effectively processes various pattern types without retraining, and consistently outperforms both commercial structured light systems and passive stereo RGB-based depth estimation methods. Code and data are available at https://github.com/Namisntimpot/NSL Qiyu Dai, Lihan Li, Praneeth Chakravarthula, He Sun 0010, Baoquan Chen, Wenzheng Chen |
SIGGRAPH Asia | 6 |
| 2025 | Neural Visibility of Point SetsabstractPoint clouds are widely used representations of 3D data, but determining the visibility of points from a given viewpoint remains a challenging problem due to their sparse nature and lack of explicit connectivity. Traditional methods, such as Hidden Point Removal (HPR), face limitations in computational efficiency, robustness to noise, and handling concave regions or low-density point clouds. In this paper, we propose a novel approach to visibility determination in point clouds by formulating it as a binary classification task. The core of our network consists of a 3D U-Net that extracts view-independent point-wise features and a shared multi-layer perceptron (MLP) that predicts point visibility using the extracted features and view direction as inputs. The network is trained end-to-end with ground-truth visibility labels generated from rendered 3D models. Our method significantly outperforms HPR in both accuracy and computational efficiency, achieving up to 126 times speedup on large point clouds. Additionally, our network demonstrates robustness to noise and varying point cloud densities and generalizes well to unseen shapes. We validate the effectiveness of our approach through extensive experiments on the ShapeNet, ABC Dataset and real-world datasets, showing substantial improvements in visibility accuracy. We also demonstrate the versatility of our method in various applications, including point cloud visualization, surface reconstruction, normal estimation, shadow rendering, and viewpoint optimization. Our code and models are available at https://github.com/octree-nn/neural-visibility. Jun-Hao Wang, Yi-Yang Tian, Baoquan Chen, Peng-Shuai Wang |
SIGGRAPH Asia | 3 |
| 2025 | FlowCapX: Physics-Grounded Flow Capture with Long-Term ConsistencyabstractAbstract We present FlowCapX , a physics‐enhanced framework for flow reconstruction from sparse video inputs, addressing the challenge of jointly optimizing complex physical constraints and sparse observational data over long time horizons. Existing methods often struggle to capture turbulent motion while maintaining physical consistency, limiting reconstruction quality and downstream tasks. Focusing on velocity inference, our approach introduces a hybrid framework that strategically separates representation and supervision across spatial scales. At the coarse level, we resolve sparse‐view ambiguities via a novel optimization strategy that aligns long‐term observation with physics‐grounded velocity fields. By emphasizing vorticity‐based physical constraints, our method enhances physical fidelity and improves optimization stability. At the fine level, we prioritize observational fidelity to preserve critical turbulent structures. Extensive experiments demonstrate state‐of‐the‐art velocity reconstruction, enabling velocity‐aware downstream tasks, e.g., accurate flow analysis, scene augmentation with tracer visualization and re‐simulation. Our implementation is released at ://github.com/taoningxiao/FlowCapX.git . Ningxiao Tao, Xingyu Ni, Mengyu Chu, Baoquan Chen |
Comput. Graph. Forum | 5 |
| 2025 | Generalized Robot Vision-Language Model via Linguistic Foreground-Aware Contrast
Kangcheng Liu, Xiaodong Han, Yong-Jin Liu 0001, Baoquan Chen |
Int. J. Comput. Vis. | 5 |
| 2025 | Correction: Generalized Robot Vision-Language Model via Linguistic Foreground-Aware Contrast
Kangcheng Liu, Xiaodong Han, Yong-Jin Liu 0001, Baoquan Chen |
Int. J. Comput. Vis. | 5 |
| 2025 | General 3D Vision-Language Model With Fast Rendering and Pre-Training Vision-Language AlignmentabstractCurrent prevailing vision-language models have achieved remarkable progress in 3D scene understanding while trained in the closed-set setting and with full labels. The major bottleneck for the current robot 3D scene recognition approach for robotic applications is that these models do not have the capacity to recognize any unseen novel classes beyond the training categories in diverse real-world robot applications such as robot manipulation as well as robot navigation. In the meantime, current state-of-the-art 3D scene understanding approaches primarily require a large number of high-quality labels to train neural networks, which merely perform well in a fully supervised manner. Therefore, we are in urgent need of a framework that can simultaneously be applicable to both 3D point cloud segmentation and detection, particularly in the circumstances where the labels are rather scarce. This work presents a generalized and straightforward framework for dealing with 3D scene understanding when the labeled scenes are quite limited. To extract knowledge for novel categories from the pre-trained vision-language models, we propose a hierarchical feature-aligned pre-training and knowledge distillation strategy to extract and distill meaningful information from large-scale vision-language models, which helps benefit the open-vocabulary scene understanding tasks. To leverage the boundary information, we propose a novel energy-based loss with boundary awareness benefiting from the region-level boundary predictions. To encourage latent instance discrimination and to guarantee efficiency, we propose the unsupervised region-level semantic contrastive learning scheme for point clouds, using confident predictions of the neural network to discriminate the intermediate feature embeddings at multiple stages. In the limited reconstruction case, our proposed approach, termed WS3D++, ranks 1st on the large-scale ScanNet benchmark on both the task of semantic segmentation and instance segmentation. Also, our proposed WS3D++ achieves state-of-the-art data-efficient learning performance on the other large-scale real-scene indoor and outdoor datasets S3DIS and SemanticKITTI. Extensive experiments with both indoor and outdoor scenes demonstrated the effectiveness of our approach in both data-efficient learning and open-world few-shot learning. Kangcheng Liu, Yong-Jin Liu 0001, Baoquan Chen |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2025 | Towards Robust Probabilistic Modeling on SO(3) via Rotation Laplace DistributionabstractEstimating the 3DoF rotation from a single RGB image is an important yet challenging problem. As a popular approach, probabilistic rotation modeling additionally carries prediction uncertainty information, compared to single-prediction rotation regression. For modeling probabilistic distribution over $\text{SO}(3)$SO(3), it is natural to use Gaussian-like Bingham distribution and matrix Fisher, however they are shown to be sensitive to outlier predictions, e.g., $180^\circ$180∘ error and thus are unlikely to converge with optimal performance. In this paper, we draw inspiration from multivariate Laplace distribution and propose a novel rotation Laplace distribution on $\text{SO}(3)$SO(3). Our rotation Laplace distribution is robust to the disturbance of outliers and enforces much gradient to the low-error region that it can improve. In addition, we show that our method also exhibits robustness to small noises and thus tolerates imperfect annotations. With this benefit, we demonstrate its advantages in semi-supervised rotation regression, where the pseudo labels are noisy. To further capture the multi-modal rotation solution space for symmetric objects, we extend our distribution to rotation Laplace mixture model and demonstrate its effectiveness. Our extensive experiments show that our proposed distribution and the mixture model achieve State-of-the-Art performance in all the rotation regression experiments over both probabilistic and non-probabilistic baselines. Yingda Yin, Jiangran Lyu, He Wang 0010, Baoquan Chen |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2025 | The Granule-In-Cell Method for Simulating Sand-Water MixturesabstractThe simulation of sand-water mixtures requires capturing the stochastic behavior of individual sand particles within a uniform, continuous fluid medium. However, most existing approaches, which only treat sand particles as markers within fluid solvers, fail to account for both the forces acting on individual sand particles and the collective feedback of the particle assemblies on the fluid. This prevents faithful reproduction of characteristic phenomena including transport, deposition, and clogging. Building upon kinetic ensemble averaging technique, we propose a physically consistent coupling strategy and introduce a novel Granule-In-Cell (GIC) method for modeling such sand-water interactions. We employ the Discrete Element Method (DEM) to capture fine-scale granule dynamics and the Particle-In-Cell (PIC) method for continuous spatial representation and density projection. To bridge these two frameworks, we treat granules as macroscopic transport flow rather than solid boundaries within the fluid domain. This bidirectional coupling allows our model to incorporate a range of interphase forces using different discretization schemes, resulting in more realistic simulations that strictly adhere to the mass conservation law. Experimental results demonstrate the effectiveness of our method in simulating complex sand-water interactions, uniquely capturing intricate physical phenomena and ensuring exact volume preservation compared to existing approaches. Yizao Tang, Yuechen Zhu, Xingyu Ni, Baoquan Chen |
ACM Trans. Graph. | 4 |
| 2024 | Control3Diff: Learning Controllable 3D Diffusion Models from Single-view ImagesabstractDiffusion models have recently become the de-facto approach for generative modeling in the 2D domain. However, extending diffusion models to 3D is challenging, due to the difficulties in acquiring 3D ground truth data for training. On the other hand, 3D GANs that integrate implicit 3D representations into GANs have shown remarkable 3D-aware generation when trained only on single-view image datasets. However, 3D GANs do not provide straightforward ways to precisely control image synthesis. To address these challenges, We present Control3Diff, a 3D diffusion model that combines the strengths of diffusion models and 3D GANs for versatile controllable 3D-aware image synthesis for single-view datasets. Control3Diff explicitly models the underlying latent distribution (optionally conditioned on external inputs), thus enabling direct control during the diffusion process. Moreover, our approach is general and applicable to any types of controlling inputs, allowing us to train it with the same diffusion objective without any auxiliary supervision. We validate the efficacy of Control3Diff on standard image generation benchmarks including FFHQ, AFHQ, and ShapeNet, using various conditioning inputs such as images, sketches, and text prompts. Jiatao Gu, Qingzhe Gao, Shuangfei Zhai, Baoquan Chen, Lingjie Liu, Joshua M. Susskind |
3DV | 4 |
| 2024 | SAI3D: Segment any Instance in 3D ScenesabstractAdvancements in 3D instance segmentation have tra-ditionally been tethered to the availability of annotated datasets, limiting their application to a narrow spectrum of object categories. Recent efforts have sought to har-ness vision-language models like CLIP for open-set semantic reasoning, yet these methods struggle to distinguish between objects of the same categories and rely on specific prompts that are not universally applicable. In this paper, we introduce SAI3D, a novel zero-shot 3D instance segmentation approach that synergistically leverages geometric priors and semantic cues derived from Segment Any-thing Model (SAM). Our method partitions a 3D scene into geometric primitives, which are then progressively merged into 3D instance segmentations that are consistent with the multi-view SAM masks. Moreover, we design a hierarchi-cal region-growing algorithm with a dynamic thresholding mechanism, which largely improves the robustness of fine-grained 3D scene parsing. Empirical evaluations on Scan-Net, Matterport3D and the more challenging ScanNet++ datasets demonstrate the superiority of our approach. No-tably, SAI3D outperforms existing open-vocabulary base-lines and even surpasses fully-supervised methods in class-agnostic segmentation on ScanNet++. Our project page is at https://yd-yin.github.io/SAI3D. Yingda Yin, Yuzheng Liu, Daniel Cohen-Or, Jingwei Huang 0001, Baoquan Chen |
CVPR | 6 |
| 2024 | Simulating Thin Shells by Bicubic Hermite Elements
Xingyu Ni, Bin Wang 0069, Baoquan Chen |
Comput. Aided Des. | 5 |
| 2024 | Cinematographic Camera Diffusion ModelabstractAbstract Designing effective camera trajectories in virtual 3D environments is a challenging task even for experienced animators. Despite an elaborate film grammar, forged through years of experience, that enables the specification of camera motions through cinematographic properties (framing, shots sizes, angles, motions), there are endless possibilities in deciding how to place and move cameras with characters. Dealing with these possibilities is part of the complexity of the problem. While numerous techniques have been proposed in the literature (optimization‐based solving, encoding of empirical rules, learning from real examples,…), the results either lack variety or ease of control. In this paper, we propose a cinematographic camera diffusion model using a transformer‐based architecture to handle temporality and exploit the stochasticity of diffusion models to generate diverse and qualitative trajectories conditioned by high‐level textual descriptions. We extend the work by integrating keyframing constraints and the ability to blend naturally between motions using latent interpolation, in a way to augment the degree of control of the designers. We demonstrate the strengths of this text‐to‐camera motion approach through qualitative and quantitative experiments and gather feedback from professional artists. The code and data are available at https://github.com/jianghd1996/Camera-control . Hongda Jiang, Xi Wang 0024, Marc Christie, Libin Liu 0002, Baoquan Chen |
Comput. Graph. Forum | 5 |
| 2024 | Embodied computational imaging: a new paradigm for observing and analyzing spatiotemporally ultrasensitive phenomena at multiple scales
Baoquan Chen, Zhouchen Lin, Peng Xi, Yebin Liu, Xiaodian Chen |
Sci. China Inf. Sci. | 1 |
| 2024 | Message from the Best Paper Award Committeeabstractthe Best Paper Award Committee to select the Best Paper.After careful deliberation, the following paper was chosen with the unanimous consensus as the winner, on the basis of its intellectual merit and potential impact:Visual attention network [1] Two other papers were awarded an Ming C. Lin, Baoquan Chen, Ying He 0001, Wenping Wang 0001, Ralph R. Martin |
Comput. Vis. Media | 2 |
| 2024 | SinGRAV: Learning a Generative Radiance Volume from a Single Natural Scene
Xuelin Chen, Baoquan Chen |
J. Comput. Sci. Technol. | 3 |
| 2024 | A Time-Dependent Inclusion-Based Method for Continuous Collision Detection between Parametric SurfacesabstractContinuous collision detection (CCD) between parametric surfaces is typically formulated as a five-dimensional constrained optimization problem. In the field of CAD and computer graphics, common approaches to solving this problem rely on linearization or sampling strategies. Alternatively, inclusion-based techniques detect collisions by employing 5D inclusion functions, which are typically designed to represent the swept volumes of parametric surfaces over a given time span, and narrowing down the earliest collision moment through subdivision in both spatial and temporal dimensions. However, when high detection accuracy is required, all these approaches significantly increases computational consumption due to the high-dimensional searching space. In this work, we develop a new time-dependent inclusion-based CCD framework that eliminates the need for temporal subdivision and can speedup conventional methods by a factor ranging from 36 to 138. To achieve this, we propose a novel time-dependent inclusion function that provides a continuous representation of a moving surface, along with a corresponding intersection detection algorithm that quickly identifies the time intervals when collisions are likely to occur. We validate our method across various primitive types, demonstrate its efficacy within the simulation pipeline and show that it significantly improves CCD efficiency while maintaining accuracy. Xingyu Ni, Mengyu Chu, Bin Wang 0069, Baoquan Chen |
ACM Trans. Graph. | 6 |
| 2024 | An Induce-on-Boundary Magnetostatic Solver for Grid-Based FerrofluidsabstractThis paper introduces a novel Induce-on-Boundary (IoB) solver designed to address the magnetostatic governing equations of ferrofluids. The IoB solver is based on a single-layer potential and utilizes only the surface point cloud of the object, offering a lightweight, fast, and accurate solution for calculating magnetic fields. Compared to existing methods, it eliminates the need for complex linear system solvers and maintains minimal computational complexities. Moreover, it can be seamlessly integrated into conventional fluid simulators without compromising boundary conditions. Through extensive theoretical analysis and experiments, we validate both the convergence and scalability of the IoB solver, achieving state-of-the-art performance. Additionally, a straightforward coupling approach is proposed and executed to showcase the solver's effectiveness when integrated into a grid-based fluid simulation pipeline, allowing for realistic simulations of representative ferrofluid instabilities. Xingyu Ni, Ruicheng Wang, Bin Wang 0069, Baoquan Chen |
ACM Trans. Graph. | 4 |
| 2024 | MiNNIE: a Mixed Multigrid Method for Real-time Simulation of Nonlinear Near-Incompressible ElasticsabstractWe propose MiNNIE, a simple yet comprehensive framework for real-time simulation of nonlinear near-incompressible elastics. To avoid the common volumetric locking issues at high Poisson's ratios of linear finite element methods (FEM), we build MiNNIE upon a mixed FEM framework and further incorporate a pressure stabilization term to ensure excellent convergence of multigrid solvers. Our pressure stabilization strategy injects bounded influence on nodal displacement which can be eliminated using a quasiNewton method. MiNNIE has a specially tailored GPU multigrid solver including a modified skinning-space interpolation scheme, a novel vertex Vanka smoother, and an efficient dense solver using Schur complement. MiNNIE supports various elastic material models and simulates them in real-time, supporting a full range of Poisson's ratios up to 0.5 while handling large deformations, element inversions, and self-collisions at the same time. Liangwang Ruan, Bin Wang 0069, Tiantian Liu 0002, Baoquan Chen |
ACM Trans. Graph. | 4 |
| 2024 | A Vortex Particle-on-Mesh Method for Soap Film SimulationabstractThis paper introduces a novel physically-based vortex fluid model for films, aimed at accurately simulating cascading vortical structures on deforming thin films. Central to our approach is a novel mechanism decomposing the film's tangential velocity into circulation and dilatation components. These components are then evolved using a hybrid particle-mesh method, enabling the effective reconstruction of three-dimensional tangential velocities and seamlessly integrating surfactant and thickness dynamics into a unified framework. By coupling with its normal component and surface-tension model, our method is particularly adept at depicting complex interactions between in-plane vortices and out-of-plane physical phenomena, such as gravity, surfactant dynamics, and solid boundary, leading to highly realistic simulations of complex thin-film dynamics, achieving an unprecedented level of vortical details and physical realism. Ningxiao Tao, Liangwang Ruan, Yitong Deng, Bo Zhu 0002, Bin Wang 0069, Baoquan Chen |
ACM Trans. Graph. | 6 |
| 2024 | MoConVQ: Unified Physics-Based Motion Control via Scalable Discrete RepresentationsabstractIn this work, we present MoConVQ, a novel unified framework for physics-based motion control leveraging scalable discrete representations. Building upon vector quantized variational autoencoders (VQ-VAE) and model-based reinforcement learning, our approach effectively learns motion embeddings from a large, unstructured dataset spanning tens of hours of motion examples. The resultant motion representation not only captures diverse motion skills but also offers a robust and intuitive interface for various applications. We demonstrate the versatility of MoConVQ through several applications: universal tracking control from various motion sources, interactive character control with latent motion representations using supervised learning, physics-based motion generation from natural language descriptions using the GPT framework, and, most interestingly, seamless integration with large language models (LLMs) with in-context learning to tackle complex and abstract tasks. Heyuan Yao, Zhenhua Song, Tenglong Ao, Baoquan Chen, Libin Liu 0002 |
ACM Trans. Graph. | 5 |
| 2024 | Semantic Gesticulator: Semantics-Aware Co-Speech Gesture SynthesisabstractIn this work, we present Semantic Gesticulator , a novel framework designed to synthesize realistic gestures accompanying speech with strong semantic correspondence. Semantically meaningful gestures are crucial for effective non-verbal communication, but such gestures often fall within the long tail of the distribution of natural human motion. The sparsity of these movements makes it challenging for deep learning-based systems, trained on moderately sized datasets, to capture the relationship between the movements and the corresponding speech semantics. To address this challenge, we develop a generative retrieval framework based on a large language model. This framework efficiently retrieves suitable semantic gesture candidates from a motion library in response to the input speech. To construct this motion library, we summarize a comprehensive list of commonly used semantic gestures based on findings in linguistics, and we collect a high-quality motion dataset encompassing both body and hand movements. We also design a novel GPT-based model with strong generalization capabilities to audio, capable of generating high-quality gestures that match the rhythm of speech. Furthermore, we propose a semantic alignment mechanism to efficiently align the retrieved semantic gestures with the GPT's output, ensuring the naturalness of the final animation. Our system demonstrates robustness in generating gestures that are rhythmically coherent and semantically explicit, as evidenced by a comprehensive collection of examples. User studies confirm the quality and human-likeness of our results, and show that our system outperforms state-of-the-art systems in terms of semantic appropriateness by a clear margin. We will release the code and dataset for academic research. Tenglong Ao, Yuyao Zhang 0003, Qingzhe Gao, Baoquan Chen, Libin Liu 0002 |
ACM Trans. Graph. | 6 |
| 2024 | Neural Novel Actor: Learning a Generalized Animatable Neural Representation for Human ActorsabstractWe propose a new method for learning a generalized animatable neural human representation from a sparse set of multi-view imagery of multiple persons. The learned representation can be used to synthesize novel view images of an arbitrary person and further animate them with the user's pose control. While most existing methods can either generalize to new persons or synthesize animations with user control, none of them can achieve both at the same time. We attribute this accomplishment to the employment of a 3D proxy for a shared multi-person human model, and further the warping of the spaces of different poses to a shared canonical pose space, in which we learn a neural field and predict the person- and pose-dependent deformations, as well as appearance with the features extracted from input images. To cope with the complexity of the large variations in body shapes, poses, and clothing deformations, we design our neural human model with disentangled geometry and appearance. Furthermore, we utilize the image features both at the spatial point and on the surface points of the 3D proxy for predicting person- and pose-dependent properties. Experiments show that our method significantly outperforms the state-of-the-arts on both tasks. Qingzhe Gao, Yiming Wang 0009, Libin Liu 0002, Lingjie Liu, Christian Theobalt, Baoquan Chen |
IEEE Trans. Vis. Comput. Graph. | 6 |
| 2024 | Optimally Ordered Orthogonal Neighbor Joining Trees for Hierarchical Cluster AnalysisabstractNJ) trees as a new way to visually explore cluster structures and outliers in multi-dimensional data. Neighbor-joining (NJ) trees are widely used in biology, and their visual representation is similar to that of dendrograms. The core difference to dendrograms, however, is that NJ trees correctly encode distances between data points, resulting in trees with varying edge lengths. We optimize NJ trees for their use in visual analysis in two ways. First, we propose to use a novel leaf sorting algorithm that helps users to better interpret adjacencies and proximities within such a tree. Second, we provide a new method to visually distill the cluster tree from an ordered NJ tree. Numerical evaluation and three case studies illustrate the benefits of this approach for exploring multi-dimensional data in areas such as biology or image analysis. Tong Ge, Yunhai Wang, Michael Sedlmair, Zhanglin Cheng, Ying Zhao 0001, Xin Liu 0007, Oliver Deussen, Baoquan Chen |
IEEE Trans. Vis. Comput. Graph. | 9 |
| 2024 | Using Multi-Level Consistency Learning for Partial-to-Partial Point Cloud RegistrationabstractPoint cloud registration is a basic task in computer vision and computer graphics. Recently, deep learning-based end-to-end methods have made great progress in this field. One of the challenges of these methods is to deal with partial-to-partial registration tasks. In this work, we propose a novel end-to-end framework called MCLNet that makes full use of multi-level consistency for point cloud registration. First, the point-level consistency is exploited to prune points located outside overlapping regions. Second, we propose a multi-scale attention module to perform consistency learning at the correspondence-level for obtaining reliable correspondences. To further improve the accuracy of our method, we propose a novel scheme to estimate the transformation based on geometric consistency between correspondences. Compared to baseline methods, experimental results show that our method performs well on smaller-scale data, especially with exact matches. The reference time and memory footprint of our method are relatively balanced, which is more beneficial for practical applications. Boyuan Tan, Hongxing Qin, Yiqun Wang 0001, Tao Xiang 0001, Baoquan Chen |
IEEE Trans. Vis. Comput. Graph. | 6 |
| 2023 | Patch-Based 3D Natural Scene Generation from a Single ExampleabstractWe target a 3D generative model for general natural scenes that are typically unique and intricate. Lacking the necessary volumes of training data, along with the difficulties of having ad hoc designs in presence of varying scene characteristics, renders existing setups intractable. Inspired by classical patch-based image models, we advocate for synthesizing 3D scenes at the patch level, given a single example. At the core of this work lies important algorithmic designs w.r.t the scene representation and generative patch nearest-neighbor module, that address unique challenges arising from lifting classical 2D patch-based framework to 3D generation. These design choices, on a collective level, contribute to a robust, effective, and efficient model that can generate high-quality general natural scenes with both realistic geometric structure and visual appearance, in large quantities and varieties, as demonstrated upon a variety of exemplar scenes. Data and code can be found at http://wyysf-98.github.io/Sin3DGen. Xuelin Chen, Jue Wang 0001, Baoquan Chen |
CVPR | 4 |
| 2023 | Delving into Discrete Normalizing Flows on SO(3) Manifold for Probabilistic Rotation ModelingabstractNormalizing flows (NFs) provide a powerful tool to construct an expressive distribution by a sequence of trackable transformations of a base distribution and form a probabilistic model of underlying data. Rotation, as an important quantity in computer vision, graphics, and robotics, can exhibit many ambiguities when occlusion and symmetry occur and thus demands such probabilistic models. Though much progress has been made for NFs in Euclidean space, there are no effective normalizing flows without discontinuity or many-to-one mapping tailored for SO(3) manifold. Given the unique non-Euclidean properties of the rotation manifold, adapting the existing NFs to SO(3) manifold is non-trivial. In this paper, we propose a novel normalizing flow on SO (3) by combining a Mobius transformation-based coupling layer and a quaternion affine transformation. With our proposed rotation normalizing flows, one can not only effectively express arbitrary distributions on SO(3), but also conditionally build the target distribution given input observations. Extensive experiments show that our rotation normalizing flows significantly outperform the baselines on both unconditional and conditional tasks. Yulin Liu 0003, Yingda Yin, Baoquan Chen, He Wang 0010 |
CVPR | 5 |
| 2023 | Neural Implicit 3D Shapes from Single Images with Spatial Patterns
Yixin Zhuang, Yunzhe Liu 0002, Baoquan Chen |
ICIG (5) | 4 |
| 2023 | A Laplace-inspired Distribution on SO(3) for Probabilistic Rotation Estimation
Yingda Yin, He Wang 0010, Baoquan Chen |
ICLR | 4 |
| 2023 | Learning Gradient Fields for Scalable and Generalizable Irregular PackingabstractThe packing problem, also known as cutting or nesting, has diverse applications in logistics, manufacturing, layout design, and atlas generation. It involves arranging irregularly shaped pieces to minimize waste while avoiding overlap. Recent advances in machine learning, particularly reinforcement learning, have shown promise in addressing the packing problem. In this work, we delve deeper into a novel machine learning-based approach that formulates the packing problem as conditional generative modeling. To tackle the challenges of irregular packing, including object validity constraints and collision avoidance, our method employs the score-based diffusion model to learn a series of gradient fields. These gradient fields encode the correlations between constraint satisfaction and the spatial relationships of polygons, learned from teacher examples. During the testing phase, packing solutions are generated using a coarse-to-fine refinement mechanism guided by the learned gradient fields. To enhance packing feasibility and optimality, we introduce two key architectural designs: multi-scale feature extraction and coarse-to-fine relation extraction. We conduct experiments on two typical industrial packing domains, considering translations only. Empirically, our approach demonstrates spatial utilization rates comparable to, or even surpassing, those achieved by the teacher algorithm responsible for training data generation. Additionally, it exhibits some level of generalization to shape variations. We are hopeful that this method could pave the way for new possibilities in solving the packing problem. Tianyang Xue, Mingdong Wu, Lin Lu 0001, Hao Dong 0003, Baoquan Chen |
SIGGRAPH Asia | 6 |
| 2023 | MILO: Multi-Bounce Inverse Rendering for Indoor Scene With Light-Emitting ObjectsabstractRecently, many advances in inverse rendering are achieved by high-dimensional lighting representations and differentiable rendering. However, multi-bounce lighting effects can hardly be handled correctly in scene editing using high-dimensional lighting representations, and light source model deviation and ambiguities exist in differentiable rendering methods. These problems limit the applications of inverse rendering. In this paper, we present a multi-bounce inverse rendering method based on Monte Carlo path tracing, to enable correct complex multi-bounce lighting effects rendering in scene editing. We propose a novel light source model that is more suitable for light source editing in indoor scenes, and design a specific neural network with corresponding disambiguation constraints to alleviate ambiguities during the inverse rendering. We evaluate our method on both synthetic and real indoor scenes through virtual object insertion, material editing, relighting tasks, and so on. The results demonstrate that our method achieves better photo-realistic quality. Bohan Yu, Xuanning Cui, Siyan Dong, Baoquan Chen, Boxin Shi |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2023 | Example-based Motion Synthesis via Generative Motion MatchingabstractWe present GenMM, a generative model that "mines" as many diverse motions as possible from a single or few example sequences. In stark contrast to existing data-driven methods, which typically require long offline training time, are prone to visual artifacts, and tend to fail on large and complex skeletons, GenMM inherits the training-free nature and the superior quality of the well-known Motion Matching method. GenMM can synthesize a high-quality motion within a fraction of a second, even with highly complex and large skeletal structures. At the heart of our generative framework lies the generative motion matching module, which utilizes the bidirectional visual similarity as a generative cost function to motion matching, and operates in a multi-stage framework to progressively refine a random guess using exemplar motion matches. In addition to diverse motion generation, we show the versatility of our generative framework by extending it to a number of scenarios that are not possible with motion matching alone, including motion completion, key frame-guided generation, infinite looping, and motion reassembly. Xuelin Chen, Peizhuo Li, Olga Sorkine-Hornung, Baoquan Chen |
ACM Trans. Graph. | 5 |
| 2023 | GARM-LS: A Gradient-Augmented Reference-Map Method for Level-Set Fluid Simulationabstractresearch-article Share on GARM-LS: A Gradient-Augmented Reference-Map Method for Level-Set Fluid Simulation Authors: Xingqiao Li School of IST & National Key Lab. of AGI, Peking University, China School of IST & National Key Lab. of AGI, Peking University, China 0000-0002-8131-6140View Profile , Xingyu Ni School of CS & National Key Lab. of AGI, Peking University, China School of CS & National Key Lab. of AGI, Peking University, China 0000-0003-1127-2848View Profile , Bo Zhu Georgia Institute of Technology, United States of America and Dartmouth College, United States of America Georgia Institute of Technology, United States of America and Dartmouth College, United States of America 0000-0002-1392-0928View Profile , Bin Wang Beijing Institute for General Artificial Intelligence, China Beijing Institute for General Artificial Intelligence, China 0000-0001-9496-772XView Profile , Baoquan Chen School of IST & National Key Lab. of AGI, Peking University, China School of IST & National Key Lab. of AGI, Peking University, China 0000-0003-4702-036XView Profile Authors Info & Claims ACM Transactions on GraphicsVolume 42Issue 6Article No.: 192pp 1–20https://doi.org/10.1145/3618377Published:05 December 2023Publication History 1citation71DownloadsMetricsTotal Citations1Total Downloads71Last 12 Months71Last 6 weeks23 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteGet Access Xingqiao Li, Xingyu Ni, Bo Zhu 0002, Bin Wang 0069, Baoquan Chen |
ACM Trans. Graph. | 5 |
| 2023 | VASCO: Volume and Surface Co-Decomposition for Hybrid ManufacturingabstractAdditive and subtractive hybrid manufacturing (ASHM) involves the alternating use of additive and subtractive manufacturing techniques, which provides unique advantages for fabricating complex geometries with otherwise inaccessible surfaces. However, a significant challenge lies in ensuring tool accessibility during both fabrication procedures, as the object shape may change dramatically, and different parts of the shape are interdependent. In this study, we propose a computational framework to optimize the planning of additive and subtractive sequences while ensuring tool accessibility. Our goal is to minimize the switching between additive and subtractive processes to achieve efficient fabrication while maintaining product quality. We approach the problem by formulating it as a Volume-And-Surface-CO-decomposition (VASCO) problem. First, we slice volumes into slabs and build a dynamic-directed graph to encode manufacturing constraints, with each node representing a slab and direction reflecting operation order. We introduce a novel geometry property called hybrid-fabricability for a pair of additive and subtractive procedures. Then, we propose a beam-guided top-down block decomposition algorithm to solve the VASCO problem. We apply our solution to a 5-axis hybrid manufacturing platform and evaluate various 3D shapes. Finally, we assess the performance of our approach through both physical and simulated manufacturing evaluations. Fanchao Zhong, Haisen Zhao, Jikai Liu, Baoquan Chen, Lin Lu 0001 |
ACM Trans. Graph. | 6 |
| 2022 | Visual Localization via Few-Shot Scene Region ClassificationabstractVisual (re)localization addresses the problem of estimating the 6-DoF (Degree of Freedom) camera pose of a query image captured in a known scene, which is a key building block of many computer vision and robotics applications. Recent advances in structure-based localization solve this problem by memorizing the mapping from image pixels to scene coordinates with neural networks to build 2D-3D correspondences for camera pose optimization. However, such memorization requires training by amounts of posed images in each scene, which is heavy and inefficient. On the contrary, few-shot images are usually sufficient to cover the main regions of a scene for a human operator to perform visual localization. In this paper, we propose a scene region classification approach to achieve fast and effective scene memorization with few-shot images. Our insight is leveraging a) pre-learned feature extractor, b) scene region classifier, and c) meta-learning strategy to accelerate training while mitigating overfitting. We evaluate our method on both indoor and outdoor benchmarks. The experiments validate the effectiveness of our method in the few-shot setting, and the training time is significantly reduced to only a few minutes.11Code available at: https://github.com/siyandong/SRC Siyan Dong, Shuzhe Wang, Yixin Zhuang, Juho Kannala, Marc Pollefeys, Baoquan Chen |
3DV | 6 |
| 2022 | Projective Manifold Gradient Layer for Deep Rotation RegressionabstractRegressing rotations on SO(3) manifold using deep neural networks is an important yet unsolved problem. The gap between the Euclidean network output space and the non-Euclidean SO(3) manifold imposes a severe challenge for neural network learning in both forward and backward passes. While several works have proposed different regression-friendly rotation representations, very few works have been devoted to improving the gradient back-propagating in the backward pass. In this paper, we propose a manifold-aware gradient that directly backpropagates into deep network weights. Leveraging Riemannian optimization to construct a novel projective gradient, our proposed regularized projective manifold gradient (RPMG) method helps networks achieve new state-of-the-art performance in a variety of rotation estimation tasks. Our proposed gradient layer can also be applied to other smooth manifolds such as the unit sphere. Our project page is at https://jychen18.github.io/RPMG. Jiayi Chen 0003, Yingda Yin, Tolga Birdal, Baoquan Chen, Leonidas J. Guibas, He Wang 0010 |
CVPR | 4 |
| 2022 | Multi-Robot Active Mapping via Neural Bipartite Graph MatchingabstractWe study the problem of multi-robot active mapping, which aims for complete scene map construction in minimum time steps. The key to this problem lies in the goal position estimation to enable more efficient robot movements. Previous approaches either choose the frontier as the goal position via a myopic solution that hinders the time efficiency, or maximize the long-term value via reinforcement learning to directly regress the goal position, but does not guarantee the complete map construction. In this paper, we propose a novel algorithm, namely NeuralCoMapping, which takes advantage of both approaches. We reduce the problem to bipartite graph matching, which establishes the node correspondences between two graphs, denoting robots and frontiers. We introduce a multiplex graph neural network (mGNN) that learns the neural distance to fill the affinity matrix for more effective graph matching. We optimize the mGNN with a differentiable linear assignment layer by maximizing the long-term values that favor time efficiency and map completeness via reinforcement learning. We compare our algorithm with several state-of-the-art multi-robot active mapping approaches and adapted reinforcement-learning baselines. Experimental results demonstrate the superior performance and exceptional generalization ability of our algorithm on various indoor scenes and unseen number of robots, when only trained with 9 indoor scenes. Kai Ye 0007, Siyan Dong, Qingnan Fan, He Wang 0010, Li Yi 0001, Fei Xia 0002, Jue Wang 0001, Baoquan Chen |
CVPR | 8 |
| 2022 | FisherMatch: Semi-Supervised Rotation Regression via Entropy-based FilteringabstractEstimating the 3DoF rotation from a single RGB image is an important yet challenging problem. Recent works achieve good performance relying on a large amount of expensive-to-obtain labeled data. To reduce the amount of supervision, we for the first time propose a general framework, FisherMatch, for semi-supervised rotation regression, without assuming any domain-specific knowledge or paired data. Inspired by the popular semi-supervised approach, FixMatch, we propose to leverage pseudo label filtering to facilitate the information flow from labeled data to unlabeled data in a teacher-student mutual learning framework. However, incorporating the pseudo label filtering mechanism into semi-supervised rotation regression is highly non-trivial, mainly due to the lack of a reliable confidence measure for rotation prediction. In this work, we propose to leverage matrix Fisher distribution to build a probabilistic model of rotation and devise a matrix Fisher-based regressor for jointly predicting rotation along with its prediction uncertainty. We then propose to use the entropy of the predicted distribution as a confidence measure, which enables us to perform pseudo label filtering for rotation regression. For supervising such distribution-like pseudo labels, we further investigate the problem of how to enforce loss between two matrix Fisher distributions. Our extensive experiments show that our method can work well even under very low labeled data ratios on different benchmarks, achieving significant and consistent performance improvement over supervised learning and other semi-supervised learning baselines. Our project page is at https://yd-yin.github.io/FisherMatch. Yingda Yin, Yingcheng Cai, He Wang 0010, Baoquan Chen |
CVPR | 4 |
| 2022 | Towards Accurate Active Camera Localization
Qihang Fang, Yingda Yin, Qingnan Fan, Fei Xia 0002, Siyan Dong, Jue Wang 0001, Leonidas J. Guibas, Baoquan Chen |
ECCV (10) | 9 |
| 2022 | MoCo-Flow: Neural Motion Consensus Flow for Dynamic Humans in Stationary Monocular CamerasabstractAbstract Synthesizing novel views of dynamic humans from stationary monocular cameras is a specialized but desirable setup. This is particularly attractive as it does not require static scenes, controlled environments, or specialized capture hardware. In contrast to techniques that exploit multi‐view observations, the problem of modeling a dynamic scene from a single view is significantly more under‐constrained and ill‐posed. In this paper, we introduce Neural Motion Consensus Flow (MoCo‐Flow), a representation that models dynamic humans in stationary monocular cameras using a 4D continuous time‐variant function. We learn the proposed representation by optimizing for a dynamic scene that minimizes the total rendering error, over all the observed images. At the heart of our work lies a carefully designed optimization scheme, which includes a dedicated initialization step and is constrained by a motion consensus regularization on the estimated motion flow. We extensively evaluate MoCo‐Flow on several datasets that contain human motions of varying complexity, and compare, both qualitatively and quantitatively, to several baselines and ablated variations of our methods, showing the efficacy and merits of the proposed approach. Pretrained model, code, and data will be released for research purposes upon paper acceptance. Xuelin Chen, Daniel Cohen-Or, Niloy J. Mitra, Baoquan Chen |
Comput. Graph. Forum | 5 |
| 2022 | Rigid Registration of Point Clouds Based on Partial Optimal TransportabstractAbstract For rigid point cloud data registration, algorithms based on soft correspondences are more robust than the traditional ICP method and its variants. However, point clouds with severe outliers and missing data may lead to imprecise many‐to‐many correspondences and consequently inaccurate registration. In this study, we propose a point cloud registration algorithm based on partial optimal transport via a hard marginal constraint. The hard marginal constraint provides an explicit parameter to adjust the ratio of points that should be accurately matched, and helps avoid incorrect many‐to‐many correspondences. Experiments show that the proposed method achieves state‐of‐the‐art registration results when dealing with point clouds with significant amount of outliers and missing points (see https://www.acm.org/publications/class‐2012 ). Hongxing Qin, Baoquan Chen |
Comput. Graph. Forum | 4 |
| 2022 | DO-Conv: Depthwise Over-Parameterized Convolutional LayerabstractConvolutional layers are the core building blocks of Convolutional Neural Networks (CNNs). In this paper, we propose to augment a convolutional layer with an additional depthwise convolution, where each input channel is convolved with a different 2D kernel. The composition of the two convolutions constitutes an over-parameterization, since it adds learnable parameters, while the resulting linear operation can be expressed by a single convolution layer. We refer to this depthwise over-parameterized convolutional layer as DO-Conv, which is a novel way of over-parameterization. We show with extensive experiments that the mere replacement of conventional convolutional layers with DO-Conv layers boosts the performance of CNNs on many classical vision tasks, such as image classification, detection, and segmentation. Moreover, in the inference phase, the depthwise convolution is folded into the conventional convolution, reducing the computation to be exactly equivalent to that of a convolutional layer without over-parameterization. As DO-Conv introduces performance gains without incurring any computational complexity increase for inference, we advocate it as an alternative to the conventional convolutional layer. We open sourced an implementation of DO-Conv in Tensorflow, PyTorch and GluonCV at https://github.com/yangyanli/DO-Conv. Jinming Cao, Yangyan Li, Mingchao Sun, Dani Lischinski, Daniel Cohen-Or, Baoquan Chen, Changhe Tu |
IEEE Trans. Image Process. | 7 |
| 2022 | Rhythmic Gesticulator: Rhythm-Aware Co-Speech Gesture Synthesis with Hierarchical Neural EmbeddingsabstractAutomatic synthesis of realistic co-speech gestures is an increasingly important yet challenging task in artificial embodied agent creation. Previous systems mainly focus on generating gestures in an end-to-end manner, which leads to difficulties in mining the clear rhythm and semantics due to the complex yet subtle harmony between speech and gestures. We present a novel co-speech gesture synthesis method that achieves convincing results both on the rhythm and semantics. For the rhythm, our system contains a robust rhythm-based segmentation pipeline to ensure the temporal coherence between the vocalization and gestures explicitly. For the gesture semantics, we devise a mechanism to effectively disentangle both low- and high-level neural embeddings of speech and motion based on linguistic theory. The high-level embedding corresponds to semantics, while the low-level embedding relates to subtle variations. Lastly, we build correspondence between the hierarchical embeddings of the speech and the motion, resulting in rhythm- and semantics-aware gesture synthesis. Evaluations with existing objective metrics, a newly proposed rhythmic metric, and human feedback show that our method outperforms state-of-the-art systems by a clear margin. Tenglong Ao, Qingzhe Gao, Yuke Lou, Baoquan Chen, Libin Liu 0002 |
ACM Trans. Graph. | 4 |
| 2022 | Simulation and optimization of magnetoelastic thin shellsabstractMagnetoelastic thin shells exhibit great potential in realizing versatile functionalities through a broad range of combination of material stiffness, remnant magnetization intensity, and external magnetic stimuli. In this paper, we propose a novel computational method for forward simulation and inverse design of magnetoelastic thin shells. Our system consists of two key components of forward simulation and backward optimization. On the simulation side, we have developed a new continuum mechanics model based on the Kirchhoff-Love thin-shell model to characterize the behaviors of a megnetolelastic thin shell under external magnetic stimuli. Based on this model, we proposed an implicit numerical simulator facilitated by the magnetic energy Hessian to treat the elastic and magnetic stresses within a unified framework, which is versatile to incorporation with other thin shell models. On the optimization side, we have devised a new differentiable simulation framework equipped with an efficient adjoint formula to accommodate various PDE-constraint, inverse design problems of magnetoelastic thin-shell structures, in both static and dynamic settings. It also encompasses applications of magnetoelastic soft robots, functional Origami, artworks, and meta-material designs. We demonstrate the efficacy of our framework by designing and simulating a broad array of magnetoelastic thin-shell objects that manifest complicated interactions between magnetic fields, materials, and control policies. Xingyu Ni, Bo Zhu 0002, Bin Wang 0069, Baoquan Chen |
ACM Trans. Graph. | 5 |
| 2022 | Computational Object-Wrapping Rope NetsabstractWrapping objects using ropes is a common practice in our daily life. However, it is difficult to design and tie ropes on a 3D object with complex topology and geometry features while ensuring wrapping security and easy operation. In this article, we propose to compute a rope net that can tightly wrap around various 3D shapes. Our computed rope net not only immobilizes the object but also maintains the load balance during lifting. Based on the key observation that if every knot of the net has four adjacent curve edges, then only a single rope is needed to construct the entire net. We reformulate the rope net computation problem into a constrained curve network optimization. We propose a discrete-continuous optimization approach, where the topological constraints are satisfied in the discrete phase and the geometrical goals are achieved in the continuous stage. We also develop a hoist planning to pick anchor points so that the rope net equally distributes the load during hoisting. Furthermore, we simulate the wrapping process and use it to guide the physical rope net construction process. We demonstrate the effectiveness of our method on 3D objects with varying geometric and topological complexity. In addition, we conduct physical experiments to demonstrate the practicability of our method. Shi-Qing Xin, Xifeng Gao, Kaihang Gao, Kai Xu 0004, Baoquan Chen, Changhe Tu |
ACM Trans. Graph. | 6 |
| 2022 | Joint neural phase retrieval and compression for energy- and computation-efficient holography on the edgeabstractRecent deep learning approaches have shown remarkable promise to enable high fidelity holographic displays. However, lightweight wearable display devices cannot afford the computation demand and energy consumption for hologram generation due to the limited onboard compute capability and battery life. On the other hand, if the computation is conducted entirely remotely on a cloud server, transmitting lossless hologram data is not only challenging but also result in prohibitively high latency and storage. In this work, by distributing the computation and optimizing the transmission, we propose the first framework that jointly generates and compresses high-quality phase-only holograms. Specifically, our framework asymmetrically separates the hologram generation process into high-compute remote encoding (on the server), and low-compute decoding (on the edge) stages. Our encoding enables light weight latent space data, thus faster and efficient transmission to the edge device. With our framework, we observed a reduction of 76% computation and consequently 83% in energy cost on edge devices, compared to the existing hologram generation methods. Our framework is robust to transmission and decoding errors, and approach high image fidelity for as low as 2 bits-per-pixel, and further reduced average bit-rates and decoding time for holographic videos. Praneeth Chakravarthula, Qi Sun 0003, Baoquan Chen |
ACM Trans. Graph. | 4 |
| 2022 | Position-Based Surface Tension FlowabstractThis paper presents a novel approach to simulating surface tension flow within a position-based dynamics (PBD) framework. We enhance the conventional PBD fluid method in terms of its surface representation and constraint enforcement to furnish support for the simulation of interfacial phenomena driven by strong surface tension and contact dynamics. The key component of our framework is an on-the-fly local meshing algorithm to build the local geometry around each surface particle. Based on this local mesh structure, we devise novel surface constraints that can be integrated seamlessly into a PBD framework to model strong surface tension effects. We demonstrate the efficacy of our approach by simulating a multitude of surface tension flow examples exhibiting intricate interfacial dynamics of films and drops, which were all infeasible for a traditional PBD method. Jingrui Xing, Liangwang Ruan, Bin Wang 0069, Bo Zhu 0002, Baoquan Chen |
ACM Trans. Graph. | 5 |
| 2022 | ControlVAE: Model-Based Learning of Generative Controllers for Physics-Based CharactersabstractIn this paper, we introduce ControlVAE, a novel model-based framework for learning generative motion control policies based on variational autoencoders (VAE). Our framework can learn a rich and flexible latent representation of skills and a skill-conditioned generative control policy from a diverse set of unorganized motion sequences, which enables the generation of realistic human behaviors by sampling in the latent space and allows high-level control policies to reuse the learned skills to accomplish a variety of downstream tasks. In the training of ControlVAE, we employ a learnable world model to realize direct supervision of the latent space and the control policy. This world model effectively captures the unknown dynamics of the simulation system, enabling efficient model-based learning of high-level downstream tasks. We also learn a state-conditional prior distribution in the VAE-based generative control policy, which generates a skill embedding that outperforms the non-conditional priors in downstream tasks. We demonstrate the effectiveness of ControlVAE using a diverse set of tasks, which allows realistic and interactive control of the simulated characters. Heyuan Yao, Zhenhua Song, Baoquan Chen, Libin Liu 0002 |
ACM Trans. Graph. | 3 |
| 2022 | MDISN: Learning multiscale deformed implicit fields from single imagesabstractWe present a multiscale deformed implicit surface network (MDISN) to reconstruct 3D objects from single images by adapting the implicit surface of the target object from coarse to fine to the input image. The basic idea is to optimize the implicit surface according to the change of consecutive feature maps from the input image. And with multi-resolution feature maps, the implicit field is refined progressively, such that lower resolutions outline the main object components, and higher resolutions reveal fine-grained geometric details. To better explore the changes in feature maps, we devise a simple field deformation module that receives two consecutive feature maps to refine the implicit field with finer geometric details. Experimental results on both synthetic and real-world datasets demonstrate the superiority of the proposed method compared to state-of-the-art methods. Yixin Zhuang, Yunzhe Liu 0002, Baoquan Chen |
Vis. Informatics | 4 |
| 2021 | Robust Neural Routing Through Space Partitions for Camera Relocalization in Dynamic Indoor EnvironmentsabstractLocalizing the camera in a known indoor environment is a key building block for scene mapping, robot navigation, AR, etc. Recent advances estimate the camera pose via optimization over the 2D/3D-3D correspondences established between the coordinates in 2D/3D camera space and 3D world space. Such a mapping is estimated with either a convolution neural network or a decision tree using only the static input image sequence, which makes these approaches vulnerable to dynamic indoor environments that are quite common yet challenging in the real world. To address the aforementioned issues, in this paper, we propose a novel outlier-aware neural tree which bridges the two worlds, deep learning and decision tree approaches. It builds on three important blocks: (a) a hierarchical space partition over the indoor scene to construct the decision tree; (b) a neural routing function, implemented as a deep classification network, employed for better 3D scene understanding; and (c) an outlier rejection module used to filter out dynamic points during the hierarchical routing process. Our proposed algorithm is evaluated on the RIO-10 benchmark developed for camera relocalization in dynamic indoor environments. It achieves robust neural routing through space partitions and outperforms the state-of-the-art approaches by around 30% on camera pose accuracy, while running comparably fast for evaluation. Siyan Dong, Qingnan Fan, He Wang 0010, Li Yi 0001, Thomas A. Funkhouser, Baoquan Chen, Leonidas J. Guibas |
CVPR | 7 |
| 2021 | CAPTRA: CAtegory-level Pose Tracking for Rigid and Articulated Objects from Point CloudsabstractIn this work, we tackle the problem of category-level online pose tracking of objects from point cloud sequences. For the first time, we propose a unified framework that can handle 9DoF pose tracking for novel rigid object instances as well as per-part pose tracking for articulated objects from known categories. Here the 9DoF pose, comprising 6D pose and 3D size, is equivalent to a 3D amodal bounding box representation with free 6D pose. Given the depth point cloud at the current frame and the estimated pose from the last frame, our novel end-to-end pipeline learns to accurately update the pose. Our pipeline is composed of three modules: 1) a pose canonicalization module that normalizes the pose of the input depth point cloud; 2) RotationNet, a module that directly regresses small interframe delta rotations; and 3) CoordinateNet, a module that predicts the normalized coordinates and segmentation, enabling analytical computation of the 3D size and translation. Leveraging the small pose regime in the pose-canonicalized point clouds, our method integrates the best of both worlds by combining dense coordinate prediction and direct rotation regression, thus yielding an end-to-end differentiable pipeline optimized for 9DoF pose accuracy (without using non-differentiable RANSAC). Our extensive experiments demonstrate that our method achieves new state-of-the-art performance on category-level rigid object pose (NOCSREAL275 [29]) and articulated object pose benchmarks (SAPIEN [34], BMVC [18]) at the fastest FPS ∼ 12. Yijia Weng, He Wang 0010, Yuzhe Qin, Yueqi Duan, Qingnan Fan, Baoquan Chen, Hao Su 0001, Leonidas J. Guibas |
ICCV | 7 |
| 2021 | Unsupervised Co-part Segmentation through AssemblyabstractCo-part segmentation is an important problem in computer vision for its rich applications. We propose an unsupervised learning approach for co-part segmentation from images. For the training stage, we leverage motion information embedded in videos and explicitly extract latent representations to segment meaningful object parts. More importantly, we introduce a dual procedure of part-assembly to form a closed loop with part-segmentation, enabling an effective self-supervision. We demonstrate the effectiveness of our approach with a host of extensive experiments, ranging from human bodies, hands, quadruped, and robot arms. We show that our approach can achieve meaningful and compact part segmentation, outperforming state-of-the-art approaches on diverse benchmarks. Qingzhe Gao, Bin Wang 0021, Libin Liu 0002, Baoquan Chen |
ICML | 4 |
| 2021 | Mid-Air Finger Sketching for Tree Modelingabstract2D sketch-based tree modeling cannot guarantee to generate plausible depth values and full 3D tree shapes. With the advent of virtual reality (VR) technologies, 3D sketching enables a new form for 3D tree modeling. However, it is labor-intensive and difficult to create realistically-looking 3D trees with complicated geometry and lots of detailed twigs with a reasonable amount of effort. In this paper, we explore the use of mid-air finger 3D sketching in VR for tree modeling. We present a hybrid approach that integrates freehand 3D sketches with an automatic population of branch geometries. The user only needs to draw a few 3D strokes in mid-air to define the envelope of the foliage (denoted as lobes) and main branches. Our algorithm then automatically generates a full 3D tree model based on these stroke inputs. Additionally, the shape of the 3D tree model can be modified by freely dragging, squeezing, or moving lobes in mid-air. We demonstrate the ease-of-use, efficiency, and flexibility in tree modeling and overall shape control. We perform user studies and show a variety of realistic tree models generated instantaneously from 3D finger sketching. Fanxing Zhang, Zhanglin Cheng, Oliver Deussen, Baoquan Chen, Yunhai Wang |
VR | 5 |
| 2021 | Towards a Neural Graphics Pipeline for Controllable Image GenerationabstractAbstract In this paper, we leverage advances in neural networks towards forming a neural rendering for controllable image generation, and thereby bypassing the need for detailed modeling in conventional graphics pipeline. To this end, we present Neural Graphics Pipeline (NGP), a hybrid generative model that brings together neural and traditional image formation models. NGP decomposes the image into a set of interpretable appearance feature maps, uncovering direct control handles for controllable image generation. To form an image, NGP generates coarse 3D models that are fed into neural rendering modules to produce view‐specific interpretable 2D maps, which are then composited into the final output image using a traditional image formation model. Our approach offers control over image generation by providing direct handles controlling illumination and camera parameters, in addition to control over shape and appearance variations. The key challenge is to learn these controls through unsupervised training that links generated coarse 3D models with unpaired real images via neural and traditional (e.g., Blinn‐Phong) rendering functions, without establishing an explicit correspondence between them. We demonstrate the effectiveness of our approach on controllable image generation of single‐object scenes. We evaluate our hybrid modeling framework, compare with neural‐only generation methods (namely, DCGAN, LSGAN, WGAN‐GP, VON, and SRNs), report improvement in FID scores against real images, and demonstrate that NGP supports direct controls common in traditional forward rendering. Code is available at http://geometry.cs.ucl.ac.uk/projects/2021/ngp . Xuelin Chen, Daniel Cohen-Or, Baoquan Chen, Niloy J. Mitra |
Comput. Graph. Forum | 3 |
| 2021 | A General Decoupled Learning Framework for Parameterized Image OperatorsabstractMany different deep networks have been used to approximate, accelerate or improve traditional image operators. Among these traditional operators, many contain parameters which need to be tweaked to obtain the satisfactory results, which we refer to as "parameterized image operators". However, most existing deep networks trained for these operators are only designed for one specific parameter configuration, which does not meet the needs of real scenarios that usually require flexible parameters settings. To overcome this limitation, we propose a new decoupled learning algorithm to learn from the operator parameters to dynamically adjust the weights of a deep network for image operators, denoted as the base network. The learned algorithm is formed as another network, namely the weight learning network, which can be end-to-end jointly trained with the base network. Experiments demonstrate that the proposed framework can be successfully applied to many traditional parameterized image operators. To accelerate the parameter tuning for practical scenarios, the proposed framework can be further extended to dynamically change the weights of only one single layer of the base network while sharing most computation cost. We demonstrate that this cheap parameter-tuning extension of the proposed decoupled learning framework even outperforms the state-of-the-art alternative approaches. Qingnan Fan, Dongdong Chen 0001, Lu Yuan 0001, Gang Hua 0001, Nenghai Yu, Baoquan Chen |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2021 | Camera keyframing with style and controlabstractWe present a novel technique that enables 3D artists to synthesize camera motions in virtual environments following a camera style , while enforcing user-designed camera keyframes as constraints along the sequence. To solve this constrained motion in-betweening problem, we design and train a camera motion generator from a collection of temporal cinematic features (camera and actor motions) using a conditioning on target keyframes. We further condition the generator with a style code to control how to perform the interpolation between the keyframes. Style codes are generated by training a second network that encodes different camera behaviors in a compact latent space, the camera style space. Camera behaviors are defined as temporal correlations between actor features and camera motions and can be extracted from real or synthetic film clips. We further extend the system by incorporating a fine control of camera speed and direction via a hidden state mapping technique. We evaluate our method on two aspects: i) the capacity to synthesize style-aware camera trajectories with user defined keyframes; and ii) the capacity to ensure that in-between motions still comply with the reference camera style while satisfying the keyframe constraints. As a result, our system is the first style-aware keyframe in-betweening technique for camera control that balances style-driven automation with precise and interactive control of keyframes. Hongda Jiang, Marc Christie, Xi Wang 0024, Libin Liu 0002, Bin Wang 0021, Baoquan Chen |
ACM Trans. Graph. | 6 |
| 2021 | Learning skeletal articulations with neural blend shapesabstractAnimating a newly designed character using motion capture (mocap) data is a long standing problem in computer animation. A key consideration is the skeletal structure that should correspond to the available mocap data, and the shape deformation in the joint regions, which often requires a tailored, pose-specific refinement. In this work, we develop a neural technique for articulating 3D characters using enveloping with a pre-defined skeletal structure which produces high quality pose dependent deformations. Our framework learns to rig and skin characters with the same articulation structure ( e.g. , bipeds or quadrupeds), and builds the desired skeleton hierarchy into the network architecture. Furthermore , we propose neural blend shapes - a set of corrective pose-dependent shapes which improve the deformation quality in the joint regions in order to address the notorious artifacts resulting from standard rigging and skinning. Our system estimates neural blend shapes for input meshes with arbitrary connectivity, as well as weighting coefficients which are conditioned on the input joint rotations. Unlike recent deep learning techniques which supervise the network with ground-truth rigging and skinning parameters, our approach does not assume that the training data has a specific underlying deformation model. Instead, during training, the network observes deformed shapes and learns to infer the corresponding rig, skin and blend shapes using indirect supervision. During inference, we demonstrate that our network generalizes to unseen characters with arbitrary mesh connectivity, including unrigged characters built by 3D artists. Conforming to standard skeletal animation models enables direct plug-and-play in standard animation software, as well as game engines. Peizhuo Li, Kfir Aberman, Rana Hanocka, Libin Liu 0002, Olga Sorkine-Hornung, Baoquan Chen |
ACM Trans. Graph. | 6 |
| 2021 | Interactive cutting and tearing in projective dynamics with progressive cholesky updatesabstractWe propose a new algorithm for updating a Cholesky factorization which speeds up Projective Dynamics simulations with topological changes. Our approach addresses an important limitation of the original Projective Dynamics, i.e., that topological changes such as cutting, fracturing, or tearing require full refactorization which compromises computation speed, especially in real-time applications. Our method progressively modifies the Cholesky factor of the system matrix in the global step instead of computing it from scratch. Only a small amount of overhead is added since most of the topological changes in typical simulations are continuous and gradual. Our method is based on the update and downdate routine in CHOLMOD, but unlike recent related work, supports dynamic sizes of the system matrix and the addition of new vertices. Our approach allows us to introduce clean cuts and perform interactive remeshing. Our experiments show that our method works particularly well in simulation scenarios involving cutting, tearing, and local remeshing operations. Tiantian Liu 0002, Ladislav Kavan, Baoquan Chen |
ACM Trans. Graph. | 4 |
| 2021 | Solid-fluid interaction with surface-tension-dominant contactabstractWe propose a novel three-way coupling method to model the contact interaction between solid and fluid driven by strong surface tension. At the heart of our physical model is a thin liquid membrane that simultaneously couples to both the liquid volume and the rigid objects, facilitating accurate momentum transfer, collision processing, and surface tension calculation. This model is implemented numerically under a hybrid Eulerian-Lagrangian framework where the membrane is modelled as a simplicial mesh and the liquid volume is simulated on a background Cartesian grid. We devise a monolithic solver to solve the interactions among the three systems of liquid, solid, and membrane. We demonstrate the efficacy of our method through an array of rigid-fluid contact simulations dominated by strong surface tension, which enables the faithful modeling of a host of new surface-tension-dominant phenomena including: objects with higher density than water that remains afloat; 'Cheerios effect' where floating objects attract one another; and surface tension weakening effect caused by surface-active constituents. Liangwang Ruan, Bo Zhu 0002, Shinjiro Sueda, Bin Wang 0069, Baoquan Chen |
ACM Trans. Graph. | 6 |
| 2021 | MotioNet: 3D Human Motion Reconstruction from Monocular Video with Skeleton ConsistencyabstractWe introduce MotioNet , a deep neural network that directly reconstructs the motion of a 3D human skeleton from a monocular video. While previous methods rely on either rigging or inverse kinematics (IK) to associate a consistent skeleton with temporally coherent joint rotations, our method is the first data-driven approach that directly outputs a kinematic skeleton, which is a complete, commonly used motion representation. At the crux of our approach lies a deep neural network with embedded kinematic priors, which decomposes sequences of 2D joint positions into two separate attributes: a single, symmetric skeleton encoded by bone lengths, and a sequence of 3D joint rotations associated with global root positions and foot contact labels. These attributes are fed into an integrated forward kinematics (FK) layer that outputs 3D positions, which are compared to a ground truth. In addition, an adversarial loss is applied to the velocities of the recovered rotations to ensure that they lie on the manifold of natural joint rotations. The key advantage of our approach is that it learns to infer natural joint rotations directly from the training data rather than assuming an underlying model, or inferring them from joint positions using a data-agnostic IK solver. We show that enforcing a single consistent skeleton along with temporally coherent joint rotations constrains the solution space, leading to a more robust handling of self-occlusions and depth ambiguities. Mingyi Shi, Kfir Aberman, Andreas Aristidou, Taku Komura, Dani Lischinski, Daniel Cohen-Or, Baoquan Chen |
ACM Trans. Graph. | 7 |
| 2021 | A material point method for nonlinearly magnetized materialsabstractWe propose a novel numerical scheme to simulate interactions between a magnetic field and nonlinearly magnetized objects immersed in it. Under our nonlinear magnetization framework, the strength of magnetic forces is effectively saturated to produce stable simulations without requiring any parameter tuning. The mathematical model of our approach is based upon Langevin's nonlinear theory of paramagnetism, which bridges microscopic structures and macroscopic equations after a statistical derivation. We devise a hybrid Eulerian-Lagrangian numerical approach to simulating this strongly nonlinear process by leveraging the discrete material points to transfer both material properties and the number density of magnetic micro-particles in the simulation domain. The magnetic equations can then be built and solved efficiently on a background Cartesian grid, followed by a finite difference method to incorporate magnetic forces. The multi-scale coupling can be processed naturally by employing the established particle-grid interpolation schemes in a conventional MLS-MPM framework. We demonstrate the efficacy of our approach with a host of simulation examples governed by magnetic-mechanical coupling effects, ranging from magnetic deformable bodies to magnetic viscous fluids with nonlinear elastic constitutive laws. Yuchen Sun 0002, Xingyu Ni, Bo Zhu 0002, Bin Wang 0069, Baoquan Chen |
ACM Trans. Graph. | 5 |
| 2021 | Implicit Multidimensional Projection of Local SubspacesabstractWe propose a visualization method to understand the effect of multidimensional projection on local subspaces, using implicit function differentiation. Here, we understand the local subspace as the multidimensional local neighborhood of data points. Existing methods focus on the projection of multidimensional data points, and the neighborhood information is ignored. Our method is able to analyze the shape and directional information of the local subspace to gain more insights into the global structure of the data through the perception of local structures. Local subspaces are fitted by multidimensional ellipses that are spanned by basis vectors. An accurate and efficient vector transformation method is proposed based on analytical differentiation of multidimensional projections formulated as implicit functions. The results are visualized as glyphs and analyzed using a full set of specifically-designed interactions supported in our efficient web-based visualization tool. The usefulness of our method is demonstrated using various multi- and high-dimensional benchmark datasets. Our implicit differentiation vector transformation is evaluated through numerical comparisons; the overall method is evaluated through exploration examples and use cases. Rongzheng Bian, Yumeng Xue, Liang Zhou 0001, Jian Zhang 0070, Baoquan Chen, Daniel Weiskopf, Yunhai Wang |
IEEE Trans. Vis. Comput. Graph. | 5 |
| 2020 | AutoRemover: Automatic Object Removal for Autonomous Driving VideosabstractMotivated by the need for photo-realistic simulation in autonomous driving, in this paper we present a video inpainting algorithm AutoRemover, designed specifically for generating street-view videos without any moving objects. In our setup we have two challenges: the first is the shadow, shadows are usually unlabeled but tightly coupled with the moving objects. The second is the large ego-motion in the videos. To deal with shadows, we build up an autonomous driving shadow dataset and design a deep neural network to detect shadows automatically. To deal with large ego-motion, we take advantage of the multi-source data, in particular the 3D data, in autonomous driving. More specifically, the geometric relationship between frames is incorporated into an inpainting deep neural network to produce high-quality structurally consistent video output. Experiments show that our method outperforms other state-of-the-art (SOTA) object removal algorithms, reducing the RMSE by over 19%. Wei Li 0111, Peng Wang 0001, Chenye Guan, Yuhang Song 0003, Baoquan Chen, Weiwei Xu 0003, Ruigang Yang |
AAAI | 8 |
| 2020 | PQ-NET: A Generative Part Seq2Seq Network for 3D ShapesabstractWe introduce PQ-NET, a deep neural network which represents and generates 3D shapes via sequential part assembly. The input to our network is a 3D shape segmented into parts, where each part is first encoded into a feature representation using a part autoencoder. The core component of PQ-NET is a sequence-to-sequence or Seq2Seq autoencoder which encodes a sequence of part features into a latent vector of fixed size, and the decoder reconstructs the 3D shape, one part at a time, resulting in a sequential assembly. The latent space formed by the Seq2Seq encoder encodes both part structure and fine part geometry. The decoder can be adapted to perform several generative tasks including shape autoencoding, interpolation, novel shape generation, and single-view 3D reconstruction, where the generated shapes are all composed of meaningful parts. Rundi Wu, Yixin Zhuang, Kai Xu 0004, Hao (Richard) Zhang, Baoquan Chen |
CVPR | 5 |
| 2020 | Multimodal Shape Completion via Conditional Generative Adversarial Networks
Rundi Wu, Xuelin Chen, Yixin Zhuang, Baoquan Chen |
ECCV (4) | 4 |
| 2020 | Unpaired Point Cloud Completion on Real Scans using Adversarial Training
Xuelin Chen, Baoquan Chen, Niloy J. Mitra |
ICLR | 2 |
| 2020 | Generative 3D Part Assembly via Dynamic Graph LearningabstractAutonomous part assembly is a challenging yet crucial task in 3D computer vision and robotics. Analogous to buying an IKEA furniture, given a set of 3D parts that can assemble a single shape, an intelligent agent needs to perceive the 3D part geometry, reason to propose pose estimations for the input parts, and finally call robotic planning and control routines for actuation. In this paper, we focus on the pose estimation subproblem from the vision side involving geometric and relational reasoning over the input part geometry. Essentially, the task of generative 3D part assembly is to predict a 6-DoF part pose, including a rigid rotation and translation, for each input part that assembles a single 3D shape as the final output. To tackle this problem, we propose an assembly-oriented dynamic graph learning framework that leverages an iterative graph neural network as a backbone. It explicitly conducts sequential part assembly refinements in a coarse-to-fine manner, exploits a pair of part relation reasoning module and part aggregation module for dynamically adjusting both part features and their relations in the part graph. We conduct extensive experiments and quantitative comparisons to three strong baseline methods, demonstrating the effectiveness of the proposed approach. Guanqi Zhan, Qingnan Fan, Kaichun Mo, Lin Shao 0002, Baoquan Chen, Leonidas J. Guibas, Hao Dong 0003 |
NeurIPS | 5 |
| 2020 | Fabricable Unobtrusive 3D-QR-Codes with Directional LightabstractAbstract QR code is a 2D matrix barcode widely used for product tracking, identification, document management and general marketing. Recently, there have been various attempts to utilize QR codes in 3D manufacturing by carving QR codes on the surface of the printed 3D shape. Nevertheless, significant shape editing and modulation may be required to allow readability of the embedded 3D‐QR‐codes with good decoding accuracy. In this paper, we introduce a novel QR code 3D fabrication framework aimed at unobtrusive embedding of 3D‐QR‐codes in the shape hence introducing minimal shape modulation. Essentially, our method computes bi‐directional carvings in the 3D shape surface to obtain the black‐and‐white QR pattern. By using a directional light source, the black‐and‐white QR pattern emerges as lighted and shadow casted blocks on the shape respectively. To account for minimal modulation and elusiveness, we optimize the QR code carving w.r.t. shape geometry, visual disparity and light source position. Our technique employs a simulation of lighting phenomena through carved modules on the shape to ensure adequate contrast of the printed 3D‐QR‐code. Hao Peng 0001, Peiqing Liu, Lin Lu 0001, Andrei Sharf, Dani Lischinski, Baoquan Chen |
Comput. Graph. Forum | 7 |
| 2020 | PointSkelCNN: Deep Learning-Based 3D Human Skeleton Extraction from Point CloudsabstractAbstract A 3D human skeleton plays important roles in human shape reconstruction and human animation. Remarkable advances have been achieved recently in 3D human skeleton estimation from color and depth images via a powerful deep convolutional neural network. However, applying deep learning frameworks to 3D human skeleton extraction from point clouds remains challenging because of the sparsity of point clouds and the high nonlinearity of human skeleton regression. In this study, we develop a deep learning‐based approach for 3D human skeleton extraction from point clouds. We convert 3D human skeleton extraction into offset vector regression and human body segmentation via deep learning‐based point cloud contraction. Furthermore, a disambiguation strategy is adopted to improve the robustness of joint points regression. Experiments on the public human pose dataset UBC3V and the human point cloud skeleton dataset 3DHumanSkeleton compiled by the authors show that the proposed approach outperforms the state‐of‐the‐art methods. Hongxing Qin, Songshan Zhang, Qihuang Liu, Baoquan Chen |
Comput. Graph. Forum | 5 |
| 2020 | Learning Elastic Constitutive Material and Damping ModelsabstractAbstract Commonly used linear and nonlinear constitutive material models in deformation simulation contain many simplifications and only cover a tiny part of possible material behavior. In this work we propose a framework for learning customized models of deformable materials from example surface trajectories. The key idea is to iteratively improve a correction to a nominal model of the elastic and damping properties of the object, which allows new forward simulations with the learned correction to more accurately predict the behavior of a given soft object. Space‐time optimization is employed to identify gentle control forces with which we extract necessary data for model inference and to finally encapsulate the material correction into a compact parametric form. Furthermore, a patch based position constraint is proposed to tackle the challenge of handling incomplete and noisy observations arising in real‐world examples. We demonstrate the effectiveness of our method with a set of synthetic examples, as well with data captured from real world homogeneous elastic objects. Bin Wang 0021, Yuanmin Deng, Paul G. Kry, Uri M. Ascher, Hui Huang 0004, Baoquan Chen |
Comput. Graph. Forum | 6 |
| 2020 | DeepPipes: Learning 3D pipelines reconstruction from point clouds
Lili Cheng, Zhuo Wei, Mingchao Sun, Shi-Qing Xin, Andrei Sharf, Yangyan Li, Baoquan Chen, Changhe Tu |
Graph. Model. | 7 |
| 2020 | Super Diffusion for Salient Object DetectionabstractOne major branch of saliency object detection methods are diffusion-based which construct a graph model on a given image and diffuse seed saliency values to the whole graph by a diffusion matrix. While their performance is sensitive to specific feature spaces and scales used for the diffusion matrix definition, little work has been published to systematically promote the robustness and accuracy of salient object detection under the generic mechanism of diffusion. In this work, we firstly present a novel view of the working mechanism of the diffusion process based on mathematical analysis, which reveals that the diffusion process is actually computing the similarity of nodes with respect to the seeds based on diffusion maps. Following this analysis, we propose super diffusion, a novel inclusive learning-based framework for salient object detection, which makes the optimum and robust performance by integrating a large pool of feature spaces, scales and even features originally computed for non-diffusion-based salient object detection. A closed-form solution of the optimal parameters for the integration is determined through supervised learning. At the local level, we propose to promote each individual diffusion before the integration. Our mathematical analysis reveals the close relationship between saliency diffusion and spectral clustering. Based on this, we propose to re-synthesize each individual diffusion matrix from the most discriminative eigenvectors and the constant eigenvector (for saliency normalization). The proposed framework is implemented and experimented on prevalently used benchmark datasets, consistently leading to state-of-the-art performance. Peng Jiang 0002, Zhiyi Pan 0001, Changhe Tu, Nuno Vasconcelos, Baoquan Chen, Jingliang Peng |
IEEE Trans. Image Process. | 5 |
| 2020 | Iterative Local-Global Collaboration Learning Towards One-Shot Video Person Re-IdentificationabstractVideo person re-identification (video Re-ID) plays an important role in surveillance video analysis and has gained increasing attention recently. However, existing supervised methods require vast labeled identities across cameras, resulting in poor scalability in practical applications. Although some unsupervised approaches have been exploited for video Re-ID, they are still in their infancy due to the complex nature of learning discriminative features on unlabelled data. In this paper, we focus on one-shot video Re-ID and present an iterative local-global collaboration learning approach to learning robust and discriminative person representations. Specifically, it jointly considers the global video information and local frame sequence information to better capture the diverse appearance of the person for feature learning and pseudo-label estimation. Moreover, as the cross-entropy loss may induce the model to focus on identity-irrelevant factors, we introduce the variational information bottleneck as a regularization term to train the model together. It can help filter undesirable information and characterize subtle differences among persons. Since accuracy cannot always be guaranteed for pseudo-labels, we adopt a dynamic selection strategy to select part of pseudo-labeled data with higher confidence to update the training set and re-train the learning model. During training, our method iteratively executes the feature learning, pseudo-label estimation, and dynamic sample selection until all the unlabeled data have been seen. Extensive experiments on two public datasets, i.e., DukeMTMC-VideoReID and MARS, have verified the superiority of our model to several cutting-edge competitors. Meng Liu 0006, Leigang Qu, Liqiang Nie, Maofu Liu, Ling-Yu Duan, Baoquan Chen |
IEEE Trans. Image Process. | 6 |
| 2020 | Neural Multimodal Cooperative Learning Toward Micro-Video UnderstandingabstractThe prevailing characteristics of micro-videos result in the less descriptive power of each modality. The micro-video representations, several pioneer efforts proposed, are limited in implicitly exploring the consistency between different modality information but ignore the complementarity. In this paper, we focus on how to explicitly separate the consistent features and the complementary features from the mixed information and harness their combination to improve the expressiveness of each modality. Toward this end, we present a neural multimodal cooperative learning (NMCL) model to split the consistent component and the complementary component by a novel relation-aware attention mechanism. Specifically, the computed attention score can be used to measure the correlation between the features extracted from different modalities. Then, a threshold is learned for each modality to distinguish the consistent and complementary features according to the score. Thereafter, we integrate the consistent parts to enhance the representations and supplement the complementary ones to reinforce the information in each modality. As to the problem of redundant information, which may cause overfitting and is hard to distinguish, we devise an attention network to dynamically capture the features which closely related the category and output a discriminative representation for prediction. The experimental results on a real-world micro-video dataset show that the NMCL outperforms the state-of-the-art methods. Further studies verify the effectiveness and cooperative effects brought by the attentive mechanism. Yinwei Wei, Xiang Wang 0010, Weili Guan, Liqiang Nie, Zhouchen Lin, Baoquan Chen |
IEEE Trans. Image Process. | 6 |
| 2020 | Skeleton-aware networks for deep motion retargetingabstractWe introduce a novel deep learning framework for data-driven motion retargeting between skeletons, which may have different structure, yet corresponding to homeomorphic graphs. Importantly, our approach learns how to retarget without requiring any explicit pairing between the motions in the training set. We leverage the fact that different homeomorphic skeletons may be reduced to a common primal skeleton by a sequence of edge merging operations, which we refer to as skeletal pooling. Thus, our main technical contribution is the introduction of novel differentiable convolution, pooling, and unpooling operators. These operators are skeleton-aware , meaning that they explicitly account for the skeleton's hierarchical structure and joint adjacency, and together they serve to transform the original motion into a collection of deep temporal features associated with the joints of the primal skeleton. In other words, our operators form the building blocks of a new deep motion processing framework that embeds the motion into a common latent space, shared by a collection of homeomorphic skeletons. Thus, retargeting can be achieved simply by encoding to, and decoding from this latent space. Our experiments show the effectiveness of our framework for motion retargeting, as well as motion processing in general, compared to existing approaches. Our approach is also quantitatively evaluated on a synthetic dataset that contains pairs of motions applied to different skeletons. To the best of our knowledge, our method is the first to perform retargeting between skeletons with differently sampled kinematic chains, without any paired examples. Kfir Aberman, Peizhuo Li, Dani Lischinski, Olga Sorkine-Hornung, Daniel Cohen-Or, Baoquan Chen |
ACM Trans. Graph. | 6 |
| 2020 | Unpaired motion style transfer from video to animationabstractTransferring the motion style from one animation clip to another, while preserving the motion content of the latter, has been a long-standing problem in character animation. Most existing data-driven approaches are supervised and rely on paired data, where motions with the same content are performed in different styles. In addition, these approaches are limited to transfer of styles that were seen during training. In this paper, we present a novel data-driven framework for motion style transfer, which learns from an unpaired collection of motions with style labels, and enables transferring motion styles not observed during training. Furthermore, our framework is able to extract motion styles directly from videos, bypassing 3D reconstruction, and apply them to the 3D input motion. Our style transfer network encodes motions into two latent codes, for content and for style, each of which plays a different role in the decoding (synthesis) process. While the content code is decoded into the output motion by several temporal convolutional layers, the style code modifies deep features via temporally invariant adaptive instance normalization (AdaIN). Moreover, while the content code is encoded from 3D joint rotations, we learn a common embedding for style from either 3D or 2D joint positions, enabling style extraction from videos. Our results are comparable to the state-of-the-art, despite not requiring paired training data, and outperform other methods when transferring previously unseen styles. To our knowledge, we are the first to demonstrate style transfer directly from videos to 3D animations - an ability which enables one to extend the set of style examples far beyond motions captured by MoCap systems. Kfir Aberman, Yijia Weng, Dani Lischinski, Daniel Cohen-Or, Baoquan Chen |
ACM Trans. Graph. | 5 |
| 2020 | Example-driven virtual cinematography by learning camera behaviorsabstractDesigning a camera motion controller that has the capacity to move a virtual camera automatically in relation with contents of a 3D animation, in a cinematographic and principled way, is a complex and challenging task. Many cinematographic rules exist, yet practice shows there are significant stylistic variations in how these can be applied. In this paper, we propose an example-driven camera controller which can extract camera behaviors from an example film clip and re-apply the extracted behaviors to a 3D animation, through learning from a collection of camera motions. Our first technical contribution is the design of a low-dimensional cinematic feature space that captures the essence of a film's cinematic characteristics (camera angle and distance, screen composition and character configurations) and which is coupled with a neural network to automatically extract these cinematic characteristics from real film clips. Our second technical contribution is the design of a cascaded deep-learning architecture trained to (i) recognize a variety of camera motion behaviors from the extracted cinematic features, and (ii) predict the future motion of a virtual camera given a character 3D animation. We propose to rely on a Mixture of Experts (MoE) gating+prediction mechanism to ensure that distinct camera behaviors can be learned while ensuring generalization. We demonstrate the features of our approach through experiments that highlight (i) the quality of our cinematic feature extractor (ii) the capacity to learn a range of behaviors through the gating mechanism, and (iii) the ability to generate a variety of camera motions by applying different behaviors extracted from film clips. Such an example-driven approach offers a high level of controllability which opens new possibilities toward a deeper understanding of cinematographic style and enhanced possibilities in exploiting real film data in virtual environments. Hongda Jiang, Bin Wang 0069, Marc Christie, Baoquan Chen |
ACM Trans. Graph. | 5 |
| 2020 | A level-set method for magnetic substance simulationabstractWe present a versatile numerical approach to simulating various magnetic phenomena using a level-set method. At the heart of our method lies a novel two-way coupling mechanism between a magnetic field and a magnetizable mechanical system, which is based on the interfacial Helmholtz force drawn from the Minkowski form of the Maxwell stress tensor. We show that a magnetic-mechanical coupling system can be solved as an interfacial problem, both theoretically and computationally. In particular, we employ a Poisson equation with a jump condition across the interface to model the mechanical-to-magnetic interaction and a Helmholtz force on the free surface to model the magnetic-to-mechanical effects. Our computational framework can be easily integrated into a standard Euler fluid solver, enabling both simulation and visualization of a complex magnetic field and its interaction with immersed magnetizable objects in a large domain. We demonstrate the efficacy of our method through an array of magnetic substance simulations that exhibit rich geometric and dynamic characteristics, encompassing ferrofluid, rigid magnetic body, deformable magnetic body, and multi-phase couplings. Xingyu Ni, Bo Zhu 0002, Bin Wang 0069, Baoquan Chen |
ACM Trans. Graph. | 4 |
| 2020 | A Recursive Subdivision Technique for Sampling Multi-class ScatterplotsabstractWe present a non-uniform recursive sampling technique for multi-class scatterplots, with the specific goal of faithfully presenting relative data and class densities, while preserving major outliers in the plots. Our technique is based on a customized binary kd-tree, in which leaf nodes are created by recursively subdividing the underlying multi-class density map. By backtracking, we merge leaf nodes until they encompass points of all classes for our subsequently applied outlier-aware multi-class sampling strategy. A quantitative evaluation shows that our approach can better preserve outliers and at the same time relative densities in multi-class scatterplots compared to the previous approaches, several case studies demonstrate the effectiveness of our approach in exploring complex and real world data. Xin Chen 0075, Tong Ge, Jian Zhang 0070, Baoquan Chen, Chi-Wing Fu, Oliver Deussen, Yunhai Wang |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2020 | Mass-Driven Topology-Aware Curve Skeleton Extraction from Incomplete Point CloudsabstractWe introduce a mass-driven curve skeleton as a curve skeleton representation for 3D point cloud data. The mass-driven curve skeleton presents geometric properties and mass distribution of a curve skeleton simultaneously. The computation of the mass-driven curve skeleton is formulated as a minimization of Wasserstein distance, with an entropic regularization term, between mass distributions of point clouds and curve skeletons. Assuming that the mass of one sampling point should be transported to a line-like structure, a topology-aware rough curve skeleton is extracted via the optimal transport plan. A Dirichlet energy regularization term is then used to obtain a smooth curve skeleton via geometric optimization. Given that rough curve skeleton extraction does not depend on complete point clouds, our algorithm can be directly applied to curve skeleton extraction from incomplete point clouds. We demonstrate that a mass-driven curve skeleton can be directly applied to an unoriented raw point scan with significant noise, outliers and large areas of missing data. In comparison with state-of-the-art methods on curve skeleton extraction, the performance of the proposed mass-driven curve skeleton is more robust in terms of extracting a correct topology. Hongxing Qin, Hui Huang 0004, Baoquan Chen |
IEEE Trans. Vis. Comput. Graph. | 5 |
| 2020 | Interactive Structure-aware Blending of Diverse Edge Bundling VisualizationsabstractMany edge bundling techniques (i.e., data simplification as a support for data visualization and decision making) exist but they are not directly applicable to any kind of dataset and their parameters are often too abstract and difficult to set up. As a result, this hinders the user ability to create efficient aggregated visualizations. To address these issues, we investigated a novel way of handling visual aggregation with a task-driven and user-centered approach. Given a graph, our approach produces a decluttered view as follows: first, the user investigates different edge bundling results and specifies areas, where certain edge bundling techniques would provide user-desired results. Second, our system then computes a smooth and structural preserving transition between these specified areas. Lastly, the user can further fine-tune the global visualization with a direct manipulation technique to remove the local ambiguity and to apply different visual deformations. In this paper, we provide details for our design rationale and implementation. Also, we show how our algorithm gives more suitable results compared to current edge bundling techniques, and in the end, we provide concrete instances of usages, where the algorithm combines various edge bundling results to support diverse data exploration and visualizations. Yunhai Wang, Mingliang Xue, Xinyuan Yan, Baoquan Chen, Chi-Wing Fu, Christophe Hurter |
IEEE Trans. Vis. Comput. Graph. | 5 |
| 2020 | Strong 3D Printing by TPMS Injectionabstract3D printed objects are rapidly becoming prevalent in science, technology and daily life. An important question is how to obtain strong and durable 3D models using standard printing techniques. This question is often translated to computing smartly designed interior structures that provide strong support and yield resistant 3D models. In this paper we suggest a combination between 3D printing and material injection to achieve strong 3D printed objects. We utilize triply periodic minimal surfaces (TPMS) to define novel interior support structures. TPMS are closed form and can be computed in a simple and straightforward manner. Since TPMS are smooth and connected, we utilize them to define channels that adequately distribute injected materials in the shape interior. To account for weak regions, TPMS channels are locally optimized according to the shape stress field. After the object is printed, we simply inject the TPMS channels with materials that solidify and yield a strong inner structure that supports the shape. Our method allows injecting a wide range of materials in an object interior in a fast and easy manner. Results demonstrate the efficiency of strong printing by combining 3D printing and injection together. Cong Rao, Lin Lu 0001, Andrei Sharf, Haisen Zhao, Baoquan Chen |
IEEE Trans. Vis. Comput. Graph. | 6 |
| 2019 | Fabricating QR codes on 3D objects using self-shadows
Hao Peng 0001, Lin Lu 0001, Andrei Sharf, Baoquan Chen |
Comput. Aided Des. | 5 |
| 2019 | Deep Video-Based Performance CloningabstractAbstract We present a new video‐based performance cloning technique. After training a deep generative network using a reference video capturing the appearance and dynamics of a target actor, we are able to generate videos where this actor reenacts other performances. All of the training data and the driving performances are provided as ordinary video segments, without motion capture or depth information. Our generative model is realized as a deep neural network with two branches, both of which train the same space‐time conditional generator, using shared weights. One branch, responsible for learning to generate the appearance of the target actor in various poses, uses paired training data, self‐generated from the reference video. The second branch uses unpaired data to improve generation of temporally coherent video renditions of unseen pose sequences. Through data augmentation, our network is able to synthesize images of the target actor in poses never captured by the reference video. We demonstrate a variety of promising results, where our method is able to generate temporally coherent videos, for challenging scenarios where the reference and driving videos consist of very different dance performances. Kfir Aberman, Mingyi Shi, Jing Liao 0001, Dani Lischinski, Baoquan Chen, Daniel Cohen-Or |
Comput. Graph. Forum | 5 |
| 2019 | Online Data Organizer: Micro-Video Categorization by Structure-Guided Multimodal Dictionary LearningabstractMicro-videos have rapidly become one of the most dominant trends in the era of social media. Accordingly, how to organize them draws our attention. Distinct from the traditional long videos that would have multi-site scenes and tolerate the hysteresis, a micro-video: 1) usually records contents at one specific venue within a few seconds. The venues are structured hierarchically regarding their category granularity. This motivates us to organize the micro-videos via their venue structure. 2) timely circulates over social networks. Thus, the timeliness of micro-videos desires effective online processing. However, only 1.22% of micro-videos are labeled with venue information when uploaded at the mobile end. To address this problem, we present a framework to organize the micro-videos online. In particular, we first build a structure-guided multi-modal dictionary learning model to learn the concept-level micro-video representation by jointly considering their venue structure and modality relatedness. We then develop an online learning algorithm to incrementally and efficiently strengthen our model, as well as categorize the micro-videos into a tree structure. Extensive experiments on a real-world data set validate our model well. In addition, we have released the codes to facilitate the research in the community. Meng Liu 0006, Liqiang Nie, Xiang Wang 0010, Qi Tian 0001, Baoquan Chen |
IEEE Trans. Image Process. | 5 |
| 2019 | Learning character-agnostic motion for motion retargeting in 2DabstractAnalyzing human motion is a challenging task with a wide variety of applications in computer vision and in graphics. One such application, of particular importance in computer animation, is the retargeting of motion from one performer to another. While humans move in three dimensions, the vast majority of human motions are captured using video, requiring 2D-to-3D pose and camera recovery, before existing retargeting approaches may be applied. In this paper, we present a new method for retargeting video-captured motion between different human performers, without the need to explicitly reconstruct 3D poses and/or camera parameters. In order to achieve our goal, we learn to extract, directly from a video, a high-level latent motion representation, which is invariant to the skeleton geometry and the camera view. Our key idea is to train a deep neural network to decompose temporal sequences of 2D poses into three components: motion, skeleton, and camera view-angle. Having extracted such a representation, we are able to re-combine motion with novel skeletons and camera views, and decode a retargeted temporal sequence, which we compare to a ground truth from a synthetic dataset. We demonstrate that our framework can be used to robustly extract human motion from videos, bypassing 3D reconstruction, and outperforming existing retargeting methods, when applied to videos in-the-wild. It also enables additional applications, such as performance cloning, video-driven cartoons, and motion retrieval. Kfir Aberman, Rundi Wu, Dani Lischinski, Baoquan Chen, Daniel Cohen-Or |
ACM Trans. Graph. | 4 |
| 2019 | Multi-robot collaborative dense scene reconstructionabstractWe present an autonomous scanning approach which allows multiple robots to perform collaborative scanning for dense 3D reconstruction of unknown indoor scenes. Our method plans scanning paths for several robots, allowing them to efficiently coordinate with each other such that the collective scanning coverage and reconstruction quality is maximized while the overall scanning effort is minimized. To this end, we define the problem as a dynamic task assignment and introduce a novel formulation based on Optimal Mass Transport (OMT). Given the currently scanned scene, a set of task views are extracted to cover scene regions which are either unknown or uncertain. These task views are assigned to the robots based on the OMT optimization. We then compute for each robot a smooth path over its assigned tasks by solving an approximate traveling salesman problem. In order to showcase our algorithm, we implement a multi-robot auto-scanning system. Since our method is computationally efficient, we can easily run it in real time on commodity hardware, and combine it with online RGB-D reconstruction approaches. In our results, we show several real-world examples of large indoor environments; in addition, we build a benchmark with a series of carefully designed metrics for quantitatively evaluating multi-robot autoscanning. Overall, we are able to demonstrate high-quality scanning results with respect to reconstruction quality and scanning efficiency, which significantly outperforms existing multi-robot exploration systems. Siyan Dong, Kai Xu 0004, Andrea Tagliasacchi, Shi-Qing Xin, Matthias Nießner, Baoquan Chen |
ACM Trans. Graph. | 7 |
| 2019 | GRAINS: Generative Recursive Autoencoders for INdoor ScenesabstractWe present a generative neural network that enables us to generate plausible 3D indoor scenes in large quantities and varieties, easily and highly efficiently. Our key observation is that indoor scene structures are inherently hierarchical . Hence, our network is not convolutional; it is a recursive neural network, or RvNN. Using a dataset of annotated scene hierarchies, we train a variational recursive autoencoder , or RvNN-VAE, which performs scene object grouping during its encoding phase and scene generation during decoding. Specifically, a set of encoders are recursively applied to group 3D objects based on support, surround, and co-occurrence relations in a scene, encoding information about objects’ spatial properties, semantics , and relative positioning with respect to other objects in the hierarchy. By training a variational autoencoder (VAE), the resulting fixed-length codes roughly follow a Gaussian distribution. A novel 3D scene can be generated hierarchically by the decoder from a randomly sampled code from the learned distribution. We coin our method GRAINS, for Generative Recursive Autoencoders for INdoor Scenes. We demonstrate the capability of GRAINS to generate plausible and diverse 3D indoor scenes and compare with existing methods for 3D scene synthesis. We show applications of GRAINS including 3D scene modeling from 2D layouts, scene editing, and semantic scene segmentation via PointNet whose performance is boosted by the large quantity and variety of 3D scenes generated by our method. Manyi Li, Akshay Gadi Patil, Kai Xu 0004, Siddhartha Chaudhuri, Owais Khan, Ariel Shamir, Changhe Tu, Baoquan Chen, Daniel Cohen-Or, Hao (Richard) Zhang |
ACM Trans. Graph. | 8 |
| 2019 | Efficient and conservative fluids using bidirectional mappingabstractIn this paper, we introduce BiMocq 2 , an unconditionally stable, pure Eulerianbased advection scheme to efficiently preserve the advection accuracy of all physical quantities for long-term fluid simulations. Our approach is built upon the method of characteristic mapping (MCM). Instead of the costly evaluation of the temporal characteristic integral, we evolve the mapping function itself by solving an advection equation for the mappings. Dual mesh characteristics (DMC) method is adopted to more accurately update the mapping. Furthermore, to avoid visual artifacts like instant blur and temporal inconsistency introduced by re-initialization, we introduce multi-level mapping and back and forth error compensation. We conduct comprehensive 2D and 3D benchmark experiments to compare against alternative advection schemes. In particular, for the vortical flow and level set experiments, our method outperforms almost all state-of-art hybrid schemes, including FLIP, PolyPic and Particle-Level-Set, at the cost of only two Semi-Lagrangian advections. Additionally, our method does not rely on the particle-grid transfer operations, leading to a highly parallelizable pipeline. As a result, more than 45× performance acceleration can be achieved via even a straightforward porting of the code from CPU to GPU. Ziyin Qu, Ming Gao 0023, Chenfanfu Jiang, Baoquan Chen |
ACM Trans. Graph. | 5 |
| 2019 | Optimizing Color Assignment for Perception of Class Separability in Multiclass ScatterplotsabstractAppropriate choice of colors significantly aids viewers in understanding the structures in multiclass scatterplots and becomes more important with a growing number of data points and groups. An appropriate color mapping is also an important parameter for the creation of an aesthetically pleasing scatterplot. Currently, users of visualization software routinely rely on color mappings that have been pre-defined by the software. A default color mapping, however, cannot ensure an optimal perceptual separability between groups, and sometimes may even lead to a misinterpretation of the data. In this paper, we present an effective approach for color assignment based on a set of given colors that is designed to optimize the perception of scatterplots. Our approach takes into account the spatial relationships, density, degree of overlap between point clusters, and also the background color. For this purpose, we use a genetic algorithm that is able to efficiently find good color assignments. We implemented an interactive color assignment system with three extensions of the basic method that incorporates top K suggestions, user-defined color subsets, and classes of interest for the optimization. To demonstrate the effectiveness of our assignment technique, we conducted a numerical study and a controlled user study to compare our approach with default color assignments; our findings were verified by two expert studies. The results show that our approach is able to support users in distinguishing cluster numbers faster and more precisely than default assignment methods. Yunhai Wang, Xin Chen 0075, Tong Ge, Chen Bao, Michael Sedlmair, Chi-Wing Fu, Oliver Deussen, Baoquan Chen |
IEEE Trans. Vis. Comput. Graph. | 8 |
| 2019 | Structure-aware Fisheye Views for Efficient Large Graph ExplorationabstractTraditional fisheye views for exploring large graphs introduce substantial distortions that often lead to a decreased readability of paths and other interesting structures. To overcome these problems, we propose a framework for structure-aware fisheye views. Using edge orientations as constraints for graph layout optimization allows us not only to reduce spatial and temporal distortions during fisheye zooms, but also to improve the readability of the graph structure. Furthermore, the framework enables us to optimize fisheye lenses towards specific tasks and design a family of new lenses: polyfocal, cluster, and path lenses. A GPU implementation lets us process large graphs with up to 15,000 nodes at interactive rates. A comprehensive evaluation, a user study, and two case studies demonstrate that our structure-aware fisheye views improve layout readability and user performance. Yunhai Wang, Yinqi Sun, Chi-Wing Fu, Michael Sedlmair, Baoquan Chen, Oliver Deussen |
IEEE Trans. Vis. Comput. Graph. | 7 |
| 2019 | Quasi-holography computational model for urban computingabstractVast amounts of data are produced with the development of smart cities and urban computing technologies. The data is often captured from multiple sensors, with heterogeneous structures and highly decentralized connections. Integrated data representation and smart computational models are required for more complex tasks in urban computing. We dwell deeply on two fundamental questions — can we provide an integrated data representation for the whole cyber–physical–social system? And, can we provide an integrated framework to choose the appropriate data for understanding a specific urban event? A holography data representation and the quasi-holography computational model have been proposed to address these problems. We describe case studies using the quasi-holography computational model, and discuss further problems to solve regarding our model. Baoquan Chen, Qiong Zeng, Zhanglin Cheng |
Vis. Informatics | 1 |
| 2018 | Revisiting Deep Intrinsic Image DecompositionsabstractWhile invaluable for many computer vision applications, decomposing a natural image into intrinsic reflectance and shading layers represents a challenging, underdetermined inverse problem. As opposed to strict reliance on conventional optimization or filtering solutions with strong prior assumptions, deep learning based approaches have also been proposed to compute intrinsic image decompositions when granted access to sufficient labeled training data. The downside is that current data sources are quite limited, and broadly speaking fall into one of two categories: either dense fully-labeled images in synthetic/narrow settings, or weakly-labeled data from relatively diverse natural scenes. In contrast to many previous learning-based approaches, which are often tailored to the structure of a particular dataset (and may not work well on others), we adopt core network structures that universally reflect loose prior knowledge regarding the intrinsic image formation process and can be largely shared across datasets. We then apply flexibly supervised loss layers that are customized for each source of ground truth labels. The resulting deep architecture achieves state-of-the-art results on all of the major intrinsic image benchmarks, and runs considerably faster than most at test time. Qingnan Fan, Jiaolong Yang, Gang Hua 0001, Baoquan Chen, David P. Wipf |
CVPR | 4 |
| 2018 | Decouple Learning for Parameterized Image Operators
Qingnan Fan, Dongdong Chen 0001, Lu Yuan 0001, Gang Hua 0001, Nenghai Yu, Baoquan Chen |
ECCV (13) | 6 |
| 2018 | SketchyScene: Richly-Annotated Scene Sketches
Changqing Zou, Qian Yu 0002, Ruofei Du, Haoran Mo, Yi-Zhe Song, Tao Xiang 0002, Chengying Gao, Baoquan Chen, Hao (Richard) Zhang |
ECCV (15) | 8 |
| 2018 | Caging Loops in Shape Embedding Space: Theory and ComputationabstractWe propose to synthesize feasible caging grasps for a target object through computing Caging Loops, a closed curve defined in the shape embedding space of the object. Different from the traditional methods, our approach decouples caging loops from the surface geometry of target objects through working in the embedding space. This enables us to synthesize caging loops encompassing multiple topological holes, instead of always tied with one specific handle which could be too small to be graspable by the robot gripper. Our method extracts caging loops through a topological analysis of the distance field defined for the target surface in the embedding space, based on a rigorous theoretical study on the relation between caging loops and the field topology. Due to the decoupling, our method can tolerate incomplete and noisy surface geometry of an unknown target object captured on-the-fly. We implemented our method with a robotic gripper and demonstrate through extensive experiments that our method can synthesize reliable grasps for objects with complex surface geometry and topology and in various scales. Shi-Qing Xin, Zengfu Gao, Kai Xu 0004, Changhe Tu, Baoquan Chen |
ICRA | 6 |
| 2018 | Cross-modal Moment Localization in VideosabstractIn this paper, we address the temporal moment localization issue, namely, localizing a video moment described by a natural language query in an untrimmed video. This is a general yet challenging vision-language task since it requires not only the localization of moments, but also the multimodal comprehension of textual-temporal information (e.g., "first" and "leaving") that helps to distinguish the desired moment from the others, especially those with the similar visual content. While existing studies treat a given language query as a single unit, we propose to decompose it into two components: the relevant cue related to the desired moment localization and the irrelevant one meaningless to the localization. This allows us to flexibly adapt to arbitrary queries in an end-to-end framework. In our proposed model, a language-temporal attention network is utilized to learn the word attention based on the temporal context information in the video. Therefore, our model can automatically select "what words to listen to" for localizing the desired moment. We evaluate the proposed model on two public benchmark datasets: DiDeMo and Charades-STA. The experimental results verify its superiority over several state-of-the-art methods. Meng Liu 0006, Xiang Wang 0010, Liqiang Nie, Qi Tian 0001, Baoquan Chen, Tat-Seng Chua |
ACM Multimedia | 5 |
| 2018 | DifNet: Semantic Segmentation by Diffusion NetworksabstractDeep Neural Networks (DNNs) have recently shown state of the art performance on semantic segmentation tasks, however, they still suffer from problems of poor boundary localization and spatial fragmented predictions. The difficulties lie in the requirement of making dense predictions from a long path model all at once since details are hard to keep when data goes through deeper layers. Instead, in this work, we decompose this difficult task into two relative simple sub-tasks: seed detection which is required to predict initial predictions without the need of wholeness and preciseness, and similarity estimation which measures the possibility of any two nodes belong to the same class without the need of knowing which class they are. We use one branch network for one sub-task each, and apply a cascade of random walks base on hierarchical semantics to approximate a complex diffusion process which propagates seed information to the whole image according to the estimated similarities. The proposed DifNet consistently produces improvements over the baseline models with the same depth and with the equivalent number of parameters, and also achieves promising performance on Pascal VOC and Pascal Context dataset. OurDifNet is trained end-to-end without complex loss functions. Peng Jiang 0002, Fanglin Gu, Yunhai Wang, Changhe Tu, Baoquan Chen |
NeurIPS | 5 |
| 2018 | PointCNN: Convolution On X-Transformed PointsabstractWe present a simple and general framework for feature learning from point cloud. The key to the success of CNNs is the convolution operator that is capable of leveraging spatially-local correlation in data represented densely in grids (e.g. images). However, point cloud are irregular and unordered, thus a direct convolving of kernels against the features associated with the points will result in deserting the shape information while being variant to the orders. To address these problems, we propose to learn a X-transformation from the input points, which is used for simultaneously weighting the input features associated with the points and permuting them into latent potentially canonical order. Then element-wise product and sum operations of typical convolution operator are applied on the X-transformed features. The proposed method is a generalization of typical CNNs into learning features from point cloud, thus we call it PointCNN. Experiments show that PointCNN achieves on par or better performance than state-of-the-art methods on multiple challenging benchmark datasets and tasks. Yangyan Li, Rui Bu, Mingchao Sun, Xinhan Di, Baoquan Chen |
NeurIPS | 6 |
| 2018 | Attentive Moment Retrieval in VideosabstractIn the past few years, language-based video retrieval has attracted a lot of attention. However, as a natural extension, localizing the specific video moments within a video given a description query is seldom explored. Although these two tasks look similar, the latter is more challenging due to two main reasons: 1) The former task only needs to judge whether the query occurs in a video and returns an entire video, but the latter is expected to judge which moment within a video matches the query and accurately returns the start and end points of the moment. Due to the fact that different moments in a video have varying durations and diverse spatial-temporal characteristics, uncovering the underlying moments is highly challenging. 2) As for the key component of relevance estimation, the former usually embeds a video and the query into a common space to compute the relevance score. However, the later task concerns moment localization where not only the features of a specific moment matter, but the context information of the moment also contributes a lot. For example, the query may contain temporal constraint words, such as "first'', therefore need temporal context to properly comprehend them. To address these issues, we develop an Attentive Cross-Modal Retrieval Network. In particular, we design a memory attention mechanism to emphasize the visual features mentioned in the query and simultaneously incorporate their context. In the light of this, we obtain the augmented moment representation. Meanwhile, a cross-modal fusion sub-network learns both the intra-modality and inter-modality dynamics, which can enhance the learning of moment-query representation. We evaluate our method on two datasets: DiDeMo and TACoS. Extensive experiments show the effectiveness of our model as compared to the state-of-the-art methods. Meng Liu 0006, Xiang Wang 0010, Liqiang Nie, Xiangnan He 0001, Baoquan Chen, Tat-Seng Chua |
SIGIR | 5 |
| 2018 | Active Assembly Guidance with Online Video ParsingabstractIn this paper, we introduce an online video-based system that actively assists users in assembly tasks. The system guides and monitors the assembly process by providing instructions and feedback on possibly erroneous operations, enabling easy and effective guidance in AR/MR applications. The core of our system is an online video-based assembly parsing method that can understand the assembly process, which is known to be extremely hard previously. Our method exploits the availability of the participating parts to significantly alleviate the problem, reducing the recognition task to an identification problem, within a constrained search space. To further constrain the search space, and understand the observed assembly activity, we introduce a tree-based global-inference technique. Our key idea is to incorporate part-interaction rules as powerful constraints which significantly regularize the search space and correctly parse the assembly video at interactive rates. Complex examples demonstrate the effectiveness of our method. Bin Wang 0035, Andrei Sharf, Yangyan Li, Fan Zhong 0001, Xueying Qin, Daniel Cohen-Or, Baoquan Chen |
VR | 8 |
| 2018 | Generating hybrid interior structure for 3D printing
Yuxin Mao, Lifang Wu, Dong-Ming Yan 0001, Jianwei Guo 0003, Chang Wen Chen, Baoquan Chen |
Comput. Aided Geom. Des. | 6 |
| 2018 | Anisotropic porous structure modeling for 3D printed objects
Jianming Ying, Lin Lu 0001, Lihao Tian, Baoquan Chen |
Comput. Graph. | 5 |
| 2018 | Class-sensitive shape dissimilarity metric
Manyi Li, Noa Fish, Lili Cheng, Changhe Tu, Daniel Cohen-Or, Hao (Richard) Zhang, Baoquan Chen |
Graph. Model. | 7 |
| 2018 | Neural best-buddies: sparse cross-domain correspondenceabstractCorrespondence between images is a fundamental problem in computer vision, with a variety of graphics applications. This paper presents a novel method for sparse cross-domain correspondence. Our method is designed for pairs of images where the main objects of interest may belong to different semantic categories and differ drastically in shape and appearance, yet still contain semantically related or geometrically similar parts. Our approach operates on hierarchies of deep features, extracted from the input images by a pre-trained CNN. Specifically, starting from the coarsest layer in both hierarchies, we search for Neural Best Buddies (NBB): pairs of neurons that are mutual nearest neighbors. The key idea is then to percolate NBBs through the hierarchy, while narrowing down the search regions at each level and retaining only NBBs with significant activations. Furthermore, in order to overcome differences in appearance, each pair of search regions is transformed into a common appearance. We evaluate our method via a user study, in addition to comparisons with alternative correspondence approaches. The usefulness of our method is demonstrated using a variety of graphics applications, including cross-domain image alignment, creation of hybrid images, automatic image morphing, and more. Kfir Aberman, Jing Liao 0001, Mingyi Shi, Dani Lischinski, Baoquan Chen, Daniel Cohen-Or |
ACM Trans. Graph. | 5 |
| 2018 | 3D fabrication with universal building blocks and pyramidal shellsabstractWe introduce a computational solution for cost-efficient 3D fabrication using universal building blocks. Our key idea is to employ a set of universal blocks, which can be massively prefabricated at a low cost, to quickly assemble and constitute a significant internal core of the target object, so that only the residual volume need to be 3D printed online. We further improve the fabrication efficiency by decomposing the residual volume into a small number of printing-friendly pyramidal pieces. Computationally, we face a coupled decomposition problem: decomposing the input object into an internal core and residual, and decomposing the residual, to fulfill a combination of objectives for efficient 3D fabrication. To this end, we formulate an optimization that jointly minimizes the residual volume, the number of pyramidal residual pieces, and the amount of support waste when printing the residual pieces. To solve the optimization in a tractable manner, we start with a maximal internal core and iteratively refine it with local cuts to minimize the cost function. Moreover, to efficiently explore the large search space, we resort to cost estimates aided by pre-computation and avoid the need to explicitly construct pyramidal decompositions for each solution candidate. Results show that our method can iteratively reduce the estimated printing time and cost, as well as the support waste, and helps to save hours of fabrication time and much material consumption. Xuelin Chen, Honghua Li, Chi-Wing Fu, Hao (Richard) Zhang, Daniel Cohen-Or, Baoquan Chen |
ACM Trans. Graph. | 6 |
| 2018 | Image smoothing via unsupervised learningabstractImage smoothing represents a fundamental component of many disparate computer vision and graphics applications. In this paper, we present a unified unsupervised (label-free) learning framework that facilitates generating flexible and high-quality smoothing effects by directly learning from data using deep convolutional neural networks (CNNs). The heart of the design is the training signal as a novel energy function that includes an edge-preserving regularizer which helps maintain important yet potentially vulnerable image structures, and a spatially-adaptive L p flattening criterion which imposes different forms of regularization onto different image regions for better smoothing quality. We implement a diverse set of image smoothing solutions employing the unified framework targeting various applications such as, image abstraction, pencil sketching, detail enhancement, texture removal and content-aware image manipulation, and obtain results comparable with or better than previous methods. Moreover, our method is extremely fast with a modern GPU (e.g, 200 fps for 1280×720 images). Qingnan Fan, Jiaolong Yang, David P. Wipf, Baoquan Chen, Xin Tong 0001 |
ACM Trans. Graph. | 4 |
| 2018 | DSCarver: decompose-and-spiral-carve for subtractive manufacturingabstractWe present an automatic algorithm for subtractive manufacturing of freeform 3D objects using high-speed machining (HSM) via CNC. A CNC machine operates a cylindrical cutter to carve off material from a 3D shape stock, following a tool path, to "expose" the target object. Our method decomposes the input object's surface into a small number of patches each of which is fully accessible and machinable by the CNC machine, in continuous fashion, under a fixed cutter-object setup configuration. This is achieved by covering the input surface with a minimum number of accessible regions and then extracting a set of machinable patches from each accessible region. For each patch obtained, we compute a continuous, space-filling, and iso-scallop tool path which conforms to the patch boundary, enabling efficient carving with high-quality surface finishing. The tool path is generated in the form of connected Fermat spirals , which have been generalized from a 2D fill pattern for layered manufacturing to work for curved surfaces. Furthermore, we develop a novel method to control the spacing of Fermat spirals based on directional surface curvature and adapt the heat method to obtain iso-scallop carving. We demonstrate automatic generation of accessible and machinable surface decompositions and iso-scallop Fermat spiral carving paths for freeform 3D objects. Comparisons are made to tool paths generated by commercial software in terms of real machining time and surface quality. Haisen Zhao, Hao (Richard) Zhang, Shi-Qing Xin, Yuanmin Deng, Changhe Tu, Wenping Wang 0001, Daniel Cohen-Or, Baoquan Chen |
ACM Trans. Graph. | 8 |
| 2018 | EdWordle: Consistency-Preserving Word Cloud EditingabstractWe present EdWordle, a method for consistently editing word clouds. At its heart, EdWordle allows users to move and edit words while preserving the neighborhoods of other words. To do so, we combine a constrained rigid body simulation with a neighborhood-aware local Wordle algorithm to update the cloud and to create very compact layouts. The consistent and stable behavior of EdWordle enables users to create new forms of word clouds such as storytelling clouds in which the position of words is carefully edited. We compare our approach with state-of-the-art methods and show that we can improve user performance, user satisfaction, as well as the layout itself. Yunhai Wang, Chen Bao, Lifeng Zhu, Oliver Deussen, Baoquan Chen, Michael Sedlmair |
IEEE Trans. Vis. Comput. Graph. | 6 |
| 2018 | A Perception-Driven Approach to Supervised Dimensionality Reduction for VisualizationabstractDimensionality reduction (DR) is a common strategy for visual analysis of labeled high-dimensional data. Low-dimensional representations of the data help, for instance, to explore the class separability and the spatial distribution of the data. Widely-used unsupervised DR methods like PCA do not aim to maximize the class separation, while supervised DR methods like LDA often assume certain spatial distributions and do not take perceptual capabilities of humans into account. These issues make them ineffective for complicated class structures. Towards filling this gap, we present a perception-driven linear dimensionality reduction approach that maximizes the perceived class separation in projections. Our approach builds on recent developments in perception-based separation measures that have achieved good results in imitating human perception. We extend these measures to be density-aware and incorporate them into a customized simulated annealing algorithm, which can rapidly generate a near optimal DR projection. We demonstrate the effectiveness of our approach by comparing it to state-of-the-art DR methods on 93 datasets, using both quantitative measure and human judgments. We also provide case studies with class-imbalanced and unlabeled data. Yunhai Wang, Kang Feng, Jian Zhang 0070, Chi-Wing Fu, Michael Sedlmair, Xiaohui Yu 0001, Baoquan Chen |
IEEE Trans. Vis. Comput. Graph. | 8 |
| 2018 | Line Graph or Scatter Plot? Automatic Selection of Methods for Visualizing Trends in Time SeriesabstractLine graphs are usually considered to be the best choice for visualizing time series data, whereas sometimes also scatter plots are used for showing main trends. So far there are no guidelines that indicate which of these visualization methods better display trends in time series for a given canvas. Assuming that the main information in a time series is its overall trend, we propose an algorithm that automatically picks the visualization method that reveals this trend best. This is achieved by measuring the visual consistency between the trend curve represented by a LOESS fit and the trend described by a scatter plot or a line graph. To measure the consistency between our algorithm and user choices, we performed an empirical study with a series of controlled experiments that show a large correspondence. In a factor analysis we furthermore demonstrate that various visual and data factors have effects on the preference for a certain type of visualization. Yunhai Wang, Fubo Han, Lifeng Zhu, Oliver Deussen, Baoquan Chen |
IEEE Trans. Vis. Comput. Graph. | 5 |
| 2018 | Revisiting Stress Majorization as a Unified Framework for Interactive Constrained Graph VisualizationabstractWe present an improved stress majorization method that incorporates various constraints, including directional constraints without the necessity of solving a constraint optimization problem. This is achieved by reformulating the stress function to impose constraints on both the edge vectors and lengths instead of just on the edge lengths (node distances). This is a unified framework for both constrained and unconstrained graph visualizations, where we can model most existing layout constraints, as well as develop new ones such as the star shapes and cluster separation constraints within stress majorization. This improvement also allows us to parallelize computation with an efficient GPU conjugant gradient solver, which yields fast and stable solutions, even for large graphs. As a result, we allow the constraint-based exploration of large graphs with 10K nodes - an approach which previous methods cannot support. Yunhai Wang, Yinqi Sun, Lifeng Zhu, Kecheng Lu 0002, Chi-Wing Fu, Michael Sedlmair, Oliver Deussen, Baoquan Chen |
IEEE Trans. Vis. Comput. Graph. | 9 |
| 2018 | Is There a Robust Technique for Selecting Aspect Ratios in Line Charts?abstractThe aspect ratio of a line chart heavily influences the perception of the underlying data. Different methods explore different criteria in choosing aspect ratios, but so far, it was still unclear how to select aspect ratios appropriately for any given data. This paper provides a guideline for the user to choose aspect ratios for any input 1D curves by conducting an in-depth analysis of aspect ratio selection methods both theoretically and experimentally. By formulating several existing methods as line integrals, we explain their parameterization invariance. Moreover, we derive a new and improved aspect ratio selection method, namely the -LOR (local orientation resolution), with a certain degree of parameterization invariance. Furthermore, we connect different methods, including AL (arc length based method), the banking to 45 principle, RV (resultant vector) and AS (average absolute slope), as well as -LOR and AO (average absolute orientation). We verify these connections by a comparative evaluation involving various data sets, and show that the selections by RV and -LOR are complementary to each other for most data. Accordingly, we propose the dual-scale banking technique that combines the strengths of RV and -LOR, and demonstrate its practicability using multiple real-world data sets. Yunhai Wang, Zeyu Wang 0005, Lifeng Zhu, Jian Zhang 0070, Chi-Wing Fu, Zhanglin Cheng, Changhe Tu, Baoquan Chen |
IEEE Trans. Vis. Comput. Graph. | 8 |
| 2018 | Classification of gait anomalies from kinect
Qiannan Li, Yafang Wang, Andrei Sharf, Ya Cao, Changhe Tu, Baoquan Chen, Shengyuan Yu |
Vis. Comput. | 6 |
| 2018 | Group optimization for multi-attribute visual embeddingabstractUnderstanding semantic similarity among images is the core of a wide range of computer graphics and computer vision applications. However, the visual context of images is often ambiguous as images that can be perceived with emphasis on different attributes. In this paper, we present a method for learning the semantic visual similarity among images, inferring their latent attributes and embedding them into multi-spaces corresponding to each latent attribute. We consider the multi-embedding problem as an optimization function that evaluates the embedded distances with respect to qualitative crowdsourced clusterings. The key idea of our approach is to collect and embed qualitative pairwise tuples that share the same attributes in clusters. To ensure similarity attribute sharing among multiple measures, image classification clusters are presented to, and solved by users. The collected image clusters are then converted into groups of tuples, which are fed into our group optimization algorithm that jointly infers the attribute similarity and multi-attribute embedding. Our multi-attribute embedding allows retrieving similar objects in different attribute spaces. Experimental results show that our approach outperforms state-of-the-art multi-embedding approaches on various datasets, and demonstrate the usage of the multi-attribute embedding in image retrieval application. Qiong Zeng, Wenzheng Chen, Zhuo Han, Mingyi Shi, Yanir Kleiman, Daniel Cohen-Or, Baoquan Chen, Yangyan Li |
Vis. Informatics | 7 |
| 2017 | A Generic Deep Architecture for Single Image Reflection Removal and Image SmoothingabstractThis paper proposes a deep neural network structure that exploits edge information in addressing representative low-level vision tasks such as layer separation and image filtering. Unlike most other deep learning strategies applied in this context, our approach tackles these challenging problems by estimating edges and reconstructing images using only cascaded convolutional layers arranged such that no handcrafted or application-specific image-processing components are required. We apply the resulting transferrable pipeline to two different problem domains that are both sensitive to edges, namely, single image reflection removal and image smoothing. For the former, using a mild reflection smoothness assumption and a novel synthetic data generation method that acts as a type of weak supervision, our network is able to solve much more difficult reflection cases that cannot be handled by previous methods. For the latter, we also exceed the state-of-the-art quantitative and qualitative results by wide margins. In all cases, the proposed framework is simple, fast, and easy to transfer across disparate domains. Qingnan Fan, Jiaolong Yang, Gang Hua 0001, Baoquan Chen, David P. Wipf |
ICCV | 4 |
| 2017 | Towards Micro-video Understanding by Joint Sequential-Sparse ModelingabstractLike the traditional long videos, micro-videos are the unity of textual, acoustic, and visual modalities. These modalities sequentially tell a real-life event from distinct angles. Yet, unlike the traditional long videos with rich content, micro-videos are very short, lasting for 6-15 seconds, and they hence usually convey one or a few high-level concepts. In the light of this, we have to characterize and jointly model the sparseness and multiple sequential structures for better micro-video understanding. To accomplish this, in this paper, we present an end-to-end deep learning model, which packs three parallel LSTMs to capture the sequential structures and a convolutional neural network to learn the sparse concept-level representations of micro-videos. We applied our model to the application of micro-video categorization. Besides, we constructed a real-world dataset for sequence modeling and released it to facilitate other researchers. Experimental results demonstrate that our model yields better performance than several state-of-the-art baselines. Meng Liu 0006, Liqiang Nie, Meng Wang 0001, Baoquan Chen |
ACM Multimedia | 4 |
| 2017 | Printable 3D TreesabstractAbstract With the growing popularity of 3D printing, different shape classes such as fibers and hair have been shown, driving research toward class‐specific solutions. Among them, 3D trees are an important class, consisting of unique structures, characteristics and botanical features. Nevertheless, trees are an especially challenging case for 3D manufacturing. They typically consist of non‐volumetric patch leaves, an extreme amount of small detail often below printable resolution and are often physically weak to be self‐sustainable. We introduce a novel 3D tree printability method which optimizes trees through a set of geometry modifications for manufacturing purposes. Our key idea is to formulate tree modifications as a minimal constrained set which accounts for the visual appearance of the model and its structural soundness. To handle non‐printable fine details, our method modifies the tree shape by gradually abstracting details of visible parts while reducing details of non‐visible parts. To guarantee structural soundness and to increase strength and stability, our algorithm incorporates a physical analysis and adjusts the tree topology and geometry accordingly while adhering to allometric rules. Our results show a variety of tree species with different complexity that are physically sound and correctly printed within reasonable time. The printed trees are correct in terms of their allometry and of high visual quality, which makes them suitable for various applications in the realm of outdoor design, modeling and manufacturing. Z. Bo, Lin Lu 0001, Andrei Sharf, Y. Xia, Oliver Deussen, Baoquan Chen |
Comput. Graph. Forum | 6 |
| 2017 | Tree Branch Level of Detail Models for Forest NavigationabstractAbstract We present a level of detail (LOD) method designed for tree branches. It can be combined with methods for processing tree foliage to facilitate navigation through large virtual forests. Starting from a skeletal representation of a tree, we fit polygon meshes of various densities to the skeleton while the mesh density is adjusted according to the required visual fidelity. For distant models, these branch meshes are gradually replaced with semi‐transparent lines until the tree recedes to a few lines. Construction of these complete LOD models is guided by error metrics to ensure smooth transitions between adjacent LOD models. We then present an instancing technique for discrete LOD branch models, consisting of polygon meshes plus semi‐transparent lines. Line models with different transparencies are instanced on the GPU by merging multiple tree samples into a single model. Our technique reduces the number of draw calls in GPU and increases rendering performance. Our experiments demonstrate that large‐scale forest scenes can be rendered with excellent detail and shadows in real time. Xiaopeng Zhang 0001, Guanbo Bao, Weiliang Meng, Marc Jaeger 0002, Hongjun Li 0002, Oliver Deussen, Baoquan Chen |
Comput. Graph. Forum | 7 |
| 2017 | Dip transform for 3D shape reconstructionabstractThe paper presents a novel three-dimensional shape acquisition and reconstruction method based on the well-known Archimedes equality between fluid displacement and the submerged volume. By repeatedly dipping a shape in liquid in different orientations and measuring its volume displacement, we generate the dip transform : a novel volumetric shape representation that characterizes the object's surface. The key feature of our method is that it employs fluid displacements as the shape sensor. Unlike optical sensors, the liquid has no line-of-sight requirements, it penetrates cavities and hidden parts of the object, as well as transparent and glossy materials, thus bypassing all visibility and optical limitations of conventional scanning devices. Our new scanning approach is implemented using a dipping robot arm and a bath of water, via which it measures the water elevation. We show results of reconstructing complex 3D shapes and evaluate the quality of the reconstruction with respect to the number of dips. Kfir Aberman, Oren Katzir, Zegang Luo, Andrei Sharf, Chen Greif, Baoquan Chen, Daniel Cohen-Or |
ACM Trans. Graph. | 7 |
| 2017 | Wasserstein Blue Noise SamplingabstractIn this article, we present a multi-class blue noise sampling algorithm by throwing samples as the constrained Wasserstein barycenter of multiple density distributions. Using an entropic regularization term, a constrained transport plan in the optimal transport problem is provided to break the partition required by the previous Capacity-Constrained Voronoi Tessellation method. The entropic regularization term cannot only control spatial regularity of blue noise sampling, but it also reduces conflicts between the desired centroids of Vornoi cells for multi-class sampling. Moreover, the adaptive blue noise property is guaranteed for each individual class, as well as their combined class. Our method can be easily extended to multi-class sampling on a point set surface. We also demonstrate applications in object distribution and color stippling. Hongxing Qin, Baoquan Chen |
ACM Trans. Graph. | 4 |
| 2016 | Synthesizing Training Images for Boosting Human 3D Pose EstimationabstractHuman 3D pose estimation from a single image is a challenging task with numerous applications. Convolutional Neural Networks (CNNs) have recently achieved superior performance on the task of 2D pose estimation from a single image, by training on images with 2D annotations collected by crowd sourcing. This suggests that similar success could be achieved for direct estimation of 3D poses. However, 3D poses are much harder to annotate, and the lack of suitable annotated training images hinders attempts towards end-to-end solutions. To address this issue, we opt to automatically synthesize training images with ground truth pose annotations. Our work is a systematic study along this road. We find that pose space coverage and texture diversity are the key ingredients for the effectiveness of synthetic training data. We present a fully automatic, scalable approach that samples the human pose space for guiding the synthesis procedure and extracts clothing textures from real images. Furthermore, we explore domain adaptation for bridging the gap between our synthetic training images and real testing photos. We demonstrate that CNNs trained with our synthetic images out-perform those trained with real photos on 3D pose estimation tasks. Wenzheng Chen, Yangyan Li, Hao Su 0001, Zhenhua Wang 0002, Changhe Tu, Dani Lischinski, Daniel Cohen-Or, Baoquan Chen |
3DV | 9 |
| 2016 | A Holistic Approach for Data-Driven Object Cutout
Huayong Xu, Yangyan Li, Wenzheng Chen, Dani Lischinski, Daniel Cohen-Or, Baoquan Chen |
ACCV (1) | 6 |
| 2016 | Mathematical foundations of arc length-based aspect ratio selectionabstractThe aspect ratio of a plot can strongly influence the perception of trends in the data. Arc length based aspect ratio selection (AL) has demonstrated many empirical advantages over previous methods. However, it is still not clear why and when this method works. In this paper, we attempt to unravel its mystery by exploring its mathematical foundation. First, we explain the rationale why this method is parameterization invariant and follow the same rationale to extend previous methods which are not parameterization invariant. As such, we propose maximizing weighted local curvature (MLC), a parameterization invariant form of local orientation resolution (LOR) and reveal the theoretical connection between average slope (AS) and resultant vector (RV). Furthermore, we establish a mathematical connection between AL and banking to 45 degrees and derive the upper and lower bounds of its average absolute slopes. Finally, we conduct a quantitative comparison that revises the understanding of aspect ratio selection methods in three aspects: (1) showing that AL, AWO and RV always perform very similarly while MS is not; (2) demonstrating the advantages in the robustness of RV over AL; (3) providing a counterexample where all previous methods produce poor results while MLC works well. Fubo Han, Yunhai Wang, Jian Zhang 0070, Oliver Deussen, Baoquan Chen |
PacificVis | 5 |
| 2016 | ShapeLearner: Towards Shape-Based Visual Knowledge HarvestingabstractThe deluge of images on the Web has led to a number of efforts to organize images semantically and mine visual knowledge. Despite enormous progress on categorizing entire images or bounding boxes, only few studies have targeted fine-grained image understanding at the level of specific shape contours. For instance, beyond recognizing that an image portrays a cat, we may wish to distinguish its legs, head, tail, and so on. To this end, we present ShapeLearner, a system that acquires such visual knowledge about object shapes and their parts in a semantic taxonomy, and then is able to exploit this hierarchy in order to analyze new kinds of objects that it has not observed before. ShapeLearner jointly learns this knowledge from sets of segmented images. The space of label and segmentation hypotheses is pruned and then evaluated using Integer Linear Programming. Experiments on a variety of shape classes show the accuracy and effectiveness of our method. Huayong Xu, Yafang Wang, Kang Feng, Gerard de Melo, Andrei Sharf, Baoquan Chen |
ECAI | 7 |
| 2016 | ShapeExplorer: Querying and Exploring Shapes using Visual Knowledge
Tong Ge, Yafang Wang, Gerard de Melo, Zengguang Hao, Andrei Sharf, Baoquan Chen |
EDBT | 6 |
| 2016 | Mobility Fitting using 4D RANSACabstractAbstract Capturing the dynamics of articulated models is becoming increasingly important. Dynamics, better than geometry, encode the functional information of articulated objects such as humans, robots and mechanics. Acquired dynamic data is noisy, sparse, and temporarily incoherent. The latter property is especially prominent for analysis of dynamics. Thus, processing scanned dynamic data is typically an ill‐posed problem. We present an algorithm that robustly computes the joints representing the dynamics of a scanned articulated object. Our key idea is to by‐pass the reconstruction of the underlying surface geometry and directly solve for motion joints. To cope with the often‐times extremely incoherent scans, we propose a space‐time fitting‐and‐voting approach in the spirit of RANSAC. We assume a restricted set of articulated motions defined by a set of joints which we fit to the 4D dynamic data and measure their fitting quality. Thus, we repeatedly select random subsets and fit with joints, searching for an optimal candidate set of mobility parameters. Without having to reconstruct surfaces as intermediate means, our approach gains the advantage of being robust and efficient. Results demonstrate the ability to reconstruct dynamics of various articulated objects consisting of a wide range of complex and compound motions. Hao Li 0015, Guowei Wan, Honghua Li, Andrei Sharf, Kai Xu 0004, Baoquan Chen |
Comput. Graph. Forum | 6 |
| 2016 | 3D attention-driven depth acquisition for object identificationabstractWe address the problem of autonomously exploring unknown objects in a scene by consecutive depth acquisitions. The goal is to reconstruct the scene while online identifying the objects from among a large collection of 3D shapes. Fine-grained shape identification demands a meticulous series of observations attending to varying views and parts of the object of interest. Inspired by the recent success of attention-based models for 2D recognition, we develop a 3D Attention Model that selects the best views to scan from, as well as the most informative regions in each view to focus on, to achieve efficient object recognition. The region-level attention leads to focus-driven features which are quite robust against object occlusion. The attention model, trained with the 3D shape collection, encodes the temporal dependencies among consecutive views with deep recurrent networks. This facilitates order-aware view planning accounting for robot movement cost. In achieving instance identification, the shape collection is organized into a hierarchy, associated with pre-trained hierarchical classifiers. The effectiveness of our method is demonstrated on an autonomous robot (PR) that explores a scene and identifies the objects to construct a 3D scene model. Kai Xu 0004, Min Liu 0019, Hui Huang 0004, Hao Su 0001, Daniel Cohen-Or, Baoquan Chen |
ACM Trans. Graph. | 9 |
| 2016 | Connected fermat spirals for layered fabricationabstractWe develop a new kind of "space-filling" curves, connected Fermat spirals , and show their compelling properties as a tool path fill pattern for layered fabrication. Unlike classical space-filling curves such as the Peano or Hilbert curves, which constantly wind and bind to preserve locality, connected Fermat spirals are formed mostly by long, low-curvature paths. This geometric property, along with continuity, influences the quality and efficiency of layered fabrication. Given a connected 2D region, we first decompose it into a set of sub-regions, each of which can be filled with a single continuous Fermat spiral. We show that it is always possible to start and end a Fermat spiral fill at approximately the same location on the outer boundary of the filled region. This special property allows the Fermat spiral fills to be joined systematically along a graph traversal of the decomposed sub-regions. The result is a globally continuous curve. We demonstrate that printing 2D layers following tool paths as connected Fermat spirals leads to efficient and quality fabrication, compared to conventional fill patterns. Haisen Zhao, Fanglin Gu, Qixing Huang, Jorge A. Garcia Galicia, Yong Chen 0017, Changhe Tu, Bedrich Benes, Hao (Richard) Zhang, Daniel Cohen-Or, Baoquan Chen |
ACM Trans. Graph. | 10 |
| 2016 | Printed Perforated Lampshades for Continuous Projective ImagesabstractWe present a technique for designing three-dimensional- (3D) printed perforated lampshades that project continuous grayscale images onto the surrounding walls. Given the geometry of the lampshade and a target grayscale image, our method computes a distribution of tiny holes over the shell, such that the combined footprints of the light emanating through the holes form the target image on a nearby diffuse surface. Our objective is to approximate the continuous tones and the spatial detail of the target image to the extent possible within the constraints of the fabrication process. To ensure structural integrity, there are lower bounds on the thickness of the shell, the radii of the holes, and the minimal distances between adjacent holes. Thus, the holes are realized as thin tubes distributed over the lampshade surface. The amount of light passing through a single tube may be controlled by the tube’s radius and by its orientation (tilt angle). The core of our technique thus consists of determining a suitable configuration of the tubes: their distribution across the relevant portion of the lampshade, as well as the parameters (radius, tilt angle) of each tube. This is achieved by computing a capacity-constrained Voronoi tessellation over a suitably defined density function and embedding a tube inside the maximal inscribed circle of each tessellation cell. Haisen Zhao, Lin Lu 0001, Dani Lischinski, Andrei Sharf, Daniel Cohen-Or, Baoquan Chen |
ACM Trans. Graph. | 7 |
| 2016 | Tree Modeling with Real Tree-Parts ExamplesabstractWe introduce a 3D tree modeling technique that utilizes examples of real trees to enhance tree creation with realistic structures and fine-level details. In contrast to previous works that use smooth generalized cylinders to represent tree branches, our method generates realistic looking tree models with complex branching geometry by employing an exemplar database consisting of real-life trees reconstructed from scanned data. These trees are sliced into representative parts (denoted as tree-cuts), representing trunk logs and branching structures. In the modeling process, tree-cuts are positioned in space in an intuitive manner, serving as efficient proxies that guide the creation of the complete tree. Allometry rules are taken into account to ensure reasonable relations between adjacent branches. Realism is further enhanced by automatically transferring geometric textures from our database onto tree branches as well as by guided growing of foliage. Our results demonstrate the complexity and variety of trees that can be generated with our method within few minutes. We carry a user study to test the effectiveness of our modeling technique. Ke Xie 0001, Feilong Yan, Andrei Sharf, Oliver Deussen, Hui Huang 0004, Baoquan Chen |
IEEE Trans. Vis. Comput. Graph. | 6 |
| 2015 | Hallucinating Stereoscopy from a Single ImageabstractAbstract We introduce a novel method for enabling stereoscopic viewing of a scene from a single pre‐segmented image. Rather than attempting full 3D reconstruction or accurate depth map recovery, we hallucinate a rough approximation of the scene's 3D model using a number of simple depth and occlusion cues and shape priors. We begin by depth‐sorting the segments, each of which is assumed to represent a separate object in the scene, resulting in a collection of depth layers. The shapes and textures of the partially occluded segments are then completed using symmetry and convexity priors. Next, each completed segment is converted to a union of generalized cylinders yielding a rough 3D model for each object. Finally, the object depths are refined using an iterative ground fitting process. The hallucinated 3D model of the scene may then be used to generate a stereoscopic image pair, or to produce images from novel viewpoints within a small neighborhood of the original view. Despite the simplicity of our approach, we show that it compares favorably with state‐of‐the‐art depth ordering methods. A user study was conducted showing that our method produces more convincing stereoscopic images than existing semi‐interactive and automatic single image depth recovery methods. Qiong Zeng, Wenzheng Chen, Changhe Tu, Daniel Cohen-Or, Dani Lischinski, Baoquan Chen |
Comput. Graph. Forum | 7 |
| 2015 | Skeleton-Intrinsic Symmetrization of ShapesabstractAbstract Enhancing the self‐symmetry of a shape is of fundamental aesthetic virtue. In this paper, we are interested in recovering the aesthetics ofintrinsicreflection symmetries, where an asymmetric shape is symmetrized while keeping its general pose and perceived dynamics. The key challenge to intrinsic symmetrization is that the input shape has only approximate reflection symmetries, possibly far from perfect. The main premise of our work is that curve skeletons provide a concise and effective shape abstraction for analyzing approximate intrinsic symmetries as well as symmetrization. By measuring intrinsic distances over a curve skeleton for symmetry analysis, symmetrizing the skeleton, and then propagating the symmetrization from skeleton to shape, our approach to shape symmetrization isskeleton‐intrinsic. Specifically, given an input shape and an extracted curve skeleton, we introduce the notion of abackboneas the path in the skeleton graph about which a self‐matching of the input shape is optimal. We define an objective function for the reflective self‐matching and develop an algorithm based on genetic programming to solve the global search problem for the backbone. The extracted backbone then guides the symmetrization of the skeleton, which in turn, guides the symmetrization of the whole shape. We show numerous intrinsic symmetrization results of hand drawn sketches and artist‐modeled or reconstructed 3D shapes, as well as several applications of skeleton‐intrinsic symmetrization of shapes. Zhuming Hao, Hui Huang 0004, Kai Xu 0004, Hao (Richard) Zhang, Daniel Cohen-Or, Baoquan Chen |
Comput. Graph. Forum | 7 |
| 2015 | Autoscanning for coupled scene reconstruction and proactive object analysisabstractDetailed scanning of indoor scenes is tedious for humans. We propose autonomous scene scanning by a robot to relieve humans from such a laborious task. In an autonomous setting, detailed scene acquisition is inevitably coupled with scene analysis at the required level of detail. We develop a framework for object-level scene reconstruction coupled with object-centric scene analysis. As a result, the autoscanning and reconstruction will be object-aware , guided by the object analysis. The analysis is, in turn, gradually improved with progressively increased object-wise data fidelity. In realizing such a framework, we drive the robot to execute an iterative analyze-and-validate algorithm which interleaves between object analysis and guided validations. The object analysis incorporates online learning into a robust graph-cut based segmentation framework, achieving a global update of object-level segmentation based on the knowledge gained from robot-operated local validation. Based on the current analysis, the robot performs proactive validation over the scene with physical push and scan refinement, aiming at reducing the uncertainty of both object-level segmentation and object-wise reconstruction. We propose a joint entropy to measure such uncertainty based on segmentation confidence and reconstruction quality, and formulate the selection of validation actions as a maximum information gain problem. The output of our system is a reconstructed scene with both object extraction and object-wise geometry fidelity. Kai Xu 0004, Hui Huang 0004, Hao Li 0015, Pinxin Long, Jianong Caichen, Baoquan Chen |
ACM Trans. Graph. | 8 |
| 2015 | Dapper: decompose-and-pack for 3D printingabstractWe pose the decompose-and-pack or DAP problem, which tightly combines shape decomposition and packing. While in general, DAP seeks to decompose an input shape into a small number of parts which can be efficiently packed, our focus is geared towards 3D printing. The goal is to optimally decompose-and-pack a 3D object into a printing volume to minimize support material, build time, and assembly cost. We present Dapper , a global optimization algorithm for the DAP problem which can be applied to both powder- and FDM-based 3D printing. The solution search is top-down and iterative. Starting with a coarse decomposition of the input shape into few initial parts, we progressively pack a pile in the printing volume, by iteratively docking parts, possibly while introducing cuts, onto the pile. Exploration of the search space is via a prioritized and bounded beam search , with breadth and depth pruning guided by local and global DAP objectives. A key feature of Dapper is that it works with pyramidal primitives, which are packing- and printing-friendly. Pyramidal shapes are also more general than boxes to reduce part counts, while still maintaining a suitable level of simplicity to facilitate DAP optimization. We demonstrate printing efficiency gains achieved by Dapper, compare to state-of-the-art alternatives, and show how fabrication criteria such as cut area and part size can be easily incorporated into our solution framework to produce more physically plausible fabrications. Xuelin Chen, Hao (Richard) Zhang, Jinjie Lin, Ruizhen Hu, Lin Lu 0001, Qixing Huang, Bedrich Benes, Daniel Cohen-Or, Baoquan Chen |
ACM Trans. Graph. | 9 |
| 2015 | JumpCut: non-successive mask transfer and interpolation for video cutoutabstractWe introduce JumpCut, a new mask transfer and interpolation method for interactive video cutout. Given a source frame for which a foreground mask is already available, we compute an estimate of the foreground mask at another, typically non-successive, target frame. Observing that the background and foreground regions typically exhibit different motions, we leverage these differences by computing two separate nearest-neighbor fields (split-NNF) from the target to the source frame. These NNFs are then used to jointly predict a coherent labeling of the pixels in the target frame. The same split-NNF is also used to aid a novel edge classifier in detecting silhouette edges (S-edges) that separate the foreground from the background. A modified level set method is then applied to produce a clean mask, based on the pixel labels and the S-edges computed by the previous two steps. The resulting mask transfer method may also be used for coherently interpolating the foreground masks between two distant source frames. Our results demonstrate that the proposed method is significantly more accurate than the existing state-of-the-art on a wide variety of video sequences. Thus, it reduces the required amount of user effort, and provides a basis for an effective interactive video object cutout tool. Qingnan Fan, Fan Zhong 0001, Dani Lischinski, Daniel Cohen-Or, Baoquan Chen |
ACM Trans. Graph. | 5 |
| 2015 | Global optimal searching for textureless 3D object tracking
Bin Wang 0035, Fan Zhong 0001, Xueying Qin, Baoquan Chen |
Vis. Comput. | 5 |
| 2014 | 2D-D Lifting for Shape ReconstructionabstractAbstract We present an algorithm for shape reconstruction from incomplete 3D scans by fusing together two acquisition modes: 2D photographs and 3D scans. The two modes exhibit complementary characteristics: scans have depth information, but are often sparse and incomplete; photographs, on the other hand, are dense and have high resolution, but lack important depth information. In this work we fuse the two modes, taking advantage of their complementary information, to enhance 3D shape reconstruction from an incomplete scan with a 2D photograph. We compute geometrical and topological shape properties in 2D photographs and use them to reconstruct a shape from an incomplete 3D scan in a principled manner. Our key observation is that shape properties such as boundaries, smooth patches and local connectivity, can be inferred with high confidence from 2D photographs. Thus, we register the 3D scan with the 2D photograph and use scanned points as 3D depth cues for lifting 2D shape structures into 3D. Our contribution is an algorithm which significantly regularizes and enhances the problem of 3D reconstruction from partial scans by lifting 2D shape structures into 3D. We evaluate our algorithm on various shapes which are loosely scanned and photographed from different views, and compare them with state‐of‐the‐art reconstruction methods. Liangliang Nan, Andrei Sharf, Baoquan Chen |
Comput. Graph. Forum | 3 |
| 2014 | Mobility-Trees for Indoor Scenes ManipulationabstractAbstract In this work, we introduce the ‘mobility‐tree’ construct for high‐level functional representation of complex 3D indoor scenes. In recent years, digital indoor scenes are becoming increasingly popular, consisting of detailed geometry and complex functionalities. These scenes often consist of objects that reoccur in various poses and interrelate with each other. In this work we analyse the reoccurrence of objects in the scene and automatically detect their functional mobilities. ‘Mobility’ analysis denotes the motion capabilities (i.e. degree of freedom) of an object and its subpart which typically relates to their indoor functionalities. We compute an object's mobility by analysing its spatial arrangement, repetitions and relations with other objects and store it in a ‘mobility‐tree’. Repetitive motions in the scenes are grouped in ‘mobility‐groups’, for which we develop a set of sophisticated controllers facilitating semantical high‐level editing operations. We show applications of our mobility analysis to interactive scene manipulation and reorganization, and present results for a variety of indoor scenes. Andrei Sharf, Hui Huang 0004, Cheng Liang 0005, Jiapei Zhang, Baoquan Chen, Minglun Gong |
Comput. Graph. Forum | 5 |
| 2014 | Inverse Procedural Modelling of TreesabstractAbstract Procedural tree models have been popular in computer graphics for their ability to generate a variety of output trees from a set of input parameters and to simulate plant interaction with the environment for a realistic placement of trees in virtual scenes. However, defining such models and their parameters is a difficult task. We propose an inverse modelling approach for stochastic trees that takes polygonal tree models as input and estimates the parameters of a procedural model so that it produces trees similar to the input. Our framework is based on a novel parametric model for tree generation and uses Monte Carlo Markov Chains to find the optimal set of parameters. We demonstrate our approach on a variety of input models obtained from different sources, such as interactive modelling systems, reconstructed scans of real trees and developmental models. Ondrej Stava, Sören Pirk, Julian Kratt, Baoquan Chen, Radomír Mech, Oliver Deussen, Bedrich Benes |
Comput. Graph. Forum | 4 |
| 2014 | Flower reconstruction from a single photoabstractAbstract We present a semi‐automatic method for reconstructing flower models from a single photograph. Such reconstruction is challenging since the 3D structure of a flower can appear ambiguous in projection. However, the flower head typically consists of petals embedded in 3D space that share similar shapes and form certain level of regular structure. Our technique employs these assumptions by first fitting a cone and subsequently a surface of revolution to the flower structure and then computing individual petal shapes from their projection in the photo. Flowers with multiple layers of petals are handled through processing different layers separately. Occlusions are dealt with both within and between petal layers. We show that our method allows users to quickly generate a variety of realistic 3D flowers from photographs and to animate an image using the underlying models reconstructed from our method. Feilong Yan, Minglun Gong, Daniel Cohen-Or, Oliver Deussen, Baoquan Chen |
Comput. Graph. Forum | 5 |
| 2014 | Build-to-last: strength to weight 3D printed objectsabstractThe emergence of low-cost 3D printers steers the investigation of new geometric problems that control the quality of the fabricated object. In this paper, we present a method to reduce the material cost and weight of a given object while providing a durable printed model that is resistant to impact and external forces. We introduce a hollowing optimization algorithm based on the concept of honeycomb-cells structure. Honeycombs structures are known to be of minimal material cost while providing strength in tension. We utilize the Voronoi diagram to compute irregular honeycomb-like volume tessellations which define the inner structure. We formulate our problem as a strength--to--weight optimization and cast it as mutually finding an optimal interior tessellation and its maximal hollowing subject to relieve the interior stress. Thus, our system allows to build-to-last 3D printed objects with large control over their strength-to-weight ratio and easily model various interior structures. We demonstrate our method on a collection of 3D objects from different categories. Furthermore, we evaluate our method by printing our hollowed models and measure their stress and weights. Lin Lu 0001, Andrei Sharf, Haisen Zhao, Qingnan Fan, Xuelin Chen, Yann Savoye, Changhe Tu, Daniel Cohen-Or, Baoquan Chen |
ACM Trans. Graph. | 10 |
| 2014 | Quality-driven poisson-guided autoscanningabstractWe present a quality-driven, Poisson-guided autonomous scanning method. Unlike previous scan planning techniques, we do not aim to minimize the number of scans needed to cover the object's surface, but rather to ensure the high quality scanning of the model. This goal is achieved by placing the scanner at strategically selected Next-Best-Views (NBVs) to ensure progressively capturing the geometric details of the object, until both completeness and high fidelity are reached. The technique is based on the analysis of a Poisson field and its geometric relation with an input scan. We generate a confidence map that reflects the quality/fidelity of the estimated Poisson iso-surface. The confidence map guides the generation of a viewing vector field, which is then used for computing a set of NBVs. We applied the algorithm on two different robotic platforms, a PR2 mobile robot and a one-arm industry robot. We demonstrated the advantages of our method through a number of autonomous high quality scannings of complex physical objects, as well as performance comparisons against state-of-the-art methods. Pinxin Long, Hui Huang 0004, Daniel Cohen-Or, Minglun Gong, Oliver Deussen, Baoquan Chen |
ACM Trans. Graph. | 8 |
| 2014 | Proactive 3D scanning of inaccessible partsabstractThe evolution of 3D scanning technologies have revolutionized the way real-world object are digitally acquired. Nowadays, high-definition and high-speed scanners can capture even large scale scenes with very high accuracy. Nevertheless, the acquisition of complete 3D objects remains a bottleneck, requiring to carefully sample the whole object's surface, similar to a coverage process. Holes and undersampled regions are common in 3D scans of complex-shaped objects with self occlusions and hidden interiors. In this paper we introduce the novel paradigm of proactive scanning , in which the user actively modifies the scene while scanning it, in order to reveal and access occluded regions. We take a holistic approach and integrate the user interaction into the continuous scanning process. Our algorithm allows for dynamic modifications of the scene as part of a global 3D scanning process. We utilize a scan registration algorithm to compute motion trajectories and separate between user modifications and other motions such as (hand-held) camera movements and small deformations. Thus, we reconstruct together the static parts into a complete unified 3D model. We evaluate our technique by scanning and reconstructing 3D objects and scenes consisting of inaccessible regions such as interiors, entangled plants and clutter. Feilong Yan, Andrei Sharf, Wenzhen Lin, Hui Huang 0004, Baoquan Chen |
ACM Trans. Graph. | 5 |
| 2014 | Morfit: interactive surface reconstruction from incomplete point clouds with curve-driven topology and geometry controlabstractWith significant data missing in a point scan, reconstructing a complete surface with sufficient geometric and topological fidelity is highly challenging. We present an interactive technique for surface reconstruction from incomplete and sparse scans of 3D objects possessing sharp features. A fundamental premise of our interaction paradigm is that directly editing data in 3D is not only counterintuitive but also ineffective, while working with 1D entities (i.e., curves) is a lot more manageable. To this end, we factor 3D editing into two "orthogonal" interactions acting on skeletal and profile curves of the underlying shape, controlling its topology and geometric features, respectively. For surface completion, we introduce a novel skeleton-driven morph-to-fit , or morfit , scheme which reconstructs the shape as an ensemble of generalized cylinders. Morfit is a hybrid operator which optimally interpolates between adjacent curve profiles (the "morph") and snaps the surface to input points (the "fit"). The interactive reconstruction iterates between user edits and morfit to converge to a desired final surface. We demonstrate various interactive reconstructions from point scans with sharp features and significant missing data. Kangxue Yin, Hui Huang 0004, Hao (Richard) Zhang, Minglun Gong, Daniel Cohen-Or, Baoquan Chen |
ACM Trans. Graph. | 6 |
| 2014 | Slippage-free background replacement for hand-held videoabstractWe introduce a method for replacing the background in a video of a moving foreground subject, when both the source video capturing the subject, and the target video capturing the new background scene, are natural videos, casually captured using a freely moving hand-held camera. We assume that the foreground subject has already been extracted, and focus on the challenging task of generating a video with a new background, such that the new background motion appears compatible with the original one. Failure to match the motion results in disturbing slippage or moonwalk artifacts, where the subject's feet appear to slide or slip over the ground. While matching the motion across the entire frame is impossible for scenes with differing geometry, we aim to match the local motion of the ground in the vicinity of the subject. This is achieved by reordering and warping the available target background frames in a manner that optimizes a suitably designed objective function. Fan Zhong 0001, Xueying Qin, Dani Lischinski, Daniel Cohen-Or, Baoquan Chen |
ACM Trans. Graph. | 6 |
| 2013 | Unsupervised co-segmentation of 3D shapes via affinity aggregation spectral clustering
Zizhao Wu, Yunhai Wang, Ruyang Shou, Baoquan Chen, Xinguo Liu |
Comput. Graph. | 4 |
| 2013 | Sketch-to-Design: Context-Based Part AssemblyabstractAbstract Designing 3D objects from scratch is difficult, especially when the user intent is fuzzy and lacks a clear target form. We facilitate design by providing reference and inspiration from existing model contexts. We rethink model design as navigating through different possible combinations of part assemblies based on a large collection of pre‐segmented 3D models. We propose an interactive sketch‐to‐design system, where the user sketches prominent features of parts to combine. The sketched strokes are analysed individually, and more importantly, in context with the other parts to generate relevant shape suggestions via adesign galleryinterface. As a modelling session progresses and more parts get selected, contextual cues become increasingly dominant, and the model quickly converges to a final form. As a key enabler, we use pre‐learned part‐based contextual information to allow the user to quickly explore different combinations of parts. Our experiments demonstrate the effectiveness of our approach for efficiently designing new variations from existing shape collections. Xiaohua Xie, Kai Xu 0004, Niloy J. Mitra, Daniel Cohen-Or, Wenyong Gong, Baoquan Chen |
Comput. Graph. Forum | 7 |
| 2013 | L1-medial skeleton of point cloudabstractWe introduce L 1 - medial skeleton as a curve skeleton representation for 3D point cloud data. The L 1 -median is well-known as a robust global center of an arbitrary set of points. We make the key observation that adapting L 1 -medians locally to a point set representing a 3D shape gives rise to a one-dimensional structure, which can be seen as a localized center of the shape. The primary advantage of our approach is that it does not place strong requirements on the quality of the input point cloud nor on the geometry or topology of the captured shape. We develop a L 1 -medial skeleton construction algorithm, which can be directly applied to an unoriented raw point scan with significant noise, outliers, and large areas of missing data. We demonstrate L 1 -medial skeletons extracted from raw scans of a variety of shapes, including those modeling high-genus 3D objects, plant-like structures, and curve networks. Hui Huang 0004, Daniel Cohen-Or, Minglun Gong, Hao (Richard) Zhang, Guiqing Li, Baoquan Chen |
ACM Trans. Graph. | 7 |
| 2013 | "Mind the gap": tele-registration for structure-driven image completionabstractConcocting a plausible composition from several non-overlapping image pieces, whose relative positions are not fixed in advance and without having the benefit of priors, can be a daunting task. Here we propose such a method, starting with a set of sloppily pasted image pieces with gaps between them. We first extract salient curves that approach the gaps from non-tangential directions, and use likely correspondences between pairs of such curves to guide a novel tele-registration method that simultaneously aligns all the pieces together. A structure-driven image completion technique is then proposed to fill the gaps, allowing the subsequent employment of standard in-painting tools to finish the job. Hui Huang 0004, Kangxue Yin, Minglun Gong, Dani Lischinski, Daniel Cohen-Or, Uri M. Ascher, Baoquan Chen |
ACM Trans. Graph. | 7 |
| 2013 | Analyzing growing plants from 4D point cloud dataabstractStudying growth and development of plants is of central importance in botany. Current quantitative are either limited to tedious and sparse manual measurements, or coarse image-based 2D measurements. Availability of cheap and portable 3D acquisition devices has the potential to automate this process and easily provide scientists with volumes of accurate data, at a scale much beyond the realms of existing methods. However, during their development, plants grow new parts (e.g., vegetative buds) and bifurcate to different components --- violating the central incompressibility assumption made by existing acquisition algorithms, which makes these algorithms unsuited for analyzing growth. We introduce a framework to study plant growth, particularly focusing on accurate localization and tracking topological events like budding and bifurcation. This is achieved by a novel forward-backward analysis, wherein we track robustly detected plant components back in time to ensure correct spatio-temporal event detection using a locally adapting threshold. We evaluate our approach on several groups of time lapse scans, often ranging from days to weeks, on a diverse set of plant species and use the results to animate static virtual plants or directly attach them to physical simulators. Yangyan Li, Xiaochen Fan, Niloy J. Mitra, Daniel A. Chamovitz, Daniel Cohen-Or, Baoquan Chen |
ACM Trans. Graph. | 6 |
| 2013 | Projective analysis for 3D shape segmentationabstractWe introduce projective analysis for semantic segmentation and labeling of 3D shapes. The analysis treats an input 3D shape as a collection of 2D projections, labels each projection by transferring knowledge from existing labeled images, and back-projects and fuses the labelings on the 3D shape. The image-space analysis involves matching projected binary images of 3D objects based on a novel bi-class Hausdorff distance . The distance is topology-aware by accounting for internal holes in the 2D figures and it is applied to piecewise-linearly warped object projections to compensate for part scaling and view discrepancies. Projective analysis simplifies the processing task by working in a lower-dimensional space, circumvents the requirement of having complete and well-modeled 3D shapes, and addresses the data challenge for 3D shape analysis by leveraging the massive available image data. A large and dense labeled set ensures that the labeling of a given projected image can be inferred from closely matched labeled images. We demonstrate semantic labeling of imperfect (e.g., incomplete or self-intersecting) 3D models which would be otherwise difficult to analyze without taking the projective analysis approach. Yunhai Wang, Minglun Gong, Tianhua Wang, Daniel Cohen-Or, Hao (Richard) Zhang, Baoquan Chen |
ACM Trans. Graph. | 6 |
| 2013 | Layered analysis of irregular facades via symmetry maximizationabstractWe present an algorithm for hierarchical and layered analysis of irregular facades, seeking a high-level understanding of facade structures. By introducing layering into the analysis, we no longer view a facade as a flat structure, but allow it to be structurally separated into depth layers, enabling more compact and natural interpretations of building facades. Computationally, we perform a symmetry-driven search for an optimal hierarchical decomposition defined by split and layering operations applied to an input facade. The objective is symmetry maximization , i.e., to maximize the sum of symmetry of the substructures resulting from recursive decomposition. To this end, we propose a novel integral symmetry measure, which behaves well at both ends of the symmetry spectrum by accounting for all partial symmetries in a discrete structure. Our analysis results in a structural representation, which can be utilized for structural editing and exploration of building facades. Hao (Richard) Zhang, Kai Xu 0004, Jinjie Lin, Daniel Cohen-Or, Baoquan Chen |
ACM Trans. Graph. | 6 |
| 2012 | Foreword to special section
Andrei Sharf, Baoquan Chen |
Comput. Graph. | 2 |
| 2012 | Sorting unorganized photo sets for urban reconstruction
Guowei Wan, Noah Snavely, Daniel Cohen-Or, Baoquan Chen, Sikun Li |
Graph. Model. | 5 |
| 2012 | Multi-scale partial intrinsic symmetry detectionabstractWe present an algorithm for multi-scale partial intrinsic symmetry detection over 2D and 3D shapes, where the scale of a symmetric region is defined by intrinsic distances between symmetric points over the region. To identify prominent symmetric regions which overlap and vary in form and scale, we decouple scale extraction and symmetry extraction by performing two levels of clustering. First, significant symmetry scales are identified by clustering sample point pairs from an input shape. Since different point pairs can share a common point, shape regions covered by points in different scale clusters can overlap. We introduce the symmetry scale matrix (SSM), where each entry estimates the likelihood two point pairs belong to symmetries at the same scale. The pair-to-pair symmetry affinity is computed based on a pair signature which encodes scales. We perform spectral clustering using the SSM to obtain the scale clusters. Then for all points belonging to the same scale cluster, we perform the second-level spectral clustering, based on a novel point-to-point symmetry affinity measure, to extract partial symmetries at that scale. We demonstrate our algorithm on complex shapes possessing rich symmetries at multiple scales. Kai Xu 0004, Hao (Richard) Zhang, Ramsay Dyer, Zhi-Quan Cheng, Ligang Liu 0001, Baoquan Chen |
ACM Trans. Graph. | 7 |
| 2012 | Active co-analysis of a set of shapesabstractUnsupervised co-analysis of a set of shapes is a difficult problem since the geometry of the shapes alone cannot always fully describe the semantics of the shape parts. In this paper, we propose a semi-supervised learning method where the user actively assists in the co-analysis by iteratively providing inputs that progressively constrain the system. We introduce a novel constrained clustering method based on a spring system which embeds elements to better respect their inter-distances in feature space together with the user-given set of constraints. We also present an active learning method that suggests to the user where his input is likely to be the most effective in refining the results. We show that each single pair of constraints affects many relations across the set. Thus, the method requires only a sparse set of constraints to quickly converge toward a consistent and error-free semantic labeling of the set. Yunhai Wang, Shmulik Asafi, Oliver van Kaick, Hao (Richard) Zhang, Daniel Cohen-Or, Baoquan Chen |
ACM Trans. Graph. | 6 |
| 2012 | Fit and diverse: set evolution for inspiring 3D shape galleriesabstractWe introduce set evolution as a means for creative 3D shape modeling, where an initial population of 3D models is evolved to produce generations of novel shapes. Part of the evolving set is presented to a user as a shape gallery to offer modeling suggestions. User preferences define the fitness for the evolution so that over time, the shape population will mainly consist of individuals with good fitness. However, to inspire the user's creativity, we must also keep the evolving set diverse. Hence the evolution is " fit and diverse ", drawing motivation from evolution theory. We introduce a novel part crossover operator which works at the finer-level part structures of the shapes, leading to significant variations and thus increased diversity in the evolved shape structures. Diversity is also achieved by explicitly compromising the fitness scores on a portion of the evolving population. We demonstrate the effectiveness of set evolution on man-made shapes. We show that selecting only models with high fitness leads to an elite population with low diversity. By keeping the population fit and diverse, the evolution can generate inspiring, and sometimes unexpected, shapes. Kai Xu 0004, Hao (Richard) Zhang, Daniel Cohen-Or, Baoquan Chen |
ACM Trans. Graph. | 4 |
| 2011 | 2D-3D fusion for layer decomposition of urban facadesabstractWe present a method for fusing two acquisition modes, 2D photographs and 3D LiDAR scans, for depth-layer decomposition of urban facades. The two modes have complementary characteristics: point cloud scans are coherent and inherently 3D, but are often sparse, noisy, and incomplete; photographs, on the other hand, are of high resolution, easy to acquire, and dense, but view-dependent and inherently 2D, lacking critical depth information. In this paper we use photographs to enhance the acquired LiDAR data. Our key observation is that with an initial registration of the 2D and 3D datasets we can decompose the input photographs into rectified depth layers. We decompose the input photographs into rectangular planar fragments and diffuse depth information from the corresponding 3D scan onto the fragments by solving a multi-label assignment problem. Our layer decomposition enables accurate repetition detection in each planar layer, using which we propagate geometry, remove outliers and enhance the 3D scan. Finally, the algorithm produces an enhanced, layered, textured model. We evaluate our algorithm on complex multi-planar building facades, where direct autocorrelation methods for repetition detection fail. We demonstrate how 2D photographs help improve the 3D scans by exploiting data redundancy, and transferring high level structural information to (plausibly) complete large missing regions. Yangyan Li, Andrei Sharf, Daniel Cohen-Or, Baoquan Chen, Niloy J. Mitra |
ICCV | 5 |
| 2011 | Structure-preserving retargeting of irregular 3D architectureabstractWe present an algorithm for interactive structure-preserving retargeting of irregular 3D architecture models, offering the modeler an easy-to-use tool to quickly generate a variety of 3D models that resemble an input piece in its structural style. Working on a more global and structural level of the input, our technique allows and even encourages replication of its structural elements, while taking into account their semantics and expected geometric interrelations such as alignments and adjacency. The algorithm performs automatic replication and scaling of these elements while preserving their structures. Instead of formulating and solving a complex constrained optimization, we decompose the input model into a set of sequences, each of which is a 1D structure that is relatively straightforward to retarget. As the sequences are retargeted in turn, they progressively constrain the retargeting of the remaining sequences. We demonstrate interactivity and variability of results from our retargeting algorithm using many examples modeled after real-world architectures exhibiting various forms of irregularity. Jinjie Lin, Daniel Cohen-Or, Hao (Richard) Zhang, Cheng Liang 0005, Andrei Sharf, Oliver Deussen, Baoquan Chen |
ACM Trans. Graph. | 7 |
| 2011 | Texture-lobes for tree modellingabstractWe present a lobe-based tree representation for modeling trees. The new representation is based on the observation that the tree's foliage details can be abstracted into canonical geometry structures, termed lobe-textures. We introduce techniques to (i) approximate the geometry of given tree data and encode it into a lobe-based representation, (ii) decode the representation and synthesize a fully detailed tree model that visually resembles the input. The encoded tree serves as a light intermediate representation, which facilitates efficient storage and transmission of massive amounts of trees, e.g., from a server to clients for interactive applications in urban environments. The method is evaluated by both reconstructing laser scanned trees (given as point sets) as well as re-representing existing tree models (given as polygons). Yotam Livny, Sören Pirk, Zhanglin Cheng, Feilong Yan, Oliver Deussen, Daniel Cohen-Or, Baoquan Chen |
ACM Trans. Graph. | 7 |
| 2011 | Conjoining Gestalt rules for abstraction of architectural drawingsabstractWe present a method for structural summarization and abstraction of complex spatial arrangements found in architectural drawings. The method is based on the well-known Gestalt rules, which summarize how forms, patterns, and semantics are perceived by humans from bits and pieces of geometric information. Although defining a computational model for each rule alone has been extensively studied, modeling a conjoint of Gestalt rules remains a challenge. In this work, we develop a computational framework which models Gestalt rules and more importantly, their complex interactions. We apply conjoining rules to line drawings, to detect groups of objects and repetitions that conform to Gestalt principles. We summarize and abstract such groups in ways that maintain structural semantics by displaying only a reduced number of repeated elements, or by replacing them with simpler shapes. We show an application of our method to line drawings of architectural models of various styles, and the potential of extending the technique to other computer-generated illustrations, and three-dimensional models. Liangliang Nan, Andrei Sharf, Ke Xie 0001, Tien-Tsin Wong, Oliver Deussen, Daniel Cohen-Or, Baoquan Chen |
ACM Trans. Graph. | 7 |
| 2010 | Consensus Skeleton for Non-rigid Space-time RegistrationabstractAbstract We introduce the notion of consensus skeletons for non‐rigid space‐time registration of a deforming shape. Instead of basing the registration on point features, which are local and sensitive to noise, we adopt the curve skeleton of the shape as a global and descriptive feature for the task. Our method uses no template and only assumes that the skeletal structure of the captured shape remains largely consistent over time. Such an assumption is generally weaker than those relying on large overlap of point features between successive frames, allowing for more sparse acquisition across time. Building our registration framework on top of the low‐dimensional skeleton‐time structure avoids heavy processing of dense point or volumetric data, while skeleton consensusization provides robust handling of incompatibilities between per‐frame skeletons. To register point clouds from all frames, we deform them by their skeletons, mirroring the skeleton registration process, to jump‐start a non‐rigid ICP. We present results for non‐rigid space‐time registration under sparse and noisy spatio‐temporal sampling, including cases where data was captured from only a single view. Andrei Sharf, Andrea Tagliasacchi, Baoquan Chen, Hao (Richard) Zhang, Alla Sheffer, Daniel Cohen-Or |
Comput. Graph. Forum | 4 |
| 2010 | Automatic reconstruction of tree skeletal structures from point cloudsabstractTrees, bushes, and other plants are ubiquitous in urban environments, and realistic models of trees can add a great deal of realism to a digital urban scene. There has been much research on modeling tree structures, but limited work on reconstructing the geometry of real-world trees -- even then, most works have focused on reconstruction from photographs aided by significant user interaction. In this paper, we perform active laser scanning of real-world vegetation and present an automatic approach that robustly reconstructs skeletal structures of trees, from which full geometry can be generated. The core of our method is a series of global optimizations that fit skeletal structures to the often sparse, incomplete, and noisy point data. A significant benefit of our approach is its ability to reconstruct multiple overlapping trees simultaneously without segmentation. We demonstrate the effectiveness and robustness of our approach on many raw scans of different tree varieties. Yotam Livny, Feilong Yan, Matt Olson, Baoquan Chen, Hao (Richard) Zhang, Jihad El-Sana |
ACM Trans. Graph. | 4 |
| 2010 | SmartBoxes for interactive urban reconstructionabstractWe introduce an interactive tool which enables a user to quickly assemble an architectural model directly over a 3D point cloud acquired from large-scale scanning of an urban scene. The user loosely defines and manipulates simple building blocks, which we call SmartBoxes, over the point samples. These boxes quickly snap to their proper locations to conform to common architectural structures. The key idea is that the building blocks are smart in the sense that their locations and sizes are automatically adjusted on-the-fly to fit well to the point data, while at the same time respecting contextual relations with nearby similar blocks. SmartBoxes are assembled through a discrete optimization to balance between two snapping forces defined respectively by a data-fitting term and a contextual term, which together assist the user in reconstructing the architectural model from a sparse and noisy point cloud. We show that a combination of the user's interactive guidance and high-level knowledge about the semantics of the underlying model, together with the snapping forces, allows the reconstruction of structures which are partially or even completely missing from the input. Liangliang Nan, Andrei Sharf, Hao (Richard) Zhang, Daniel Cohen-Or, Baoquan Chen |
ACM Trans. Graph. | 5 |
| 2010 | Non-local scan consolidation for 3D urban scenesabstractRecent advances in scanning technologies, in particular devices that extract depth through active sensing, allow fast scanning of urban scenes. Such rapid acquisition incurs imperfections: large regions remain missing, significant variation in sampling density is common, and the data is often corrupted with noise and outliers. However, buildings often exhibit large scale repetitions and self-similarities. Detecting, extracting, and utilizing such large scale repetitions provide powerful means to consolidate the imperfect data. Our key observation is that the same geometry, when scanned multiple times over reoccurrences of instances, allow application of a simple yet effective non-local filtering. The multiplicity of the geometry is fused together and projected to abase-geometrydefined by clustering corresponding surfaces. Denoising is applied by separating the process into off-plane and in-plane phases. We show that the consolidation of the reoccurrences provides robust denoising and allow reliable completion of missing parts. We present evaluation results of the algorithm on several LiDAR scans of buildings of varying complexity and styles. Andrei Sharf, Guowei Wan, Yangyan Li, Niloy J. Mitra, Daniel Cohen-Or, Baoquan Chen |
ACM Trans. Graph. | 7 |
| 2008 | Energy-Based Hierarchical Edge Clustering of GraphsabstractEffectively visualizing complex node-link graphs which depict relationships among data nodes is a challenging task due to the clutter and occlusion resulting from an excessive amount of edges. In this paper, we propose a novel energy-based hierarchical edge clustering method for node-link graphs. Taking into the consideration of the graph topology, our method first samples graph edges into segments using Delaunay triangulation to generate the control points, which are then hierarchically clustered by energy-based optimization. The edges are grouped according to their positions and directions to improve comprehensibility through abstraction and thus reduce visual clutter. The experimental results demonstrate the effectiveness of our proposed method in clustering edges and providing good high level abstractions of complex graphs. Hong Zhou 0004, Xiaoru Yuan, Weiwei Cui 0001, Huamin Qu, Baoquan Chen |
PacificVis | 5 |
| 2008 | Efficient and Dynamic Simplification of Line DrawingsabstractAbstract In this paper we present a pipeline for rendering dynamic 2D/3D line drawings efficiently. Our main goal is to create efficient static renditions and coherent animations of line drawings in a setting where lines can be added, deleted and arbitrarily transformed on‐the‐fly. Such a dynamic setting enables us to handle interactively sketched 2D line data, as well as arbitrarily transformed 3D line data in a unified manner. We evaluate the proximity of screen projected strokes to simplify them while preserving their continuity. We achieve this by using a special data structure that facilitates efficient proximity calculations in a dynamic setting. This on‐the‐fly proximity evaluation also facilitates generation of appropriate visibility cues to mitigate depth ambiguities and visual clutter for 3D line data. As we perform all these operations using only line data, we can create line drawings from 3D models without any surface information. We demonstrate the effectiveness and applicability of our approach by showing several examples with initial line representations obtained from a variety of sources: 2D and 3D hand‐drawn sketches and 3D salient geometry lines obtained from 3D surface representations. Amit Shesh, Baoquan Chen |
Comput. Graph. Forum | 2 |
| 2008 | Peek-in-the-Pic: Flying Through Architectural Scenes From a Single ImageabstractAbstract Many casually taken ‘tourist’ photographs comprise of architectural objects like houses, buildings, etc. Reconstructing such 3D scenes captured in a single photograph is a very challenging problem. We propose a novel approach to reconstruct such architectural scenes with minimal and simple user interaction, with the goal of providing 3D navigational capability to an image rather than acquiring accurate geometric detail. Our system, Peek‐in‐the‐Pic, is based on a sketch‐based geometry reconstruction paradigm. Given an image, the user simply traces out objects from it. Our system regards these as perspective line drawings, automatically completes them and reconstructs geometry from them. We make basic assumptions about the structure of traced objects and provide simple gestures for placing additional constraints. We also provide a simple sketching tool to progressively complete parts of the reconstructed buildings that are not visible in the image and cannot be automatically completed. Finally, we fill holes created in the original image when reconstructed buildings are removed from it, by automatic texture synthesis. Users can spend more time using interactive texture synthesis for further refining the image. Thus, instead of looking at flat images, a user can fly through them after some simple processing. Minimal manual work, ease of use and interactivity are the salient features of our approach. Amit Shesh, Baoquan Chen |
Comput. Graph. Forum | 2 |
| 2008 | Visual Clustering in Parallel CoordinatesabstractAbstract Parallel coordinates have been widely applied to visualize high‐dimensional and multivariate data, discerning patterns within the data through visual clustering. However, the effectiveness of this technique on large data is reduced by edge clutter. In this paper, we present a novel framework to reduce edge clutter, consequently improving the effectiveness of visual clustering. We exploit curved edges and optimize the arrangement of these curved edges by minimizing their curvature and maximizing the parallelism of adjacent edges. The overall visual clustering is improved by adjusting the shape of the edges while keeping their relative order. The experiments on several representative datasets demonstrate the effectiveness of our approach. Hong Zhou 0004, Xiaoru Yuan, Huamin Qu, Weiwei Cui 0001, Baoquan Chen |
Comput. Graph. Forum | 5 |
| 2008 | Architectural Modeling from Sparsely Scanned Range Data
Jie Chen 0007, Baoquan Chen |
Int. J. Comput. Vis. | 2 |
| 2007 | Special section on the joint Symposium on Point-based Graphics and Volume Graphics 2006
Mario Botsch, Baoquan Chen, Raghu Machiraju, Torsten Möller |
Comput. Graph. | 2 |
| 2007 | Simple Reconstruction of Tree Branches from a Single Range Image
Zhanglin Cheng, Xiaopeng Zhang 0001, Baoquan Chen |
J. Comput. Sci. Technol. | 3 |
| 2007 | Knowledge and heuristic-based modeling of laser-scanned treesabstractWe present a semi-automatic and efficient method for producing full polygonal models of range scanned trees, which are initially represented as sparse point clouds. First, a skeleton of the trunk and main branches of the tree is produced based on the scanned point clouds. Due to the unavoidable incompleteness of the point clouds produced by range scans of trees, steps are taken to synthesize additional branches to produce plausible support for the tree crown. Appropriate dimensions for each branch section are estimated using allometric theory. Using this information, a mesh is produced around the full skeleton. Finally, leaves are positioned, oriented and connected to nearby branches. Our process requires only minimal user interaction, and the full process including scanning and modeling can be completed within minutes. Hui Xu 0002, Nathan Gossett, Baoquan Chen |
ACM Trans. Graph. | 3 |
| 2006 | HDR VolVis: High Dynamic Range Volume VisualizationabstractIn this paper, we present an interactive high dynamic range volume visualization framework (HDR VolVis) for visualizing volumetric data with both high spatial and intensity resolutions. Volumes with high dynamic range values require high precision computing during the rendering process to preserve data precision. Furthermore, it is desirable to render high resolution volumes with low opacity values to reveal detailed internal structures, which also requires high precision compositing. High precision rendering will result in a high precision intermediate image (also known as high dynamic range image). Simply rounding up pixel values to regular display scales will result in loss of computed details. Our method performs high precision compositing followed by dynamic tone mapping to preserve details on regular display devices. Rendering high precision volume data requires corresponding resolution in the transfer function. To assist the users in designing a high resolution transfer function on a limited resolution display device, we propose a novel transfer function specification interface with nonlinear magnification of the density range and logarithmic scaling of the color/ opacity range. By leveraging modern commodity graphics hardware, multiresolution rendering techniques and out-of-core acceleration, our system can effectively produce an interactive visualization of large volume data, such as 2,048(3). Xiaoru Yuan, Minh X. Nguyen, Baoquan Chen, David H. Porter |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2005 | Stippling and Silhouettes Rendering in Geometry-Image Space
Xiaoru Yuan, Minh X. Nguyen, Nan Zhang 0011, Baoquan Chen |
Rendering Techniques | 4 |
| 2005 | High Dynamic Range Volume VisualizationabstractHigh resolution volumes require high precision compositing to preserve detailed structures. This is even more desirable for volumes with high dynamic range values. After the high precision intermediate image has been computed, simply rounding up pixel values to regular display scales loses the computed details. In this paper, we present a novel high dynamic range volume visualization method for rendering volume data with both high spatial and intensity resolutions. Our method performs high precision volume rendering followed by dynamic tone mapping to preserve details on regular display devices. By leveraging available high dynamic range image display algorithms, this dynamic tone mapping can be automatically adjusted to enhance selected features for the final display. We also present a novel transfer function design interface with nonlinear magnification of the density range and logarithmic scaling of the color/opacity range to facilitate high dynamic range volume visualization. By leveraging modern commodity graphics hardware and out-of-core acceleration, our system can produce an effective visualization of huge volume data. Xiaoru Yuan, Minh X. Nguyen, Baoquan Chen, David H. Porter |
IEEE Visualization | 3 |
| 2005 | Geometry completion and detail generation by texture synthesis
Minh X. Nguyen, Xiaoru Yuan, Baoquan Chen |
Vis. Comput. | 3 |
| 2005 | Volume cutout
Xiaoru Yuan, Nan Zhang 0011, Minh X. Nguyen, Baoquan Chen |
Vis. Comput. | 4 |
| 2004 | SMARTPAPER: An Interactive and User Friendly Sketching SystemabstractAbstract This paper describes an interactive sketching system for 3D design/modeling that diverts from the conventional menu‐and‐button interfaces of CAD tools. The system, dubbed SMARTPAPER, offers a unified sketching environment that supports direct sketching as well as gestured sketching with more emphasis on the former to encourage natural sketching styles. SMARTPAPER also provides a unified 2D and 3D drawing domain by allowing the user to sketch directly on a 3D model in addition to the usual 2D sketching from scratch. A natural sketching experience is offered by supporting casual sketching consisting of wiggly, discontinuous, overlapping strokes. The system is empowered by an array of seamlessly integrated 2D and 3D features such as 2D sketch cleaning, 3D reconstruction from 2D sketch, 3D transformations, sketching on 3D, and conventional 3D CSG operations like cutting and joining. The key to the success of SMARTPAPER is efficient and robust 3D reconstruction from a single freehand 2D sketch with minimal hints. We have employed and improved Lipson's optimization method, originally designed for offline reconstruction of engineering drawings, in our interactive system by leveraging additional clues obtained by interaction during sketching. Categories and Subject Descriptors (according to ACM CCS): I.3.3 [Computer Graphics]: Interaction Techniques, Pen‐based Interaction Amit Shesh, Baoquan Chen |
Comput. Graph. Forum | 2 |
| 2004 | Footprint Area Sampled TexturingabstractWe study texture projection based on a four region subdivision: magnification, minification, and two mixed regions. We propose improved versions of existing techniques by providing exact filtering methods which reduce both aliasing and overblurring, especially in the mixed regions. We further present a novel texture mapping algorithm called FAST (Footprint Area Sampled Texturing), which not only delivers high quality, but also is efficient. By utilizing coherence between neighboring pixels, performing prefiltering, and applying an area sampling scheme, we guarantee a minimum number of samples sufficient for effective antialiasing. Unlike existing methods (e.g., MIP-map, Feline), our method adapts the sampling rate in each chosen MIP-map level separately to avoid undersampling in the lower level l for effective antialiasing and to avoid oversampling in the higher level l + 1 for efficiency. Our method has been shown to deliver superior image quality to Feline and other methods while retaining the same efficiency. We also provide implementation trade offs to apply a variable degree of accuracy versus speed. Baoquan Chen, Frank Dachille, Arie E. Kaufman |
IEEE Trans. Vis. Comput. Graph. | 1 |
| 2003 | INSPIRE: An Interactive Image Assisted Non-Photorealistic Rendering SystemabstractWe present a GPU supported interactive non-photorealistic rendering system, INSPIRE, which performs feature extraction in both image space, on intermediately rendered images, and object space, on models of various representations, e.g., point, polygon, or hybrid models, without needing connectivity information. INSPIRE obtains interactive NPR rendering with most styles of existing NPR systems, but offers more flexibility on model representations and compromises little on rendering speed. Minh X. Nguyen, Hui Xu 0002, Xiaoru Yuan, Baoquan Chen |
PG | 4 |
| 2001 | POP: A Hybrid Point and Polygon Rendering System for Large DataabstractWe introduce a simple but effective extension to the existing pure point rendering systems. Rather than using only points, we use both points and polygons to represent and render large mesh models. We start from triangles as leaf nodes and build up a hierarchical tree structure with intermediate nodes as points. During the rendering, the system determines whether to use a point (of a certain intermediate level node) or a triangle (of a leaf node) for display depending on the screen contribution of each node. While points are used to speedup the rendering of distant objects, triangles are used to ensure the quality of close objects. Our method can accelerate the rendering of large models, compromising little in image quality. Baoquan Chen, Minh Xuan Nguye |
IEEE Visualization | 1 |
| 2000 | 3D Volume Rotation Using Shear Transformations
Baoquan Chen, Arie E. Kaufman |
Graph. Model. | 1 |
| 1999 | Forward Image MappingabstractWe present a novel forward image mapping algorithm, which speeds up perspective warping, as in texture mapping. It processes the source image in a special scanline order instead of the normal raster scanline order. This special scanline has the property of preserving parallelism when projecting to the target image. The algorithm reduces the complexity of perspective-correct image warping by eliminating the division per pixel and replacing it with a division per scanline. The method also corrects the perspective distortion in Gouraud shading with negligible overhead. Furthermore, the special scanline order is suitable for antialiasing using a more accurate antialiasing conic filter, with minimum additional cost. The algorithm is highlighted by incremental calculations and optimized memory bandwidth by reading each source pixel only once, suggesting a potential hardware implementation. Baoquan Chen, Frank Dachille, Arie E. Kaufman |
IEEE Visualization | 1 |
| 1999 | LOD-Sprite Technique for Accelerated Terrain Rendering
Baoquan Chen, J. Edward Swan II, Eddy Kuo, Arie E. Kaufman |
IEEE Visualization | 1 |
| 1999 | Navigating through sparse viewsabstractThis paper presents an image-based walkthrough technique where reference images are sparsely sampled along a path. The technique relies on a simple user interface for rapid modeling. Simple meshes are drawn to model and represent the underlying scene in each of the reference images. The meshes, consisting of only few polygons for each image are then registered by drawing a single line on each image, called model registration line, to form an aligned 3D model. To synthesize a novel view, two nearby reference images are mapped back onto their models by projective texture-mapping. Since the simple meshes are a crude approximation to the real model in the scene, image feature lines are drawn and used as aligning anchors to further register and blend two views together and form a final novel view. The simplicity of the technique yields rapid, “home-made” image-based walkthroughs. We have produced walkthroughs from a set of photographs to show the effectiveness of the technique. Shachar Fleishman, Baoquan Chen, Arie E. Kaufman, Daniel Cohen-Or |
VRST | 2 |
| 1997 | Multiresolution tetrahedral framework for visualizing regular volume dataabstractThe authors present a multiresolution framework, called Multi-Tetra framework, that approximates volume data with different levels-of-detail tetrahedra. The framework is generated through a recursive subdivision of the volume data and is represented by binary trees. Instead of using a certain level of the Multi-Tetra framework for approximation, an error-based model (EBM) is generated by recursively fusing a sequence of tetrahedra from different levels of the Multi-Tetra framework. The EBM significantly reduces the number of voxels required to model an object, while preserving the original topology. The approach provides continuous distribution of rendered intensity or generated isosurfaces along boundaries of different levels-of-detail thus solving the crack problem. The model supports typical rendering approaches, such as marching cubes, direct volume projection, and splatting. Experimental results demonstrate the strengths of the approach. Yong Zhou 0001, Baoquan Chen, Arie E. Kaufman |
IEEE Visualization | 2 |