VLDB 2026 Research / reviewers in the wild / expert
Weiwei Xu 0003
dblp:56/3321-3
· DBLP profile ↗
121ranked-venue papers
7as first author
61since 2021 · last 2026
0000-0003-3756-3539ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 105 · 6 first-author · 52 since 2021Artificial intelligence and machine learning · 29 · 24 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 5 · 4 since 2021Systems, architecture and hardware · 1 · 1 since 2021Computer networks · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Mimic-X: A Large-Scale Motion Dataset via Fast Physics-Based Controller AdaptationabstractLarge and high-quality motion datasets are essential for advancing human motion modeling. However, limitations of existing motion datasets, such as insufficient scale or inadequate quality, significantly hinder the progress of this field. To address these limitations, we introduce Mimic-X, a large-scale (52 hours), physically plausible 3D human motion dataset. To construct Mimic-X, we develop an adaptive option framework that controls a physically simulated character to imitate low-quality motions extracted from a vast collection of online videos. Specifically, we first apply hierarchical clustering to group motions into clusters, and then train option policies to mimic motions sampled from these clusters. Considering the noisy nature of low-quality motions, we utilize a separate encoder for each cluster to map the noisy motions within the cluster into a compact latent space. This significantly enhances the quality of the imitated motions while accelerating the learning process. Subsequently, we employ dynamic programming as a meta-policy to efficiently organize the option policies to generate complete motion clips. Finally, we perform fine-tuning to each motion sequence to further refine motion quality. The proposed adaptive option framework outperforms state-of-the-art human motion recovery methods across various evaluation metrics, demonstrating that motions in Mimic-X exhibit higher quality and greater physical plausibility. Furthermore, experimental results show that Mimic-X enhances the performance of motion generation methods, verifying its effectiveness for motion modeling tasks. Hongyu Tao, Shuaiying Hou, Junheng Fang, Mingyao Shi, Weiwei Xu 0003 |
AAAI | 5 |
| 2026 | FashionMAC: Deformation-Free Fashion Image Generation with Fine-Grained Model Appearance CustomizationabstractGarment-centric fashion image generation aims to synthesize realistic and controllable human models dressing a given garment, which has attracted growing interest due to its practical applications in e-commerce. The key challenges of the task lie in two aspects: (1) faithfully preserving the garment details, and (2) gaining fine-grained controllability over the model's appearance. Existing methods typically require performing garment deformation in the generation process, which often leads to garment texture distortions. Also, they fail to control the fine-grained attributes of the generated models, due to the lack of specifically designed mechanisms. To address these issues, we propose FashionMAC, a novel diffusion-based deformation-free framework that achieves high-quality and controllable fashion showcase image generation. The core idea of our framework is to eliminate the need for performing garment deformation and directly outpaint the garment segmented from a dressed person, which enables faithful preservation of the intricate garment details. Moreover, we propose a novel region-adaptive decoupled attention (RADA) mechanism along with a chained mask injection strategy to achieve fine-grained appearance controllability over the synthesized human models. Specifically, RADA adaptively predicts the generated regions for each fine-grained text attribute and enforces the text attribute to focus on the predicted regions by a chained mask injection strategy, significantly enhancing the visual fidelity and the controllability. Extensive experiments validate the superior performance of our framework compared to existing state-of-the-art methods. Jinxiao Li, Jingnan Wang, Zhiwen Zuo, Jianfeng Dong, Wei Li 0111, Chi Wang 0004, Weiwei Xu 0003, Xun Wang 0007 |
AAAI | 8 |
| 2026 | VividAnimator: An End-to-End Audio and Pose-driven Half-Body Human Animation FrameworkabstractExisting for audio- and pose-driven human animation methods often struggle with stiff head movements and blurry hands, primarily due to the weak correlation between audio and head movements and the structural complexity of hands. To address these issues, we propose VividAnimator, an end-to-end framework for generating high-quality, half-body human animations driven by audio and sparse hand pose conditions. Our framework introduces three key innovations. First, to overcome the instability and high cost of online codebook training, we pre-train a Hand Clarity Codebook (HCC) that encodes rich, high-fidelity hand texture priors, significantly mitigating hand degradation. Second, we design a Dual-Stream Audio-Aware Module (DSAA) to model lip synchronization and natural head pose dynamics separately while enabling interaction. Third, we introduce a Pose Calibration Trick (PCT) that refines and aligns pose conditions by relaxing rigid constraints, ensuring smooth and natural gesture transitions. Extensive experiments demonstrate that Vivid Animator achieves state-of-the-art performance, producing videos with superior hand detail, gesture realism, and identity consistency, validated by both quantitative metrics and qualitative evaluations. Donglin Huang, Yongyuan Li, Tianhang Liu, Junming Huang 0002, Xiaoda Yang, Chi Wang 0004, Weiwei Xu 0003 |
WACV | 7 |
| 2026 | GAN-Based Domain Adaptation for Image-Aware Layout Generation in Advertising Poster DesignabstractLayout plays a crucial role in graphic design and poster generation. Recently, the application of deep learning models for layout generation has gained significant attention. This paper focuses on using a GAN-based model conditioned on images to generate advertising poster graphic layouts, requiring a dataset of paired product images and layouts. To address this task, we introduce the Content-aware Graphic Layout Dataset (CGL-Dataset), consisting of 60,548 paired inpainted posters with annotations and 121,000 clean product images. The inpainting artifacts introduce a domain gap between the inpainted posters and clean images. To bridge this gap, we design two GAN-based models. The first model, CGL-GAN, uses Gaussian blur on the inpainted regions to generate layouts. The second model combines unsupervised domain adaptation by introducing a GAN with a pixel-level discriminator (PD), abbreviated as PDA-GAN, to generate image-aware layouts based on the visual texture of input images. The PD is connected to shallow-level feature maps and computes the GAN loss for each input-image pixel. Additionally, we propose three novel content-aware metrics to assess the model's ability to capture the intricate relationships between graphic elements and image content. Quantitative and qualitative evaluations demonstrate that PDA-GAN achieves state-of-the-art performance and generates high-quality image-aware layouts. Tiezheng Ge, Weiwei Xu 0003 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2026 | JGS2-GQ: Training-free 2nd Jacobi with Gaussian QuadratureabstractJGS2 is a Jacobi-like GPU simulation algorithm. It avoids the overshooting issue by augmenting each subproblem with a perturbation subspace that predicts the global influence of the local solve. The efficiency of JGS2 is due to Cubature-based subspace integration at each subproblem. Being a data-driven method, Cubature requires a set of representative deformed poses that cover deformations likely to occur in the simulation. This requirement is unlikely for high-resolution deformation with rich local details. Therefore, simulation performance and convergence degenerate when Cubature extrapolates. This paper proposes a training-free subspace integration algorithm based on classic Gaussian quadrature (GQ). We leverage the fact that the subproblem's subspace bases can be well-approximated by a low-degree multivariable polynomial, which suggests GQ an excellent candidate for Cubature substitute. To this end, we introduce a novel algorithm that adaptively generates the integration region for each subproblem. As a result, GQ integration can be analytically retrieved without cumbersome data generation and training. We also show how to handle frictional contact by modifying the pre-computed perturbation subspace. The resulting JGS2-GQ framework is more versatile than the vanilla JGS2 method. It is more stable for large and novel deformations, and is free of data generation and expensive training, while maintaining a near second-order convergence that is comparable to Newton. Performance-wise, JGS2-GQ is as efficient as JGS2, which is three orders faster than classic CPU methods and up to two orders faster than classic GPU algorithms. When novel deformation occurs, JGS2-GQ outperforms JGS2 over 50%. Dewen Guo, Yuqi Meng, Lei Lan, Weiwei Xu 0003, Chenfanfu Jiang, Yin Yang 0002 |
ACM Trans. Graph. | 7 |
| 2026 | From architecture to evaluation: A comprehensive review of video generation techniquesabstractThe rapid developments of artificial intelligence have significantly impacted daily life and content production modes. In the field of video generation, researchers are now exploring this emerging technique with innovative approaches, aiming to produce videos of higher quality, longer duration, and greater diversity. Currently, numerous video generation algorithms have been developed using different architecture designs. Unlike image generation, video generation requires maintaining consistency across both spatial and temporal dimensions while ensuring aesthetic quality and dynamic coherence, making it a more challenging task. In this survey, we provide a systematic review of existing video generation methods, tracing their evolution across different architectural paradigms. We further categorize recent models by their control conditions (e.g., text-to-video, image-to-video, multi-modal guidance) and summarize their unique theoretical foundations, architectural designs, and algorithmic innovations. In the meantime, we review the commonly used video datasets and analyze their applicability to different tasks. We also present evaluations of representative models to offer a more comprehensive perspective. Our goal is to provide a clear and concise overview of these algorithms, offering insights to support future breakthroughs in video generation. Chi Wang 0004, Guojun Lei, Weiwei Xu 0003 |
Virtual Real. Intell. Hardw. | 4 |
| 2025 | OmniSR: Shadow Removal Under Direct and Indirect LightingabstractShadows can originate from occlusions in both direct and indirect illumination. Although most current shadow removal research focuses on shadows caused by direct illumination, shadows from indirect illumination are often just as pervasive, particularly in indoor scenes. A significant challenge in removing shadows from indirect illumination is obtaining shadow-free images to train the shadow removal network. To overcome this challenge, we propose a novel rendering pipeline for generating shadowed and shadow-free images under direct and indirect illumination, and create a comprehensive synthetic dataset that contains over 30,000 image pairs, covering various object types and lighting conditions. We also propose an innovative shadow removal network that explicitly integrates semantic and geometric priors through concatenation and attention mechanisms. The experiments show that our method outperforms state-of-the-art shadow removal techniques and can effectively generalize to indoor and outdoor scenes under various lighting conditions, enhancing the overall effectiveness and applicability of shadow removal methods. Jiamin Xu, Renshu Gu, Weiwei Xu 0003, Gang Xu 0001 |
AAAI | 6 |
| 2025 | AnimateAnything: Consistent and Controllable Animation for Video GenerationabstractWe present a unified controllable video generation approach AnimateAnything that facilitates precise and consistent video manipulation across various conditions, including camera trajectories, text prompts, and user motion annotations. Specifically, we carefully design a multiscale control feature fusion network to construct a common motion representation for different conditions. It explicitly converts all control information into frame-by-frame optical flows. Then we incorporate the optical flows as motion priors to guide the final video generation. In addition, to reduce the flickering issues caused by large-scale motion, we propose a frequency-based stabilization module. It can enhance temporal coherence by ensuring the video's frequency domain consistency. Experiments demonstrate that our method outperforms the state-of-the-art approaches. For more details and videos, please refer to the anonymous webpage: https://yu-shaonian.github.io/Animate_Anything/. Guojun Lei, Chi Wang 0004, Hong Li 0016, Weiwei Xu 0003 |
CVPR | 6 |
| 2025 | DecoupledGaussian: Object-Scene Decoupling for Physics-Based InteractionabstractWe present DecoupledGaussian, a novel system that decouples static objects from their contacted surfaces captured in-the-wild videos, a key prerequisite for realistic Newtonian-based physical simulations. Unlike prior methods focused on synthetic data or elastic jittering along the contact surface, which prevent objects from fully detaching or moving independently, DecoupledGaussian allows for significant positional changes without being constrained by the initial contacted surface. Recognizing the limitations of current 2D inpainting tools for restoring 3D locations, our approach proposes joint Poisson fields to repair and expand the Gaussians of both objects and contacted scenes after separation. This is complemented by a multi-carve strategy to refine the object’s geometry. Our system enables realistic simulations of decoupling motions, collisions, and fractures driven by user-specified impulses, supporting complex interactions within and across multiple scenes. We validate DecoupledGaussian through a comprehensive user study and quantitative benchmarks. This system enhances digital interaction with objects and scenes in real-world environments, benefiting industries such as VR, robotics, and autonomous driving. Our project page is at: https://wangmiaowei.github.io/DecoupledGaussian.github.io/. Miaowei Wang, Weiwei Xu 0003, Rui Ma 0011, Changqing Zou, Daniel D. Morris |
CVPR | 3 |
| 2025 | Detail-Preserving Latent Diffusion for Stable Shadow RemovalabstractAchieving high-quality shadow removal with strong generalizability is challenging in scenes with complex global illumination. Due to the limited diversity in shadow removal datasets, current methods are prone to overfitting training data, often leading to reduced performance on unseen cases. To address this, we leverage the rich visual priors of a pre-trained Stable Diffusion (SD) model and propose a two-stage fine-tuning pipeline to adapt the SD model for stable and efficient shadow removal. In the first stage, we fix the VAE and fine-tune the denoiser in latent space, which yields substantial shadow removal but may lose some high-frequency details. To resolve this, we introduce a second stage, called the detail injection stage. This stage selectively extracts features from the VAE encoder to modulate the decoder, injecting fine details into the final results. Experimental results show that our method outperforms state-of-the-art shadow removal techniques. The cross-dataset evaluation further demonstrates that our method generalizes effectively to unseen data, enhancing the applicability of shadow removal methods. Jiamin Xu, Chi Wang 0004, Renshu Gu, Weiwei Xu 0003, Gang Xu 0001 |
CVPR | 6 |
| 2025 | Neural Shell Texture Splatting: More Details and Fewer Primitives
Anpei Chen, Jincheng Xiong, Pinxuan Dai, Yujun Shen, Weiwei Xu 0003 |
ICCV | 6 |
| 2025 | SpatialCrafter: Unleashing the Imagination of Video Diffusion Models for Scene Reconstruction from Limited ObservationsabstractNovel view synthesis (NVS) boosts immersive experiences in computer vision and graphics. Existing techniques, though progressed, rely on dense multi-view observations, restricting their application. This work takes on the challenge of reconstructing photorealistic 3D scenes from sparse or single-view inputs. We introduce SpatialCrafter, a framework that leverages the rich knowledge in video diffusion models to generate plausible additional observations, thereby alleviating reconstruction ambiguity. Through a trainable camera encoder and an epipolar attention mechanism for explicit geometric constraints, we achieve precise camera control and 3D consistency, further reinforced by a unified scale estimation strategy to handle scale discrepancies across datasets. Furthermore, by integrating monocular depth priors with semantic features in the video latent space, our framework directly regresses 3D Gaussian primitives and efficiently processes long-sequence features using a hybrid network structure. Extensive experiments show our method enhances sparse view reconstruction and restores the realistic appearance of 3D scenes. Songchun Zhang, Huiyao Xu, Sitong Guo, Zhongwei Xie, Hujun Bao, Weiwei Xu 0003, Changqing Zou |
ICCV | 6 |
| 2025 | MotionFlow: Learning Implicit Motion Flow for Complex Camera Trajectory Control in Video GenerationabstractGenerating videos guided by camera trajectories poses significant challenges in achieving consistency and generalizability, particularly when active objects are present. Existing approaches often attempt to learn object motions separately, which may lead to confusion regarding the relative motion between the camera and the objects. To address this challenge, we propose a novel approach that integrates both camera trajectory and object semantics and converts them into the reference motion of the corresponding pixels. Utilizing a stable diffusion network, we effectively extract reference motion maps in relation to the specified camera trajectory. These maps, along with an extracted semantic object prior, are then fed into an image-to-video network to generate the desired video that can accurately follow the designated camera trajectory while maintaining the consistency of object. Extensive experiments verify that our model outperforms SOTA methods by a large margin. Project Page: https://yu-shaonian.github.io/MotionFlow/ Guojun Lei, Chi Wang 0004, Hong Li 0016, Weiwei Xu 0003 |
ICME | 6 |
| 2025 | Intensity-Augmented LiDAR-Visual-Inertial Odometry and MeshingabstractThis paper presents a tightly-coupled LiDAR-Visual-Inertial Odometry (LIVO) system that integrates both LIO and VIO subsystems. The system jointly estimates the state by fusing LiDAR or visual data with Inertial Measurement Units (IMUs). It employs point-to-mesh tracking to optimize LiDAR poses and leverages intensity information from LiDAR point clouds to refine camera pose estimation. The optimized camera pose, derived from VIO, plays a crucial role in texture mapping and 3D geometry synthesis (3DGS) rendering. Our experiments demonstrate a significant improvement in average Peak Signal-to-Noise Ratio (PSNR) compared to existing methods, including R3LIVE, SR-LIVO, and FAST-LIVO. Furthermore, the system features a real-time mapping module implemented on the GPU, utilizing Truncated Signed Distance Function (TSDF) fields for global map maintenance and the Marching Cubes algorithm for mesh extraction. This approach ensures rapid and precise tracking and reconstruction capabilities. Additionally, our system supports real-time remeshing of the global map upon detecting loop closures, thereby enhancing the robustness and accuracy of the overall SLAM process. Yunfeng Hua, Qinyu Liu, Zhongwei Lin, Bintao Zhao, Tengfei Jiang, Shengjun Shi, Weiwei Xu 0003 |
IROS | 9 |
| 2025 | Enhanced Velocity Field Modeling for Gaussian Video ReconstructionabstractHigh-fidelity 3D video reconstruction is essential for enabling real-time rendering of dynamic scenes with realistic motion in VR/AR. The deformation field paradigm of 3D Gaussian splatting has achieved near-photorealistic results in video reconstruction due to the great representation capability of deep deformation networks. However, in videos with complex motion and significant scale variations, deformation networks often overfit to irregular Gaussian trajectories, leading to suboptimal visual quality. Moreover, the gradient-based densification strategy designed for static scene reconstruction proves inadequate to address the absence of dynamic content. In light of these challenges, we propose a flow-empowered velocity field modeling scheme tailored for Gaussian video reconstruction, dubbed FlowGaussian-VR. It consists of two core components: a velocity field rendering (VFR) pipeline which enables optical flow-based optimization, and a flow-assisted adaptive densification (FAD) strategy that adjusts the number and size of Gaussians in dynamic regions. We also explore a temporal velocity refinement (TVR) post-processing algorithm to further estimate and correct noise in Gaussian trajectories via extended Kalman filtering. We validate our model's effectiveness on multi-view dynamic reconstruction and novel view synthesis with real-world datasets containing challenging motion scenarios, demonstrating not only notable visual improvements (over 2.5 dB gain in PSNR) and less blurry artifacts in dynamic textures, but also regularized and trackable per-Gaussian trajectories. Xiaoyang Bai, Tongchen Zhang, Pengfei Shen, Weiwei Xu 0003, Yifan Peng 0001 |
ISMAR | 5 |
| 2025 | UniTransfer: Video Concept Transfer via Progressive Spatio-Temporal DecompositionabstractRecent advancements in video generation models have enabled the creation of diverse and realistic videos, with promising applications in advertising and film production. However, as one of the essential tasks of video generation models, video concept transfer remains significantly challenging.
Existing methods generally model video as an entirety, leading to limited flexibility and precision when solely editing specific regions or concepts. To mitigate this dilemma, we propose a novel architecture UniTransfer, which introduces both spatial and diffusion timestep decomposition in a progressive paradigm, achieving precise and controllable video concept transfer. Specifically, in terms of spatial decomposition, we decouple videos into three key components: the foreground subject, the background, and the motion flow. Building upon this decomposed formulation, we further introduce a dual-to-single-stream DiT-based architecture for supporting fine-grained control over different components in the videos. We also introduce a self-supervised pretraining strategy based on random masking to enhance the decomposed representation learning from large-scale unlabeled video data. Inspired by the Chain-of-Thought reasoning paradigm, we further revisit the denoising diffusion process and propose a Chain-of-Prompt (CoP) mechanism to achieve the timestep decomposition. We decompose the denoising process into three stages of different granularity and leverage large language models (LLMs) for stage-specific instructions to guide the generation progressively. We also curate an animal-centric video dataset called OpenAnimal to facilitate the advancement and benchmarking of research in video concept transfer.
Extensive experiments demonstrate that our method achieves high-quality and controllable video concept transfer across diverse reference images and scenes, surpassing existing baselines in both visual fidelity and editability. Guojun Lei, Tianhang Liu, Hong Li 0016, Chi Wang 0004, Weiwei Xu 0003 |
NeurIPS | 7 |
| 2025 | Efficient Object Reconstruction with Differentiable Area Light ShadingabstractIn 3D object reconstruction from photographs, estimating material properties is challenging. We propose an inverse rendering method that uses active area lighting: as this provides a wider range of lighting angles per photo than point lighting, material reconstruction can be more accurate for the same number of photos. We compare area light shading with point lighting. With either mesh or 3D Gaussian splatting pipelines, area lighting can improve BRDF reconstruction and leads to +3 dB relighting PSNR over point lights, or need only \(\nicefrac {1}{5}\) of the input photos for the same quality. We also compare area light shading with Monte Carlo ray tracing and with differential linearly transformed cosines (LTC) plus shadow visibility weighting. LTC can be faster, improving optimization times by 25%. In SOTA method-level comparisons, our approach improves material reconstruction, particularly for material roughness, leading to superior relighting quality. Yaoan Gao, Jiamin Xu, James Tompkin 0001, Qi Wang 0111, Hujun Bao, Yujun Shen, Huamin Wang 0001, Changqing Zou, Weiwei Xu 0003 |
SIGGRAPH Asia | 10 |
| 2025 | Parametric 3D human modeling with biharmonic SMPL
Yin Chen 0003, Yuping Ye, Weiwei Xu 0003, Qiliang Yang, Qizhen Zhou |
Comput. Graph. | 3 |
| 2025 | FS-control: Artistic fashion design with discriminated and conditional diffusion model
Jionghang Wu, Zhengkui Chen, Weiwei Xu 0003 |
Comput. Graph. | 4 |
| 2025 | Alternating optimization for bundle adjustment with closed form solutions
Chengzhe Meng, Yiwen Jiang, Weiwei Xu 0003 |
Sci. China Inf. Sci. | 3 |
| 2025 | Hybrid Mesh-Neural Representation for 3D Transparent Object ReconstructionabstractIn this study, we propose a novel method to reconstruct the 3D shapes of transparent objects using images captured by handheld cameras under natural lighting conditions. It combines the advantages of an explicit mesh and multi-layer perceptron (MLP) network as a hybrid representation to simplify the capture settings used in recent studies. After obtaining an initial shape through multi-view silhouettes, we introduced surface-based local MLPs to encode the vertex displacement field (VDF) for reconstructing surface details. The design of local MLPs allowed representation of the VDF in a piecewise manner using two-layer MLP networks to support the optimization algorithm. Defining local MLPs on the surface instead of on the volume also reduced the search space. Such a hybrid representation enabled us to relax the ray-pixel correspondences that represent the light path constraint to our designed ray-cell correspondences, which significantly simplified the implementation of a single-image-based environment-matting algorithm. We evaluated our representation and reconstruction algorithm on several transparent objects based on ground truth models. The experimental results show that our method produces high-quality reconstructions that are superior to those of state-of-the-art methods using a simplified data-acquisition setup. Jiamin Xu, Zihan Zhu, Hujun Bao, Weiwei Xu 0003 |
Comput. Vis. Media | 4 |
| 2025 | Accurate LiDAR-camera calibration using feature edges
Yunfeng Hua, Qinyu Liu, Tengfei Jiang, Weiwei Xu 0003 |
Image Vis. Comput. | 5 |
| 2025 | FaTNET: Feature-alignment transformer network for human pose transfer
Chengzhi Yuan, Lin Gao 0004, Weiwei Xu 0003, Xiaosong Yang, Pengjie Wang 0001 |
Pattern Recognit. | 4 |
| 2025 | Unsupervised Salient Object Detection on Light Field With High-Quality Synthetic LabelsabstractMost current Light Field Salient Object Detection (LFSOD) methods require full supervision with labor-intensive pixel-level annotations. Unsupervised Light Field Salient Object Detection (ULFSOD) has gained attention due to this limitation. However, existing methods use traditional handcrafted techniques to generate noisy pseudo-labels, which degrades the performance of models trained on them. To mitigate this issue, we present a novel learning-based approach to synthesize labels for ULFSOD. We introduce a prominent focal stack identification module that utilizes light field information (focal stack, depth map, and RGB color image) to generate high-quality pixel-level pseudo-labels, aiding network training. Additionally, we propose a novel model architecture for LFSOD, combining a multi-scale spatial attention module for focal stack information with a cross fusion module for RGB and focal stack integration. Through extensive experiments, we demonstrate that our pseudo-label generation method significantly outperforms existing methods in label quality. Our proposed model, trained with our labels, shows significant improvement on ULFSOD, achieving new state-of-the-art scores across public benchmarks. Yanfeng Zheng, Zhong Luo, Ying Cao 0001, Xiaosong Yang, Weiwei Xu 0003, Zheng Lin 0005, Pengjie Wang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2025 | 4D Gaussian Videos with Motion LayeringabstractOnline free-view navigation in volumetric videos requires high-quality rendering and real-time streaming in order to provide immersive user experiences. However, existing methods ( e.g. , dynamic NeRF and 3DGS) may not handle dynamic scenes with complex motions, and their models may not be streamable due to storage and bandwidth constraints. In this paper, we propose a novel 4D Gaussian Video (4DGV) approach that enables the creation and streaming of photorealistic, volumetric videos for dynamic scenes over the Internet. The core of our 4DGV is a novel streamable group of Gaussians (GOG) representation based on motion layering. Each GOG consists of static and dynamic points obtained via lifting 2D segmentation into 3D in motion layering, where the deformation of each dynamic point is represented as the temporal offset of its attributes. We also adaptively convert static points back to dynamic points to handle the appearance change, (e.g. , moving shadows and reflections), of static objects through optimization. To support real-time streaming of 4DGVs, we show that by applying quantization on Gaussian attributes and H.265 encoding on deformation offsets, our GOG representation can be significantly compressed (to around 6% of the original model size) without sacrificing the accuracy (PSNR loss less than 0.01dB). Extensive experiments on standard benchmarks demonstrate that our method outperforms state-of-the-art volumetric video approaches, with superior rendering quality and minimum storage overheads. Pinxuan Dai, Peiquan Zhang, Ke Xu 0010, Yifan Peng 0001, Dandan Ding, Yujun Shen, Yin Yang 0002, Xinguo Liu, Rynson W. H. Lau, Weiwei Xu 0003 |
ACM Trans. Graph. | 11 |
| 2025 | JGS2: Near Second-order Converging Jacobi/Gauss-Seidel for GPU ElastodynamicsabstractIn parallel simulation, convergence and parallelism are often seen as inherently conflicting objectives. Improved parallelism typically entails lighter local computation and weaker coupling, which unavoidably slow the global convergence. This paper presents a novel GPU algorithm that achieves convergence rates comparable to fullspace Newton's method while maintaining good parallelizability just like the Jacobi method. Our approach is built on a key insight into the phenomenon of overshoot. Overshoot occurs when a local solver aggressively minimizes its local energy without accounting for the global context, resulting in a local update that undermines global convergence. To address this, we derive a theoretically second-order optimal solution to mitigate overshoot. Furthermore, we adapt this solution into a pre-computable form. Leveraging Cubature sampling, our runtime cost is only marginally higher than the Jacobi method, yet our algorithm converges nearly quadratically as Newton's method. We also introduce a novel full-coordinate formulation for more efficient pre-computation. Our method integrates seamlessly with the incremental potential contact method and achieves second-order convergence for both stiff and soft materials. Experimental results demonstrate that our approach delivers high-quality simulations and outperforms state-of-the-art GPU methods with 50× to 100× better convergence. Lei Lan, Chun Yuan 0001, Weiwei Xu 0003, Hao Su 0001, Huamin Wang 0001, Chenfanfu Jiang, Yin Yang 0002 |
ACM Trans. Graph. | 4 |
| 2025 | Interpretable procedural material graph generation via diffusion models from reference images
Xiaoyu Lv, Zizhao Wu, Jiamin Xu, Xiaoling Gu, Ming Zeng 0008, Weiwei Xu 0003 |
Vis. Comput. | 6 |
| 2025 | Image-aware layout generation with user constraints for poster design
Kaixin Han, Weiwei Xu 0003 |
Vis. Comput. | 3 |
| 2024 | Text-Guided 3D Face Synthesis - From Generation to EditingabstractText-guided 3D face synthesis has achieved remarkable results by leveraging text-to-image (T2I) diffusion models. However, most existing works focus solely on the direct gen-eration, ignoring the editing, restricting them from synthe-sizing customized 3D faces through iterative adjustments. In this paper, we propose a unified text-guided framework from face generation to editing. In the generation stage, we propose a geometry-texture decoupled generation to miti-gate the loss of geometric details caused by coupling. Be-sides, decoupling enables us to utilize the generated geom-etry as a condition for texture generation, yielding highly geometry-texture aligned results. We further employ a fine-tuned texture diffusion model to enhance texture quality in both RGB and YUV space. In the editing stage, we first em-ploy a pre-trained diffusion model to update facial geometry or texture based on the texts. To enable sequential editing, we introduce a UV domain consistency preservation reg-ularization, preventing unintentional changes to irrelevant facial attributes. Besides, we propose a self-guided consis-tency weight strategy to improve editing efficacy while pre-serving consistency. Through comprehensive experiments, we showcase our method's superiority in face synthesis. Project page: https://faceg2e.github.io/. Yunjie Wu, Yapeng Meng, Zhipeng Hu, Lincheng Li, Haoqian Wu, Kun Zhou 0001, Weiwei Xu 0003, Xin Yu 0002 |
CVPR | 7 |
| 2024 | 3D-SceneDreamer: Text-Driven 3D-Consistent Scene GenerationabstractText-driven 3D scene generation techniques have made rapid progress in recent years. Their success is mainly at-tributed to using existing generative models to iteratively perform image warping and inpainting to generate 3D scenes. However, these methods heavily rely on the out-puts of existing models, leading to error accumulation in geometry and appearance that prevent the models from being used in various scenarios (e.g., outdoor and unreal sce-narios). To address this limitation, we generatively refine the newly generated local views by querying and aggregating global 3D information, and then progressively generate the 3D scene. Specifically, we employ a tri-plane features-based NeRF as a unified representation of the 3D scene to constrain global 3D consistency, and propose a generative refinement network to synthesize new contents with higher quality by exploiting the natural image prior from 2D dif-fusion model as well as the global 3D information of the current scene. Our extensive experiments demonstrate that, in comparison to previous methods, our approach supports wide variety of scene generation and arbitrary camera tra-jectories with improved visual quality and 3D consistency. Songchun Zhang, Quan Zheng 0004, Rui Ma 0011, Wei Hua 0002, Hujun Bao, Weiwei Xu 0003, Changqing Zou |
CVPR | 7 |
| 2024 | Local Gaussian Density Mixtures for Unstructured Lumigraph RenderingabstractPSNR 29.47 PSNR 28.80 PSNR 27.28 PSNR 30.84 PSNR 26.69 PSNR 26. Xiuchao Wu, Jiamin Xu, Chi Wang 0004, Yifan Peng 0001, Qixing Huang, James Tompkin 0001, Weiwei Xu 0003 |
SIGGRAPH Asia | 7 |
| 2024 | Guest Editorial: Proceedings of SPM 2024 Symposium
Lucia Romani, Weiwei Xu 0003 |
Comput. Aided Des. | 3 |
| 2024 | Gaussian Surfel Splatting for Live Human Performance CaptureabstractHigh-quality real-time rendering using user-affordable capture rigs is an essential property of human performance capture systems for real-world applications. However, state-of-the-art performance capture methods may not yield satisfactory rendering results under a very sparse (e.g., four) capture setting. Specifically, neural radiance field (NeRF)-based methods and 3D Gaussian Splatting (3DGS)-based methods tend to produce local geometry errors for unseen performers, while occupancy field (PIFu)-based methods often produce unrealistic rendering results. In this paper, we propose a novel generalizable neural approach to reconstruct and render the performers from very sparse RGBD streams in high quality. The core of our method is a novel point-based generalizable human (PGH) representation conditioned on the pixel-aligned RGBD features. The PGH representation learns a surface implicit function for the regression of surface points and a Gaussian implicit function for parameterizing the radiance fields of the regressed surface points with 2D Gaussian surfels, and uses surfel splatting for fast rendering. We learn this hybrid human representation via two novel networks. First, we propose a novel point-regressing network (PRNet) with a depth-guided point cloud initialization (DPI) method to regress an accurate surface point cloud based on the denoised depth information. Second, we propose a novel neural blending-based surfel splatting network (SPNet) to render high-quality geometries and appearances in novel views based on the regressed surface points and high-resolution RGBD features of adjacent views. Our method produces free-view human performance videos of 1K resolution at 12 fps on average. Experiments on two benchmarks show that our method outperforms state-of-the-art human performance capture methods. Ke Xu 0010, Yaoan Gao, Hujun Bao, Weiwei Xu 0003, Rynson W. H. Lau |
ACM Trans. Graph. | 5 |
| 2024 | Automatic Digital Garment Initialization from Sewing PatternsabstractThe rapid advancement of digital fashion and generative AI technology calls for an automated approach to transform digital sewing patterns into well-fitted garments on human avatars. When given a sewing pattern with its associated sewing relationships, the primary challenge is to establish an initial arrangement of sewing pieces that is free from folding and intersections. This setup enables a physics-based simulator to seamlessly stitch them into a digital garment, avoiding undesirable local minima. To achieve this, we harness AI classification, heuristics, and numerical optimization. This has led to the development of an innovative hybrid system that minimizes the need for user intervention in the initialization of garment pieces. The seeding process of our system involves the training of a classification network for selecting seed pieces, followed by solving an optimization problem to determine their positions and shapes. Subsequently, an iterative selection-arrangement procedure automates the selection of pattern pieces and employs a phased initialization approach to mitigate local minima associated with numerical optimization. Our experiments confirm the reliability, efficiency, and scalability of our system when handling intricate garments with multiple layers and numerous pieces. According to our findings, 68 percent of garments can be initialized with zero user intervention, while the remaining garments can be easily corrected through user operations. Chen Liu 0012, Weiwei Xu 0003, Yin Yang 0002, Huamin Wang 0001 |
ACM Trans. Graph. | 2 |
| 2024 | A Two-Part Transformer Network for Controllable Motion SynthesisabstractAlthough part-based motion synthesis networks have been investigated to reduce the complexity of modeling heterogeneous human motions, their computational cost remains prohibitive in interactive applications. To this end, we propose a novel two-part transformer network that aims to achieve high-quality, controllable motion synthesis results in real-time. Our network separates the skeleton into the upper and lower body parts, reducing the expensive cross-part fusion operations, and models the motions of each part separately through two streams of auto-regressive modules formed by multi-head attention layers. However, such a design might not sufficiently capture the correlations between the parts. We thus intentionally let the two parts share the features of the root joint and design a consistency loss to penalize the difference in the estimated root features and motions by these two auto-regressive modules, significantly improving the quality of synthesized motions. After training on our motion dataset, our network can synthesize a wide range of heterogeneous motions, like cartwheels and twists. Experimental and user study results demonstrate that our network is superior to state-of-the-art human motion synthesis networks in the quality of generated motions. Shuaiying Hou, Hongyu Tao, Hujun Bao, Weiwei Xu 0003 |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2024 | Intelligent CAD 2.0abstractIntegrating modern artificial intelligence (AI) techniques, particularly generative AI, holds the promise of revolutionizing computer-aided design (CAD) tools and the engineering design process. However, the direction of “AI+CAD” remains unclear: how will the current generation of intelligent CAD (ICAD) differ from its predecessor in the 1980s and 1990s, what strategic pathways should researchers and engineers pursue for its implementation, and what potential technical challenges might arise?As an attempt to address these questions, this paper investigates the transformative role of modern AI techniques in advancing CAD towards ICAD. It first analyzes the design process and reconsiders the roles AI techniques can assume in this process, highlighting how they can restructure the path humans, computers, and designs interact with each other. The primary conclusion is that ICAD systems should assume an intensional rather than an extensional role in the design process. This offers insights into the evaluation of the previous generation of ICAD (ICAD 1.0) and outlines a prospective framework and trajectory for the next generation of ICAD (ICAD 2.0). Qiang Zou 0007, Yingcai Wu, Weiwei Xu 0003, Shuming Gao |
Vis. Informatics | 4 |
| 2023 | CF-Font: Content Fusion for Few-Shot Font GenerationabstractContent and style disentanglement is an effective way to achieve few-shot font generation. It allows to transfer the style of the font image in a source domain to the style defined with a few reference images in a target domain. However, the content feature extracted using a representative font might not be optimal. In light of this, we propose a content fusion module (CFM) to project the content feature into a linear space defined by the content features of basis fonts, which can take the variation of content features caused by different fonts into consideration. Our method also allows to optimize the style representation vector of reference images through a lightweight iterative style-vector refinement (ISR) strategy. Moreover, we treat the 1D projection of a character image as a probability distribution and leverage the distance between two distributions as the reconstruction loss (namely projected character loss, PCL). Compared to L2 or L1 reconstruction loss, the distribution distance pays more attention to the global shape of characters. We have evaluated our method on a dataset of 300 fonts with 6.5k characters each. Experimental results verify that our method outperforms existing state-of-the-art few-shot font generation methods by a large margin. The source code can be found at https://github.com/wangchi95/CF-Font. Chi Wang 0004, Tiezheng Ge, Yuning Jiang 0001, Hujun Bao, Weiwei Xu 0003 |
CVPR | 6 |
| 2023 | Unsupervised Domain Adaption with Pixel-Level Discriminator for Image-Aware Layout GenerationabstractLayout is essential for graphic design and poster generation. Recently, applying deep learning models to generate layouts has attracted increasing attention. This paper focuses on using the GAN-based model conditioned on image contents to generate advertising poster graphic layouts, which requires an advertising poster layout dataset with paired product images and graphic layouts. However, the paired images and layouts in the existing dataset are collected by inpainting and annotating posters, respectively. There exists a domain gap between inpainted posters (source domain data) and clean product images (target domain data). Therefore, this paper combines unsupervised domain adaption techniques to design a GAN with a novel pixel-level discriminator (PD), called PDA-GAN, to generate graphic layouts according to image contents. The PD is connected to the shallow level feature map and computes the GAN loss for each input-image pixel. Both quantitative and qualitative evaluations demonstrate that PDAGAN can achieve state-of-the-art performances and generate high-quality image-aware graphic layouts for advertising posters. Tiezheng Ge, Yuning Jiang 0001, Weiwei Xu 0003 |
CVPR | 5 |
| 2023 | Weakly Supervised Image Matting via Patch Clustering
Yunke Zhang, Chi Wang 0004, Hujun Bao, Weiwei Xu 0003 |
ICIG (1) | 5 |
| 2023 | Neural Motion GraphabstractDeep learning techniques have been employed to design a controllable human motion synthesizer. Despite their potential, however, designing a neural network-based motion synthesis that enables flexible user interaction, fine-grained controllability, and the support of new types of motions at reduced time and space consumption costs remains a challenge. In this paper, we propose a novel approach, a neural motion graph, that addresses the challenge by enabling scalability to new motions while using compact neural networks. Our approach represents each type of motion with a separate neural node to reduce the cost of adding new motion types. In addition, designing a separate neural node for each motion type enables task-specific control strategies and has greater potential to achieve a high-quality synthesis of complex motions, such as the Mongolian dance. Furthermore, a single transition network, which acts as neural edges, is used to model the transition between two motion nodes. The transition network is designed with a lightweight control module to achieve a fine-grained response to user control signals. Overall, the design choice makes the neural motion graph highly controllable and scalable. In addition to being fully flexible to user interaction through high-level and fine-grained user-control signals, our experimental and subjective evaluation results demonstrate that our proposed approach, neural motion graph, outperforms state-of-the-art human motion synthesis methods in terms of the quality of controlled motion generation. Hongyu Tao, Shuaiying Hou, Changqing Zou, Hujun Bao, Weiwei Xu 0003 |
SIGGRAPH Asia | 5 |
| 2023 | A causal convolutional neural network for multi-subject motion modeling and generationabstractInspired by the success of WaveNet in multi-subject speech synthesis, we propose a novel neural network based on causal convolutions for multi-subject motion modeling and generation. The network can capture the intrinsic characteristics of the motion of different subjects, such as the influence of skeleton scale variation on motion style. Moreover, after fine-tuning the network using a small motion dataset for a novel skeleton that is not included in the training dataset, it is able to synthesize high-quality motions with a personalized style for the novel skeleton. The experimental results demonstrate that our network can model the intrinsic characteristics of motions well and can be applied to various motion modeling and synthesis tasks. Shuaiying Hou, Congyi Wang, Wenlin Zhuang, Yangang Wang 0001, Hujun Bao, Jinxiang Chai, Weiwei Xu 0003 |
Comput. Vis. Media | 8 |
| 2023 | Fast and Robust Non-Rigid Registration Using Accelerated Majorization-MinimizationabstractNon-rigid 3D registration, which deforms a source 3D shape in a non-rigid way to align with a target 3D shape, is a classical problem in computer vision. Such problems can be challenging because of imperfect data (noise, outliers and partial overlap) and high degrees of freedom. Existing methods typically adopt the$\ell _{p}$type robust norm to measure the alignment error and regularize the smoothness of deformation, and use a proximal algorithm to solve the resulting non-smooth optimization problem. However, the slow convergence of such algorithms limits their wide applications. In this paper, we propose a formulation for robust non-rigid registration based on a globally smooth robust norm for alignment and regularization, which can effectively handle outliers and partial overlaps. The problem is solved using the majorization-minimization algorithm, which reduces each iteration to a convex quadratic problem with a closed-form solution. We further apply Anderson acceleration to speed up the convergence of the solver, enabling the solver to run efficiently on devices with limited compute capability. Extensive experiments demonstrate the effectiveness of our method for non-rigid alignment between two shapes with outliers and partial overlaps, with quantitative evaluation showing that it outperforms state-of-the-art methods in terms of registration accuracy and computational speed. The source code is available athttps://github.com/yaoyx689/AMM_NRR. Yuxin Yao 0001, Bailin Deng, Weiwei Xu 0003, Juyong Zhang |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2023 | SAILOR: Synergizing Radiance and Occupancy Fields for Live Human Performance CaptureabstractImmersive user experiences in live VR/AR performances require a fast and accurate free-view rendering of the performers. Existing methods are mainly based on Pixel-aligned Implicit Functions (PIFu) or Neural Radiance Fields (NeRF). However, while PIFu-based methods usually fail to produce photorealistic view-dependent textures, NeRF-based methods typically lack local geometry accuracy and are computationally heavy ( e.g. , dense sampling of 3D points, additional fine-tuning, or pose estimation). In this work, we propose a novel generalizable method, named SAILOR, to create high-quality human free-view videos from very sparse RGBD live streams. To produce view-dependent textures while preserving locally accurate geometry, we integrate PIFu and NeRF such that they work synergistically by conditioning the PIFu on depth and then rendering view-dependent textures through NeRF. Specifically, we propose a novel network, named SRONet, for this hybrid representation. SRONet can handle unseen performers without fine-tuning. Besides, a neural blending-based ray interpolation approach, a tree-based voxel-denoising scheme, and a parallel computing pipeline are incorporated to reconstruct and render live free-view videos at 10 fps on average. To evaluate the rendering performance, we construct a real-captured RGBD benchmark from 40 performers. Experimental results show that SAILOR outperforms existing human reconstruction and performance capture methods. Ke Xu 0010, Yaoan Gao, Qilin Sun 0001, Hujun Bao, Weiwei Xu 0003, Rynson W. H. Lau |
ACM Trans. Graph. | 6 |
| 2023 | ScaNeRF: Scalable Bundle-Adjusting Neural Radiance Fields for Large-Scale Scene RenderingabstractHigh-quality large-scale scene rendering requires a scalable representation and accurate camera poses. This research combines tile-based hybrid neural fields with parallel distributive optimization to improve bundle-adjusting neural radiance fields. The proposed method scales with a divide-and-conquer strategy. We partition scenes into tiles, each with a multi-resolution hash feature grid and shallow chained diffuse and specular multilayer perceptrons (MLPs). Tiles unify foreground and background via a spatial contraction function that allows both distant objects in outdoor scenes and planar reflections as virtual images outside the tile. Decomposing appearance with the specular MLP allows a specular-aware warping loss to provide a second optimization path for camera poses. We apply the alternating direction method of multipliers (ADMM) to achieve consensus among camera poses while maintaining parallel tile optimization. Experimental results show that our method outperforms state-of-the-art neural scene rendering method quality by 5%--10% in PSNR, maintaining sharp distant objects and view-dependent reflections across six indoor and outdoor scenes. Xiuchao Wu, Jiamin Xu, Hujun Bao, Qixing Huang, Yujun Shen, James Tompkin 0001, Weiwei Xu 0003 |
ACM Trans. Graph. | 8 |
| 2022 | Active Boundary Loss for Semantic SegmentationabstractThis paper proposes a novel active boundary loss for semantic segmentation. It can progressively encourage the alignment between predicted boundaries and ground-truth boundaries during end-to-end training, which is not explicitly enforced in commonly used cross-entropy loss. Based on the predicted boundaries detected from the segmentation results using current network parameters, we formulate the boundary alignment problem as a differentiable direction vector prediction problem to guide the movement of predicted boundaries in each iteration. Our loss is model-agnostic and can be plugged in to the training of segmentation networks to improve the boundary details. Experimental results show that training with the active boundary loss can effectively improve the boundary F-score and mean Intersection-over-Union on challenging image and video object segmentation datasets. Chi Wang 0004, Yunke Zhang, Miaomiao Cui, Peiran Ren, Yin Yang 0002, Xuansong Xie, Xian-Sheng Hua 0001, Hujun Bao, Weiwei Xu 0003 |
AAAI | 9 |
| 2022 | NICE-SLAM: Neural Implicit Scalable Encoding for SLAMabstractNeural implicit representations have recently shown encouraging results in various domains, including promising progress in simultaneous localization and mapping (SLAM). Nevertheless, existing methods produce over- smoothed scene reconstructions and have difficulty scaling up to large scenes. These limitations are mainly due to their simple fully-connected network architecture that does not incorporate local information in the observations. In this paper, we present NICE-SLAM, a dense SLAM system that incorporates multi-level local information by introducing a hierarchical scene representation. Optimizing this representation with pre-trained geometric priors enables detailed reconstruction on large indoor scenes. Compared to recent neural implicit SLAM systems, our approach is more scalable, efficient, and robust. Experiments on five challenging datasets demonstrate competitive results of NICE-SLAM in both mapping and tracking quality. Project page: https://pengsongyou.github.io/nice-slam. Zihan Zhu, Songyou Peng, Viktor Larsson, Weiwei Xu 0003, Hujun Bao, Zhaopeng Cui, Martin R. Oswald, Marc Pollefeys |
CVPR | 4 |
| 2022 | Composition-aware Graphic Layout GAN for Visual-Textual Presentation DesignsabstractIn this paper, we study the graphic layout generation problem of producing high-quality visual-textual presentation designs for given images. We note that image compositions, which contain not only global semantics but also spatial information, would largely affect layout results. Hence, we propose a deep generative model, dubbed as composition-aware graphic layout GAN (CGL-GAN), to synthesize layouts based on the global and spatial visual contents of input images. To obtain training images from images that already contain manually designed graphic layout data, previous work suggests masking design elements (e.g., texts and embellishments) as model inputs, which inevitably leaves hint of the ground truth. We study the misalignment between the training inputs (with hint masks) and test inputs (without masks), and design a novel domain alignment module (DAM) to narrow this gap. For training, we built a large-scale layout dataset which consists of 60,548 advertising posters with annotated layout information. To evaluate the generated layouts, we propose three novel metrics according to aesthetic intuitions. Through both quantitative and qualitative evaluations, we demonstrate that the proposed model can synthesize high-quality graphic layouts according to image compositions. The data and code will be available at https://github.com/minzhouGithub/CGL-GAN. Tiezheng Ge, Yuning Jiang 0001, Weiwei Xu 0003 |
IJCAI | 6 |
| 2022 | Geometry-aware Two-scale PIFu Representation for Human ReconstructionabstractAlthough PIFu-based 3D human reconstruction methods are popular, the quality of recovered details is still unsatisfactory. In a sparse (e.g., 3 RGBD sensors) capture setting, the depth noise is typically amplified in the PIFu representation, resulting in flat facial surfaces and geometry-fallible bodies. In this paper, we propose a novel geometry-aware two-scale PIFu for 3D human reconstruction from sparse, noisy inputs. Our key idea is to exploit the complementary properties of depth denoising and 3D reconstruction, for learning a two-scale PIFu representation to reconstruct high-frequency facial details and consistent bodies separately. To this end, we first formulate depth denoising and 3D reconstruction as a multi-task learning problem. The depth denoising process enriches the local geometry information of the reconstruction features, while the reconstruction process enhances depth denoising with global topology information. We then propose to learn the two-scale PIFu representation using two MLPs based on the denoised depth and geometry-aware features. Extensive experiments demonstrate the effectiveness of our approach in reconstructing facial details and bodies of different poses and its superiority over state-of-the-art methods. Ke Xu 0010, Ziheng Duan, Hujun Bao, Weiwei Xu 0003, Rynson W. H. Lau |
NeurIPS | 5 |
| 2022 | Erroneous pixel prediction for semantic image segmentationabstractWe consider semantic image segmentation. Our method is inspired by Bayesian deep learning which improves image segmentation accuracy by modeling the uncertainty of the network output. In contrast to uncertainty, our method directly learns to predict the erroneous pixels of a segmentation network, which is modeled as a binary classification problem. It can speed up training comparing to the Monte Carlo integration often used in Bayesian deep learning. It also allows us to train a branch to correct the labels of erroneous pixels. Our method consists of three stages: (i) predict pixel-wise error probability of the initial result, (ii) redetermine new labels for pixels with high error probability, and (iii) fuse the initial result and the redetermined result with respect to the error probability. We formulate the error-pixel prediction problem as a classification task and employ an error-prediction branch in the network to predict pixel-wise error probabilities. We also introduce a detail branch to focus the training process on the erroneous pixels. We have experimentally validated our method on the Cityscapes and ADE20K datasets. Our model can be easily added to various advanced segmentation networks to improve their performance. Taking DeepLabv3+ as an example, our network can achieve 82.88% of mIoU on Cityscapes testing dataset and 45.73% on ADE20K validation dataset, improving corresponding DeepLabv3+ results by 0.74% and 0.13% respectively. Lixue Gong, Yunke Zhang, Yin Yang 0002, Weiwei Xu 0003 |
Comput. Vis. Media | 5 |
| 2022 | Learning-Based Bending Stiffness Parameter Estimation by a Drape TesterabstractReal-world fabrics often possess complicated nonlinear, anisotropic bending stiffness properties. Measuring the physical parameters of such properties for physics-based simulation is difficult yet unnecessary, due to the persistent existence of numerical errors in simulation technology. In this work, we propose to adopt a simulation-in-the-loop strategy: instead of measuring the physical parameters, we estimate the simulation parameters to minimize the discrepancy between reality and simulation. This strategy offers good flexibility in test setups, but the associated optimization problem is computationally expensive to solve by numerical methods. Our solution is to train a regression-based neural network for inferring bending stiffness parameters, directly from drape features captured in the real world. Specifically, we choose the Cusick drape test method and treat multiple-view depth images as the feature vector. To effectively and efficiently train our network, we develop a highly expressive and physically validated bending stiffness model, and we use the traditional cantilever test to collect the parameters of this model for 618 real-world fabrics. Given the whole parameter data set, we then construct a parameter subspace, generate new samples within the sub-space, and finally simulate and augment synthetic data for training purposes. The experiment shows that our trained system can replace cantilever tests for quick, reliable and effective estimation of simulation-ready parameters. Thanks to the use of the system, our simulator can now faithfully simulate bending effects comparable to those in the real world. Xudong Feng, Weiwei Xu 0003, Huamin Wang 0001 |
ACM Trans. Graph. | 3 |
| 2022 | Automatic quantization for physics-based simulationabstractQuantization has proven effective in high-resolution and large-scale simulations, which benefit from bit-level memory saving. However, identifying a quantization scheme that meets the requirement of both precision and memory efficiency requires trial and error. In this paper, we propose a novel framework to allow users to obtain a quantization scheme by simply specifying either an error bound or a memory compression rate. Based on the error propagation theory, our method takes advantage of auto-diff to estimate the contributions of each quantization operation to the total error. We formulate the task as a constrained optimization problem, which can be efficiently solved with analytical formulas derived for the linearized objective function. Our workflow extends the Taichi compiler and introduces dithering to improve the precision of quantized simulations. We demonstrate the generality and efficiency of our method via several challenging examples of physics-based simulation, which achieves up to 2.5× memory compression without noticeable degradation of visual quality in the results. Our code and data are available at https://github.com/Hanke98/AutoQantizer. Jiafeng Liu, Haoyang Shi, Yin Yang 0002, Chongyang Ma, Weiwei Xu 0003 |
ACM Trans. Graph. | 6 |
| 2022 | Scalable neural indoor scene renderingabstractWe propose a scalable neural scene reconstruction and rendering method to support distributed training and interactive rendering of large indoor scenes. Our representation is based on tiles. Tile appearances are trained in parallel through a background sampling strategy that augments each tile with distant scene information via a proxy global mesh. Each tile has two low-capacity MLPs: one for view-independent appearance (diffuse color and shading) and one for view-dependent appearance (specular highlights, reflections). We leverage the phenomena that complex view-dependent scene reflections can be attributed to virtual lights underneath surfaces at the total ray distance to the source. This lets us handle sparse samplings of the input scene where reflection highlights do not always appear consistently in input images. We show interactive free-viewpoint rendering results from five scenes, one of which covers an area of more than 100 m 2 . Experimental results show that our method produces higher-quality renderings than a single large-capacity MLP and five recent neural proxy-geometry and voxel-based baseline methods. Our code and data are available at project webpage https://xchaowu.github.io/papers/scalable-nisr. Xiuchao Wu, Jiamin Xu, Zihan Zhu, Hujun Bao, Qixing Huang, James Tompkin 0001, Weiwei Xu 0003 |
ACM Trans. Graph. | 7 |
| 2021 | Location-aware Single Image Reflection RemovalabstractThis paper proposes a novel location-aware deep-learning-based single image reflection removal method. Our network has a reflection detection module to regress a probabilistic reflection confidence map, taking multi-scale Laplacian features as inputs. This probabilistic map tells if a region is reflection-dominated or transmission-dominated, and it is used as a cue for the network to control the feature flow when predicting the reflection and transmission layers. We design our network as a recurrent network to progressively refine reflection removal results at each iteration. The novelty is that we leverage Laplacian kernel parameters to emphasize the boundaries of strong reflections. It is beneficial to strong reflection detection and substantially improves the quality of reflection removal results. Extensive experiments verify the superior performance of the proposed method over state-of-the-art approaches. Our code and the pre-trained model can be found at https://github.com/zdlarr/Location-aware-SIRR. Ke Xu 0010, Yin Yang 0002, Hujun Bao, Weiwei Xu 0003, Rynson W. H. Lau |
ICCV | 5 |
| 2021 | Image Re-composition via Regional Content-Style DecouplingabstractTypical image composition harmonizes regions from different images to a single plausible image. We extend the idea of image composition by introducing the content-style decomposition and combination to form the concept of image re-composition. In other words, our image re-composition could arbitrarily combine those contents and styles decomposed from different images to generate more diverse images in a unified framework. In the decomposition stage, we incorporate the whitening normalization to obtain a more thorough content-style decoupling, which substantially improves the re-composition results. Moreover, to handle the variation of structure and texture of different objects in an image, we design the network to support regional feature representation and achieve region-aware content-style decomposition. Regarding the composition stage, we propose a cycle consistency loss to constrain the network preserving the content and style information during the composition. Our method can produce diverse re-composition results, including content-content, content-style and style-style. Our experimental results demonstrate a large improvement over the current state-of-the-art methods. Wei Li 0111, Hong Zhang 0009, Ruigang Yang, Weiwei Xu 0003 |
ACM Multimedia | 7 |
| 2021 | Attention-guided Temporally Coherent Video Object MattingabstractThis paper proposes a novel deep learning-based video object matting method that can achieve temporally coherent matting results. Its key component is an attention-based temporal aggregation module that maximizes image matting networks' strength for video matting networks. This module computes temporal correlations for pixels adjacent to each other along the time axis in feature space, which is robust against motion noises. We also design a novel loss term to train the attention weights, which drastically boosts the video matting performance. Besides, we show how to effectively solve the trimap generation problem by fine-tuning a state-of-the-art video object segmentation network with a sparse set of user-annotated keyframes. To facilitate video matting and trimap generation networks' training, we construct a large-scale video matting dataset with 80 training and 28 validation foreground video clips with ground-truth alpha mattes. Experimental results show that our method can generate high-quality alpha mattes for various videos featuring appearance change, occlusion, and fast motion. Our code and dataset can be found at: https://github.com/yunkezhang/TCVOM Yunke Zhang, Chi Wang 0004, Miaomiao Cui, Peiran Ren, Xuansong Xie, Xian-Sheng Hua 0001, Hujun Bao, Qixing Huang, Weiwei Xu 0003 |
ACM Multimedia | 9 |
| 2021 | ScPnP: A non-iterative scale compensation solution for PnP problems
Chengzhe Meng, Weiwei Xu 0003 |
Image Vis. Comput. | 2 |
| 2021 | QuanTaichi: a compiler for quantized simulationsabstractHigh-resolution simulations can deliver great visual quality, but they are often limited by available memory, especially on GPUs. We present a compiler for physical simulation that can achieve both high performance and significantly reduced memory costs, by enabling flexible and aggressive quantization. Low-precision ("quantized") numerical data types are used and packed to represent simulation states, leading to reduced memory space and bandwidth consumption. Quantized simulation allows higher resolution simulation with less memory, which is especially attractive on GPUs. Implementing a quantized simulator that has high performance and packs the data tightly for aggressive storage reduction would be extremely labor-intensive and error-prone using a traditional programming language. To make the creation of quantized simulation practical, we have developed a new set of language abstractions and a compilation system. A suite of tailored domain-specific optimizations ensure quantized simulators often run as fast as the full-precision simulators, despite the overhead of encoding-decoding the packed quantized data types. Our programming language and compiler, based on Taichi , allow developers to effortlessly switch between different full-precision and quantized simulators, to explore the full design space of quantization schemes, and ultimately to achieve a good balance between space and precision. The creation of quantized simulation with our system has large benefits in terms of memory consumption and performance, on a variety of hardware, from mobile devices to workstations with high-end GPUs. We can simulate with levels of resolution that were previously only achievable on systems with much more memory, such as multiple GPUs. For example, on a single GPU, we can simulate a Game of Life with 20 billion cells (8× compression per pixel), an Eulerian fluid system with 421 million active voxels (1.6× compression per voxel), and a hybrid Eulerian-Lagrangian elastic object simulation with 235 million particles (1.7× compression per particle). At the same time, quantized simulations create physically plausible results. Our quantization techniques are complementary to existing acceleration approaches of physical simulation: they can be used in combination with these existing approaches, such as sparse data structures, for even higher scalability and performance. Yuanming Hu, Jiafeng Liu, Xuanda Yang, Mingkuan Xu, Ye Kuang, Weiwei Xu 0003, William T. Freeman, Frédo Durand |
ACM Trans. Graph. | 6 |
| 2021 | Scalable image-based indoor scene rendering with reflectionsabstractThis paper proposes a novel scalable image-based rendering (IBR) pipeline for indoor scenes with reflections. We make substantial progress towards three sub-problems in IBR, namely, depth and reflection reconstruction, view selection for temporally coherent view-warping, and smooth rendering refinements. First, we introduce a global-mesh-guided alternating optimization algorithm that robustly extracts a two-layer geometric representation. The front and back layers encode the RGB-D reconstruction and the reflection reconstruction, respectively. This representation minimizes the image composition error under novel views, enabling accurate renderings of reflections. Second, we introduce a novel approach to select adjacent views and compute blending weights for smooth and temporal coherent renderings. The third contribution is a supersampling network with a motion vector rectification module that refines the rendering results to improve the final output's temporal coherence. These three contributions together lead to a novel system that produces highly realistic rendering results with various reflections. The rendering quality outperforms state-of-the-art IBR or neural rendering algorithms considerably. Jiamin Xu, Xiuchao Wu, Zihan Zhu, Qixing Huang, Yin Yang 0002, Hujun Bao, Weiwei Xu 0003 |
ACM Trans. Graph. | 7 |
| 2021 | ART-UP: A Novel Method for Generating Scanning-Robust Aesthetic QR CodesabstractQuick response (QR) codes are usually scanned in different environments, so they must be robust to variations in illumination, scale, coverage, and camera angles. Aesthetic QR codes improve the visual quality, but subtle changes in their appearance may cause scanning failure. In this article, a new method to generate scanning-robust aesthetic QR codes is proposed, which is based on a module-based scanning probability estimation model that can effectively balance the tradeoff between visual quality and scanning robustness. Our method locally adjusts the luminance of each module by estimating the probability of successful sampling. The approach adopts the hierarchical, coarse-to-fine strategy to enhance the visual quality of aesthetic QR codes, which sequentially generate the following three codes: a binary aesthetic QR code, a grayscale aesthetic QR code, and the final color aesthetic QR code. Our approach also can be used to create QR codes with different visual styles by adjusting some initialization parameters. User surveys and decoding experiments were adopted for evaluating our method compared with state-of-the-art algorithms, which indicates that the proposed approach has excellent performance in terms of both visual quality and scanning robustness. Mingliang Xu 0001, Qingfeng Li 0004, Jianwei Niu 0002, Hao Su 0001, Xiting Liu, Weiwei Xu 0003, Pei Lv, Bing Zhou 0003, Yi Yang 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 6 |
| 2021 | Computational Design of Skinned Quad-RobotsabstractWe present a computational design system that assists users to model, optimize, and fabricate quad-robots with soft skins. Our system addresses the challenging task of predicting their physical behavior by fully integrating the multibody dynamics of the mechanical skeleton and the elastic behavior of the soft skin. The developed motion control strategy uses an alternating optimization scheme to avoid expensive full space time-optimization, interleaving space-time optimization for the skeleton, and frame-by-frame optimization for the full dynamics. The output are motor torques to drive the robot to achieve a user prescribed motion trajectory. We also provide a collection of convenient engineering tools and empirical manufacturing guidance to support the fabrication of the designed quad-robot. We validate the feasibility of designs generated with our system through physics simulations and with a physically-fabricated prototype. Xudong Feng, Jiafeng Liu, Huamin Wang 0001, Yin Yang 0002, Hujun Bao, Bernd Bickel, Weiwei Xu 0003 |
IEEE Trans. Vis. Comput. Graph. | 7 |
| 2021 | AgentDress: Realtime Clothing Synthesis for Virtual Agents using Plausible DeformationsabstractWe present a CPU-based real-time cloth animation method for dressing virtual humans of various shapes and poses. Our approach formulates the clothing deformation as a high-dimensional function of body shape parameters and pose parameters. In order to accelerate the computation, our formulation factorizes the clothing deformation into two independent components: the deformation introduced by body pose variation (Clothing Pose Model) and the deformation from body shape variation (Clothing Shape Model). Furthermore, we sample and cluster the poses spanning the entire pose space and use those clusters to efficiently calculate the anchoring points. We also introduce a sensitivity-based distance measurement to both find nearby anchoring points and evaluate their contributions to the final animation. Given a query shape and pose of the virtual agent, we synthesize the resulting clothing deformation by blending the Taylor expansion results of nearby anchoring points. Compared to previous methods, our approach is general and able to add the shape dimension to any clothing pose model. Furthermore, we can animate clothing represented with tens of thousands of vertices at 50+ FPS on a CPU. We also conduct a user evaluation and show that our method can improve a user's perception of dressed virtual agents in an immersive virtual environment (IVE) compared to a realtime linear blend skinning method. Qianwen Chao, Yanzhen Chen, Weiwei Xu 0003, Chen Liu 0012, Dinesh Manocha, Wenxin Sun, Xinran Yao, Xiaogang Jin 0001 |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2020 | AutoRemover: Automatic Object Removal for Autonomous Driving VideosabstractMotivated by the need for photo-realistic simulation in autonomous driving, in this paper we present a video inpainting algorithm AutoRemover, designed specifically for generating street-view videos without any moving objects. In our setup we have two challenges: the first is the shadow, shadows are usually unlabeled but tightly coupled with the moving objects. The second is the large ego-motion in the videos. To deal with shadows, we build up an autonomous driving shadow dataset and design a deep neural network to detect shadows automatically. To deal with large ego-motion, we take advantage of the multi-source data, in particular the 3D data, in autonomous driving. More specifically, the geometric relationship between frames is incorporated into an inpainting deep neural network to produce high-quality structurally consistent video output. Experiments show that our method outperforms other state-of-the-art (SOTA) object removal algorithms, reducing the RMSE by over 19%. Wei Li 0111, Peng Wang 0001, Chenye Guan, Yuhang Song 0003, Baoquan Chen, Weiwei Xu 0003, Ruigang Yang |
AAAI | 9 |
| 2020 | Quasi-Newton Solver for Robust Non-Rigid RegistrationabstractImperfect data (noise, outliers and partial overlap) and high degrees of freedom make non-rigid registration a classical challenging problem in computer vision. Existing methods typically adopt the l_p type robust estimator to regularize the fitting and smoothness, and the proximal operator is used to solve the resulting non-smooth problem. However, the slow convergence of these algorithms limits its wide applications. In this paper, we propose a formulation for robust non-rigid registration based on a globally smooth robust estimator for data fitting and regularization, which can handle outliers and partial overlaps. We apply the majorization-minimization algorithm to the problem, which reduces each iteration to solving a simple least-squares problem with L-BFGS. Extensive experiments demonstrate the effectiveness of our method for non-rigid alignment between two shapes with outliers and partial overlap. with quantitative evaluation showing that it outperforms state-of-the-art methods in terms of registration accuracy and computational speed. The source code is available at https://github.com/Juyong/Fast_RNRR. Yuxin Yao 0001, Bailin Deng, Weiwei Xu 0003, Juyong Zhang |
CVPR | 3 |
| 2020 | EditorialabstractThis special issue contains 28 full papers selected from the Computer Animation and Social Agents 2020 Conference (CASA2020). This conference was founded by the Computer Graphics Society in 1988 in Geneva and is the oldest conference on Computer Animation in the world. It has been held in many countries around the world and in recent years in Beijing, China (2018), Paris, France (2019) and this year in Bournemouth, United Kingdom. Because of the Covid-19 pandemic, this year, the conference will be held online through the Youtube Channel. The best paper award will be announced on the conference website after the conference. We would like to thank the authors for sharing their research findings by submitting papers to CASA2020. We are very grateful to the Program Committee members for reviewing the papers and to all the people who have contributed to the success of CASA2020 in Bournemouth. The conference is organized by Bournemouth University under the guidance of the Computer Graphics Society (CGS). Conference co-chairs Jian Jun Zhang (Bournemouth University, UK) Nadia Magnenat Thalmann (University of Geneva, Switzerland and Nanyang Technological University, Singapore) Program co-chairs Daniel Thalmann (EPFL, Switzerland) Xiaosong Yang (Bournemouth University, UK) Weiwei Xu (Zhejiang University, China) Publicity chair Jian Chang (Bournemouth University, UK) Local chair Feng Tian (Bournemouth University, UK) International program committee Nadine Aburumman, Brunel University, UK Norman Badler, University of Pennsylvania, USA Selim Balcisoy, Sabanci University, Turkey Loic Barthe, IRIT—Université de Toulouse, France Jan Bender, RWTH Aachen University, Germany Raphaëlle Chaine, LIRIS Université Lyon 1, France Jian Chang, Bournemouth University, UK Fred Charles, Bournemouth University, UK Parag Chaudhuri, Indian Institute of Technology, Bombay, India Marc Christie, INRIA, France Justin Dauwels, Nanyang Technological University, Singapore Shujie Deng, King's College London, UK Zhigang Deng, University of Houston, USA Etienne de Sevin, SANPSY University of Bordeaux, France Petros Faloutsos, York University, Canada Christos Gatzidis, Bournemouth University, UK Ugur Gudukbay, Bilkent University, Turkey Shihui Guo, Xiamen University, China Xiaohu Guo, The University of Texas at Dallas, USA James Hahn, George Washington University, USA Carlo Harvey, Birmingham City University, UK Gaoqi He, East China Normal University, China Ying He, Nanyang Technological University, Singapore Kemao Qian, Nanyang Technological University, Singapore Ruizhen Hu, Shenzhen University, China Jinyuan Jia, Tongji University, China Tao Jiang, University of Surrey, UK Xiaogang Jin, Zhejiang University, China Marcelo Kallmann, University of California, Merced, USA Prem Kalra, IIT Delhi, India Dongwann Kang, Seoul National University of Science and Technology, Korea Mubbasir Kapadia, Rutgers University, USA Min H. Kim, Korea Advanced Institute of Science and Technology, Korea Scott King, Texas A&M University—Corpus Christi, USA Taesoo Kwon, Hanyang University, China Sung-Hee Lee, Korea Advanced Institute of Science and Technology, Korea Wonsook Lee, University of Ottawa, Canada Tsai-Yen Li, National Chengchi University, Taiwan Guoliang Luo, East China Jiaotong University, China Chongyang Ma, Snap Inc., USA Anderson Maciel, Universidade Federal do Rio Grande do Sul, Brazil Nadia Magnenat Thalmann, University Of Geneva, Switzerland Shigeo Morishima, Waseda University, Japan Soraia Musse, Pontificia Universidade Catolica do Roi Grande do Sul, PUCRS, Brazil Rahul Narain, Indian Institute of Technology, Delhi, India Junjun Pan, Beihang University, China Nuria Pelechano, Universitat Politècnica de Catalunya, Spain Julien Pettre, INRIA, France Nicolas Pronost, Université Claude Bernard Lyon 1, France Kun Qian, King's College London, UK Craig Schroeder, University of California, Riverside, USA Ari Shapiro, Embody Digital, USA Hubert P. H. Shum, Northumbria University, UK Shinjiro Sueda, Texas A&M University, USA Daniel Thalmann, Ecole Polytechnique Fédérale de Lausanne, Switzerland Feng Tian, Bournemouth University, UK Yiying Tong, Michigan State University, USA Meili Wang, Northwest A&F University, China Zhao Wang, Zhejiang University, China Enhua Wu, University of Macau & ISCAS, China Zhongke Wu, Beijing Normal University, China Weiwei Xu, Zhejiang University, China Yachun Fan, Beijing Normal University, China Bailin Yang, Zhejiang Gongshang University, China Yin Yang, University of New Mexico, USA Xiaosong Yang, Bournemouth University, UK Yuting Ye, Oculus Research, USA Lihua You, Bournemouth University, UK Hongchuan Yu, Bournemouth University, UK Zerrin Yumak, Utrecht University, Netherlands Wenshu Zhang, Cardiff Metropolitan University Jian Zhang, Bournemouth University, UK Jianmin Zheng, Nanyang Technological University, Singapore Jian J. Zhang 0001, Nadia Magnenat-Thalmann, Daniel Thalmann, Xiaosong Yang, Weiwei Xu 0003, Jian Chang 0001, Feng Tian 0009 |
Comput. Animat. Virtual Worlds | 5 |
| 2020 | Medial Elastics: Efficient and Collision-Ready Deformation via Medial Axis TransformabstractWe propose a framework for the interactive simulation of nonlinear deformable objects. The primary feature of our system is the seamless integration of deformable simulation and collision culling, which are often independently handled in existing animation systems. The bridge connecting them is the medial axis transform (MAT), a high-fidelity volumetric approximation of complex 3D shapes. From the physics simulation perspective, MAT leads to an expressive and compact reduced nonlinear model. We employ a semireduced projective dynamics formulation, which well captures high-frequency local deformations of high-resolution models while retaining a low computation cost. Our key observation is that the most compelling (nonlinear) deformable effects are enabled by the local constraints projection, which should not be aggressively reduced, and only apply model reduction at the global stage. From the collision detection (CD)/collision culling (CC) perspective, MAT is geometrically versatile using linear-interpolated spheres (i.e., the so-called medial primitives (MPs)) to approximate the boundary of the input model. The intersection test between two MPs is formulated as a quadratically constrained quadratic program problem. We give an algorithm to solve this problem exactly, which returns the deepest penetration between a pair of intersecting MPs. When coupled with spatial hashing, collision (including self-collision) can be efficiently identified on the GPU within a few milliseconds even for massive simulations. We have tested our system on a variety of geometrically complex and high-resolution deformable objects, and our system produces convincing animations with all of the collisions/self-collisions well handled at an interactive rate. Lei Lan, Ran Luo 0001, Marco Fratarcangeli, Weiwei Xu 0003, Huamin Wang 0001, Xiaohu Guo, Junfeng Yao, Yin Yang 0002 |
ACM Trans. Graph. | 4 |
| 2020 | NNWarp: Neural Network-Based Nonlinear DeformationabstractNNWarp is a highly re-usable and efficient neural network (NN) based nonlinear deformable simulation framework. Unlike other machine learning applications such as image recognition, where different inputs have a uniform and consistent format (e.g., an array of all the pixels in an image), the input for deformable simulation is quite variable, high-dimensional, and parametrization-unfriendly. Consequently, even though the neural network is known for its rich expressivity of nonlinear functions, directly using an NN to reconstruct the force-displacement relation for general deformable simulation is nearly impossible. NNWarp obviates this difficulty by partially restoring the force-displacement relation via warping the nodal displacement simulated using a simplistic constitutive model-the linear elasticity. In other words, NNWarp yields an incremental displacement fix per mesh node based on a simplified (therefore incorrect) simulation result other than synthesizing the unknown displacement directly. We introduce a compact yet effective feature vector including geodesic, potential and digression to sort training pairs of per-node linear and nonlinear displacement. NNWarp is robust under different model shapes and tessellations. With the assistance of deformation substructuring, one NN training is able to handle a wide range of 3D models of various geometries. Thanks to the linear elasticity and its constant system matrix, the underlying simulator only needs to perform one pre-factorized matrix solve at each time step, which allows NNWarp to simulate large models in real time. Ran Luo 0001, Tianjia Shao, Huamin Wang 0001, Weiwei Xu 0003, Xiang Chen 0001, Kun Zhou 0001, Yin Yang 0002 |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2019 | A Late Fusion CNN for Digital MattingabstractThis paper studies the structure of a deep convolutional neural network to predict the foreground alpha matte by taking a single RGB image as input. Our network is fully convolutional with two decoder branches for the foreground and background classification respectively. Then a fusion branch is used to integrate the two classification results which gives rise to alpha values as the soft segmentation result. This design provides more degrees of freedom than a single decoder branch for the network to obtain better alpha values during training. The network can implicitly produce trimaps without user interaction, which is easy to use for novices without expertise in digital matting. Experimental results demonstrate that our network can achieve high-quality alpha mattes for various types of objects and outperform the state-of-the-art CNN-based image matting methods on the human image matting task. Yunke Zhang, Lixue Gong, Lubin Fan, Peiran Ren, Qixing Huang, Hujun Bao, Weiwei Xu 0003 |
CVPR | 7 |
| 2019 | Parametric 3D modeling of a symmetric human body
Yin Chen 0003, Zhan Song, Weiwei Xu 0003, Ralph R. Martin, Zhi-Quan Cheng |
Comput. Graph. | 3 |
| 2019 | A survey on fast simulation of elastic objects
Jin Huang 0001, Jiong Chen 0001, Weiwei Xu 0003, Hujun Bao |
Frontiers Comput. Sci. | 3 |
| 2019 | Disassembling a 3D mechanism for efficient packingabstractAbstract This paper introduces a disassemble‐and‐pack algorithm to disassemble a mechanical 3D model in groups that can be efficiently packed within a box, with the objective of reassembling them easily after delivery. Its key feature is that, mostly, the mechanism can be disassembled at the joint and each part can be an adjusted motion structure based on its joint type. Our system consists of two steps: disassembling the mechanical object into a group set and packing them within a box efficiently. The first step consists in the creation of a hierarchy of possible group set of parts that can be tightly packed within their minimum bounding boxes. Use the breadth‐first search algorithm to traverse the hierarchy of possible group set in order to disconnect the joints and get the group set. In the second step, according to the reverse order of volume, each group in the set is inserted into the specified box. The fact that mechanism disassembly and shape packing are both an NP‐complete problem justifies finding approximated solutions according to efficacy and efficiency. Experimental results show that our approach can really efficiently pack a range of mechanisms from a simple model to complex objects. Xiaoheng Jiang, Ning-Bo Gu, Weiwei Xu 0003, Junxiao Xue, Bing Zhou 0003, Mingliang Xu 0001 |
Comput. Animat. Virtual Worlds | 4 |
| 2019 | Accelerated complex-step finite difference for expedient deformable simulationabstractIn deformable simulation, an important computing task is to calculate the gradient and derivative of the strain energy function in order to infer the corresponding internal force and tangent stiffness matrix. The standard numerical routine is the finite difference method, which evaluates the target function multiple times under a small real-valued perturbation. Unfortunately, the subtractive cancellation prevents us from setting this perturbation sufficiently small, and the regular finite difference is doomed for computing problems requiring a high-accuracy derivative evaluation. In this paper, we graft a new finite difference scheme, namely the complex-step finite difference (CSFD), with physics-based animation. CSFD is based on the complex Taylor series expansion, which avoids subtractions in first-order derivative approximation. As a result, one can use a very small perturbation to calculate the numerical derivative that is as accurate as its analytic counterpart. We accelerate the original CSFD method so that it is also as efficient as the analytic derivative. This is achieved by discarding high-order error terms, decoupling real and imaginary calculations, replacing costly functions based on the theory of equivalent infinitesimal, and isolating the propagation of the perturbation in composite/nesting functions. CSFD can be further augmented with multicomplex Taylor expansion and Cauchy-Riemann formula to handle higher-order derivatives and tensor-valued functions. We demonstrate the accuracy, convenience, and efficiency of this new numerical routine in the context of deformable simulation - one can easily deploy a robust simulator for general hyperelastic materials, including user-crafted ones to cater specific needs in different applications. Higher-order derivatives of the energy can be readily computed to construct modal derivative bases for reduced real-time simulation. Inverse simulation problems can also be conveniently solved using gradient/Hessian-based optimization procedures. Ran Luo 0001, Weiwei Xu 0003, Tianjia Shao, Yin Yang 0002 |
ACM Trans. Graph. | 2 |
| 2019 | Direct design to stress mapping for cellular structuresabstractThis paper aims to instantly predict within any accuracy the stress distribution of cellular structures under parametric design, including the shapes or distributions of the cell geometries, or the magnitudes of external loadings. A classical model reduction technique has to balance the simulation accuracy and interaction speed, and has difficulty achieving this goal. We achieve this by computing offline a design-to-stress mapping that ultimately expresses the stress distribution as an explicit function in terms of its design parameters. The mapping is determined as a solution to an extended finite element analysis problem in a high-dimension space, including both the spatial coordinates and the design parameters. The well-known curse of dimensionality intrinsic to the high-dimension problem is (partly) resolved through a spatial separation using two main techniques. First, the target mapping takes a reduced form as a sum of the products of separated one-variable functions, extending the proper generalized decomposition technique. Second, the simulation problem in a varied computation domain is reformulated as that in a fixed-domain, taking an integration function as the sum of the products of separated one-variable functions, in combination with high-order singular value decomposition. Extensive 2D and 3D examples are shown to demonstrate the approach’s performance. Liangchao Zhu, Weiwei Xu 0003 |
Vis. Informatics | 3 |
| 2019 | Survey of 3D modeling using depth camerasabstractThree-dimensional (3D) modeling is an important topic in computer graphics and computer vision. In recent years, the introduction of consumer-grade depth cameras has resulted in profound advances in 3D modeling. Starting with the basic data structure, this survey reviews the latest developments of 3D modeling based on depth cameras, including research works on camera tracking, 3D object and scene reconstruction, and high-quality texture reconstruction. We also discuss the future work and possible solutions for 3D modeling based on the depth camera. Hantong Xu, Jiamin Xu, Weiwei Xu 0003 |
Virtual Real. Intell. Hardw. | 3 |
| 2018 | Stress-aware large-scale mesh editing using a domain-decomposed multigrid solver
Weiwei Xu 0003, Yin Yang 0002, Yiduo Wang 0004, Kun Zhou 0001 |
Comput. Aided Geom. Des. | 1 |
| 2018 | Online Global Non-rigid Registration for 3D Object Reconstruction Using Consumer-level Depth CamerasabstractAbstract We investigate how to obtain high‐quality 360‐degree 3D reconstructions of small objects using consumer‐level depth cameras. For many homeware objects such as shoes and toys with dimensions around 0.06 – 0.4 meters, their whole projections, in the hand‐held scanning process, occupy fewer than 20% pixels of the camera's image. We observe that existing 3D reconstruction algorithms like KinectFusion and other similar methods often fail in such cases even under the close‐range depth setting. To achieve high‐quality 3D object reconstruction results at this scale, our algorithm relies on an online global non‐rigid registration, where embedded deformation graph is employed to handle the drifting of camera tracking and the possible nonlinear distortion in the captured depth data. We perform an automatic target object extraction from RGBD frames to remove the unrelated depth data so that the registration algorithm can focus on minimizing the geometric and photogrammetric distances of the RGBD data of target objects. Our algorithm is implemented using CUDA for a fast non‐rigid registration. The experimental results show that the proposed method can reconstruct high‐quality 3D shapes of various small objects with textures. Jiamin Xu, Weiwei Xu 0003, Yin Yang 0002, Zhigang Deng 0001, Hujun Bao |
Comput. Graph. Forum | 2 |
| 2018 | Efficient voxelization using projected optimal scanline
Steven Garcia, Weiwei Xu 0003, Tianjia Shao, Yin Yang 0002 |
Graph. Model. | 3 |
| 2018 | Automatic unpaired shape deformation transferabstractTransferring deformation from a source shape to a target shape is a very useful technique in computer graphics. State-of-the-art deformation transfer methods require either point-wise correspondences between source and target shapes, or pairs of deformed source and target shapes with corresponding deformations. However, in most cases, such correspondences are not available and cannot be reliably established using an automatic algorithm. Therefore, substantial user effort is needed to label the correspondences or to obtain and specify such shape sets. In this work, we propose a novel approach to automatic deformation transfer between two unpaired shape sets without correspondences. 3D deformation is represented in a high-dimensional space. To obtain a more compact and effective representation, two convolutional variational autoencoders are learned to encode source and target shapes to their latent spaces. We exploit a Generative Adversarial Network (GAN) to map deformed source shapes to deformed target shapes, both in the latent spaces, which ensures the obtained shapes from the mapping are indistinguishable from the target shapes. This is still an under-constrained problem, so we further utilize a reverse mapping from target shapes to source shapes and incorporate cycle consistency loss, i.e. applying both mappings should reverse to the input shape. This VAE-Cycle GAN (VC-GAN) architecture is used to build a reliable mapping between shape spaces. Finally, a similarity constraint is employed to ensure the mapping is consistent with visual similarity, achieved by learning a similarity neural network that takes the embedding vectors from the source and target latent spaces and predicts the light field distance between the corresponding shapes. Experimental results show that our fully automatic method is able to obtain high-quality deformation transfer results with unpaired data sets, comparable or better than existing methods where strict correspondences are required. Lin Gao 0004, Jie Yang 0038, Yi-Ling Qiao, Yukun Lai, Paul L. Rosin, Weiwei Xu 0003, Shihong Xia |
ACM Trans. Graph. | 6 |
| 2018 | Physics-Based Quadratic Deformation Using Elastic WeightingabstractThis paper presents a spatial reduction framework for simulating nonlinear deformable objects interactively. This reduced model is built using a small number of overlapping quadratic domains as we notice that incorporating high-order degrees of freedom (DOFs) is important for the simulation quality. Departing from existing multi-domain methods in graphics, our method interprets deformed shapes as blended quadratic transformations from nearby domains. Doing so avoids expensive safeguards against the domain coupling and improves the numerical robustness under large deformations. We present an algorithm that efficiently computes weight functions for reduced DOFs in a physics-aware manner. Inspired by the well-known multi-weight enveloping technique, our framework also allows subspace tweaking based on a few representative deformation poses. Such elastic weighting mechanism significantly extends the expressivity of the reduced model with light-weight computational efforts. Our simulator is versatile and can be well interfaced with many existing techniques. It also supports local DOF adaption to incorporate novel deformations (i.e., induced by the collision). The proposed algorithm complements state-of-the-art model reduction and domain decomposition methods by seeking for good trade-offs among animation quality, numerical robustness, pre-computation complexity, and simulation efficiency from an alternative perspective. Ran Luo 0001, Weiwei Xu 0003, Huamin Wang 0001, Kun Zhou 0001, Yin Yang 0002 |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2017 | Acoustic VR in the mouth: A real-time speech-driven visual tongue systemabstractWe propose an acoustic-VR system that converts acoustic signals of human language (Chinese) to realistic 3D tongue animation sequences in real time. It is known that directly capturing the 3D geometry of the tongue at a frame rate that matches the tongue's swift movement during the language production is challenging. This difficulty is handled by utilizing the electromagnetic articulography (EMA) sensor as the intermediate medium linking the acoustic data to the simulated virtual reality. We leverage Deep Neural Networks to train a model that maps the input acoustic signals to the positional information of pre-defined EMA sensors based on 1,108 utterances. Afterwards, we develop a novel reduced physics-based dynamics model for simulating the tongue's motion. Unlike the existing methods, our deformable model is nonlinear, volume-preserving, and accommodates collision between the tongue and the oral cavity (mostly with the jaw). The tongue's deformation could be highly localized which imposes extra difficulties for existing spectral model reduction methods. Alternatively, we adopt a spatial reduction method that allows an expressive subspace representation of the tongue's deformation. We systematically evaluate the simulated tongue shapes with real-world shapes acquired by MRI/CT. Our experiment demonstrates that the proposed system is able to deliver a realistic visual tongue animation corresponding to a user's speech signal. Ran Luo 0001, Qiang Fang 0003, Jianguo Wei, Wenhuan Lu, Weiwei Xu 0003, Yin Yang 0002 |
VR | 5 |
| 2017 | Modeling, Evaluation and Optimization of Interlocking Shell PiecesabstractAbstract While the 3D printing technology has become increasingly popular in recent years, it suffers from two critical limitations: expensive printing material and long printing time. An effective solution is to hollow the 3D model into a shell and print the shell by parts. Unfortunately, making shell pieces tightly assembled and easy to disassemble seem to be two contradictory conditions, and there exists no easy way to satisfy them at the same time yet. In this paper, we present a computational system to design an interlocking structure of a partitioned shell model, which uses only male and female connectors to lock shell pieces in the assembled configuration. Given a mesh segmentation input, our system automatically finds an optimal installation plan specifying both the installation order and the installation directions of the pieces, and then builds the models of the shell pieces using optimized shell thickness and connector sizes. To find the optimal installation plan, we develop simulation‐based and data‐driven metrics, and we incorporate them into an optimal plan search algorithm with fast pruning and local optimization strategies. The whole system is automatic, except for the shape design of the key piece. The interlocking structure does not introduce new gaps on the outer surface, which would become noticeable inevitably due to limited printer precision. Our experiment shows that the assembled object is strong against separation, yet still easy to disassemble. Miaojun Yao, Weiwei Xu 0003, Huamin Wang 0001 |
Comput. Graph. Forum | 3 |
| 2017 | Stress-Constrained Thickness Optimization for Shell Object FabricationabstractAbstract We present an approach to fabricate shell objects with thickness parameters, which are computed to maintain the user‐specified structural stability. Given a boundary surface and user‐specified external forces, we optimize the thickness parameters according to stress constraints to extrude the surface. Our approach mainly consists of two technical components: First, we develop a patch‐based shell simulation technique to efficiently support the static simulation of extruded shell objects using finite element methods. Second, we analytically compute the derivative of stress required in the sensitivity analysis technique to turn the optimization into a sequential linear programming problem. Experimental results demonstrate that our approach can optimize the thickness parameters for arbitrary surfaces in a few minutes and well predict the physical properties, such as the deformation and stress of the fabricated object. Haiming Zhao, Weiwei Xu 0003, Kun Zhou 0001, Yin Yang 0002, Xiaogang Jin 0001, Hongzhi Wu |
Comput. Graph. Forum | 2 |
| 2017 | Mechanical Assembly Packing Problem Using Joint Constraints
Mingliang Xu 0001, Ning-Bo Gu, Weiwei Xu 0003, Junxiao Xue, Bing Zhou 0003 |
J. Comput. Sci. Technol. | 3 |
| 2016 | Fast Nearest Neighbor Search in the Hamming Space
Zhansheng Jiang, Lingxi Xie, Xiaotie Deng, Weiwei Xu 0003, Jingdong Wang 0001 |
MMM (1) | 4 |
| 2016 | Make it swing: Fabricating personalized roly-poly toys
Haiming Zhao, Chengkuan Hong, Juncong Lin, Xiaogang Jin 0001, Weiwei Xu 0003 |
Comput. Aided Geom. Des. | 5 |
| 2016 | View-Aware Image Object Compositing and Synthesis from Multiple Sources
Xiang Chen 0001, Weiwei Xu 0003, Sai-Kit Yeung, Kun Zhou 0001 |
J. Comput. Sci. Technol. | 2 |
| 2016 | Error resilience video coding parameters and mechanisms selection with End-to-End rate-distortion analysis at frame level
Weiwei Xu 0003, Yaowu Chen |
Multim. Tools Appl. | 1 |
| 2016 | Interactive mechanism modeling from multi-view imagesabstractIn this paper, we present an interactive system for mechanism modeling from multi-view images. Its key feature is that the generated 3D mechanism models contain not only geometric shapes but also internal motion structures: they can be directly animated through kinematic simulation. Our system consists of two steps: interactive 3D modeling and stochastic motion parameter estimation. At the 3D modeling step, our system is designed to integrate the sparse 3D points reconstructed from multi-view images and a sketching interface to achieve accurate 3D modeling of a mechanism. To recover the motion parameters, we record a video clip of the mechanism motion and adopt stochastic optimization to recover its motion parameters by edge matching. Experimental results show that our system can achieve the 3D modeling of a range of mechanisms from simple mechanical toys to complex mechanism objects. Weiwei Xu 0003, Zhigang Deng 0001, Yin Yang 0002, Kun Zhou 0001 |
ACM Trans. Graph. | 3 |
| 2016 | All-hex meshing using closed-form induced polycubeabstractThe polycube-based hexahedralization methods are robust to generate all-hex meshes without internal singularities. They avoid the difficulty to control the global singularity structure for a valid hexahedralization in frame-field based methods. To thoroughly utilize this advantage, we propose to use a frame field without internal singularities to guide the polycube construction. Theoretically, our method extends the vector fields associated with the polycube from exact forms to closed forms, which are curl free everywhere but may be not globally integrable. The closed forms give additional degrees of freedom to deal with the topological structure of high-genus models, and also provide better initial axis alignment for subsequent polycube generation. We demonstrate the advantages of our method on various models, ranging from genus-zero models to high-genus ones, and from single-boundary models to multiple-boundary ones. Xianzhong Fang, Weiwei Xu 0003, Hujun Bao, Jin Huang 0001 |
ACM Trans. Graph. | 2 |
| 2015 | Lightweight wrinkle synthesis for 3D facial modeling and animation
Jun Li 0042, Weiwei Xu 0003, Zhi-Quan Cheng, Kai Xu 0004, Reinhard Klein |
Comput. Aided Des. | 2 |
| 2015 | Agile structural analysis for fabrication-aware shape editing
Weiwei Xu 0003, Yin Yang 0002, Xiaohu Guo, Kun Zhou 0001 |
Comput. Aided Geom. Des. | 2 |
| 2015 | A Suggestive Interface for Sketch-based Character PosingabstractWe present a user-friendly suggestive interface for sketch-based character posing. Our interface provides suggestive information on the sketching canvas in succession by combining image retrieval technique with 3D character posing, while the user is drawing. The system highlights the canvas region where the user should draw on and constrains the user's sketches in a reasonable solution space. This is based on an efficient image descriptor, which is used to measure the distance between the user's sketch and 2D views of 3D poses. In order to achieve faster query response, local sensitive hashing is involved in our system. In addition, sampling-based optimization algorithm is adopted to synthesize and optimize the retrieved 3D pose to match the user's sketches the best. Experiments show that our interface can provide smooth suggestive information to improve the reality of sketching poses and shorten the time required for 3D posing. Pei Lv, Pengjie Wang 0001, Weiwei Xu 0003, Jinxiang Chai |
Comput. Graph. Forum | 3 |
| 2015 | Interactive design and simulation of tubular supporting structure
Ran Luo 0001, Lifeng Zhu, Weiwei Xu 0003, Patrick Gage Kelley, Vanessa Svihla, Yin Yang 0002 |
Graph. Model. | 3 |
| 2015 | Integrating 3D structure into traffic scene understanding with RGB-D data
Yingjie Xia, Weiwei Xu 0003, Xingmin Shi, Kuang Mao |
Neurocomputing | 2 |
| 2015 | Recognizing multi-view objects with occlusions using a deep architecture
Yingjie Xia, Weiwei Xu 0003, Zhenyu Shan, Yuncai Liu |
Inf. Sci. | 3 |
| 2015 | Boundary-dominant flower blooming simulationabstractAbstract This paper presents a new physics‐based simulation method for flower blossom, which is based on biological observations that flower opening is usually driven by a boundary‐dominant morphological transition in a curved petal. We use an elastic triangular mesh representing a flower petal and adopt in‐plane expansion to induce global bending. Out‐of‐plane curl plays an auxiliary role in reducing the curvatures of cross‐sections. We also propose to adapt semi‐implicit Euler time integrator for fast simulation results, which has intrinsic damping and at least one order precision. Our system allows users to control the blossoming process by simply specifying a growth curve, which is easy to design because of the boundary‐dominant property. Experimental results show that our physics‐based system runs faster and generates more realistic and convincing blossom results than the existing simulation methods. Copyright © 2015 John Wiley & Sons, Ltd. Jianfang Li 0001, Weiwei Xu 0003, Haiyi Liang |
Comput. Animat. Virtual Worlds | 3 |
| 2015 | Expediting precomputation for reduced deformable simulationabstractModel reduction has popularized itself for simulating elastic deformation for graphics applications. While these techniques enjoy orders-of-magnitude speedups at runtime simulation, the efficiency of precomputing reduced subspaces remains largely over-looked. We present a complete system of precomputation pipeline as a faster alternative to the classic linear and nonlinear modal analysis. We identify three bottlenecks in the traditional model reduction precomputation, namely modal matrix construction, cubature training, and training dataset generation, and accelerate each of them. Even with complex deformable models, our method has achieved orders-of-magnitude speedups over the traditional precomputation steps, while retaining comparable runtime simulation quality. Yin Yang 0002, Dingzeyu Li, Weiwei Xu 0003, Yuan Tian 0002, Changxi Zheng |
ACM Trans. Graph. | 3 |
| 2015 | Online Structure Analysis for Real-Time Indoor Scene ReconstructionabstractWe propose a real-time approach for indoor scene reconstruction. It is capable of producing a ready-to-use 3D geometric model even while the user is still scanning the environment with a consumer depth camera. Our approach features explicit representations of planar regions and nonplanar objects extracted from the noisy feed of the depth camera, via an online structure analysis on the dynamic, incomplete data. The structural information is incorporated into the volumetric representation of the scene, resulting in a seamless integration with KinectFusion's global data structure and an efficient implementation of the whole reconstruction process. Moreover, heuristics based on rectilinear shapes in typical indoor scenes effectively eliminate camera tracking drift and further improve reconstruction accuracy. The instantaneous feedback enabled by our on-the-fly structure analysis, including repeated object recognition, allows the user to selectively scan the scene and produce high-fidelity large-scale models efficiently. We demonstrate the capability of our system with real-life examples. Weiwei Xu 0003, Yiying Tong, Kun Zhou 0001 |
ACM Trans. Graph. | 2 |
| 2014 | Automatic 3D Indoor Scene Updating with RGBD CamerasabstractAbstract Since indoor scenes are frequently changed in daily life, such as re‐layout of furniture, the 3D reconstructions for them should be flexible and easy to update. We present an automatic 3D scene update algorithm to indoor scenes by capturing scene variation with RGBD cameras. We assume an initial scene has been reconstructed in advance in manual or other semi‐automatic way before the change, and automatically update the reconstruction according to the newly captured RGBD images of the real scene update. It starts with an automatic segmentation process without manual interaction, which benefits from accurate labeling training from the initial 3D scene. After the segmentation, objects captured by RGBD camera are extracted to form a local updated scene. We formulate an optimization problem to compare to the initial scene to locate moved objects. The moved objects are then integrated with static objects in the initial scene to generate a new 3D scene. We demonstrate the efficiency and robustness of our approach by updating the 3D scene of several real‐world scenes. Zhenbao Liu, Sicong Tang, Weiwei Xu 0003, Shuhui Bu, Junwei Han 0001, Kun Zhou 0001 |
Comput. Graph. Forum | 3 |
| 2014 | Transductive 3D Shape Segmentation using Sparse ReconstructionabstractAbstract We propose a transductive shape segmentation algorithm, which can transfer prior segmentation results in database to new shapes without explicitly specification of prior category information. Our method first partitions an input shape into a set of segmentations as a data preparation, and then a linear integer programming algorithm is used to select segments from them to form the final optimal segmentation. The key idea is to maximize the segment similarity between the segments in the input shape and the segments in database, where the segment similarity is computed through sparse reconstruction error. The segment‐level similarity enables to handle a large amount of shapes with significant topology or shape variations with a small set of segmented example shapes. Experimental results show that our algorithm can generate high quality segmentation and semantic labeling results in the Princeton segmentation benchmark. Weiwei Xu 0003, Zhouxu Shi, Kun Zhou 0001, Jingdong Wang 0001, JinRong Wang 0002, Zhenming Yuan |
Comput. Graph. Forum | 1 |
| 2014 | An asymptotic numerical method for inverse elastic shape designabstractInverse shape design for elastic objects greatly eases the design efforts by letting users focus on desired target shapes without thinking about elastic deformations. Solving this problem using classic iterative methods (e.g., Newton-Raphson methods), however, often suffers from slow convergence toward a desired solution. In this paper, we propose an asymptotic numerical method that exploits the underlying mathematical structure of specific nonlinear material models, and thus runs orders of magnitude faster than traditional Newton-type methods. We apply this method to compute rest shapes for elastic fabrication, where the rest shape of an elastic object is computed such that after physical fabrication the real object deforms into a desired shape. We illustrate the performance and robustness of our method through a series of elastic fabrication experiments. Xiang Chen 0001, Changxi Zheng, Weiwei Xu 0003, Kun Zhou 0001 |
ACM Trans. Graph. | 3 |
| 2014 | Imagining the unseen: stability-based cuboid arrangements for scene understandingabstractMissing data due to occlusion is a key challenge in 3D acquisition, particularly in cluttered man-made scenes. Such partial information about the scenes limits our ability to analyze and understand them. In this work we abstract such environments as collections of cuboids and hallucinate geometry in the occluded regions by globally analyzing the physical stability of the resultant arrangements of the cuboids. Our algorithm extrapolates the cuboids into the un-seen regions to infer both their corresponding geometric attributes (e.g., size, orientation) and how the cuboids topologically interact with each other (e.g., touch or fixed). The resultant arrangement provides an abstraction for the underlying structure of the scene that can then be used for a range of common geometry processing tasks. We evaluate our algorithm on a large number of test scenes with varying complexity, validate the results on existing benchmark datasets, and demonstrate the use of the recovered cuboid-based structures towards object retrieval, scene completion, etc. Tianjia Shao, Áron Monszpart, Youyi Zheng, Bongjin Koo, Weiwei Xu 0003, Kun Zhou 0001, Niloy J. Mitra |
ACM Trans. Graph. | 5 |
| 2014 | Sensitivity-optimized rigging for example-based real-time clothing synthesisabstractWe present a real-time solution for generating detailed clothing deformations from pre-computed clothing shape examples. Given an input pose, it synthesizes a clothing deformation by blending skinned clothing deformations of nearby examples controlled by the body skeleton. Observing that cloth deformation can be well modeled with sensitivity analysis driven by the underlying skeleton, we introduce a sensitivity based method to construct a pose-dependent rigging solution from sparse examples. We also develop a sensitivity based blending scheme to find nearby examples for the input pose and evaluate their contributions to the result. Finally, we propose a stochastic optimization based greedy scheme for sampling the pose space and generating example clothing shapes. Our solution is fast, compact and can generate realistic clothing animation results for various kinds of clothes in real time. Weiwei Xu 0003, Nobuyuki Umetani, Qianwen Chao, Jie Mao, Xiaogang Jin 0001, Xin Tong 0001 |
ACM Trans. Graph. | 1 |
| 2013 | Supervised Kernel Descriptors for Visual RecognitionabstractIn visual recognition tasks, the design of low level image feature representation is fundamental. The advent of local patch features from pixel attributes such as SIFT and LBP, has precipitated dramatic progresses. Recently, a kernel view of these features, called kernel descriptors (KDES), generalizes the feature design in an unsupervised fashion and yields impressive results. In this paper, we present a supervised framework to embed the image level label information into the design of patch level kernel descriptors, which we call supervised kernel descriptors (SKDES). Specifically, we adopt the broadly applied bag-of-words (BOW) image classification pipeline and a large margin criterion to learn the low-level patch representation, which makes the patch features much more compact and achieve better discriminative ability than KDES. With this method, we achieve competitive results over several public datasets comparing with state-of-the-art methods. Peng Wang 0001, Jingdong Wang 0001, Weiwei Xu 0003, Hongbin Zha, Shipeng Li 0001 |
CVPR | 4 |
| 2013 | As-Rigid-AsPossible Distance Field MetamorphosisabstractAbstract Widely used for morphing between objects with arbitrary topology, distance field interpolation (DFI) handles topological transition naturally without the need for correspondence or remeshing, unlike surface‐based interpolation approaches. However, lack of correspondence in DFI also leads to ineffective control over the morphing process. In particular, unless the user specifies a dense set of landmarks, it is not even possible to measure the distortion of intermediate shapes during interpolation, let alone control it. To remedy such issues, we introduce an approach for establishing correspondence between the interior of two arbitrary objects, formulated as an optimal mass transport problem with a sparse set of landmarks. This correspondence enables us to compute non‐rigid warping functions that better align the source and target objects as well as to incorporate local rigidity constraints to perform as‐rigid‐aspossible DFI. We demonstrate how our approach helps achieve flexible morphing results with a small number of landmarks. Yanlin Weng, Menglei Chai, Weiwei Xu 0003, Yiying Tong, Kun Zhou 0001 |
Comput. Graph. Forum | 3 |
| 2013 | Interpreting concept sketchesabstractConcept sketches are popularly used by designers to convey pose and function of products. Understanding such sketches, however, requires special skills to form a mental 3D representation of the product geometry by linking parts across the different sketches and imagining the intermediate object configurations. Hence, the sketches can remain inaccessible to many, especially non-designers. We present a system to facilitate easy interpretation and exploration of concept sketches. Starting from crudely specified incomplete geometry, often inconsistent across the different views, we propose a globally-coupled analysis to extract part correspondence and inter-part junction information that best explain the different sketch views. The user can then interactively explore the abstracted object to gain better understanding of the product functions. Our key technical contribution is performing shape analysis without access to any coherent 3D geometric model by reasoning in the space of inter-part relations. We evaluate our system on various concept sketches obtained from popular product design books and websites. Tianjia Shao, Wilmot Li, Kun Zhou 0001, Weiwei Xu 0003, Baining Guo, Niloy J. Mitra |
ACM Trans. Graph. | 4 |
| 2013 | Boundary-Aware Multidomain Subspace DeformationabstractIn this paper, we propose a novel framework for multidomain subspace deformation using node-wise corotational elasticity. With the proper construction of subspaces based on the knowledge of the boundary deformation, we can use the Lagrange multiplier technique to impose coupling constraints at the boundary without overconstraining. In our deformation algorithm, the number of constraint equations to couple two neighboring domains is not related to the number of the nodes on the boundary but is the same as the number of the selected boundary deformation modes. The crack artifact is not present in our simulation result, and the domain decomposition with loops can be easily handled. Experimental results show that the single-core implementation of our algorithm can achieve real-time performance in simulating deformable objects with around quarter million tetrahedral elements. Yin Yang 0002, Weiwei Xu 0003, Xiaohu Guo, Kun Zhou 0001, Baining Guo |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2012 | All-hex meshing using singularity-restricted fieldabstractDecomposing a volume into high-quality hexahedral cells is a challenging task in geometric modeling and computational geometry. Inspired by the use of cross field in quad meshing and the CubeCover approach in hex meshing, we present a complete all-hex meshing framework based on singularity-restricted field that is essential to induce a valid all-hex structure. Given a volume represented by a tetrahedral mesh, we first compute a boundary-aligned 3D frame field inside it, then convert the frame field to be singularity-restricted by our effective topological operations. In our all-hex meshing framework, we apply the CubeCover method to achieve the volume parametrization. For reducing degenerate elements appearing in the volume parametrization, we also propose novel tetrahedral split operations to preprocess singularity-restricted frame fields. Experimental results show that our algorithm generates high-quality all-hex meshes from a variety of 3D volumes robustly and efficiently. Yang Liu 0014, Weiwei Xu 0003, Wenping Wang 0001, Baining Guo |
ACM Trans. Graph. | 3 |
| 2012 | An interactive approach to semantic modeling of indoor scenes with an RGBD cameraabstractWe present an interactive approach to semantic modeling of indoor scenes with a consumer-level RGBD camera. Using our approach, the user first takes an RGBD image of an indoor scene, which is automatically segmented into a set of regions with semantic labels. If the segmentation is not satisfactory, the user can draw some strokes to guide the algorithm to achieve better results. After the segmentation is finished, the depth data of each semantic region is used to retrieve a matching 3D model from a database. Each model is then transformed according to the image depth to yield the scene. For large scenes where a single image can only cover one part of the scene, the user can take multiple images to construct other parts of the scene. The 3D models built for all images are then transformed and unified into a complete scene. We demonstrate the efficiency and robustness of our approach by modeling several real-world scenes. Tianjia Shao, Weiwei Xu 0003, Kun Zhou 0001, Jingdong Wang 0001, Dongping Li, Baining Guo |
ACM Trans. Graph. | 2 |
| 2012 | Diffusion curve textures for resolution independent texture mappingabstractWe introduce a vector representation called diffusion curve textures for mapping diffusion curve images (DCI) onto arbitrary surfaces. In contrast to the original implicit representation of DCIs [Orzan et al. 2008], where determining a single texture value requires iterative computation of the entire DCI via the Poisson equation, diffusion curve textures provide an explicit representation from which the texture value at any point can be solved directly, while preserving the compactness and resolution independence of diffusion curves. This is achieved through a formulation of the DCI diffusion process in terms of Green's functions. This formulation furthermore allows the texture value of any rectangular region (e.g. pixel area) to be solved in closed form, which facilitates anti-aliasing. We develop a GPU algorithm that renders anti-aliased diffusion curve textures in real time, and demonstrate the effectiveness of this method through high quality renderings with detailed control curves and color variations. Xin Sun 0014, Guofu Xie, Yue Dong 0001, Stephen Lin 0001, Weiwei Xu 0003, Xin Tong 0001, Baining Guo |
ACM Trans. Graph. | 5 |
| 2012 | Motion-guided mechanical toy modelingabstractWe introduce a new method to synthesize mechanical toys solely from the motion of their features. The designer specifies the geometry and a time-varying rotation and translation of each rigid feature component. Our algorithm automatically generates a mechanism assembly located in a box below the feature base that produces the specified motion. Parts in the assembly are selected from a parameterized set including belt-pulleys, gears, crank-sliders, quick-returns, and various cams (snail, ellipse, and double-ellipse). Positions and parameters for these parts are optimized to generate the specified motion, minimize a simple measure of complexity, and yield a well-distributed layout of parts over the driving axes. Our solution uses a special initialization procedure followed by simulated annealing to efficiently search the complex configuration space for an optimal assembly. Lifeng Zhu, Weiwei Xu 0003, John M. Snyder, Yang Liu 0014, Baining Guo |
ACM Trans. Graph. | 2 |
| 2011 | Discriminative Sketch-based 3D Model Retrieval via Robust Shape MatchingabstractAbstract We propose a sketch‐based 3D shape retrieval system that is substantially more discriminative and robust than existing systems, especially for complex models. The power of our system comes from a combination of a contour‐based 2D shape representation and a robust sampling‐based shape matching scheme. They are defined over discriminative local features and applicable for partial sketches; robust to noise and distortions in hand drawings; and consistent when strokes are added progressively. Our robust shape matching, however, requires dense sampling and registration and incurs a high computational cost. We thus devise critical acceleration methods to achieve interactive performance: precomputing kNN graphs that record transformations between neighboring contour images and enable fast online shape alignment; pruning sampling and shape registration strategically and hierarchically; and parallelizing shape matching on multi‐core platforms or GPUs. We demonstrate the effectiveness of our system through various experiments, comparisons, and user studies. Tianjia Shao, Weiwei Xu 0003, KangKang Yin, Jingdong Wang 0001, Kun Zhou 0001, Baining Guo |
Comput. Graph. Forum | 2 |
| 2011 | New Technique: Sketch-based rotation editingabstractWe present a sketch-based rotation editing system for enriching rotational motion in keyframe animations. Given a set of keyframe orientations of a rigid object, the user first edits its angular velocity trajectory by sketching curves, and then the system computes the altered rotational motion by solving a variational curve fitting problem. The solved rotational motion not only satisfies the orientation constraints at the keyframes, but also fits well the user-specified angular velocity trajectory. Our system is simple and easy to use. We demonstrate its usefulness by adding interesting and realistic rotational details to several keyframe animations. Weiwei Xu 0003, Yizhou Yu, Yanlin Weng |
J. Zhejiang Univ. Sci. C | 2 |
| 2011 | General planar quadrilateral mesh design using conjugate direction fieldabstractWe present a novel method to approximate a freeform shape with a planar quadrilateral (PQ) mesh for modeling architectural glass structures. Our method is based on the study of conjugate direction fields (CDF) which allow the presence of ±κ/4(κ ε Z) singularities. Starting with a triangle discretization of a freeform shape, we first compute an as smooth as possible conjugate direction field satisfying the user's directional and angular constraints, then apply mixed-integer quadrangulation and planarization techniques to generate a PQ mesh which approximates the input shape faithfully. We demonstrate that our method is effective and robust on various 3D models. Yang Liu 0014, Weiwei Xu 0003, Lifeng Zhu, Baining Guo, Falai Chen |
ACM Trans. Graph. | 2 |
| 2010 | Deformation Transfer to Multi-Component ObjectsabstractAbstract We present a simple and effective algorithm to transfer deformation between surface meshes with multiple components. The algorithm automatically computes spatial relationships between components of the target object, builds correspondences between source and target, and finally transfers deformation of the source onto the target while preserving cohesion between the target's components. We demonstrate the versatility of our approach on various complex models. Kun Zhou 0001, Weiwei Xu 0003, Yiying Tong, Mathieu Desbrun |
Comput. Graph. Forum | 2 |
| 2010 | Sampling-based contact-rich motion controlabstractHuman motions are the product of internal and external forces, but these forces are very difficult to measure in a general setting. Given a motion capture trajectory, we propose a method to reconstruct its open-loop control and the implicit contact forces. The method employs a strategy based on randomized sampling of the control within user-specified bounds, coupled with forward dynamics simulation. Sampling-based techniques are well suited to this task because of their lack of dependence on derivatives, which are difficult to estimate in contact-rich scenarios. They are also easy to parallelize, which we exploit in our implementation on a compute cluster. We demonstrate reconstruction of a diverse set of captured motions, including walking, running, and contact rich tasks such as rolls and kip-up jumps. We further show how the method can be applied to physically based motion transformation and retargeting, physically plausible motion variations, and reference-trajectory-free idling motions. Alongside the successes, we point out a number of limitations and directions for future work. Libin Liu 0002, KangKang Yin, Michiel van de Panne, Tianjia Shao, Weiwei Xu 0003 |
ACM Trans. Graph. | 5 |
| 2009 | Gradient Domain Mesh Deformation - A Survey
Weiwei Xu 0003, Kun Zhou 0001 |
J. Comput. Sci. Technol. | 1 |
| 2009 | Joint-aware manipulation of deformable modelsabstractComplex mesh models of man-made objects often consist of multiple components connected by various types of joints. We propose a joint-aware deformation framework that supports the direct manipulation of an arbitrary mix of rigid and deformable components. First we apply slippable motion analysis to automatically detect multiple types of joint constraints that are implicit in model geometry. For single-component geometry or models with disconnected components, we support user-defined virtual joints. Then we integrate manipulation handle constraints, multiple components, joint constraints, joint limits, and deformation energies into a single volumetric-cell-based space deformation problem. An iterative, parallelized Gauss-Newton solver is used to solve the resulting nonlinear optimization. Interactive deformable manipulation is demonstrated on a variety of geometric models while automatically respecting their multi-component nature and the natural behavior of their joints. Weiwei Xu 0003, KangKang Yin, Kun Zhou 0001, Michiel van de Panne, Falai Chen, Baining Guo |
ACM Trans. Graph. | 1 |
| 2007 | Gradient domain editing of deforming mesh sequencesabstractMany graphics applications, including computer games and 3D animated films, make heavy use of deforming mesh sequences. In this paper, we generalize gradient domain editing to deforming mesh sequences. Our framework is keyframe based. Given sparse and irregularly distributed constraints at unevenly spaced keyframes, our solution first adjusts the meshes at the keyframes to satisfy these constraints, and then smoothly propagate the constraints and deformations at keyframes to the whole sequence to generate new deforming mesh sequence. To achieve convenient keyframe editing, we have developed an efficient alternating least-squares method. It harnesses the power of subspace deformation and two-pass linear methods to achieve high-quality deformations. We have also developed an effective algorithm to define boundary conditions for all frames using handle trajectory editing. Our deforming mesh editing framework has been successfully applied to a number of editing scenarios with increasing complexity, including footprint editing, path editing, temporal filtering, handle-based deformation mixing, and spacetime morphing. Weiwei Xu 0003, Kun Zhou 0001, Yizhou Yu, Qifeng Tan, Qunsheng Peng 0001, Baining Guo |
ACM Trans. Graph. | 1 |
| 2007 | Direct manipulation of subdivision surfaces on GPUsabstractWe present an algorithm for interactive deformation of subdivision surfaces, including displaced subdivision surfaces and subdivision surfaces with geometric textures. Our system lets the user directly manipulate the surface using freely-selected surface points as handles. During deformation the control mesh vertices are automatically adjusted such that the deforming surface satisfies the handle position constraints while preserving the original surface shape and details. To best preserve surface details, we develop a gradient domain technique that incorporates the handle position constraints and detail preserving objectives into the deformation energy. For displaced subdivision surfaces and surfaces with geometric textures, the deformation energy is highly nonlinear and cannot be handled with existing iterative solvers. To address this issue, we introduce a shell deformation solver, which replaces each numerically unstable iteration step with two stable mesh deformation operations. Our deformation algorithm only uses local operations and is thus suitable for GPU implementation. The result is a real-time deformation system running orders of magnitude faster than the state-of-the-art multigrid mesh deformation solver. We demonstrate our technique with a variety of examples, including examples of creating visually pleasing character animations in real-time by driving a subdivision surface with motion capture data. Kun Zhou 0001, Weiwei Xu 0003, Baining Guo, Harry Shum |
ACM Trans. Graph. | 3 |
| 2006 | 2D shape deformation using nonlinear least squares optimization
Yanlin Weng, Weiwei Xu 0003, Yanchen Wu, Kun Zhou 0001, Baining Guo |
Vis. Comput. | 2 |
| 2003 | Easybowling: a small bowling machine based on virtual simulation
Weiwei Xu 0003, Jin Huang 0001, Jiaoying Shi |
Comput. Graph. | 2 |