EDBT 2026 Demo / reviewers in the wild / expert
Song-Hai Zhang
dblp:45/6733
· DBLP profile ↗
115ranked-venue papers
13as first author
87since 2021 · last 2026
0000-0003-0460-1586ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 95 · 9 first-author · 74 since 2021Artificial intelligence and machine learning · 29 · 2 first-author · 26 since 2021Applied, interdisciplinary, general and emerging computing · 11 · 4 first-author · 4 since 2021Human-computer interaction and ubiquitous computing · 10 · 10 since 2021Systems, architecture and hardware · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Incorporating strafing gain into redirected walking with pose score guidance
Jin-Feng Li, Sen-Zhe Xu 0001, Qiang Tong 0001, Peng-Hui Yuan, Ling-Long Zou, Er-Xia Luo, Qi Wen Gan, Song-Hai Zhang |
Comput. Graph. | 8 |
| 2026 | Automatic Planning of Urban Green SpacesabstractUrban green spaces such as parks and gardens are indispensable in both virtual and real-world environments. Therefore, planning such spaces is highly valuable. While scene synthesis literature has limited interest in this topic, many existing parametric design and procedural content generation approaches can be adapted to generate urban green spaces. However, these approaches heavily rely on manual work or are prone to producing monotonously repeated objects. This paper presents a framework that can automatically plan urban green spaces. Tailored to urban green space design, our framework comprises three steps: road system generation, region type planning, and model placement. First, it constructs undirected graphs to generate a sound road system for an empty site and divides the space into separate regions. Then it applies a genetic algorithm to plan suitable surface and vegetation for every region. Finally, it places landscape models based on various patterns and adds embellishments to complete an appealing urban green space. Our framework enables the automatic production of urban green spaces. Through extensive experiments, we demonstrate that the generated results are plausible and reasonable. Jia-Hong Liu, Shao-Kui Zhang, Qiaochu Liu, Chuyue Zhang, Song-Hai Zhang |
Comput. Vis. Media | 7 |
| 2026 | StoreSketcher: An Interactive Framework for Planning Commercial Retail Scene LayoutabstractRetail space planning, arranging store sections and product placements to optimize customer flow and stimulate purchases helps retailers to increase sales and enhances the customer shopping experience. It can be challenging for retailers to arrange numerous products within limited shelf space. This paper introduces StoreSketcher, an interactive tool that assists retailers in planning retail layouts efficiently at macro and micro levels by providing intelligent suggestions. We have extracted commercial relationships between products and categories, built spatial rules for commercial objects, and developed an interactive framework for synthesizing retail layouts. When the user points to shelf space in the layout, StoreSketcher evaluates the spatial significance of the location and its commercial relation to the surrounding context to present appropriate suggestions. Quantitative experiments demonstrate that StoreSketcher significantly assists in planning well-organized retail layouts. The suggestions provided by StoreSketcher not only boost cross-selling and impulse purchasing for retailers, but also enhance product findability for customers. Hou Tam, Shao-Kui Zhang, Yulin Jin, Hanxi Zhu, Song-Hai Zhang |
Comput. Vis. Media | 5 |
| 2026 | DiffPano++: Scalable and Consistent Multi-View Panorama Generation with Spherical Epipolar-Aware Diffusion
Chenhao Ji, Weicai Ye, Zheng Chen 0016, Junyao Gao 0002, Xiaoshui Huang, Xuekuan Wang, Guofeng Zhang 0001, Song-Hai Zhang, Tong He 0001, Wanli Ouyang, Cairong Zhao |
Int. J. Comput. Vis. | 8 |
| 2026 | Gait-Synced Translation Gain for Naturalistic VR MotionabstractTranslation gain is a key Redirected Walking (RDW) technique in Virtual Reality (VR) that enables users to navigate virtual environments (VEs) larger than the available physical space. The technique was originally developed to scale users' walking distance in the VE and is typically applied continuously, regardless of the user's motion state. We introduce Gait-Synced Translation Gain (GSTG), a novel approach that adapts translation gain by synchronizing it with the user's gait cycle. GSTG leverages the single-limb support phase of walking-when users are less stable and thus less sensitive to external disturbances-to apply higher levels of gain. This approach allows greater manipulation while preserving natural walking sensations and avoiding additional cybersickness. A user study comparing GSTG with continuous translation gain demonstrates significant improvements in perceived naturalness and comfort. Our results highlight the potential of gait-synchronized gain to enhance immersion, offering new possibilities for more realistic and comfortable VR locomotion. Fiona Xiao Yu Chen, Sen-Zhe Xu 0001, Kui Huang, Ariel Shamir, Song-Hai Zhang |
IEEE Trans. Vis. Comput. Graph. | 5 |
| 2026 | SceneCluster: Interactive Scene Synthesis by Clustering Groups of Furniture ObjectsabstractScene synthesis is crucial to computer graphics. However, the current interactive scene synthesis methods usually cost the user too much time and interactions to edit objects. This paper presents a new interactive scene synthesis method that alleviates the designer from interacting with the 3D scene and its objects. Designers only need to select an object group in an independent panel through coarse clustering and fine clustering. Then, the furniture objects will be automatically added to the scene. This paper proposes a two-level clustering that applies the Affinity Propagation Algorithm (APA) to groups of furniture objects such that the object groups can even be clustered without linear representations, latent encoding, etc. To fully apply the APA, we also propose quantitatively measuring how different the two layouts are, i.e., how quantitatively the arrangements of two object groups differ. Experiments first show that our method is more user-friendly and interactively efficient than other interactive synthesis methods. By comparing our method with recent automatic scene synthesis methods, we demonstrate that our methods still have competitive plausibility. We also verify that our method does not harm the diversity and generalization of 3D scenes. Shao-Kui Zhang, Hanxi Zhu, Song-Hai Zhang |
IEEE Trans. Vis. Comput. Graph. | 6 |
| 2026 | Synchronizing virtual backgrounds with monocular camera motion: A novel panoramic framework for video conferencingabstractWe present a novel 360° panoramic video conferencing system that dynamically synchronizes virtual backgrounds with real-time camera motion, addressing the limitations of static backgrounds in conventional systems. By integrating robust human segmentation, monocular visual odometry (VO), and virtual environment rendering, our method achieves seamless alignment between foreground participants and immersive 3D virtual scenes. Unlike prior approaches that suffer from foreground-background desynchronization during camera rotations or user movements, our framework estimates camera rotation in 3-DoF using a hybrid pipeline combining feature-based patch tracking and pose smoothing, while ignoring translation artifacts to maintain stability. This work bridges the gap between computational efficiency and MR-driven telepresence, offering a practical solution for next-generation virtual collaboration. Sen-Zhe Xu 0001, Zian Zhou, Song-Hai Zhang |
Virtual Real. Intell. Hardw. | 4 |
| 2025 | Splatter-360: Generalizable 360 Gaussian Splatting for Wide-baseline Panoramic ImagesabstractWide-baseline panoramic images are frequently used in applications like VR and simulations to minimize capturing labor costs and storage needs. However, synthesizing novel views from these panoramic images in real time remains a significant challenge, especially due to panoramic imagery’s high resolution and inherent distortions. Although existing 3D Gaussian splatting (3DGS) methods can produce photo-realistic views under narrow baselines, they often overfit the training views when dealing with wide-baseline panoramic images due to the difficulty in learning precise geometry from sparse 360° views. This paper presents Splatter-360, a novel end-to-end generalizable 3DGS framework designed to handle wide-baseline panoramic images. Unlike previous approaches, Splatter-360 performs multi-view matching directly in the spherical domain by constructing a spherical cost volume through a spherical sweep algorithm, enhancing the network’s depth perception and geometry estimation. Additionally, we introduce a 3D-aware bi-projection encoder to mitigate the distortions inherent in panoramic images and integrate cross-view attention to improve feature interactions across multiple viewpoints. This enables robust 3D-aware feature representations and real-time rendering capabilities. Experimental results on the HM3D [26] and Replica [27] demonstrate that Splatter-360 significantly outperforms state-of-the-art NeRF and 3DGS methods (e.g., PanoGRF, MVSplat, DepthSplat, and HiSplat) in both synthesis quality and generalization performance for wide-baseline panoramic images. Code and trained models are available at https://3d-aigc.github.io/Splatter-360/. Zheng Chen 0016, Chenming Wu, Zhelun Shen, Chen Zhao 0011, Weicai Ye, Haocheng Feng, Errui Ding, Song-Hai Zhang |
CVPR | 8 |
| 2025 | MeGA: Hybrid Mesh-Gaussian Head Avatar for High-Fidelity Rendering and Head EditingabstractCreating high-fidelity head avatars from multi-view videos is essential for many AR/VR applications. However, current methods often struggle to achieve high-quality renderings across all head components (e.g., skin vs. hair) due to the limitations of using one single representation for elements with varying characteristics. In this paper, we introduce a Hybrid Mesh-Gaussian Head Avatar (MeGA) that models different head components with more suitable representations. Specifically, we employ an enhanced FLAME mesh for the facial representation and predict a UV displacement map to provide per-vertex offsets for improved personalized geometric details. To achieve photorealistic rendering, we use deferred neural rendering to obtain facial colors and decompose neural textures into three meaningful parts. For hair modeling, we first build a static canonical hair using 3D Gaussian Splatting. A rigid transformation and an MLP-based deformation field are further applied to handle complex dynamic expressions. Combined with our occlusion-aware blending, MeGA generates higher-fidelity renderings for the whole head and naturally supports diverse downstream tasks. Experiments on the NeRSemble dataset validate the effectiveness of our designs, outperforming previous state-of-the-art methods and enabling versatile editing capabilities, including hairstyle alteration and texture editing. The code is released in https://github.com/conallwang/MeGA. Cong Wang 0045, Heyi Sun, Shen-Han Qian, Linchao Bao, Song-Hai Zhang |
CVPR | 7 |
| 2025 | DepthSync: Diffusion Guidance-Based Depth Synchronization for Scale- and Geometry-Consistent Video Depth EstimationabstractDiffusion-based video depth estimation methods have achieved remarkable success with strong generalization ability. However, predicting depth for long videos remains challenging. Existing methods typically split videos into overlapping sliding windows, leading to accumulated scale discrepancies across different windows, particularly as the number of windows increases. Additionally, these methods rely solely on 2D diffusion priors, overlooking the inherent 3D geometric structure of video depths, which results in geometrically inconsistent predictions. In this paper, we propose DepthSync, a novel, training-free framework using diffusion guidance to achieve scale- and geometry-consistent depth predictions for long videos. Specifically, we introduce scale guidance to synchronize the depth scale across windows and geometry guidance to enforce geometric alignment within windows based on the inherent 3D constraints in video depths. These two terms work synergistically, steering the denoising process toward consistent depth predictions. Experiments on various datasets validate the effectiveness of our method in producing depth estimates with improved scale and geometry consistency, particularly for long videos. Yuejiang Dong, Wang Zhao 0001, Ying Shan, Song-Hai Zhang |
ICCV | 5 |
| 2025 | NeuFrameQ: Neural Frame Fields for Scalable and Generalizable Anisotropic Quadrangulation
Ying-Tian Liu, Xin Yu 0004, Yan-Pei Cao 0001, Ding Liang, Ariel Shamir, Song-Hai Zhang |
ICCV | 9 |
| 2025 | SVG-Head: Hybrid Surface-Volumetric Gaussians for High-Fidelity Head Reconstruction and Real-Time Editing
Heyi Sun, Cong Wang 0045, Tian-Xing Xu, Chunchao Guo, Song-Hai Zhang |
ICCV | 7 |
| 2025 | Geometrycrafter: Consistent Geometry Estimation for Open-World Videos With Diffusion PriorsabstractDespite remarkable advancements in video depth estimation, existing methods exhibit inherent limitations in achieving geometric fidelity through the affine-invariant predictions, limiting their applicability in reconstruction and other metrically grounded downstream tasks. We propose GeometryCrafter, a novel framework that recovers high-fidelity point map sequences with temporal coherence from open-world videos, enabling accurate 3D/4D reconstruction, camera parameter estimation, and other depth-based applications. At the core of our approach lies a point map Variational Autoencoder (VAE) that learns a latent space agnostic to video latent distributions for effective point map encoding and decoding. Leveraging the VAE, we train a video diffusion model to model the distribution of point map sequences conditioned on the input videos. Extensive evaluations on diverse datasets demonstrate that GeometryCrafter achieves state-of-the-art 3D accuracy, temporal consistency, and generalization capability. Tian-Xing Xu, Xiangjun Gao, Wenbo Hu 0002, Xiaoyu Li 0002, Song-Hai Zhang, Ying Shan |
ICCV | 5 |
| 2025 | Perceiving Safety: User Preferences and Perception of Multimodal Virtual Boundary Cues in VR
Sen-Zhe Xu 0001, Yang-Fu Ren, Qi Wen Gan, Song-Hai Zhang |
ICXR | 5 |
| 2025 | The Effects of Head Pitch on Translation Gain in Virtual Reality Environments
Hong-Ru Ji, Sen-Zhe Xu 0001, Yang-Fu Ren, Song-Hai Zhang |
ICXR | 5 |
| 2025 | XR Headset-Empowered Sketch Art Education: Enhancing Learning Outcomes Through Spatial Perception and Layered Teaching
Lijia Li, Zeya Wang, Sen-Zhe Xu 0001, Qi Wen Gan, Song-Hai Zhang |
ICXR | 5 |
| 2025 | Safeteleport: Potential Field-Guided Teleportation for Personal Space Protection in Social VRabstractIn social virtual reality (VR), maintaining appropriate interpersonal distance is essential for user comfort and privacy. However, most existing locomotion methods provide limited support for respecting personal space, leaving users vulnerable to unintentional or socially inappropriate intrusions. To address this issue, we propose potential field-guided teleportation, a proactive locomotion framework consisting of two method implementations that dynamically adjust teleportation targets based on real-time interpersonal proximity, preventing entry into others' personal spaces without explicit user intervention. We evaluate our technique through two user studies: a preliminary study exploring energy-based constraint parameters, followed by a comparative study against conventional and negotiated teleportation methods. Experiments were conducted in socially interactive VR scenarios populated with simulated users exhibiting human-like behaviors. Results demonstrate that our methods reduce perceived social anxiety while maintaining locomotion efficiency and usability. This work presents a socially-aware locomotion strategy that balances personal space protection with effective and socially appropriate movement in shared virtual environments. Yijun Li 0006, Sen-Zhe Xu 0001, Wentong Shu, Hao-Zhong Yang, Zinan Han, Miao Wang 0004, Song-Hai Zhang |
ISMAR | 7 |
| 2025 | Can the Perceived Capability of Your Virtual Avatar Enhance Exercise Performance?abstractThe rise of Virtual Reality (VR) sports has been driven by evolving work patterns, limited access to physical exercise spaces, and a growing focus on health and wellness. Beyond the enjoyment and convenience offered by VR exercise, we aim to enhance users' athletic performance through this medium. Prior research on the Proteus Effect in VR sports has demonstrated the potential of customized, stronger, or younger avatars to improve exercise outcomes. However, such effects are limited for users who already perceive themselves as strong and youthful. To address this, we propose a more general approach: representing perceived avatar capability through facial expressions and vocal cues to influence user performance. In this study, we examined how manipulating the perceived capability level of virtual avatars (high, neutral, low) during dumbbell lateral raises in a VR gym affected exercise performance. Results indicated that avatars with high-capability expressions significantly enhanced performance and motivation compared to low-capability avatars. These findings underscore the promise of using facial and auditory cues to represent perceived capability, offering new directions for designing emotionally intelligent fitness applications. Sen-Zhe Xu 0001, Bo-Sheng Huang, Zian Zhou, Run-Yu Li, Song-Hai Zhang, Xu-Cheng Yin |
ISMAR | 5 |
| 2025 | Synthesizing 3D Scenes via Diffusion Model that Incorporates Indoor Scene CharacteristicsabstractDiffusion model has been used in indoor scene synthesis and has made significant progress. Current works encode an indoor scene as a top-down view of the room, a list of objects, and their world co-ordinates and orientation. In this paper, we develop a diffusion-based training and synthetic method which incorporates indoor scene ''characteristics''. Firstly, we calculate the relative transformations among objects to capture the local characteristics of the scene. We send this relative transformation into the self-attention layer of the denoising network as ''relative positional encoding''. Secondly, we use room guidance to guide the objects to fit the room's geometry. This improvement uses the room's characteristics to solve the physical collision problem occurring in former diffusion-based works, while preserving plausibilities. Experiments show that our improvements improve the scene variety and quality. Shao-Kui Zhang, Yi-Tao Chen, Zirui Zhou, Song-Hai Zhang |
ACM Multimedia | 6 |
| 2025 | Foregrounding collaboration in CAVE systems: A survey across domains, interaction, and system design
Fiona Xiao Yu Chen, Sen-Zhe Xu 0001, Song-Hai Zhang |
Comput. Graph. | 3 |
| 2025 | Foreword to chinagraph 2024 special section
Song-Hai Zhang, Juyong Zhang |
Comput. Graph. | 2 |
| 2025 | Look at that distractor: Dynamic translation gain under low perceptual load in virtual reality
Ling-Long Zou, Qiang Tong 0001, Er-Xia Luo, Sen-Zhe Xu 0001, Song-Hai Zhang |
Comput. Graph. | 5 |
| 2025 | DIFF: A dataset for indoor flexible furnitureabstractRecently, indoor scene synthesis has gathered significant attention, leading to the development of numerous indoor datasets. However, existing datasets only address static furniture and scenes, ignoring the need for dynamic interior design scenarios that emphasize flexible functionalities. Addressing this gap, we present DIFF (Dataset for Indoor Flexible Furniture), featuring expertly crafted and labeled furniture modules capable of inter-transforming between different states, e.g., a cabinet can be inter-transformed to a desk. Each module exhibits flexibility in shifting to multiple shapes and functionalities. Additionally, we propose a method that adapts our dataset to generate flexible layouts. By matching our flexible objects to objects from existing datasets, we use a graph-based approach to migrate the spatial relation priors for optimizing a layout; subsequent layouts are then generated by minimizing a transition-cost function. Analyses and user studies validate the quality of our modules and demonstrate the plausibility of the proposed method. Jia-Hong Liu, Shao-Kui Zhang, Shuran Sun, Song-Hai Zhang |
Graph. Model. | 5 |
| 2025 | Strategies for reducing motion sickness in virtual reality through improved handheld controller movementsabstractAs technology advances, user demand for immersive and authentic information presentation rises. Traditional 2D displays and interactions fail to meet modern standards, while virtual reality (VR) is gaining attention for its immersive experience. However, using a controller for VR movement can cause dizziness due to mismatched visual and vestibular cues, impacting the VR experience. This paper analyzes the main causes of VR-induced vertigo and develops improved handheld controller movement strategies. These strategies adjust the user’s pitch angle and field of view in real time or map the user’s real-world head acceleration to the virtual character. By intelligently adjusting the controller-to-VR display mapping, these methods reduce vertigo. In addition, this paper also verified the actual effects of these designs through a series of experiments, and conducted detailed data analysis on the degree of user vertigo. The experimental results showed that using a specific improved handheld controller movement design can significantly improve the user’s comfort in the VR environment, effectively reducing the occurrence of vertigo and discomfort. Khang Yeu Tang, Juhong Wang, Yu He 0001, Sen-Zhe Xu 0001, Song-Hai Zhang |
Graph. Model. | 6 |
| 2025 | SceneExplorer: An Interactive System for Expanding, Scheduling, and Organizing Transformable LayoutsabstractNowadays, 3D scenes are not merely static arrangements of objects. With the development of transformable modules, furniture objects can be translated, rotated, and even reshaped to achieve scenes with different functions (e.g., from a bedroom to a living room). Transformable domestic space, therefore, studies how a layout can change its function by reshaping and rearranging transformable modules, resulting in various transformable layouts. In practice, a rearrangement is dynamically conducted by reshaping/translating/rotating furniture objects with proper schedules, which can consume more time for designers than static scene design. Due to changes in objects' functions, potential transformable layouts may also be extensive, making it hard to explore desired layouts. We present a system for exploring transformable layouts. Given a single input scene consisting of transformable modules, our system first attempts to derive more layouts by reshaping and rearranging the modules. The derived scenes are organized into a graph-like hierarchy according to their functions, where edges represent functional evolutions (e.g., a living room can be reshaped to a bedroom), and nodes represent layouts that are dynamically transformable through translating/rotating/reshaping modules. The resulting hierarchy lets scene designers interactively explore possible scene variants and preview the animated rearrangement process. Experiments show that our system is efficient for generating transformable layouts, sensible for organizing functional hierarchies, and inspiring for providing ideas during interactions. Shao-Kui Zhang, Jia-Hong Liu, Junkai Huang 0003, Ziwei Chi, Hou Tam, Yongliang Yang 0002, Song-Hai Zhang |
IEEE Trans. Vis. Comput. Graph. | 7 |
| 2025 | GP-Recon: Online Monocular Neural 3D Reconstruction With Geometric PriorabstractHigh-fidelity online 3D scene reconstruction from monocular videos continues to be challenging, especially for coherent and fine-grained geometry reconstruction. The previous learning-based online 3D reconstruction approaches with neural implicit representations have shown a promising ability for coherent scene reconstruction, but often fail to consistently reconstruct fine-grained geometric details during online reconstruction. This paper presents a new on-the-fly monocular 3D reconstruction approach, named GP-Recon, to perform high-fidelity online neural 3D reconstruction with fine-grained geometric details. We incorporate geometric prior (GP) into a scene's neural geometry learning to better capture its geometric details and, more importantly, propose an online volume rendering optimization to reconstruct and maintain geometric details during the online reconstruction task. The extensive comparisons with state-of-the-art approaches show that our GP-Recon consistently generates more accurate and complete reconstruction results with much better fine-grained details, both quantitatively and qualitatively. Zixin Zou, Shi-Sheng Huang, Yan-Pei Cao 0001, Tai-Jiang Mu, Ying Shan, Hongbo Fu 0001, Song-Hai Zhang |
IEEE Trans. Vis. Comput. Graph. | 7 |
| 2024 | PPEA-Depth: Progressive Parameter-Efficient Adaptation for Self-Supervised Monocular Depth EstimationabstractSelf-supervised monocular depth estimation is of significant importance with applications spanning across autonomous driving and robotics. However, the reliance on self-supervision introduces a strong static-scene assumption, thereby posing challenges in achieving optimal performance in dynamic scenes, which are prevalent in most real-world situations. To address these issues, we propose PPEA-Depth, a Progressive Parameter-Efficient Adaptation approach to transfer a pre-trained image model for self-supervised depth estimation. The training comprises two sequential stages: an initial phase trained on a dataset primarily composed of static scenes, succeeded by an expansion to more intricate datasets involving dynamic scenes. To facilitate this process, we design compact encoder and decoder adapters to enable parameter-efficient tuning, allowing the network to adapt effectively. They not only uphold generalized patterns from pre-trained image models but also retain knowledge gained from the preceding phase into the subsequent one. Extensive experiments demonstrate that PPEA-Depth achieves state-of-the-art performance on KITTI, CityScapes and DDAD datasets. Yue-Jiang Dong, Ying-Tian Liu, Song-Hai Zhang |
AAAI | 5 |
| 2024 | Sparse3D: Distilling Multiview-Consistent Diffusion for Object Reconstruction from Sparse ViewsabstractReconstructing 3D objects from extremely sparse views is a long-standing and challenging problem. While recent techniques employ image diffusion models for generating plausible images at novel viewpoints or for distilling pre-trained diffusion priors into 3D representations using score distillation sampling (SDS), these methods often struggle to simultaneously achieve high-quality, consistent, and detailed results for both novel-view synthesis (NVS) and geometry. In this work, we present Sparse3D, a novel 3D reconstruction method tailored for sparse view inputs. Our approach distills robust priors from a multiview-consistent diffusion model to refine a neural radiance field. Specifically, we employ a controller that harnesses epipolar features from input views, guiding a pre-trained diffusion model, such as Stable Diffusion, to produce novel-view images that maintain 3D consistency with the input. By tapping into 2D priors from powerful image diffusion models, our integrated model consistently delivers high-quality results, even when faced with open-world objects. To address the blurriness introduced by conventional SDS, we introduce the category-score distillation sampling (C-SDS) to enhance detail. We conduct experiments on CO3DV2 which is a multi-view dataset of real-world objects. Both quantitative and qualitative evaluations demonstrate that our approach outperforms previous state-of-the-art works on the metrics regarding NVS and geometry reconstruction. Zixin Zou, Weihao Cheng 0002, Yan-Pei Cao 0001, Shi-Sheng Huang, Ying Shan, Song-Hai Zhang |
AAAI | 6 |
| 2024 | Single-Video Temporal Consistency Enhancement with Rolling Guidance
Xiaonan Fang 0001, Song-Hai Zhang |
CVM (2) | 2 |
| 2024 | Walking Telescope: Exploring the Zooming Effect in Expanding Detection Threshold Range for Translation Gain
Er-Xia Luo, Khang Yeu Tang, Sen-Zhe Xu 0001, Qiang Tong 0001, Song-Hai Zhang |
CVM (1) | 5 |
| 2024 | PI3D: Efficient Text-to-3D Generation with Pseudo-Image DiffusionabstractDiffusion models trained on large-scale text-image datasets have demonstrated a strong capability of con-trollable high-quality image generation from arbitrary text prompts. However, the generation quality and general-ization ability of 3D diffusion models is hindered by the scarcity of high-quality and large-scale 3D datasets. In this paper, we present PI3D, a framework that fully lever-ages the pre-trained text-to-image diffusion models' abil-ity to generate high-quality 3D shapes from text prompts in minutes. The core idea is to connect the 2D and 3D domains by representing a 3D shape as a set of Pseudo RGB Images. We fine-tune an existing text-to-image dif-fusion model to produce such pseudo-images using a small number of text-3D pairs. Surprisingly, we find that it can al-ready generate meaningful and consistent 3D shapes given complex text descriptions. We further take the generated shapes as the starting point for a lightweight iterative re-finement using score distillation sampling to achieve high-quality generation under a low budget. PI3D generates a single 3D shape from text in only 3 minutes and the quality is validated to outperform existing 3D generative models by a large margin. Ying-Tian Liu, Guan Luo, Heyi Sun, Song-Hai Zhang |
CVPR | 6 |
| 2024 | Wonder3D: Single Image to 3D Using Cross-Domain DiffusionabstractIn this work, we introduce Wonder3D, a novel method for efficiently generating high-fidelity textured meshes from single-view images. Recent methods based on Score Distillation Sampling (SDS) have shown the potential to recover 3D geometry from 2D diffusion priors, but they typically suffer from time-consuming per-shape optimization and inconsistent geometry. In contrast, certain works di-rectly produce 3D information via fast network inferences, but their results are often of low quality and lack geometric details. To holistically improve the quality, consistency, and efficiency of single-view reconstruction tasks, we pro-pose a cross-domain diffusion model that generates multi-view normal maps and the corresponding color images. To ensure the consistency of generation, we employ a multi-view cross-domain attention mechanism that facilitates information exchange across views and modalities. Lastly, we introduce a geometry-aware normal fusion algorithm that extracts high-quality surfaces from the multi-view 2D representations in only 2 r-;» 3 minutes. Our extensive evaluations demonstrate that our method achieves high-quality reconstruction results, robust generalization, and good efficiency compared to prior works. Xiaoxiao Long, Cheng Lin 0001, Yuan Liu 0025, Zhiyang Dou, Lingjie Liu, Yuexin Ma, Song-Hai Zhang, Marc Habermann, Christian Theobalt, Wenping Wang 0001 |
CVPR | 8 |
| 2024 | DreamComposer: Controllable 3D Object Generation via Multi-View ConditionsabstractUtilizing pretrained 2D large-scale generative models, recent works are capable of generating high-quality novel views from a single in-the-wild image. However, due to the lack of information from multiple views, these works encounter difficulties in generating controllable novel views. In this paper, we present DreamComposer, a flexible and scalable framework that can enhance existing view-aware diffusion models by injecting multi-view conditions. Specifically, DreamComposer first uses a view-aware 3D lifting module to obtain 3D representations of an object from multiple views. Then, it renders the latent features of the target view from 3D representations with the multi-view feature fusion module. Finally the target view features extracted from multi-view inputs are injected into a pretrained diffusion model. Experiments show that DreamComposer is compatible with state-of-the-art diffusion models for zero-shot novel view synthesis, further enhancing them to generate high-fidelity novel view images with multi-view conditions, ready for controllable 3D object reconstruction and various other applications. Yunhan Yang, Xiaoyang Wu 0002, Song-Hai Zhang, Hengshuang Zhao, Tong He 0001, Xihui Liu |
CVPR | 5 |
| 2024 | Triplane Meets Gaussian Splatting: Fast and Generalizable Single-View 3D Reconstruction with TransformersabstractRecent advancements in 3D reconstruction from single images have been driven by the evolution of generative models. Prominent among these are methods based on Score Distillation Sampling (SDS) and the adaptation ofdiffusion models in the 3D domain. Despite their progress, these techniques often face limitations due to slow optimization or rendering processes, leading to extensive training and optimization times. In this paper, we introduce a novel approach for single-view reconstruction that efficiently generates a 3D model from a single image via feed-forward inference. Our method utilizes two transformer-based networks, namely a point decoder and a triplane decoder, to reconstruct 3D objects using a hybrid Triplane-Gaussian intermediate representation. This hybrid representation strikes a balance, achieving a faster rendering speed compared to implicit representations while simultaneously delivering superior rendering quality than explicit representations. The point decoder is designed for generating point clouds from single images, offering an explicit representation which is then utilized by the triplane decoder to query Gaussian features for each point. This design choice addresses the challenges associated with directly regressing explicit 3D Gaussian attributes characterized by their non-structural nature. Subsequently, the 3D Gaussians are decoded by an MLP to enable rapid rendering through splatting. Both decoders are built upon a scalable, transformer-based architecture and have been efficiently trained on large-scale 3D datasets. The evaluations conducted on both synthetic datasets and real-world images demonstrate that our method not only achieves higher quality but also ensures a faster runtime in comparison to previous state-of-the-art techniques. Please see our project page at https://zouzx.github.io/TriplaneGaussian/ Zixin Zou, Yangguang Li 0001, Ding Liang, Yan-Pei Cao 0001, Song-Hai Zhang |
CVPR | 7 |
| 2024 | Texture-GS: Disentangling the Geometry and Texture for 3D Gaussian Splatting Editing
Tian-Xing Xu, Wenbo Hu 0002, Yukun Lai, Ying Shan, Song-Hai Zhang |
ECCV (25) | 5 |
| 2024 | Text-to-3D with Classifier Score DistillationabstractText-to-3D generation has made remarkable progress recently, particularly with methods based on Score Distillation Sampling (SDS) that leverages pre-trained 2D diffusion models. While the usage of classifier-free guidance is well acknowledged to be crucial for successful optimization, it is considered an auxiliary trick rather than the most essential component. In this paper, we re-evaluate the role of classifier-free guidance in score distillation and discover a surprising finding: the guidance alone is enough for effective text-to-3D generation tasks.
We name this method Classifier Score Distillation (CSD), which can be interpreted as using an implicit classification model for generation. This new perspective reveals new insights for understanding existing techniques. We validate the effectiveness of CSD across a variety of text-to-3D tasks including shape generation, texture synthesis, and shape editing, achieving results superior to those of state-of-the-art methods. Our project page is https://xinyu-andy.github.io/Classifier-Score-Distillation Xin Yu 0004, Yangguang Li 0001, Ding Liang, Song-Hai Zhang, Xiaojuan Qi 0001 |
ICLR | 5 |
| 2024 | MAL: Motion-Aware Loss with Temporal and Distillation Hints for Self-Supervised Depth EstimationabstractDepth perception is crucial for a wide range of robotic applications. Multi-frame self-supervised depth estimation methods have gained research interest due to their ability to leverage large-scale, unlabeled real-world data. However, the self-supervised methods often rely on the assumption of a static scene and their performance tends to degrade in dynamic environments. To address this issue, we present Motion-Aware Loss, which leverages the temporal relation among consecutive input frames and a novel distillation scheme between the teacher and student networks in the multi-frame self-supervised depth estimation methods. Specifically, we associate the spatial locations of moving objects with the temporal order of input frames to eliminate errors induced by object motion. Meanwhile, we enhance the original distillation scheme in multi-frame methods to better exploit the knowledge from a teacher network. MAL is a novel, plug-and-play module designed for seamless integration into multi-frame self-supervised monocular depth estimation methods. Adding MAL into previous state-of-the-art methods leads to a reduction in depth estimation errors by up to 4.2% and 10.8% on KITTI and CityScapes benchmarks, respectively. Yue-Jiang Dong, Song-Hai Zhang |
ICRA | 3 |
| 2024 | Controllable Procedural Generation of Landscapes
Jia-Hong Liu, Shao-Kui Zhang, Chuyue Zhang, Song-Hai Zhang |
ACM Multimedia | 4 |
| 2024 | 3D Gaussian Editing with A Single ImageabstractThe modeling and manipulation of 3D scenes captured from the real world are pivotal in various applications, attracting growing research interest. While previous works on editing have achieved interesting results through manipulating 3D meshes, they often require accurately reconstructed meshes to perform editing, which limits their application in 3D content generation. To address this gap, we introduce a novel single-image-driven 3D scene editing approach based on 3D Gaussian Splatting, enabling intuitive manipulation via directly editing the content on a 2D image plane. Our method learns to optimize the 3D Gaussians to align with an edited version of the image rendered from a user-specified viewpoint of the original scene. To capture long-range object deformation, we introduce positional loss into the optimization process of 3D Gaussian Splatting and enable gradient propagation through reparameterization. To handle occluded 3D Gaussians when rendering from the specified viewpoint, we build an anchor-based structure and employ a coarse-to-fine optimization strategy capable of handling long-range deformation while maintaining structural stability. Furthermore, we design a novel masking strategy to adaptively identify non-rigid deformation regions for fine-scale modeling. Extensive experiments show the effectiveness of our method in handling geometric details, long-range, and non-rigid deformation, demonstrating superior editing flexibility and quality compared to previous approaches. Guan Luo, Tian-Xing Xu, Ying-Tian Liu, Xiaoxiong Fan, Song-Hai Zhang |
ACM Multimedia | 6 |
| 2024 | SceneExpander: Real-Time Scene Synthesis for Interactive Floor Plan EditingabstractScene synthesis has gained significant attention recently, and interactive scene synthesis focuses on yielding scenes according to user preferences. Existing literature either generates floor plans or scenes according to the floor plans. The system proposed in this paper generates scenes over floor plans in real-time. Given an initial scene, the only interaction a user needs is changing the room shapes. Our framework splits/merges rooms and adds/rearranges/removes objects for each transient moment during interactions. A systematic pipeline achieves our framework by compressing objects' arrangements over modified room shapes in a transient moment, thus enabling real-time performances. We also propose elastic boxes that indicate how objects should be arranged according to their continuously changed contexts, such as room shapes and other objects. Through a few interactions, a floor plan filled with object layouts is generated concerning user preferences on floor plans and object layouts according to floor plans. Experiments show that our framework is efficient at user interactions and plausible for synthesizing 3D scenes. Shao-Kui Zhang, Junkai Huang 0003, Jia-Tong Zhang, Jia-Hong Liu, Yukun Lai, Song-Hai Zhang |
ACM Multimedia | 7 |
| 2024 | ScenePhotographer: Object-Oriented Photography for Residential Scenes
Shao-Kui Zhang, Hanxi Zhu, Jinghuan Chen, Zhike Peng, Yongliang Yang 0002, Song-Hai Zhang |
ACM Multimedia | 8 |
| 2024 | DiffPano: Scalable and Consistent Text to Panorama Generation with Spherical Epipolar-Aware DiffusionabstractDiffusion-based methods have achieved remarkable achievements in 2D image or 3D object generation, however, the generation of 3D scenes and even $360^{\circ}$ images remains constrained, due to the limited number of scene datasets, the complexity of 3D scenes themselves, and the difficulty of generating consistent multi-view images. To address these issues, we first establish a large-scale panoramic video-text dataset containing millions of consecutive panoramic keyframes with corresponding panoramic depths, camera poses, and text descriptions. Then, we propose a novel text-driven panoramic generation framework, termed DiffPano, to achieve scalable, consistent, and diverse panoramic scene generation. Specifically, benefiting from the powerful generative capabilities of stable diffusion, we fine-tune a single-view text-to-panorama diffusion model with LoRA on the established panoramic video-text dataset. We further design a spherical epipolar-aware multi-view diffusion model to ensure the multi-view consistency of the generated panoramic images. Extensive experiments demonstrate that DiffPano can generate scalable, consistent, and diverse panoramic images with given unseen text descriptions and camera poses. Weicai Ye, Chenhao Ji, Zheng Chen 0016, Junyao Gao 0002, Xiaoshui Huang, Song-Hai Zhang, Wanli Ouyang, Tong He 0001, Cairong Zhao, Guofeng Zhang 0001 |
NeurIPS | 6 |
| 2024 | SafeRDW: Keep VR Users Safe When Jumping Using Redirected WalkingabstractRedirected Walking (RDW) is an important intermediary layer in virtual reality (VR) interaction systems. It addresses the issues of spatial restrictions in VR exploration by imperceptibly remapping the virtual environment’s movement to the physical environment, enhancing the user’s immersive experience. In VR, jumping is also a noteworthy motion besides walking. However, existing redirected walking (RDW) algorithms typically focus on reducing collisions between users and obstacles during walking but overlook the safety when users perform significant actions such as jumping. This oversight can pose serious risks to users during VR exploration, especially when there are physical obstacles or boundaries near the virtual locations that require user jumping. We propose SafeRDW, the first RDW algorithm that takes the user’s jumping safety into consideration. The proposed method considers both walking and jumping actions in the virtual environment, reducing physical resets and redirecting users to safer locations when a jump is required in the virtual space, ensuring user safety. Simulation experiments and user study results both show that our method not only reduces the number of resets, but also significantly ensures user safety when they reach the jumping points in the virtual scene. Sen-Zhe Xu 0001, Kui Huang, Cheng-Wei Fan, Song-Hai Zhang |
VR | 4 |
| 2024 | Exploring the Impact of Visual Scene Characteristics and Adaptation Effects on Rotation Gain Perception in VRabstractRotation gain is a subtle manipulation technique commonly employed in Redirected Walking (RDW) methods due to its superior capability to alter a user’s virtual trajectory. Previous studies have reported that the imperceptible ranges of rotation gains are influenced by various factors, resulting in different detection threshold values, which may alter RDW performance. In this study, we focus on the effects of scene visual characteristics on the rotation gain and rotation gain thresholds (RGTs), which have been less explored in this area. In our experiments, we focus on three visual characteristics: visual density, spatial size, and realism. Each characteristic is tested at two different levels, resulting in a design of eight distinct VR scenes. Through extensive statistical analysis, we find that spatial size may influence user perception of rotation gain in different virtual environments (VEs), though the effect appears to be small. No significant results of sensitivity differences were found for visual density and realism. We show that the short-term temporal effect is another predominant factor influencing user perception of rotation gain, even when users experience different visual stimuli in VEs, such as different scene visual characteristic settings in our study. This result indicates that users’ adaptation effects on rotation gain can occur in as short a time as overnight intervals, rather than over weeks. Qi Wen Gan, Sen-Zhe Xu 0001, Song-Hai Zhang |
VRST | 4 |
| 2024 | AdaPIP: Adaptive picture-in-picture guidance for 360° film watchingabstract360° videos enable viewers to watch freely from different directions but inevitably prevent them from perceiving all the helpful information. To mitigate this problem, picture-in-picture (PIP) guidance was proposed using preview windows to show regions of interest (ROIs) outside the current view range. We identify several drawbacks of this representation and propose a new method for 360° film watching called AdaPIP. AdaPIP enhances traditional PIP by adaptively arranging preview windows with changeable view ranges and sizes. In addition, AdaPIP incorporates the advantage of arrow-based guidance by presenting circular windows with arrows attached to them to help users locate the corresponding ROIs more efficiently. We also adapted AdaPIP and Outside-In to HMD-based immersive virtual reality environments to demonstrate the usability of PIP-guided approaches beyond 2D screens. Comprehensive user experiments on 2D screens, as well as in VR environments, indicate that AdaPIP is superior to alternative methods in terms of visual experiences while maintaining a comparable degree of immersion. Yi-Xiao Li, Guan Luo, Yi-Ke Xu, Yu He 0001, Song-Hai Zhang |
Comput. Vis. Media | 6 |
| 2024 | Overcoming Spatial Constraints in VR: A Survey of Redirected Walking Techniques
Jia-Hong Liu, Yang-Fu Ren, Qi Wen Gan, Kui Huang, Fiona Xiao Yu Chen, Er-Xia Luo, Khang Yeu Tang, Yue-Yao Fu, Cheng-Wei Fan, Sen-Zhe Xu 0001, Song-Hai Zhang |
J. Comput. Sci. Technol. | 11 |
| 2024 | ScenePalette: Contextually Exploring Object Collections Through Multiplex Relations in 3D Scenes
Shao-Kui Zhang, Weiyu Xie, Chen Wang 0049, Song-Hai Zhang |
J. Comput. Sci. Technol. | 4 |
| 2024 | BiRD: Using Bidirectional Rotation Gain Differences to Redirect Users during Back-and-forth Head Turns in WalkingabstractRedirected walking (RDW) facilitates user navigation within expansive virtual spaces despite the constraints of limited physical spaces. It employs discrepancies between human visual-proprioceptive sensations, known as gains, to enable the remapping of virtual and physical environments. In this paper, we explore how to apply rotation gain while the user is walking. We propose to apply a rotation gain to let the user rotate by a different angle when reciprocating from a previous head rotation, to achieve the aim of steering the user to a desired direction. To apply the gains imperceptibly based on such a Bidirectional Rotation gain Difference (BiRD), we conduct both measurement and verification experiments on the detection thresholds of the rotation gain for reciprocating head rotations during walking. Unlike previous rotation gains which are measured when users are turning around in place (standing or sitting), BiRD is measured during users' walking. Our study offers a critical assessment of the acceptable range of rotational mapping differences for different rotational orientations across the user's walking experience, contributing to an effective tool for redirecting users in virtual environments. Sen-Zhe Xu 0001, Fiona Xiao Yu Chen, Song-Hai Zhang |
IEEE Trans. Vis. Comput. Graph. | 5 |
| 2024 | Spatial Contraction Based on Velocity Variation for Natural Walking in Virtual RealityabstractVirtual Reality (VR) offers an immersive 3D digital environment, but enabling natural walking sensations without the constraints of physical space remains a technological challenge. Previous VR locomotion methods, including game controller, teleportation, treadmills, walking-in-place, and redirected walking (RDW), have made strides towards overcoming this challenge. However, these methods also face limitations such as possible unnaturalness, additional hardware requirements, or motion sickness risks. This paper introduces "Spatial Contraction (SC)", an innovative VR locomotion method inspired by the phenomenon of Lorentz contraction in Special Relativity. Similar to the Lorentz contraction, our SC contracts the virtual space along the user's velocity direction in response to velocity variation. The virtual space contracts more when the user's speed is high, whereas minimal or no contraction happens at low speeds. We provide a virtual space transformation method for spatial contraction and optimize the user experience in smoothness and stability. Through SC, VR users can effectively traverse a longer virtual distance with a shorter physical walking. Different from locomotion gains, the spatial contraction effect is observable by the user and aligns with their intentions, so there is no inconsistency between the user's proprioception and visual perception. SC is a general locomotion method that has no special requirements for VR scenes. The experimental results of our live user studies in various virtual scenarios demonstrate that SC has a significant effect in reducing both the number of resets and the physical walking distance users need to cover. Furthermore, experiments have also demonstrated that SC has the potential for integration with existing locomotion techniques such as RDW. Sen-Zhe Xu 0001, Kui Huang, Cheng-Wei Fan, Song-Hai Zhang |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2024 | Multi-User Redirected Walking in Separate Physical Spaces for Online VR ScenariosabstractWith the recent rise of Metaverse, online multiplayer VR applications are becoming increasingly prevalent worldwide. However, as multiple users are located in different physical environments, different reset frequencies and timings can lead to serious fairness issues for online collaborative/competitive VR applications. For the fairness of online VR apps/games, an ideal online RDW strategy must make the locomotion opportunities of different users equal, regardless of different physical environment layouts. The existing RDW methods lack the scheme to coordinate multiple users in different PEs, and thus have the issue of triggering too many resets for all the users under the locomotion fairness constraint. We propose a novel multi-user RDW method that is able to significantly reduce the overall reset number and give users a better immersive experience by providing a fair exploration. Our key idea is to first find out the "bottleneck" user that may cause all users to be reset and estimate the time to reset given the users' next targets, and then redirect all the users to favorable poses during that maximized bottleneck time to ensure the subsequent resets can be postponed as much as possible. More particularly, we develop methods to estimate the time of possibly encountering obstacles and the reachable area for a specific pose to enable the prediction of the next reset caused by any user. Our experiments and user study found that our method outperforms existing RDW methods in online VR applications. Sen-Zhe Xu 0001, Jia-Hong Liu, Miao Wang 0004, Song-Hai Zhang |
IEEE Trans. Vis. Comput. Graph. | 5 |
| 2024 | SceneDirector: Interactive Scene Synthesis by Simultaneously Editing Multiple Objects in Real-TimeabstractIntelligent tools for creating synthetic scenes have been developed significantly in recent years. Existing techniques on interactive scene synthesis only incorporate a single object at every interaction, i.e., crafting a scene through a sequence of single-object insertions with user preferences. These techniques suggest objects by considering existent objects in the scene instead of fully picturing the eventual result, which is inherently problematic since the sets of objects to be inserted are seldom fixed during interactive processes. In this article, we introduce SceneDirector, a novel interactive scene synthesis tool to help users quickly picture various potential synthesis results by simultaneously editing groups of objects. Specifically, groups of objects are rearranged in real-time with respect to a position of an object specified by a mouse cursor or gesture, i.e., a movement of a single object would trigger the rearrangement of the existing object group, the insertions of potentially appropriate objects, and the removal of redundant objects. To achieve this, we first propose an idea of coherent group set which expresses various concepts of layout strategies. Subsequently, we present layout attributes, where users can adjust how objects are arranged by tuning the weights of the attributes. Thus, our method gives users intuitive control of both how to arrange groups of objects and where to place them. Through extensive experiments and two applications, we demonstrate the potentiality of our framework and how it enables concurrently effective and efficient interactions of editing groups of objects. Shao-Kui Zhang, Hou Tam, Ke-Xin Ren, Hongbo Fu 0001, Song-Hai Zhang |
IEEE Trans. Vis. Comput. Graph. | 6 |
| 2023 | DualVector: Unsupervised Vector Font Synthesis with Dual-Part RepresentationabstractAutomatic generation of fonts can be an important aid to typeface design. Many current approaches regard glyphs as pixelated images, which present artifacts when scaling and inevitable quality losses after vectorization. On the other hand, existing vector font synthesis methods either fail to represent the shape concisely or require vector supervision during training. To push the quality of vector font synthesis to the next level, we propose a novel dual-part representation for vector glyphs, where each glyph is modeled as a collection of closed “positive” and “negative” path pairs. The glyph contour is then obtained by boolean operations on these paths. We first learn such a representation only from glyph images and devise a subsequent contour refinement step to align the contour with an image representation to further enhance details. Our method, named DualVector, outperforms state-of-the-art methods in vector font synthesis both quantitatively and qualitatively. Our synthesized vector fonts can be easily converted to common digital font formats like TrueType Font for practical use. The code is released at https://github.com/thuliu-yt16/dualvector. Ying-Tian Liu, Matthew Fisher, Song-Hai Zhang |
CVPR | 6 |
| 2023 | CXTrack: Improving 3D Point Cloud Tracking with Contextual Informationabstract3D single object tracking plays an essential role in many applications, such as autonomous driving. It remains a challenging problem due to the large appearance variation and the sparsity of points caused by occlusion and lim-ited sensor capabilities. Therefore, contextual information across two consecutive frames is crucial for effective object tracking. However, points containing such useful information are often overlooked and cropped out in existing methods, leading to insufficient use of important contextual knowledge. To address this issue, we propose CXTrack, a novel transformer-based network for 3D object tracking, which exploits ConteXtual information to improve the tracking results. Specifically, we design a target-centric transformer network that directly takes point features from two consecutive frames and the previous bounding box as input to explore contextual information and implicitly propagate target cues. To achieve accurate localization for objects of all sizes, we propose a transformer-based localization head with a novel center embedding module to distinguish the target from distractors. Extensive experiments on three large-scale datasets, KITTI, nuScenes and Waymo Open Dataset, show that CXTrack achieves state-of-the-art tracking performance while running at 34 FPS. Tian-Xing Xu, Yukun Lai, Song-Hai Zhang |
CVPR | 4 |
| 2023 | Joint Implicit Neural Representation for High-fidelity and Compact Vector FontsabstractExisting vector font generation approaches either struggle to preserve high-frequency corner details of the glyph or produce vector shapes that have redundant segments, which hinders their applications in practical scenarios. In this paper, we propose to learn vector fonts from pixelated font images utilizing a joint neural representation that consists of a signed distance field (SDF) and a probabilistic corner field (CF) to capture shape corner details. To achieve smooth shape interpolation on the learned shape manifold, we establish connections between the two fields for better alignment. We further design a vectorization process to extract high-quality and compact vector fonts from our joint neural representation. Experiments demonstrate that our method can generate more visually appealing vector fonts with a higher level of compactness compared to existing alternatives. Chia-Hao Chen, Ying-Tian Liu, Song-Hai Zhang |
ICCV | 5 |
| 2023 | MBPTrack: Improving 3D Point Cloud Tracking with Memory networks and Box Priorsabstract3D single object tracking has been a crucial problem for decades with numerous applications such as autonomous driving. Despite its wide-ranging use, this task remains challenging due to the significant appearance variation caused by occlusion and size differences among tracked targets. To address these issues, we present MBPTrack, which adopts a Memory mechanism to utilize past information and formulates localization in a coarse-to-fine scheme using Box Priors given in the first frame. Specifically, past frames with targetness masks serve as an extenral memory, and a transformer-based module propagates tracked target cues from the memory to the current frame. To precisely localize objects of all sizes, MBPTrack first predicts the target center via Hough voting. By leveraging box priors given in the first frame, we adaptively sample reference points around the target center that roughly cover the target of different sizes. Then, we obtain dense feature maps by aggregating point features into the reference points, where localization can be performed more effectively. Extensive experiments demonstrate that MBPTrack achieves state-of-the-art performance on KITTI, nuScenes and Waymo Open Dataset, while running at 50 FPS on a single RTX3090 GPU. Tian-Xing Xu, Yukun Lai, Song-Hai Zhang |
ICCV | 4 |
| 2023 | Automatic Generation of Commercial ScenesabstractCommercial scenes such as markets and shops are everyday scenes for both virtual scenes and real-world interior designs. However, existing literature on interior scene synthesis mainly focuses on formulating and optimizing residential scenes such as bedrooms, living rooms, etc. Existing literature typically presents a set of relations among objects. It recognizes each furniture object as the smallest unit while optimizing a residential room. However, object relations become less critical in commercial scenes since shelves are often placed next to each other so pre-calculated relations of objects are less needed. Instead, interior designers resort to evaluating how groups of objects perform in commercial scenes, i.e., the smallest unit to be evaluated is a group of objects. This paper presents a system automatically synthesizes market-like commercial scenes in virtual environments. Following the rules of commercial layout design, we parameterize groups of objects as "patterns" contributing to a scene. Each pattern directly yields a human-centric routine locally, provides potential connectivity with other routines, and derives the arrangements of objects concerning itself according to the assigned parameters. In order to optimize a scene, the patterns are iteratively multiplexed to insert new routines or modify existing ones under a set of constraints derived from commercial layout designs. Through extensive experiments, we demonstrate the ability of our framework to generate plausible and practical commercial scenes. Shao-Kui Zhang, Jia-Hong Liu, Tianyi Xiong, Ke-Xin Ren, Hongbo Fu 0001, Song-Hai Zhang |
ACM Multimedia | 7 |
| 2023 | PanoGRF: Generalizable Spherical Radiance Fields for Wide-baseline PanoramasabstractAchieving an immersive experience enabling users to explore virtual environments with six degrees of freedom (6DoF) is essential for various applications such as virtual reality (VR). Wide-baseline panoramas are commonly used in these applications to reduce network bandwidth and storage requirements. However, synthesizing novel views from these panoramas remains a key challenge. Although existing neural radiance field methods can produce photorealistic views under narrow-baseline and dense image captures, they tend to overfit the training views when dealing with wide-baseline panoramas due to the difficulty in learning accurate geometry from sparse $360^{\circ}$ views. To address this problem, we propose PanoGRF, Generalizable Spherical Radiance Fields for Wide-baseline Panoramas, which construct spherical radiance fields incorporating $360^{\circ}$ scene priors. Unlike generalizable radiance fields trained on perspective images, PanoGRF avoids the information loss from panorama-to-perspective conversion and directly aggregates geometry and appearance features of 3D sample points from each panoramic view based on spherical projection. Moreover, as some regions of the panorama are only visible from one view while invisible from others under wide baseline settings, PanoGRF incorporates $360^{\circ}$ monocular depth priors into spherical depth estimation to improve the geometry features. Experimental results on multiple panoramic datasets demonstrate that PanoGRF significantly outperforms state-of-the-art generalizable view synthesis methods for wide-baseline panoramas (e.g., OmniSyn) and perspective images (e.g., IBRNet, NeuRay). Zheng Chen 0016, Yan-Pei Cao 0001, Chen Wang 0049, Ying Shan, Song-Hai Zhang |
NeurIPS | 6 |
| 2023 | VMesh: Hybrid Volume-Mesh Representation for Efficient View SynthesisabstractWith the emergence of neural radiance fields (NeRFs), view synthesis quality has reached an unprecedented level. Compared to traditional mesh-based assets, this volumetric representation is more powerful in expressing scene geometry but inevitably suffers from high rendering costs and can hardly be involved in further processes like editing, posing significant difficulties in combination with the existing graphics pipeline. In this paper, we present a hybrid volume-mesh representation, VMesh, which depicts an object with a textured mesh along with an auxiliary sparse volume. VMesh retains the advantages of mesh-based assets, such as efficient rendering and compact storage, while also incorporating the ability to represent subtle geometric structures provided by the volumetric counterpart. VMesh can be obtained from multi-view images of an object and renders at 2K 60FPS on common consumer devices with high fidelity, unleashing new opportunities for real-time immersive applications. Yan-Pei Cao 0001, Chen Wang 0049, Yu He 0001, Ying Shan, Song-Hai Zhang |
SIGGRAPH Asia | 6 |
| 2023 | Neural Point-based Volumetric Avatar: Surface-guided Neural Points for Efficient and Photorealistic Volumetric Head AvatarabstractRendering photorealistic and dynamically moving human heads is crucial for ensuring a pleasant and immersive experience in AR/VR and video conferencing applications. However, existing methods often struggle to model challenging facial regions (e.g., mouth interior, eyes, and beard), resulting in unrealistic and blurry results. In this paper, we propose Neural Point-based Volumetric Avatar (NPVA), a method that adopts the neural point representation as well as the neural volume rendering process and discards the predefined connectivity and hard correspondence imposed by mesh-based approaches. Specifically, the neural points are strategically constrained around the surface of the target expression via a high-resolution UV displacement map, achieving increased modeling capacity and more accurate control. We introduce three technical innovations to improve the rendering and training efficiency: a patch-wise depth-guided (shading point) sampling strategy, a lightweight radiance decoding process, and a Grid-Error-Patch (GEP) ray sampling strategy during training. By design, our NPVA is better equipped to handle topologically changing regions and thin structures while also ensuring accurate expression control when animating avatars. Experiments conducted on three subjects from the Multiface dataset demonstrate the effectiveness of our designs, outperforming previous state-of-the-art methods, especially in handling challenging facial regions. Cong Wang 0045, Yan-Pei Cao 0001, Linchao Bao, Ying Shan, Song-Hai Zhang |
SIGGRAPH Asia | 6 |
| 2023 | Redirected Walking Based on Historical User Walking DataabstractWith redirected walking (RDW) technology, people can explore large virtual worlds in smaller physical spaces. RDW controls the trajectory of the user's walking in the physical space through subtle adjustments, so as to minimize the collision between the user and the physical space. Previous predictive algorithms place constraints on the user's path according to the spatial layouts of the virtual environment and work well when applicable, while reactive algorithms are more general for scenarios involving free exploration or uncon-strained movements. However, even in relatively free environments, we can predict the user's walking to a certain extent by analyzing the user's historical walking data, which can help the decision-making of reactive algorithms. This paper proposes a novel RDW method that improves the effect of real-time unrestricted RDW by analyzing and utilizing the user's historical walking data. In this method, the physical space is discretized by considering the user's location and orientation in the physical space. Using the weighted directed graph obtained from the user's historical walking data, we dynamically update the scores of different reachable poses in the physical space during the user's walking. We rank the scores and choose the optimal target position and orientation to guide the user to the best pose. Since simulation experiments have been shown to be effective in many previous RDW studies, we also provide a method to simulate user walking trajectories and generate a dataset. Experiments show that our method outperforms multiple state-of-the-art methods in various environments of different sizes and spatial layouts. Cheng-Wei Fan, Sen-Zhe Xu 0001, Song-Hai Zhang |
VR | 5 |
| 2023 | DeepPortraitDrawing: Generating human body images from freehand sketches
Xian Wu 0004, Chen Wang 0049, Hongbo Fu 0001, Ariel Shamir, Song-Hai Zhang |
Comput. Graph. | 5 |
| 2023 | Focusing on your subject: Deep subject-aware image composition recommendation networksabstractPhoto composition is one of the most important factors in the aesthetics of photographs. As a popular application, composition recommendation for a photo focusing on a specific subject has been ignored by recent deep-learning-based composition recommendation approaches. In this paper, we propose a subject-aware image composition recommendation method, SAC-Net, which takes an RGB image and a binary subject window mask as input, and returns good compositions as crops containing the subject. Our model first determines candidate scores for all possible coarse cropping windows. The crops with high candidate scores are selected and further refined by regressing their corner points to generate the output recommended cropping windows. The final scores of the refined crops are predicted by a final score regression module. Unlike existing methods that need to preset several cropping windows, our network is able to automatically regress cropping windows with arbitrary aspect ratios and sizes. We propose novel stability losses for maximizing smoothness when changing cropping windows along with view changes. Experimental results show that our method outperforms state-of-the-art methods not only on the subject-aware image composition recommendation task, but also for general purpose composition recommendation. We also have designed a multistage labeling scheme so that a large amount of ranked pairs can be produced economically. We use this scheme to propose the first subject-aware composition dataset SACD, which contains 2777 images, and more than 5 million composition ranked pairs. The SACD dataset is publicly available at https://cg.cs.tsinghua.edu.cn/SACD/ . Guo-Ye Yang, Wenyang Zhou, Song-Hai Zhang |
Comput. Vis. Media | 4 |
| 2023 | Learning Local Contrast for Crisp Edge Detection
Xiaonan Fang 0001, Song-Hai Zhang |
J. Comput. Sci. Technol. | 2 |
| 2023 | StructNeRF: Neural Radiance Fields for Indoor Scenes With Structural HintsabstractNeural Radiance Fields (NeRF) achieve photo-realistic view synthesis with densely captured input images. However, the geometry of NeRF is extremely under-constrained given sparse views, resulting in significant degradation of novel view synthesis quality. Inspired by self-supervised depth estimation methods, we propose StructNeRF, a solution to novel view synthesis for indoor scenes with sparse inputs. StructNeRF leverages the structural hints naturally embedded in multi-view inputs to handle the unconstrained geometry issue in NeRF. Specifically, it tackles the texture and non-texture regions respectively: a patch-based multi-view consistent photometric loss is proposed to constrain the geometry of textured regions; for non-textured ones, we explicitly restrict them to be 3D consistent planes. Through the dense self-supervised depth constraints, our method improves both the geometry and the view synthesis performance of NeRF without any additional training on external data. Extensive experiments on several real-world datasets demonstrate that StructNeRF shows superior or comparable performance compared to state-of-the-art methods (e.g. NeRF, DSNeRF, RegNeRF, Dense Depth Priors, MonoSDF, etc.) for indoor scenes with sparse inputs both quantitatively and qualitatively. Zheng Chen 0016, Chen Wang 0049, Song-Hai Zhang |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2023 | Learning Implicit Glyph Shape RepresentationabstractAutomatic generation of fonts can greatly facilitate the font design process, and provide prototypes where designers can draw inspiration from. Existing generation methods are mainly built upon rasterized glyph images to utilize the successful convolutional architecture, but ignore the vector nature of glyph shapes. We present an implicit representation, modeling each glyph as shape primitives enclosed by several quadratic curves. This structured implicit representation is shown to be better suited for glyph modeling, and enables rendering glyph images at arbitrary high resolutions. Our representation gives high-quality glyph reconstruction and interpolation results, and performs well on the challenging one-shot font style transfer task comparing to other alternatives both qualitatively and quantitatively. Ying-Tian Liu, Yi-Xiao Li, Chen Wang 0049, Song-Hai Zhang |
IEEE Trans. Vis. Comput. Graph. | 5 |
| 2023 | On Rotation Gains Within and Beyond Perceptual Limitations for Seated VRabstractHead tracking in head-mounted displays (HMDs) enables users to explore a 360-degree virtual scene with free head movements. However, for seated use of HMDs such as users sitting on a chair or a couch, physically turning around 360-degree is not possible. Redirection techniques decouple tracked physical motion and virtual motion, allowing users to explore virtual environments with more flexibility. In seated situations with only head movements available, the difference of stimulus might cause the detection thresholds of rotation gains to differ from that of redirected walking. Therefore we present an experiment with a two-alternative forced-choice (2AFC) design to compare the thresholds for seated and standing situations. Results indicate that users are unable to discriminate rotation gains between 0.89 and 1.28, a smaller range compared to the standing condition. We further treated head amplification as an interaction technique and found that a gain of 2.5, though not a hard threshold, was near the largest gain that users consider applicable. Overall, our work aims to better understand human perception of rotation gains in seated VR and the results provide guidance for future design choices of its applications. Chen Wang 0049, Song-Hai Zhang, Yizhuo Zhang 0001, Stefanie Zollmann, Shi-Min Hu 0001 |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2023 | One-Step Out-of-Place Resetting for Redirected Walking in VRabstractRedirected walking (RDW) allows users to explore virtual environments in limited physical spaces by imperceptibly steering them away from obstacles and space boundaries. However, even with those techniques, the risk of collision cannot always be avoided. For such situations, resetting techniques have been proposed to provide an immediate adjustment of the physical walking direction of a user. Existing resetting techniques are either applied in-place, where the user changes orientation but stays in the same position or out-of-place methods where the user is guided to move from the current position to a safe location all while freezing the movement in the virtual world. While out-of-place methods have the potential to provide more freedom to user movements after resetting, current out-of-place methods do not provide enough guidance for the users to move to optimal locations. In this work, we propose a novel out-of-place resetting strategy that guides users to optimal physical locations with the most potential for free movement and a smaller amount of resetting required for their further movements. For this purpose, we calculate a heat map of the walking area according to the average walking distance using a simulation of the currently used RDW algorithm. Based on this heat map, we identify the most suitable position for a one-step reset within a predefined searching range and use this one as the reset point. Our results show that our method increases the average moving distance within one cycle of resetting. Furthermore, our resetting method can be applied to any physical area with obstacles. That means that RDW methods that were not suitable for such environments (e.g., Steer to Center) combined with our resetting can also be extended to such complex walking areas. In addition, we present a user interface to provide a similar visual experience between these methods, using a two-arrows indicator to help users adjust their position and direction. Song-Hai Zhang, Chia-Hao Chen, Stefanie Zollmann |
IEEE Trans. Vis. Comput. Graph. | 1 |
| 2023 | Adaptive Optimization Algorithm for Resetting Techniques in Obstacle-Ridden EnvironmentsabstractRedirected Walking (RDW) algorithms aim to impose several types of gains on users immersed in Virtual Reality and distort their walking paths in the real world, thus enabling them to explore a larger space. Since collision with physical boundaries is inevitable, a reset strategy needs to be provided to allow users to reset when they hit the boundary. However, most reset strategies are based on simple heuristics by choosing a seemingly suitable solution, which may not perform well in practice. In this article, we propose a novel optimization-based reset algorithm adaptive to different RDW algorithms. Inspired by the approach of finite element analysis, our algorithm splits the boundary of the physical world by a set of endpoints. Each endpoint is assigned a reset vector to represent the optimized reset direction when hitting the boundary. The reset vectors on the edge will be determined by the interpolation between two neighbouring endpoints. We conduct simulation-based experiments for three RDW algorithms with commonly used reset algorithms to compare with. The results demonstrate that the proposed algorithm significantly reduces the number of resets. Song-Hai Zhang, Chia-Hao Chen, Fu Zheng, Yongliang Yang 0002, Shi-Min Hu 0001 |
IEEE Trans. Vis. Comput. Graph. | 1 |
| 2023 | SceneViewer: Automating Residential Photography in Virtual EnvironmentsabstractSelecting views is one of the most common but overlooked procedures in topics related to 3D scenes. Typically, existing applications and researchers manually select views through a trial-and-error process or "preset" a direction, such as the top-down views. For example, literature for scene synthesis requires views for visualizing scenes. Research on panorama and VR also require initial placements for cameras, etc. This article presents SceneViewer, an integrated system for automatic view selections. Our system is achieved by applying rules of interior photography, which guides potential views and seeks better views. Through experiments and applications, we show the potentiality and novelty of the proposed method. Shao-Kui Zhang, Hou Tam, Yi-Xiao Li, Tai-Jiang Mu, Song-Hai Zhang |
IEEE Trans. Vis. Comput. Graph. | 5 |
| 2022 | NeRFReN: Neural Radiance Fields with ReflectionsabstractNeural Radiance Fields (NeRF) has achieved unprece-dented view synthesis quality using coordinate-based neu-ral scene representations. However, NeRF's view depen-dency can only handle simple reflections like highlights but cannot deal with complex reflections such as those from glass and mirrors. In these scenarios, NeRF models the virtual image as real geometries which leads to inaccurate depth estimation, and produces blurry renderings when the multi-view consistency is violated as the reflected objects may only be seen under some of the viewpoints. To over-come these issues, we introduce NeRFReN, which is built upon NeRF to model scenes with reflections. Specifically, we propose to split a scene into transmitted and reflected components, and model the two components with separate neural radiance fields. Considering that this decomposition is highly under-constrained, we exploit geometric priors and apply carefully-designed training strategies to achieve reasonable decomposition results. Experiments on various self-captured scenes show that our method achieves high-quality novel view synthesis and physically sound depth es-timation results while enabling scene editing applications. Linchao Bao, Yu He 0001, Song-Hai Zhang |
CVPR | 5 |
| 2022 | Face2Faceρ: Real-Time High-Resolution One-Shot Face Reenactment
Daoliang Guo, Song-Hai Zhang |
ECCV (13) | 4 |
| 2022 | NeRF-SR: High Quality Neural Radiance Fields using SupersamplingabstractWe present NeRF-SR, a solution for high-resolution (HR) novel view synthesis with mostly low-resolution (LR) inputs. Our method is built upon Neural Radiance Fields (NeRF) that predicts per-point density and color with a multi-layer perceptron. While producing images at arbitrary scales, NeRF struggles with resolutions that go beyond observed images. Our key insight is that NeRF benefits from 3D consistency, which means an observed pixel absorbs information from nearby views. We first exploit it by a super-sampling strategy that shoots multiple rays at each image pixel, which further enforces multi-view constraint at a sub-pixel level. Then, we show that NeRF-SR can further boost the performance of super-sampling by a refinement network that leverages the estimated depth at hand to hallucinate details from related patches on only one HR reference image. Experiment results demonstrate that NeRF-SR generates high-quality results for novel view synthesis at HR on both synthetic and real-world datasets without any external information. Project page: https://cwchenwang.github.io/NeRF-SR Chen Wang 0049, Xian Wu 0004, Song-Hai Zhang, Yu-Wing Tai, Shi-Min Hu 0001 |
ACM Multimedia | 4 |
| 2022 | Optimal Pose Guided Redirected Walking with Pose Score PrecomputationabstractRedirected walking (RDW) aims to reduce the collisions in the physical space for VR applications. However, most of the previous RDW methods do not consider future possibilities of collisions after imperceptibly redirecting users. In this paper, we combine the subtle RDW methods and reset strategy in our method design and propose a novel solution for RDW that can make better use of physical space and trigger fewer resets. The key idea of our method is to discretize the representation of possible user positions and orientations by a series of standard poses and rate them based on the possibilities of hitting obstacles of their reachable poses. A transfer path algorithm is proposed to measure the accessibility among standard poses and is used to support the calculation of the scores of standard poses. Using our method, the user can be redirected imperceptibly to the optimal pose with the best score among all the reachable poses from the user’s current pose during walking. Experiments demonstrate that our method outperforms state-of-the-art methods in various environment sizes and obstacle layouts. Sen-Zhe Xu 0001, Tian Lv, Guangrong He, Chia-Hao Chen, Song-Hai Zhang |
VR | 6 |
| 2022 | Attention mechanisms in computer vision: A surveyabstractHumans can naturally and effectively find salient regions in complex scenes. Motivated by this observation, attention mechanisms were introduced into computer vision with the aim of imitating this aspect of the human visual system. Such an attention mechanism can be regarded as a dynamic weight adjustment process based on features of the input image. Attention mechanisms have achieved great success in many visual tasks, including image classification, object detection, semantic segmentation, video understanding, image generation, 3D vision, multimodal tasks, and self-supervised learning. In this survey, we provide a comprehensive review of various attention mechanisms in computer vision and categorize them according to approach, such as channel attention, spatial attention, temporal attention, and branch attention; a related repository https://github.com/MenghaoGuo/Awesome-Vision-Attentions is dedicated to collecting related work. We also suggest future directions for attention mechanism research. Menghao Guo 0001, Tian-Xing Xu, Jiang-Jiang Liu 0001, Zheng-Ning Liu, Peng-Tao Jiang, Tai-Jiang Mu, Song-Hai Zhang, Ralph R. Martin, Ming-Ming Cheng, Shi-Min Hu 0001 |
Comput. Vis. Media | 7 |
| 2022 | Smoothness preserving layout for dynamic labels by hybrid optimizationabstractStable label movement and smooth label trajectory are critical for effective information understanding. Sudden label changes cannot be avoided by whatever forced directed methods due to the unreliability of resultant force or global optimization methods due to the complex trade-off on the different aspects. To solve this problem, we proposed a hybrid optimization method by taking advantages of the merits of both approaches. We first detect the spatial-temporal intersection regions from whole trajectories of the features, and initialize the layout by optimization in decreasing order by the number of the involved features. The label movements between the spatial-temporal intersection regions are determined by force directed methods. To cope with some features with high speed relative to neighbors, we introduced a force from future, called temporal force, so that the labels of related features can elude ahead of time and retain smooth movements. We also proposed a strategy by optimizing the label layout to predict the trajectories of features so that such global optimization method can be applied to streaming data. ELECTRONIC SUPPLEMENTARY MATERIAL: Supplementary material is available in the online version of this article at 10.1007/s41095-021-0231-y. Yu He 0001, Song-Hai Zhang |
Comput. Vis. Media | 3 |
| 2022 | Deep image synthesis from intuitive user input: A review and perspectivesabstractIn many applications of computer graphics, art, and design, it is desirable for a user to provide intuitive non-image input, such as text, sketch, stroke, graph, or layout, and have a computer system automatically generate photo-realistic images according to that input. While classically, works that allow such automatic image content generation have followed a framework of image retrieval and composition, recent advances in deep generative models such as generative adversarial networks (GANs), variational autoencoders (VAEs), and flow-based methods have enabled more powerful and versatile image generation approaches. This paper reviews recent works for image synthesis given intuitive user input, covering advances in input versatility, image generation methodology, benchmark datasets, and evaluation metrics. This motivates new perspectives on input representation and interactivity, cross fertilization between major image generation paradigms, and evaluation and comparison of generation methods. Yuan Xue 0002, Han Zhang 0010, Tao Xu 0029, Song-Hai Zhang, Sharon X. Huang |
Comput. Vis. Media | 5 |
| 2022 | Local Homography Estimation on User-Specified Textureless Regions
Zheng Chen 0016, Xiaonan Fang 0001, Song-Hai Zhang |
J. Comput. Sci. Technol. | 3 |
| 2022 | User-Guided Deep Human Image Matting Using Arbitrary TrimapsabstractImage matting is widely studied for accurate foreground extraction. Most algorithms, including deep-learning based solutions, require a carefully edited trimap. Recent works attempt to combine the segmentation stage and matting stage in one CNN model, but errors occurring at the segmentation stage lead to unsatisfactory matte. We propose a user-guided approach for practical human matting. More precisely, we provide a good automatic initial matting and a natural way of interaction that reduces the workload of drawing trimaps and allows users to guide the matting in ambiguous situation. We also combine the segmentation and matting stage in an end-to-end CNN architecture and introduce a residual-learning module to support convenient stroke-based interaction. The proposed model learns to propagate the input trimap and modify the deep image features, which can efficiently correct the segmentation errors. Our model supports arbitrary forms of trimaps from carefully edited to totally unknown maps. Our model also allows users to choose from different foreground estimations according to their preference. We collected a large human matting dataset consisting of 12K real-world human images with complex background and human-object relations. The proposed model is trained on the new dataset with a novel trimap generation strategy that enables the model to tackle different test situations and highly improves the interaction efficiency. Our method outperforms other state-of-the-art automatic methods and achieve competitive accuracy when high-quality trimaps are provided. Experiments indicate that our interactive matting strategy is superior to separately estimating the trimap and alpha matte using two models. Xiaonan Fang 0001, Song-Hai Zhang, Tao Chen 0015, Xian Wu 0004, Ariel Shamir, Shi-Min Hu 0001 |
IEEE Trans. Image Process. | 2 |
| 2022 | Making Resets away from Targets: POI aware Redirected WalkingabstractRapidly developing Redirected Walking (ROW) technologies have enabled VR applications to immerse users in large virtual environments (VE) while actually walking in relatively small physical environments (PE). When an unavoidable collision emerges in a PE, the ROW controller suspends the user's immersive experience and resets the user to a new direction in PE. Existing ROW methods mainly aim to reduce the number of resets. However, from the perspective of the user experience, when users are about to reach a point of interest (POI) in a VE, reset interruptions are more likely to have an impact on user experience. In this paper, we propose a new ROW method, aiming to keep resets occurring at a longer distance from the virtual target, as well as to reduce the number of resets. Simulation experiments and real user studies demonstrate that our method outperforms state-of-the-art ROW methods in the number of resets and dramatically increases the distance between the reset locations and the virtual targets. Sen-Zhe Xu 0001, Jia-Hong Liu, Stefanie Zollmann, Song-Hai Zhang |
IEEE Trans. Vis. Comput. Graph. | 5 |
| 2022 | Towards Better Caption Supervision for Object DetectionabstractAs training high-performance object detectors requires expensive bounding box annotations, recent methods resort to free-available image captions. However, detectors trained on caption supervision perform poorly because captions are usually noisy and cannot provide precise location information. To tackle this issue, we present a visual analysis method, which tightly integrates caption supervision with object detection to mutually enhance each other. In particular, object labels are first extracted from captions, which are utilized to train the detectors. Then, the objects detected from images are fed into caption supervision for further improvement. To effectively loop users into the object detection process, a node-link-based set visualization supported by a multi-type relational co-clustering algorithm is developed to explain the relationships between the extracted labels and the images with detected objects. The co-clustering algorithm clusters labels and images simultaneously by utilizing both their representations and their relationships. Quantitative evaluations and a case study are conducted to demonstrate the efficiency and effectiveness of the developed method in improving the performance of object detectors. Changjian Chen, Jing Wu 0004, Shouxing Xiang, Song-Hai Zhang, Qifeng Tang, Shixia Liu |
IEEE Trans. Vis. Comput. Graph. | 5 |
| 2022 | Context-Consistent Generation of Indoor Virtual Environments Based on Geometry ConstraintsabstractIn this article, we propose a system that can automatically generate immersive and interactive virtual reality (VR) scenes by taking real-world geometric constraints into account. Our system can not only help users avoid real-world obstacles in virtual reality experiences, but also provide context-consistent contents to preserve their sense of presence. To do so, our system first identifies the positions and bounding boxes of scene objects as well as a set of interactive planes from 3D scans. Then context-consistent virtual objects that have similar geometric properties to the real ones can be automatically selected and placed into the virtual scene, based on learned object association relations and layout patterns from large amounts of indoor scene configurations. We regard virtual object replacement as a combinatorial optimization problem, considering both geometric and contextual consistency constraints. Quantitative and qualitative results show that our system can generate plausible interactive virtual scenes that highly resemble real environments, and have the ability to keep the sense of presence for users in their VR experiences. Yu He 0001, Ying-Tian Liu, Yi-Han Jin, Song-Hai Zhang, Yukun Lai, Shi-Min Hu 0001 |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2022 | Fast 3D Indoor Scene Synthesis by Learning Spatial Relation Priors of ObjectsabstractWe present a framework for fast synthesizing indoor scenes, given a room geometry and a list of objects with learnt priors. Unlike existing data-driven solutions, which often learn priors by co-occurrence analysis and statistical model fitting, our method measures the strengths of spatial relations by tests for complete spatial randomness (CSR), and learns discrete priors based on samples with the ability to accurately represent exact layout patterns. With the learnt priors, our method achieves both acceleration and plausibility by partitioning the input objects into disjoint groups, followed by layout optimization using position-based dynamics (PBD) based on the Hausdorff metric. Experiments show that our framework is capable of measuring more reasonable relations among objects and simultaneously generating varied arrangements in seconds compared with the state-of-the-art works. Song-Hai Zhang, Shao-Kui Zhang, Weiyu Xie, Yongliang Yang 0002, Hongbo Fu 0001 |
IEEE Trans. Vis. Comput. Graph. | 1 |
| 2021 | Sketch2Model: View-Aware 3D Modeling From Single Free-Hand SketchesabstractWe investigate the problem of generating 3D meshes from single free-hand sketches, aiming at fast 3D modeling for novice users. It can be regarded as a single-view reconstruction problem, but with unique challenges, brought by the variation and conciseness of sketches. Ambiguities in poorly-drawn sketches could make it hard to determine how the sketched object is posed. In this paper, we address the importance of viewpoint specification for overcoming such ambiguities, and propose a novel view-aware generation approach. By explicitly conditioning the generation process on a given viewpoint, our method can generate plausible shapes automatically with predicted viewpoints, or with specified viewpoints to help users better express their intentions. Extensive evaluations on various datasets demonstrate the effectiveness of our view-aware design in solving sketch ambiguities and improving reconstruction quality. Song-Hai Zhang, Qing-Wen Gu |
CVPR | 1 |
| 2021 | MageAdd: Real-Time Interaction Simulation for Scene SynthesisabstractWhile recent researches on computational 3D scene synthesis have achieved impressive results, automatically synthesized scenes do not guarantee satisfaction of end users. On the other hand, manual scene modelling can always ensure high quality, but requires a cumbersome trial-and-error process. In this paper, we bridge the above gap by presenting a data-driven 3D scene synthesis framework that can intelligently infer objects to the scene by incorporating and simulating user preferences with minimum input. While the cursor is moved and clicked in the scene, our framework automatically selects and transforms suitable objects into scenes in real time. This is based on priors learnt from the dataset for placing different types of objects, and updated according to the current scene context. Through extensive experiments we demonstrate that our framework outperforms the state-of-the-art on result aesthetics, and enables effective and efficient user interactions. Shao-Kui Zhang, Yi-Xiao Li, Yu He 0001, Yongliang Yang 0002, Song-Hai Zhang |
ACM Multimedia | 5 |
| 2021 | Geometry-Based Layout Generation with Hyper-Relations AMONG Objects
Shao-Kui Zhang, Weiyu Xie, Song-Hai Zhang |
Graph. Model. | 3 |
| 2021 | ChoreoMaster: choreography-oriented music-driven dance synthesisabstractDespite strong demand in the game and film industry, automatically synthesizing high-quality dance motions remains a challenging task. In this paper, we present ChoreoMaster, a production-ready music-driven dance motion synthesis system. Given a piece of music, ChoreoMaster can automatically generate a high-quality dance motion sequence to accompany the input music in terms of style, rhythm and structure. To achieve this goal, we introduce a novel choreography-oriented choreomusical embedding framework, which successfully constructs a unified choreomusical embedding space for both style and rhythm relationships between music and dance phrases. The learned choreomusical embedding is then incorporated into a novel choreography-oriented graph-based motion synthesis framework, which can robustly and efficiently generate high-quality dance motions following various choreographic rules. Moreover, as a production-ready system, ChoreoMaster is sufficiently controllable and comprehensive for users to produce desired results. Experimental results demonstrate that dance motions generated by ChoreoMaster are accepted by professional artists. Jin Lei, Song-Hai Zhang, Shi-Min Hu 0001 |
ACM Trans. Graph. | 4 |
| 2021 | MoCap-solver: a neural solver for optical motion capture dataabstractIn a conventional optical motion capture (MoCap) workflow, two processes are needed to turn captured raw marker sequences into correct skeletal animation sequences. Firstly, various tracking errors present in the markers must be fixed ( cleaning or refining ). Secondly, an agent skeletal mesh must be prepared for the actor/actress, and used to determine skeleton information from the markers ( re-targeting or solving ). The whole process, normally referred to as solving MoCap data, is extremely time-consuming, labor-intensive, and usually the most costly part of animation production. Hence, there is a great demand for automated tools in industry. In this work, we present MoCap-Solver, a production-ready neural solver for optical MoCap data. It can directly produce skeleton sequences and clean marker sequences from raw MoCap markers, without any tedious manual operations. To achieve this goal, our key idea is to make use of neural encoders concerning three key intrinsic components: the template skeleton, marker configuration and motion, and to learn to predict these latent vectors from imperfect marker sequences containing noise and errors. By decoding these components from latent vectors, sequences of clean markers and skeletons can be directly recovered. Moreover, we also provide a novel normalization strategy based on learning a pose-dependent marker reliability function, which greatly improves system robustness. Experimental results demonstrate that our algorithm consistently outperforms the state-of-the-art on both synthetic and real-world datasets. Yupan Wang, Song-Hai Zhang, Sen-Zhe Xu 0001, Shi-Min Hu 0001 |
ACM Trans. Graph. | 3 |
| 2020 | Visualization of COVID-19 spread based on spread and extinction indexes
Song-Hai Zhang |
Sci. China Inf. Sci. | 1 |
| 2020 | What and where: A context-based recommendation system for object insertionabstractWe propose a novel problem revolving around two tasks: (i) given a scene, recommend objects to insert, and (ii) given an object category, retrieve suitable background scenes. A bounding box for the inserted object is predicted in both tasks, which helps downstream applications such as semiautomated advertising and video composition. The major challenge lies in the fact that the target object is neither present nor localized in the input, and furthermore, available datasets only provide scenes with existing objects. To tackle this problem, we build an unsupervised algorithm based on object-level contexts, which explicitly models the joint probability distribution of object categories and bounding boxes using a Gaussian mixture model. Experiments on our own annotated test set demonstrate that our system outperforms existing baselines on all sub-tasks, and does so using a unified framework. Future extensions and applications are suggested. Song-Hai Zhang, Zhengping Zhou, Xi Dong, Peter Hall 0001 |
Comput. Vis. Media | 1 |
| 2020 | A new dataset of dog breed images and a benchmark for finegrained classificationabstractIn this paper, we introduce an image dataset for fine-grained classification of dog breeds: the Tsinghua Dogs Dataset. It is currently the largest dataset for fine-grained classification of dogs, including 130 dog breeds and 70,428 real-world images. It has only one dog in each image and provides annotated bounding boxes for the whole body and head. In comparison to previous similar datasets, it contains more breeds and more carefully chosen images for each breed. The diversity within each breed is greater, with between 200 and 7000+ images for each breed. Annotation of the whole body and head makes the dataset not only suitable for the improvement of finegrained image classification models based on overall features, but also for those locating local informative parts. We show that dataset provides a tough challenge by benchmarking several state-of-the-art deep neural models. The dataset is available for academic purposes at https://cg.cs.tsinghua.edu.cn/ThuDogs/ . Ding-Nan Zou, Song-Hai Zhang, Tai-Jiang Mu |
Comput. Vis. Media | 2 |
| 2019 | Example-Guided Style-Consistent Image Synthesis From Semantic LabelingabstractExample-guided image synthesis aims to synthesize an image from a semantic label map and an exemplary image indicating style. We use the term "style" in this problem to refer to implicit characteristics of images, for example: in portraits "style" includes gender, racial identity, age, hairstyle; in full body pictures it includes clothing; in street scenes it refers to weather and time of day and such like. A semantic label map in these cases indicates facial expression, full body pose, or scene segmentation. We propose a solution to the example-guided image synthesis problem using conditional generative adversarial networks with style consistency. Our key contributions are (i) a novel style consistency discriminator to determine whether a pair of images are consistent in style; (ii) an adaptive semantic consistency loss; and (iii) a training data sampling strategy, for synthesizing style-consistent results to the exemplar. We demonstrate the efficiency of our method on face, dance and street view synthesis tasks. Miao Wang 0004, Guo-Ye Yang, Ruilong Li, Runze Liang, Song-Hai Zhang, Peter Hall 0001, Shi-Min Hu 0001 |
CVPR | 5 |
| 2019 | Pose2Seg: Detection Free Human Instance SegmentationabstractThe standard approach to image instance segmentation is to perform the object detection first, and then segment the object from the detection bounding-box. More recently, deep learning methods like Mask R-CNN perform them jointly. However, little research takes into account the uniqueness of the "human" category, which can be well defined by the pose skeleton. Moreover, the human pose skeleton can be used to better distinguish instances with heavy occlusion than using bounding-boxes. In this paper, we present a brand new pose-based instance segmentation framework for humans which separates instances based on human pose, rather than proposal region detection. We demonstrate that our pose-based framework can achieve better accuracy than the state-of-art detection-based approach on the human instance segmentation problem, and can moreover better handle occlusion. Furthermore, there are few public datasets containing many heavily occluded humans along with comprehensive annotations, which makes this a challenging problem seldom noticed by researchers. Therefore, in this paper we introduce a new benchmark "Occluded Human (OCHuman)", which focuses on occluded humans with comprehensive annotations including bounding-box, human pose and instance masks. This dataset contains 8110 detailed annotated human instances within 4731 images. With an average 0.67 MaxIoU for each person, OCHuman is the most complex and challenging dataset related to human instance segmentation. Through this dataset, we want to emphasize occlusion as a challenging problem for researchers to study. Song-Hai Zhang, Ruilong Li, Paul L. Rosin, Zixi Cai, Dingcheng Yang, Hao-Zhi Huang 0001, Shi-Min Hu 0001 |
CVPR | 1 |
| 2019 | PortraitNet: Real-time portrait segmentation network for mobile device
Song-Hai Zhang, Ruilong Li, Yongliang Yang 0002 |
Comput. Graph. | 1 |
| 2019 | A Survey of 3D Indoor Scene Synthesis
Song-Hai Zhang, Shao-Kui Zhang, Peter Hall 0001 |
J. Comput. Sci. Technol. | 1 |
| 2019 | Learning guidelines for automatic indoor scene design
Song-Hai Zhang, Ralph R. Martin |
Multim. Tools Appl. | 2 |
| 2019 | Deep Online Video Stabilization With Multi-Grid Warping Transformation LearningabstractVideo stabilization techniques are essential for most hand-held captured videos due to high-frequency shakes. Several 2D, 2.5D and 3D-based stabilization techniques have been presented previously, but to our knowledge, no solutions based on deep neural networks had been proposed to date. The main reason for this omission is shortage in training data as well as the challenge of modeling the problem using neural networks. In this paper, we present a video stabilization technique using a convolutional neural network. Previous works usually propose an offline algorithm that smoothes a holistic camera path based on feature matching. Instead, we focus on low-latency, real-time camera path smoothing, that does not explicitly represent the camera path, and does not use future frames. Our neural network model, called StabNet, learns a set of mesh-grid transformations progressively for each input frame from the previous set of stabalized camera frames, and creates stable corresponding latent camera paths implicitly. To train the network, we collect a dataset of synchronized steady and unsteady video pairs via a specially designed hand-held hardware. Experimental results show that our proposed online method performs comparatively to traditional offline video stabilization methods without using future frames, while running about 10× faster. More importantly, our proposed StabNet is able to handle low-quality videos such as night-scene videos, watermarked videos, blurry videos and noisy videos, where existing methods fail in feature extraction or matching. Miao Wang 0004, Guo-Ye Yang, Jin-Kun Lin, Song-Hai Zhang, Ariel Shamir, Shao-Ping Lu, Shi-Min Hu 0001 |
IEEE Trans. Image Process. | 4 |
| 2018 | Knowledge graph construction with structure and parameter learning for indoor scene designabstractWe consider the problem of learning a representation of both spatial relations and dependencies between objects for indoor scene design. We propose a novel knowledge graph framework based on the entity-relation model for representation of facts in indoor scene design, and further develop a weaklysupervised algorithm for extracting the knowledge graph representation from a small dataset using both structure and parameter learning. The proposed framework is flexible, transferable, and readable. We present a variety of computer-aided indoor scene design applications using this representation, to show the usefulness and robustness of the proposed framework. Song-Hai Zhang, Yukun Lai, Tai-Jiang Mu |
Comput. Vis. Media | 3 |
| 2018 | Traffic signal detection and classification in street views using an attention modelabstractDetecting small objects is a challenging task. We focus on a special case: the detection and classification of traffic signals in street views. We present a novel framework that utilizes a visual attention model to make detection more efficient, without loss of accuracy, and which generalizes. The attention model is designed to generate a small set of candidate regions at a suitable scale so that small targets can be better located and classified. In order to evaluate our method in the context of traffic signal detection, we have built a traffic light benchmark with over 15,000 traffic light instances, based on Tencent street view panoramas. We have tested our method both on the dataset we have built and the Tsinghua-Tencent 100K (TT100K) traffic sign benchmark. Experiments show that our method has superior detection performance and is quicker than the general faster RCNN object detection framework on both datasets. It is competitive with state-of-the-art specialist traffic sign detectors on TT100K, but is an order of magnitude faster. To show generality, we tested it on the LISA dataset without tuning, and obtained an average precision in excess of 90%. Yi-Fan Lu, Jiaming Lu, Song-Hai Zhang, Peter Hall 0001 |
Comput. Vis. Media | 3 |
| 2018 | Hyper-Lapse From Multiple Spatially-Overlapping VideosabstractHyper-lapse video with high speed-up rate is an efficient way to overview long videos, such as a human activity in first-person view. Existing hyper-lapse video creation methods produce a fast-forward video effect using only one video source. In this paper, we present a novel hyper-lapse video creation approach based on multiple spatially-overlapping videos. We assume the videos share a common view or location, and find transition points where jumps from one video to another may occur. We represent the collection of videos using a hyper-lapse transition graph; the edges between nodes represent possible hyper-lapse frame transitions. To create a hyper-lapse video, a shortest path search is performed on this digraph to optimize frame sampling and assembly simultaneously. Finally, we render the hyper-lapse results using video stabilization and appearance smoothing techniques on the selected frames. Our technique can synthesize novel virtual hyper-lapse routes, which may not exist originally. We show various application results on both indoor and outdoor video collections with static scenes, moving objects, and crowds. Miao Wang 0004, Jun-Bang Liang, Song-Hai Zhang, Shao-Ping Lu, Ariel Shamir, Shi-Min Hu 0001 |
IEEE Trans. Image Process. | 3 |
| 2018 | A Comparative Study of Algorithms for Realtime Panoramic Video BlendingabstractUnlike image blending algorithms, video blending algorithms have been little studied. In this paper, we investigate 6 popular blending algorithms-feather blending, multi-band blending, modified Poisson blending, mean value coordinate blending, multi-spline blending and convolution pyramid blending. We consider their application to blending realtime panoramic videos, a key problem in various virtual reality tasks. To evaluate the performances and suitabilities of the 6 algorithms for this problem, we have created a video benchmark with several videos captured under various conditions. We analyze the time and memory needed by the above 6 algorithms, for both CPU and GPU implementations (where readily parallelizable). The visual quality provided by these algorithms is also evaluated both objectively and subjectively. The video benchmark and algorithm implementations are publicly available1. Zhe Zhu, Jiaming Lu, Minxuan Wang, Song-Hai Zhang, Ralph R. Martin, Hantao Liu, Shi-Min Hu 0001 |
IEEE Trans. Image Process. | 4 |
| 2018 | PhotoRecomposer: Interactive Photo Recomposition by CroppingabstractWe present a visual analysis method for interactively recomposing a large number of photos based on example photos with high-quality composition. The recomposition method is formulated as a matching problem between photos. The key to this formulation is a new metric for accurately measuring the composition distance between photos. We have also developed an earth-mover-distance-based online metric learning algorithm to support the interactive adjustment of the composition distance based on user preferences. To better convey the compositions of a large number of example photos, we have developed a multi-level, example photo layout method to balance multiple factors such as compactness, aspect ratio, composition distance, stability, and overlaps. By introducing an EulerSmooth-based straightening method, the composition of each photos is clearly displayed. The effectiveness and usefulness of the method has been demonstrated by the experimental results, user study, and case studies. Xiting Wang, Song-Hai Zhang, Shi-Min Hu 0001, Shixia Liu |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2017 | Avoiding bleeding in image blendingabstractThough elegant in mathematical formulation, gradient-domain image blending suffers from bleeding artefacts in real world applications. We propose an image blending algorithm that avoids bleeding artefacts while preserving the good properties of gradient-domain blending, such as smooth transitions between the candidate regions. Our key idea to finesse the non-smooth boundary difference calculation that causes bleeding artefacts is to use local patch differences. While most previous gradient-domain blending algorithms change one region to fit the other, to further reduce bleeding, we perform bidirectional blending so that both regions change simultaneously. Our blending algorithm is fast: when applied to image stitching, it can achieve 20 fps at 4K resolution. Source code and test images are publicly available. Minxuan Wang, Zhe Zhu, Song-Hai Zhang, Ralph R. Martin, Shi-Min Hu 0001 |
ICIP | 3 |
| 2017 | Practical automatic background substitution for live videoabstractIn this paper we present a novel automatic background substitution approach for live video. The objective of background substitution is to extract the foreground from the input video and then combine it with a new background. In this paper, we use a color line model to improve the Gaussian mixture model in the background cut method to obtain a binary foreground segmentation result that is less sensitive to brightness differences. Based on the high quality binary segmentation results, we can automatically create a reliable trimap for alpha matting to refine the segmentation boundary. To make the composition result more realistic, an automatic foreground color adjustment step is added to make the foreground look consistent with the new background. Compared to previous approaches, our method can produce higher quality binary segmentation results, and to the best of our knowledge, this is the first time such an automatic and integrated background substitution system has been proposed which can run in real time, which makes it practical for everyday applications. Hao-Zhi Huang 0001, Xiaonan Fang 0001, Yufei Ye 0001, Song-Hai Zhang, Paul L. Rosin |
Comput. Vis. Media | 4 |
| 2017 | Intelligent Visual Media Processing: When Graphics Meets Vision
Ming-Ming Cheng, Qibin Hou, Song-Hai Zhang, Paul L. Rosin |
J. Comput. Sci. Technol. | 3 |
| 2016 | Traffic-Sign Detection and Classification in the WildabstractAlthough promising results have been achieved in the areas of traffic-sign detection and classification, few works have provided simultaneous solutions to these two tasks for realistic real world images. We make two contributions to this problem. Firstly, we have created a large traffic-sign benchmark from 100000 Tencent Street View panoramas, going beyond previous benchmarks. It provides 100000 images containing 30000 traffic-sign instances. These images cover large variations in illuminance and weather conditions. Each traffic-sign in the benchmark is annotated with a class label, its bounding box and pixel mask. We call this benchmark Tsinghua-Tencent 100K. Secondly, we demonstrate how a robust end-to-end convolutional neural network (CNN) can simultaneously detect and classify trafficsigns. Most previous CNN image processing solutions target objects that occupy a large proportion of an image, and such networks do not work well for target objects occupying only a small fraction of an image like the traffic-signs here. Experimental results show the robustness of our network and its superiority to alternatives. The benchmark, source code and the CNN model introduced in this paper is publicly available1. Zhe Zhu, Dun Liang, Song-Hai Zhang, Sharon X. Huang, Baoli Li 0004, Shi-Min Hu 0001 |
CVPR | 3 |
| 2016 | Comfort-driven disparity adjustment for stereoscopic videoabstractPixel disparity—the offset of corresponding pixels between left and right views—is a crucial parameter in stereoscopic three-dimensional (S3D) video, as it determines the depth perceived by the human visual system (HVS). Unsuitable pixel disparity distribution throughout an S3D video may lead to visual discomfort. We present a unified and extensible stereoscopic video disparity adjustment framework which improves the viewing experience for an S3D video by keeping the perceived 3D appearance as unchanged as possible while minimizing discomfort. We first analyse disparity and motion attributes of S3D video in general, then derive a wide-ranging visual discomfort metric from existing perceptual comfort models. An objective function based on this metric is used as the basis of a hierarchical optimisation method to find a disparity mapping function for each input video frame. Warping-based disparity manipulation is then applied to the input video to generate the output video, using the desired disparity mappings as constraints. Our comfort metric takes into account disparity range, motion , and stereoscopic window violation ; the framework could easily be extended to use further visual comfort models. We demonstrate the power of our approach using both animated cartoons and real S3D videos. Miao Wang 0004, Xi-Jin Zhang, Jun-Bang Liang, Song-Hai Zhang, Ralph R. Martin |
Comput. Vis. Media | 4 |
| 2016 | Multi-Task Learning for Food Identification and Analysis with Deep Convolutional Neural Networks
Xi-Jin Zhang, Yi-Fan Lu, Song-Hai Zhang |
J. Comput. Sci. Technol. | 3 |
| 2014 | Learning Natural Colors for Image RecoloringabstractAbstract We present a data‐driven method for automatically recoloring a photo to enhance its appearance or change a viewer's emotional response to it. A compact representation called a RegionNet summarizes color and geometric features of image regions, and geometric relationships between them. Correlations between color property distributions and geometric features of regions are learned from a database of well‐colored photos. A probabilistic factor graph model is used to summarize distributions of color properties and generate an overall probability distribution for color suggestions. Given a new input image, we can generate multiple recolored results which unlike previous automatic results, are both natural and artistic, and compatible with their spatial arrangements. Hao-Zhi Huang 0001, Song-Hai Zhang, Ralph R. Martin, Shi-Min Hu 0001 |
Comput. Graph. Forum | 2 |
| 2013 | Timeline Editing of Objects in VideoabstractWe present a video editing technique based on changing the timelines of individual objects in video, which leaves them in their original places but puts them at different times. This allows the production of object-level slow motion effects, fast motion effects, or even time reversal. This is more flexible than simply applying such effects to whole frames, as new relationships between objects can be created. As we restrict object interactions to the same spatial locations as in the original video, our approach can produce highquality results using only coarse matting of video objects. Coarse matting can be done efficiently using automatic video object segmentation, avoiding tedious manual matting. To design the output, the user interactively indicates the desired new life spans of objects, and may also change the overall running time of the video. Our method rearranges the timelines of objects in the video whilst applying appropriate object interaction constraints. We demonstrate that, while this editing technique is somewhat restrictive, it still allows many interesting results. Shao-Ping Lu, Song-Hai Zhang, Shi-Min Hu 0001, Ralph R. Martin |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2011 | Painting patches: Reducing flicker in painterly re-rendering of video
Song-Hai Zhang, Qiang Tong 0001, Shi-Min Hu 0001, Ralph R. Martin |
Sci. China Inf. Sci. | 1 |
| 2011 | Saliency-Based Fidelity Adaptation Preprocessing for Video Coding
Shao-Ping Lu, Song-Hai Zhang |
J. Comput. Sci. Technol. | 2 |
| 2011 | Online Video Stream Abstraction and StylizationabstractThis paper gives an automatic method for online video stream abstraction, producing a temporally coherent output video stream, in a style with large regions of constant color and highlighted bold edges. Our system includes two novel components. Firstly, to provide coherent and simplified output, we segment frames, and use optical flow to propagate segmentation information from frame to frame; an error control strategy is used to help ensure that the propagated information is reliable. Secondly, to achieve coherent and attractive coloring of the output, we use a color scheme replacement algorithm specifically designed for an online video stream. We demonstrate real-time performance for CIF videos, allowing our approach to be used for live communication and other related applications. Song-Hai Zhang, Xian-Ying Li, Shi-Min Hu 0001, Ralph R. Martin |
IEEE Trans. Multim. | 1 |
| 2009 | Video-based running water animation in Chinese painting style
Song-Hai Zhang, Tao Chen 0015, Yi-Fei Zhang, Shi-Min Hu 0001, Ralph R. Martin |
Sci. China Ser. F Inf. Sci. | 1 |
| 2009 | Vectorizing Cartoon AnimationsabstractWe present a system for vectorizing 2D raster format cartoon animations. The output animations are visually flicker free, smaller in file size, and easy to edit. We identify decorative lines separately from colored regions. We use an accurate and semantically meaningful image decomposition algorithm, supporting an arbitrary color model for each region. To ensure temporal coherence in the output, we reconstruct a universal background for all frames and separately extract foreground regions. Simple user-assistance is required to complete the background. Each region and decorative line is vectorized and stored together with their motions from frame to frame. The contributions of this paper are: 1) the new trapped-ball segmentation method, which is fast, supports nonuniformly colored regions, and allows robust region segmentation even in the presence of imperfectly linked region edges, 2) the separate handling of decorative lines as special objects during image decomposition, avoiding results containing multiple short, thin oversegmented regions, and 3) extraction of a single patch-based background for all frames, which provides a basis for consistent, flicker-free animations. Song-Hai Zhang, Tao Chen 0015, Yi-Fei Zhang, Shi-Min Hu 0001, Ralph R. Martin |
IEEE Trans. Vis. Comput. Graph. | 1 |
| 2002 | An extension algorithm for B-splines by curve unclamping
Shi-Min Hu 0001, Chiew-Lan Tai, Song-Hai Zhang |
Comput. Aided Des. | 3 |