EDBT 2026 Demo / reviewers in the wild / expert
Miao Wang 0004
dblp:80/5294-4
· DBLP profile ↗
52ranked-venue papers
18as first author
35since 2021 · last 2026
0000-0002-4102-8582ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 47 · 17 first-author · 30 since 2021Human-computer interaction and ubiquitous computing · 10 · 2 first-author · 9 since 2021Artificial intelligence and machine learning · 6 · 2 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Negotiating without turning: Exploring rear-space interaction for negotiated teleportation in VR
Hao-Zhong Yang, Wentong Shu, Yijun Li 0006, Miao Wang 0004 |
Comput. Graph. | 4 |
| 2026 | Artificial intelligence for virtual reality: a review
Lili Wang 0006, Yebin Liu, Miao Wang 0004, Xubo Yang, Lan Xu 0003, Zhangyao Tan, Runze Fan, Hongwen Zhang 0001, Yijian Wen, Haozhong Yang, Jian Wu 0033, Jiahui Fan, Hui Wang 0045, Qixuan Zhang, Yongtian Wang, Qinping Zhao |
Sci. China Inf. Sci. | 4 |
| 2026 | Dynamic scene representation in the era of neural rendering: from NeRFs to 3DGSs
Chengye Su, Fan-Yi Zeng, Miao Wang 0004 |
Frontiers Comput. Sci. | 5 |
| 2026 | Language Embedded 3D Gaussians for Open-Vocabulary Scene QueryingabstractOpen-vocabulary querying in 3D space is challenging but essential for scene understanding tasks such as object localization and segmentation. Language embedded scene representations have made progress by incorporating language features into 3D spaces. However, their efficacy heavily depends on neural networks that are resource-intensive in training and rendering. Although recent 3D Gaussians offer efficient and high-quality novel view synthesis, directly embedding language features in them leads to prohibitive memory usage and decreased performance. In this work, we introduce Language Embedded 3D Gaussians, a novel scene representation for open-vocabulary query tasks. Instead of embedding high-dimensional raw semantic features on 3D Gaussians, we propose a dedicated quantization scheme that drastically alleviates the memory requirement, and a novel embedding procedure that achieves smoother yet high accuracy query, countering the multi-view feature inconsistencies and the high-frequency inductive bias in point-based representations. Our comprehensive experiments show that our representation achieves the best visual quality and language querying accuracy across current language embedded representations, while maintaining real-time rendering frame rates on a single desktop GPU. Miao Wang 0004, Jin-Chuan Shi, Shao-Hua Guan, Hao-Bin Duan |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2026 | DGM-RDW: Redirected Walking With Dynamic Geometric Mapping Between EnvironmentsabstractRedirected walking (RDW) subtly adjusts the user's visual perspective on head-mounted displays during natural walking to reduce forced resets, thus enlarging the size of the virtual environment that can be explored beyond that of the physical environment. Alignment-based RDW controllers aim to minimize spatial discrepancies by optimizing the alignment between the user's physical and virtual environments. We introduce a novel alignment-based method that dynamically calculates mapping functions between physical and virtual geometries to enhance the algorithm's awareness of the RDW environments. To achieve this, we first construct an abstract model defining a mapping function between physical and virtual geometries and establish feasibility constraints in differential form. We then concretize this mapping, optimize it, and develop a practical implementation for dynamic geometric mapping in RDW. Our approach distinguishes itself by determining dense spatial mappings around the user, rather than aligning environments according to limited metrics. Through extensive testing, our algorithm has proven to markedly decrease reset incidents in natural walking, surpassing existing RDW controllers. The introduction of dynamic geometric mapping provides a fresh perspective, contributing significant insights and advancing the field. Miao Wang 0004, Yijun Li 0006 |
IEEE Trans. Vis. Comput. Graph. | 1 |
| 2026 | Generating Distance-Aware Human-to-Human Interaction Motions From Text GuidanceabstractThe growing demand for diverse and realistic character animations in video games and films has driven the development of natural language-controlled motion generation systems. While recent advances in text-driven 3D human motion synthesis have made significant progress, generating realistic multi-person interactions remains a major challenge. Existing methods, such as denoising diffusion models and autoregressive frameworks, have explored interaction dynamics using attention mechanisms and causal modeling. However, they consistently overlook a critical physical constraint: the explicit spatial distance between interacting body parts, which is essential for producing semantically accurate and physically plausible interactions. To address this limitation, we propose InterDist, a novel masked generative Transformer model operating in a discrete state space. Our key idea is to decompose two-person motion into three components: two independent, interaction-agnostic single-person motion sequences and a separate interaction distance sequence. This formulation enables direct learning of both individual motion and dynamic spatial relationships from text prompts. We implement this via a VQ-VAE that jointly encodes independent motions and relative distances into discrete codebooks, followed by a bidirectional masked generative Transformer that models their joint distribution conditioned on text. To better align motion and language, we also introduce a cross-modal interaction module to enhance text-motion association. Our approach ensures the generated motions exhibit both semantic alignment with textual descriptions and preserving plausible inter-character distances, setting a new benchmark for text-driven multi-person interaction generation. Jia-Qi Zhang, Miao Wang 0004 |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2025 | GPAvatar: High-fidelity Head Avatars by Learning Efficient Gaussian ProjectionsabstractExisting radiance field-based head avatar methods have mostly relied on pre-computed explicit priors (e.g., mesh, point) or neural implicit representations, making it challenging to achieve high fidelity with both computational efficiency and low memory consumption. To overcome this, we present GPAvatar, a novel and efficient Gaussian splatting-based method for reconstructing high-fidelity dynamic 3D head avatars from monocular videos. We extend Gaussians in 3D space to a high-dimensional embedding space encompassing Gaussian’s spatial position and avatar expression, enabling the representation of the head avatar with arbitrary pose and expression. To enable splatting-based rasterization, a linear transformation is learned to project each high-dimensional Gaussian back to the 3D space, which is sufficient to capture expression variations instead of using complex neural networks. Furthermore, we propose an adaptive densification strategy that dynamically allocates Gaussians to regions with high expression variance, improving the facial detail representation. Experimental results on three datasets show that our method outperforms existing state-of-the-art methods in rendering quality and speed while reducing memory usage in training and rendering. Wei-Qi Feng, Ze-Kang Zhou, Shunkai Li, Pengfei Wan 0001, Di Zhang 0026, Miao Wang 0004 |
CVPR | 8 |
| 2025 | MySpace: Metaphor Design of Personal Space Visualization in Social VR
Wen-Tong Shu, Yijun Li 0006, Miao Wang 0004 |
ICXR | 3 |
| 2025 | Safeteleport: Potential Field-Guided Teleportation for Personal Space Protection in Social VRabstractIn social virtual reality (VR), maintaining appropriate interpersonal distance is essential for user comfort and privacy. However, most existing locomotion methods provide limited support for respecting personal space, leaving users vulnerable to unintentional or socially inappropriate intrusions. To address this issue, we propose potential field-guided teleportation, a proactive locomotion framework consisting of two method implementations that dynamically adjust teleportation targets based on real-time interpersonal proximity, preventing entry into others' personal spaces without explicit user intervention. We evaluate our technique through two user studies: a preliminary study exploring energy-based constraint parameters, followed by a comparative study against conventional and negotiated teleportation methods. Experiments were conducted in socially interactive VR scenarios populated with simulated users exhibiting human-like behaviors. Results demonstrate that our methods reduce perceived social anxiety while maintaining locomotion efficiency and usability. This work presents a socially-aware locomotion strategy that balances personal space protection with effective and socially appropriate movement in shared virtual environments. Yijun Li 0006, Sen-Zhe Xu 0001, Wentong Shu, Hao-Zhong Yang, Zinan Han, Miao Wang 0004, Song-Hai Zhang |
ISMAR | 6 |
| 2025 | Exploring the Influence of Crowd Size Across Different Tasks on User Performance, Experience and Social Presence in Shared Virtual EnvironmentsabstractShared virtual environments are becoming essential platforms for collaborative interaction and immersive entertainment, enabling users to be co-located and engage in activities together. The presence of surrounding virtual humans forms an environmental crowd, serving as a component of ambient stimuli in these environments. However, it remains unclear how crowd size affects users under different cognitive and motor demands. This study investigates the influence of crowd size on user performance, experience and social presence across three fundamental VR tasks: Spatial Locomotion, Memory Search, and Motor Coordination. We conducted a controlled within-subjects experiment, manipulating each task's crowd size at Small, Medium, and Large levels. Our results show that crowd size significantly impacts user performance, experience, and social presence, but these effects are task-dependent. While Medium size can enhance performance, Large size in cognitively demanding tasks may induce attentional blindness and diminish sensitivity to social cues. Task functionality further shapes how users perceive and respond to virtual crowds. Additionally, users' preferences for crowd size varied across different tasks, and most participants expressed a desire for control over the number of visible avatars. These findings provide novel insights into human crowd perception mechanisms, revealing cross-task perceptual variations that pave the way for further exploring crowd perception in shared virtual environments. Hao-Zhong Yang, Yijun Li 0006, Zi-Nan Han, Wen-Tong Shu, Miao Wang 0004 |
ISMAR | 5 |
| 2025 | Semantics-Aware Avatar Locomotion Adaption for Indoor Cross-Scene AR TelepresenceabstractGeographically dispersed users often rely on virtual avatars as intermediaries to facilitate interactive communication and collaboration. However, existing methods for augmented reality (AR) telepresence applications exhibit limitations, including restricted movement within confined sub-areas, lack of smooth transitions, and the necessity for manually establishing object mapping between dissimilar environments. We present a novel interactive AR framework for virtual avatar locomotion adaption while preserving semantic coherence across dissimilar indoor scenes. Initially, we conduct a preliminary user study to identify key attributes influencing preferred avatar movement. These attributes are quantified as features, and a dataset of user annotations on avatar movements is created. Based on the user interaction and scene configurations, we employ a deep reinforcement learning neural network to guide the avatar to the ideal position while maximizing semantic coherence. We validate our proposed framework through simulations and user studies by implementing an AR-based 3D telepresence prototype, demonstrating the efficacy of our framework in conveying user intentions across dissimilar environments, enabling natural and immersive 3D telepresence interactions. Yijun Li 0006, Hao-Zhong Yang, Wentong Shu, Miao Wang 0004 |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2025 | Negotiated User-to-Group Teleportations in Social VRabstractThe locomotion and interaction of multi-user groups are critical components of social virtual reality (VR), where users collaboratively navigate shared spaces defined by their relationships. As metaverse and social VR platforms evolve, safeguarding group spatial integrity becomes paramount. While teleportation-widely adopted for efficient navigation-enhances communication, it risks unintended intrusion into group spaces, compromising privacy and security. This paper presents two novel negotiated user-to-group teleportation techniques, paired with dynamic zone computation methods, to address these challenges. Our approach enables guest users to join groups through spatially aware teleportation, mediated by real-time negotiation. The negotiation interaction process is designed to facilitate users to negotiate teleportation locations efficiently and smoothly. To validate our techniques, we conducted a user study with 36 participants in a VR art museum environment, where they performed collaborative social-tour tasks. The findings demonstrate that our techniques significantly enhance group privacy protection, effectively support user-to-group negotiation of teleportation joining requirements, and alleviate anxiety associated with unwanted proximity within social VR groups. Wentong Shu, Yijun Li 0006, Hao-Zhong Yang, Zi-Nan Han, Frank Steinicke, Miao Wang 0004 |
IEEE Trans. Vis. Comput. Graph. | 6 |
| 2025 | Can I Get There? Negotiated User-to-User Teleportations in Social VRabstractThe growing adoption of social virtual reality (VR) platforms underscores the importance of safeguarding personal VR space to maintain user privacy and security. Teleportation, a prevalent instantaneous locomotion method in VR, facilitates user engagement but can also inadvertently intrude upon personal VR space, thereby raising privacy concerns. This paper introduces three innovative negotiated teleportation techniques designed to secure user-to-user teleportation and protect personal space privacy, all under a unified small-group development framework. We have designed and evaluated three types of negotiated teleportation techniques: Sector technique for directional control, Distance technique for minimum social distance control, and Area technique for defining circular permissible teleportation areas. These techniques foster a collaborative approach to selecting teleportation points that respect personal space. To evaluate the efficacy of these techniques, we conducted a user study with 20 participants who performed social tasks within a virtual campus environment. The findings demonstrate that our techniques significantly enhance privacy protection and alleviate anxiety associated with unwanted proximity in social VR. Miao Wang 0004, Wentong Shu, Yijun Li 0006, Wanwan Li |
IEEE Trans. Vis. Comput. Graph. | 1 |
| 2025 | Skinned Motion Retargeting With Preservation of Body Part RelationshipsabstractMotion retargeting is an active research area in computer graphics and animation, allowing for the transfer of motion from one character to another, thereby creating diverse animated character data. While this technology has numerous applications in animation, games, and movies, current methods often produce unnatural or semantically inconsistent motion when applied to characters with different shapes or joint counts. This is primarily due to a lack of consideration for the geometric and spatial relationships between the body parts of the source and target characters. To tackle this challenge, we introduce a novel spatially-preserving Skinned Motion Retargeting Network (SMRNet) capable of handling motion retargeting for characters with varying shapes and skeletal structures while maintaining semantic consistency. By learning a hybrid representation of the character's skeleton and shape in a rest pose, SMRNet transfers the rotation and root joint position of the source character's motion to the target character through embedded rest pose feature alignment. Additionally, it incorporates a differentiable loss function to further preserve the spatial consistency of body parts between the source and target. Comprehensive quantitative and qualitative evaluations demonstrate the superiority of our approach over existing alternatives, particularly in preserving spatial relationships more effectively. Jia-Qi Zhang, Miao Wang 0004, Fu-Cheng Zhang |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2024 | Neural 3D Strokes: Creating Stylized 3D Scenes with Vectorized 3D StrokesabstractWe present Neural 3D Strokes, a novel technique to gen-erate stylized images of a 3D scene at arbitrary novel views from multi-view 2D images. Different from existing methods which apply stylization to trained neural radiance fields at the voxel level, our approach draws inspiration from image-to-painting methods, simulating the progressive painting process of human artwork with vector strokes. We develop a palette of stylized 3D strokes from basic primitives and splines, and consider the 3D scene stylization task as a multi-view reconstruction process based on these 3D stroke primitives. Instead of directly searching for the parame-ters of these 3D strokes, which would be too costly, we introduce a differentiable renderer that allows optimizing stroke parameters using gradient descent, and propose a training scheme to alleviate the vanishing gradient issue. The extensive evaluation demonstrates that our approach effectively synthesizes 3D scenes with significant geomet-ric and aesthetic stylization while maintaining a consis-tent appearance across different views. Our method can be further integrated with style loss and image-text con-trastive models to extend its applications, including color transfer and text-driven 3D scene drawing. Results and code are available at http://buaavrcg.github.io/Neura13DStrokes. Hao-Bin Duan, Miao Wang 0004, Yan-Xun Li |
CVPR | 2 |
| 2024 | Exploring Regional Clues in CLIP for Zero-Shot Semantic SegmentationabstractCLIP has demonstrated marked progress in visual recognition due to its powerful pre-training on large-scale image-text pairs. However, it still remains a critical challenge: how to transfer image-level knowledge into pixel-level understanding tasks such as semantic segmentation. In this paper, to solve the mentioned challenge, we analyze the gap between the capability of the CLIP model and the requirement of the zero-shot semantic segmentation task. Based on our analysis and observations, we propose a novel method for zero-shot semantic segmentation, dubbed CLIP-RC (CLIP with Regional Clues), bringing two main insights. On the one hand, a region-level bridge is necessary to provide fine-grained semantics. On the other hand, over-fitting should be mitigated during the training stage. Benefiting from the above discoveries, CLIP-RC achieves state-of-the-art performance on various zero-shot semantic segmentation benchmarks, including PASCAL VOC, PASCAL Context, and COCO-Stuff 164K. Code will be available at https://github.com/Jittor/JSeg. Yi Zhang 0099, Menghao Guo 0001, Miao Wang 0004, Shi-Min Hu 0001 |
CVPR | 3 |
| 2024 | Transferable Virtual-Physical Environmental Alignment With Redirected WalkingabstractSeveral advanced redirected walking techniques have been proposed in recent years to improve natural walking in virtual environments. One active and important research challenge of redirected walking focuses on the alignment of virtual and physical environments by redirection gains. If both environments are aligned, physical objects appear at the same positions as their virtual counterparts. When a user arrives at such a virtual object, she can touch the corresponding physical object providing passive haptic feedback. When multiple transferable virtual or physical target positions exist, the alignment can exploit multiple options, but the process requires more complicated solutions. In this paper, we study the problem of virtual-physical environmental alignment at multiple transferable target positions, and introduce a novel reinforcement learning-based redirected walking method. We design a novel comprehensive reward function that dynamically determines virtual-physical target matching and updates virtual target weights for reward computation. We evaluate our method through various simulated experiments as well as real user tests. The results show that our method obtains less physical distance error for environmental alignment and requires fewer resets than state-of-the-art techniques. Miao Wang 0004, Ze-Yin Chen, Wen-Chuan Cai, Frank Steinicke |
IEEE Trans. Vis. Comput. Graph. | 1 |
| 2024 | SceneFusion: Room-Scale Environmental Fusion for Efficient Traveling Between Separate Virtual EnvironmentsabstractTraveling between scenes has become a major requirement for navigation in numerous virtual reality (VR) social platforms and game applications, allowing users to efficiently explore multiple virtual environments (VEs). To facilitate scene transition, prevalent techniques such as instant teleportation and virtual portals have been extensively adopted. However, these techniques exhibit limitations when there is a need for frequent travel between separate VEs, particularly within indoor environments, resulting in low efficiency. In this article, we first analyze the design rationale for a novel navigation method supporting efficient travel between virtual indoor scenes. Based on the analysis, we introduce the SceneFusion technique that fuses separate virtual rooms into an integrated environment. SceneFusion enables users to perceive rich visual information from both rooms simultaneously, achieving high visual continuity and spatial awareness. While existing teleportation techniques passively transport users, SceneFusion allows users to actively access the fused environment using short-range locomotion techniques. User experiments confirmed that SceneFusion outperforms instant teleportation and virtual portal techniques in terms of efficiency, workload, and preference for both single-user exploration and multi-user collaboration tasks in separate VEs. Thus, SceneFusion presents an effective solution for seamless traveling between virtual indoor scenes. Miao Wang 0004, Yijun Li 0006, Jin-Chuan Shi, Frank Steinicke |
IEEE Trans. Vis. Comput. Graph. | 1 |
| 2024 | Studying the Effect of Material and Geometry on Perceptual Outdoor IlluminationabstractUnderstanding and modeling perceived properties of sky-dome illumination is an important but challenging problem due to the interplay of several factors such as the materials and geometries of the objects present in the scene being observed. Existing models of sky-dome illumination focus on the physical properties of the sky. However, these parametric models often do not align well with the properties perceived by a human observer. In this work, drawing inspiration from the Hosek-Wilkie sky-dome model, we investigate the perceptual properties of outdoor illumination. For this purpose, we perform a large-scale user study via crowdsourcing to collect a dataset of perceived illumination properties (scattering, glare, and brightness) for different combinations of geometries and materials under a variety of outdoor illuminations, totaling 5,000 distinct images. We perform a thorough statistical analysis of the collected data which reveals several interesting effects. For instance, our analysis shows that when there are objects in the scene made of rough materials, the perceived scattering of the sky increases. Furthermore, we utilize our extensive collection of images and their corresponding perceptual attributes to train a predictor. This predictor, when provided with a single image as input, generates an estimation of perceived illumination properties that align with human perceptual judgments. Accurately estimating perceived illumination properties can greatly enhance the overall quality of integrating virtual objects into real scene photographs. Consequently, we showcase various applications of our predictor. For instance, we demonstrate its utility as a luminance editing tool for showcasing virtual objects in outdoor scenes. Miao Wang 0004, Jinchao Zhou, Wei-Qi Feng, Yu-Zhu Jiang, Ana Serrano |
IEEE Trans. Vis. Comput. Graph. | 1 |
| 2024 | Multi-User Redirected Walking in Separate Physical Spaces for Online VR ScenariosabstractWith the recent rise of Metaverse, online multiplayer VR applications are becoming increasingly prevalent worldwide. However, as multiple users are located in different physical environments, different reset frequencies and timings can lead to serious fairness issues for online collaborative/competitive VR applications. For the fairness of online VR apps/games, an ideal online RDW strategy must make the locomotion opportunities of different users equal, regardless of different physical environment layouts. The existing RDW methods lack the scheme to coordinate multiple users in different PEs, and thus have the issue of triggering too many resets for all the users under the locomotion fairness constraint. We propose a novel multi-user RDW method that is able to significantly reduce the overall reset number and give users a better immersive experience by providing a fair exploration. Our key idea is to first find out the "bottleneck" user that may cause all users to be reset and estimate the time to reset given the users' next targets, and then redirect all the users to favorable poses during that maximized bottleneck time to ensure the subsequent resets can be postponed as much as possible. More particularly, we develop methods to estimate the time of possibly encountering obstacles and the reachable area for a specific pose to enable the prediction of the next reset caused by any user. Our experiments and user study found that our method outperforms existing RDW methods in online VR applications. Sen-Zhe Xu 0001, Jia-Hong Liu, Miao Wang 0004, Song-Hai Zhang |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2023 | HoughLaneNet: Lane detection with deep hough transform and dynamic convolution
Jia-Qi Zhang, Hao-Bin Duan, Ariel Shamir, Miao Wang 0004 |
Comput. Graph. | 5 |
| 2023 | BakedAvatar: Baking Neural Fields for Real-Time Head Avatar SynthesisabstractSynthesizing photorealistic 4D human head avatars from videos is essential for VR/AR, telepresence, and video game applications. Although existing Neural Radiance Fields (NeRF)-based methods achieve high-fidelity results, the computational expense limits their use in real-time applications. To overcome this limitation, we introduce BakedAvatar , a novel representation for real-time neural head avatar synthesis, deployable in a standard polygon rasterization pipeline. Our approach extracts deformable multi-layer meshes from learned isosurfaces of the head and computes expression-, pose-, and view-dependent appearances that can be baked into static textures for efficient rasterization. We thus propose a three-stage pipeline for neural head avatar synthesis, which includes learning continuous deformation, manifold, and radiance fields, extracting layered meshes and textures, and fine-tuning texture details with differential rasterization. Experimental results demonstrate that our representation generates synthesis results of comparable quality to other state-of-the-art methods while significantly reducing the inference time required. We further showcase various head avatar synthesis results from monocular videos, including view synthesis, face reenactment, expression editing, and pose editing, all at interactive frame rates on commodity devices. Source codes and demos are available on our project page. Hao-Bin Duan, Miao Wang 0004, Jin-Chuan Shi, Xu-Chuan Chen, Yan-Pei Cao 0001 |
ACM Trans. Graph. | 2 |
| 2022 | Rendering-Aware HDR Environment Map Prediction from a Single ImageabstractHigh dynamic range (HDR) illumination estimation from a single low dynamic range (LDR) image is a significant task in computer vision, graphics, and augmented reality. We present a two-stage deep learning-based method to predict an HDR environment map from a single narrow field-of-view LDR image. We first learn a hybrid parametric representation that sufficiently covers high- and low-frequency illumination components in the environment. Taking the estimated illuminations as guidance, we build a generative adversarial network to synthesize an HDR environment map that enables realistic rendering effects. We specifically consider the rendering effect by supervising the networks using rendering losses in both stages, on the predicted environment map as well as the hybrid illumination representation. Quantitative and qualitative experiments demonstrate that our approach achieves lower relighting errors for virtual object insertion and is preferred by users compared to state-of-the-art methods. Jun-Peng Xu, Chenyu Zuo, Miao Wang 0004 |
AAAI | 4 |
| 2022 | Bullet Comments for 360°VideoabstractTime-anchored on-screen comments, as known as bullet comments, are a popular feature for online video streaming. Bullet comments reflect audiences’ feelings and opinions at specific video timings, which have been shown to be beneficial to video content understanding and social connection level. In this paper, we for the first time investigate the problem of bullet comment display and insertion for 360° video via head-mounted display and controller. We design four bullet comment display methods and evaluate their effects on 360° video experiences. We further propose two controller-based methods for bullet comment insertion. Combining the display and insertion methods, the user can experience 360° videos with bullet comments, and interactively post new ones by selecting among existing comments. User study results revealed how the factors of display and insertion methods affect 360° video experience. With the experiment findings, we also discuss useful design insights for 360° video bullet comments. Yijun Li 0006, Jin-Chuan Shi, Miao Wang 0004 |
VR | 4 |
| 2022 | A Comprehensive Review of Redirected Walking Techniques: Taxonomy, Methods, and Future Directions
Yijun Li 0006, Frank Steinicke, Miao Wang 0004 |
J. Comput. Sci. Technol. | 3 |
| 2022 | Efficient Flower Text Entry in Virtual RealityabstractText entry is a frequently used task in virtual reality (VR) applications, and controller is the most common interactive device in current VR systems. However, in terms of typing speed, there is still a gap between the existing controller-based text entry techniques and using a physical keyboard in reality, so it is important to improve the efficiency of the controller-based text entry. In this paper, we introduce Flower Text Entry, a single-controller text entry method based on a newly designed flower-shaped keyboard using hand 3D translation interaction for letters selection. We conduct user studies to optimize the keyboard design and the mapping between the interaction and selection, so as to evaluate our method. The results show that our method has high typing speed, lower error rate, and is very friendly to novices compared with the state-of-the-art controller-based text entry methods. After a short training, the novice group can type at 17.65 words per minute (WPM), and the potential expert group can type at 22.97 WPM. The highest typing speed is up to 30.80 WPM achieved by a potential expert participant. Jiaye Leng, Lili Wang 0006, Xuehuai Shi, Miao Wang 0004 |
IEEE Trans. Vis. Comput. Graph. | 5 |
| 2021 | A Reinforcement Learning Approach to Redirected Walking with Passive Haptic FeedbackabstractVarious redirected walking (RDW) techniques have been proposed, which unwittingly manipulate the mapping from the user’s physical locomotion to motions of the virtual camera. Thereby, RDW techniques guide users on physical paths with the goal to keep them inside a limited tracking area, whereas users perceive the illusion of being able to walk infinitely in the virtual environment. However, the inconsistency between the user’s virtual and physical location hinders passive haptic feedback when the user interacts with virtual objects, which are represented by physical props in the real environment.In this paper, we present a novel reinforcement learning approach towards RDW with passive haptics. With a novel dense reward function, our method learns to jointly consider physical boundary avoidance and consistency of user-object positioning between virtual and physical spaces. The weights of reward and penalty terms in the reward function are dynamically adjusted to adaptively balance term impacts during the walking process. Experimental results demonstrate the advantages of our technique in comparison to previous approaches. Finally, the code of our technique is provided as an open-source solution. Ze-Yin Chen, Yijun Li 0006, Miao Wang 0004, Frank Steinicke, Qinping Zhao |
ISMAR | 3 |
| 2021 | OpenRDW: A Redirected Walking Library and Benchmark with Multi-User, Learning-based Functionalities and State-of-the-art AlgorithmsabstractRedirected walking (RDW) is a locomotion technique that guides users on virtual paths, which might vary from the paths they physically walk in the real world. Thereby, RDW enables users to explore a virtual space that is larger than the physical counterpart with near-natural walking experiences. Several approaches have been proposed and developed; each using individual platforms and evaluated on a custom dataset, making it challenging to compare between methods. However, there are seldom public toolkits and recognized benchmarks in this field. In this paper, we introduce OpenRDW, an open-source library and benchmark for developing, deploying and evaluating a variety of methods for walking path redirection. The OpenRDW library provides application program interfaces to access the attributes of scenes, to customize the RDW controllers, to simulate and visualize the navigation process, to export multiple formats of the results, and to evaluate RDW techniques. It also supports the deployment of multi-user real walking, as well as reinforcement learning-based models exported from TensorFlow or PyTorch. The OpenRDW benchmark includes multiple testing conditions, such as walking in size varied tracking spaces or shape varied tracking spaces with obstacles, multiple user walking, etc. On the other hand, procedurally generated paths and walking paths collected from user experiments are provided for a comprehensive evaluation. It also contains several classic and state-of-the-art RDW techniques, which include the above mentioned functionalities. Yijun Li 0006, Miao Wang 0004, Frank Steinicke, Qinping Zhao |
ISMAR | 2 |
| 2021 | PAVAL: Position-Aware Virtual Agent Locomotion for Assisted Virtual Reality NavigationabstractVirtual agents are typical assistance tools for navigation and interaction in Virtual Reality (VR) tour, training, education, etc. It has been demonstrated that the gaits, gestures, gazes, and positions of virtual agents are major factors that affect the user’s perception and experience for seated and standing VR. In this paper, we present a novel position-aware virtual agent locomotion method, called PAVAL, that can perform virtual agent positioning (position+orientation) in real time for room-scale VR navigation assistance. We first analyze design guidelines for virtual agent locomotion and model the problem using the positions of the user and the surrounding virtual objects. Then we conduct a one-off preliminary study to collect subjective data and present a model for virtual agent positioning prediction with fixed user position. Based on the model, we propose an algorithm to optimize the object of interest, virtual agent position, and virtual agent orientation in sequence for virtual agent locomotion. As a result, during user navigation in a virtual scene, the virtual agent automatically moves in real time and introduces virtual object information to the user. We evaluate PAVAL and two alternative methods via a user study with humanoid virtual agents in various scenes, including virtual museum, factory, and school gym. The results reveal that our method is superior to the baseline condition. Zi-Ming Ye, Miao Wang 0004 |
ISMAR | 3 |
| 2021 | Detection Thresholds with Joint Horizontal and Vertical Gains in Redirected JumpingabstractRedirected jumping (RDJ) is a locomotion technique that allows users to explore a virtual space that is larger than the available physical space by imperceptibly manipulating users' virtual viewpoints according to different gains. In previous redirected jumping work, different types of gains were imposed separately, without considering the possible interaction effects of horizontal and vertical gains on the jumping distance perception. To figure out how humans perceive distance manipulation when more than one gain is used, in this paper, we explored joint horizontal and vertical gains that manipulate horizontal and vertical distances at the same time during two-legged takeoff jumping in the virtual space. We estimated and analyzed horizontal and vertical detection thresholds by conducting a user study, fitting the data to two-dimensional psychometric functions, and visualizing the fitted 3D plots. We provided quantitative insights into the effects of joint gains on detection thresholds, where the imperceptible range for one gain can be affected by the variation of the other gain. Finally, we designed redirected jumping-based games as applications with joint horizontal and vertical gains and demonstrated the effectiveness of the redirected jumping technique. Yijun Li 0006, De-Rong Jin, Miao Wang 0004, Frank Steinicke, Shi-Min Hu 0001, Qinping Zhao |
VR | 3 |
| 2021 | Scene-Context-Aware Indoor Object Selection and Movement in VRabstractVirtual reality (VR) applications such as interior design typically require accurate and efficient selection and movement of indoor objects. In this paper, we present an indoor object selection and movement approach by taking into account scene contexts such as object semantics and interrelations. This provides more intelligence and guidance to the interaction, and greatly enhances user experience. We evaluate our proposals by comparing them with traditional approaches in different interaction modes based on controller, head pose, and eye gaze. Extensive user studies on a variety of selection and movement tasks are conducted to validate the advantages of our approach. We demonstrate our findings via a furniture arrangement application. Miao Wang 0004, Zi-Ming Ye, Jin-Chuan Shi, Yang-Liang Yang |
VR | 1 |
| 2021 | Write-An-Animation: High-level Text-based Animation Editing with Character-Scene InteractionabstractAbstract 3D animation production for storytelling requires essential manual processes of virtual scene composition, character creation, and motion editing, etc. Although professional artists can favorably create 3D animations using software, it remains a complex and challenging task for novice users to handle and learn such tools for content creation. In this paper, we present Write‐An‐Animation, a 3D animation system that allows novice users to create, edit, preview, and render animations, all through text editing. Based on the input texts describing virtual scenes and human motions in natural languages, our system first parses the texts as semantic scene graphs, then retrieves 3D object models for virtual scene composition and motion clips for character animation. Character motion is synthesized with the combination of generative locomotions using neural state machine as well as template action motions retrieved from the dataset. Moreover, to make the virtual scene layout compatible with character motion, we propose an iterative scene layout and character motion optimization algorithm that jointly considers character‐object collision and interaction. We demonstrate the effectiveness of our system with customized texts and public film scripts. Experimental results indicate that our system can generate satisfactory animations from texts. Jia-Qi Zhang, Zhi-Meng Shen, Zehuan Huang, Yan-Pei Cao 0001, Pengfei Wan 0001, Miao Wang 0004 |
Comput. Graph. Forum | 8 |
| 2021 | Prominent Structures for Video Analysis and EditingabstractWe present prominent structures in video, a representation of visually strong, spatially sparse and temporally stable structural units, for use in video analysis and editing. With a novel quality measurement of prominent structures in video, we develop a general framework for prominent structure computation, and an efficient hierarchical structure alignment algorithm between a pair of videos. The prominent structural unit map is proposed to encode both binary prominence guidance and numerical strength and geometry details for each video frame. Even though the detailed appearance of videos could be visually different, the proposed alignment algorithm can find matched prominent structure sub-volumes. Prominent structures in video support a wide range of video analysis and editing applications including graphic match-cut between successive videos, instant cut editing, finding transition portals from a video collection, structure-aware video re-ranking, visualizing human action differences, etc. Miao Wang 0004, Xiaonan Fang 0001, Ariel Shamir, Shi-Min Hu 0001 |
IEEE Trans. Vis. Comput. Graph. | 1 |
| 2021 | Effects of virtual environment and self-representations on perception and physical performance in redirected jumpingabstractRedirected jumping (RDJ) allows users to explore virtual environments (VEs) naturally by scaling a small real-world jump to a larger virtual jump with virtual camera motion manipulation, thereby addressing the problem of limited physical space in VR applications. Previous RDJ studies have mainly focused on detection threshold estimation. However, the effect VE or selfrepresentation (SR) has on the perception or performance of RDJs remains unclear. In this paper, we report experiments to measure the perception (detection thresholds for gains, presence, embodiment, intrinsic motivation, and cybersickness) and physical performance (heart rate intensity, preparation time, and actual jumping distance) of redirected forward jumping under six different combinations of VE (low and high visual richness) and SRs (invisible, shoes, and human-like). Our results indicated that the detection threshold ranges for horizontal translation gains were significantly smaller in the VE with high rather than low visual richness. When different SRs were applied, our results did not suggest significant differences in detection thresholds, but it did report longer actual jumping distances in the invisible body case compared with the other two SRs. In the high visual richness VE, the preparation time for jumping with a human-like avatar was significantly longer than that with other SRs. Finally, some correlations were found between perception and physical performance measures. All these findings suggest that both VE and SRs influence users' perception and performance in RDJ and must be considered when designing locomotion techniques. Yijun Li 0006, Miao Wang 0004, De-Rong Jin, Frank Steinicke, Shi-Min Hu 0001, Qinping Zhao |
Virtual Real. Intell. Hardw. | 2 |
| 2021 | Locomotion perception and redirection
Miao Wang 0004, Songhai Zhang, Shi-Min Hu 0001 |
Virtual Real. Intell. Hardw. | 1 |
| 2020 | Transitioning360: Content-aware NFoV Virtual Camera Paths for 360° Video PlaybackabstractDespite the increasing number of head-mounted displays, many 360° VR videos are still being viewed by users on existing 2D displays. To this end, a subset of the 360° video content is often shown inside a manually or semi-automatically selected normal-field-of-view (NFoV) window. However, during the playback, simply watching an NFoV video can easily miss concurrent off-screen content. We present Transitioning360, a tool for 360° video navigation and playback on 2D displays by transitioning between multiple NFoV views that track potentially interesting targets or events. Our method computes virtual NFoV camera paths considering content awareness and diversity in an offline preprocess. During playback, the user can watch any NFoV view corresponding to a precomputed camera path. Moreover, our interface shows other candidate views, providing a sense of concurrent events. At any time, the user can transition to other candidate views for fast navigation and exploration. Experimental results including a user study demonstrate that the viewing experience using our method is more enjoyable and convenient than previous methods. Miao Wang 0004, Yijun Li 0006, Christian Richardt, Shi-Min Hu 0001 |
ISMAR | 1 |
| 2020 | 3D computational modeling and perceptual analysis of kinetic depth effectsabstractHumans have the ability to perceive kinetic depth effects , i.e., to perceived 3D shapes from 2D projections of rotating 3D objects. This process is based on a variety of visual cues such as lighting and shading effects. However, when such cues are weak or missing, perception can become faulty, as demonstrated by the famous silhouette illusion example of the spinning dancer . Inspired by this, we establish objective and subjective evaluation models of rotated 3D objects by taking their projected 2D images as input. We investigate five different cues: ambient luminance, shading, rotation speed, perspective, and color difference between the objects and background. In the objective evaluation model, we first apply 3D reconstruction algorithms to obtain an objective reconstruction quality metric, and then use quadratic stepwise regression analysis to determine weights of depth cues to represent the reconstruction quality. In the subjective evaluation model, we use a comprehensive user study to reveal correlations with reaction time and accuracy, rotation speed, and perspective. The two evaluation models are generally consistent, and potentially of benefit to inter-disciplinary research into visual perception and 3D reconstruction. Mengyao Cui 0001, Shao-Ping Lu, Miao Wang 0004, Yongliang Yang 0002, Yukun Lai, Paul L. Rosin |
Comput. Vis. Media | 3 |
| 2020 | VR content creation and exploration with deep learning: A surveyabstractVirtual reality (VR) offers an artificial, computer generated simulation of a real life environment. It originated in the 1960s and has evolved to provide increasing immersion, interactivity, imagination, and intelligence. Because deep learning systems are able to represent and compose information at various levels in a deep hierarchical fashion, they can build very powerful models which leverage large quantities of visual media data. Intelligence of VR methods and applications has been significantly boosted by the recent developments in deep learning techniques. VR content creation and exploration relates to image and video analysis, synthesis and editing, so deep learning methods such as fully convolutional networks and general adversarial networks are widely employed, designed specifically to handle panoramic images and video and virtual 3D scenes. This article surveys recent research that uses such deep learning methods for VR content creation and exploration. It considers the problems involved, and discusses possible future directions in this active and emerging research area. Miao Wang 0004, Xu-Quan Lyu, Yijun Li 0006 |
Comput. Vis. Media | 1 |
| 2020 | Photorealistic Audio-driven Video PortraitsabstractVideo portraits are common in a variety of applications, such as videoconferencing, news broadcasting, and virtual education and training. We present a novel method to synthesize photorealistic video portraits for an input portrait video, automatically driven by a person's voice. The main challenge in this task is the hallucination of plausible, photorealistic facial expressions from input speech audio. To address this challenge, we employ a parametric 3D face model represented by geometry, facial expression, illumination, etc., and learn a mapping from audio features to model parameters. The input source audio is first represented as a high-dimensional feature, which is used to predict facial expression parameters of the 3D face model. We then replace the expression parameters computed from the original target video with the predicted one, and rerender the reenacted face. Finally, we generate a photorealistic video portrait from the reenacted synthetic face sequence via a neural face renderer. One appealing feature of our approach is the generalization capability for various input speech audio, including synthetic speech audio from text-to-speech software. Extensive experimental results show that our approach outperforms previous general-purpose audio-driven video portrait methods. This includes a user study demonstrating that our results are rated as more realistic than previous methods. Miao Wang 0004, Christian Richardt, Ze-Yin Chen, Shi-Min Hu 0001 |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2019 | Example-Guided Style-Consistent Image Synthesis From Semantic LabelingabstractExample-guided image synthesis aims to synthesize an image from a semantic label map and an exemplary image indicating style. We use the term "style" in this problem to refer to implicit characteristics of images, for example: in portraits "style" includes gender, racial identity, age, hairstyle; in full body pictures it includes clothing; in street scenes it refers to weather and time of day and such like. A semantic label map in these cases indicates facial expression, full body pose, or scene segmentation. We propose a solution to the example-guided image synthesis problem using conditional generative adversarial networks with style consistency. Our key contributions are (i) a novel style consistency discriminator to determine whether a pair of images are consistent in style; (ii) an adaptive semantic consistency loss; and (iii) a training data sampling strategy, for synthesizing style-consistent results to the exemplar. We demonstrate the efficiency of our method on face, dance and street view synthesis tasks. Miao Wang 0004, Guo-Ye Yang, Ruilong Li, Runze Liang, Song-Hai Zhang, Peter Hall 0001, Shi-Min Hu 0001 |
CVPR | 1 |
| 2019 | Learning Explicit Smoothing Kernels for Joint Image FilteringabstractAbstract Smoothing noises while preserving strong edges in images is an important problem in image processing. Image smoothing filters can be either explicit (based on local weighted average) or implicit (based on global optimization). Implicit methods are usually time‐consuming and cannot be applied to joint image filtering tasks, i.e., leveraging the structural information of a guidance image to filter a target image. Previous deep learning based image smoothing filters are all implicit and unavailable for joint filtering. In this paper, we propose to learn explicit guidance feature maps as well as offset maps from the guidance image and smoothing parameter that can be utilized to smooth the input itself or to filter images in other target domains. We design a deep convolutional neural network consisting of a fully‐convolution block for guidance and offset maps extraction together with a stacked spatially varying deformable convolution block for joint image filtering. Our models can approximate several representative image smoothing filters with high accuracy comparable to state‐of‐the‐art methods, and serve as general tools for other joint image filtering tasks, such as color interpolation, depth map upsampling, saliency map upsampling, flash/non‐flash image denoising and RGB/NIR image denoising. Xiaonan Fang 0001, Miao Wang 0004, Ariel Shamir, Shi-Min Hu 0001 |
Comput. Graph. Forum | 2 |
| 2019 | ShadowGAN: Shadow synthesis for virtual objects with conditional adversarial networksabstractWe introduce ShadowGAN , a generative adversarial network (GAN) for synthesizing shadows for virtual objects inserted in images. Given a target image containing several existing objects with shadows, and an input source object with a specified insertion position, the network generates a realistic shadow for the source object. The shadow is synthesized by a generator; using the proposed local adversarial and global adversarial discriminators, the synthetic shadow’s appearance is locally realistic in shape, and globally consistent with other objects’ shadows in terms of shadow direction and area. To overcome the lack of training data, we produced training samples based on public 3D models and rendering technology. Experimental results from a user study show that the synthetic shadowed results look natural and authentic. Runze Liang, Miao Wang 0004 |
Comput. Vis. Media | 3 |
| 2019 | Deep Online Video Stabilization With Multi-Grid Warping Transformation LearningabstractVideo stabilization techniques are essential for most hand-held captured videos due to high-frequency shakes. Several 2D, 2.5D and 3D-based stabilization techniques have been presented previously, but to our knowledge, no solutions based on deep neural networks had been proposed to date. The main reason for this omission is shortage in training data as well as the challenge of modeling the problem using neural networks. In this paper, we present a video stabilization technique using a convolutional neural network. Previous works usually propose an offline algorithm that smoothes a holistic camera path based on feature matching. Instead, we focus on low-latency, real-time camera path smoothing, that does not explicitly represent the camera path, and does not use future frames. Our neural network model, called StabNet, learns a set of mesh-grid transformations progressively for each input frame from the previous set of stabalized camera frames, and creates stable corresponding latent camera paths implicitly. To train the network, we collect a dataset of synchronized steady and unsteady video pairs via a specially designed hand-held hardware. Experimental results show that our proposed online method performs comparatively to traditional offline video stabilization methods without using future frames, while running about 10× faster. More importantly, our proposed StabNet is able to handle low-quality videos such as night-scene videos, watermarked videos, blurry videos and noisy videos, where existing methods fail in feature extraction or matching. Miao Wang 0004, Guo-Ye Yang, Jin-Kun Lin, Song-Hai Zhang, Ariel Shamir, Shao-Ping Lu, Shi-Min Hu 0001 |
IEEE Trans. Image Process. | 1 |
| 2019 | Write-a-video: computational video montage from themed textabstractWe present Write-A-Video , a tool for the creation of video montage using mostly text-editing. Given an input themed text and a related video repository either from online websites or personal albums, the tool allows novice users to generate a video montage much more easily than current video editing tools. The resulting video illustrates the given narrative, provides diverse visual content, and follows cinematographic guidelines. The process involves three simple steps: (1) the user provides input, mostly in the form of editing the text, (2) the tool automatically searches for semantically matching candidate shots from the video repository, and (3) an optimization method assembles the video montage. Visual-semantic matching between segmented text and shots is performed by cascaded keyword matching and visual-semantic embedding, that have better accuracy than alternative solutions. The video assembly is formulated as a hybrid optimization problem over a graph of shots, considering temporal constraints, cinematography metrics such as camera movement and tone, and user-specified cinematography idioms. Using our system, users without video editing experience are able to generate appealing videos. Miao Wang 0004, Shi-Min Hu 0001, Shing-Tung Yau, Ariel Shamir |
ACM Trans. Graph. | 1 |
| 2018 | Synthesis of Shaking Video Using Motion Capture Data and Dynamic 3D Scene ModelingabstractImportant video processing methods such as video stabilization and deblurring often do not have ground-truth data available. This poses a great challenge in the development and parameter tunning of such methods. Synthetic shaken video is very useful to generate well-defined ground-truth datasets. Existing shaking video synthesis methods simulate shaky camera motion by performing 2D view warping using only a single 2D video, which does not always correspond to realistic 3D motions. In this paper, we introduce a novel shaking video synthesis approach. The proposed framework constructs the camera motion trajectory by making use of human motion information that is captured in the real-world. Moreover, we render the shaken video from man-made dynamic 3D scenes with detailed camera pose information. Our novel approach provides both accurate 2D visual content and camera motion trajectory in the 3D scene, which allows for evaluating the visual distortion as well as the offsets of the recovered camera trajectory. The proposed synthesis method of shaking video will benefit and ease future research on 3D-aware video stabilization. Shao-Ping Lu, Beerend Ceulemans, Miao Wang 0004, Adrian Munteanu 0001 |
ICIP | 4 |
| 2018 | Deep Video Stabilization Using Adversarial NetworksabstractAbstract Video stabilization is necessary for many hand‐held shot videos. In the past decades, although various video stabilization methods were proposed based on the smoothing of 2D, 2.5D or 3D camera paths, hardly have there been any deep learning methods to solve this problem. Instead of explicitly estimating and smoothing the camera path, we present a novel online deep learning framework to learn the stabilization transformation for each unsteady frame, given historical steady frames. Our network is composed of a generative network with spatial transformer networks embedded in different layers, and generates a stable frame for the incoming unstable frame by computing an appropriate affine transformation. We also introduce an adversarial network to determine the stability of apiece of video. The network is trained directly using the pair of steady and unsteady videos. Experiments show that our method can produce similar results as traditional methods, moreover, it is capable of handling challenging unsteady video of low quality, where traditional methods fail, such as video with heavy noise or multiple exposures. Our method runs in real time, which is much faster than traditional methods. Sen-Zhe Xu 0001, Miao Wang 0004, Tai-Jiang Mu, Shi-Min Hu 0001 |
Comput. Graph. Forum | 3 |
| 2018 | Hyper-Lapse From Multiple Spatially-Overlapping VideosabstractHyper-lapse video with high speed-up rate is an efficient way to overview long videos, such as a human activity in first-person view. Existing hyper-lapse video creation methods produce a fast-forward video effect using only one video source. In this paper, we present a novel hyper-lapse video creation approach based on multiple spatially-overlapping videos. We assume the videos share a common view or location, and find transition points where jumps from one video to another may occur. We represent the collection of videos using a hyper-lapse transition graph; the edges between nodes represent possible hyper-lapse frame transitions. To create a hyper-lapse video, a shortest path search is performed on this digraph to optimize frame sampling and assembly simultaneously. Finally, we render the hyper-lapse results using video stabilization and appearance smoothing techniques on the selected frames. Our technique can synthesize novel virtual hyper-lapse routes, which may not exist originally. We show various application results on both indoor and outdoor video collections with static scenes, moving objects, and crowds. Miao Wang 0004, Jun-Bang Liang, Song-Hai Zhang, Shao-Ping Lu, Ariel Shamir, Shi-Min Hu 0001 |
IEEE Trans. Image Process. | 1 |
| 2018 | BiggerSelfie: Selfie Video Expansion With Hand-Held CameraabstractSelfie photography from the hand-held camera is becoming a popular media type. Although being convenient and flexible, it suffers from low camera motion stability, small field of view, and limited background content. These limitations can annoy users, especially, when touring a place of interest and taking selfie videos. In this paper, we present a novel method to create what we call a BiggerSelfie that deals with these shortcomings. Using a video of the environment that has partial content overlap with the selfie video, we stitch plausible frames selected from the environment video to the original selfie frames and stabilize the composed video content with a portrait-preserving constraint. Using the proposed method, one can easily obtain a stable selfie video with expanded background content by merely capturing some background shots. We show various results and several evaluations to demonstrate the applicability of our method. Miao Wang 0004, Ariel Shamir, Guo-Ye Yang, Jin-Kun Lin, Shao-Ping Lu, Shi-Min Hu 0001 |
IEEE Trans. Image Process. | 1 |
| 2016 | Comfort-driven disparity adjustment for stereoscopic videoabstractPixel disparity—the offset of corresponding pixels between left and right views—is a crucial parameter in stereoscopic three-dimensional (S3D) video, as it determines the depth perceived by the human visual system (HVS). Unsuitable pixel disparity distribution throughout an S3D video may lead to visual discomfort. We present a unified and extensible stereoscopic video disparity adjustment framework which improves the viewing experience for an S3D video by keeping the perceived 3D appearance as unchanged as possible while minimizing discomfort. We first analyse disparity and motion attributes of S3D video in general, then derive a wide-ranging visual discomfort metric from existing perceptual comfort models. An objective function based on this metric is used as the basis of a hierarchical optimisation method to find a disparity mapping function for each input video frame. Warping-based disparity manipulation is then applied to the input video to generate the output video, using the desired disparity mappings as constraints. Our comfort metric takes into account disparity range, motion , and stereoscopic window violation ; the framework could easily be extended to use further visual comfort models. We demonstrate the power of our approach using both animated cartoons and real S3D videos. Miao Wang 0004, Xi-Jin Zhang, Jun-Bang Liang, Song-Hai Zhang, Ralph R. Martin |
Comput. Vis. Media | 1 |
| 2014 | BiggerPicture: data-driven image extrapolation using graph matchingabstractFilling a small hole in an image with plausible content is well studied. Extrapolating an image to give a distinctly larger one is much more challenging---a significant amount of additional content is needed which matches the original image, especially near its boundaries. We propose a data-driven approach to this problem. Given a source image, and the amount and direction(s) in which it is to be extrapolated, our system determines visually consistent content for the extrapolated regions using library images. As well as considering low-level matching, we achieve consistency at a higher level by using graph proxies for regions of source and library images. Treating images as graphs allows us to find candidates for image extrapolation in a feasible time. Consistency of subgraphs in source and library images is used to find good candidates for the additional content; these are then further filtered. Region boundary curves are aligned to ensure consistency where image parts are joined using a photomontage method. We demonstrate the power of our method in image editing applications. Miao Wang 0004, Yukun Lai, Ralph R. Martin, Shi-Min Hu 0001 |
ACM Trans. Graph. | 1 |
| 2013 | Aesthetic Image Enhancement by Dependence-Aware Object RecompositionabstractThis paper proposes an image-enhancement method to optimize photograph composition by rearranging foreground objects in the photograph. To adjust objects' positions while keeping the original scene content, we first perform a novel structure dependence analysis on the image to obtain the dependencies between all background regions. To determine the optimal positions for foreground objects, we formulate an optimization problem based on widely used heuristics for aesthetically pleasing pictures. Semantic relations between foreground objects are also taken into account during optimization. The final output is produced by moving foreground objects, together with their dependent regions, to optimal positions. The results show that our approach can effectively optimize photographs with single or multiple foreground objects without compromising the original photograph content. Miao Wang 0004, Shi-Min Hu 0001 |
IEEE Trans. Multim. | 2 |
| 2013 | PatchNet: a patch-based image representation for interactive library-driven image editingabstractWe introduce PatchNets , a compact, hierarchical representation describing structural and appearance characteristics of image regions, for use in image editing. In a PatchNet, an image region with coherent appearance is summarized by a graph node, associated with a single representative patch, while geometric relationships between different regions are encoded by labelled graph edges giving contextual information. The hierarchical structure of a PatchNet allows a coarse-to-fine description of the image. We show how this PatchNet representation can be used as a basis for interactive, library-driven, image editing. The user draws rough sketches to quickly specify editing constraints for the target image. The system then automatically queries an image library to find semantically-compatible candidate regions to meet the editing goal. Contextual image matching is performed using the PatchNet representation, allowing suitable regions to be found and applied in a few seconds, even from a library containing thousands of images. Shi-Min Hu 0001, Miao Wang 0004, Ralph R. Martin, Jue Wang 0001 |
ACM Trans. Graph. | 3 |