Feng Liu 0015

dblp:77/1318-15 · DBLP profile ↗
← Back
97ranked-venue papers
16as first author
25since 2021 · last 2026
0000-0002-5399-6214ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 79 · 13 first-author · 22 since 2021Artificial intelligence and machine learning · 39 · 5 first-author · 17 since 2021Human-computer interaction and ubiquitous computing · 12 · 1 first-author · 2 since 2021Computer networks · 5 · 1 first-authorSecurity and privacy · 1Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
YearPublicationVenuePosition
2026 CineVerse: Consistent Keyframe Synthesis for Cinematic Scene Composition
abstract
Multi-shot generation requires preserving the identity of characters and settings across frames. Cinematic scene composition goes beyond standard multi-shot generation, introducing additional challenges such as expressing complex interactions among multiple characters and visual effects to convey creative narratives—challenges existing datasets cannot fully address. We present CineVerse, a large-scale dataset of diverse movie scenes labeled with shot-level annotations tailored for filmmaking. CineVerse includes refined scene descriptions, shot-type information, and newly extracted shot, character, setting descriptions. We validate our dataset by developing a baseline framework that first generates a scene plan containing detailed information for the overall scene and each individual shot, then produces a set of coherent keyframes. Our results show significant improvements in controlling and synthesizing cinematic content through the added context provided by CineVerse.
Quynh Phung, Long Mai, Fabian Caba Heilbron, Feng Liu 0015, Jia-Bin Huang 0001, Cusuh Ham
WACV4
2025 Move-in-2D: 2D-Conditioned Human Motion Generation
abstract
Generating realistic human videos remains a challenging task, with the most effective methods currently relying on a human motion sequence as a control signal. Existing approaches often use existing motion extracted from other videos, which restricts applications to specific motion types and global scene matching. We propose Move-in-2D, a novel approach to generate human motion sequences conditioned on a scene image, allowing for diverse motion that adapts to different scenes. Our approach utilizes a diffusion model that accepts both a scene image and text prompt as inputs, producing a motion sequence tailored to the scene. To train this model, we collect a large-scale video dataset featuring single-human activities, annotating each video with the corresponding human motion as the target output. Experiments demonstrate that our method effectively predicts human motion that aligns with the scene image after projection. Furthermore, we show that the generated motion sequence improves human motion quality in video synthesis tasks.
Hsin-Ping Huang, Yang Zhou 0009, Jui-Hsien Wang, Difan Liu, Feng Liu 0015, Ming-Hsuan Yang 0001
CVPR5
2025 Visual Persona: Foundation Model for Full-Body Human Customization
abstract
We introduce Visual Persona, a foundation model for text-to-image full-body human customization that, given a single in-the-wild human image, generates diverse images of the individual guided by text descriptions. Unlike prior methods that focus solely on preserving facial identity, our approach captures detailed full-body appearance, aligning with text descriptions for body structure and scene variations. Training this model requires large-scale paired human data, consisting of multiple images per individual with consistent full-body identities, which is notoriously difficult to obtain. To address this, we propose a data curation pipeline leveraging vision-language models to evaluate full-body appearance consistency, resulting in Visual Persona-500K—a dataset of 580k paired human images across 100k unique identities. For precise appearance transfer, we introduce a transformer encoder-decoder architecture adapted to a pre-trained text-to-image diffusion model, which augments the input image into distinct body regions, encodes these regions as local appearance features, and projects them into dense identity embeddings independently to condition the diffusion model for synthesizing customized images. Visual Persona consistently surpasses existing approaches, generating high-quality, customized images from in-the-wild inputs. Extensive ablation studies validate design choices, and we demonstrate the versatility of Visual Persona across various downstream tasks.
Jisu Nam, Soowon Son, Jing Shi 0005, Difan Liu, Feng Liu 0015, Seungryong Kim, Yang Zhou 0009
CVPR6
2025 VideoGigaGAN: Towards Detail-rich Video Super-Resolution
abstract
Video super-resolution (VSR) models achieve temporal consistency but often produce blurrier results than their image-based counterparts due to limited generative capacity. This prompts the question: can we adapt a generative image upsampler for VSR while preserving temporal consistency? We introduce VideoGigaGAN, a new generative VSR model that combines high-frequency detail with temporal stability, building on the large-scale GigaGAN image upsampler. Simple adaptations of GigaGAN for VSR led to flickering issues, so we propose techniques to enhance temporal consistency. We validate the effectiveness of VideoGigaGAN by comparing it with state-of-the-art VSR models on public datasets and showcasing video results with 8× upsampling.
Taesung Park, Richard Zhang 0001, Yang Zhou 0009, Eli Shechtman, Feng Liu 0015, Jia-Bin Huang 0001, Difan Liu
CVPR6
2025 Progressive Growing of Video Tokenizers for Temporally Compact Latent Spaces
Aniruddha Mahapatra, Long Mai, David Bourgin, Feng Liu 0015
ICCV5
2025 REGEN: Learning Compact Video Embedding with (Re-)Generative Decoder
abstract
We present a novel perspective on learning video embedders for generative modeling: rather than requiring an exact reproduction of an input video, an effective embedder should focus on synthesizing visually plausible reconstructions. This relaxed criterion enables substantial improvements in compression ratios without compromising the quality of downstream generative models. Specifically, we propose replacing the conventional encoder-decoder video embedder with an encoder-generator framework that employs a diffusion transformer (DiT) to synthesize missing details from a compact latent space. Therein, we develop a dedicated latent conditioning module to condition the DiT decoder on the encoded video latent embedding. Our experiments demonstrate that our approach enables superior encoding-decoding performance compared to state-of-the-art methods, particularly as the compression ratio increases. To demonstrate the efficacy of our approach, we report results from our video embedders achieving a temporal compression ratio of up to 32x (8x higher than leading video embedders) and validate the robustness of this ultra-compact latent space for text-to-video generation, providing a significant efficiency boost in latent diffusion model training and inference.
Long Mai, Aniruddha Mahapatra, David Bourgin, Yicong Hong, Jonah Casebeer, Feng Liu 0015, Yun Fu 0001
ICCV7
2024 SNED: Superposition Network Architecture Search for Efficient Video Diffusion Model
abstract
While AI-generated content has garnered significant attention, achieving photo-realistic video synthesis remains a formidable challenge. Despite the promising advances in diffusion models for video generation quality, the complex model architecture and substantial computational demands for both training and inference create a significant gap between these models and real-world applications. This paper presents SNED, a superposition network architecture search method for efficient video diffusion model. Our method employs a supernet training paradigm that targets various model cost and resolution options using a weight-sharing method. Moreover, we propose the supernet training sampling warm-up for fast training optimization. To showcase the flexibility of our method, we conduct experiments involving both pixel-space and latent-space video diffusion models. The results demonstrate that our framework consistently produces comparable results across different model options with high efficiency. According to the experiment for the pixel-space video diffusion model, we can achieve consistent video generation results simultaneously across 64×64 to 256×256 resolutions with a large range of model sizes from 640M to 1.6B number of parameters for pixel-space video diffusion models.
Zhengang Li 0001, Yuchen Liu 0002, Difan Liu, Tobias Hinz, Feng Liu 0015, Yanzhi Wang 0001
CVPR6
2024 HARIVO: Harnessing Text-to-Image Models for Video Generation
Mingi Kwon, Seoung Wug Oh, Yang Zhou 0009, Difan Liu, Joon-Young Lee, Haoran Cai, Baqiao Liu, Feng Liu 0015, Youngjung Uh
ECCV (53)8
2024 Fast View Synthesis of Casual Videos with Soup-of-Planes
Yao-Chih Lee, Zhoutong Zhang, Kevin Matzen, Simon Niklaus, Jianming Zhang 0001, Jia-Bin Huang 0001, Feng Liu 0015
ECCV (38)7
2024 Customize-A-Video: One-Shot Motion Customization of Text-to-Video Diffusion Models
Yixuan Ren, Yang Zhou 0009, Jimei Yang, Jing Shi 0005, Difan Liu, Feng Liu 0015, Mingi Kwon, Abhinav Shrivastava
ECCV (89)6
2024 LRM: Large Reconstruction Model for Single Image to 3D
abstract
We propose the first Large Reconstruction Model (LRM) that predicts the 3D model of an object from a single input image within just 5 seconds. In contrast to many previous methods that are trained on small-scale datasets such as ShapeNet in a category-specific fashion, LRM adopts a highly scalable transformer-based architecture with 500 million learnable parameters to directly predict a neural radiance field (NeRF) from the input image. We train our model in an end-to-end manner on massive multi-view data containing around 1 million objects, including both synthetic renderings from Objaverse and real captures from MVImgNet. This combination of a high-capacity model and large-scale training data empowers our model to be highly generalizable and produce high-quality 3D reconstructions from various testing inputs, including real-world in-the-wild captures and images created by generative models. Video demos and interactable 3D meshes can be found on our LRM project webpage: https://yiconghong.me/LRM.
Yicong Hong, Kai Zhang 0045, Jiuxiang Gu, Sai Bi, Yang Zhou 0009, Difan Liu, Feng Liu 0015, Kalyan Sunkavalli, Trung Bui, Hao Tan 0002
ICLR7
2024 Auxiliary Features-Guided Super Resolution for Monte Carlo Rendering
abstract
Abstract This paper investigates super‐resolution to reduce the number of pixels to render and thus speed up Monte Carlo rendering algorithms. While great progress has been made to super‐resolution technologies, it is essentially an ill‐posed problem and cannot recover high‐frequency details in renderings. To address this problem, we exploit high‐resolution auxiliary features to guide super‐resolution of low‐resolution renderings. These high‐resolution auxiliary features can be quickly rendered by a rendering engine and at the same time provide valuable high‐frequency details to assist super‐resolution. To this end, we develop a cross‐modality transformer network that consists of an auxiliary feature branch and a low‐resolution rendering branch. These two branches are designed to fuse high‐resolution auxiliary features with the corresponding low‐resolution rendering. Furthermore, we design Residual Densely Connected Swin Transformer groups to learn to extract representative features to enable high‐quality super‐resolution. Our experiments show that our auxiliary features‐guided super‐resolution method outperforms both super‐resolution methods and Monte Carlo denoising methods in producing high‐quality renderings.
Qiqi Hou, Feng Liu 0015
Comput. Graph. Forum2
2024 Scalable video transformer for full-frame video prediction
Feng Liu 0015
Comput. Vis. Image Underst.2
2023 AnyFlow: Arbitrary Scale Optical Flow with Implicit Neural Representation
abstract
To apply optical flow in practice, it is often necessary to resize the input to smaller dimensions in order to reduce computational costs. However, downsizing inputs makes the estimation more challenging because objects and motion ranges become smaller. Even though recent approaches have demonstrated high-quality flow estimation, they tend to fail to accurately model small objects and precise boundaries when the input resolution is lowered, restricting their applicability to high-resolution inputs. In this paper, we introduce AnyFlow, a robust network that estimates accurate flow from images of various resolutions. By representing optical flow as a continuous coordinate-based representation, AnyFlow generates outputs at arbitrary scales from low-resolution inputs, demonstrating superior performance over prior works in capturing tiny objects with detail preservation on a wide range of scenes. We establish a new state-of-the-art performance of cross-dataset generalization on the KITTI dataset, while achieving comparable accuracy on the online benchmarks to other SOTA methods.
Hyunyoung Jung 0001, Zhuo Hui, Feng Liu 0015, Sungjoo Yoo, Denis Demandolx
CVPR5
2023 View Synthesis of Dynamic Scenes Based on Deep 3D Mask Volume
abstract
Image view synthesis has seen great success in reconstructing photorealistic visuals, thanks to deep learning and various novel representations. The next key step in immersive virtual experiences is view synthesis of dynamic scenes. However, several challenges exist due to the lack of high-quality training datasets, and the additional time dimension for videos of dynamic scenes. To address this issue, we introduce a multi-view video dataset, captured with a custom 10-camera rig in 120FPS. The dataset contains 96 high-quality scenes showing various visual effects and human interactions in outdoor scenes. We develop a new algorithm, Deep 3D Mask Volume, which enables temporally-stable view extrapolation from binocular videos of dynamic scenes, captured by static cameras. Our algorithm addresses the temporal inconsistency of disocclusions by identifying the error-prone areas with a 3D mask volume, and replaces them with static background observed throughout the video. Our method enables manipulation in 3D space as opposed to simple 2D masks, We demonstrate better temporal stability than frame-by-frame static view synthesis methods, or those that use 2D masks. The resulting view synthesis videos show minimal flickering artifacts and allow for larger translational movements.
Kai-En Lin, Lei Xiao 0014, Feng Liu 0015, Ravi Ramamoorthi
IEEE Trans. Pattern Anal. Mach. Intell.4
2022 Motion-Adjustable Neural Implicit Video Representation
abstract
Implicit neural representation (INR) has been successful in representing static images. Contemporary image-based INR, with the use of Fourier-based positional encoding, can be viewed as a mapping from sinusoidal patterns with different frequencies to image content. Inspired by that view, we hypothesize that it is possible to generate temporally varying content with a single image-based INR model by displacing its input sinusoidal patterns over time. By exploiting the relation between the phase information in sinusoidal functions and their displacements, we incorporate into the conventional image-based INR model a phase-varying positional encoding module, and couple it with a phase-shift generation module that determines the phase-shift values at each frame. The model is trained end-to-end on a video to jointly determine the phase-shift values at each time with the mapping from the phase-shifted sinusoidal functions to the corresponding frame, enabling an implicit video representation. Experiments on a wide range of videos suggest that such a model is capable of learning to interpret phase-varying positional embeddings into the corresponding time-varying content. More importantly, we found that the learned phase-shift vectors tend to capture meaningful temporal and motion information from the video. In particular, manipulating the phase-shift vectors induces meaningful changes in the temporal dynamics of the resulting video, enabling non-trivial temporal and motion editing effects such as temporal interpolation, motion magnification, motion smoothing, and video loop detection.
Long Mai, Feng Liu 0015
CVPR2
2022 Shift-Tolerant Perceptual Similarity Metric
Abhijay Ghildyal, Feng Liu 0015
ECCV (18)2
2022 A Perceptual Quality Metric for Video Frame Interpolation
Qiqi Hou, Abhijay Ghildyal, Feng Liu 0015
ECCV (15)3
2022 Future Frame Synthesis for Fast Monte Carlo Rendering
Carl S. Marshall, Deepak S. Vembar, Feng Liu 0015
Graphics Interface4
2022 Sequential Frame-Interpolation and DCT-based Video Compression Framework
abstract
Video data is ubiquitous; capturing, transferring, and storing even compressed video data is challenging because it requires substantial resources. With the large amount of video traffic being transmitted on the internet, any improvement in compressing such data, even small, can drastically impact resource consumption. In this paper, we present a hybrid video compression framework that unites the advantages of both DCT-based and interpolation-based video compression methods in a single framework. We show that our work can deliver the same visual quality or, in some cases, improve visual quality while reducing the bandwidth by 10--20%.
Yeganeh Jalalpour, Wu-chi Feng, Feng Liu 0015
MMAsia3
2022 Are High-quality Photos More Popular Than Low-quality Ones? A Quantitative Study
abstract
Social media popularity has increased over the years. Millions of images are uploaded and shared on social media daily. Image popularity prediction is extensively studied over past years. Several research works have shown that image content and social context play an important role in predicting image popularity. However, the impact of image aesthetics and quality on popularity has not been studied in detail until now. In this paper, we report a detailed study and analysis for finding the correlation between image quality/aesthetics and image popularity. To solidify the understanding of the impact, we present experimental results that are conducted on SMP 2020 and ICIP image popularity 2020 datasets to investigate if a high quality image or an aesthetically pleasing image is likely to be popular. We found that image aesthetics is moderately correlated with image popularity while quality is less so.
Priyanka Mudgal, Feng Liu 0015
MMSP2
2022 SNeRF: stylized neural implicit representations for 3D scenes
abstract
This paper presents a stylized novel view synthesis method. Applying state-of-the-art stylization methods to novel views frame by frame often causes jittering artifacts due to the lack of cross-view consistency. Therefore, this paper investigates 3D scene stylization that provides a strong inductive bias for consistent novel view synthesis. Specifically, we adopt the emerging neural radiance fields (NeRF) as our choice of 3D scene representation for their capability to render high-quality novel views for a variety of scenes. However, as rendering a novel view from a NeRF requires a large number of samples, training a stylized NeRF requires a large amount of GPU memory that goes beyond an off-the-shelf GPU capacity. We introduce a new training method to address this problem by alternating the NeRF and stylization optimization steps. Such a method enables us to make full use of our hardware memory capacity to both generate images at higher resolution and adopt more expressive image style transfer methods. Our experiments show that our method produces stylized NeRFs for a wide range of content, including indoor, outdoor and dynamic scenes, and synthesizes high-quality novel views with cross-view consistency.
Thu Nguyen-Phuoc, Feng Liu 0015, Lei Xiao 0014
ACM Trans. Graph.2
2021 Fast Monte Carlo Rendering via Multi-Resolution Sampling
Qiqi Hou, Carl S. Marshall, Selvakumar Panneer, Feng Liu 0015
Graphics Interface5
2021 Deep 3D Mask Volume for View Synthesis of Dynamic Scenes
abstract
Image view synthesis has seen great success in reconstructing photorealistic visuals, thanks to deep learning and various novel representations. The next key step in immersive virtual experiences is view synthesis of dynamic scenes. However, several challenges exist due to the lack of high-quality training datasets, and the additional time dimension for videos of dynamic scenes. To address this issue, we introduce a multi-view video dataset, captured with a custom 10-camera rig in 120FPS. The dataset contains 96 high-quality scenes showing various visual effects and human interactions in outdoor scenes. We develop a new algorithm, Deep 3D Mask Volume, which enables temporally-stable view extrapolation from binocular videos of dynamic scenes, captured by static cameras. Our algorithm addresses the temporal inconsistency of disocclusions by identifying the error-prone areas with a 3D mask volume, and replaces them with static background observed throughout the video. Our method enables manipulation in 3D space as opposed to simple 2D masks, We demonstrate better temporal stability than frame-by-frame static view synthesis methods, or those that use 2D masks. The resulting view synthesis videos show minimal flickering artifacts and allow for larger translational movements.
Kai-En Lin, Lei Xiao 0014, Feng Liu 0015, Ravi Ramamoorthi
ICCV3
2021 Learned Dual-View Reflection Removal
abstract
Traditional reflection removal algorithms either use a single image as input, which suffers from intrinsic ambiguities, or use multiple images from a moving camera, which is inconvenient for users. We instead propose a learning-based dereflection algorithm that uses stereo images as input. This is an effective trade-off between the two extremes: the parallax between two views provides cues to remove reflections, and two views are easy to capture due to the adoption of stereo cameras in smartphones. Our model consists of a learning-based reflection-invariant flow model for dual-view registration, and a learned synthesis model for combining aligned image pairs. Because no dataset for dual-view reflection removal exists, we render a synthetic dataset of dual-views with and without reflections for use in training. Our evaluation on an additional real-world dataset of stereo pairs shows that our algorithm outperforms existing single-image and multi-image dereflection approaches.
Simon Niklaus, Xuaner Cecilia Zhang, Jonathan T. Barron, Neal Wadhwa, Rahul Garg 0002, Feng Liu 0015, Tianfan Xue
WACV6
2020 Deep Homography Estimation for Dynamic Scenes
abstract
Homography estimation is an important step in many computer vision problems. Recently, deep neural network methods have shown to be favorable for this problem when compared to traditional methods. However, these new methods do not consider dynamic content in input images. They train neural networks with only image pairs that can be perfectly aligned using homographies. This paper investigates and discusses how to design and train a deep neural network that handles dynamic scenes. We first collect a large video dataset with dynamic content. We then develop a multi-scale neural network and show that when properly trained using our new dataset, this neural network can already handle dynamic scenes to some extent. To estimate a homography of a dynamic scene in a more principled way, we need to identify the dynamic content. Since dynamic content detection and homography estimation are two tightly coupled tasks, we follow the multi-task learning principles and augment our multi-scale network such that it jointly estimates the dynamics masks and homographies. Our experiments show that our method can robustly estimate homography for challenging scenarios with dynamic scenes, blur artifacts, or lack of textures.
Hoang Le, Feng Liu 0015, Aseem Agarwala
CVPR2
2020 Softmax Splatting for Video Frame Interpolation
abstract
Differentiable image sampling in the form of backward warping has seen broad adoption in tasks like depth estimation and optical flow prediction. In contrast, how to perform forward warping has seen less attention, partly due to additional challenges such as resolving the conflict of mapping multiple pixels to the same target location in a differentiable way. We propose softmax splatting to address this paradigm shift and show its effectiveness on the application of frame interpolation. Specifically, given two input frames, we forward-warp the frames and their feature pyramid representations based on an optical flow estimate using softmax splatting. In doing so, the softmax splatting seamlessly handles cases where multiple source pixels map to the same target location. We then use a synthesis network to predict the interpolation result from the warped representations. Our softmax splatting allows us to not only interpolate frames at an arbitrary time but also to fine tune the feature pyramid and the optical flow. We show that our synthesis approach, empowered by softmax splatting, achieves new state-of-the-art results for video frame interpolation.
Simon Niklaus, Feng Liu 0015
CVPR2
2020 FID: Frame Interpolation and DCT-based Video Compression
abstract
In this paper, we present a hybrid video compression technique that combines the advantages of residual coding techniques found in traditional DCT-based video compression and learning-based video frame interpolation to reduce the amount of residual data that needs to be compressed. Learning-based frame interpolation techniques use machine learning algorithms to predict frames but have difficulty with uncovered areas and non-linear motion. This approach uses DCT-based residual coding only on areas that are difficult for video interpolation and provides tunable compression for such areas through an adaptive selection of data to be encoded. Experimental data for both PSNR and the newer video multi-method assessment fusion (VMAF) metrics are provided. Our results show that we can reduce the amount of data required to represent a video stream compared with traditional video coding while outperforming video frame interpolation techniques in quality.
Yeganeh Jalalpour, Li-Yun Wang, Wu-chi Feng, Feng Liu 0015
ISM4
2019 Context-Aware Image Matting for Simultaneous Foreground and Alpha Estimation
abstract
Natural image matting is an important problem in computer vision and graphics. It is an ill-posed problem when only an input image is available without any external information. While the recent deep learning approaches have shown promising results, they only estimate the alpha matte. This paper presents a context-aware natural image matting method for simultaneous foreground and alpha matte estimation. Our method employs two encoder networks to extract essential information for matting. Particularly, we use a matting encoder to learn local features and a context encoder to obtain more global context information. We concatenate the outputs from these two encoders and feed them into decoder networks to simultaneously estimate the foreground and alpha matte. To train this whole deep neural network, we employ both the standard Laplacian loss and the feature loss: the former helps to achieve high numerical performance while the latter leads to more perceptually plausible results. We also report several data augmentation strategies that greatly improve the network's generalization performance. Our qualitative and quantitative experiments show that our method enables high-quality matting for a single natural image.
Qiqi Hou, Feng Liu 0015
ICCV2
2019 Appearance Flow Completion for Novel View Synthesis
abstract
Abstract Novel view synthesis from sparse and unstructured input views faces challenges like the difficulty with dense 3D reconstruction and large occlusion. This paper addresses these problems by estimating proper appearance flows from the target to input views to warp and blend the input views. Our method first estimates a sparse set 3D scene points using an off‐the‐shelf 3D reconstruction method and calculates sparse flows from the target to input views. Our method then performs appearance flow completion to estimate the dense flows from the corresponding sparse ones. Specifically, we design a deep fully convolutional neural network that takes sparse flows and input views as input and outputs the dense flows. Furthermore, we estimate the optical flows between input views as references to guide the estimation of dense flows between the target view and input views. Besides the dense flows, our network also estimates the masks to blend multiple warped inputs to render the target view. Experiments on the KITTI benchmark show that our method can generate high quality novel views from sparse and unstructured input views.
Hoang Le, Feng Liu 0015
Comput. Graph. Forum2
2019 3D Ken Burns effect from a single image
abstract
The Ken Burns effect allows animating still images with a virtual camera scan and zoom. Adding parallax, which results in the 3D Ken Burns effect, enables significantly more compelling results. Creating such effects manually is time-consuming and demands sophisticated editing skills. Existing automatic methods, however, require multiple input images from varying viewpoints. In this paper, we introduce a framework that synthesizes the 3D Ken Burns effect from a single image, supporting both a fully automatic mode and an interactive mode with the user controlling the camera. Our framework first leverages a depth prediction pipeline, which estimates scene depth that is suitable for view synthesis tasks. To address the limitations of existing depth estimation methods such as geometric distortions, semantic distortions, and inaccurate depth boundaries, we develop a semantic-aware neural network for depth prediction, couple its estimate with a segmentation-based depth adjustment process, and employ a refinement neural network that facilitates accurate depth predictions at object boundaries. According to this depth estimate, our framework then maps the input image to a point cloud and synthesizes the resulting video frames by rendering the point cloud from the corresponding camera positions. To address disocclusions while maintaining geometrically and temporally coherent synthesis results, we utilize context-aware color- and depth-inpainting to fill in the missing information in the extreme views of the camera path, thus extending the scene geometry of the point cloud. Experiments with a wide variety of image content show that our method enables realistic synthesis results. Our study demonstrates that our system allows users to achieve better results while requiring little effort compared to existing solutions for the 3D Ken Burns effect creation.
Simon Niklaus, Long Mai, Jimei Yang, Feng Liu 0015
ACM Trans. Graph.4
2019 Joint Stabilization and Direction of 360° Videos
abstract
Three-hundred-sixty-degree (360°) video provides an immersive experience for viewers, allowing them to freely explore the world by turning their head. However, creating high-quality 360° video content can be challenging, as viewers may miss important events by looking in the wrong direction, or they may see things that ruin the immersion, such as stitching artifacts and the film crew. We take advantage of the fact that not all directions are equally likely to be observed; most viewers are more likely to see content located at “true north,” i.e., in front of them, due to ergonomic constraints. We therefore propose 360° video direction, where the video is jointly optimized to orient important events to the front of the viewer and visual clutter behind them, while producing smooth camera motion. Unlike traditional video, viewers can still explore the space as desired, but with the knowledge that the most important content is likely to be in front of them. Constraints can be user guided, either added directly on the equirectangular projection or by recording “guidance” viewing directions while watching the video in a VR headset or automatically computed, such as via visual saliency or forward-motion direction. To accomplish this, we propose a new motion estimation technique specifically designed for 360° video that outperforms the commonly used five-point algorithm on wide-angle video. We additionally formulate the direction problem as an optimization where a novel parametrization of spherical warping allows us to correct for some degree of parallax effects. We compare our approach to recent methods that address stabilization-only and converting 360° video to narrow field-of-view video. Our pipeline can also enable the viewing of wide-angle non-360° footage in a spherical 360° space, giving an immersive “virtual cinema” experience for a wide range of existing content filmed with first-person cameras.
Chengzhou Tang, Oliver Wang, Feng Liu 0015, Ping Tan 0002
ACM Trans. Graph.3
2018 Depth Conflict Reduction for Stereo VR Video Interfaces
abstract
Applications for viewing and editing 360° video often render user interface (UI) elements on top of the video. For stereoscopic video, in which the perceived depth varies over the image, the perceived depth of the video can conflict with that of the UI elements, creating discomfort and making it hard to shift focus. To address this problem, we explore two new techniques that adjust the UI rendering based on the video content. The first technique dynamically adjusts the perceived depth of the UI to avoid depth conflict, and the second blurs the video in a halo around the UI. We conduct a user study to assess the effectiveness of these techniques in two stereoscopic VR video tasks: video watching with subtitles, and video search.
Cuong Nguyen 0003, Stephen DiVerdi, Aaron Hertzmann, Feng Liu 0015
CHI4
2018 Context-Aware Synthesis for Video Frame Interpolation
abstract
Video frame interpolation algorithms typically estimate optical flow or its variations and then use it to guide the synthesis of an intermediate frame between two consecutive original frames. To handle challenges like occlusion, bidirectional flow between the two input frames is often estimated and used to warp and blend the input frames. However, how to effectively blend the two warped frames still remains a challenging problem. This paper presents a context-aware synthesis approach that warps not only the input frames but also their pixel-wise contextual information and uses them to interpolate a high-quality intermediate frame. Specifically, we first use a pre-trained neural network to extract per-pixel contextual information for input frames. We then employ a state-of-the-art optical flow algorithm to estimate bidirectional flow between them and pre-warp both input frames and their context maps. Finally, unlike common approaches that blend the pre-warped frames, our method feeds them and their context maps to a video frame synthesis neural network to produce the interpolated frame in a context-aware fashion. Our neural network is fully convolutional and is trained end to end. Our experiments show that our method can handle challenging scenarios such as occlusion and large motion and outperforms representative state-of-the-art approaches.
Simon Niklaus, Feng Liu 0015
CVPR2
2018 Interactive Boundary Prediction for Object Selection
Hoang Le, Long Mai, Brian L. Price, Scott Cohen, Hailin Jin, Feng Liu 0015
ECCV (14)6
2017 Vremiere: In-Headset Virtual Reality Video Editing
abstract
Creative professionals are creating Virtual Reality (VR) experiences today by capturing spherical videos, but video editing is still done primarily in traditional 2D desktop GUI applications such as Premiere. These interfaces provide limited capabilities for previewing content in a VR headset or for directly manipulating the spherical video in an intuitive way. As a result, editors must alternate between editing on the desktop and previewing in the headset, which is tedious and interrupts the creative process. We demonstrate an application that enables a user to directly edit spherical video while fully immersed in a VR headset. We first interviewed professional VR filmmakers to understand current practice and derived a suitable workflow for in-headset VR video editing. We then developed a prototype system implementing this new workflow. Our system is built upon a familiar timeline design, but is enhanced with custom widgets to enable intuitive editing of spherical video inside the headset. We conducted an expert review study and found that with our prototype, experts were able to edit videos entirely within the headset. Experts also found our interface and widgets useful, providing intuitive controls for their editing needs.
Cuong Nguyen 0003, Stephen DiVerdi, Aaron Hertzmann, Feng Liu 0015
CHI4
2017 Spatial-Semantic Image Search by Visual Feature Synthesis
abstract
The performance of image retrieval has been improved tremendously in recent years through the use of deep feature representations. Most existing methods, however, aim to retrieve images that are visually similar or semantically relevant to the query, irrespective of spatial configuration. In this paper, we develop a spatial-semantic image search technology that enables users to search for images with both semantic and spatial constraints by manipulating concept text-boxes on a 2D query canvas. We train a convolutional neural network to synthesize appropriate visual features that captures the spatial-semantic constraints from the user canvas query. We directly optimize the retrieval performance of the visual features when training our deep neural network. These visual features then are used to retrieve images that are both spatially and semantically relevant to the user query. The experiments on large-scale datasets such as MS-COCO and Visual Genome show that our method outperforms other baseline and state-of-the-art methods in spatial-semantic image search.
Long Mai, Hailin Jin, Zhe Lin 0001, Jonathan Brandt, Feng Liu 0015
CVPR6
2017 Video Frame Interpolation via Adaptive Convolution
abstract
Video frame interpolation typically involves two steps: motion estimation and pixel synthesis. Such a two-step approach heavily depends on the quality of motion estimation. This paper presents a robust video frame interpolation method that combines these two steps into a single process. Specifically, our method considers pixel synthesis for the interpolated frame as local convolution over two input frames. The convolution kernel captures both the local motion between the input frames and the coefficients for pixel synthesis. Our method employs a deep fully convolutional neural network to estimate a spatially-adaptive convolution kernel for each pixel. This deep neural network can be directly trained end to end using widely available video data without any difficult-to-obtain ground-truth data like optical flow. Our experiments show that the formulation of video interpolation as a single convolution process allows our method to gracefully handle challenges like occlusion, blur, and abrupt brightness change and enables high-quality video frame interpolation.
Simon Niklaus, Long Mai, Feng Liu 0015
CVPR3
2017 Content and Surface Aware Projection
Long Mai, Hoang Le, Feng Liu 0015
Graphics Interface3
2017 Video Frame Interpolation via Adaptive Separable Convolution
abstract
Standard video frame interpolation methods first estimate optical flow between input frames and then synthesize an intermediate frame guided by motion. Recent approaches merge these two steps into a single convolution process by convolving input frames with spatially adaptive kernels that account for motion and re-sampling simultaneously. These methods require large kernels to handle large motion, which limits the number of pixels whose kernels can be estimated at once due to the large memory demand. To address this problem, this paper formulates frame interpolation as local separable convolution over input frames using pairs of 1D kernels. Compared to regular 2D kernels, the 1D kernels require significantly fewer parameters to be estimated. Our method develops a deep fully convolutional neural network that takes two input frames and estimates pairs of 1D kernels for all pixels simultaneously. Since our method is able to estimate kernels and synthesizes the whole video frame at once, it allows for the incorporation of perceptual loss to train the neural network to produce visually pleasing frames. This deep neural network is trained end-to-end using widely available video data without any human annotation. Both qualitative and quantitative experiments show that our method provides a practical solution to high-quality video frame interpolation.
Simon Niklaus, Long Mai, Feng Liu 0015
ICCV3
2017 Detecting Good Surface for Improvisatory Visual Projection
abstract
A projector is usually coupled with a dedicated projection surface to properly display visual information. This prevents the application of projection in places where a dedicated projection surface is not readily available. This paper presents a method for automatically detecting a good surface in a daily living and working space to support improvisatory projection without a pre-installed projection surface. Our method uses a projector-camera system that scans an environment and evaluates the quality of the environment surface for visual projection in two steps. Our method first excludes non-planar or highly-textured surface through epipolar geometry analysis and texture analysis. For a surface that passes the first test, our method further evaluates its quality for visual projection by quickly projecting the sampled projection content onto the surface and measuring the quality of the projected visual content. Our experiment shows that our method can reliably identify a good surface in a daily environment for high-quality visual projection.
Hoang Le, Thong Doan, Carl S. Marshall, Selvakumar Panneer, Feng Liu 0015
ISM5
2017 CollaVR: Collaborative In-Headset Review for VR Video
abstract
Collaborative review and feedback is an important part of conventional filmmaking and now Virtual Reality (VR) video production as well. However, conventional collaborative review practices do not easily translate to VR video because VR video is normally viewed in a headset, which makes it difficult to align gaze, share context, and take notes. This paper presents CollaVR, an application that enables multiple users to review a VR video together while wearing headsets. We interviewed VR video professionals to distill key considerations in reviewing VR video. Based on these insights, we developed a set of networked tools that enable filmmakers to collaborate and review video in real-time. We conducted a preliminary expert study to solicit feedback from VR video professionals about our system and assess their usage of the system with and without collaboration features.
Cuong Nguyen 0003, Stephen DiVerdi, Aaron Hertzmann, Feng Liu 0015
UIST4
2016 Gaze-based Notetaking for Learning from Lecture Videos
abstract
Taking notes has been shown helpful for learning. This activity, however, is not well supported when learning from watching lecture videos. The conventional video interface does not allow users to quickly locate and annotate important content in the video as notes. Moreover, users sometimes need to manually pause the video while taking notes, which is often distracting. In this paper, we develop a gaze-based system to assist a user in notetaking while watching lecture videos. Our system has two features to support notetaking. First, our system integrates offline video analysis and online gaze analysis to automatically detect and highlight key content from the lecture video for notetaking. Second, our system provides adaptive video control that automatically reduces the video playback speed or pauses it while a user is taking notes to minimize the user's effort in controlling video. Our study shows that our system enables users to take notes more easily and with better quality than the traditional video interface.
Cuong Nguyen 0003, Feng Liu 0015
CHI2
2016 Composition-Preserving Deep Photo Aesthetics Assessment
abstract
Photo aesthetics assessment is challenging. Deep convolutional neural network (ConvNet) methods have recently shown promising results for aesthetics assessment. The performance of these deep ConvNet methods, however, is often compromised by the constraint that the neural network only takes the fixed-size input. To accommodate this requirement, input images need to be transformed via cropping, scaling, or padding, which often damages image composition, reduces image resolution, or causes image distortion, thus compromising the aesthetics of the original images. In this paper, we present a composition-preserving deep Con-vNet method that directly learns aesthetics features from the original input images without any image transformations. Specifically, our method adds an adaptive spatial pooling layer upon the regular convolution and pooling layers to directly handle input images with original sizes and aspect ratios. To allow for multi-scale feature extraction, we develop the Multi-Net Adaptive Spatial Pooling ConvNet architecture which consists of multiple sub-networks with different adaptive spatial pooling sizes and leverage a scene-based aggregation layer to effectively combine the predictions from multiple sub-networks. Our experiments on the large-scale aesthetics assessment benchmark (AVA [29]) demonstrate that our method can significantly improve the state-of-the-art results in photo aesthetics assessment.
Long Mai, Hailin Jin, Feng Liu 0015
CVPR3
2016 Understanding the Impact of Compression on Feature Detection and Matching in Computer Vision
abstract
As video-based sensor networks continue to scale and become more ubiquitous, it is becoming increasingly important to focus systems research on techniques that support content-based decisions in real-time towards the edge of the network. While some prior work has focused on high-level image and video quality's effect on computer vision (e.g., object recognition). We are unaware of any work that focuses on the low-level details of why. This paper explores the impact of compression on underlying computer vision techniques. Specifically, this paper focuses on understanding the fundamental impact of compression on SIFT feature detection and matching. We show how reduced resolution or frame quality can negatively impact feature detection and tracking.
Wu-chi Feng, Ryan Feng, Paul Wyatt, Feng Liu 0015
ISM4
2016 Introduction to Special Issue MMSys/NOSSDAV 2015
abstract
No abstract available.
Feng Liu 0015, Wu-chi Feng, Michael Zink
ACM Trans. Multim. Comput. Commun. Appl.1
2016 Erratum to: Geometry-shader-based real-time voxelization and applications
Shu-Huai Chang, Yu-Chi Lai, Chih-Yuan Yao, Kai-Lung Hua, Yuzhen Niu, Feng Liu 0015
Vis. Comput.6
2015 Making Software Tutorial Video Responsive
abstract
Tutorial videos are widely available to help people use software. These videos, however, are viewed by users as captured and offer little direct interaction between users and software. This paper presents a video navigation method that allows users to interact with software tutorial video as if they were using the software. To make the tutorial video responsive, our method records the user interaction events like mouse click and drag during capturing the video. Our method then analyzes, selects, and visualizes these user interaction events at the event locations. When a user directly interacts with an event visualization, our method automatically navigates to the proper video frame to provide the visual feedback as if the software were responding to the user input. Thus, our method provides the experience of interacting with the software through directly manipulating the tutorial video. Our study shows our method can better help users follow tutorial videos to complete tasks than the baseline timeline interface.
Cuong Nguyen 0003, Feng Liu 0015
CHI2
2015 Kernel fusion for better image deblurring
abstract
Kernel estimation for image deblurring is a challenging task and a large number of algorithms have been developed. Our hypothesis is that while individual kernels estimated using different methods alone are sometimes inadequate, they often complement each other. This paper addresses the problem of fusing multiple kernels estimated using different methods into a more accurate one that can better support image deblurring than each individual kernel. In this paper, we develop a data-driven approach to kernel fusion that learns how each kernel contributes to the final kernel and how they interact with each other. We discuss various kernel fusion models and find that kernel fusion using Gaussian Conditional Random Fields performs best. This Gaussian Conditional Random Fields-based kernel fusion method not only models how individual kernels are fused at each kernel element but also the interaction of kernel fusion among multiple kernel elements. Our experiments show that our method can significantly improve image deblurring by combining kernels from multiple methods into a better one.
Long Mai, Feng Liu 0015
CVPR2
2015 Casual stereoscopic panorama stitching
abstract
This paper presents a method for stitching stereoscopic panoramas from stereo images casually taken using a stereo camera. This method addresses three challenges of stereoscopic image stitching: how to handle parallax, how to stitch the left- and right-view panorama consistently, and how to take care of disparity during stitching. This method addresses these challenges using a three-step approach. First, we employ a state-of-the-art stitching algorithm that handles parallax well to stitch the left views of input stereo images and create the left view of the final stereoscopic panorama. Second, we stitch the input disparity maps to obtain the target disparity map for the stereoscopic panorama by solving a Poisson's equation. This target disparity map is optimized to avoid vertical disparities and preserve the original perceived depth distribution. Finally, we warp the right views of the input stereo images and stitch them into the right-view panorama according to the target disparity map. The stitching of the right views is formulated as a labeling problem that is constrained by the stitching of the left views to make the left- and right-view panorama consistent to avoid “retinal rivalry”. Our experiments show that our method can effectively stitch casually taken stereo images and produce high-quality stereo panoramas that deliver a pleasant stereoscopic 3D viewing experience.
Feng Liu 0015
CVPR2
2015 Thumbnail-Preserving Encryption for JPEG
abstract
With more and more data being stored in the cloud, securing multimedia data is becoming increasingly important. Use of existing encryption methods with cloud services is possible, but makes many web-based applications difficult or impossible to use. In this paper, we propose a new image encryption scheme specially designed to protect JPEG images in cloud photo storage services. Our technique allows efficient reconstruction of an accurate low-resolution thumbnail from the ciphertext image, but aims to prevent the extraction of any more detailed information. This will allow efficient storage and retrieval of image data in the cloud but protect its contents from outside hackers or snooping cloud administrators. Experiments of the proposed approach using an online selfie database show that it can achieve a good balance of privacy, utility, image quality, and file size.
Charles V. Wright, Wu-chi Feng, Feng Liu 0015
IH&MMSec3
2015 A multi-lens stereoscopic synthetic video dataset
abstract
This paper describes a synthetically-generated, multi-lens stereoscopic video dataset and associated 3D models. Creating a multi-lens video stream requires small inter-lens spacing. While such cameras can be built out of off-the-shelf parts, they are not "professional" enough to allow for necessary requirements such as zoom-lens control or synchronization between cameras. Other dedicated devices exist but do not have sufficient resolution per image. This dataset provides 20 synthetic models, each with an associated multi-lens walkthrough, and the uncompressed video from its generation. This dataset can be used for multi-view compression, multi-view streaming, view-interpolation, or other computer graphics related research.
Wu-chi Feng, Feng Liu 0015
MMSys3
2014 Depth Enhancement via Low-Rank Matrix Completion
abstract
Depth captured by consumer RGB-D cameras is often noisy and misses values at some pixels, especially around object boundaries. Most existing methods complete the missing depth values guided by the corresponding color image. When the color image is noisy or the correlation between color and depth is weak, the depth map cannot be properly enhanced. In this paper, we present a depth map enhancement algorithm that performs depth map completion and de-noising simultaneously. Our method is based on the observation that similar RGB-D patches lie in a very low-dimensional subspace. We can then assemble the similar patches into a matrix and enforce this low-rank subspace constraint. This low-rank subspace constraint essentially captures the underlying structure in the RGB-D patches and enables robust depth enhancement against the noise or weak correlation between color and depth. Based on this subspace constraint, our method formulates depth map enhancement as a low-rank matrix completion problem. Since the rank of a matrix changes over matrices, we develop a data-driven method to automatically determine the rank number for each matrix. The experiments on both public benchmarks and our own captured RGB-D images show that our method can effectively enhance depth maps.
Si Lu, Xiaofeng Ren, Feng Liu 0015
CVPR3
2014 Parallax-Tolerant Image Stitching
abstract
Parallax handling is a challenging task for image stitching. This paper presents a local stitching method to handle parallax based on the observation that input images do not need to be perfectly aligned over the whole overlapping region for stitching. Instead, they only need to be aligned in a way that there exists a local region where they can be seamlessly blended together. We adopt a hybrid alignment model that combines homography and content-preserving warping to provide flexibility for handling parallax and avoiding objectionable local distortion. We then develop an efficient randomized algorithm to search for a homography, which, combined with content-preserving warping, allows for optimal stitching. We predict how well a homography enables plausible stitching by finding a plausible seam and using the seam cost as the quality metric. We develop a seam finding method that estimates a plausible seam from only roughly aligned images by considering both geometric alignment and image content. We then pre-align input images using the optimal homography and further use content-preserving warping to locally refine the alignment. We finally compose aligned images together using a standard seam-cutting algorithm and a multi-band blending algorithm. Our experiments show that our method can effectively stitch images with large parallax that are difficult for existing methods.
Feng Liu 0015
CVPR2
2014 Comparing Salient Object Detection Results without Ground Truth
Long Mai, Feng Liu 0015
ECCV (3)2
2014 Direct manipulation video navigation on touch screens
abstract
Direct Manipulation Video Navigation (DMVN) systems allow a user to directly drag an object of interest along its motion trajectory and have been shown effective for space-centric video browsing tasks. This paper designs touch-based interface techniques to support DMVN on touchscreen devices. While touch screens can suit DMVN systems naturally and enhance the directness during video navigation, the fat finger problems, such as precise selection and occlusion handling, must be properly addressed. In this paper, we discuss the effect of the fat finger problems on DMVN and develop three touch-based object dragging techniques for DMVN on touch screens, namely Offset Drag, Window Drag, and Drag Anywhere. We conduct user studies to evaluate our techniques as well as two baseline solutions on a smartphone and a desktop touch screen. Our studies show that two of our techniques can support DMVN on touch screen devices well and perform better than the baseline solutions.
Cuong Nguyen 0003, Yuzhen Niu, Feng Liu 0015
Mobile HCI3
2014 Oscillation analysis for salient object detection
Yang Liu 0009, Lei Wang 0184, Yuzhen Niu, Feng Liu 0015
Multim. Tools Appl.5
2014 Geometry-shader-based real-time voxelization and applications
Hsu-Huai Chang, Yu-Chi Lai, Chin-Yuan Yao, Kai-Lung Hua, Yuzhen Niu, Feng Liu 0015
Vis. Comput.6
2013 Direct manipulation video navigation in 3D
abstract
Direct Manipulation Video Navigation (DMVN) systems allow a user to navigate a video by dragging an object along its motion trajectory. These systems have been shown effective for space-centric video browsing. Their performance, however, is often limited by temporal ambiguities in a video with complex motion, such as recurring motion, self-intersecting motion, and pauses. The ambiguities come from reducing the 3D spatial-temporal motion (x, y, t) to the 2D spatial motion (x, y) in visualizing the motion and dragging the object. In this paper, we present a 3D DMVN system that maps the spatial-temporal motion (x, y, t) to 3D space (x, y, z) by mapping time t to depth z, visualizes the motion and video frame in 3D, and allows to navigate the video by spatial-temporally manipulating the object in 3D. We show that since our 3D DMVN system preserves all the motion information, it resolves the temporal ambiguities and supports intuitive navigation on challenging videos with complex motion.
Cuong Nguyen 0003, Yuzhen Niu, Feng Liu 0015
CHI3
2013 Saliency Aggregation: A Data-Driven Approach
abstract
A variety of methods have been developed for visual saliency analysis. These methods often complement each other. This paper addresses the problem of aggregating various saliency analysis methods such that the aggregation result outperforms each individual one. We have two major observations. First, different methods perform differently in saliency analysis. Second, the performance of a saliency analysis method varies with individual images. Our idea is to use data-driven approaches to saliency aggregation that appropriately consider the performance gaps among individual methods and the performance dependence of each method on individual images. This paper discusses various data-driven approaches and finds that the image-dependent aggregation method works best. Specifically, our method uses a Conditional Random Field (CRF) framework for saliency aggregation that not only models the contribution from individual saliency map but also the interaction between neighboring pixels. To account for the dependence of aggregation on an individual image, our approach selects a subset of images similar to the input image from a training data set and trains the CRF aggregation model only using this subset instead of the whole training set. Our experiments on public saliency benchmarks show that our aggregation method outperforms each individual saliency method and is robust with the selection of aggregated methods.
Long Mai, Yuzhen Niu, Feng Liu 0015
CVPR3
2013 Joint Subspace Stabilization for Stereoscopic Video
abstract
Shaky stereoscopic video is not only unpleasant to watch but may also cause 3D fatigue. Stabilizing the left and right view of a stereoscopic video separately using a monocular stabilization method tends to both introduce undesirable vertical disparities and damage horizontal disparities, which may destroy the stereoscopic viewing experience. In this paper, we present a joint subspace stabilization method for stereoscopic video. We prove that the low-rank subspace constraint for monocular video [10] also holds for stereoscopic video. Particularly, the feature trajectories from the left and right video share the same subspace. Based on this proof, we develop a stereo subspace stabilization method that jointly computes a common subspace from the left and right video and uses it to stabilize the two videos simultaneously. Our method meets the stereoscopic constraints without 3D reconstruction or explicit left-right correspondence. We test our method on a variety of stereoscopic videos with different scene content and camera motion. The experiments show that our method achieves high-quality stabilization for stereoscopic video in a robust and efficient way.
Feng Liu 0015, Yuzhen Niu, Hailin Jin
ICCV1
2013 Making stereo photo cropping easy
abstract
The increasing popularity of stereoscopic 3D brings the demand for tools for editing and authoring stereoscopic images and videos. This paper shows that even a simple task like cropping is difficult for amateur users with little stereoscopic photography knowledge. Unlike regular monocular (2D) images, cropping a stereoscopic image needs to be carefully executed to avoid stereoscopic violations, which otherwise cause an unpleasant stereoscopic viewing experience. In this paper, we present a system that assists in stereoscopic photo cropping by automatically measuring the stereoscopic photography violations and alerting users with the potential violations. Our study shows that compared to a popular stereoscopic photo editing system, our system makes stereoscopic photo cropping easier even for amateur users with little stereoscopic photography knowledge and provides a good user experience.
Yuzhen Niu, Feng Liu 0015
ICME3
2013 Eye Blink Detection for Smart Glasses
abstract
Eye blink is a quick action of closing and opening of the eyelids. Eye blink detection has a wide range of applications in human computer interaction and human vision health care research. Existing approaches to eye blink detection often cannot suit well resource-limited eye blink detection platforms like Smart Glasses, which have limited energy supply and typically cannot afford strong imaging and computational capabilities. In this paper, we present an efficient and robust eye blink detection method for Smart Glasses. Our method first employs an eigen-eye approach to detect closing-eye in individual video frames. Our method then learns eye blink patterns based on the closing-eye detection results and detects eye blinks using a Gradient Boosting method. Our method further uses a non-maximum suppression algorithm to remove repeated detection of the same eye-blink action among consecutive video frames. Experiments with our prototyped smart glasses equipped with a low-power camera and an embedded processor show an accurate detection result (with more than 96% accuracy) on video frames of a small size of 16 × 12 at 96 fps, which enables a number of applications in health care, driving safety, and human computer interaction.
Hoang Le, Thanh Dang, Feng Liu 0015
ISM3
2013 Addressing the semantic gap between video sensors and applications
abstract
In this paper, we propose a framework to support the bridging of applications and computer-vision based sensor networks. We argue that the semantic gap, the difference between the data collected in a sensor network and the information needed by the application, in video-based sensor networks can only be addressed by providing systems support in such a way that allows users and computing systems to meet in the middle. We first outline the vision of the system that we are working towards. We then describe initial experiments that we have conducted using a functional component of the system applied to real-world video data that is being collected by intelligent transportation systems researchers.
Wu-chi Feng, Feng Liu 0015, Thanh Dang
NOSSDAV3
2013 Casual Stereoscopic Photo Authoring
abstract
Stereoscopic 3D displays become more and more popular these years. However, authoring high-quality stereoscopic 3D content remains challenging. In this paper, we present a method for easy stereoscopic photo authoring with a regular (monocular) camera. Our method takes two images or video frames using a monocular camera as input and transforms them into a stereoscopic image pair that provides a pleasant viewing experience. The key technique of our method is a perceptual-plausible image rectification algorithm that warps the input image pairs to meet the stereoscopic geometric constraint while avoiding noticeable visual distortion. Our method uses spatially-varying mesh-based image warps. Our warping method encodes a variety of constraints to best meet the stereoscopic geometric constraint and minimize visual distortion. Since each energy term is quadratic, our method eventually formulates the warping problem as a quadratic energy minimization which is solved efficiently using a sparse linear solver. Our method also allows both local and global adjustments of the disparities, an important property for adapting resulting stereoscopic images to different viewing conditions. Our experiments demonstrate that our spatially-varying warping technique can better support casual stereoscopic photo authoring than existing methods and our results and user study show that our method can effectively use casually-taken photos to create high-quality stereoscopic photos that deliver a pleasant 3D viewing experience.
Feng Liu 0015, Yuzhen Niu, Hailin Jin
IEEE Trans. Multim.1
2013 Spatially and Temporally Optimized Video Stabilization
abstract
Properly handling parallax is important for video stabilization. Existing methods that achieve the aim require either 3D reconstruction or long feature trajectories to enforce the subspace or epipolar geometry constraints. In this paper, we present a robust and efficient technique that works on general videos. It achieves high-quality camera motion on videos where 3D reconstruction is difficult or long feature trajectories are not available. We represent each trajectory as a Bézier curve and maintain the spatial relations between trajectories by preserving the original offsets of neighboring curves. Our technique formulates stabilization as a spatial-temporal optimization problem that finds smooth feature trajectories and avoids visual distortion. The Bézier representation enables strong smoothness of each feature trajectory and reduces the number of variables in the optimization problem. We also stabilize videos in a streaming fashion to achieve scalability. The experiments show that our technique achieves high-quality camera motion on a variety of challenging videos that are difficult for existing methods.
Yu-Shuen Wang, Feng Liu 0015, Pu-Sheng Hsu, Tong-Yee Lee
IEEE Trans. Vis. Comput. Graph.2
2012 Video summagator: an interface for video summarization and navigation
abstract
This paper presents Video Summagator (VS), a volume-based interface for video summarization and navigation. VS models a video as a space-time cube and visualizes the video cube using real-time volume rendering techniques. VS empowers a user to interactively manipulate the video cube. We show that VS can quickly summarize both the static and dynamic video content by visualizing the space-time information in 3D. We demonstrate that VS enables a user to quickly look into the video cube, understand the content, and navigate to the content of interest.
Cuong Nguyen 0003, Yuzhen Niu, Feng Liu 0015
CHI3
2012 Leveraging stereopsis for saliency analysis
abstract
Stereopsis provides an additional depth cue and plays an important role in the human vision system. This paper explores stereopsis for saliency analysis and presents two approaches to stereo saliency detection from stereoscopic images. The first approach computes stereo saliency based on the global disparity contrast in the input image. The second approach leverages domain knowledge in stereoscopic photography. A good stereoscopic image takes care of its disparity distribution to avoid 3D fatigue. Particularly, salient content tends to be positioned in the stereoscopic comfort zone to alleviate the vergence-accommodation conflict. Accordingly, our method computes stereo saliency of an image region based on the distance between its perceived location and the comfort zone. Moreover, we consider objects popping out from the screen salient as these objects tend to catch a viewer's attention. We build a stereo saliency analysis benchmark dataset that contains 1000 stereoscopic images with salient object masks. Our experiments on this dataset show that stereo saliency provides a useful complement to existing visual saliency analysis and our method can successfully detect salient content from images that are difficult for monocular saliency analysis methods.
Yuzhen Niu, Yujie Geng, Feng Liu 0015
CVPR4
2012 Detecting rule of simplicity from photos
abstract
Simplicity refers to one of the most important photography composition rules. Simplicity states that simplifying the image background can draw viewers' attention to the subject of interest in a photograph and help them better comprehend and appreciate it. Understanding whether a photo respects photography rules or not facilitates photo quality assessment. In this paper, we present a method to automatically detect whether a photo is composed according to the rule of simplicity. We design features according to the definition, implementation and effect of the rule. First, we make use of saliency analysis to infer the subject of interest in a photo and measure its compactness. Second, we segment an image into background and foreground and measure the homogeneity within the background as another feature. Third, when looking at an image created with the rule of simplicity, different viewers tend to agree on what the subject of interest is in this photo. We accordingly measure the consistency among various saliency detection results as a feature. We experiment with these features in a range of machine learning methods. Our experiments show that our methods, together with these features, provide an encouraging result in detecting the rule of simplicity in a photo.
Long Mai, Hoang Le, Yuzhen Niu, Yu-Chi Lai, Feng Liu 0015
ACM Multimedia5
2012 Understanding the impact of inter-lens and temporal stereoscopic video compression
abstract
As we move toward more ubiquitous stereoscopic video, particularly with multiple (> 2) lenses, the need to understand the efficiency of compression will become increasingly important. In this paper, we explore the impact of spatial (between lenses) and temporal (over time) compression for stereoscopic video images. In particular, because stereoscopic images are taken at the same time, there is expected to be a high correlation between pixels in the horizontal direction due to the fixed nature of the multiple lenses. We propose a vertically reduced search window in order to take advantage of this correlation. Starting with multiple stereoscopic video sequences shot using a production studio 3D camera, we explore the effectiveness of temporal and inter-lens motion compensation for stereoscopic video. Furthermore, the experiments use exhaustive search to remove the effects of heuristic-based motion-compensation techniques.
Wu-chi Feng, Feng Liu 0015
NOSSDAV2
2012 Image resizing via non-homogeneous warping
Yuzhen Niu, Feng Liu 0015, Michael Gleicher
Multim. Tools Appl.2
2012 What Makes a Professional Video? A Computational Aesthetics Approach
abstract
Understanding the characteristics of high-quality professional videos is important for video classification, video quality measurement, and video enhancement. A professional video is good not only for its interesting story but also for its high visual quality. In this paper, we study what makes a professional video from the perspective of aesthetics. We discuss how a professional video is created and correspondingly design a variety of features that distinguish professional videos from amateur ones. We study general aesthetics features that are applied to still photos and extend them to videos. We design a variety of features that are particularly relevant to videos. We examined the performance of these features in the problem of professional and amateur video classification. Our experiments show that with these features, 97.3% professional and amateur shot classification accuracy rate is achieved on our own data set and 91.2% professional video detection rate is achieved on a public professional video set. Our experiments also show that the features that are particularly for videos are shown most effective for this task.
Yuzhen Niu, Feng Liu 0015
IEEE Trans. Circuits Syst. Video Technol.2
2012 Aesthetics-Based Stereoscopic Photo Cropping for Heterogeneous Displays
abstract
Stereoscopic displays are becoming ubiquitous, ranging from large 3-D TVs to small mobile phones. Stereoscopic photos need to be carefully adapted to be effectively viewed on the displays other than originally intended. In this paper, we present a method that can automatically crop and scale an existing stereoscopic photo to a variety of displays while preserving its aesthetic value. We formulate stereoscopic photo adaptation as an optimization problem that aims to preserve the aesthetic value of the input photo. We define a wide range of energy terms to preserve the stereoscopic photo aesthetics by borrowing rules from stereoscopic photography. Our experiments on a wide variety of stereoscopic photos demonstrate that our method can robustly produce display-dependent stereoscopic photos that deliver pleasant viewing experiences.
Yuzhen Niu, Feng Liu 0015, Wu-chi Feng, Hailin Jin
IEEE Trans. Multim.2
2012 Enabling warping on stereoscopic images
abstract
Warping is one of the basic image processing techniques. Directly applying existing monocular image warping techniques to stereoscopic images is problematic as it often introduces vertical disparities and damages the original disparity distribution. In this paper, we show that these problems can be solved by appropriately warping both the disparity map and the two images of a stereoscopic image. We accordingly develop a technique for extending existing image warping algorithms to stereoscopic images. This technique divides stereoscopic image warping into three steps. Our method first applies the user-specified warping to one of the two images. Our method then computes the target disparity map according to the user specified warping. The target disparity map is optimized to preserve the perceived 3D shape of image content after image warping. Our method finally warps the other image using a spatially-varying warping method guided by the target disparity map. Our experiments show that our technique enables existing warping methods to be effectively applied to stereoscopic images, ranging from parametric global warping to non-parametric spatially-varying warping.
Yuzhen Niu, Wu-chi Feng, Feng Liu 0015
ACM Trans. Graph.3
2011 Rule of Thirds Detection from Photograph
abstract
The rule of thirds is one of the most important composition rules used by photographers to create high-quality photos. The rule of thirds states that placing important objects along the imagery thirds lines or around their intersections often produces highly aesthetic photos. In this paper, we present a method to automatically determine whether a photo respects the rule of thirds. Detecting the rule of thirds from a photo requires semantic content understanding to locate important objects, which is beyond the state of the art. This paper makes use of the recent saliency and generic objectness analysis as an alternative and accordingly designs a range of features. Our experiment with a variety of saliency and generic objectness methods shows that an encouraging performance can be achieved in detecting the rule of thirds from photos.
Long Mai, Hoang Le, Yuzhen Niu, Feng Liu 0015
ISM4
2011 Systems support for stereoscopic video compression
abstract
In this paper, we propose a content-based, threaded stereoscopic video compression algorithm. We believe that it is likely future stereoscopic imaging systems will contain more than two lenses, allowing the display system to optimize the stereoscopic viewing experience. Furthermore, it will allow for the adjustment of framing (scene) composition errors that can arise from stereoscopic capture. Our proposed system uses 10 linearly aligned lenses. Upon capture, the images are run through a feature detection and matching algorithm in order to determine disparity between the stereoscopic images. During compression, the disparity measurements can be used to drive the selection of key frames within the image sets to provide better retrieval of data for display.
Wu-chi Feng, Feng Liu 0015, Yuzhen Niu, Scott Price
NOSSDAV2
2011 Subspace video stabilization
abstract
We present a robust and efficient approach to video stabilization that achieves high-quality camera motion for a wide range of videos. In this article, we focus on the problem of transforming a set of input 2D motion trajectories so that they are both smooth and resemble visually plausible views of the imaged scene; our key insight is that we can achieve this goal by enforcing subspace constraints on feature trajectories while smoothing them. Our approach assembles tracked features in the video into a trajectory matrix, factors it into two low-rank matrices, and performs filtering or curve fitting in a low-dimensional linear space. In order to process long videos, we propose a moving factorization that is both efficient and streamable. Our experiments confirm that our approach can efficiently provide stabilization results comparable with prior 3D methods in cases where those methods succeed, but also provides smooth camera motions in cases where such approaches often fail, such as videos that lack parallax. The presented approach offers the first method that both achieves high-quality video stabilization and is practical enough for consumer applications.
Feng Liu 0015, Michael Gleicher, Jue Wang 0001, Hailin Jin, Aseem Agarwala
ACM Trans. Graph.1
2010 Warp propagation for video resizing
abstract
This paper presents a video resizing approach that provides both efficiency and temporal coherence. Prior approaches either sacrifice temporal coherence (resulting in jitter), or require expensive spatio-temporal optimization. By assessing the requirements for video resizing we observe a fundamental tradeoff between temporal coherence in the background and shape preservation for the moving objects. Understanding this tradeoff enables us to devise a novel approach that is efficient, because it warps each frame independently, yet can avoid introducing jitter. Like previous approaches, our method warps frames so that the background are distorted similarly to prior frames while avoiding distortion of the moving objects. However, our approach introduces a motion history map that propagates information about the moving objects between frames, allowing for graceful tradeoffs between temporal coherence in the background and shape preservation for the moving objects. The approach can handle scenes with significant camera and object motion and avoid jitter, yet warp each frame sequentially for efficiency. Experiments with a variety of videos demonstrate that our approach can efficiently produce high-quality video resizing results.
Yuzhen Niu, Feng Liu 0015, Michael Gleicher
CVPR2
2010 Multi-View Video Summarization
abstract
Previous video summarization studies focused on monocular videos, and the results would not be good if they were applied to multi-view videos directly, due to problems such as the redundancy in multiple views. In this paper, we present a method for summarizing multi-view videos. We construct a spatio-temporal shot graph and formulate the summarization problem as a graph labeling task. The spatio-temporal shot graph is derived from a hypergraph, which encodes the correlations with different attributes among multi-view video shots in hyperedges. We then partition the shot graph and identify clusters of event-centered shots with similar contents via random walks. The summarization result is generated through solving a multi-objective optimization problem based on shot importance evaluated using a Gaussian entropy fusion scheme. Different summarization objectives, such as minimum summary length and maximum information coverage, can be accomplished in the framework. Moreover, multi-level summarization can be achieved easily by configuring the optimization parameters. We also propose the multi-view storyboard and event board for presenting multi-view summaries. The storyboard naturally reflects correlations among multi-view summarized shots that describe the same important event. The event-board serially assembles event-centered multi-view shots in temporal order. Single video summary which facilitates quick browsing of the summarized multi-view video can be easily generated based on the event board representation.
Yanwei Fu 0001, Yanwen Guo 0001, Yanshu Zhu, Feng Liu 0015, Chuanming Song 0001, Zhi-Hua Zhou
IEEE Trans. Multim.4
2010 Animation rendering with Population Monte Carlo image-plane sampler
Yu-Chi Lai, Stephen Chenney, Feng Liu 0015, Yuzhen Niu, Shaohua Fan
Vis. Comput.3
2009 Learning color and locality cues for moving object detection and segmentation
abstract
This paper presents an algorithm for automatically detecting and segmenting a moving object from a monocular video. Detecting and segmenting a moving object from a video with limited object motion is challenging. Since existing automatic algorithms rely on motion to detect the moving object, they cannot work well when the object motion is sparse and insufficient. In this paper, we present an unsupervised algorithm to learn object color and locality cues from the sparse motion information. We first detect key frames with reliable motion cues and then estimate moving sub-objects based on these motion cues using a Markov Random Field (MRF) framework. From these sub-objects, we learn an appearance model as a color Gaussian Mixture Model. To avoid the false classification of background pixels with similar color to the moving objects, the locations of these sub-objects are propagated to neighboring frames as locality cues. Finally, robust moving object segmentation is achieved by combining these learned color and locality cues with motion cues in a MRF framework. Experiments on videos with a variety of object and camera motion demonstrate the effectiveness of this algorithm.
Feng Liu 0015, Michael Gleicher
CVPR1
2009 Using Web Photos for Measuring Video Frame Interestingness
Feng Liu 0015, Yuzhen Niu, Michael Gleicher
IJCAI1
2009 Visual-Quality Optimizing Super Resolution
abstract
Abstract In this paper, we propose a robust image super‐resolution (SR) algorithm that aims to maximize the overall visual quality of SR results. We consider a good SR algorithm to be fidelity preserving, image detail enhancing and smooth. Accordingly, we define perception‐based measures for these visual qualities. Based on these quality measures, we formulate image SR as an optimization problem aiming to maximize the overall quality. Since the quality measures are quadratic, the optimization can be solved efficiently. Experiments on a large image set and subjective user study demonstrate the effectiveness of the perception‐based quality measures and the robustness and efficiency of the presented method.
Feng Liu 0015, Jinjun Wang, Shenghuo Zhu, Michael Gleicher, Yihong Gong
Comput. Graph. Forum1
2009 Image Retargeting Using Mesh Parametrization
abstract
Image retargeting aims to adapt images to displays of small sizes and different aspect ratios. Effective retargeting requires emphasizing the important content while retaining surrounding context with minimal visual distortion. In this paper, we present such an effective image retargeting method using saliency-based mesh parametrization. Our method first constructs a mesh image representation that is consistent with the underlying image structures. Such a mesh representation enables easy preservation of image structures during retargeting since it captures underlying image structures. Based on this mesh representation, we formulate the problem of retargeting an image to a desired size as a constrained image mesh parametrization problem that aims at finding a homomorphous target mesh with desired size. Specifically, to emphasize salient objects and minimize visual distortion, we associate image saliency into the image mesh and regard image structure as constraints for mesh parametrization. Through a stretch-based mesh parametrization process we obtain the homomorphous target mesh, which is then used to render the target image by texture mapping. The effectiveness of our algorithm is demonstrated by experiments.
Yanwen Guo 0001, Feng Liu 0015, Zhi-Hua Zhou, Michael Gleicher
IEEE Trans. Multim.2
2009 Content-preserving warps for 3D video stabilization
abstract
We describe a technique that transforms a video from a hand-held video camera so that it appears as if it were taken with a directed camera motion. Our method adjusts the video to appear as if it were taken from nearby viewpoints, allowing 3D camera movements to be simulated. By aiming only for perceptual plausibility, rather than accurate reconstruction, we are able to develop algorithms that can effectively recreate dynamic scenes from a single source video. Our technique first recovers the original 3D camera motion and a sparse set of 3D, static scene points using an off-the-shelf structure-from-motion system. Then, a desired camera path is computed either automatically (e.g., by fitting a linear or quadratic path) or interactively. Finally, our technique performs a least-squares optimization that computes a spatially-varying warp from each input video frame into an output frame. The warp is computed to both follow the sparse displacements suggested by the recovered 3D structure, and avoid deforming the content in the video frame. Our experiments on stabilizing challenging videos of dynamic scenes demonstrate the effectiveness of our technique.
Feng Liu 0015, Michael Gleicher, Hailin Jin, Aseem Agarwala
ACM Trans. Graph.1
2008 Texture-Consistent Shadow Removal
Feng Liu 0015, Michael Gleicher
ECCV (4)1
2008 Discovering panoramas in web videos
abstract
While methods for stitching panoramas have been successful given proper source images, providing these source images still remains a burden. In this paper, we present a method to discover panoramic source images within widely available web videos. The challenge comes from the fact that many of these videos are not recorded intentionally for stitching panoramas. Our method aims to find segments within a video that work as panorama sources. Specifically, we determine a video segment to be a valid panorama source according to the following three criteria. First, its camera motion should cover a wide field-of-view of the scene. Second, its frames should be "mosaicable", which states that the inter-frame motion should observe the underlying conditions for stitching a panorama. Third, its frames should have good image quality. Based on these criteria, we formulate discovering panoramas in a video as an optimization problem that aims to find an optimal set of video segments as panorama sources. After discovering these panorama sources, we synthesize regular scene panoramas using them. When significant dynamics is detected in the sources, we fuse the dynamics into the scene panoramas to make activity synopses to convey the dynamics. Our experiment of querying panoramas from YouTube confirms the feasibility of using web videos as panorama sources and demonstrates the effectiveness of our method.
Feng Liu 0015, Yu Hen Hu, Michael Gleicher
ACM Multimedia1
2008 Noisy video super-resolution
abstract
Low-quality videos often not only have limited resolution, but also suffer from noise. Directly up-sampling a video without considering noise could deteriorate its visual quality due to magnifying noise. This paper addresses this problem with a unified framework that achieves simultaneous de-noising and super-resolution. This framework formulates noisy video super-resolution as an optimization problem, aiming to maximize the visual quality of the result. We consider a good quality result to be fidelity-preserving, detailpreserving and smooth. Accordingly, we propose measures for these qualities in the scenario of de-noising and superresolution. The experiments on a variety of noisy videos demonstrate the effectiveness of the presented algorithm.
Feng Liu 0015, Jinjun Wang, Shenghuo Zhu, Michael Gleicher, Yihong Gong
ACM Multimedia1
2008 Re-cinematography: Improving the camerawork of casual video
abstract
This article presents an approach to postprocessing casually captured videos to improve apparent camera movement. Re-cinematography transforms each frame of a video such that the video better follows cinematic conventions. The approach breaks a video into shorter segments. Segments of the source video where there is no intentional camera movement are made to appear as if the camera is completely static. For segments with camera motions, camera paths are keyframed automatically and interpolated with matrix logarithms to give velocity-profiled movements that appear intentional and directed. Closeups are inserted to provide compositional variety in otherwise uniform segments. The approach automatically balances the tradeoff between motion smoothness and distortion to the original imagery. Results from our prototype show improvements to poor quality home videos.
Michael Gleicher, Feng Liu 0015
ACM Trans. Multim. Comput. Commun. Appl.2
2007 Re-cinematography: improving the camera dynamics of casual video
abstract
This paper presents an approach to post-processing casually captured videos to improve apparent camera movement. Re-cinematography transforms each frame of a video such that the video better follows cinematic conventions. The approach breaks videos into shorter segments. For segments of the source video where the camera is relatively static, re-cinematography uses image stabilization to make the result look locked-down. For segments with camera motions, camera paths are keyframed automatically and interpolated with matrix logarithms to give velocity-profiled movements that appear intentional and directed. The approach automatically balances the tradeoff between motion smoothness and distortion to the original imagery. Results from our prototype show improvements to poor quality home videos.
Michael Gleicher, Feng Liu 0015
ACM Multimedia2
2006 Region Enhanced Scale-Invariant Saliency Detection
abstract
Saliency measures the low-level stimuli to human vision, and serves as an alternative to semantic image understanding. This paper presents a region enhanced scale-invariant saliency detection method. Our method constructs a scale-invariant saliency map from an image, segments the image into regions, and enhances the saliency map with the region information. Compared with previous methods, our method has advantages in providing robust scale-invariant saliency, giving meaningful region information for applications, and eliminating misleading high-contrast edges
Feng Liu 0015, Michael Gleicher
ICME1
2006 Video retargeting: automating pan and scan
abstract
When a video is displayed on a smaller display than originally intended, some of the information in the video is necessarily lost. In this paper, we introduce Video Retargeting that adapts video to better suit the target display, minimizing the important information lost. We define a framework that measures the preservation of the source material, and methods for estimating the important information in the video. Video retargeting crops each frame and scales it to fit the target display. An optimization process minimizes information loss by balancing the loss of detail due to scaling with the loss of content and composition due to cropping. The cropping window can be moved during a shot to introduce virtual pans and cuts, subject to constraints that ensure cinematic plausibility. We demonstrate results of adapting a variety of source videos to small display sizes.
Feng Liu 0015, Michael Gleicher
ACM Multimedia1
2005 Automatic image retargeting with fisheye-view warping
abstract
Image retargeting is the problem of adapting images for display on devices different than originally intended. This paper presents a method for adapting large images, such as those taken with a digital camera, for a small display, such as a cellular telephone. The method uses a non-linear fisheye-view warp that emphasizes parts of an image while shrinking others. Like previous methods, fisheye-view warping uses image information, such as low-level salience and high-level object recognition to find important regions of the source image. However, unlike prior approaches, a non-linear image warping function emphasizes the important aspects of the image while retaining the surrounding context. The method has advantages in preserving information content, alerting the viewer to missing information and providing robustness.
Feng Liu 0015, Michael Gleicher
UIST1
2003 Intelligent Crowd Simulation
Feng Liu 0015, Ronghua Liang
ICCSA (1)1
2003 3D motion retrieval with motion index tree
Feng Liu 0015, Yueting Zhuang, Fei Wu 0001, Yunhe Pan
Comput. Vis. Image Underst.1
2002 Incomplete motion feature tracking algorithm in video sequences
abstract
To effectively track incomplete motion features, a novel feature tracking algorithm for motion capture is presented. According to feature attributes and relationship among features, extracted features are classified as four types of features. Then different strategies are applied to track different kinds of features. To verify the tracks, cross correlation test and predicted 3D model based test are used to test and remove outliers. Experimental results demonstrate the effectiveness of our algorithm.
Zhongxiang Luo, Yueting Zhuang, Feng Liu 0015, Yunhe Pan
ICIP (3)3
2002 Multiple animated characters motion fusion
abstract
Abstract One of the major problems of the motion capture‐based computer animation technique is the relatively high cost of equipment and low reuse rate of data. To overcome this problem, many motion‐editing methods have been developed. However, most of them can only handle one character whose motions are preset, and hence cannot interact with its environment automatically. In this paper, we construct a new architecture of multiple animated character motion fusion, which not only enables the characters to perceive and respond to the virtual environment, but also allows them to interact with each other. We will also discuss in detail the key issues, such as motion planning, coordination of multiple animated characters and generation of vivid continuous motions. Our experimental results will further testify to the effectiveness of the new methodology. Copyright © 2002 John Wiley & Sons, Ltd.
Zhongxiang Luo, Yueting Zhuang, Feng Liu 0015, Yunhe Pan
Comput. Animat. Virtual Worlds3