EDBT 2026 Demo / reviewers in the wild / expert
Shih-En Wei
dblp:119/7860
· DBLP profile ↗
25ranked-venue papers
5as first author
15since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 17 · 2 first-author · 11 since 2021Artificial intelligence and machine learning · 11 · 1 first-author · 8 since 2021Computer networks · 3 · 2 first-authorSystems, architecture and hardware · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | REWIND: Real-Time Egocentric Whole-Body Motion Diffusion with Exemplar-Based Identity ConditioningabstractWe present REWIND (Real-Time Egocentric Whole-Body Motion Diffusion), a one-step diffusion model for real-time, high-fidelity human motion estimation from egocentric image inputs. While an existing method for egocentric whole-body (i.e., body and hands) motion estimation is non-real-time and acausal due to diffusion-based iterative motion refinement to capture correlations between body and hand poses, REWIND operates in a fully causal and real-time manner. To enable real-time inference, we introduce (1) cascaded body-hand denoising diffusion, which effectively models the correlation between egocentric body and hand motions in a fast, feed-forward manner, and (2) diffusion distillation, which enables high-quality motion estimation with a single denoising step. Our denoising diffusion model is based on a modified Transformer architecture, designed to causally model output motions while enhancing generalizability to unseen motion lengths. Additionally, REWIND optionally supports identity-conditioned motion estimation when identity prior is available. To this end, we propose a novel identity conditioning method based on a small set of pose exemplars of the target identity, which further enhances motion estimation quality. Through extensive experiments, we demonstrate that REWIND significantly outperforms the existing baselines both with and without exemplar-based identity conditioning. Weipeng Xu, Alexander Richard, Shih-En Wei, Shunsuke Saito, Shaojie Bai, Te-Li Wang, Minhyuk Sung, Tae-Kyun Kim 0001, Jason M. Saragih |
CVPR | 4 |
| 2025 | Audio Driven Real-Time Facial Animation for Social TelepresenceabstractWe present an audio-driven real-time system for animating photorealistic 3D facial avatars with minimal latency, designed for social interactions in virtual reality for anyone. Central to our approach is an encoder model that transforms audio signals into latent facial expression sequences in real time, which are then decoded as photorealistic 3D facial avatars. Leveraging the generative capabilities of diffusion models, we capture the rich spectrum of facial expressions necessary for natural communication while achieving real-time performance (<15ms GPU time). Our novel architecture minimizes latency through two key innovations: an online transformer that eliminates dependency on future inputs and a distillation pipeline that accelerates iterative denoising into a single step. We further address critical design challenges in live scenarios for processing continuous audio signals frame-by-frame while maintaining consistent animation quality. The versatility of our framework extends to multimodal applications, including semantic modalities such as emotion conditions and multimodal sensors with head-mounted eye cameras on VR headsets. Experimental results demonstrate significant improvements in facial animation accuracy over existing offline state-of-the-art baselines, achieving 100 to 1000 × faster inference speed. We validate our approach through live VR demonstrations and across various scenarios such as multilingual speeches. Jiye Lee 0001, Chenghui Li, Shih-En Wei, Jason M. Saragih, Alexander Richard, Hanbyul Joo, Shaojie Bai |
SIGGRAPH Asia | 4 |
| 2025 | Generative Head-Mounted Camera Captures for Photorealistic AvatarsabstractEnabling photorealistic avatar animations in virtual and augmented reality (VR/AR) has been challenging because of the difficulty of obtaining ground truth state of faces. It is physically impossible to obtain synchronized images from head-mounted cameras (HMC) sensing input, which has partial observations in infrared (IR), and an array of outside-in dome cameras, which have full observations that match avatars' appearance. Prior works relying on analysis-by-synthesis methods could generate accurate ground truth, but suffer from imperfect disentanglement between expression and style in their personalized training. The reliance of extensive paired captures (HMC and dome) for the same subject makes it operationally expensive to collect large-scale datasets, which cannot be reused for different HMC viewpoints and lighting. In this work, we propose a novel generative approach, Generative HMC (GenHMC), that leverages large unpaired HMC captures , which are much easier to collect, to directly generate high-quality synthetic HMC images given any conditioning avatar state from dome captures. We show that our method is able to properly disentangle the input conditioning signal that specifies facial expression and viewpoint, from facial appearance, leading to more accurate ground truth. Furthermore, our method can generalize to unseen identities, removing the reliance on the paired captures. We demonstrate these breakthroughs by both evaluating synthetic HMC images and universal face encoders trained from these new HMC-avatar correspondences, which achieve better data efficiency and state-of-the-art accuracy. Shaojie Bai, Seunghyeon Seo, Chenghui Li, Owen Wang, Te-Li Wang, Tianyang Ma, Jason M. Saragih, Shih-En Wei, Nojun Kwak, Hyung Jun(John) Kim |
ACM Trans. Graph. | 9 |
| 2024 | Fast Registration of Photorealistic Avatars for VR Facial Animation
Chaitanya Patel, Shaojie Bai, Te-Li Wang, Jason M. Saragih, Shih-En Wei |
ECCV (62) | 5 |
| 2024 | Content-based Power-saving Design for Augmented Reality Applications on Mobile DevicesabstractWith the growing appeal of real-time interactions between physical views and virtual objects in augmented reality (AR) applications among contemporary users, optimizing power consumption is crucial for extending the battery life of mobile devices. This paper delves into methods for reducing power consumption specifically in mobile OLED devices when running AR applications, depending on user behaviors. Initially, we introduce a detection algorithm designed to accurately identify user status using cost-effective sensors. Subsequently, we present two dynamic configurations aimed at adjusting the display of physical views and virtual objects based on user visual attention of AR applications. The results of extensive experiments conducted on a commercial smartphone using an open-source AR application to assess the performance of the proposed power-saving design are highly promising. Ping-Han Chou, Shih-En Wei, Chun-Han Lin |
ISLPED | 2 |
| 2024 | Codec Avatar Studio: Paired Human Captures for Complete, Driveable, and Generalizable AvatarsabstractTo build photorealistic avatars that users can embody, human modelling must be complete (cover the full body), driveable (able to reproduce the current motion and appearance from the user), and generalizable (i.e., easily adaptable to novel identities).Towards these goals, paired captures, that is, captures of the same subject obtained from systems of diverse quality and availability, are crucial.However, paired captures are rarely available to researchers outside of dedicated industrial labs: Codec Avatar Studio is our proposal to close this gap.Towards generalization and driveability, we introduce a dataset of 256 subjects captured in two modalities: high resolution multi-view scans of their heads, and video from the internal cameras of a headset.Towards completeness, we introduce a dataset of 4 subjects captured in eight modalities: high quality relightable multi-view captures of heads and hands, full body multi-view captures with minimal and regular clothes, and corresponding head, hands and body phone captures.Together with our data, we also provide code and pre-trained models for different state-of-the-art human generation models.Our datasets and code are available at https://github.com/facebookresearch/ava-256 and https://github.com/facebookresearch/goliath. Julieta Martinez 0001, Emily Kim, Javier Romero 0002, Timur M. Bagautdinov, Shunsuke Saito, Shoou-I Yu, Michael Zollhöfer, Te-Li Wang, Shaojie Bai, Chenghui Li, Shih-En Wei, Rohan Joshi, Wyatt Borsos, Tomas Simon, Jason M. Saragih, Paul Theodosis, Alexander Greene, Anjani Josyula, Silvio Maeta, Andrew Jewett, Simion Venshtain, Christopher Heilman, Yueh-Tung Chen, Sidi Fu, Mohamed Elshaer, Tingfang Du, Longhua Wu, Shen-Chi Chen, Youssef Emad, Steven Longay, Ashley Brewer, Hitesh Shah, Taylor Koska, Kayla Haidle, Matthew Andromalos, Joanna Hsu, Thomas Dauer, Peter Selednik, Timothy Godisart, Scott Ardisson, Matthew Cipperly, Ben Humberston, Lon Farr, Bob Hansen, Peihong Guo, Dave Braun, Steven Krenn, He Wen 0001, Lucas Evans, Natalia Fadeeva, Matthew Stewart, Gabriel Schwartz, Divam Gupta, Gyeongsik Moon, Takaaki Shiratori, Fabian Prada, Bernardo Pires, Julia Buffalini, Autumn Trimble, Kevyn McPhail, Melissa Schoeller, Yaser Sheikh |
NeurIPS | 12 |
| 2024 | Universal Facial Encoding of Codec Avatars from VR HeadsetsabstractFaithful real-time facial animation is essential for avatar-mediated telepresence in Virtual Reality (VR). To emulate authentic communication, avatar animation needs to be efficient and accurate: able to capture both extreme and subtle expressions within a few milliseconds to sustain the rhythm of natural conversations. The oblique and incomplete views of the face, variability in the donning of headsets, and illumination variation due to the environment are some of the unique challenges in generalization to unseen faces. In this paper, we present a method that can animate a photorealistic avatar in realtime from head-mounted cameras (HMCs) on a consumer VR headset. We present a self-supervised learning approach, based on a cross-view reconstruction objective, that enables generalization to unseen users. We present a lightweight expression calibration mechanism that increases accuracy with minimal additional cost to run-time efficiency. We present an improved parameterization for precise ground-truth generation that provides robustness to environmental variation. The resulting system produces accurate facial animation for unseen users wearing VR headsets in realtime. We compare our approach to prior face-encoding methods demonstrating significant improvements in both quantitative metrics and qualitative results. Shaojie Bai, Te-Li Wang, Chenghui Li, Akshay Venkatesh, Tomas Simon, Chen Cao 0001, Gabriel Schwartz, Jason M. Saragih, Yaser Sheikh, Shih-En Wei |
ACM Trans. Graph. | 10 |
| 2022 | Robust Egocentric Photo-realistic Facial Expression Transfer for Virtual RealityabstractSocial presence, the feeling of being there with a “real” person, will fuel the next generation of communication systems driven by digital humans in virtual reality (VR). The best 3D video-realistic VR avatars that minimize the uncanny effect rely on person-specific (PS) models. However, these PS models are time-consuming to build and are typically trained with limited data variability, which results in poor generalization and robustness. Major sources of variability that affects the accuracy of facial expression transfer algorithms include using different VR headsets (e.g., camera configuration, slop of the headset), facial appearance changes over time (e.g., beard, make-up), and environmental factors (e.g., lighting, backgrounds). This is a major drawback for the scalability of these models in VR. This paper makes progress in overcoming these limitations by proposing an end-to-end multi-identity architecture (MIA) trained with specialized augmentation strategies. MIA drives the shape component of the avatar from three cameras in the VR headset (two eyes, one mouth), in untrained subjects, using minimal personalized information (i.e., neutral 3D mesh shape). Similarly, if the PS texture decoder is available, MIA is able to drive the full avatar (shape + texture) robustly outperforming PS models in challenging scenarios. Our key contribution to improve robustness and generalization, is that our method implicitly decouples, in an unsupervised manner, the facial expression from nuisance factors (e.g., headset, environment, facial appearance). We demonstrate the superior performance and robustness of the proposed method versus state-of-the-art PS approaches in a variety of experiments. Amin Jourabloo, Fernando De la Torre, Jason M. Saragih, Shih-En Wei, Stephen Lombardi, Te-Li Wang, Danielle Belko, Autumn Trimble, Hernán Badino |
CVPR | 4 |
| 2022 | LiP-Flow: Learning Inference-Time Priors for Codec Avatars via Normalizing Flows in Latent Space
Emre Aksan, Shugao Ma, Akin Caliskan, Stanislav Pidhorskyi, Alexander Richard, Shih-En Wei, Jason M. Saragih, Otmar Hilliges |
ECCV (26) | 6 |
| 2022 | Authentic volumetric avatars from a phone scanabstractCreating photorealistic avatars of existing people currently requires extensive person-specific data capture, which is usually only accessible to the VFX industry and not the general public. Our work aims to address this drawback by relying only on a short mobile phone capture to obtain a drivable 3D head avatar that matches a person's likeness faithfully. In contrast to existing approaches, our architecture avoids the complex task of directly modeling the entire manifold of human appearance, aiming instead to generate an avatar model that can be specialized to novel identities using only small amounts of data. The model dispenses with low-dimensional latent spaces that are commonly employed for hallucinating novel identities, and instead, uses a conditional representation that can extract person-specific information at multiple scales from a high resolution registered neutral phone scan. We achieve high quality results through the use of a novel universal avatar prior that has been trained on high resolution multi-view video captures of facial performances of hundreds of human subjects. By fine-tuning the model using inverse rendering we achieve increased realism and personalize its range of motion. The output of our approach is not only a high-fidelity 3D head avatar that matches the person's facial shape and appearance, but one that can also be driven using a jointly discovered shared global expression space with disentangled controls for gaze direction. Via a series of experiments we demonstrate that our avatars are faithful representations of the subject's likeness. Compared to other state-of-the-art methods for lightweight avatar creation, our approach exhibits superior visual quality and animateability. Chen Cao 0001, Tomas Simon, Jin Kyu Kim, Gabe Schwartz, Michael Zollhöfer, Shunsuke Saito, Stephen Lombardi, Shih-En Wei, Danielle Belko, Shoou-I Yu, Yaser Sheikh, Jason M. Saragih |
ACM Trans. Graph. | 8 |
| 2021 | SimPoE: Simulated Character Control for 3D Human Pose EstimationabstractAccurate estimation of 3D human motion from monocular video requires modeling both kinematics (body motion without physical forces) and dynamics (motion with physical forces). To demonstrate this, we present SimPoE, a Simulation-based approach for 3D human Pose Estimation, which integrates image-based kinematic inference and physics-based dynamics modeling. SimPoE learns a policy that takes as input the current-frame pose estimate and the next image frame to control a physically-simulated character to output the next-frame pose estimate. The policy contains a learnable kinematic pose refinement unit that uses 2D keypoints to iteratively refine its kinematic pose estimate of the next frame. Based on this refined kinematic pose, the policy learns to compute dynamics-based control (e.g., joint torques) of the character to advance the current-frame pose estimate to the pose estimate of the next frame. This design couples the kinematic pose refinement unit with the dynamics-based control generation unit, which are learned jointly with reinforcement learning to achieve accurate and physically-plausible pose estimation. Furthermore, we propose a meta-control mechanism that dynamically adjusts the character’s dynamics parameters based on the character state to attain more accurate pose estimates. Experiments on large-scale motion datasets demonstrate that our approach establishes the new state of the art in pose accuracy while ensuring physical plausibility. Ye Yuan 0007, Shih-En Wei, Tomas Simon, Kris Makoto Kitani, Jason M. Saragih |
CVPR | 2 |
| 2021 | OpenPose: Realtime Multi-Person 2D Pose Estimation Using Part Affinity FieldsabstractRealtime multi-person 2D pose estimation is a key component in enabling machines to have an understanding of people in images and videos. In this work, we present a realtime approach to detect the 2D pose of multiple people in an image. The proposed method uses a nonparametric representation, which we refer to as Part Affinity Fields (PAFs), to learn to associate body parts with individuals in the image. This bottom-up system achieves high accuracy and realtime performance, regardless of the number of people in the image. In previous work, PAFs and body part location estimation were refined simultaneously across training stages. We demonstrate that a PAF-only refinement rather than both PAF and body part location refinement results in a substantial increase in both runtime performance and accuracy. We also present the first combined body and foot keypoint detector, based on an internal annotated foot dataset that we have publicly released. We show that the combined detector not only reduces the inference time compared to running them sequentially, but also maintains the accuracy of each component individually. This work has culminated in the release of OpenPose, the first open-source realtime system for multi-person 2D pose detection, including body, foot, hand, and facial keypoints. Zhe Cao 0003, Gines Hidalgo, Tomas Simon, Shih-En Wei, Yaser Sheikh |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2021 | Supervision by Registration and Triangulation for Landmark DetectionabstractWe present supervision by registration and triangulation (SRT), an unsupervised approach that utilizes unlabeled multi-view video to improve the accuracy and precision of landmark detectors. Being able to utilize unlabeled data enables our detectors to learn from massive amounts of unlabeled data freely available and not be limited by the quality and quantity of manual human annotations. To utilize unlabeled data, there are two key observations: (I) The detections of the same landmark in adjacent frames should be coherent with registration, i.e., optical flow. (II) The detections of the same landmark in multiple synchronized and geometrically calibrated views should correspond to a single 3D point, i.e., multi-view consistency. Registration and multi-view consistency are sources of supervision that do not require manual labeling, thus it can be leveraged to augment existing training data during detector training. End-to-end training is made possible by differentiable registration and 3D triangulation modules. Experiments with 11 datasets and a newly proposed metric to measure precision demonstrate accuracy and precision improvements in landmark detection on both images and video. Xuanyi Dong, Yi Yang 0001, Shih-En Wei, Xinshuo Weng, Yaser Sheikh, Shoou-I Yu |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2021 | Driving-signal aware full-body avatarsabstractWe present a learning-based method for building driving-signal aware full-body avatars. Our model is a conditional variational autoencoder that can be animated with incomplete driving signals, such as human pose and facial keypoints, and produces a high-quality representation of human geometry and view-dependent appearance. The core intuition behind our method is that better drivability and generalization can be achieved by disentangling the driving signals and remaining generative factors, which are not available during animation. To this end, we explicitly account for information deficiency in the driving signal by introducing a latent space that exclusively captures the remaining information, thus enabling the imputation of the missing factors required during full-body animation, while remaining faithful to the driving signal. We also propose a learnable localized compression for the driving signal which promotes better generalization, and helps minimize the influence of global chance-correlations often found in real datasets. For a given driving signal, the resulting variational model produces a compact space of uncertainty for missing factors that allows for an imputation strategy best suited to a particular application. We demonstrate the efficacy of our approach on the challenging problem of full-body animation for virtual telepresence with driving signals acquired from minimal sensors placed in the environment and mounted on a VR-headset. Timur M. Bagautdinov, Chenglei Wu, Tomas Simon, Fabian Prada, Takaaki Shiratori, Shih-En Wei, Weipeng Xu, Yaser Sheikh, Jason M. Saragih |
ACM Trans. Graph. | 6 |
| 2021 | Deep relightable appearance models for animatable facesabstractWe present a method for building high-fidelity animatable 3D face models that can be posed and rendered with novel lighting environments in real-time. Our main insight is that relightable models trained to produce an image lit from a single light direction can generalize to natural illumination conditions but are computationally expensive to render. On the other hand, efficient, high-fidelity face models trained with point-light data do not generalize to novel lighting conditions. We leverage the strengths of each of these two approaches. We first train an expensive but generalizable model on point-light illuminations, and use it to generate a training set of high-quality synthetic face images under natural illumination conditions. We then train an efficient model on this augmented dataset, reducing the generalization ability requirements. As the efficacy of this approach hinges on the quality of the synthetic data we can generate, we present a study of lighting pattern combinations for dynamic captures and evaluate their suitability for learning generalizable relightable models. Towards achieving the best possible quality, we present a novel approach for generating dynamic relightable faces that exceeds state-of-the-art performance. Our method is capable of capturing subtle lighting effects and can even generate compelling near-field relighting despite being trained exclusively with far-field lighting data. Finally, we motivate the utility of our model by animating it with images captured from VR-headset mounted cameras, demonstrating the first system for face-driven interactions in VR that uses a photorealistic relightable face model. Sai Bi, Stephen Lombardi, Shunsuke Saito, Tomas Simon, Shih-En Wei, Kevyn McPhail, Ravi Ramamoorthi, Yaser Sheikh, Jason M. Saragih |
ACM Trans. Graph. | 5 |
| 2020 | The eyes have it: an integrated eye and face model for photorealistic facial animationabstractInteracting with people across large distances is important for remote work, interpersonal relationships, and entertainment. While such face-to-face interactions can be achieved using 2D video conferencing or, more recently, virtual reality (VR), telepresence systems currently distort the communication of eye contact and social gaze signals. Although methods have been proposed to redirect gaze in 2D teleconferencing situations to enable eye contact, 2D video conferencing lacks the 3D immersion of real life. To address these problems, we develop a system for face-to-face interaction in VR that focuses on reproducing photorealistic gaze and eye contact. To do this, we create a 3D virtual avatar model that can be animated by cameras mounted on a VR headset to accurately track and reproduce human gaze in VR. Our primary contributions in this work are a jointly-learnable 3D face and eyeball model that better represents gaze direction and upper facial expressions, a method for disentangling the gaze of the left and right eyes from each other and the rest of the face allowing the model to represent entirely unseen combinations of gaze and expression, and a gaze-aware model for precise animation from headset-mounted cameras. Our quantitative experiments show that our method results in higher reconstruction quality, and qualitative results show our method gives a greatly improved sense of presence for VR avatars. Gabriel Schwartz, Shih-En Wei, Te-Li Wang, Stephen Lombardi, Tomas Simon, Jason M. Saragih, Yaser Sheikh |
ACM Trans. Graph. | 2 |
| 2019 | VR facial animation via multiview image translationabstractA key promise of Virtual Reality (VR) is the possibility of remote social interaction that is more immersive than any prior telecommunication media. However, existing social VR experiences are mediated by inauthentic digital representations of the user (i.e., stylized avatars). These stylized representations have limited the adoption of social VR applications in precisely those cases where immersion is most necessary (e.g., professional interactions and intimate conversations). In this work, we present a bidirectional system that can animate avatar heads of both users' full likeness using consumer-friendly headset mounted cameras (HMC). There are two main challenges in doing this: unaccommodating camera views and the image-to-avatar domain gap. We address both challenges by leveraging constraints imposed by multiview geometry to establish precise image-to-avatar correspondence, which are then used to learn an end-to-end model for real-time tracking. We present designs for a training HMC, aimed at data-collection and model building, and a tracking HMC for use during interactions in VR. Correspondence between the avatar and the HMC-acquired images are automatically found through self-supervised multiview image translation, which does not require manual annotation or one-to-one correspondence between domains. We evaluate the system on a variety of users and demonstrate significant improvements over prior work. Shih-En Wei, Jason M. Saragih, Tomas Simon, Adam W. Harley, Stephen Lombardi, Michal Perdoch, Alexander Hypes, Hernán Badino, Yaser Sheikh |
ACM Trans. Graph. | 1 |
| 2018 | Supervision-by-Registration: An Unsupervised Approach to Improve the Precision of Facial Landmark DetectorsabstractIn this paper, we present supervision-by-registration, an unsupervised approach to improve the precision of facial landmark detectors on both images and video. Our key observation is that the detections of the same landmark in adjacent frames should be coherent with registration, i.e., optical flow. Interestingly, coherency of optical flow is a source of supervision that does not require manual labeling, and can be leveraged during detector training. For example, we can enforce in the training loss function that a detected landmark at framet-1followed by optical flow tracking from framet-1to frametshould coincide with the location of the detection at framet. Essentially, supervision-by-registration augments the training loss function with a registration loss, thus training the detector to have output that is not only close to the annotations in labeled images, but also consistent with registration on large amounts of unlabeled videos. End-to-end training with the registration loss is made possible by a differentiable Lucas-Kanade operation, which computes optical flow registration in the forward pass, and back-propagates gradients that encourage temporal coherency in the detector. The output of our method is a more precise image-based facial landmark detector, which can be applied to single images or video. With supervision-by-registration, we demonstrate (1) improvements in facial landmark detection on both images (300W, ALFW) and video (300VW, Youtube-Celebrities), and (2) significant reduction of jittering in video detections. Xuanyi Dong, Shoou-I Yu, Xinshuo Weng, Shih-En Wei, Yi Yang 0001, Yaser Sheikh |
CVPR | 4 |
| 2017 | Realtime Multi-person 2D Pose Estimation Using Part Affinity FieldsabstractWe present an approach to efficiently detect the 2D pose of multiple people in an image. The approach uses a nonparametric representation, which we refer to as Part Affinity Fields (PAFs), to learn to associate body parts with individuals in the image. The architecture encodes global context, allowing a greedy bottom-up parsing step that maintains high accuracy while achieving realtime performance, irrespective of the number of people in the image. The architecture is designed to jointly learn part locations and their association via two branches of the same sequential prediction process. Our method placed first in the inaugural COCO 2016 keypoints challenge, and significantly exceeds the previous state-of-the-art result on the MPII Multi-Person benchmark, both in performance and efficiency. Zhe Cao 0003, Tomas Simon, Shih-En Wei, Yaser Sheikh |
CVPR | 3 |
| 2016 | Convolutional Pose MachinesabstractPose Machines provide a sequential prediction framework for learning rich implicit spatial models. In this work we show a systematic design for how convolutional networks can be incorporated into the pose machine framework for learning image features and image-dependent spatial models for the task of pose estimation. The contribution of this paper is to implicitly model long-range dependencies between variables in structured prediction tasks such as articulated pose estimation. We achieve this by designing a sequential architecture composed of convolutional networks that directly operate on belief maps from previous stages, producing increasingly refined estimates for part locations, without the need for explicit graphical model-style inference. Our approach addresses the characteristic difficulty of vanishing gradients during training by providing a natural learning objective function that enforces intermediate supervision, thereby replenishing back-propagated gradients and conditioning the learning procedure. We demonstrate state-of-the-art performance and outperform competing methods on standard benchmarks including the MPII, LSP, and FLIC datasets. Shih-En Wei, Varun Ramakrishna, Takeo Kanade, Yaser Sheikh |
CVPR | 1 |
| 2015 | Robust Action Recognition via Borrowing Information Across Video ModalitiesabstractThe recent advances in imaging devices have opened the opportunity of better solving the tasks of video content analysis and understanding. Next-generation cameras, such as the depth or binocular cameras, capture diverse information, and complement the conventional 2D RGB cameras. Thus, investigating the yielded multimodal videos generally facilitates the accomplishment of related applications. However, the limitations of the emerging cameras, such as short effective distances, expensive costs, or long response time, degrade their applicability, and currently make these devices not online accessible in practical use. In this paper, we provide an alternative scenario to address this problem, and illustrate it with the task of recognizing human actions. In particular, we aim at improving the accuracy of action recognition in RGB videos with the aid of one additional RGB-D camera. Since RGB-D cameras, such as Kinect, are typically not applicable in a surveillance system due to its short effective distance, we instead offline collect a database, in which not only the RGB videos but also the depth maps and the skeleton data of actions are available jointly. The proposed approach can adapt the interdatabase variations, and activate the borrowing of visual knowledge across different video modalities. Each action to be recognized in RGB representation is then augmented with the borrowed depth and skeleton features. Our approach is comprehensively evaluated on five benchmark data sets of action recognition. The promising results manifest that the borrowed information leads to remarkable boost in recognition accuracy. Nick C. Tang, Yen-Yu Lin, Ju-Hsuan Hua, Shih-En Wei, Ming-Fang Weng, Hong-Yuan Mark Liao |
IEEE Trans. Image Process. | 4 |
| 2014 | Optimizing Small Cell Deployment in Arbitrary Wireless Networks with Minimum Service Rate ConstraintsabstractFemtocell technology has shifted beyond indoor residential applications to cover a wider range of scenarios including metropolitan and rural areas. The term “small cell” has hence been used to denote such low-power transmission points deployed for enhancing macrocell coverage and/or capacity. While deployment of femto BSs has typically followed the bottom-up paradigm driven by the ad hoc demand of users, more and more studies have prompted a move toward a more managed deployment model for better tradeoff between performance and cost. In this paper, we investigate an optimization problem for femtocell deployment in a dense network with arbitrary topology. The goal is to determine deployment locations and operation parameters of femtocells for maximizing the number of customers supported with QoS constraints. Since the formulated problem belongs to mixed-integer non-linear programming (MINLP), we propose an anytime algorithm that transforms the joint problem into a cluster formation sub-problem (involving location selection and cell coverage) and a resource management sub-problem (involving power control and resource allocation) for effectively solving all optimization variables in an iterative fashion. Compared with other approaches for femtocell deployment, our evaluation results show that the proposed algorithm can effectively solve the target problem while striking a better performance tradeoff between computation complexity and solution quality. Hung-Yun Hsieh, Shih-En Wei, Cheng-Pang Chien |
IEEE Trans. Mob. Comput. | 2 |
| 2012 | Joint optimization of cluster formation and power control for interference-limited machine-to-machine communicationsabstractClustered communication has been considered as one key technology for supporting machine-to-machine (M2M) wireless networks with a large number of communicating devices. Unlike related work that focuses on clustering with simple or no wireless interference model at the physical layer, in this paper we investigate the optimization problem of cluster formation and power control for interference-limited M2M communications. We consider a scenario where machines that form in clusters are allowed to reuse the spectrum occupied by human devices through proper transmission power control. To maximize the number of machines that can communicate while meeting the data rate constraints of human devices and machines themselves, we formulate a mixed-integer non-linear programming (MINLP) problem. Since the MINLP problem becomes too complex when the number of machines increases, we propose an algorithm that transforms the problem into a coalition structure generation sub-problem embedded with a linear power control sub-problem. The proposed algorithm is an anytime algorithm and hence the length of the running time can be arbitrarily controlled while yielding a feasible solution with the desired quality. Compared with other approaches for solving the original MINLP problem, we show through numerical results that the proposed algorithm can effectively solve the target problem and allow machines to achieve better spatial reuse with human devices in interference-limited M2M communications. Shih-En Wei, Hung-Yun Hsieh, Hsuan-Jung Su |
GLOBECOM | 1 |
| 2012 | Formulating and solving the femtocell deployment problem in two-tier heterogeneous networksabstractRecently, there has been an increasing interest in the deployment and management of femto base stations (BSs) to optimize the overall system performance in macro-femto heterogeneous networks. While deployment of femto BSs is typically not as planned as that of pico BSs, given a number of femto BSs to be distributed to candidate customer sites, questions regarding the optimal deployment locations and transmission configurations still need to be answered. In this paper, we formulate a joint optimization problem involving deployment location, cell selection, and power control to maximize the number of users that can be supported for a given number of femto BSs to be deployed in the macro cell. Since the formulated problem belongs to mixed-integer non-linear programming (MINLP), we propose an anytime algorithm that can yield a desirable solution within proper time limit. Specifically, based on the concept of coalition structure generation, the algorithm decouples the problem into the cluster formation sub-problem and power control sub-problem to find the optimal cluster head (femto BS location), cluster membership (cell selection), and transmission power in an iterative fashion. Evaluation results presented in this paper show that the proposed algorithm can effectively solve the problem with better complexity-optimality tradeoffs compared to baseline approaches. Shih-En Wei, Chih-Hua Chang, You-En Lin, Hung-Yun Hsieh, Hsuan-Jung Su |
ICC | 1 |
| 2012 | Enabling dense machine-to-machine communications through interference-controlled clusteringabstractClustering of machines for better spatial reuse has been considered as one key technology for supporting machine-to-machine (M2M) communications with a large number of communicating devices. Unlike related work that focuses on greedy clustering algorithms without interference control, in this paper we consider a scenario where machines through joint cluster formation and power control are allowed to opportunistically use the spectrum occupied by human devices for interference-limited M2M communications. To maximize the number of machines that can communicate without violating the QoS constraint of the human device, we formulate a mixed-integer non-linear programming (MINLP) problem to determine the optimal cluster structure and power control. We then propose an anytime algorithm based on simulated annealing to solve the MINLP problem under a high density of machines. Compared with the approach of directly solving the MINLP problem and the approach of separately performing cluster formation and power control, we show through numerical results that the proposed algorithm can effectively solve the target problem while striking a better performance tradeoff between complexity and optimality. Shih-En Wei, Hung-Yun Hsieh, Hsuan-Jung Su |
IWCMC | 1 |