EDBT 2026 Demo / reviewers in the wild / expert
Tony Tung
dblp:54/1097
· DBLP profile ↗
45ranked-venue papers
17as first author
16since 2021 · last 2025
0000-0002-0824-0960ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 37 · 14 first-author · 15 since 2021Artificial intelligence and machine learning · 30 · 10 first-author · 14 since 2021Systems, architecture and hardware · 3Computer networks · 2 · 1 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | DGH: Dynamic Gaussian HairabstractThe creation of photorealistic dynamic hair remains a major challenge in digital human modeling because of the complex motions, occlusions, and light scattering. Existing methods often resort to static capture and physics-based models that do not scale as they require manual parameter fine-tuning to handle the diversity of hairstyles and motions, and heavy computation to obtain high-quality appearance. In this paper, we present Dynamic Gaussian Hair (DGH), a novel framework that efficiently learns hair dynamics and appearance. We propose: (1) a coarse-to-fine model that learns temporally coherent hair motion dynamics across diverse hairstyles; (2) a strand-guided optimization module that learns a dynamic 3D Gaussian representation for hair appearance with support for differentiable rendering, enabling gradient-based learning of view-consistent appearance under motion. Unlike prior simulation-based pipelines, our approach is fully data-driven, scales with training data, and generalizes across various hairstyles and head motion sequences. Additionally, DGH can be seamlessly integrated into a 3D Gaussian avatar framework, enabling realistic, animatable hair for high-fidelity avatar representation. DGH achieves promising geometry and appearance results, providing a scalable, data-driven alternative to physics-based simulation and rendering. Yuanlu Xu, Edith Tretschk, Anastasia Ianina, Aljaz Bozic, Ulrich Neumann, Tony Tung |
NeurIPS | 8 |
| 2025 | Generating High-Fidelity Clothed Human Dynamics with Temporal DiffusionabstractClothed human modeling plays a crucial role in multimedia research, with applications spanning virtual reality, gaming, and fashion design. The goal is to learn clothed human dynamics from observations and then generate humans with high-fidelity clothing details for motion animation. Despite tremendous advancements in clothing shape analysis by existing approaches, the community still faces challenges in generating convincing visual effects of cloth dynamics, maintaining temporally smooth clothing details, and handling diverse clothing patterns. To address these challenges, we introduce ClothDiffuse, a temporal diffusion model that seamlessly integrates three key components into this task—temporal dynamics modeling, iterative refinement, and diversified generation. Our approach begins by using an encoder to extract high-level temporal features from input human body motions. These features are combined with a learnable pixel-aligned garment feature, serving as prior conditions for the shape decoder. The decoder then iteratively denoise Gaussian noise to produce clothing deformations over time on the input unclothed human bodies. To ensure that the results align with observations and adhere to physical plausibility for clothing shape inference, we propose two physics-inspired loss functions that preserve the intra-frame distances and inter-frame forces of clothing points. Additionally, the stochastic nature of the denoising process allows for the generation of diverse and plausible clothing shapes. Experiments show that our approach outperforms state-of-the-art methods in chamfer distance and visual effects, particularly for loose clothing such as dresses and skirts. Furthermore, our approach effectively adapts to out-of-domain clothing types and generates realistic clothes dynamics. Shihao Zou, Yuanlu Xu, Nikolaos Sarafianos, Federica Bogo, Tony Tung, Weixin Si, Li Cheng 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 5 |
| 2024 | HISR: Hybrid Implicit Surface Representation for Photorealistic 3D Human ReconstructionabstractNeural reconstruction and rendering strategies have demonstrated state-of-the-art performances due, in part, to their ability to preserve high level shape details. Existing approaches, however, either represent objects as implicit surface functions or neural volumes and still struggle to recover shapes with heterogeneous materials, in particular human skin, hair or clothes. To this aim, we present a new hybrid implicit surface representation to model human shapes. This representation is composed of two surface layers that represent opaque and translucent regions on the clothed human body. We segment different regions automatically using visual cues and learn to reconstruct two signed distance functions (SDFs). We perform surface-based rendering on opaque regions (e.g., body, face, clothes) to preserve high-fidelity surface normals and volume rendering on translucent regions (e.g., hair). Experiments demonstrate that our approach obtains state-of-the-art results on 3D human reconstructions, and also shows competitive performances on other objects. Angtian Wang, Yuanlu Xu, Nikolaos Sarafianos, Robert Maier 0001, Edmond Boyer, Alan L. Yuille, Tony Tung |
AAAI | 7 |
| 2024 | ANIM: Accurate Neural Implicit Model for Human Reconstruction from a Single RGB-D ImageabstractRecent progress in human shape learning, shows that neural implicit models are effective in generating 3D hu-man surfaces from limited number of views, and even from a single RGB image. However, existing monocular approaches still struggle to recover fine geometric details such as face, hands or cloth wrinkles. They are also easily prone to depth ambiguities that result in distorted geome-tries along the camera optical axis. In this paper, we ex-plore the benefits of incorporating depth observations in the reconstruction process by introducing ANIM, a novel method that reconstructs arbitrary 3D human shapes from single-view RGB-D images with an unprecedented level of accuracy. Our model learns geometric details from both multi-resolution pixel-aligned and voxel-aligned features to leverage depth information and enable spatial relation-ships, mitigating depth ambiguities. We further enhance the quality of the reconstructed shape by introducing a depth-supervision strategy, which improves the accuracy of the signed distance field estimation of points that lie on the re-constructed surface. Experiments demonstrate that ANIM outperforms state-of-the-art works that use RGB, surface normals, point cloud or RGB-D data as input. In addition, we introduce ANIM-Real, a new multi-modal dataset comprising highquality scans paired with consumer-grade RGB-D camera, and our protocol to fine-tune ANIM, enabling highquality reconstruction from real-world human capture. https://marcopesavento.github.io/Anim/ Marco Pesavento, Yuanlu Xu, Nikolaos Sarafianos, Robert Maier 0001, Chun-Han Yao, Marco Volino, Edmond Boyer, Adrian Hilton 0001, Tony Tung |
CVPR | 10 |
| 2024 | SplatFields: Neural Gaussian Splats for Sparse 3D and 4D Reconstruction
Marko Mihajlovic, Sergey Prokudin, Siyu Tang 0001, Robert Maier 0001, Federica Bogo, Tony Tung, Edmond Boyer |
ECCV (2) | 6 |
| 2023 | VIVE3D: Viewpoint-Independent Video Editing using 3D-Aware GANsabstractWe introduce VIVE3D, a novel approach that extends the capabilities of image-based 3D GANs to video editing and is able to represent the input video in an identity-preserving and temporally consistent way. We propose two new building blocks. First, we introduce a novel GAN inversion technique specifically tailored to 3D GANs by jointly embedding multiple frames and optimizing for the camera parameters. Second, besides traditional semantic face edits (e.g. for age and expression), we are the first to demonstrate edits that show novel views of the head enabled by the inherent prop-erties of 3D GANs and our optical flow-guided compositing technique to combine the head with the background video. Our experiments demonstrate that VIVE3D generates high-fidelity face edits at consistent quality from a range of camera viewpoints which are composited with the original video in a temporally and spatially consistent manner. Anna Frühstück, Nikolaos Sarafianos, Yuanlu Xu, Peter Wonka, Tony Tung |
CVPR | 5 |
| 2023 | Multi-View Reconstruction Using Signed Ray Distance Functions (SRDF)abstractIn this paper, we investigate a new optimization framework for multi-view 3D shape reconstructions. Recent differentiable rendering approaches have provided breakthrough performances with implicit shape representations though they can still lack precision in the estimated geometries. On the other hand multi-view stereo methods can yield pixel wise geometric accuracy with local depth predictions along viewing rays. Our approach bridges the gap between the two strategies with a novel volumetric shape representation that is implicit but parameterized with pixel depths to better materialize the shape surface with consistent signed distances along viewing rays. The approach retains pixel-accuracy while benefiting from volumetric integration in the optimization. To this aim, depths are optimized by evaluating, at each 3D location within the volumetric discretization, the agreement between the depth prediction consistency and the photometric consistency for the corresponding pixels. The optimization is agnostic to the associated photo-consistency term which can vary from a median-based baseline to more elaborate criteria, e.g. learned functions. Our experiments demonstrate the benefit of the volumetric integration with depth predictions. They also show that our approach outperforms existing approaches over standard 3D benchmarks with better geometry estimations. Pierre Zins, Yuanlu Xu, Edmond Boyer, Stefanie Wuhrer, Tony Tung |
CVPR | 5 |
| 2023 | NSF: Neural Surface Fields for Human Modeling from Monocular DepthabstractObtaining personalized 3D animatable avatars from a monocular camera has several real world applications in gaming, virtual try-on, animation, and VR/XR, etc. However, it is very challenging to model dynamic and finegrained clothing deformations from such sparse data. Existing methods for modeling 3D humans from depth data have limitations in terms of computational efficiency, mesh coherency, and flexibility in resolution and topology. For instance, reconstructing shapes using implicit functions and extracting explicit meshes per frame is computationally expensive and cannot ensure coherent meshes across frames. Moreover, predicting per-vertex deformations on a predesigned human template with a discrete surface lacks flexibility in resolution and topology. To overcome these limitations, we propose a novel method 'NSF : Neural Surface Fields’ for modeling 3D clothed humans from monocular depth. NSF defines a neural field solely on the base surface which models a continuous and flexible displacement field. NSF can be adapted to the base surface with different resolution and topology without retraining at inference time. Compared to existing approaches, our method eliminates the expensive per-frame surface extraction while maintaining mesh coherency, and is capable of reconstructing meshes with arbitrary resolution without retraining. To foster research in this direction, we release our code in project page at: https://yuxuan-xue.com/nsf. Yuxuan Xue 0001, Bharat Lal Bhatnagar, Riccardo Marin, Nikolaos Sarafianos, Yuanlu Xu, Gerard Pons-Moll, Tony Tung |
ICCV | 7 |
| 2022 | BodyMap: Learning Full-Body Dense Correspondence MapabstractDense correspondence between humans carries powerful semantic information that can be utilized to solve fundamental problems for full-body understanding such as in-the-wild surface matching, tracking and reconstruction. In this paper we present BodyMap, a new framework for obtaining high-definition full-body and continuous dense correspondence between in-the-wild images of clothed humans and the surface of a 3D template model. The correspondences cover fine details such as hands and hair, while capturing regions far from the body surface, such as loose clothing. Prior methods for estimating such dense surface correspondence i) cut a 3D body into parts which are unwrapped to a 2D UV space, producing discontinuities along part seams, or ii) use a single surface for representing the whole body, but none handled body details. Here, we introduce a novel network architecture with Vision Transformers that learn fine-level features on a continuous body surface. BodyMap outperforms prior work on various metrics and datasets, including DensePose-COCO by a large margin. Furthermore, we show various applications ranging from multi-layer dense cloth correspondence, neural rendering with novel-view synthesis and appearance swapping. Anastasia Ianina, Nikolaos Sarafianos, Yuanlu Xu, Ignacio Rocco, Tony Tung |
CVPR | 5 |
| 2022 | SPAMs: Structured Implicit Parametric ModelsabstractParametric 3D models have formed a fundamental role in modeling deformable objects, such as human bodies, faces, and hands; however, the construction of such parametric models requires significant manual intervention and domain expertise. Recently, neural implicit 3D representations have shown great expressibility in capturing 3D shape geometry. We observe that deformable object motion is often semantically structured, and thus propose to learn Structured-implicit PArametric Models (SPAMs) as a deformable object representation that structurally decomposes non-rigid object motion into part-based disentangled representations of shape and pose, with each being represented by deep implicit functions. This enables a structured characterization of object movement, with part decomposition characterizing a lower-dimensional space in which we can establish coarse motion correspondence. In particular, we can leverage the part decompositions at test time to fit to new depth sequences of unobserved shapes, by establishing part correspondences between the input observation and our learned part spaces; this guides a robust joint optimization between the shape and pose of all parts, even under dramatic motion sequences. Experiments demonstrate that our part-aware shape and pose understanding lead to state-of-the-art performance in reconstruction and tracking of depth sequences of complex deforming object motion. Pablo R. Palafox, Nikolaos Sarafianos, Tony Tung, Angela Dai |
CVPR | 3 |
| 2022 | Free-Viewpoint RGB-D Human Performance Capture and Rendering
Phong Nguyen 0001, Nikolaos Sarafianos, Christoph Lassner, Janne Heikkilä, Tony Tung |
ECCV (16) | 5 |
| 2022 | Pose-NDF: Modeling Human Pose Manifolds with Neural Distance Fields
Garvita Tiwari, Dimitrije Antic, Jan Eric Lenssen, Nikolaos Sarafianos, Tony Tung, Gerard Pons-Moll |
ECCV (5) | 5 |
| 2021 | Data-Driven 3D Reconstruction of Dressed Humans From Sparse ViewsabstractRecently, data-driven single-view reconstruction methods have shown great progress in modeling 3D dressed humans. However, such methods suffer heavily from depth ambiguities and occlusions inherent to single view inputs. In this paper, we tackle this problem by considering a small set of input views and investigate the best strategy to suitably exploit information from these views. We propose a data-driven end-to-end approach that reconstructs an implicit 3D representation of dressed humans from sparse camera views. Specifically, we introduce three key components: first a spatially consistent reconstruction that allows for arbitrary placement of the person in the input views using a perspective camera model; second an attention-based fusion layer that learns to aggregate visual information from several viewpoints; and third a mechanism that encodes local 3D patterns under the multi-view context. In the experiments, we show the proposed approach outperforms the state of the art on standard data both quantitatively and qualitatively. To demonstrate the spatially consistent reconstruction, we apply our approach to dynamic scenes. Additionally, we apply our method on real data acquired with a multi-camera platform and demonstrate our approach can obtain results comparable to multi-view stereo with dramatically less views. Pierre Zins, Yuanlu Xu, Edmond Boyer, Stefanie Wuhrer, Tony Tung |
3DV | 5 |
| 2021 | Semi-Supervised Synthesis of High-Resolution Editable Textures for 3D HumansabstractWe introduce a novel approach to generate diverse high fidelity texture maps for 3D human meshes in a semi-supervised setup. Given a segmentation mask defining the layout of the semantic regions in the texture map, our network generates high-resolution textures with a variety of styles, that are then used for rendering purposes. To accomplish this task, we propose a Region-adaptive Adversarial Variational AutoEncoder (ReAVAE) that learns the probability distribution of the style of each region individually so that the style of the generated texture can be controlled by sampling from the region-specific distributions. In addition, we introduce a data generation technique to augment our training set with data lifted from single-view RGB inputs. Our training strategy allows the mixing of reference image styles with arbitrary styles for different regions, a property which can be valuable for virtual try-on AR/VR applications. Experimental results show that our method synthesizes better texture maps compared to prior work while enabling independent layout and style controllability. Bindita Chaudhuri, Nikolaos Sarafianos, Linda G. Shapiro, Tony Tung |
CVPR | 4 |
| 2021 | ARCH++: Animation-Ready Clothed Human Reconstruction RevisitedabstractWe present ARCH++, an image-based method to reconstruct 3D avatars with arbitrary clothing styles. Our reconstructed avatars are animation-ready and highly realistic, in both the visible regions from input views and the unseen regions. While prior work shows great promise of reconstructing animatable clothed humans with various topologies, we observe that there exist fundamental limitations resulting in sub-optimal reconstruction quality. In this paper, we revisit the major steps of image-based avatar reconstruction and address the limitations with ARCH++. First, we introduce an end-to-end point based geometry encoder to better describe the semantics of the underlying 3D human body, in replacement of previous hand-crafted features. Second, in order to address the occupancy ambiguity caused by topological changes of clothed humans in the canonical pose, we propose a co-supervising framework with cross-space consistency to jointly estimate the occupancy in both the posed and canonical spaces. Last, we use image-to-image translation networks to further refine detailed geometry and texture on the reconstructed surface, which improves the fidelity and consistency across arbitrary viewpoints. In the experiments, we demonstrate improvements over the state of the art on both public benchmarks and user studies in reconstruction quality and realism. Tong He 0002, Yuanlu Xu, Shunsuke Saito, Stefano Soatto, Tony Tung |
ICCV | 5 |
| 2021 | Neural-GIF: Neural Generalized Implicit Functions for Animating People in ClothingabstractWe present Neural Generalized Implicit Functions (Neural-GIF), to animate people in clothing as a function of the body pose. Given a sequence of scans of a subject in various poses, we learn to animate the character for new poses. Existing methods have relied on template-based representations of the human body (or clothing). However such models usually have fixed and limited resolutions, require difficult data pre-processing steps and cannot be used with complex clothing. We draw inspiration from template-based methods, which factorize motion into articulation and nonrigid deformation, but generalize this concept for implicit shape learning to obtain a more flexible model. We learn to map every point in the space to a canonical space, where a learned deformation field is applied to model non-rigid effects, before evaluating the signed distance field. Our formulation allows the learning of complex and non-rigid deformations of clothing and soft tissue, without computing a template registration as it is common with current approaches. Neural-GIF can be trained on raw 3D scans and reconstructs detailed complex surface geometry and deformations. Moreover, the model can generalize to new poses. We evaluate our method on a variety of characters from different public datasets in diverse clothing styles and show significant improvements over baseline methods, quantitatively and qualitatively. We also extend our model to multiple shape setting. To stimulate further research, we will make the model, code and data publicly available at [1]. Garvita Tiwari, Nikolaos Sarafianos, Tony Tung, Gerard Pons-Moll |
ICCV | 3 |
| 2020 | ARCH: Animatable Reconstruction of Clothed HumansabstractIn this paper, we propose ARCH (Animatable Reconstruction of Clothed Humans), a novel end-to-end framework for accurate reconstruction of animation-ready 3D clothed humans from a monocular image. Existing approaches to digitize 3D humans struggle to handle pose variations and recover details. Also, they do not produce models that are animation ready. In contrast, ARCH is a learned pose-aware model that produces detailed 3D rigged full-body human avatars from a single unconstrained RGB image. A Semantic Space and a Semantic Deformation Field are created using a parametric 3D body estimator. They allow the transformation of 2D/3D clothed humans into a canonical space, reducing ambiguities in geometry caused by pose variations and occlusions in training data. Detailed surface geometry and appearance are learned using an implicit function representation with spatial local features. Furthermore, we propose additional per-pixel supervision on the 3D reconstruction using opacity-aware differentiable rendering. Our experiments indicate that ARCH increases the fidelity of the reconstructed humans. We obtain more than 50% lower reconstruction errors for standard metrics compared to state-of-the-art methods on public datasets. We also show numerous qualitative examples of animated, high-quality reconstructed avatars unseen in the literature so far. Zeng Huang, Yuanlu Xu, Christoph Lassner, Hao Li 0015, Tony Tung |
CVPR | 5 |
| 2020 | SIZER: A Dataset and Model for Parsing 3D Clothing and Learning Size Sensitive 3D Clothing
Garvita Tiwari, Bharat Lal Bhatnagar, Tony Tung, Gerard Pons-Moll |
ECCV (3) | 3 |
| 2020 | TexMesh: Reconstructing Detailed Human Texture and Geometry from RGB-D Video
Tiancheng Zhi, Christoph Lassner, Tony Tung, Carsten Stoll, Srinivasa G. Narasimhan, Minh Vo |
ECCV (10) | 3 |
| 2019 | DenseRaC: Joint 3D Pose and Shape Estimation by Dense Render-and-CompareabstractWe present DenseRaC, a novel end-to-end framework for jointly estimating 3D human pose and body shape from a monocular RGB image. Our two-step framework takes the body pixel-to-surface correspondence map (i.e., IUV map) as proxy representation and then performs estimation of parameterized human pose and shape. Specifically, given an estimated IUV map, we develop a deep neural network optimizing 3D body reconstruction losses and further integrating a render-and-compare scheme to minimize differences between the input and the rendered output, i.e., dense body landmarks, body part masks, and adversarial priors. To boost learning, we further construct a large-scale synthetic dataset (MOCA) utilizing web-crawled Mocap sequences, 3D scans and animations. The generated data covers diversified camera views, human actions and body shapes, and is paired with full ground truth. Our model jointly learns to represent the 3D human body from hybrid datasets, mitigating the problem of unpaired training data. Our experiments show that DenseRaC obtains superior performance against state of the art on public benchmarks of various human-related tasks. Yuanlu Xu, Song-Chun Zhu, Tony Tung |
ICCV | 3 |
| 2018 | DeepWrinkles: Accurate and Realistic Clothing Modeling
Zorah Lähner, Daniel Cremers, Tony Tung |
ECCV (4) | 3 |
| 2017 | Pano2CAD: Room Layout from a Single Panorama ImageabstractThis paper presents a method of estimating the geometry of a room and the 3D pose of objects from a single 360 panorama image. Assuming ManhattanWorld geometry, we formulate the task as an inference problem in which we estimate positions and orientations of walls and objects. The method combines surface normal estimation, 2D object detection and 3D object pose estimation. Quantitative results are presented on a dataset of synthetically generated 3D rooms containing objects, as well as on a subset of handlabeled images from the public SUN360 dataset. Jiu Xu, Björn Stenger, Tommi Kerola, Tony Tung |
WACV | 4 |
| 2015 | Invariant shape descriptor for 3D video encoding
Tony Tung, Takashi Matsuyama |
Vis. Comput. | 1 |
| 2014 | A 3D Shape Descriptor for Segmentation of Unstructured Meshes into Segment-Wise Coherent Mesh SeriesabstractThis paper presents a novel shape descriptor for topology-based segmentation of 3D video sequence. 3D video is a series of 3D meshes without temporal correspondences which benefit for applications including compression, motion analysis, and kinematic editing. In 3D video, both 3D mesh connectivities and the global surface topology can change frame by frame. This characteristic prevents from making accurate temporal correspondences through the entire 3D mesh series. To overcome this difficulty, we propose a two-step strategy which decomposes the entire sequence into a series of topologically coherent segments using our new shape descriptor, and then estimates temporal correspondences on a per-segment basis. We demonstrate the robustness and accuracy of the shape descriptor on real data which consist of large non-rigid motion and reconstruction errors. Tomoyuki Mukasa, Shohei Nobuhara, Tony Tung, Takashi Matsuyama |
3DV | 3 |
| 2014 | Timing-Based Local Descriptor for Dynamic SurfacesabstractIn this paper, we present the first local descriptor designed for dynamic surfaces. A dynamic surface is a surface that can undergo non-rigid deformation (e.g., human body surface). Using state-of-the-art technology, details on dynamic surfaces such as cloth wrinkle or facial expression can be accurately reconstructed. Hence, various results (e.g., surface rigidity, or elasticity) could be derived by microscopic categorization of surface elements. We propose a timing-based descriptor to model local spatiotemporal variations of surface intrinsic properties. The low-level descriptor encodes gaps between local event dynamics of neighboring keypoints using timing structure of linear dynamical systems (LDS). We also introduce the bag-of-timings (BoT) paradigm for surface dynamics characterization. Experiments are performed on synthesized and real-world datasets. We show the proposed descriptor can be used for challenging dynamic surface classification and segmentation with respect to rigidity at surface keypoints. Tony Tung, Takashi Matsuyama |
CVPR | 1 |
| 2014 | On Mean Pose and Variability of 3D Deformable Models
Benjamin Allain, Jean-Sébastien Franco, Edmond Boyer, Tony Tung |
ECCV (2) | 4 |
| 2014 | Geodesic Mapping for Dynamic Surface AlignmentabstractThis paper presents a novel approach that achieves dynamic surface alignment by geodesing mapping. The surfaces are 3D manifold meshes representing non-rigid objects in motion (e.g., humans) which can be obtained by multiview stereo reconstruction. The proposed framework consists of a geodesic mapping (i.e., geodesic diffeomorphism) between surfaces which carry a distance function (namely the global geodesic distance), and a geodesic-based coordinate system (namely the global geodesic coordinates) defined similarly to generalized barycentric coordinates. The coordinates are used to recursively choose correspondence points in non-ambiguous regions using a coarse-to-fine strategy to reliably locate all surface points and define a discrete mapping. Complete point-to-point surface alignment with smooth mapping is then derived by optimizing a piecewise objective function within a probabilistic framework. The proposed technique only relies on surface intrinsic geometrical properties, and does not require prior knowledge on surface appearance (e.g., color or texture), shape (e.g., topology) or parameterization (e.g., mesh connectivity or complexity). The method can be used for numerous applications, such as visual information (e.g., texture) transfer between surface models representing different objects, dense motion flow estimation of 3D dynamic surfaces, wide-timeframe matching, etc. Experiments show compelling results on challenging publicly available real-world datasets. Tony Tung, Takashi Matsuyama |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2014 | Multiparty Interaction Understanding Using Smart Multimodal Digital SignageabstractThis paper presents a novel multimodal system designed for multi-party human-human interaction analysis. The design of human-machine interfaces for multiple users is challenging because simultaneous processing of actions and reactions have to be consistent. The proposed system consists of a large display equipped with multiple sensing devices: microphone array, HD video cameras, and depth sensors. Multiple users positioned in front of the panel freely interact using voice or gesture while looking at the displayed content, without wearing any particular devices (such as motion capture sensors or head mounted devices). Acoustic and visual information is captured and processed jointly using established and state-of-the-art techniques to obtain individual speech and gaze direction. Furthermore, a new framework is proposed to model A/V multimodal interaction between verbal and nonverbal communication events. Dynamics of audio signals obtained from speaker diarization and head poses extracted from video images are modeled using hybrid dynamical systems (HDS). We show that HDS temporal structure characteristics can be used for multimodal interaction level estimation, which is useful feedback that can help to improve multi-party communication experience. Experimental results using synthetic and real-world datasets of group communication such as poster presentations show the feasibility of the proposed multimodal system. Tony Tung, Randy Gomez, Tatsuya Kawahara, Takashi Matsuyama |
IEEE Trans. Hum. Mach. Syst. | 1 |
| 2013 | Intrinsic Characterization of Dynamic SurfacesabstractThis paper presents a novel approach to characterize deformable surface using intrinsic property dynamics. 3D dynamic surfaces representing humans in motion can be obtained using multiple view stereo reconstruction methods or depth cameras. Nowadays these technologies have become capable to capture surface variations in real-time, and give details such as clothing wrinkles and deformations. Assuming repetitive patterns in the deformations, we propose to model complex surface variations using sets of linear dynamical systems (LDS) where observations across time are given by surface intrinsic properties such as local curvatures. We introduce an approach based on bags of dynamical systems, where each surface feature to be represented in the codebook is modeled by a set of LDS equipped with timing structure. Experiments are performed on datasets of real-world dynamical surfaces and show compelling results for description, classification and segmentation. Tony Tung, Takashi Matsuyama |
CVPR | 1 |
| 2013 | Scaling Memcache at Facebook
Rajesh Nishtala, Hans Fugal, Steven Grimm, Marc Kwiatkowski, Herman Lee, Harry C. Li, Ryan McElroy, Mike Paleczny, Daniel Peek, Paul Saab, David Stafford, Tony Tung, Venkateshwaran Venkataramani |
NSDI | 12 |
| 2012 | Invariant Surface-Based Shape Descriptor for Dynamic Surface Encoding
Tony Tung, Takashi Matsuyama |
ACCV (1) | 1 |
| 2012 | Topology Dictionary for 3D Video UnderstandingabstractThis paper presents a novel approach that achieves 3D video understanding. 3D video consists of a stream of 3D models of subjects in motion. The acquisition of long sequences requires large storage space (2 GB for 1 min). Moreover, it is tedious to browse data sets and extract meaningful information. We propose the topology dictionary to encode and describe 3D video content. The model consists of a topology-based shape descriptor dictionary which can be generated from either extracted patterns or training sequences. The model relies on 1) topology description and classification using Reeb graphs, and 2) a Markov motion graph to represent topology change states. We show that the use of Reeb graphs as the high-level topology descriptor is relevant. It allows the dictionary to automatically model complex sequences, whereas other strategies would require prior knowledge on the shape and topology of the captured subjects. Our approach serves to encode 3D video sequences, and can be applied for content-based description and summarization of 3D video sequences. Furthermore, topology class labeling during a learning process enables the system to perform content-based event recognition. Experiments were carried out on various 3D videos. We showcase an application for 3D video progressive summarization using the topology dictionary. Tony Tung, Takashi Matsuyama |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2010 | Dynamic surface matching by geodesic mapping for 3D animation transferabstractThis paper presents a novel approach that achieves complete matching of 3D dynamic surfaces. Surfaces are captured from multi-view video data and represented by sequences of 3D manifold meshes in motion (3D videos). We propose to perform dense surface matching between 3D video frames using geodesic diffeomorphisms. Our algorithm uses a coarse-to-fine strategy to derive a robust correspondence map, then a probabilistic formulation is coupled with a voting scheme in order to obtain local unicity of matching candidates and a smooth mapping. The significant advantage of the proposed technique compared to existing approaches is that it does not rely on a color-based feature extraction process. Hence, our method does not lose accuracy in poorly textured regions and is not bounded to be used on video sequences of a unique subject. Therefore our complete surface mapping can be applied to: (1) texture transfer between surface models extracted from different sequences, (2) dense motion flow estimation in 3D video, and (3) motion transfer from a 3D video to an unanimated 3D model. Experiments are performed on challenging publicly available real-world datasets and show compelling results. Tony Tung, Takashi Matsuyama |
CVPR | 1 |
| 2010 | 3D video performance segmentationabstractWe present a novel approach that achieves segmentation of subject body parts in 3D videos. 3D video consists in a free-viewpoint video of real-world subjects in motion immersed in a virtual world. Each 3D video frame is composed of one or several 3D models. A topology dictionary is used to cluster 3D video sequences with respect to the model topology and shape. The topology is characterized using Reeb graph-based descriptors and no prior explicit model on the subject shape is necessary to perform the clustering process. In this framework, the dictionary consists in a set of training input poses with a priori segmentation and labels. As a consequence, all identified frames of 3D video sequences can be automatically segmented. Finally, motion flows computed between consecutive frames are used to transfer segmented region labels to unidentified frames. Our method allows us to perform robust body part segmentation and tracking in 3D cinema sequences. Tony Tung, Takashi Matsuyama |
ICIP | 1 |
| 2009 | Topology dictionary with Markov model for 3D video content-based skimming and descriptionabstractThis paper presents a novel approach to skim and describe 3D videos. 3D video is an imaging technology which consists in a stream of 3D models in motion captured by a synchronized set of video cameras. Each frame is composed of one or several 3D models, and therefore the acquisition of long sequences at video rate requires massive storage devices. In order to reduce the storage cost while keeping relevant information, we propose to encode 3D video sequences using a topology-based shape descriptor dictionary. This dictionary is either generated from a set of extracted patterns or learned from training input sequences with semantic annotations. It relies on an unsupervised 3D shape-based clustering of the dataset by Reeb graphs, and features a Markov network to characterize topological changes. The approach allows content-based compression and skimming with accurate recovery of sequences and can handle complex topological changes. Redundancies are detected and skipped based on a probabilistic discrimination process. Semantic description of video sequences is then automatically performed. In addition, forthcoming frame encoding is achieved using a multiresolution matching scheme and allows action recognition in 3D. Our experiments were performed on complex 3D video sequences. We demonstrate the robustness and accuracy of the 3D video skimming with dramatic low bitrate coding and high compression ratio. Tony Tung, Takashi Matsuyama |
CVPR | 1 |
| 2009 | Complete multi-view reconstruction of dynamic scenes from probabilistic fusion of narrow and wide baseline stereoabstractThis paper presents a novel approach to achieve accurate and complete multi-view reconstruction of dynamic scenes (or 3D videos). 3D videos consist in sequences of 3D models in motion captured by a surrounding set of video cameras. To date 3D videos are reconstructed using multiview wide baseline stereo (MVS) reconstruction techniques. However it is still tedious to solve stereo correspondence problems: reconstruction accuracy falls when stereo photo-consistency is weak, and completeness is limited by self-occlusions. Most MVS techniques were indeed designed to deal with static objects in a controlled environment and therefore cannot solve these issues. Hence we propose to take advantage of the image content stability provided by each single-view video to recover any surface regions visible by at least one camera. In particular we present an original probabilistic framework to derive and predict the true surface of models. We propose to fuse multi-view structure-from-motion with robust 3D features obtained by MVS in order to significantly improve reconstruction completeness and accuracy. A min-cut problem where all exact features serve as priors is solved in a final step to reconstruct the 3D models. In addition, experimental results were conducted on synthetic and challenging real world datasets to illustrate the robustness and accuracy of our method. Tony Tung, Shohei Nobuhara, Takashi Matsuyama |
ICCV | 1 |
| 2009 | Minimal 3D videoabstractWe present a new concept that achieves the 3D reconstruction of dynamic scenes from multi-view video cameras (or 3D videos) using a minimal number of cameras, as opposed to the present state of the art approaches which require either several tens of cameras or high definition devices. A 3D video consists of a sequence of 3D models in motion captured by a surrounding set of video cameras. The result is a video where observers can choose freely their viewpoints. It is a markerless motion capture system where subjects do not need to wear special equipment. Hence, this system suits to a very wide range of applications (e.g. entertainment, medicine, sports, and so on). The 3D models are obtained using image-based multi-view stereo reconstruction techniques (or MVS). The performance of MVS relies on the quality and quantity of images taken from different viewpoints. As stereo correspondences have to be found between the images, the reconstruction fails in the case of weak stereo photo-consistency due to lack of camera views or lighting variations: consistent information is necessary. Tony Tung, Takashi Matsuyama |
SIGGRAPH ASIA Sketches | 1 |
| 2008 | Simultaneous super-resolution and 3D video using graph-cutsabstractThis paper presents a new method to increase the quality of 3D video, a new media developed to represent 3D objects in motion. This representation is obtained from multi-view reconstruction techniques that require images recorded simultaneously by several video cameras. All cameras are calibrated and placed around a dedicated studio to fully surround the models. The limited quality and quantity of cameras may produce inaccurate 3D model reconstruction with low quality texture. To overcome this issue, first we propose super-resolution (SR) techniques for 3D video: SR on multi-view images and SR on single-view video frames. Second, we propose to combine both super-resolution and dynamic 3D shape reconstruction problems into a unique Markov Random Field (MRF) energy formulation. The MRF minimization is performed using graph-cuts. Thus, we jointly compute the optimal solution for super-resolved texture and 3D shape model reconstruction. Moreover, we propose a coarse-to-fine strategy to iteratively produce 3D video with increasing quality. Our experiments show the accuracy and robustness of the proposed technique on challenging 3D video sequences. Tony Tung, Shohei Nobuhara, Takashi Matsuyama |
CVPR | 1 |
| 2008 | SHREC'08 entry: Shape retrieval of noisy watertight models using aMRGabstractThis paper presents an evaluation of the stability of the augmented multiresolution Reeb graph (aMRG) for shape retrieval of 3D watertight models. The method is based on a Reeb graph construction which is a well-known topology based shape descriptor. Using multiresolution property and additive geometrical and topological informations, aMRG has shown its efficiency and robustness to retrieve high quality 3D models. The SHREC - SHape REtrieval Contest 2008 Stability on Watertight Models Track data collection B is composed of 1500 models: 15 classes of 100 models. Each class contains original models and models with noise. We propose to evaluate the robustness of the aMRG with embedded topological features with regards to the following perturbations: Gaussian noise, uneven re-sampling, small protrusions, and topological noise. Tony Tung, Francis J. M. Schmitt |
Shape Modeling International | 1 |
| 2007 | Topology matching for 3D video compressionabstractThis paper presents a new technique to reduce the storage cost of high quality 3D video. In 3D video, a sequence of 3D objects represents scenes in motion. Every frame is composed by one or several accurate 3D meshes with attached high fidelity properties such as color and texture. Each frame is acquired at video rate. The entire video sequence requires a huge amount of free disk space. To overcome this issue, we propose an original approach using Reeb graphs, which are well-known topology based shape descriptors. In particular, we take advantage of the augmented multiresolution Reeb graph properties to store the relevant information of the 3D model of each frame. This graph structure has shown its efficiency as a motion descriptor, being able to track similar nodes all along the 3D video sequence. Therefore we can describe and reconstruct the 3D models of all frames with a very low-cost data size. The algorithm has been implemented as a fully automatic 3D video compression system. Our experiments show the robustness and accuracy of the proposed technique by comparing reconstructed sequences against challenging real ones. Tony Tung, Francis J. M. Schmitt, Takashi Matsuyama |
CVPR | 1 |
| 2004 | Augmented Reeb Graphs for Content-Based Retrieval of 3D Mesh ModelsabstractThis work presents an improved method of 3D mesh models indexing for content-based retrieval in database with shape similarity and appearance queries. The approach is based on the multiresolutional Reeb graph matching presented by Hilaga et al. (2001). The original method only takes into account topological information what is often not sufficient for effective matchings. Therefore we proposed to augment this graph with geometrical attributes. We also provide a new topological coherence condition to improve the graph matching. Moreover 2D appearance attributes and 3D features are extracted and merged to improve the estimation of the similarity between models. Besides, all these new attributes are user-dependent as they can be weighted by variable terms. We obtain a flexible multiresolutional and multicriteria representation called augmented Reeb graph (ARG). Good preliminary results have been obtained in a shape-based matching framework compared to existing methods based only on statistical measures. In addition, our study lead us to an innovative part matching scheme based on the same approach as our augmented Reeb graph matching. Tony Tung, Francis J. M. Schmitt |
SMI | 1 |
| 2004 | Augmented Reeb Graphs for Content-Based Retrieval of 3D Mesh Models (Figures 2 and 15)
Tony Tung, Francis J. M. Schmitt |
SMI | 1 |
| 2001 | Performance characterization of a hardware mechanism for dynamic optimizationabstractWe evaluate the rePLay microarchitecture as a means for reducing application execution time by facilitating dynamic optimization. The framework contains a programmable optimization engine coupled with a hardware-based recovery mechanism. The optimization engine enables the dynamic optimizer to run concurrently with program execution. The recovery mechanism enables the optimizer to make speculative optimizations without requiring recovery code. We demonstrate that a rePLay configuration performing a small suite of simple optimizations on Alpha code attains an average of 13% reduction in execution cycles on the SPEC2000 integer benchmarks over a rePLay configuration not performing optimizations, and a 21% reduction over an aggressive standard superscalar microarchitecture. Brian Fahs, Satarupa Bose, Matthew M. Crum, Brian Slechta, Francesco Spadini, Tony Tung, Sanjay J. Patel, Steven S. Lumetta |
MICRO | 6 |
| 2000 | Increasing the size of atomic instruction blocks using control flow assertionsabstractFor a variety of reasons, branch-less regions of instructions are desirable for high-performance execution. In this paper we propose a means for increasing the dynamic length of branch-less regions of instructions for the purposes of dynamic program optimization. We call these atomic regions frames and we construct them by replacing original branch instructions with assertions. Assertion instructions check if the original branching conditions still hold. If they hold, no action is taken. If they do not, then the entire region is undone. In this manner an assertion has no explicit control flow. We demonstrate that using branch correlation to decide when a branch should be converted into an assertion results in atomic regions that average over 100 instructions in length, with a probability of completion of 97%, and that constitute over 80% of the dynamic instruction stream. We demonstrate both static and dynamic means for constructing frames. When frames are built dynamically using finite sized hardware, they average 80 instructions in length and have good caching properties. Sanjay J. Patel, Tony Tung, Satarupa Bose, Matthew M. Crum |
MICRO | 2 |
| 1999 | HSRA: High-Speed, Hierarchical Synchroous Reconfigurable ArrayabstractThere is no inherent characteristic forcing Field Programmable Gate Array (FPGA) or Reconfigurable Computing (RC) Array cycle times to be greater than processors in the same process. Modern FPGAs seldom achieve application clock rates close to their processor cousins because (1) resources in the FPGAs are not balanced appropriately for high-speed operation, (2) FPGA CAD does not automatically provide the requisite transforms to support this operation, and (3) interconnect delays can be large and vary almost continuously, complicating high frequency mapping. We introduce a novel reconfigurable computing array, the High-Speed, Hierarchical Synchronous Reconfigurable Array (HSRA), and its supporting tools. This packagedemonstrates that computing arrays can achieve efficient, high-speedoperation. We have designedand implemented a prototype component in a 0.4 m logic design on a DRAM process which will support 250MHz operation for CAD mapped designs. William Tsu, Kip Macy, Atul Joshi, Randy Huang, Norman Walker, Tony Tung, Omid Rowhani, George Varghese, John Wawrzynek, André DeHon |
FPGA | 6 |