Tae-Kyun Kim 0001

dblp:28/787-1 · DBLP profile ↗
← Back
149ranked-venue papers
24as first author
42since 2021 · last 2025
0000-0002-7587-6053ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 116 · 20 first-author · 33 since 2021Graphics, computer vision, multimedia, augmented reality and games · 104 · 14 first-author · 28 since 2021Systems, architecture and hardware · 8 · 3 since 2021Human-computer interaction and ubiquitous computing · 3Applied, interdisciplinary, general and emerging computing · 2
YearPublicationVenuePosition
2025 REWIND: Real-Time Egocentric Whole-Body Motion Diffusion with Exemplar-Based Identity Conditioning
abstract
We present REWIND (Real-Time Egocentric Whole-Body Motion Diffusion), a one-step diffusion model for real-time, high-fidelity human motion estimation from egocentric image inputs. While an existing method for egocentric whole-body (i.e., body and hands) motion estimation is non-real-time and acausal due to diffusion-based iterative motion refinement to capture correlations between body and hand poses, REWIND operates in a fully causal and real-time manner. To enable real-time inference, we introduce (1) cascaded body-hand denoising diffusion, which effectively models the correlation between egocentric body and hand motions in a fast, feed-forward manner, and (2) diffusion distillation, which enables high-quality motion estimation with a single denoising step. Our denoising diffusion model is based on a modified Transformer architecture, designed to causally model output motions while enhancing generalizability to unseen motion lengths. Additionally, REWIND optionally supports identity-conditioned motion estimation when identity prior is available. To this end, we propose a novel identity conditioning method based on a small set of pose exemplars of the target identity, which further enhances motion estimation quality. Through extensive experiments, we demonstrate that REWIND significantly outperforms the existing baselines both with and without exemplar-based identity conditioning.
Weipeng Xu, Alexander Richard, Shih-En Wei, Shunsuke Saito, Shaojie Bai, Te-Li Wang, Minhyuk Sung, Tae-Kyun Kim 0001, Jason M. Saragih
CVPR9
2025 Joint Learning of Pose Regression and Denoising Diffusion with Score Scaling Sampling for Category-Level 6D Pose Estimation
abstract
Latest diffusion models have shown promising results in category-level 6D object pose estimation by modeling the conditional pose distribution with depth image input. The existing methods, however, suffer from slow convergence during training, learning its encoder with the diffusion denoising network in end-to-end fashion, and require an additional network that evaluates sampled pose hypotheses to filter out low-quality pose candidates. In this paper, we propose a novel pipeline that tackles these limitations by two key components. First, the proposed method pretrains the encoder with the direct pose regression head, and jointly learns the networks via the regression head and the denoising diffusion head, significantly accelerating training convergence while achieving higher accuracy. Second, sampling guidance via time-dependent score scaling is proposed s.t. the exploration-exploitation trade-off is effectively taken, eliminating the need for the additional evaluation network. The sampling guidance maintains multi-modal characteristics of symmetric objects at early denoising steps while ensuring high-quality pose generation at final steps. Extensive experiments on multiple benchmarks including REAL275, HouseCat6D, and ROPE, demonstrate that the proposed method, simple yet effective, achieves state-of-the-art accuracies even with single-pose inference, while being more efficient in both training and inference.
Tae-Kyun Kim 0001
ICCV2
2025 Body-Hand Modality Expertized Networks with Cross-attention for Fine-grained Skeleton Action Recognition
abstract
Skeleton-based Human Action Recognition (HAR) is a vital technology in robotics and human–robot interaction. However, most existing methods concentrate primarily on full-body movements and often overlook subtle hand motions that are critical for distinguishing fine-grained actions. Recent work leverages a unified graph representation that combines body, hand, and foot keypoints to capture detailed body dynamics. Yet, these models often blur fine hand details due to the disparity between body and hand action characteristics and the loss of subtle features during the spatial-pooling. In this paper, we propose BHaRNet (Body–Hand action Recognition Network), a novel framework that augments a typical body-expert model with a hand-expert model. Our model jointly trains both streams with an ensemble loss that fosters cooperative specialization, functioning in a manner reminiscent of a Mixture-of-Experts (MoE). Moreover, cross-attention is employed via an expertized branch method and a pooling-attention module to enable feature-level interactions and selectively fuse complementary information. Inspired by MMNet, we also demonstrate the applicability of our approach to multi-modal tasks by leveraging RGB information, where body features guide RGB learning to capture richer contextual cues. Experiments on large-scale benchmarks (NTU RGB+D 60, NTU RGB+D 120, PKU-MMD, and Northwestern-UCLA) demonstrate that BHaRNet achieves SOTA accuracies—improving from 86.4% to 93.0% in hand-intensive actions—while maintaining fewer GFLOPs and parameters than the relevant unified methods.
Seungyeon Cho, Tae-Kyun Kim 0001
IROS2
2025 SRHand: Super-Resolving Hand Images and 3D Shapes via View/Pose-aware Neural Image Representations and Explicit Meshes
abstract
Reconstructing detailed hand avatars plays a crucial role in various applications. While prior works have focused on capturing high-fidelity hand geometry, they heavily rely on high-resolution multi-view image inputs and struggle to generalize on low-resolution images. Multi-view image super-resolution methods have been proposed to enforce 3D view consistency. These methods, however, are limited to static objects/scenes with fixed resolutions and are not applicable to articulated deformable hands. In this paper, we propose SRHand (Super-Resolution Hand), the method for reconstructing detailed 3D geometry as well as textured images of hands from low-resolution images. SRHand leverages the advantages of implicit image representation with explicit hand meshes. Specifically, we introduce a geometric-aware implicit image function (GIIF) that learns detailed hand prior by upsampling the coarse input images. By jointly optimizing the implicit image function and explicit 3D hand shapes, our method preserves multi-view and pose consistency among upsampled hand images, and achieves fine-detailed 3D reconstruction (wrinkles, nails). In experiments using the InterHand2.6M and Goliath datasets, our method significantly outperforms state-of-the-art image upsampling methods adapted to hand datasets, and 3D hand reconstruction methods, quantitatively and qualitatively. The code will be publicly available.
Tae-Kyun Kim 0001
NeurIPS2
2025 MPMAvatar: Learning 3D Gaussian Avatars with Accurate and Robust Physics-Based Dynamics
abstract
While there has been significant progress in the field of 3D avatar creation from visual observations, modeling physically plausible dynamics of humans with loose garments remains a challenging problem. Although a few existing works address this problem by leveraging physical simulation, they suffer from limited accuracy or robustness to novel animation inputs. In this work, we present MPMAvatar, a framework for creating 3D human avatars from multi-view videos that supports highly realistic, robust animation, as well as photorealistic rendering from free viewpoints. For accurate and robust dynamics modeling, our key idea is to use a Material Point Method-based simulator, which we carefully tailor to model garments with complex deformations and contact with the underlying body by incorporating an anisotropic constitutive model and a novel collision handling algorithm. We combine this dynamics modeling scheme with our canonical avatar that can be rendered using 3D Gaussian Splatting with quasi-shadowing, enabling high-fidelity rendering for physically realistic animations. In our experiments, we demonstrate that MPMAvatar significantly outperforms the existing state-of-the-art physics-based avatar in terms of (1) dynamics modeling accuracy, (2) rendering accuracy, and (3) robustness and efficiency. Additionally, we present a novel application in which our avatar generalizes to unseen interactions in a zero-shot manner—which was not achievable with previous learning-based methods due to their limited simulation generalizability. Our code will be publicly available.
Tae-Kyun Kim 0001
NeurIPS3
2025 LLDiffusion: Learning degradation representations in diffusion models for low-light image enhancement
Tao Wang 0052, Kaihao Zhang, Yong Zhang 0034, Wenhan Luo, Björn Stenger, Tong Lu 0002, Tae-Kyun Kim 0001, Wei Liu 0005
Pattern Recognit.7
2025 BP-SGCN: Behavioral Pseudo-Label Informed Sparse Graph Convolution Network for Pedestrian and Heterogeneous Trajectory Prediction
abstract
Trajectory prediction allows better decision-making in applications of autonomous vehicles (AVs) or surveillance by predicting the short-term future movement of traffic agents. It is classified into pedestrian or heterogeneous trajectory prediction. The former exploits the relatively consistent behavior of pedestrians, but is limited in real-world scenarios with heterogeneous traffic agents such as cyclists and vehicles. The latter typically relies on extra class label information to distinguish the heterogeneous agents, but such labels are costly to annotate and cannot be generalized to represent different behaviors within the same class of agents. In this work, we introduce the behavioral pseudo-labels that effectively capture the behavior distributions of pedestrians and heterogeneous agents solely based on their motion features, significantly improving the accuracy of trajectory prediction. To implement the framework, we propose the behavioral pseudo-label informed sparse graph convolution network (BP-SGCN) that learns pseudo-labels and informs to a trajectory predictor. For optimization, we propose a cascaded training scheme, in which we first learn the pseudo-labels in an unsupervised manner, and then perform end-to-end fine-tuning on the labels in the direction of increasing the trajectory prediction accuracy. Experiments show that our pseudo-labels effectively model different behavior clusters and improve trajectory prediction. Our proposed BP-SGCN outperforms existing methods using both pedestrian [ETH/UCY, pedestrian-only Stanford Drone Dataset (SDD)] and heterogeneous agent datasets (SDD and Argoverse 1).
Ruochen Li 0002, Stamos Katsigiannis, Tae-Kyun Kim 0001, Hubert P. H. Shum
IEEE Trans. Neural Networks Learn. Syst.3
2024 Prompt Augmentation for Self-supervised Text-guided Image Manipulation
abstract
Text-guided image editing finds applications in various creative and practical fields. While recent studies in image generation have advanced the field, they often struggle with the dual challenges of coherent image transformation and context preservation. In response, our work introduces prompt augmentation, a method amplifying a single input prompt into several target prompts, strengthening textual context and enabling localised image editing. Specifically, we use the augmented prompts to delineate the intended manipulation area. We propose a Contrastive Loss tailored to driving effective image editing by displacing edited areas and drawing preserved regions closer. Acknowledging the continuous nature of image manipulations, we further refine our approach by incorporating the similarity concept, creating a Soft Contrastive Loss. The new losses are incorporated to the diffusion model, demonstrating improved or competitive image editing results on public datasets and generated images over state-of-the-art approaches.
Rumeysa Bodur, Binod Bhattarai, Tae-Kyun Kim 0001
CVPR3
2024 Arbitrary-Scale Image Generation and Upsampling Using Latent Diffusion Model and Implicit Neural Decoder
abstract
Super-resolution (SR) and image generation are important tasks in computer vision and are widely adopted in real-world applications. Most existing methods, however, generate images only at fixed-scale magnification and suffer from over-smoothing and artifacts. Additionally, they do not offer enough diversity of output images nor image consistency at different scales. Most relevant work applied Implicit Neural Representation (INR) to the denoising diffusion model to obtain continuous-resolution yet diverse and high-quality SR results. Since this model operates in the image space, the larger the resolution of image is produced, the more memory and inference time is required, and it also does not maintain scale-specific consistency. We propose a novel pipeline that can super-resolve an input image or generate from a random noise a novel image at arbitrary scales. The method consists of a pre-trained auto-encoder, a latent diffusion model, and an implicit neural decoder, and their learning strategies. The proposed method adopts diffusion processes in a latent space, thus efficient, yet aligned with output image space decoded by MLPs at arbitrary scales. More specifically, our arbitrary-scale decoder is designed by the symmetric decoder w/o up-scaling from the pre-trained auto-encoder, and Local Implicit Image Function (LIIF) in series. The latent diffusion process is learnt by the denoising and the alignment losses jointly. Errors in output images are backpropagated via the fixed decoder, improving the quality of output images. In the extensive experiments using multiple public benchmarks on the two tasks i.e. image super-resolution and novel image generation at arbitrary scales, the proposed method outperforms relevant methods in metrics of image quality, diversity and scale consistency. It is significantly better than the relevant prior-art in the inference speed and memory usage.
Tae-Kyun Kim 0001
CVPR2
2024 BiTT: Bi-Directional Texture Reconstruction of Interacting Two Hands from a Single Image
abstract
Creating personalized hand avatars is important to offer a realistic experience to users on AR / VR platforms. While most prior studies focused on reconstructing 3D hand shapes, some recent work has tackled the reconstruction of hand textures on top of shapes. However, these methods are often limited to capturing pixels on the visible side of a hand, requiring diverse views of the hand in a video or multiple images as input. In this paper, we propose a novel method, BiTT(Bi-directional Texture reconstruction of Two hands), which is the first end-to-end trainable method for relightable, pose-free texture reconstruction of two interacting hands taking only a single RGB image, by three novel components: 1) bi-directional (left$\leftrightarrow$right) texture reconstruction using the texture symmetry of left / right hands, 2) utilizing a texture parametric model for hand texture recovery, and 3) the overall coarse-to-fine stage pipeline for reconstructing personalized texture of two interacting hands. BiTT first estimates the scene light condition and albedo image from an input image, then reconstructs the texture of both hands through the texture parametric model and bi-directional texture reconstructor. In experiments using InterHand2.6M and RGB2Hands datasets, our method significantly outperforms state-of-the-art hand texture reconstruction methods quantitatively and qualitatively. The code is available at https://github.com/yunminjin2/BiTT.
Tae-Kyun Kim 0001
CVPR2
2024 InterHandGen: Two-Hand Interaction Generation via Cascaded Reverse Diffusion
abstract
We present InterHandGen, a novel framework that learns the generative prior of two-hand interaction. Sampling from our model yields plausible and diverse two-hand shapes in close interaction with or without an object. Our prior can be incorporated into any optimization or learning methods to reduce ambiguity in an ill-posed setup. Our key observation is that directly modeling the joint distribution of multiple instances imposes high learning complexity due to its combinatorial nature. Thus, we propose to decom-pose the modeling of joint distribution into the modeling of factored unconditional and conditional single instance distribution. In particular, we introduce a diffusion model that learns the single-hand distribution unconditional and conditional to another hand via conditioning dropout. For sampling, we combine anti-penetration and classifier-free guidance to enable plausible generation. Furthermore, we establish the rigorous evaluation protocol of two-hand synthesis, where our method significantly outperforms baseline generative models in terms of plausibility and diversity. We also demonstrate that our diffusion prior can boost the performance of two-hand reconstruction from monocular in-the-wild images, achieving new state-of-the-art accuracy.
Shunsuke Saito, Giljoo Nam, Minhyuk Sung, Tae-Kyun Kim 0001
CVPR5
2024 Dense Hand-Object (HO) GraspNet with Full Grasping Taxonomy and Dynamics
Woojin Cho 0002, Minjae Yi, Taeyun Woo, Taewook Ha, Hyokeun Lee, Je-Hwan Ryu, Woontack Woo, Tae-Kyun Kim 0001
ECCV (82)11
2024 Semi-Supervised 3D Object Detection With Channel Augmentation Using Transformation Equivariance
abstract
Accurate 3D object detection is crucial for autonomous vehicles and robots to navigate and interact with the environment safely and effectively. Meanwhile, the performance of 3D detector relies on the data size and annotation which is expensive. Consequently, the demand of training with limited labeled data is growing. We explore a novel teacher-student framework employing channel augmentation for 3D semisupervised object detection. The teacher-student SSL typically adopts a weak augmentation and strong augmentation to teacher and student, respectively. In this work, we apply multiple channel augmentations to both networks using the transformation equivariance detector (TED). The TED allows us to explore different combinations of augmentation on point clouds and efficiently aggregates multi-channel transformation equivariance features. In principle, by adopting fixed channel augmentations for the teacher network, the student can train stably on reliable pseudo-labels. Adopting strong channel augmentations can enrich the diversity of data, fostering robustness to transformations and enhancing generalization performance of the student network. We use SOTA hierarchical supervision as a baseline and adapt its dual-threshold to TED, which is called channel IoU consistency. We evaluate our method with KITTI dataset, and achieved a significant performance leap, surpassing SOTA 3D semi-supervised object detection models.
Minju Kang, Taehun Kong, Tae-Kyun Kim 0001
ICIP3
2024 Hand-Object Reconstruction Via Interaction-Aware Graph Attention Mechanism
abstract
Estimating the poses of both a hand and an object has become an important area of research due to the growing need for advanced vision computing. The primary challenge involves understanding and reconstructing how hands and objects interact, such as contact and physical plausibility. Existing approaches often adopt a graph neural network to incorporate spatial information of hand and object meshes. However, these approaches have not fully exploited the potential of graphs without modification of edges within and between hand- and object-graphs. We propose a graph-based refinement method that incorporates an interaction-aware graph-attention mechanism to account for hand-object interactions. Using edges, we establish connections among closely correlated nodes, both within individual graphs and across different graphs. Experiments demonstrate the effectiveness of our proposed method with notable improvements in the realm of physical plausibility.
Taeyun Woo, Tae-Kyun Kim 0001, Jinah Park
ICIP2
2024 Multi-hypotheses Conditioned Point Cloud Diffusion for 3D Human Reconstruction from Occluded Images
abstract
3D human shape reconstruction under severe occlusion due to human-object or human-human interaction is a challenging problem. While implicit function methods capture detailed clothed shapes, they require aligned shape priors and or are weak at inpainting occluded regions given an image input. Parametric models i.e. SMPL, instead offer whole body shapes, however, are often misaligned with images. In this work, we propose a novel pipeline composed of a probabilistic SMPL model and point cloud diffusion for pixel-aligned detailed 3D human reconstruction under occlusion. Multiple hypotheses generated by the probabilistic SMPL method are conditioned via continuous 3D shape representations. Point cloud diffusion refines the distribution of 3D points fitted to both the multi-hypothesis shape condition and pixel-aligned image features, offering detailed clothed shapes and inpainting occluded parts of human bodies. In the experiments using the CAPE, MultiHuman and Hi4D datasets, the proposed method outperforms various SOTA methods based on SMPL, implicit functions, point cloud diffusion, and their combined, under synthetic and real occlusions. Our code is publicly available at https://donghwankim0101.github.io/projects/mhcdiff.
Tae-Kyun Kim 0001
NeurIPS2
2024 GridFormer: Residual Dense Transformer with Grid Structure for Image Restoration in Adverse Weather Conditions
Tao Wang 0052, Kaihao Zhang, Ziqian Shao, Wenhan Luo, Björn Stenger, Tong Lu 0002, Tae-Kyun Kim 0001, Wei Liu 0005, Hongdong Li
Int. J. Comput. Vis.7
2024 Dream360: Diverse and Immersive Outdoor Virtual Scene Creation via Transformer-Based 360° Image Outpainting
abstract
360° images, with a field-of-view (FoV) of $180^{\circ}\times 360^{\circ}$, provide immersive and realistic environments for emerging virtual reality (VR) applications, such as virtual tourism, where users desire to create diverse panoramic scenes from a narrow FoV photo they take from a viewpoint via portable devices. It thus brings us to a technical challenge: 'How to allow the users to freely create diverse and immersive virtual scenes from a narrow FoV image with a specified viewport?' To this end, we propose a transformer-based 360° image outpainting framework called Dream360, which can generate diverse, high-fidelity, and high-resolution panoramas from user-selected viewports, considering the spherical properties of 360° images. Compared with existing methods, e.g., [3], which primarily focus on inputs with rectangular masks and central locations while overlooking the spherical property of 360° images, our Dream360 offers higher outpainting flexibility and fidelity based on the spherical representation. Dream360 comprises two key learning stages: (I) codebook-based panorama outpainting via Spherical-VQGAN (S-VQGAN), and (II) frequency-aware refinement with a novel frequency-aware consistency loss. Specifically, S-VQGAN learns a sphere-specific codebook from spherical harmonic (SH) values, providing a better representation of spherical data distribution for scene modeling. The frequency-aware refinement matches the resolution and further improves the semantic consistency and visual fidelity of the generated results. Our Dream360 achieves significantly lower Frechet Inception Distance (FID) scores and better visual fidelity than existing methods. We also conducted a user study involving 15 participants to interactively evaluate the quality of the generated results in VR, demonstrating the flexibility and superiority of our Dream360 framework.
Hao Ai, Zidong Cao, Haonan Lu, Chen Chen 0015, Jian Ma 0010, Peng Yuan Zhou, Tae-Kyun Kim 0001, Pan Hui 0001, Lin Wang 0025
IEEE Trans. Vis. Comput. Graph.7
2023 Unsupervised Contour Tracking of Live Cells by Mechanical and Cycle Consistency Losses
abstract
Analyzing the dynamic changes of cellular morphology is important for understanding the various functions and characteristics of live cells, including stem cells and metastatic cancer cells. To this end, we need to track all points on the highly deformable cellular contour in every frame of live cell video. Local shapes and textures on the contour are not evident, and their motions are complex, often with expansion and contraction of local contour features. The prior arts for optical flow or deep point set tracking are unsuited due to the fluidity of cells, and previous deep contour tracking does not consider point correspondence. We propose the first deep learning-based tracking of cellular (or more generally viscoelastic materials) contours with point correspondence by fusing dense representation between two contours with cross attention. Since it is impractical to manually label dense tracking points on the contour, unsupervised learning comprised of the mechanical and cyclical consistency losses is proposed to train our contour tracker. The mechanical loss forcing the points to move perpendicular to the contour effectively helps out. For quantitative evaluation, we labeled sparse tracking points along the contour of live cells from two live cell datasets taken with phase contrast and confocal fluorescence microscopes. Our contour tracker quantitatively outperforms compared methods and produces qualitatively more favorable results. Our code and data are publicly available at https://github.com/JunbongJang/contour-tracking/.
Junbong Jang, Kwonmoo Lee, Tae-Kyun Kim 0001
CVPR3
2023 Im2Hands: Learning Attentive Implicit Representation of Interacting Two-Hand Shapes
abstract
We present Implicit Two Hands (Im2Hands), the first neural implicit representation of two interacting hands. Unlike existing methods on two-hand reconstruction that rely on a parametric hand model and/or low-resolution meshes, Im2Hands can produce fine-grained geometry of two hands with high hand-to-hand and hand-to-image coherency. To handle the shape complexity and interaction context between two hands, Im2Hands models the occupancy volume of two hands – conditioned on an RGB image and coarse 3D keypoints – by two novel attention-based modules responsible for (1) initial occupancy estimation and (2) context-aware occupancy refinement, respectively. Im2Hands first learns per-hand neural articulated occupancy in the canonical space designed for each hand using query-image attention. It then refines the initial two-hand occupancy in the posed space to enhance the coherency between the two hand shapes using query-anchor attention. In addition, we introduce an optional keypoint refinement module to enable robust two-hand shape estimation from predicted hand keypoints in a single-image reconstruction scenario. We experimentally demonstrate the effectiveness of Im2Hands on two-hand reconstruction in comparison to related methods, where ours achieves state-of-the-art results. Our code is publicly available at https://github.com/jyunlee/Im2Hands.
Minhyuk Sung, Honggyu Choi, Tae-Kyun Kim 0001
CVPR4
2023 Joint Training of Hierarchical GANs and Semantic Segmentation for Expression Translation
abstract
Manipulating images by changing only specific attributes has been a long-standing research problem. Existing methods that rely solely on a global generator often suffer from changing unwanted attributes along with the desired attributes. Although hierarchical networks consisting of global and local networks have shown success, they extract local regions using bounding boxes and are non-differential, inaccurate, and unrealistic. As a result, the solution becomes suboptimal and introduces unwanted artifacts. A recent study has shown a strong correlation between facial attributes and local regions. To exploit this correlation, we have designed a unified architecture that combines semantic segmentation and hierarchical GANs. One advantage of our end-to-end differential framework is that the segmentation network conditions the GANs during the forward pass, and gradients from the GANs are propagated to the segmentation network during the backward pass, allowing both architectures to benefit from each other. We evaluated our method on two challenging expression translation benchmarks, AffectNet and RaFD, and a segmentation benchmark, CelebAMask-HQ, validating its effectiveness over existing methods.
Rumeysa Bodur, Binod Bhattarai, Tae-Kyun Kim 0001
ICASSP3
2023 3D Distillation: Improving Self-Supervised Monocular Depth Estimation on Reflective Surfaces
abstract
Self-supervised monocular depth estimation (SSMDE) aims at predicting the dense depth maps of monocular images, by learning to minimize a photometric loss using spatially neighboring image pairs during training. While SSMDE offers a significant scalability advantage over supervised approaches, it performs poorly on reflective surfaces as the photometric constancy assumption of the photometric loss is violated. We note that the appearance of reflective surfaces is view-dependent and often there are views of such surfaces in the training data that are not contaminated by strong specular reflections. Thus, reflective surfaces can be accurately reconstructed by aggregating the predicted depth of these views. Motivated by this observation, we propose 3D distillation: a novel training framework that utilizes the projected depth of reconstructed reflective surfaces to generate reasonably accurate depth pseudo-labels. To identify those surfaces automatically, we employ an uncertainty-guided depth fusion method, combining the smoother and more accurate projected depth on reflective surfaces and the detailed predicted depth elsewhere. In our experiments using the ScanNet and 7-Scenes datasets, we show that 3D distillation not only significantly improves the prediction accuracy, especially on the problematic surfaces, but also that it generalizes well over various underlying network architectures and to new datasets.
Xuepeng Shi, Georgi Dikov, Gerhard Reitmayr, Tae-Kyun Kim 0001, Mohsen Ghafoorian
ICCV4
2023 MAPConNet: Self-supervised 3D Pose Transfer with Mesh and Point Contrastive Learning
abstract
3D pose transfer is a challenging generation task that aims to transfer the pose of a source geometry onto a target geometry with the target identity preserved. Many prior methods require keypoint annotations to find correspondence between the source and target. Current pose transfer methods allow end-to-end correspondence learning but require the desired final output as ground truth for supervision. Unsupervised methods have been proposed for graph convolutional models but they require ground truth correspondence between the source and target inputs. We present a novel self-supervised framework for 3D pose transfer which can be trained in unsupervised, semi-supervised, or fully supervised settings without any correspondence labels. We introduce two contrastive learning constraints in the latent space: a mesh-level loss for disentangling global patterns including pose and identity, and a point-level loss for discriminating local semantics. We demonstrate quantitatively and qualitatively that our method achieves state-of-the-art results in supervised 3D pose transfer, with comparable results in unsupervised and semi-supervised settings. Our method is also generalisable to unseen human and animal data with complex topologies†.
Jiaze Sun, Zhixiang Chen 0003, Tae-Kyun Kim 0001
ICCV3
2023 FourierHandFlow: Neural 4D Hand Representation Using Fourier Query Flow
abstract
Recent 4D shape representations model continuous temporal evolution of implicit shapes by (1) learning query flows without leveraging shape and articulation priors or (2) decoding shape occupancies separately for each time value. Thus, they do not effectively capture implicit correspondences between articulated shapes or regularize jittery temporal deformations. In this work, we present FourierHandFlow, which is a spatio-temporally continuous representation for human hands that combines a 3D occupancy field with articulation-aware query flows represented as Fourier series. Given an input RGB sequence, we aim to learn a fixed number of Fourier coefficients for each query flow to guarantee smooth and continuous temporal shape dynamics. To effectively model spatio-temporal deformations of articulated hands, we compose our 4D representation based on two types of Fourier query flow: (1) pose flow that models query dynamics influenced by hand articulation changes via implicit linear blend skinning and (2) shape flow that models query-wise displacement flow. In the experiments, our method achieves state-of-the-art results on video-based 4D reconstruction while being computationally more efficient than the existing 3D/4D implicit shape representations. We additionally show our results on motion inter- and extrapolation and texture transfer using the learned correspondences of implicit shapes. To the best of our knowledge, FourierHandFlow is the first neural 4D continuous hand representation learned from RGB videos. The code will be publicly accessible.
Junbong Jang, Minhyuk Sung, Tae-Kyun Kim 0001
NeurIPS5
2023 CRT-6D: Fast 6D Object Pose Estimation with Cascaded Refinement Transformers
abstract
Learning based 6D object pose estimation methods rely on computing large intermediate pose representations and/or iteratively refining an initial estimation with a slow render-compare pipeline. This paper introduces a novel method we call Cascaded Pose Refinement Transformers, or CRT-6D. We replace the commonly used dense intermediate representation with a sparse set of features sampled from the feature pyramid we call OSKFs(Object Surface Keypoint Features) where each element corresponds to an object keypoint. We employ lightweight deformable transformers and chain them together to iteratively refine proposed poses over the sampled OSKFs. We achieve inference runtimes 2× faster than the closest real-time state of the art methods while supporting up to 21 objects on a single model. We demonstrate the effectiveness of CRT-6D by performing extensive experiments on the LM-O and YCBV datasets. Compared to real-time methods, we achieve state of the art on LM-O and YCB-V, falling slightly behind methods with inference runtimes one order of magnitude higher. The source code is available at: https://github.com/PedroCastro/CRT-6D
Tae-Kyun Kim 0001
WACV2
2023 Multivariate Probabilistic Monocular 3D Object Detection
abstract
In autonomous driving, monocular 3D object detection is an important but challenging task. Towards accurate monocular 3D object detection, some recent methods recover the distance of objects from the physical height and visual height of objects. Such decomposition framework can introduce explicit constraints on the distance prediction, thus improving its accuracy and robustness. However, the inaccurate physical height and visual height prediction still may exacerbate the inaccuracy of the distance prediction. In this paper, we improve the framework by multivariate probabilistic modeling. We explicitly model the joint probability distribution of the physical height and visual height. This is achieved by learning a full covariance matrix of the physical height and visual height during training, with the guide of a multivariate likelihood. Such explicit joint probability distribution modeling not only leads to robust distance prediction when both the predicted physical height and visual height are inaccurate, but also brings learned covariance matrices with expected behaviors. The experimental results on the challenging Waymo Open and KITTI datasets show the effectiveness of our framework1.
Xuepeng Shi, Zhixiang Chen 0003, Tae-Kyun Kim 0001
WACV3
2022 MoBYv2AL: Self-supervised Active Learning for Image Classification
Razvan Caramalau, Binod Bhattarai, Danail Stoyanov, Tae-Kyun Kim 0001
BMVC4
2022 Semi-Supervised Object Detection with Object-wise Contrastive Learning and Regression Uncertainty
Honggyu Choi, Zhixiang Chen 0003, Xuepeng Shi, Tae-Kyun Kim 0001
BMVC4
2022 Pop-Out Motion: 3D-Aware Image Deformation via Learning the Shape Laplacian
abstract
We propose a framework that can deform an object in a 2D image as it exists in 3D space. Most existing methods for 3D-aware image manipulation are limited to (1) only changing the global scene information or depth, or (2) manipulating an object of specific categories. In this paper, we present a 3D-aware image deformation method with minimal restrictions on shape category and deformation type. While our framework leverages 2D-to-3D reconstruction, we argue that reconstruction is not sufficient for realistic deformations due to the vulnerability to topological errors. Thus, we propose to take a supervised learning-based approach to predict the shape Laplacian of the underlying volume of a 3D reconstruction represented as a point cloud. Given the deformation energy calculated using the predicted shape Laplacian and user-defined deformation handles (e.g., keypoints), we obtain bounded biharmonic weights to model plausible handle-based image deformation. In the experiments, we present our results of deforming 2D character and clothed human images. We also quantitatively show that our approach can produce more accurate deformation weights compared to alternative methods (i.e., mesh reconstruction and point cloud Laplacian methods).
Minhyuk Sung, Tae-Kyun Kim 0001
CVPR4
2022 Label Geometry Aware Discriminator for Conditional Generative Adversarial Networks
abstract
Multi-domain image-to-image translation with conditional Generative Adversarial Networks (GANs) can generate highly photo realistic images with desired target classes, yet these synthetic images have not always been helpful to improve downstream supervised tasks such as image classification. Improving downstream tasks with synthetic examples requires generating images with high fidelity to the unknown conditional distribution of the target class, which many labeled conditional GANs attempt to achieve by adding soft-max cross-entropy loss based auxiliary classifier in the discriminator. As recent studies suggest that the soft-max loss in Euclidean space of deep feature does not leverage their intrinsic angular distribution, we propose to replace this loss in auxiliary classifier with an additive angular margin (AAM) loss that takes benefit of the intrinsic angular distribution, and promotes intra-class compactness and inter-class separation to help generator synthesize high fidelity images. We validate on RaFD and CIFAR-100, two challenging face expression and image classification data set. Our method outperforms state-of-the-art methods in several different evaluation criteria including recently proposed GAN-train and GAN-test metrics designed to assess the impact of synthetic data on downstream classification task, assessing the usefulness in data augmentation for supervised tasks with prediction accuracy score and average confidence score.
Suman Sapkota, Bidur Khanal, Binod Bhattarai, Bishesh Khanal, Tae-Kyun Kim 0001
ICPR5
2022 Modular Adaptive Policy Selection for Multi- Task Imitation Learning through Task Division
abstract
Deep imitation learning requires many expert demonstrations, which can be hard to obtain, especially when many tasks are involved. However, different tasks often share similarities, so learning them jointly can greatly benefit them and alleviate the need for many demonstrations. But, joint multi-task learning often suffers from negative transfer, sharing information that should be task-specific. In this work, we introduce a method to perform multi-task imitation while allowing for task-specific features. This is done by using proto-policies as modules to divide the tasks into simple sub-behaviours that can be shared. The proto-policies operate in parallel and are adaptively chosen by a selector mechanism that is jointly trained with the modules. Experiments on different sets of tasks show that our method improves upon the accuracy of single agents, task-conditioned and multi-headed multi-task agents, as well as state-of-the-art meta learning agents. We also demonstrate its ability to autonomously divide the tasks into both shared and task-specific sub-behaviours.
Dafni Antotsiou, Carlo Ciliberto, Tae-Kyun Kim 0001
ICRA3
2022 SHAQ: Incorporating Shapley Value Theory into Multi-Agent Q-Learning
abstract
Value factorisation is a useful technique for multi-agent reinforcement learning (MARL) in global reward game, however, its underlying mechanism is not yet fully understood. This paper studies a theoretical framework for value factorisation with interpretability via Shapley value theory. We generalise Shapley value to Markov convex game called Markov Shapley value (MSV) and apply it as a value factorisation method in global reward game, which is obtained by the equivalence between the two games. Based on the properties of MSV, we derive Shapley-Bellman optimality equation (SBOE) to evaluate the optimal MSV, which corresponds to an optimal joint deterministic policy. Furthermore, we propose Shapley-Bellman operator (SBO) that is proved to solve SBOE. With a stochastic approximation and some transformations, a new MARL algorithm called Shapley Q-learning (SHAQ) is established, the implementation of which is guided by the theoretical results of SBO and MSV. We also discuss the relationship between SHAQ and relevant value factorisation methods. In the experiments, SHAQ exhibits not only superior performances on all tasks but also the interpretability that agrees with the theoretical analysis. The implementation of this paper is placed on https://github.com/hsvgbkhgbv/shapley-q-learning.
Yuan Zhang 0027, Yunjie Gu, Tae-Kyun Kim 0001
NeurIPS4
2022 Unsupervised random forest for affinity estimation
abstract
This paper presents an unsupervised clustering random-forest-based metric for affinity estimation in large and high-dimensional data. The criterion used for node splitting during forest construction can handle rank-deficiency when measuring cluster compactness. The binary forest-based metric is extended to continuous metrics by exploiting both the common traversal path and the smallest shared parent node. The proposed forest-based metric efficiently estimates affinity by passing down data pairs in the forest using a limited number of decision trees. A pseudo-leaf-splitting (PLS) algorithm is introduced to account for spatial relationships, which regularizes affinity measures and overcomes inconsistent leaf assign-ments. The random-forest-based metric with PLS facilitates the establishment of consistent and point-wise correspondences. The proposed method has been applied to automatic phrase recognition using color and depth videos and point-wise correspondence. Extensive experiments demonstrate the effectiveness of the proposed method in affinity estimation in a comparison with the state-of-the-art.
Yunai Yi, Diya Sun, Peixin Li, Tae-Kyun Kim 0001, Tianmin Xu, Yuru Pei
Comput. Vis. Media4
2022 Joint Framework for Single Image Reconstruction and Super-Resolution With an Event Camera
abstract
Event cameras sense brightness changes in each pixel and yield asynchronous event streams instead of producing intensity images. They have distinct advantages over conventional cameras, such as a high dynamic range (HDR) and no motion blur. To take advantage of event cameras with existing image-based algorithms, a few methods have been proposed to reconstruct images from event streams. However, the output images have a low resolution (LR) and are unrealistic. Low-quality outputs stem from broader applications of event cameras, where high-quality and high-resolution (HR) images are needed. In this work, we consider the problem of reconstructing and super-resolving images from LR events when no ground truth (GT) HR images and degradation models are available. We propose a novel end-to-end joint framework for single image reconstruction and super-resolution from LR event data. Our method is primarily unsupervised to handle the absence of real inputs from GT and deploys adversarial learning. To train our framework, we constructed an open dataset, including simulated events and real-world images. The use of the dataset boosts the network performance, and the network architectures and various loss functions in each phase help improve the quality of the resulting image. Various experiments showed that our method surpasses the state-of-the-art LR image reconstruction methods for real-world and synthetic datasets. The experiments for super-resolution (SR) image reconstruction also substantiate the effectiveness of the proposed method. We further extended our method to more challenging problems of HDR, sharp image reconstruction, and color events. In addition, we demonstrate that the reconstruction and super-resolution results serve as intermediate representations of events for high-level tasks, such as semantic segmentation, object recognition, and detection. We further examined how events affect the outputs of the three phases and analyze our method's efficacy through an ablation study.
Lin Wang 0025, Tae-Kyun Kim 0001, Kuk-Jin Yoon
IEEE Trans. Pattern Anal. Mach. Intell.2
2021 Sequential Graph Convolutional Network for Active Learning
abstract
We propose a novel pool-based Active Learning framework constructed on a sequential Graph Convolution Network (GCN). Each images feature from a pool of data represents a node in the graph and the edges encode their similarities. With a small number of randomly sampled images as seed labelled examples, we learn the parameters of the graph to distinguish labelled vs unlabelled nodes by minimising the binary cross-entropy loss. GCN performs message-passing operations between the nodes, and hence, induces similar representations of the strongly associated nodes. We exploit these characteristics of GCN to select the unlabelled examples which are sufficiently different from labelled ones. To this end, we utilise the graph node embeddings and their confidence scores and adapt sampling techniques such as CoreSet and uncertainty-based methods to query the nodes. We flip the label of newly queried nodes from unlabelled to labelled, re-train the learner to optimise the downstream task and the graph to minimise its modified objective. We continue this process within a fixed budget. We evaluate our method on 6 different benchmarks: 4 real image classification, 1 depth-based hand pose estimation and 1 synthetic RGB image classification datasets. Our method outperforms several competitive baselines such as VAAL, Learning Loss, CoreSet and attains the new state-of-the-art performance on multiple applications.
Razvan Caramalau, Binod Bhattarai, Tae-Kyun Kim 0001
CVPR3
2021 Learning Feature Aggregation for Deep 3D Morphable Models
abstract
3D morphable models are widely used for the shape representation of an object class in computer vision and graphics applications. In this work, we focus on deep 3D morphable models that directly apply deep learning on 3D mesh data with a hierarchical structure to capture information at multiple scales. While great efforts have been made to design the convolution operator, how to best aggregate vertex features across hierarchical levels deserves further attention. In contrast to resorting to mesh decimation, we propose an attention based module to learn mapping matrices for better feature aggregation across hierarchical levels. Specifically, the mapping matrices are generated by a compatibility function of the keys and queries. The keys and queries are trainable variables, learned by optimizing the target objective, and shared by all data samples of the same object class. Our proposed module can be used as a train-only drop-in replacement for the feature aggregation in existing architectures for both downsampling and upsampling. Our experiments show that through the end-to-end training of the mapping matrices, we achieve state-of-the-art results on a variety of 3D shape datasets in comparison to existing morphable models.
Zhixiang Chen 0003, Tae-Kyun Kim 0001
CVPR2
2021 EvDistill: Asynchronous Events To End-Task Learning via Bidirectional Reconstruction-Guided Cross-Modal Knowledge Distillation
abstract
Event cameras sense per-pixel intensity changes and produce asynchronous event streams with high dynamic range and less motion blur, showing advantages over the conventional cameras. A hurdle of training event-based models is the lack of large qualitative labeled data. Prior works learning end-tasks mostly rely on labeled or pseudo-labeled datasets obtained from the active pixel sensor (APS) frames; however, such datasets’ quality is far from rivaling those based on the canonical images. In this paper, we propose a novel approach, called EvDistill, to learn a student network on the unlabeled and unpaired event data (target modality) via knowledge distillation (KD) from a teacher network trained with large-scale, labeled image data (source modality). To enable KD across the unpaired modalities, we first propose a bidirectional modality reconstruction (BMR) module to bridge both modalities and simultaneously exploit them to distill knowledge via the crafted pairs, causing no extra computation in the inference. The BMR is improved by the end-tasks and KD losses in an end-to-end manner. Second, we leverage the structural similarities of both modalities and adapt the knowledge by matching their distributions. Moreover, as most prior feature KD methods are uni-modality and less applicable to our problem, we propose an affinity graph KD loss to boost the distillation. Our extensive experiments on semantic segmentation and object recognition demonstrate that EvDistill achieves significantly better results than the prior works and KD with only events and APS frames.
Lin Wang 0025, Yujeong Chae, Sung-Hoon Yoon 0001, Tae-Kyun Kim 0001, Kuk-Jin Yoon
CVPR4
2021 Geometry-based Distance Decomposition for Monocular 3D Object Detection
abstract
Monocular 3D object detection is of great significance for autonomous driving but remains challenging. The core challenge is to predict the distance of objects in the absence of explicit depth information. Unlike regressing the distance as a single variable in most existing methods, we propose a novel geometry-based distance decomposition to recover the distance by its factors. The decomposition factors the distance of objects into the most representative and stable variables, i.e. the physical height and the projected visual height in the image plane. Moreover, the decomposition maintains the self-consistency between the two heights, leading to robust distance prediction when both predicted heights are inaccurate. The decomposition also enables us to trace the causes of the distance uncertainty for different scenarios. Such decomposition makes the distance prediction interpretable, accurate, and robust. Our method directly predicts 3D bounding boxes from RGB images with a compact architecture, making the training and inference simple and efficient. The experimental results show that our method achieves the state-of-the-art performance on the monocular 3D Object Detection and Bird’s Eye View tasks of the KITTI dataset, and can generalize to images with different camera intrinsics1.
Xuepeng Shi, Qi Ye 0001, Xiaozhi Chen, Chuangrong Chen, Zhixiang Chen 0003, Tae-Kyun Kim 0001
ICCV6
2021 Modelling Hierarchical Structure between Dialogue Policy and Natural Language Generator with Option Framework for Task-oriented Dialogue System
Yuan Zhang 0027, Tae-Kyun Kim 0001, Yunjie Gu
ICLR3
2021 Adversarial Imitation Learning with Trajectorial Augmentation and Correction
abstract
Deep Imitation Learning requires a large number of expert demonstrations, which are not always easy to obtain, especially for complex tasks. A way to overcome this shortage of labels is through data augmentation. However, this cannot be easily applied to control tasks due to the sequential nature of the problem. In this work, we introduce a novel augmentation method which preserves the success of the augmented trajectories. To achieve this, we introduce a semi-supervised correction network that aims to correct distorted expert actions. To adequately test the abilities of the correction network, we develop an adversarial data augmented imitation architecture to train an imitation agent using synthetic experts. Additionally, we introduce a metric to measure diversity in trajectory datasets. Experiments show that our data augmentation strategy can improve accuracy and convergence time of adversarial imitation while preserving the diversity between the generated and real trajectories.
Dafni Antotsiou, Carlo Ciliberto, Tae-Kyun Kim 0001
ICRA3
2021 3D Dense Geometry-Guided Facial Expression Synthesis by Adversarial Learning
abstract
Manipulating facial expressions is a challenging task due to fine-grained shape changes produced by facial muscles and the lack of input-output pairs for supervised learning. Unlike previous methods using Generative Adversarial Networks (GAN), which rely on cycle-consistency loss or sparse geometry (landmarks) loss for expression synthesis, we propose a novel GAN framework to exploit 3D dense (depth and surface normals) information for expression manipulation. However, a large-scale dataset containing RGB images with expression annotations and their corresponding depth maps is not available. To this end, we propose to use an off-the-shelf state-of-the-art 3D reconstruction model to estimate the depth and create a large-scale RGB-Depth dataset after a manual data clean-up process. We utilise this dataset to minimise the novel depth consistency loss via adversarial learning (note we do not have ground truth depth maps for generated face images) and the depth categorical loss of synthetic data on the discriminator. In addition, to improve the generalisation and lower the bias of the depth parameters, we propose to use a novel confidence regulariser on the discriminator side of the framework. We extensively performed both quantitative and qualitative evaluations on two publicly available challenging facial expression benchmarks: AffectNet and RaFD. Our experiments demonstrate that the proposed method outperforms the competitive baseline and existing arts by a large margin.
Rumeysa Bodur, Binod Bhattarai, Tae-Kyun Kim 0001
WACV3
2021 Active Learning for Bayesian 3D Hand Pose Estimation
abstract
We propose a Bayesian approximation to a deep learning architecture for 3D hand pose estimation. Through this framework, we explore and analyse the two types of uncertainties that are influenced either by data or by the learning capability. Furthermore, we draw comparisons against the standard estimator over three popular benchmarks. The first contribution lies in outperforming the baseline while in the second part we address the active learning application. We also show that with a newly proposed acquisition function, our Bayesian 3D hand pose estimator obtains lowest errors with the least amount of data. The underlying code is publicly available at: https://github.com/razvancaramalau/al_bhpe.
Razvan Caramalau, Binod Bhattarai, Tae-Kyun Kim 0001
WACV3
2021 Multiple object tracking: A literature review
Wenhan Luo, Junliang Xing, Anton Milan, Xiaoqin Zhang 0002, Wei Liu 0005, Tae-Kyun Kim 0001
Artif. Intell.6
2020 Introducing Pose Consistency and Warp-Alignment for Self-Supervised 6D Object Pose Estimation in Color Images
abstract
Most successful approaches to estimate the 6D pose of an object typically train a neural network by supervising the learning with annotated poses in real world images. These annotations are generally expensive to obtain and a common workaround is to generate and train on synthetic scenes, with the drawback of limited generalisation when the model is deployed in the real world. In this work, a two-stage 6D object pose estimator framework that can be applied on top of existing neural-network-based approaches and that does not require pose annotations on real images is proposed. The first self-supervised stage enforces the pose consistency between rendered predictions and real input images, narrowing the gap between the two domains. The second stage fine-tunes the previously trained model by enforcing the photometric consistency between pairs of different object views, where one image is warped and aligned to match the view of the other and thus enabling their comparison. In the absence of both real image annotations and depth information, applying the proposed framework on top of two recent approaches results in state-of-the-art performance when compared to methods trained only on synthetic data, domain adaptation baselines and a concurrent self-supervised approach on LINEMOD, LINEMOD OCCLUSION and HomebrewedDB datasets.
Juil Sock, Guillermo Garcia-Hernando, Anil Armagan, Tae-Kyun Kim 0001
3DV4
2020 Shapley Q-Value: A Local Reward Approach to Solve Global Reward Games
abstract
Cooperative game is a critical research area in the multi-agent reinforcement learning (MARL). Global reward game is a subclass of cooperative games, where all agents aim to maximize the global reward. Credit assignment is an important problem studied in the global reward game. Most of previous works stood by the view of non-cooperative-game theoretical framework with the shared reward approach, i.e., each agent being assigned a shared global reward directly. This, however, may give each agent an inaccurate reward on its contribution to the group, which could cause inefficient learning. To deal with this problem, we i) introduce a cooperative-game theoretical framework called extended convex game (ECG) that is a superset of global reward game, and ii) propose a local reward approach called Shapley Q-value. Shapley Q-value is able to distribute the global reward, reflecting each agent's own contribution in contrast to the shared reward approach. Moreover, we derive an MARL algorithm called Shapley Q-value deep deterministic policy gradient (SQDDPG), using Shapley Q-value as the critic for each agent. We evaluate SQDDPG on Cooperative Navigation, Prey-and-Predator and Traffic Junction, compared with the state-of-the-art algorithms, e.g., MADDPG, COMA, Independent DDPG and Independent A2C. In the experiments, SQDDPG shows a significant improvement on the convergence rate. Finally, we plot Shapley Q-value and validate the property of fair credit assignment.
Yuan Zhang 0027, Tae-Kyun Kim 0001, Yunjie Gu
AAAI3
2020 MatchGAN: A Self-supervised Semi-supervised Conditional Generative Adversarial Network
Jiaze Sun, Binod Bhattarai, Tae-Kyun Kim 0001
ACCV (4)3
2020 EventSR: From Asynchronous Events to Image Reconstruction, Restoration, and Super-Resolution via End-to-End Adversarial Learning
abstract
Event cameras sense intensity changes and have many advantages over conventional cameras. To take advantage of event cameras, some methods have been proposed to reconstruct intensity images from event streams. However, the outputs are still in low resolution (LR), noisy, and unrealistic. The low-quality outputs stem broader applications of event cameras, where high spatial resolution (HR) is needed as well as high temporal resolution, dynamic range, and no motion blur. We consider the problem of reconstructing and super-resolving intensity images from pure events, when no ground truth (GT) HR images and down-sampling kernels are available. To tackle the challenges, we propose a novel end-to-end pipeline that reconstructs LR images from event streams, enhances the image qualities and upsamples the enhanced images, called EventSR. For the absence of real GT images, our method is primarily unsupervised, deploying adversarial learning. To train EventSR, we create an open dataset including both real-world and simulated scenes. The use of both datasets boosts up the network performance, and the network architectures and various loss functions in each phase help improve the image qualities. The whole pipeline is trained in three phases. While each phase is mainly for one of the three tasks, the networks in earlier phases are fine-tuned by respective loss functions in an end-to-end manner. Experimental results show that EventSR generates high-quality SR images from events for both simulated and real-world data.
Lin Wang 0025, Tae-Kyun Kim 0001, Kuk-Jin Yoon
CVPR2
2020 Weakly-Supervised Domain Adaptation via GAN and Mesh Model for Estimating 3D Hand Poses Interacting Objects
abstract
Despite recent successes in hand pose estimation, there yet remain challenges on RGB-based 3D hand pose estimation (HPE) under hand-object interaction (HOI) scenarios where severe occlusions and cluttered backgrounds exhibit. Recent RGB HOI benchmarks have been collected either in real or synthetic domain, however, the size of datasets is far from enough to deal with diverse objects combined with hand poses, and 3D pose annotations of real samples are lacking, especially for occluded cases. In this work, we propose a novel end-to-end trainable pipeline that adapts the hand-object domain to the single hand-only domain, while learning for HPE. The domain adaption occurs in image space via 2D pixel-level guidance by Generative Adversarial Network (GAN) and 3D mesh guidance by mesh renderer (MR). Via the domain adaption in image space, not only 3D HPE accuracy is improved, but also HOI input images are translated to segmented and de-occluded hand-only images. The proposed method takes advantages of both the guidances: GAN accurately aligns hands, while MR effectively fills in occluded pixels. The experiments using Dexter-Object, Ego-Dexter and HO3D datasets show that our method significantly outperforms state-of-the-arts trained by hand-only data and is comparable to those supervised by HOI data. Note our method is trained primarily by hand-only images with pose labels, and HOI images without pose labels.
Seungryul Baek, Kwang In Kim, Tae-Kyun Kim 0001
CVPR3
2020 Measuring Generalisation to Unseen Viewpoints, Articulations, Shapes and Objects for 3D Hand Pose Estimation Under Hand-Object Interaction
Anil Armagan, Guillermo Garcia-Hernando, Seungryul Baek, Shreyas Hampali, Mahdi Rad, Shipeng Xie, Mingxiu Chen, Boshen Zhang, Fu Xiong, Yang Xiao 0007, Zhiguo Cao 0001, Junsong Yuan 0001, Pengfei Ren 0001, Weiting Huang, Haifeng Sun 0001, Marek Hrúz, Jakub Kanis, Zdenek Krnoul, Qingfu Wan, Shile Li, Linlin Yang 0001, Dongheui Lee, Angela Yao, Weiguo Zhou, Sijia Mei, Adrian Spurr, Umar Iqbal 0001, Pavlo Molchanov 0001, Philippe Weinzaepfel, Romain Brégier, Grégory Rogez, Vincent Lepetit, Tae-Kyun Kim 0001
ECCV (23)35
2020 Inducing Optimal Attribute Representations for Conditional GANs
Binod Bhattarai, Tae-Kyun Kim 0001
ECCV (7)2
2020 Unsupervised Learning of Optical Flow with Deep Feature Similarity
Woobin Im, Tae-Kyun Kim 0001, Sung-Eui Yoon
ECCV (24)2
2020 Distance-Normalized Unified Representation for Monocular 3D Object Detection
Xuepeng Shi, Zhixiang Chen 0003, Tae-Kyun Kim 0001
ECCV (29)3
2020 Sampling Strategies for GAN Synthetic Data
abstract
Generative Adversarial Networks (GANs) have been used widely to generate large volumes of synthetic data. This data is being utilised for augmenting with real examples in order to train deep Convolutional Neural Networks (CNNs). Studies have shown that the generated examples lack sufficient realism to train deep CNNs and are poor in diversity. Unlike previous studies of randomly augmenting the synthetic data with real data, we present our simple, effective and easy to implement synthetic data sampling methods to train deep CNNs more efficiently and accurately. To this end, we propose to maximally utilise the parameters learned during training of the GAN itself. These include discriminator's realism confidence score and the confidence on the target label of the synthetic data. In addition to this, we explore reinforcement learning (RL) to automatically search a subset of meaningful synthetic examples from a large pool of GAN synthetic data. We evaluate our method on two challenging face attribute classification data sets viz. AffectNet and CelebA. Our extensive experiments clearly demonstrate the need of sampling synthetic data before augmentation, which also improves the performance of one of the state-of-the-art deep CNNs in vitro.
Binod Bhattarai, Seungryul Baek, Rumeysa Bodur, Tae-Kyun Kim 0001
ICASSP4
2020 Auglabel: Exploiting Word Representations to Augment Labels for Face Attribute Classification
abstract
Augmenting data in image space (eg. flipping, cropping etc) and activation space (eg. dropout) are being widely used to regularise deep neural networks and have been successfully applied on several computer vision tasks. Unlike previous works, which are mostly focused on doing augmentation in the aforementioned domains, we propose to do augmentation in label space. In this paper, we present a simple, yet effective novel method to generate fixed dimensional labels with continuous values for images by exploiting the word2vec - semantic representations - of the existing categorical labels. We then append these representations to existing categorical labels and train the model. We validated our idea on two challenging face attribute classification data sets viz. CelebA and LFWA. Our extensive experiments show that the augmented labels improve the performance of the competitive deep learning baseline and also attains the state-of-the-art performance.
Binod Bhattarai, Rumeysa Bodur, Tae-Kyun Kim 0001
ICASSP3
2020 Accurate 6D Object Pose Estimation by Pose Conditioned Mesh Reconstruction
abstract
Current 6D object pose estimation methods consist of Deep Convo-lutional Neural Networks fully optimized for a single object but with its architecture standardized among objects with different shapes. In contrast to previous works, we explicitly exploit each object's distinct topological information with an automated process and prior to any post-processing refinement stage. In order to achieve this, we propose a learning framework in which a Graph Convolutional Neural Network reconstructs a Pose Conditioned 3D mesh of the object. A robust estimation of the allocentric orientation of the target object is recovered by computing, in a differentiable manner, the Procrustes' alignment between the canonical and reconstructed dense 3D meshes. Our method is capable of self validating its pose estimation by measuring the quality of the reconstructed mesh, which is invaluable in real life applications. In our experiments on the LINEMOD, OCCLUSION and YCB-Video benchmarks, the proposed method outperforms state-of-the-arts.
Anil Armagan, Tae-Kyun Kim 0001
ICASSP3
2020 Physics-Based Dexterous Manipulations with Estimated Hand Poses and Residual Reinforcement Learning
abstract
Dexterous manipulation of objects in virtual environments with our bare hands, by using only a depth sensor and a state-of-the-art 3D hand pose estimator (HPE), is challenging. While virtual environments are ruled by physics, e.g. object weights and surface frictions, the absence of force feedback makes the task challenging, as even slight inaccuracies on finger tips or contact points from HPE may make the interactions fail. Prior arts simply generate contact forces in the direction of the fingers' closures, when finger joints penetrate virtual objects. Although useful for simple grasping scenarios, they cannot be applied to dexterous manipulations such as inhand manipulation. Existing reinforcement learning (RL) and imitation learning (IL) approaches train agents that learn skills by using task-specific rewards, without considering any online user input. In this work, we propose to learn a model that maps noisy input hand poses to target virtual poses, which introduces the needed contacts to accomplish the tasks on a physics simulator. The agent is trained in a residual setting by using a model-free hybrid RL+IL approach. A 3D hand pose estimation reward is introduced leading to an improvement on HPE accuracy when the physics-guided corrected target poses are remapped to the input space. As the model corrects HPE errors by applying minor but crucial joint displacements for contacts, this helps to keep the generated motion visually close to the user input. Since HPE sequences performing successful virtual interactions do not exist, a data generation scheme to train and evaluate the system is proposed. We test our framework in two applications that use hand pose estimates for dexterous manipulations: hand-object interactions in VR and hand-object motion reconstruction in-the-wild. Experiments show that the proposed method outperforms various RL/IL baselines and the simple prior art of enforcing hand closure, both in task success and hand pose accuracy.
Guillermo Garcia-Hernando, Edward Johns, Tae-Kyun Kim 0001
IROS3
2020 Active 6D Multi-Object Pose Estimation in Cluttered Scenarios with Deep Reinforcement Learning
abstract
In this work, we explore how a strategic selection of camera movements can facilitate the task of 6D multi-object pose estimation in cluttered scenarios while respecting real-world constraints such as time and distance travelled, important in robotics and augmented reality applications. In the proposed framework, multiple object hypotheses inferred by an object pose estimator are accumulated both in space and time with a fusion function. At each time step, this fusion function makes use of a verification score to quantify the quality of the hypotheses in the absence of ground-truth annotations and passes this information to an agent. The agent reasons about these hypotheses, directing its attention to the object which it is most uncertain about, moving the camera towards such an object. Unlike previous works that propose short-sighted policies, our agent is trained in simulated scenarios using reinforcement learning, attempting to learn the camera moves that produce the most accurate object poses hypotheses for a given temporal and spatial budget, without the need of viewpoints rendering during inference. Our experiments show that the proposed approach successfully estimates the 6D object pose of a stack of objects in both challenging cluttered synthetic and real scenarios, showing superior performance compared to other baselines.
Juil Sock, Guillermo Garcia-Hernando, Tae-Kyun Kim 0001
IROS3
2020 3D Hand Pose Estimation with a Single Infrared Camera via Domain Transfer Learning
abstract
Previous methods successfully estimated 3D hand poses from unblurred depth images with slow and smooth hand motions. However, the performance drops when the depth images are contaminated by motion blur due to fast hand motion. In this paper, we exploit an infrared (IR) image input, which is weakly blurred under fast hand motion. The proposed method is based on domain transfer learning from depth to infrared images. Note we do not have IR images with hand skeletons, thus proposing self-supervision rather than direct supervision using the skeleton labels. We train a Hand Image Generator (HIG) and two Hand Pose Estimators (HPEs) on paired depth and infrared images via self-supervision using a consistency loss, guided by an existing HPE trained on paired depth and hand skeleton entries. The IR-based HPE is then refined on the weakly blurred infrared images. The qualitative and quantitative experiments demonstrate that the proposed method accurately estimates 3D hand poses under motion blur by fast hand motion, while existing depth-based methods fail. Our solution therefore supports fast 3D manipulation of virtual objects for augmented reality applications. Our model and dataset are publicly available for future research.1
Gabyong Park, Tae-Kyun Kim 0001, Woontack Woo
ISMAR2
2020 DeepFisheye: Near-Surface Multi-Finger Tracking Technology Using Fisheye Camera
abstract
Near-surface multi-finger tracking (NMFT) technology expands the input space of touchscreens by enabling novel interactions such as mid-air and finger-aware interactions. We present DeepFisheye, a practical NMFT solution for mobile devices, that utilizes a fisheye camera attached at the bottom of a touchscreen. DeepFisheye acquires the image of an interacting hand positioned above the touchscreen using the camera and employs deep learning to estimate the 3D position of each fingertip. We created two new hand pose datasets comprising fisheye images, on which our network was trained. We evaluated DeepFisheye's performance for three device sizes. DeepFisheye showed average errors with approximate value of 20 mm for fingertip tracking across the different device sizes. Additionally, we created simple rule-based classifiers that estimate the contact finger and hand posture from DeepFisheye's output. The contact finger and hand posture classifiers showed accuracy of approximately 83 and 90%, respectively, across the device sizes.
Keun-Woo Park, Sunbum Kim, Youngwoo Yoon, Tae-Kyun Kim 0001, Geehyuk Lee
UIST4
2020 A review on object pose recovery: From 3D bounding box detectors to full 6D pose estimators
Caner Sahin, Guillermo Garcia-Hernando, Juil Sock, Tae-Kyun Kim 0001
Image Vis. Comput.4
2019 Pushing the Envelope for RGB-Based Dense 3D Hand Pose Estimation via Neural Rendering
abstract
Estimating 3D hand meshes from single RGB images is challenging, due to intrinsic 2D-3D mapping ambiguities and limited training data. We adopt a compact parametric 3D hand model that represents deformable and articulated hand meshes. To achieve the model fitting to RGB images, we investigate and contribute in three ways: 1) Neural rendering: inspired by recent work on human body, our hand mesh estimator (HME) is implemented by a neural network and a differentiable renderer, supervised by 2D segmentation masks and 3D skeletons. HME demonstrates good performance for estimating diverse hand shapes and improves pose estimation accuracies. 2) Iterative testing refinement: Our fitting function is differentiable. We iteratively refine the initial estimate using the gradients, in the spirit of iterative model fitting methods like ICP. The idea is supported by the latest research on human body. 3) Self-data augmentation: collecting sized RGB-mesh (or segmentation mask)-skeleton triplets for training is a big hurdle. Once the model is successfully fitted to input RGB images, its meshes i.e. shapes and articulations, are realistic, and we augment view-points on top of estimated dense hand poses. Experiments using three RGB-based benchmarks show that our framework offers beyond state-of-the-art accuracy in 3D pose estimation, as well as recovers dense 3D hand shapes. Each technical component above meaningfully improves the accuracy in the ablation study.
Seungryul Baek, Kwang In Kim, Tae-Kyun Kim 0001
CVPR3
2019 Special Issue on Machine Vision
Tae-Kyun Kim 0001, Stefanos Zafeiriou, Ben Glocker, Stefan Leutenegger
Int. J. Comput. Vis.1
2019 Editorial: Special Issue on Deep Learning for Face Analysis
Chen Change Loy, Xiaoming Liu 0002, Tae-Kyun Kim 0001, Fernando De la Torre, Rama Chellappa
Int. J. Comput. Vis.3
2019 Opening the Black Box: Hierarchical Sampling Optimization for Hand Pose Estimation
abstract
Hand pose estimation, formulated as an inverse problem, is typically optimized by an energy function over pose parameters using a 'black box' image generation procedure, knowing little about either the relationships between the parameters or the form of the energy function. In this paper, we show significant improvement upon such black box optimization by exploiting high-level knowledge of the parameter structure and using a local surrogate energy function. Our new framework, hierarchical sampling optimization (HSO), consists of a sequence of discriminative predictors organized into a kinematic hierarchy. Each predictor is conditioned on its ancestors, and generates a set of samples over a subset of the pose parameters, with only one selected by the highly-efficient surrogate energy. The selected partial poses are concatenated to generate a full-pose hypothesis. Repeating the same process, several hypotheses are generated and the full energy function selects the best result. Under the same kinematic hierarchy, two methods based on decision forest and convolutional neural network are proposed to generate the samples and two optimization methods are studied when optimizing these samples. Experimental evaluations on three publicly available datasets show that our method is particularly impressive in low-compute scenarios where it significantly outperforms all other state-of-the-art methods.
Danhang Tang, Qi Ye 0001, Shanxin Yuan, Jonathan Taylor 0001, Pushmeet Kohli, Cem Keskin, Tae-Kyun Kim 0001, Jamie Shotton
IEEE Trans. Pattern Anal. Mach. Intell.7
2019 Trajectories as Topics: Multi-Object Tracking by Topic Discovery
abstract
This paper proposes a new approach to multi-object tracking by semantic topic discovery. We dynamically cluster frame-by-frame detections and treat objects as topics, allowing the application of the Dirichlet process mixture model. The tracking problem is cast as a topic-discovery task, where the video sequence is treated analogously to a document. It addresses tracking issues such as object exclusivity constraints as well as tracking management without the need for heuristic thresholds. Variation of object appearance is modeled as the dynamics of word co-occurrence and handled by updating the cluster parameters across the sequence in the dynamical clustering procedure. We develop two kinds of visual representation based on super-pixel and deformable part model and integrate them into the model of automatic topic discovery for tracking rigid and non-rigid objects, respectively. In experiments on public data sets, we demonstrate the effectiveness of the proposed algorithm.
Wenhan Luo, Björn Stenger, Tae-Kyun Kim 0001
IEEE Trans. Image Process.4
2018 Perceiving, Learning, and Recognizing 3D Objects: An Approach to Cognitive Service Robots
abstract
There is growing need for robots that can interact with people in everyday situations. For service robots, it is not reasonable to assume that one can pre-program all object categories. Instead, apart from learning from a batch of labelled training data, robots should continuously update and learn new object categories while working in the environment. This paper proposes a cognitive architecture designed to create a concurrent 3D object category learning and recognition in an interactive and open-ended manner. In particular, this cognitive architecture provides automatic perception capabilities that will allow robots to detect objects in highly crowded scenes and learn new object categories from the set of accumulated experiences in an incremental and open-ended way. Moreover, it supports constructing the full model of an unknown object in an on-line manner and predicting next best view for improving object detection and manipulation performance. We provide extensive experimental results demonstrating system performance in terms of recognition, scalability, next-best-view prediction and real-world robotic applications.
Hamidreza Kasaei 0001, Juil Sock, Luís Seabra Lopes, Ana Maria Tomé, Tae-Kyun Kim 0001
AAAI5
2018 Multi-Task Deep Networks for Depth-Based 6D Object Pose and Joint Registration in Crowd Scenarios
Juil Sock, Kwang In Kim, Caner Sahin, Tae-Kyun Kim 0001
BMVC4
2018 Augmented Skeleton Space Transfer for Depth-Based Hand Pose Estimation
abstract
Crucial to the success of training a depth-based 3D hand pose estimator (HPE) is the availability of comprehensive datasets covering diverse camera perspectives, shapes, and pose variations. However, collecting such annotated datasets is challenging. We propose to complete existing databases by generating new database entries. The key idea is to synthesize data in the skeleton space (instead of doing so in the depth-map space) which enables an easy and intuitive way of manipulating data entries. Since the skeleton entries generated in this way do not have the corresponding depth map entries, we exploit them by training a separate hand pose generator (HPG) which synthesizes the depth map from the skeleton entries. By training the HPG and HPE in a single unified optimization framework enforcing that 1) the HPE agrees with the paired depth and skeleton entries; and 2) the HPG-HPE combination satisfies the cyclic consistency (both the input and the output of HPG-HPE are skeletons) observed via the newly generated unpaired skeletons, our algorithm constructs a HPE which is robust to variations that go beyond the coverage of the existing database. Our training algorithm adopts the generative adversarial networks (GAN) training process. As a by-product, we obtain a hand pose discriminator (HPD) that is capable of picking out realistic hand poses. Our algorithm exploits this capability to refine the initial skeleton estimates in testing, further improving the accuracy. We test our algorithm on four challenging benchmark datasets (ICVL, MSRA, NYU and Big Hand 2.2M datasets) and demonstrate that our approach outperforms or is on par with state-of-the-art methods quantitatively and qualitatively.
Seungryul Baek, Kwang In Kim, Tae-Kyun Kim 0001
CVPR3
2018 First-Person Hand Action Benchmark With RGB-D Videos and 3D Hand Pose Annotations
abstract
In this work we study the use of 3D hand poses to recognize first-person dynamic hand actions interacting with 3D objects. Towards this goal, we collected RGB-D video sequences comprised of more than 100K frames of 45 daily hand action categories, involving 26 different objects in several hand configurations. To obtain hand pose annotations, we used our own mo-cap system that automatically infers the 3D location of each of the 21 joints of a hand model via 6 magnetic sensors and inverse kinematics. Additionally, we recorded the 6D object poses and provide 3D object models for a subset of hand-object interaction sequences. To the best of our knowledge, this is the first benchmark that enables the study of first-person hand actions with the use of 3D hand poses. We present an extensive experimental evaluation of RGB-D and pose-based action recognition by 18 baselines/state-of-the-art approaches. The impact of using appearance features, poses, and their combinations are measured, and the different training/testing protocols are evaluated. Finally, we assess how ready the 3D hand pose estimation field is when hands are severely occluded by objects in egocentric views and its influence on action recognition. From the results, we see clear benefits of using hand pose as a cue for action recognition compared to other data modalities. Our dataset and experiments can be of interest to communities of 3D hand pose estimation, 6D object pose, and robotics as well as action recognition.
Guillermo Garcia-Hernando, Shanxin Yuan, Seungryul Baek, Tae-Kyun Kim 0001
CVPR4
2018 Depth-Based 3D Hand Pose Estimation: From Current Achievements to Future Goals
abstract
In this paper, we strive to answer two questions: What is the current state of 3D hand pose estimation from depth images? And, what are the next challenges that need to be tackled? Following the successful Hands In the Million Challenge (HIM2017), we investigate the top 10 state-of-the-art methods on three tasks: single frame 3D pose estimation, 3D hand tracking, and hand pose estimation during object interaction. We analyze the performance of different CNN structures with regard to hand shape, joint visibility, view point and articulation distributions. Our findings include: (1) isolated 3D hand pose estimation achieves low mean errors (10 mm) in the view point range of [70, 120] degrees, but it is far from being solved for extreme view points; (2) 3D volumetric representations outperform 2D CNNs, better capturing the spatial structure of the depth data; (3) Discriminative methods still generalize poorly to unseen hand shapes; (4) While joint occlusions pose a challenge for most methods, explicit modeling of structure constraints can significantly narrow the gap between errors on visible and occluded joints.
Shanxin Yuan, Guillermo Garcia-Hernando, Björn Stenger, Gyeongsik Moon, Ju Yong Chang, Kyoung Mu Lee, Pavlo Molchanov 0001, Jan Kautz, Sina Honari, Liuhao Ge, Junsong Yuan 0001, Xinghao Chen 0001, Guijin Wang, Fan Yang 0032, Kai Akiyama, Yang Wu 0001, Qingfu Wan, Meysam Madadi, Sergio Escalera, Shile Li, Dongheui Lee, Iasonas Oikonomidis, Antonis A. Argyros, Tae-Kyun Kim 0001
CVPR24
2018 Semi-supervised Adversarial Learning to Generate Photorealistic Face Images of New Identities from 3D Morphable Model
Baris Gecer, Binod Bhattarai, Josef Kittler, Tae-Kyun Kim 0001
ECCV (11)4
2018 BOP: Benchmark for 6D Object Pose Estimation
Tomas Hodan, Frank Michel 0002, Eric Brachmann, Wadim Kehl, Anders Glent Buch, Dirk Kraft, Bertram Drost, Joel Vidal, Stephan Ihrke, Xenophon Zabulis, Caner Sahin, Fabian Manhardt, Federico Tombari, Tae-Kyun Kim 0001, Jiri Matas, Carsten Rother
ECCV (10)14
2018 Occlusion-Aware Hand Pose Estimation Using Hierarchical Mixture Density Network
Qi Ye 0001, Tae-Kyun Kim 0001
ECCV (10)2
2018 Spatio-temporal elastic cuboid trajectories for efficient fight recognition using Hough forests
Ismael Serrano, Oscar Déniz-Suárez, Gloria Bueno García, Guillermo Garcia-Hernando, Tae-Kyun Kim 0001
Mach. Vis. Appl.5
2018 Latent-Class Hough Forests for 6 DoF Object Pose Estimation
abstract
In this paper we present Latent-Class Hough Forests, a method for object detection and 6 DoF pose estimation in heavily cluttered and occluded scenarios. We adapt a state of the art template matching feature into a scale-invariant patch descriptor and integrate it into a regression forest using a novel template-based split function. We train with positive samples only and we treat class distributions at the leaf nodes as latent variables. During testing we infer by iteratively updating these distributions, providing accurate estimation of background clutter and foreground occlusions and, thus, better detection rate. Furthermore, as a by-product, our Latent-Class Hough Forests can provide accurate occlusion aware segmentation masks, even in the multi-instance scenario. In addition to an existing public dataset, which contains only single-instance sequences with large amounts of clutter, we have collected two, more challenging, datasets for multiple-instance detection containing heavy 2D and 3D clutter as well as foreground occlusions. We provide extensive experiments on the various parameters of the framework such as patch size, number of trees and number of iterations to infer class distributions at test time. We also evaluate the Latent-Class Hough Forests on all datasets where we outperform state of the art methods.
Alykhan Tejani, Rigas Kouskouridas, Andreas Doumanoglou, Danhang Tang, Tae-Kyun Kim 0001
IEEE Trans. Pattern Anal. Mach. Intell.5
2018 Latent Bi-Constraint SVM for Video-Based Object Recognition
abstract
We address the task of recognizing objects from video input. This important problem is relatively unexplored, compared with image-based object recognition. To this end, we make the following contributions. First, we introduce two comprehensive data sets for video-based object recognition. Second, we propose latent bi-constraint SVM (LBSVM), a maximum-margin framework for video-based object recognition. LBSVM is based on structured-output SVM, but extends it to handle noisy video data and ensure consistency of the output decision throughout time. We apply LBSVM to recognize office objects and museum sculptures, and we demonstrate its benefits over image-based, set-based, and other video-based object recognition.
Yang Liu 0020, Minh Hoai, Mang Shao, Tae-Kyun Kim 0001
IEEE Trans. Circuits Syst. Video Technol.4
2018 Spatially Consistent Supervoxel Correspondences of Cone-Beam Computed Tomography Images
abstract
Establishing dense correspondences of cone-beam computed tomography (CBCT) images is a crucial step for the attribute transfer and morphological variation assessment in clinical orthodontics. In this paper, a novel method, unsupervised spatially consistent clustering forest, is proposed to tackle the challenges for automatic supervoxel-wise correspondences of CBCT images. A complexity analysis of the proposed method with respect to the clustering hypotheses is provided with a data-dependent learning guarantee. The learning bound considers both the sequential tree traversals determined by questions stored in branch nodes and the clustering compactness of leaf nodes. A novel tree-pruning algorithm, guided by the learning bound, is also proposed to remove locally inconsistent leaf nodes. The resulting forest yields spatially consistent affinity estimations, thanks to the pruning penalizing trees with inconsistent leaf assignments and the combinational contextual feature channels used to learn the forest. A forest-based metric is utilized to derive the pairwise affinities and dense correspondences of CBCT images. The proposed method has been applied to the label propagation of clinically captured CBCT images. In the experiments, the method outperforms variants of both supervised and unsupervised forest-based methods and state-of-the-art label-propagation methods, achieving the mean dice similarity coefficients of 0.92, 0.89, 0.94, and 0.93 for the mandible, the maxilla, the zygoma arch, and the teeth data, respectively.
Yuru Pei, Yunai Yi, Gengyu Ma, Tae-Kyun Kim 0001, Yuke Guo, Tianmin Xu, Hongbin Zha
IEEE Trans. Medical Imaging4
2017 Kinematic-Layout-aware Random Forests for Depth-based Action Recognition
Seungryul Baek, Zhiyuan Shi 0001, Masato Kawade, Tae-Kyun Kim 0001
BMVC4
2017 Transition Forests: Learning Discriminative Temporal Transitions for Action Recognition and Detection
abstract
A human action can be seen as transitions between ones body poses over time, where the transition depicts a temporal relation between two poses. Recognizing actions thus involves learning a classifier sensitive to these pose transitions as well as to static poses. In this paper, we introduce a novel method called transitions forests, an ensemble of decision trees that both learn to discriminate static poses and transitions between pairs of two independent frames. During training, node splitting is driven by alternating two criteria: the standard classification objective that maximizes the discrimination power in individual frames, and the proposed one in pairwise frame transitions. Growing the trees tends to group frames that have similar associated transitions and share same action label incorporating temporal information that was not available otherwise. Unlike conventional decision trees where the best split in a node is determined independently of other nodes, the transition forests try to find the best split of nodes jointly (within a layer) for incorporating distant node transitions. When inferring the class label of a new frame, it is passed down the trees and the prediction is made based on previous frame predictions and the current one in an efficient and online manner. We apply our method on varied skeleton action recognition and online detection datasets showing its suitability over several baselines and state-of-the-art approaches.
Guillermo Garcia-Hernando, Tae-Kyun Kim 0001
CVPR2
2017 Learning and Refining of Privileged Information-Based RNNs for Action Recognition from Depth Sequences
abstract
Existing RNN-based approaches for action recognition from depth sequences require either skeleton joints or hand-crafted depth features as inputs. An end-to-end manner, mapping from raw depth maps to action classes, is non-trivial to design due to the fact that: 1) single channel map lacks texture thus weakens the discriminative power, 2) relatively small set of depth training data. To address these challenges, we propose to learn an RNN driven by privileged information (PI) in three-steps: An encoder is pre-trained to learn a joint embedding of depth appearance and PI (i.e. skeleton joints). The learned embedding layers are then tuned in the learning step, aiming to optimize the network by exploiting PI in a form of multi-task loss. However, exploiting PI as a secondary task provides little help to improve the performance of a primary task (i.e. classification) due to the gap between them. Finally, a bridging matrix is defined to connect two tasks by discovering latent PI in the refining step. Our PI-based classification loss maintains a consistency between latent PI and predicted distribution. The latent PI and network are iteratively estimated and updated in an expectation-maximization procedure. The proposed learning process provides greater discriminative power to model subtle depth difference, while helping avoid overfitting the scarcer training data. Our experiments show significant performance gains over state-of-the-art methods on three public benchmark datasets and our newly collected Blanket dataset.
Zhiyuan Shi 0001, Tae-Kyun Kim 0001
CVPR2
2017 BigHand2.2M Benchmark: Hand Pose Dataset and State of the Art Analysis
abstract
In this paper we introduce a large-scale hand pose dataset, collected using a novel capture method. Existing datasets are either generated synthetically or captured using depth sensors: synthetic datasets exhibit a certain level of appearance difference from real depth images, and real datasets are limited in quantity and coverage, mainly due to the difficulty to annotate them. We propose a tracking system with six 6D magnetic sensors and inverse kinematics to automatically obtain 21-joints hand pose annotations of depth maps captured with minimal restriction on the range of motion. The capture protocol aims to fully cover the natural hand pose space. As shown in embedding plots, the new dataset exhibits a significantly wider and denser range of hand poses compared to existing benchmarks. Current state-of-the-art methods are evaluated on the dataset, and we demonstrate significant improvements in cross-benchmark performance. We also show significant improvements in egocentric hand pose estimation with a CNN trained on the new dataset.
Shanxin Yuan, Qi Ye 0001, Björn Stenger, Siddhant Jain, Tae-Kyun Kim 0001
CVPR5
2017 Pose Guided RGBD Feature Learning for 3D Object Pose Estimation
abstract
In this paper we examine the effects of using object poses as guidance to learning robust features for 3D object pose estimation. Previous works have focused on learning feature embeddings based on metric learning with triplet comparisons and rely only on the qualitative distinction of similar and dissimilar pose labels. In contrast, we consider the exact pose differences between the training samples, and aim to learn embeddings such that the distances in the pose label space are proportional to the distances in the feature space. However, since it is less desirable to force the pose-feature correlation when objects are symmetric, we discuss the use of weights that reflect object symmetry when measuring the pose distances. Furthermore, end-to-end pose regression is investigated and is shown to further boost the discriminative power of feature learning, improving pose recognition accuracies. Experimental results show that the features that are learnt guided by poses, are significantly more discriminative than the ones learned in the traditional way, outperforming state-of-the-art works. Finally, we measure the generalisation capacity of pose guided feature learning in previously unseen scenes containing objects under different occlusion levels, and we show that it adapts well to novel tasks.
Vassileios Balntas, Andreas Doumanoglou, Caner Sahin, Juil Sock, Rigas Kouskouridas, Tae-Kyun Kim 0001
ICCV6
2017 Real-Time Online Action Detection Forests Using Spatio-Temporal Contexts
abstract
Online action detection (OAD) is challenging since 1) robust yet computationally expensive features cannot be straightforwardly used due to the real-time processing requirements and 2) the localization and classification of actions have to be performed even before they are fully observed. We propose a new random forest (RF)-based online action detection framework that addresses these challenges. Our algorithm uses computationally efficient skeletal joint features. High accuracy is achieved by using robust convolutional neural network (CNN)-based features which are extracted from the raw RGBD images, plus the temporal relationships between the current frame of interest, and the past and futures frames. While these high-quality features are not available in real-time testing scenario, we demonstrate that they can be effectively exploited in training RF classifiers: We use these spatio-temporal contexts to craft RF's new split functions improving RFs' leaf node statistics. Experiments with challenging MSRAction3D, G3D, and OAD datasets demonstrate that our algorithm significantly improves the accuracy over the state-of-the-art on-line action detection algorithms while achieving the real-time efficiency of existing skeleton-based RF classifiers.
Seungryul Baek, Kwang In Kim, Tae-Kyun Kim 0001
WACV3
2017 A learning-based variable size part extraction architecture for 6D object pose recovery in depth images
Caner Sahin, Rigas Kouskouridas, Tae-Kyun Kim 0001
Image Vis. Comput.3
2017 Near-lighting Photometric Stereo for unknown scene distance and medium attenuation
Chourmouzios Tsiotsios, Andrew J. Davison, Tae-Kyun Kim 0001
Image Vis. Comput.3
2017 Latent Regression Forest: Structured Estimation of 3D Hand Poses
abstract
In this paper we present the latent regression forest (LRF), a novel framework for real-time, 3D hand pose estimation from a single depth image. Prior discriminative methods often fall into two categories: holistic and patch-based. Holistic methods are efficient but less flexible due to their nearest neighbour nature. Patch-based methods can generalise to unseen samples by consider local appearance only. However, they are complex because each pixel need to be classified or regressed during testing. In contrast to these two baselines, our method can be considered as a structured coarse-to-fine search, starting from the centre of mass of a point cloud until locating all the skeletal joints. The searching process is guided by a learnt latent tree model which reflects the hierarchical topology of the hand. Our main contributions can be summarised as follows: (i) Learning the topology of the hand in an unsupervised, data-driven manner. (ii) A new forest-based, discriminative framework for structured search in images, as well as an error regression step to avoid error accumulation. (iii) A new multi-view hand pose dataset containing 180 K annotated images from 10 different subjects. Our experiments on two datasets show that the LRF outperforms baselines and prior arts in both accuracy and efficiency.
Danhang Tang, Hyung Jin Chang, Alykhan Tejani, Tae-Kyun Kim 0001
IEEE Trans. Pattern Anal. Mach. Intell.4
2017 Metaphoric Hand Gestures for Orientation-Aware VR Object Manipulation With an Egocentric Viewpoint
abstract
We present a novel natural user interface framework, called Meta-Gesture, for selecting and manipulating rotatable virtual reality (VR) objects in egocentric viewpoint. Meta-Gesture uses the gestures of holding and manipulating the tools of daily use. Specifically, the holding gesture is used to summon a virtual object into the palm, and the manipulating gesture to trigger the function of the summoned virtual tool. Our contributions are broadly threefold: 1) Meta-Gesture is the first to perform bare hand-gesture-based orientation-aware selection and manipulation of very small (nail-sized) VR objects, which has become possible by combining a stable 3-D palm pose estimator (publicly available) with the proposed static-dynamic (SD) gesture estimator; 2) the proposed novel SD random forest, as an SD gesture estimator can classify a 3-D static gesture and its action status hierarchically, in a single classifier; and 3) our novel voxel coding scheme, called layered shape pattern, which is configured by calculating the fill rate of point clouds (raw source of data) in each voxel on the top of the palm pose estimation, allows for dispensing with the need for preceding hand skeletal tracking or joint classification while defining a gesture. Experimental results show that the proposed method can deliver promising performance, even under frequent occlusions, during orientation-aware selection and manipulation of objects in VR space by wearing head-mounted display with an attached egocentric-depth camera (see the supplementary video available at: http://ieeexplore.ieee.org).
Youngkyoon Jang, Ikbeom Jeon, Tae-Kyun Kim 0001, Woontack Woo
IEEE Trans. Hum. Mach. Syst.3
2016 Recovering 6D Object Pose and Predicting Next-Best-View in the Crowd
abstract
Object detection and 6D pose estimation in the crowd (scenes with multiple object instances, severe foreground occlusions and background distractors), has become an important problem in many rapidly evolving technological areas such as robotics and augmented reality. Single shotbased 6D pose estimators with manually designed features are still unable to tackle the above challenges, motivating the research towards unsupervised feature learning and next-best-view estimation. In this work, we present a complete framework for both single shot-based 6D object pose estimation and next-best-view prediction based on Hough Forests, the state of the art object pose estimator that performs classification and regression jointly. Rather than using manually designed features we a) propose an unsupervised feature learnt from depth-invariant patches using a Sparse Autoencoder and b) offer an extensive evaluation of various state of the art features. Furthermore, taking advantage of the clustering performed in the leaf nodes of Hough Forests, we learn to estimate the reduction of uncertainty in other views, formulating the problem of selecting the next-best-view. To further improve pose estimation, we propose an improved joint registration and hypotheses verification module as a final refinement step to reject false detections. We provide two additional challenging datasets inspired from realistic scenarios to extensively evaluate the state of the art and our framework. One is related to domestic environments and the other depicts a bin-picking scenario mostly found in industrial settings. We show that our framework significantly outperforms state of the art both on public and on our datasets.
Andreas Doumanoglou, Rigas Kouskouridas, Sotiris Malassiotis, Tae-Kyun Kim 0001
CVPR4
2016 Spatial Attention Deep Net with Partial PSO for Hierarchical Hybrid Hand Pose Estimation
Qi Ye 0001, Shanxin Yuan, Tae-Kyun Kim 0001
ECCV (8)3
2016 Iterative Hough Forest with Histogram of Control Points for 6 DoF object registration from depth images
abstract
State-of-the-art techniques proposed for 6D object pose recovery depend on occlusion-free point clouds to accurately register objects in 3D space. To reduce this dependency, we introduce a novel architecture called Iterative Hough Forest with Histogram of Control Points that is capable of estimating occluded and cluttered objects' 6D pose given a candidate 2D bounding box. Our Iterative Hough Forest is learnt using patches extracted only from the positive samples. These patches are represented with Histogram of Control Points (HoCP), a “scale-variant” implicit volumetric description, which we derive from recently introduced Implicit B-Splines (IBS). The rich discriminative information provided by this scale-variance is leveraged during inference, where the initial pose estimation of the object is iteratively refined based on more discriminative control points by using our Iterative Hough Forest. We conduct experiments on several test objects of a publicly available dataset to test our architecture and to compare with the state-of-the-art.
Caner Sahin, Rigas Kouskouridas, Tae-Kyun Kim 0001
IROS3
2016 Transition Hough forest for trajectory-based action recognition
abstract
In this paper, we propose a new discriminative framework based on Hough forests that enables us to efficiently recognize and localize sequential data in the form of spatio-temporal trajectories. Contrary to traditional decision forest-based methods where predictions are made independently of its output temporal context, we introduce the concept of "transition", which enforces the temporal coherence of estimations and further enhances the discrimination between action classes. We start applying our proposed framework to the problem of recognizing and localizing fingertip written trajectories in mid-air using an egocentric camera. To this purpose, we present a new challenging dataset that allows us to evaluate and compare our method with previous approaches. Finally, we apply our framework to general human action recognition using local spatio-temporal trajectories obtaining comparable to state-of-the-art performance on a public benchmark.
Guillermo Garcia-Hernando, Hyung Jin Chang, Ismael Serrano, Oscar Déniz-Suárez, Tae-Kyun Kim 0001
WACV5
2016 Spatio-Temporal Hough Forest for efficient detection-localisation-recognition of fingerwriting in egocentric camera
Hyung Jin Chang, Guillermo Garcia-Hernando, Danhang Tang, Tae-Kyun Kim 0001
Comput. Vis. Image Underst.4
2016 Model effectiveness prediction and system adaptation for photometric stereo in murky water
Chourmouzios Tsiotsios, Tae-Kyun Kim 0001, Andrew J. Davison, Srinivasa G. Narasimhan
Comput. Vis. Image Underst.2
2016 A comparative study of video-based object recognition from an egocentric viewpoint
Mang Shao, Danhang Tang, Yang Liu 0020, Tae-Kyun Kim 0001
Neurocomputing4
2016 Convolutional Fusion Network for Face Verification in the Wild
abstract
Part-based methods have seen popular applications for face verification in the wild, since they are more robust to local variations in terms of pose, illumination, and so on. However, most of the part-based approaches are built on hand-crafted features, which may not be suitable for the specific face verification purpose. In this paper, we propose to learn a part-based feature representation under the supervision of face identities through a deep model that ensures that the generated representations are more robust and suitable for face verification. The proposed framework consists of the following two deliberate components: 1) a deep mixture model (DMM) to find accurate patch correspondence and 2) a convolutional fusion network (CFN) to extract the part-based facial features. Specifically, DMM robustly depicts the spatial-appearance distribution of patch features over the faces via several Gaussian mixtures, which provide more accurate patch correspondence even in the presence of local distortions. Then, DMM only feeds the patches which preserve the identity information to the following CFN. The proposed CFN is a two-layer cascade of convolutional neural networks: 1) a local layer built on face patches to deal with local variations and 2) a fusion layer integrating the responses from the local layer. CFN jointly learns and fuses multiple local responses to optimize the verification performance. The composite representation obtained possesses certain robustness to pose and illumination variations and shows comparable performance with the state-of-the-art methods on two benchmark data sets.
Luoqi Liu, Shuicheng Yan, Tae-Kyun Kim 0001
IEEE Trans. Circuits Syst. Video Technol.5
2016 Folding Clothes Autonomously: A Complete Pipeline
abstract
This work presents a complete pipeline for folding a pile of clothes using a dual-armed robot. This is a challenging task both from the viewpoint of machine vision and robotic manipulation. The presented pipeline is comprised of the following parts: isolating and picking up a single garment from a pile of crumpled garments, recognizing its category, unfolding the garment using a series of manipulations performed in the air, placing the garment roughly flat on a work table, spreading it, and, finally, folding it in several steps. The pile is segmented into separate garments using color and texture information, and the ideal grasping point is selected based on the features computed from a depth map. The recognition and unfolding of the hanging garment are performed in an active manner, utilizing the framework of active random forests to detect grasp points, while optimizing the robot actions. The spreading procedure is based on the detection of deformations of the garment's contour. The perception for folding employs fitting of polygonal models to the contour of the observed garment, both spread and already partially folded. We have conducted several experiments on the full pipeline producing very promising results. To our knowledge, this is the first work addressing the complete unfolding and folding pipeline on a variety of garments, including T-shirts, towels, and shorts.
Andreas Doumanoglou, Jan Stria, Georgia Peleka, Ioannis Mariolis, Vladimír Petrík, Andreas Kargakos, Libor Wagner, Václav Hlavác, Tae-Kyun Kim 0001, Sotiris Malassiotis
IEEE Trans. Robotics9
2015 Automatic Topic Discovery for Multi-Object Tracking
abstract
This paper proposes a new approach to multi-object tracking by semantic topic discovery. We dynamically cluster frame-by-frame detections and treat objects as topics, allowing the application of the Dirichlet Process Mixture Model (DPMM). The tracking problem is cast as a topic-discovery task where the video sequence is treated analogously to a document. This formulation addresses tracking issues such as object exclusivity constraints as well as cannot-link constraints which are integrated without the need for heuristic thresholds. The video is temporally segmented into epochs to model the dynamics of word (superpixel) co-occurrences and to model the temporal damping effect. In experiments on public data sets we demonstrate the effectiveness of the proposed algorithm.
Wenhan Luo, Björn Stenger, Tae-Kyun Kim 0001
AAAI4
2015 Opening the Black Box: Hierarchical Sampling Optimization for Estimating Human Hand Pose
abstract
We address the problem of hand pose estimation, formulated as an inverse problem. Typical approaches optimize an energy function over pose parameters using a 'black box' image generation procedure. This procedure knows little about either the relationships between the parameters or the form of the energy function. In this paper, we show that we can significantly improving upon black box optimization by exploiting high-level knowledge of the structure of the parameters and using a local surrogate energy function. Our new framework, hierarchical sampling optimization, consists of a sequence of predictors organized into a kinematic hierarchy. Each predictor is conditioned on its ancestors, and generates a set of samples over a subset of the pose parameters. The highly-efficient surrogate energy is used to select among samples. Having evaluated the full hierarchy, the partial pose samples are concatenated to generate a full-pose hypothesis. Several hypotheses are generated using the same procedure, and finally the original full energy function selects the best result. Experimental evaluation on three publically available datasets show that our method is particularly impressive in low-compute scenarios where it significantly outperforms all other state-of-the-art methods.
Danhang Tang, Jonathan Taylor 0001, Pushmeet Kohli, Cem Keskin, Tae-Kyun Kim 0001, Jamie Shotton
ICCV5
2015 Conditional Convolutional Neural Network for Modality-Aware Face Recognition
abstract
Faces in the wild are usually captured with various poses, illuminations and occlusions, and thus inherently multimodally distributed in many tasks. We propose a conditional Convolutional Neural Network, named as c-CNN, to handle multimodal face recognition. Different from traditional CNN that adopts fixed convolution kernels, samples in c-CNN are processed with dynamically activated sets of kernels. In particular, convolution kernels within each layer are only sparsely activated when a sample is passed through the network. For a given sample, the activations of convolution kernels in a certain layer are conditioned on its present intermediate representation and the activation status in the lower layers. The activated kernels across layers define the sample-specific adaptive routes that reveal the distribution of underlying modalities. Consequently, the proposed framework does not rely on any prior knowledge of modalities in contrast with most existing methods. To substantiate the generic framework, we introduce a special case of c-CNN via incorporating the conditional routing of the decision tree, which is evaluated with two problems of multimodality - multi-view face identification and occluded face verification. Extensive experiments demonstrate consistent improvements over the counterparts unaware of modalities.
Danhang Tang, Jayashree Karlekar, Shuicheng Yan, Tae-Kyun Kim 0001
ICCV6
2015 STARE: Spatio-Temporal Attention Relocation for Multiple Structured Activities Detection
abstract
We present a spatio-temporal attention relocation (STARE) method, an information-theoretic approach for efficient detection of simultaneously occurring structured activities. Given multiple human activities in a scene, our method dynamically focuses on the currently most informative activity. Each activity can be detected without complete observation, as the structure of sequential actions plays an important role on making the system robust to unattended observations. For such systems, the ability to decide where and when to focus is crucial to achieving high detection performances under resource bounded condition. Our main contributions can be summarized as follows: 1) information-theoretic dynamic attention relocation framework that allows the detection of multiple activities efficiently by exploiting the activity structure information and 2) a new high-resolution data set of temporally-structured concurrent activities. Our experiments on applications show that the STARE method performs efficiently while maintaining a reasonable level of accuracy.
Kyuhwa Lee, Dimitri Ognibene, Hyung Jin Chang, Tae-Kyun Kim 0001, Yiannis Demiris
IEEE Trans. Image Process.4
2015 3D Finger CAPE: Clicking Action and Position Estimation under Self-Occlusions in Egocentric Viewpoint
abstract
In this paper we present a novel framework for simultaneous detection of click action and estimation of occluded fingertip positions from egocentric viewed single-depth image sequences. For the detection and estimation, a novel probabilistic inference based on knowledge priors of clicking motion and clicked position is presented. Based on the detection and estimation results, we were able to achieve a fine resolution level of a bare hand-based interaction with virtual objects in egocentric viewpoint. Our contributions include: (i) a rotation and translation invariant finger clicking action and position estimation using the combination of 2D image-based fingertip detection with 3D hand posture estimation in egocentric viewpoint. (ii) a novel spatio-temporal random forest, which performs the detection and estimation efficiently in a single framework. We also present (iii) a selection process utilizing the proposed clicking action detection and position estimation in an arm reachable AR/VR space, which does not require any additional device. Experimental results show that the proposed method delivers promising performance under frequent self-occlusions in the process of selecting objects in AR/VR space whilst wearing an egocentric-depth camera-attached HMD.
Youngkyoon Jang, Seungtak Noh, Hyung Jin Chang, Tae-Kyun Kim 0001, Woontack Woo
IEEE Trans. Vis. Comput. Graph.4
2014 Bi-label Propagation for Generic Multiple Object Tracking
abstract
In this paper, we propose a label propagation framework to handle the multiple object tracking (MOT) problem for a generic object type (cf. pedestrian tracking). Given a target object by an initial bounding box, all objects of the same type are localized together with their identities. We treat this as a problem of propagating bi-labels, i.e. a binary class label for detection and individual object labels for tracking. To propagate the class label, we adopt clustered Multiple Task Learning (cMTL) while enforcing spatio-temporal consistency and show that this improves the performance when given limited training data. To track objects, we propagate labels from trajectories to detections based on affinity using appearance, motion, and context. Experiments on public and challenging new sequences show that the proposed method improves over the current state of the art on this task.
Wenhan Luo, Tae-Kyun Kim 0001, Björn Stenger, Roberto Cipolla
CVPR2
2014 Latent Regression Forest: Structured Estimation of 3D Articulated Hand Posture
abstract
In this paper we present the Latent Regression Forest (LRF), a novel framework for real-time, 3D hand pose estimation from a single depth image. In contrast to prior forest-based methods, which take dense pixels as input, classify them independently and then estimate joint positions afterwards, our method can be considered as a structured coarse-to-fine search, starting from the centre of mass of a point cloud until locating all the skeletal joints. The searching process is guided by a learnt Latent Tree Model which reflects the hierarchical topology of the hand. Our main contributions can be summarised as follows: (i) Learning the topology of the hand in an unsupervised, data-driven manner. (ii) A new forest-based, discriminative framework for structured search in images, as well as an error regression step to avoid error accumulation. (iii) A new multi-view hand pose dataset containing 180K annotated images from 10 different subjects. Our experiments show that the LRF out-performs state-of-the-art methods in both accuracy and efficiency.
Danhang Tang, Hyung Jin Chang, Alykhan Tejani, Tae-Kyun Kim 0001
CVPR4
2014 Backscatter Compensated Photometric Stereo with 3 Sources
abstract
Photometric stereo offers the possibility of object shape reconstruction via reasoning about the amount of light reflected from oriented surfaces. However, in murky media such as sea water, the illuminating light interacts with the medium and some of it is backscattered towards the camera. Due to this additive light component, the standard Photometric Stereo equations lead to poor quality shape estimation. Previous authors have attempted to reformulate the approach but have either neglected backscatter entirely or disregarded its non-uniformity on the sensor when camera and lights are close to each other. We show that by compensating effectively for the backscatter component, a linear formulation of Photometric Stereo is allowed which recovers an accurate normal map using only 3 lights. Our backscatter compensation method for point-sources can be used for estimating the uneven backscatter directly from single images without any prior knowledge about the characteristics of the medium or the scene. We compare our method with previous approaches through extensive experimental results, where a variety of objects are imaged in a big water tank whose turbidity is systematically increased, and show reconstruction quality which degrades little relative to clean water results even with a very significant scattering level.
Chourmouzios Tsiotsios, Maria E. Angelopoulou, Tae-Kyun Kim 0001, Andrew J. Davison
CVPR3
2014 Unified Face Analysis by Iterative Multi-output Random Forests
abstract
In this paper, we present a unified method for joint face image analysis, i.e., simultaneously estimating head pose, facial expression and landmark positions in real-world face images. To achieve this goal, we propose a novel iterative Multi-Output Random Forests (iMORF) algorithm, which explicitly models the relations among multiple tasks and iteratively exploits such relations to boost the performance of all tasks. Specifically, a hierarchical face analysis forest is learned to perform classification of pose and expression at the top level, while performing landmark positions regression at the bottom level. On one hand, the estimated pose and expression provide strong shape prior to constrain the variation of landmark positions. On the other hand, more discriminative shape-related features could be extracted from the estimated landmark positions to further improve the predictions of pose and expression. This relatedness of face analysis tasks is iteratively exploited through several cascaded hierarchical face analysis forests until convergence. Experiments conducted on publicly available real-world face datasets demonstrate that the performance of all individual tasks are significantly improved by the proposed iMORF algorithm. In addition, our method outperforms state-of-the-arts for all three face analysis tasks.
Tae-Kyun Kim 0001, Wenhan Luo
CVPR2
2014 Active Random Forests: An Application to Autonomous Unfolding of Clothes
Andreas Doumanoglou, Tae-Kyun Kim 0001, Sotiris Malassiotis
ECCV (5)2
2014 Latent-Class Hough Forests for 3D Object Detection and Pose Estimation
Alykhan Tejani, Danhang Tang, Rigas Kouskouridas, Tae-Kyun Kim 0001
ECCV (6)4
2014 Enhanced Random Forest with Image/Patch-Level Learning for Image Understanding
abstract
Image understanding is an important research domain in the computer vision due to its wide real-world applications. For an image understanding framework that uses the Bag-of-Words model representation, the visual codebook is an essential part. Random forest (RF) as a tree-structure discriminative codebook has been a popular choice. However, the performance of the RF can be degraded if the local patch labels are poorly assigned. In this paper, we tackle this problem by a novel way to update the RF codebook learning for a more discriminative codebook with the introduction of the soft class labels, estimated from the pLSA model based on a feedback scheme. The feedback scheme is performed on both the image and patch levels respectively, which is in contrast to the state-of-the-art RF codebook learning that focused on either image or patch level only. Experiments on 15-Scene and C-Pascal datasets had shown the effectiveness of the proposed method in image understanding task.
Wai Lam Hoo, Tae-Kyun Kim 0001, Yuru Pei, Chee Seng Chan
ICPR2
2014 Autonomous active recognition and unfolding of clothes using random decision forests and probabilistic planning
abstract
We present a novel approach to the problem of autonomously recognizing and unfolding articles of clothing using a dual manipulator. The problem consists of grasping an article from a random point, recognizing it and then bringing it into an unfolded state. We propose a data-driven method for clothes recognition from depth images using Random Decision Forests. We also propose a method for unfolding an article of clothing after estimating and grasping two key-points, using Hough forests. Both methods are implemented into a POMDP framework allowing the robot to interact optimally with the garments, taking into account uncertainty in the recognition and point estimation process. This active recognition and unfolding makes our system very robust to noisy observations. Our methods were tested on regular-sized clothes using a dual-arm manipulator and an Xtion depth sensor. We achieved 100% accuracy in active recognition and 93.3% unfolding success rate, while our system operates faster compared to the state of the art.
Andreas Doumanoglou, Andreas Kargakos, Tae-Kyun Kim 0001, Sotiris Malassiotis
ICRA3
2014 Adaptive Learning for Celebrity Identification With Video Context
abstract
In this paper, we propose a novel semi-supervised learning strategy to address the problem of celebrity identification. The video context information is explored to facilitate the learning process based on the assumption that faces in the same video track share the same identity. Once a frame within a track is recognized confidently, the label can be propagated through the whole track, referred to as the confident track. More specifically, given a few static images and vast face videos, an initial weak classifier is trained and gradually evolves by iteratively promoting the confident tracks into the “labeled” set. The iterative selection process enriches the diversity of the “labeled” set such that the performance of the classifier is gradually improved. This learning theme may suffer from semantic drifting caused by errors in selecting the confident tracks. To address this issue, we propose to treat the selected frames as related samples-an intermediate state between labeled and unlabeled instead of labeled as in the traditional approach. To evaluate the performance, we construct a new dataset, which includes 3000 static images and 2700 face tracks of 30 celebrities. Comprehensive evaluations on this dataset and a public video dataset indicate significant improvement of our approach over established baseline methods.
Guangyu Gao, Zhengjun Zha, Shuicheng Yan, Huadong Ma, Tae-Kyun Kim 0001
IEEE Trans. Multim.6
2013 Generic Object Crowd Tracking by Multi-Task Learning
abstract
We address Multiple Object Tracking (MOT) in crowds, where the type of target objects is generic and not limited to pedestrians as in most previous work. Following the popular tracking-by-detection strategy, we decompose this problem into two main tasks, detection and tracking, and formulate them under the Multiple Task Learning (MTL) framework. A binary detector is learnt to detect objects in images, whilst multiple trackers are learnt on top of the detector by MTL to trace detected objects in subsequent frames. The detector is utilised to anchor the trackers, helping them not drift away from targets. The trackers are jointly learnt by sharing common features. To further improve the performance, we use a smoothness term which considers all labelled and unlabelled data globally. Experiments on challenging new generic object sequences as well as a publicly available sequence show that the proposed method significantly outperforms the state-of-the-art methods.
Wenhan Luo, Tae-Kyun Kim 0001
BMVC2
2013 Unconstrained Monocular 3D Human Pose Estimation by Action Detection and Cross-Modality Regression Forest
abstract
This work addresses the challenging problem of unconstrained 3D human pose estimation (HPE) from a novel perspective. Existing approaches struggle to operate in realistic applications, mainly due to their scene-dependent priors, such as background segmentation and multi-camera network, which restrict their use in unconstrained environments. We therfore present a framework which applies action detection and 2D pose estimation techniques to infer 3D poses in an unconstrained video. Action detection offers spatiotemporal priors to 3D human pose estimation by both recognising and localising actions in space-time. Instead of holistic features, e.g. silhouettes, we leverage the flexibility of deformable part model to detect 2D body parts as a feature to estimate 3D poses. A new unconstrained pose dataset has been collected to justify the feasibility of our method, which demonstrated promising results, significantly outperforming the relevant state-of-the-arts.
Tsz-Ho Yu, Tae-Kyun Kim 0001, Roberto Cipolla
CVPR2
2013 Unsupervised Random Forest Manifold Alignment for Lipreading
abstract
Lip reading from visual channels remains a challenging topic considering the various speaking characteristics. In this paper, we address an efficient lip reading approach by investigating the unsupervised random forest manifold alignment (RFMA). The density random forest is employed to estimate affinity of patch trajectories in speaking facial videos. We propose novel criteria for node splitting to avoid the rank-deficiency in learning density forests. By virtue of the hierarchical structure of random forests, the trajectory affinities are measured efficiently, which are used to find embeddings of the speaking video clips by a graph-based algorithm. Lip reading is formulated as matching between manifolds of query and reference video clips. We employ the manifold alignment technique for matching, where the L∞-norm-based manifold-to-manifold distance is proposed to find the matching pairs. We apply this random forest manifold alignment technique to various video data sets captured by consumer cameras. The experiments demonstrate that lip reading can be performed effectively, and outperform state-of-the-arts.
Yuru Pei, Tae-Kyun Kim 0001, Hongbin Zha
ICCV2
2013 Real-Time Articulated Hand Pose Estimation Using Semi-supervised Transductive Regression Forests
abstract
This paper presents the first semi-supervised transductive algorithm for real-time articulated hand pose estimation. Noisy data and occlusions are the major challenges of articulated hand pose estimation. In addition, the discrepancies among realistic and synthetic pose data undermine the performances of existing approaches that use synthetic data extensively in training. We therefore propose the Semi-supervised Transductive Regression (STR) forest which learns the relationship between a small, sparsely labelled realistic dataset and a large synthetic dataset. We also design a novel data-driven, pseudo-kinematic technique to refine noisy or occluded joints. Our contributions include: (i) capturing the benefits of both realistic and synthetic data via transductive learning, (ii) showing accuracies can be improved by considering unlabelled data, and (iii) introducing a pseudo-kinematic technique to refine articulations efficiently. Experimental results show not only the promising performance of our method with respect to noise and occlusions, but also its superiority over state-of-the-arts in accuracy, robustness and speed.
Danhang Tang, Tsz-Ho Yu, Tae-Kyun Kim 0001
ICCV3
2012 Fast Pedestrian Detection by Cascaded Random Forest with Dominant Orientation Templates
abstract
In this paper, we present a new pedestrian detection method combining Random Forest and Dominant Orientation Templates(DOT) to achieve state-of-the-art accuracy and, more importantly, to accelerate run-time speed. DOT can be considered as a binary version of Histogram of Oriented Gradients(HOG) and therefore provides time-efficient properties. However, since discarding magnitude information, it degrades the detection rate, when it is directly incorporated. We propose a novel template-matching split function using DOT for Random Forest. It divides a feature space in a non-linear manner, but has a very low complexity up to binary bit-wise operations. Experiments demonstrate that our method provides much superior speed with comparable accuracy to state-ofthe-art pedestrian detectors. By combining a holistic and a patch-based detectors in a cascade manner, we accelerate the detection speed of Hough Forest, a prior-art using Random Forest and HOG, by about 20 times. The obtained speed is 5 frames per second for 640×480 images with 24 scales.
Danhang Tang, Yang Liu 0020, Tae-Kyun Kim 0001
BMVC3
2012 Set-based label propagation of face images
abstract
Graph-based Semi-Supervised Learning (SSL) has proven to be an effective tool for label propagation, however, its accuracy is highly dependent on how to form the data weight matrix, in which each element is obtained as the similarity between every pair of data points. Inspired by the success of set-based recognition methods, a novel approach is brought up to incorporate the set-to-set matching as well as single-to-single matching when building up the weight matrix. Canonical Correlation Analysis (CCA), which measures the principal angles between two manifolds, is adopted to compute the set similarity. Moreover, Local Binary Pattern, an effective texture descriptor, is investigated as a data representation to further improve the label propagation performance. The proposed approach is evaluated on two public face image data sets, and shown to significantly outperform the standard SSL methods in terms of accuracy.
Tae-Kyun Kim 0001
ICIP2
2012 Learning action symbols for hierarchical grammar induction
Kyuhwa Lee, Tae-Kyun Kim 0001, Yiannis Demiris
ICPR2
2012 Learning reusable task components using hierarchical activity grammars with uncertainties
abstract
We present a novel learning method using activity grammars capable of learning reusable task components from a reasonably small number of samples under noisy conditions. Our linguistic approach aims to extract the hierarchical structure of activities which can be recursively applied to help recognize unforeseen, more complicated tasks that share the same underlying structures. To achieve this goal, our method 1) actively searches for frequently occurring action symbols that are subset of input samples to effectively discover the hierarchy, and 2) explicitly takes into account the uncertainty values associated with input symbols due to the noise inherent in low-level detectors. In addition to experimenting with a synthetic dataset to systematically analyze the algorithm's performance, we apply our method in human-led imitation learning environment where a robot learns reusable components of the task from short demonstrations to correctly imitate more complicated, longer demonstrations of the same task category. The results suggest that under reasonable amount of noise, our method is capable to capture the reusable structures of tasks and generalize to cope with recursions.
Kyuhwa Lee, Tae-Kyun Kim 0001, Yiannis Demiris
ICRA2
2012 Making a Shallow Network Deep: Conversion of a Boosting Classifier into a Decision Tree by Boolean Optimisation
Tae-Kyun Kim 0001, Ignas Budvytis, Roberto Cipolla
Int. J. Comput. Vis.1
2011 Silhouette-based object phenotype recognition using 3D shape priors
abstract
This paper tackles the novel challenging problem of 3D object phenotype recognition from a single 2D silhouette. To bridge the large pose (articulation or deformation) and camera viewpoint changes between the gallery images and query image, we propose a novel probabilistic inference algorithm based on 3D shape priors. Our approach combines both generative and discriminative learning. We use latent probabilistic generative models to capture 3D shape and pose variations from a set of 3D mesh models. Based on these 3D shape priors, we generate a large number of projections for different phenotype classes, poses, and camera viewpoints, and implement Random Forests to efficiently solve the shape and pose inference problems. By model selection in terms of the silhouette coherency between the query and the projections of 3D shapes synthesized using the galleries, we achieve the phenotype recognition result as well as a fast approximate 3D reconstruction of the query. To verify the efficacy of the proposed approach, we present new datasets which contain over 500 images of various human and shark phenotypes and motions. The experimental results clearly show the benefits of using the 3D priors in the proposed method over previous 2D-based methods.
Yu Chen 0009, Tae-Kyun Kim 0001, Roberto Cipolla
ICCV2
2011 Incremental Linear Discriminant Analysis Using Sufficient Spanning Sets and Its Applications
Tae-Kyun Kim 0001, Björn Stenger, Josef Kittler, Roberto Cipolla
Int. J. Comput. Vis.1
2010 Randomised Manifold Forests for Principal Angle-Based Face Recognition
Ujwal D. Bonde, Tae-Kyun Kim 0001, K. R. Ramakrishnan
ACCV (4)2
2010 Making a Shallow Network Deep: Growing a Tree from Decision Regions of a Boosting Classifier
abstract
This paper presents a novel way to speed up the classification time of a boosting classifier. We make the shallow (flat) network deep (hierarchical) by growing a tree from the decision regions of a given boosting classifier. This provides many short paths for speeding up and preserves the reasonably smooth decision regions of the boosting classifier for good generalisation. We express the conversion as a Boolean optimisation problem, which has been previously studied for circuit design but limited to a small number of binary variables. In this work, a novel optimisation method is proposed for several tens of variables, i.e. weak-learners of a boosting classifier. The method is then used in a two stage cascade allowing the speed-up of a boosting classifier with any larger number of weak-learners. Experiments on the synthetic and face image data sets show that the obtained tree significantly speeds up both a standard boosting classifier and Fast-exit, a prior-art for fast boosting classification, at the same accuracy. The proposed method as a general meta-algorithm is also shown useful for a boosting cascade, since it speeds up individual stage classifiers by different gains. The proposed method is further demonstrated for rapid object tracking and segmentation problems.
Tae-Kyun Kim 0001, Ignas Budvytis, Roberto Cipolla
BMVC1
2010 Real-time Action Recognition by Spatiotemporal Semantic and Structural Forests
abstract
This paper presents a novel real-time action recogniser by utilising both local appearance and structural information. Our method is able to recognise actions continuously in real-time while achieving comparably high accuracy over state-of-the-arts. Run-time speed is of vital importance in real-world action recognition systems, but existing methods seldom take computational complexity into full consideration. A class label is assigned after an entire query video is analysed, or a large lookahead is required to recognise an action. In addition, the “bag of words”(BOW) has proven effective for action recognition [5]. However, the standard BOW model ignores the spatiotemporal relationships among feature descriptors, which are useful for describing actions. Addressing these challenges, we present a novel approach for action recognition. The major contributions include the followings: Efficient Spatiotemporal Codebook Learning: We extend the use of semantic texton forests [6] (STFs) from 2D image segmentation to spatiotemporal analysis. As well as being much faster than a traditional flat codebook such as k-means clustering, STFs achieve high accuracy comparable to that of existing approaches. STFs are ensembles of random decision trees that textonise input video patches into semantic textons. Since only a small number of simple features are used to traverse the trees, STFs are extremely fast to evaluate. They also serve a powerful discriminative codebook by multiple decision trees. Figure 1 illustrates how visual codewords are generated using STFs in the proposed method. Combined Structural and Appearance Information: We propose a richer description of features, hence actions can be classified in very short video sequences. Based on [3], we introduce the pyramidal spatiotemporal relationship match (PSRM) to encapsulate both local appearance and structural information efficiently. Subsequences are sampled from an input video in short intervals (e.g. ≤ 10 frames). After spatiotemporal interest points are localised, the trained STFs assign visual codewords to the features. A set of pairwise spatiotemporal associations are designed to capture the structural relationships among features (i.e. pairwise distances along space-time axes). All possible pairs in the bag of features are analysed by the association rules and stored in the 3-D histogram. PSRM leverages the properties of semantic trees and pyramidal match kernels. Multiple pyramidal histograms are then combined to classify a query video. Figure 2 illustrates how the relationship histograms are constructed and matched using PSRM. For each tree in STFs, the threedimensional histogram is constructed according to their spatiotemporal structures (see figure 2 (left)). Its hierarchical structure offers a time efficient way to perform the pyramid match kernel [1] for codeword matching (figure 2 (right)). Enhanced Efficiency and Combined Classification: Several techniques are employed to improve the recognition speed and accuracy. A novel spatiotemporal interest point detector, called V-FAST, is designed based on the FAST 2D corners [2]. The recognition accuracy is enhanced by adaptively combining PSRM and the bag of semantic texton (BOST) method [6]: the k-means forest classifier is learned using PSRM as a matching kernel. The task of action recognition is performed separately Spatiotemporal Relationship Match of visual codewords from Semantic Texton Forest Pyramid Match Kernel is utilised to match the histograms Feature Extraction Feature Matching
Tsz-Ho Yu, Tae-Kyun Kim 0001, Roberto Cipolla
BMVC2
2010 Inferring 3D Shapes and Deformations from Single Views
Yu Chen 0009, Tae-Kyun Kim 0001, Roberto Cipolla
ECCV (3)2
2010 On-line Learning of Mutually Orthogonal Subspaces for Face Recognition by Image Sets
abstract
We address the problem of face recognition by matching image sets. Each set of face images is represented by a subspace (or linear manifold) and recognition is carried out by subspace-to-subspace matching. In this paper, 1) a new discriminative method that maximises orthogonality between subspaces is proposed. The method improves the discrimination power of the subspace angle based face recognition method by maximizing the angles between different classes. 2) We propose a method for on-line updating the discriminative subspaces as a mechanism for continuously improving recognition accuracy. 3) A further enhancement called locally orthogonal subspace method is presented to maximise the orthogonality between competing classes. Experiments using 700 face image sets have shown that the proposed method outperforms relevant prior art and effectively boosts its accuracy by online learning. It is shown that the method for online learning delivers the same solution as the batch computation at far lower computational cost and the locally orthogonal method exhibits improved accuracy. We also demonstrate the merit of the proposed face recognition method on portal scenarios of multiple biometric grand challenge.
Tae-Kyun Kim 0001, Josef Kittler, Roberto Cipolla
IEEE Trans. Image Process.1
2009 Image mosaicing via quadric surface estimation with priors for tunnel inspection
abstract
In this paper, a system which constructs a mosaic image of the tunnel surface with little distortion is presented. The tunnel surface is typically composed of a roughly cylindrical surface and protuberant regions containing objects such as pipes, pans and tunnel ridges. Since the true surface is neither planar nor quadric, existing mosaicing methods, which assume either homography or quadratic motion models, suffer from distortion. The proposed system obtains a sparse 3D model of the tunnel by multi-view reconstruction. Then, the Support Vector Machine (SVM) classifier is applied in order to separate image features lying on the cylindrical surface from those of the non-surface. The reconstructed 3D points are reprojected into images to retrieve the priors given by the SVM classifier for accurate cylindrical surface estimation. The final mosaic image is obtained by flattening the estimated textured surface onto a plane. The results suggest that the mosaic quality depends critically on the surface estimation accuracy and the proposed system is able to produce the mosaic image that preserves all physical sense, e.g. line parallelism and straightness, which is important for tunnel inspection.
Krisada Chaiyasarn, Tae-Kyun Kim 0001, Fabio Viola, Roberto Cipolla, Kenichi Soga
ICIP2
2009 Canonical Correlation Analysis of Video Volume Tensors for Action Categorization and Detection
abstract
This paper addresses a spatiotemporal pattern recognition problem. The main purpose of this study is to find a right representation and matching of action video volumes for categorization. A novel method is proposed to measure video-to-video volume similarity by extending Canonical Correlation Analysis (CCA), a principled tool to inspect linear relations between two sets of vectors, to that of two multiway data arrays (or tensors). The proposed method analyzes video volumes as inputs avoiding the difficult problem of explicit motion estimation required in traditional methods and provides a way of spatiotemporal pattern matching that is robust to intraclass variations of actions. The proposed matching is demonstrated for action classification by a simple Nearest Neighbor classifier. We, moreover, propose an automatic action detection method, which performs 3D window search over an input video with action exemplars. The search is speeded up by dynamic learning of subspaces in the proposed CCA. Experiments on a public action data set (KTH) and a self-recorded hand gesture data showed that the proposed method is significantly better than various state-of-the-art methods with respect to accuracy. Our method has low time complexity and does not require any major tuning parameters.
Tae-Kyun Kim 0001, Roberto Cipolla
IEEE Trans. Pattern Anal. Mach. Intell.1
2008 AIDIA - Adaptive Interface for Display InterAction
abstract
This paper presents a vision-based system for interaction with a display via hand pointing. An attention mechanism based on face and hand detection allows users in the camera’s field of view to take control of the interface. Face recognition is used for identification and customisation. The system allows the user to control the screen pointer by tracking their fist. On-screen items can be selected using one of four activation mechanisms. Current sample applications include browsing image and video collections as well as viewing a gallery of 3D objects. In experiments we demonstrate the performance of the vision components in challenging conditions and compare it to that of other systems. 1
Björn Stenger, Thomas Woodley, Tae-Kyun Kim 0001, Carlos Hernández 0002, Roberto Cipolla
BMVC3
2008 MCBoost: Multiple Classifier Boosting for Perceptual Co-clustering of Images and Visual Features
abstract
We present a new co-clustering problem of images and visual features. The problem involves a set of non-object images in addition to a set of object images and features to be co-clustered. Co-clustering is performed in a way of maximising discrimination of object images from non-object images, thus emphasizing discriminative features. This provides a way of obtaining perceptual joint-clusters of object images and features. We tackle the problem by simultaneously boosting multiple strong classifiers which compete for images by their expertise. Each boosting classifier is an aggregation of weak-learners, i.e. simple visual features. The obtained classifiers are useful for multi-category and multi-view object detection tasks. Experiments on a set of pedestrian images and a face data set demonstrate that the method yields intuitive image clusters with associated features and is much superior to conventional boosting classifiers in object detection tasks.
Tae-Kyun Kim 0001, Roberto Cipolla
NIPS1
2007 Gesture Recognition Under Small Sample Size
Tae-Kyun Kim 0001, Roberto Cipolla
ACCV (1)1
2007 Tensor Canonical Correlation Analysis for Action Classification
abstract
We introduce a new framework, namely tensor canonical correlation analysis (TCCA) which is an extension of classical canonical correlation analysis (CCA) to multidimensional data arrays (or tensors) and apply this for action/gesture classification in videos. By tensor CCA, joint space-time linear relationships of two video volumes are inspected to yield flexible and descriptive similarity features of the two videos. The TCCA features are combined with a discriminative feature selection scheme and a nearest neighbor classifier for action classification. In addition, we propose a time-efficient action detection method based on dynamic learning of subspaces for tensor CCA for the case that actions are not aligned in the space-time domain. The proposed method delivered significantly better accuracy and comparable detection speed over state-of-the-art methods on the KTH action data set as well as self-recorded hand gesture data sets.
Tae-Kyun Kim 0001, Shu-Fai Wong, Roberto Cipolla
CVPR1
2007 Incremental Linear Discriminant Analysis Using Sufficient Spanning Set Approximations
abstract
This paper presents a new incremental learning solution for linear discriminant analysis (LDA). We apply the concept of the sufficient spanning set approximation in each update step, i.e. for the between-class scatter matrix, the projected data matrix as well as the total scatter matrix. The algorithm yields a more general and efficient solution to incremental LDA than previous methods. It also significantly reduces the computational complexity while providing a solution which closely agrees with the batch LDA result. The proposed algorithm has a time complexity of O(Nd2) and requires O(Nd) space, where d is the reduced subspace dimension and N the data dimension. We show two applications of incremental LDA: First, the method is applied to semi-supervised learning by integrating it into an EM framework. Secondly, we apply it to the task of merging large databases which were collected during MPEG standardization for face image retrieval.
Tae-Kyun Kim 0001, Shu-Fai Wong, Björn Stenger, Josef Kittler, Roberto Cipolla
CVPR1
2007 Learning Motion Categories using both Semantic and Structural Information
abstract
Current approaches to motion category recognition typically focus on either full spatiotemporal volume analysis (holistic approach) or analysis of the content of spatiotemporal interest points (part-based approach). Holistic approaches tend to be more sensitive to noise e.g. geometric variations, while part-based approaches usually ignore structural dependencies between parts. This paper presents a novel generative model, which extends probabilistic latent semantic analysis (pLSA), to capture both semantic (content of parts) and structural (connection between parts) information for motion category recognition. The structural information learnt can also be used to infer the location of motion for the purpose of motion detection. We test our algorithm on challenging datasets involving human actions, facial expressions and hand gestures and show its performance is better than existing unsupervised methods in both tasks of motion localisation and recognition.
Shu-Fai Wong, Tae-Kyun Kim 0001, Roberto Cipolla
CVPR2
2007 Discriminative Learning and Recognition of Image Set Classes Using Canonical Correlations
abstract
We address the problem of comparing sets of images for object recognition, where the sets may represent variations in an object's appearance due to changing camera pose and lighting conditions. Canonical Correlations (also known as principal or canonical angles), which can be thought of as the angles between two d-dimensional subspaces, have recently attracted attention for image set matching. Canonical correlations offer many benefits in accuracy, efficiency, and robustness compared to the two main classical methods: parametric distribution-based and nonparametric sample-based matching of sets. Here, this is first demonstrated experimentally for reasonably sized data sets using existing methods exploiting canonical correlations. Motivated by their proven effectiveness, a novel discriminative learning method over sets is proposed for set classification. Specifically, inspired by classical Linear Discriminant Analysis (LDA), we develop a linear discriminant function that maximizes the canonical correlations of within-class sets and minimizes the canonical correlations of between-class sets. Image sets transformed by the discriminant function are then compared by the canonical correlations. Classical orthogonal subspace method (OSM) is also investigated for the similar purpose and compared with the proposed method. The proposed method is evaluated on various object recognition problems using face image sets with arbitrary motion captured under different illuminations and image sets of 500 general objects taken at different views. The method is also applied to object category recognition using ETH-80 database. The proposed method is shown to outperform the state-of-the-art methods in terms of accuracy and efficiency.
Tae-Kyun Kim 0001, Josef Kittler, Roberto Cipolla
IEEE Trans. Pattern Anal. Mach. Intell.1
2007 Boosted manifold principal angles for image set-based recognition
Tae-Kyun Kim 0001, Ognjen Arandjelovic, Roberto Cipolla
Pattern Recognit.1
2006 Incremental Learning of Locally Orthogonal Subspaces for Set-based Object Recognition
abstract
Orthogonal subspaces are effective models to represent object image sets (generally any high-dimensional vector sets). Canonical correlation analysis of the orthogonal subspaces provides a good solution to discriminate objects with sets of images. In such a recognition task involving image sets, an efficient learning over a large volume of image sets, which may be increasing over time, is important. In this paper, an incremental learning method of orthogonal subspaces is proposed by updating the principal components of the class correlation and total correlation matrices separately, yielding the same solution as the batch computation with far lower computational cost. A novel concept of local orthogonality is further proposed to cope with non-linear manifolds of data vectors and find a more optimal solution of orthogonal subspaces for a certain neighbouring object image sets. In the experiments using 700 face image sets, the locally orthogonal subspaces outperformed the orthogonal subspaces as well as relevant state-of-the-art methods in accuracy. Note that the locally orthogonal subspaces are also amenable to incremental updating due to their linear property. 1
Tae-Kyun Kim 0001, Josef Kittler, Roberto Cipolla
BMVC1
2006 Learning Discriminative Canonical Correlations for Object Recognition with Image Sets
Tae-Kyun Kim 0001, Josef Kittler, Roberto Cipolla
ECCV (3)1
2006 Design and Fusion of Pose-Invariant Face-Identification Experts
abstract
We address the problem of pose-invariant face recognition based on a single model image. To cope with novel view face images, a model of the effect of pose changes on face appearance must be available. Face images at an arbitrary pose can be mapped to a reference pose by the model yielding view-invariant representation. Such a model typically relies on dense correspondences of different view face images, which are difficult to establish in practice. Errors in the correspondences seriously degrade the accuracy of any recognizer. Therefore, we assume only the minimal possible set of correspondences, given by the corresponding eye positions. We investigate a number of approaches to pose-invariant face recognition exploiting such a minimal set of facial features correspondences. Four different methods are proposed as pose-invariant face recognition "experts" and combined in a single framework of expert fusion. Each expert explicitly or implicitly realizes the three sequential functions jointly required to capture the nonlinear manifolds of face pose changes: representation, view transformation, and class discriminative feature extraction. Within this structure, the experts are designed for diversity. We compare a design in which the three stages are sequentially optimized with two methods which employ an overall single nonlinear function learnt from different view face images. We also propose an approach exploiting a three-dimensional face data. A lookup table storing facial feature correspondences between different pose images, found by 3-D face models, is constructed. The designed experts are different in their nature owing to different sources of information and architectures used. The proposed fusion architecture of the pose-invariant face experts achieves an impressive accuracy gain by virtue of the individual experts diversity. It is experimentally shown that the individual experts outperform the classical linear discriminant analysis (LDA) method on the XM2VTS face data set consisting of about 300 face classes. Further impressive performance gains are obtained by combining the outputs of the experts using different fusion strategies
Tae-Kyun Kim 0001, Josef Kittler
IEEE Trans. Circuits Syst. Video Technol.1
2005 Learning over Sets using Boosted Manifold Principal Angles (BoMPA)
abstract
In this paper we address the problem of classifying vector sets. We motivate and introduce a novel method based on comparisons between corresponding vector subspaces. In particular, there are two main areas of novelty: (i) we extend the concept of principal angles between linear subspaces to manifolds with arbitrary nonlinearities; (ii) it is demonstrated how boosting can be used for application-optimal principal angle fusion. The strengths of the proposed method are empirically demonstrated on the task of automatic face recognition (AFR), in which it is shown to outperform state-of-the-art methods in the literature.
Tae-Kyun Kim 0001, Ognjen Arandjelovic, Roberto Cipolla
BMVC1
2005 Component-based LDA face description for image retrieval and MPEG-7 standardisation
Tae-Kyun Kim 0001, Wonjun Hwang, Josef Kittler
Image Vis. Comput.1
2005 Locally Linear Discriminant Analysis for Multimodally Distributed Classes for Face Recognition with a Single Model Image
abstract
We present a novel method of nonlinear discriminant analysis involving a set of locally linear transformations called "Locally Linear Discriminant Analysis (LLDA)." The underlying idea is that global nonlinear data structures are locally linear and local structures can be linearly aligned. Input vectors are projected into each local feature space by linear transformations found to yield locally linearly transformed classes that maximize the between-class covariance while minimizing the within-class covariance. In face recognition, linear discriminant analysis (LDA) has been widely adopted owing to its efficiency, but it does not capture nonlinear manifolds of faces which exhibit pose variations. Conventional nonlinear classification methods based on kernels such as generalized discriminant analysis (GDA) and support vector machine (SVM) have been developed to overcome the shortcomings of the linear method, but they have the drawback of high computational cost of classification and overfitting. Our method is for multiclass nonlinear discrimination and it is computationally highly efficient as compared to GDA. The method does not suffer from overfitting by virtue of the linear base structure of the solution. A novel gradient-based learning algorithm is proposed for finding the optimal set of local linear bases. The optimization does not exhibit a local-maxima problem. The transformation functions facilitate robust face recognition in a low-dimensional subspace, under pose variations, using a single model image. The classification results are given for both synthetic and real face data.
Tae-Kyun Kim 0001, Josef Kittler
IEEE Trans. Pattern Anal. Mach. Intell.1
2004 Independent component analysis in a local facial residue space for face recognition
Tae-Kyun Kim 0001, Wonjun Hwang, Josef Kittler
Pattern Recognit.1
2003 Discriminant Analysis by Locally Linear Transformations
abstract
We present a novel discriminant analysis learning method which is applicable to non-linear data structures. The method can deal with pattern classification problems which have a multi-modal distribution for each class and samples of other classes may be closer to a class than those of the class itself. Conventional linear discriminant analysis (LDA) and LDA mixture model can not solve this linearly non-separable problem. Several local linear transformations are considered to yield locally transformed classes that maximize the between-class covariance and minimize the within-class covariance. The method invloves a novel gradient based algorithm for finding the optimal set of local linear bases. It does not have a local-maxima problem and stably converges to the global maximum point. The method is computationally efficienct as compared to the previous non-linear discriminant analysis based on the kernel approach. The method does not suffer from an overfitting problem by virtue of the linear base structure of the solution. The classification results are given for both simulated data and real face data. 1
Tae-Kyun Kim 0001, Josef Kittler, Seok-Cheol Kee
BMVC1
2003 Independent Component Analysis in a Facial Local Residue Space
abstract
In this paper, we propose an ICA (Independent Component Analysis) based face recognition algorithm, which is robust to illumination and pose variation. Generally, it is well known that the first few eigenfaces represent illumination variation rather than identity. Most PCA (Principal Component Analysis)-based methods have overcome illumination variation by discarding the projection to a few leading eigenfaces. The space spanned after removing a few leading eigenfaces is called the "residual face space". We found that ICA in the residual face space provides more efficient encoding in terms of redundancy reduction and robustness to pose variation as well as illumination variation, owing to its ability to represent non-Gaussian statistics. Moreover, a face image is separated into several facial components, local spaces, and each local space is represented by the ICA bases (independent components) of its corresponding residual space. The statistical models of face images in local spaces are relatively simple and facilitate classification by a linear encoding. Various experimental results show that the accuracy of face recognition is significantly improved by the proposed method under large illumination and pose variations.
Tae-Kyun Kim 0001, Wonjun Hwang, Seok-Cheol Kee, Josef Kittler
CVPR (1)1
2003 Face description based on decomposition and combining of a facial space with LDA
abstract
We propose a method of efficient face description for facial image retrieval from a large data set. The novel descriptor is obtained by decomposing the face image into several components and then combining the component features. The decomposition combined with LDA (linear discriminant analysis) provides discriminative facial features that are less sensitive to light and pose changes. Each component is represented in its Fisher space and another LDA is then applied to compactly combine the features of the components. To enhance retrieval accuracy further, a simple pose classification and transformation technique is performed, followed by recursive matching. The experimental results obtained on the MPEG-7 data set show an impressive accuracy of our algorithm as compared with the conventional PCA/ICA/LDA methods.
Tae-Kyun Kim 0001, Wonjun Hwang, Seok-Cheol Kee, Josef Kittler
ICIP (3)1
2002 Component-based LDA Face Descriptor for Image Retrieval
abstract
We present a component-based face descriptor with LDA (Linear Discriminant Analysis) and a simple pose classification. Our algorithm has been developed to deal with face image retrieval in huge database such as those in internet environments. Such retrieval requires a compact face descriptor and an efficient recognition algorithm that is robust to variations in lighting and facial poses. Partitioning of a face image into components facilitates the development of an efficient and robust algorithm as follows. First, compensation for light and pose variations is much more easily done on individual components than on the whole image. Second, pose variation is compensated by classifying facial pose and aligning facial components. Finally, LDA is more effective at the component level which has simplified statistics than the whole image. Experimental results on MPEG-7 database show an impressive accuracy of our algorithm compared with conventional LDA methods. 1.
Tae-Kyun Kim 0001, Wonjun Hwang, Seok-Cheol Kee, Jong Ha Lee
BMVC1
2002 Hybrid and parallel face classifier based on artificial neural networks and principal component analysis
abstract
Presents a hybrid and parallel system based on artificial neural networks for a face invariant classifier and general pattern recognition problems. A set of face features is extracted by using the eigenpaxel method, which is based on principal component analysis (PCA) of a group of pixels, that is called a paxel. To classify subjects, multi-layer perceptron neural networks (NNs) are trained for each eigenpaxel. These parallel NN kernels provide sage, fast and efficient classification. To combine the results of parallel NNs, a novel judge analyzer is proposed based on bond rating classification and prediction. The proposed judge strategy can detect distinguishable face features even in arguable situations. The proposed method was evaluated on Olivetti and HongIk university (HIU) face databases and it yields a top recognition rate of 95.5% and 94.11% respectively, which are better results than the previous eigenpaxel and NN approach.
Peter V. Bazanov, Tae-Kyun Kim 0001, Seok-Cheol Kee, Sang Uk Lee
ICIP (1)2
2002 Learning a decision boundary for face detection
abstract
Describes a pattern classification approach for detecting frontal-view faces via learning a decision boundary. The classification can be achieved either by explicit estimation of density functions of two classes, face and non-face or by direct learning of a classification function (decision boundary). The latter is a more effective approach, when the number of training available examples is small, compared to the dimensionality of image space. The proposed method consists of a implicit modeling of both face and near-face classes using Independent Component Analysis (ICA), and a subsequent classification stage based on the decision boundary estimation using Support Vector Machine (SVM). Multiple nonlinear SVMs are trained for local subspaces, considering the general non-Gaussian and multi-modal characteristic of face space. This parallelization of SVMs reduces computational cost of on-line classification, since the locally trained SVM has small number of support vectors compared to the SVM trained on entire data space. We showed that the proposed algorithm is superior to the simple combination of ICA and SVM, both in accuracy and computational burden.
Tae-Kyun Kim 0001, Donggeon Kong, Sang Ryong Kim
ICIP (1)1
1979 A Method of Recognition and Representation of Korean Characters by Tree Grammars
abstract
A syntactic method is applied to the recognition of Korean characters (Hangul). Since they develop into complex characters by the sequential addition of fundamental characters under positioning rules, there are a large amount of characters and consequently there exist many similar characters. Therefore, the sequential extraction, according to the positioning rules, of fundamental characters composing Korean characters is effective for automatic recognition. As a structural analysis, a production process of fundamental characters is represented by tree grammars, and the extraction algorithm of fundamental characters and the results of computer simulation are described.
Takeshi Agui, Masayuki Nakajima 0001, Tae-Kyun Kim 0001, Eduardo T. Takahashi
IEEE Trans. Pattern Anal. Mach. Intell.3