VLDB 2026 Research / reviewers in the wild / expert
Hyung Jin Chang
dblp:96/3551
· DBLP profile ↗
94ranked-venue papers
8as first author
58since 2021 · last 2026
0000-0001-7495-9677ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 75 · 7 first-author · 47 since 2021Graphics, computer vision, multimedia, augmented reality and games · 70 · 5 first-author · 44 since 2021Systems, architecture and hardware · 6 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Force-Aware 3D Contact Modeling for Stable Grasp GenerationabstractContact-based grasp generation plays a crucial role in various applications. Recent methods typically focus on the geometric structure of objects, producing grasps with diverse hand poses and plausible contact points. However, these approaches often overlook the physical attributes of the grasp, specifically the contact force, leading to reduced stability of the grasp. In this paper, we focus on stable grasp generation using explicit contact force predictions. First, we define a force-aware contact representation by transforming the normal force value into discrete levels and encoding it using a one-hot vector. Next, we introduce force-aware stability constraints. We define the stability problem as an acceleration minimization task and explicitly relate stability with contact geometry by formulating the underlying physical constraints. Finally, we present a pose optimizer that systematically integrates our contact representation and stability constraints to enable stable grasp generation. We show that these constraints can help identify key contact points for stability which provide effective initialization and guidance for optimization towards a stable grasp. Experiments are carried out on two public benchmarks, showing that our method brings about 20% improvement in stability metrics and adapts well to novel objects. Zhuo Chen 0028, Zhongqun Zhang, Yihua Cheng, Ales Leonardis, Hyung Jin Chang |
AAAI | 5 |
| 2026 | RTGaze: Real-Time 3D-Aware Gaze Redirection from a Single ImageabstractGaze redirection methods aim to generate realistic human face images with controllable eye movement. However, recent methods often struggle with 3D consistency, efficiency, or quality, limiting their practical applications. In this work, we propose RTGaze, a real-time and high-quality gaze redirection method. Our approach learns a gaze-controllable facial representation from face images and gaze prompts, then decodes this representation via neural rendering for gaze redirection. Additionally, we distill face geometric priors from a pretrained 3D portrait generator to enhance generation quality. We evaluate RTGaze both qualitatively and quantitatively, demonstrating state-of-the-art performance in efficiency, redirection accuracy, and image quality across multiple datasets. Our system achieves real-time, 3D-aware gaze redirection with a feedforward network (~0.06 sec/image), making it 800× faster than the previous state-of-the-art 3D-aware methods. Hengfei Wang, Zhongqun Zhang, Yihua Cheng, Hyung Jin Chang |
AAAI | 4 |
| 2026 | SimForce: Force and Surface Electromyography from Full Body Video with Graph Neural NetsabstractWe propose a novel framework, named SimForce, for simultaneously estimating skeletal pose, ground reaction force and surface electromyography from an input video. Simforce predicts the more biomechanically accurate 3D human pose and shape of a given subject, along with their proposed muscle activations and resultant ground reaction force which leads to their input motion. Previous research has either focused on estimating these attributes singly and not treated them as a related task by taking into account the inherent shared motion between the three. In contrast, SimForce is designed to take advantage of the shared biological structure of the human body and its intrinsic connections to infer these attributes jointly using past, current, and future frames. SimForce features a newly introduced temporal and attention aware GCN-based architecture. To learn the subtle links between the body parts and how it affects the distribution of weight on the muscles over time, we introduce the Spatially Aware Attention Module. Esha Dasgupta, Boeun Kim, Sang Hoon Yeo, Hyung Jin Chang |
WACV | 4 |
| 2026 | On the design fundamentals of diffusion models: A surveyabstractDiffusion models are learning pattern-learning systems to model and sample from data distributions with three functional components namely the forward process, the reverse process, and the sampling process. The components of diffusion models have gained significant attention with many design factors being considered in common practice. Existing reviews have primarily focused on higher-level solutions, covering less on the design fundamentals of components. This study seeks to address this gap by providing a comprehensive and coherent review of seminal designable factors within each functional component of diffusion models. This provides a finer-grained perspective of diffusion models, benefiting future studies in the analysis of individual components, the design factors for different purposes, and the implementation of diffusion models. • A literature review of design fundamentals of diffusion models. • Identified the three functional components of diffusion models. • Examined the key factors and their functionalities of the components. • Discussed the popular designs and seminal works for each factor. Ziyi Chang, George Alex Koulieris, Hyung Jin Chang, Hubert P. H. Shum |
Pattern Recognit. | 3 |
| 2026 | Bidirectional regression for monocular 6DoF head pose estimation and reference system alignment
Sungho Chun, Boeun Kim, Hyung Jin Chang, Ju Yong Chang |
Pattern Recognit. | 3 |
| 2026 | Video-based human pose estimation via feature decoupling and multi-hypothesis calibration
Runyang Feng, Tze Ho Elden Tse, Haoming Chen, Hyung Jin Chang, Haifeng Zhong, Yixing Gao 0001 |
Pattern Recognit. | 4 |
| 2025 | Collaborative Learning for 3D Hand-Object Reconstruction and Compositional Action Recognition from Egocentric RGB Videos Using SuperquadricsabstractWith the availability of egocentric 3D hand-object interaction datasets, there is increasing interest in developing unified models for hand-object pose estimation and action recognition. However, existing methods still struggle to recognise seen actions on unseen objects due to the limitations in representing object shape and movement using 3D bounding boxes. Additionally, the reliance on object templates at test time limits their generalisability to unseen objects. To address these challenges, we propose to leverage superquadrics as an alternative 3D object representation to bounding boxes and demonstrate their effectiveness on both template-free object reconstruction and action recognition tasks. Moreover, as we find that pure appearance-based methods can outperform the unified methods, the potential benefits from 3D geometric information remain unclear. Therefore, we study the compositionality of actions by considering a more challenging task where the training combinations of verbs and nouns do not overlap with the testing split. We extend H2O and FPHA datasets with compositional splits and design a novel collaborative learning framework that can explicitly reason about the geometric relations between hands and the manipulated object. Through extensive quantitative and qualitative evaluations, we demonstrate significant improvements over the state-of-the-arts in (compositional) action recognition. Tze Ho Elden Tse, Runyang Feng, Linfang Zheng, Yixing Gao 0001, Jihie Kim, Ales Leonardis, Hyung Jin Chang |
AAAI | 8 |
| 2025 | Single-view Image to Novel-view Generation for Hand-Object InteractionsabstractHand-object interaction modeling from a single RGB image is a significantly challenging task. Previous works typically reconstruct hand-object interactions as texture-less meshes, ignoring photo-realistic image generation. In this work, we introduce the HO123, a novel method to synthesize novel-view hand-object interaction images from a single image. To this end, we first train a 2D diffusion prior. Given the camera pose in novel views, our approach transfers the camera information into explicit hand representations, including hand depth and skeleton images. We propose a global hand embedding to control the diffusion model based on these hand representations. We then learn a 3D Gaussian splatting for novel-view rendering using the diffusion prior. However, occluded objects present a persistent challenge. To address this issue, we further introduce local hand embedding, where a contact field is defined in the 3D Gaussian Splatting. We leverage contact information to guide the rendering in the contact field. Extensive experiments on the HO3D and DexYCB datasets demonstrate that our method significantly outperforms state-of-the-art novel-view synthesis for hand-object interactions. Zhongqun Zhang, Yihua Cheng, Eduardo Pérez-Pellitero, Yiren Zhou, Jiankang Deng, Hyung Jin Chang, Jifei Song |
AAAI | 6 |
| 2025 | 3D Prior Is All You Need: Cross-Task Few-shot 2D Gaze Estimationabstract3D and 2D gaze estimation share the fundamental objective of capturing eye movements but are traditionally treated as two distinct research domains. In this paper, we introduce a novel cross-task few-shot 2D gaze estimation approach, aiming to adapt a pre-trained 3D gaze estimation network for 2D gaze prediction on unseen devices using only a few training images. This task is highly challenging due to the domain gap between 3D and 2D gaze, unknown screen poses, and limited training data. To address these challenges, we propose a novel framework that bridges the gap between 3D and 2D gaze. Our framework contains a physics-based differentiable projection module with learnable parameters to model screen poses and project 3D gaze into 2D gaze. The framework is fully differentiable and can integrate into existing 3D gaze networks without modifying their original architecture. Additionally, we introduce a dynamic pseudo-labelling strategy for flipped images, which is particularly challenging for 2D labels due to unknown screen poses. To overcome this, we reverse the projection process by converting 2D labels to 3D space, where flipping is performed. Notably, this 3D space is not aligned with the camera coordinate system, so we learn a dynamic transformation matrix to compensate for this misalignment. We evaluate our method on MPIIGaze, EVE, and GazeCapture datasets, collected respectively on laptops, desktop computers, and mobile devices. The superior performance highlights the effectiveness of our approach, and demonstrates its strong potential for real-world applications. Yihua Cheng, Hengfei Wang, Zhongqun Zhang, Boeun Kim, Feng Lu 0005, Hyung Jin Chang |
CVPR | 7 |
| 2025 | PoseBH: Prototypical Multi-Dataset Training Beyond Human Pose EstimationabstractWe study multi-dataset training (MDT) for pose estimation, where skeletal heterogeneity presents a unique challenge that existing methods have yet to address. In traditional domains, e.g. regression and classification, MDT typically relies on dataset merging or multi-head supervision. However, the diversity of skeleton types and limited cross-dataset supervision complicate integration in pose estimation. To address these challenges, we introduce PoseBH, a new MDT framework that tackles keypoint heterogeneity and limited supervision through two key techniques. First, we propose nonparametric keypoint prototypes that learn within a unified embedding space, enabling seamless integration across skeleton types. Second, we develop a cross-type self-supervision mechanism that aligns keypoint predictions with keypoint embedding prototypes, providing supervision without relying on teacher-student models or additional augmentations. PoseBH substantially improves generalization across whole-body and animal pose datasets, including COCO-WholeBody, AP-10K, and APT-36K, while preserving performance on standard human pose benchmarks (COCO, MPII, and AIC). Furthermore, our learned key-point embeddings transfer effectively to hand shape estimation (InterHand2.6M) and human body shape estimation (3DPW). The code for PoseBH is available at: https://github.com/uyoung-jeong/PoseBH. Uyoung Jeong, Jonathan Freer, Seungryul Baek, Hyung Jin Chang, Kwang In Kim |
CVPR | 4 |
| 2025 | PersonaBooth: Personalized Text-to-Motion GenerationabstractThis paper introduces Motion Personalization, a new task that generates personalized motions aligned with text descriptions using several basic motions containing Persona. To support this novel task, we introduce a new large-scale motion dataset called PerMo (PersonaMotion), which captures the unique personas of multiple actors. We also propose a multi-modal finetuning method of a pretrained motion diffusion model called PersonaBooth. PersonaBooth addresses two main challenges: i) A significant distribution gap between the persona-focused PerMo dataset and the pretraining datasets, which lack persona-specific data, and ii) the difficulty of capturing a consistent persona from the motions vary in content (action type). To tackle the dataset distribution gap, we introduce a persona token to accept new persona features and perform multi-modal adaptation for both text and visuals during finetuning. To capture a consistent persona, we incorporate a contrastive learning technique to enhance intra-cohesion among samples with the same persona. Furthermore, we introduce a context-aware fusion mechanism to maximize the integration of persona cues from multiple input motions. PersonaBooth outperforms state-of-the-art motion style transfer methods, establishing a new benchmark for motion personalization. Boeun Kim, Hea In Jeong, JungHoon Sung, Yihua Cheng, Jeongmin Lee 0007, Ju Yong Chang, Sang-Il Choi, Younggeun Choi 0001, Saim Shin, Hyung Jin Chang |
CVPR | 11 |
| 2025 | High-Resolution Spatiotemporal Modeling with Global-Local State Space Models for Video-Based Human Pose EstimationabstractModeling high-resolution spatiotemporal representations, including both global dynamic contexts (e.g., holistic human motion tendencies) and local motion details (e.g., high-frequency changes of keypoints), is essential for video-based human pose estimation (VHPE). Current state-of-the-art methods typically unify spatiotemporal learning within a single type of modeling structure (convolution or attention-based blocks), which inherently have difficulties in balancing global and local dynamic modeling and may bias the network to one of them, leading to suboptimal performance. Moreover, existing VHPE models suffer from quadratic complexity when capturing global dependencies, limiting their applicability especially for high-resolution sequences. Recently, the state space models (known as Mamba) have demonstrated significant potential in modeling long-range contexts with linear complexity; however, they are restricted to 1D sequential data. In this paper, we present a novel framework that extends Mamba from two aspects to separately learn global and local high-resolution spatiotemporal representations for VHPE. Specifically, we first propose a Global Spatiotemporal Mamba, which performs 6D selective space-time scan and spatial- and temporal-modulated scan merging to efficiently extract global representations from high-resolution sequences. We further introduce a windowed space-time scan-based Local Refinement Mamba to enhance the high-frequency details of localized keypoint motions. Extensive experiments on four benchmark datasets demonstrate that the proposed model outperforms state-of-the-art VHPE approaches while achieving better computational trade-offs. Runyang Feng, Hyung Jin Chang, Tze Ho Elden Tse, Boeun Kim, Yi Chang 0001, Yixing Gao 0001 |
ICCV | 2 |
| 2025 | Instruction-Grounded Visual Projectors for Continual Learning of Generative Vision-Language ModelsabstractContinual learning enables pre-trained generative vision-language models (VLMs) to incorporate knowledge from new tasks without retraining data from previous ones. Recent methods update a visual projector to translate visual information for new tasks, connecting pre-trained vision encoders with large language models. However, such adjustments may cause the models to prioritize visual inputs over language instructions, particularly learning tasks with repetitive types of textual instructions. To address the neglect of language instructions, we propose a novel framework that grounds the translation of visual information on instructions for language models. We introduce a mixture of visual projectors, each serving as a specialized visual-to-language translation expert based on the given instruction context to adapt to new tasks. To avoid using experts for irrelevant instruction contexts, we propose an expert recommendation strategy that reuses experts for tasks similar to those previously learned. Additionally, we introduce expert pruning to alleviate interference from the use of experts that cumulatively activated in previous tasks. Extensive experiments on diverse vision-language tasks demonstrate that our method outperforms existing continual learning approaches by generating instruction-following responses. Hyundong Jin, Hyung Jin Chang, Eunwoo Kim |
ICCV | 2 |
| 2025 | AMDANet: Attention-Driven Multi-Perspective Discrepancy Alignment for RGB-Infrared Image Fusion and SegmentationabstractThe challenge of multimodal semantic segmentation lies in establishing semantically consistent and segmentable multimodal fusion features under conditions of significant visual feature discrepancies. Existing methods commonly construct cross-modal self-attention fusion frameworks or introduce additional multimodal fusion loss functions to establish fusion features. However, these approaches often overlook the challenge caused by feature discrepancies between modalities during the fusion process. To achieve precise segmentation, we propose an Attention-Driven Multimodal Discrepancy Alignment Network (AMDANet). AMDANet reallocates weights to reduce the saliency of discrepant features and utilizes low-weight features as cues to mitigate discrepancies between modalities, thereby achieving multimodal feature alignment. Furthermore, to simplify the feature alignment process, a semantic consistency inference mechanism is introduced to reveal the network’s inherent bias toward specific modalities, thereby compressing cross-modal feature discrepancies from the foundational level. Extensive experiments on the FMB, MFNet, and PST900 datasets demonstrate that AMDANet achieves mIoU improvements of 3.6%, 3.0%, and 1.6%, respectively, significantly outperforming state-of-the-art methods. The code is available at https://github.com/Zhonghaifeng6/AMDANet Haifeng Zhong, Fan Tang, Zhuo Chen 0028, Hyung Jin Chang, Yixing Gao 0001 |
ICCV | 4 |
| 2025 | Multi-Hypothesis 3D Hand Mesh Recovering from a Single Blurry ImageabstractRecovery of 3D hand mesh from blurry hand images is challenging due to the ambiguity. Most existing works attempt to solve this issue by exploiting physical and temporal constraints. However, those works ignore the fact that multiple feasible solutions exist. In this paper, we propose a two-stage Multi-Hypothesis Hand Mesh Recovery network, consisting of a generation and selection model. In the first stage, the generation model explicitly extracts the temporal information with an unfolder. Then, a multi-hypothesis Transformer generates multiple diverse hypotheses with a lightweight hypothesis embedding set. In the second stage, the selection model selects a subset of good-quality hypotheses. We additionally combine the classifying and ranking loss to better align with the target of the selection model. Extensive experiments show that the proposed method produces much more accurate results on blurry images. Source code is available at https://github.com/RandSF/Multi_Hypothesis_BlurHandNet. Rongyu Chen, Zhongqun Zhang, Yihua Cheng, Hyung Jin Chang |
ICME | 5 |
| 2025 | DarkSeg: Infrared-Driven Semantic Segmentation for Garment Grasping Detection in Low-Light ConditionsabstractGarment grasping in low-light environments is a critical challenge for domestic intelligent robots, yet existing research has not sufficiently addressed this issue. In low-light conditions, the scarcity of visual features due to insufficient illumination causes different categories of garments to exhibit ambiguous feature similarities, thereby hindering the robot’s ability to detect the categories of different garments. Although traditional methods can compensate for visual deficiencies in low-light scenarios by applying preprocessing strategies that fuse infrared multimodal features, their complex computational processes incur significant computational overhead. To address this limitation, we propose a low-light garment detection model based on the student-teacher model. The innovation of DarkSeg lies in its replacement of complex multimodal feature fusion with an indirect feature alignment mechanism between the student and teacher models, thereby circumventing high computational demands. Through feature alignment, DarkSeg enables the student model to learn illumination-invariant structural representations from the infrared features provided by the teacher model, effectively correcting structural deficiencies in low-light environments. Furthermore, to evaluate DarkSeg’s feasibility for low-light clothing grasping, we propose a depth-perceptive grasping strategy and build a low-light multimodal garment detection dataset, DarkClothes. Extensive experiments deploying DarkSeg on a Baxter robot demonstrate that DarkSeg achieves a 22% improvement in the grasping success rate while reducing the model parameters by 99.08 million compared to traditional methods, validating the practical viability of DarkSeg for robotic garment grasping in low-light conditions. The code and dataset are available at https://github.com/Zhonghaifeng6/Darkseg Haifeng Zhong, Fan Tang, Hyung Jin Chang, Xingyu Zhu 0014, Yixing Gao 0001 |
IROS | 3 |
| 2025 | Roll Your Eyes: Gaze Redirection via Explicit 3D Eyeball RotationabstractWe propose a novel 3D gaze redirection framework that leverages an explicit 3D eyeball structure. Existing gaze redirection methods are typically based on neural radiance fields, which employ implicit neural representations via volume rendering. Unlike these NeRF-based approaches, where the rotation and translation of 3D representations are not explicitly modeled, we introduce a dedicated 3D eyeball structure to represent the eyeballs with 3D Gaussian Splatting (3DGS). Our method generates photorealistic images that faithfully reproduce the desired gaze direction by explicitly rotating and translating the 3D eyeball structure. In addition, we propose an adaptive deformation module that enables the replication of subtle muscle movements around the eyes. Through experiments conducted on the ETH-XGaze dataset, we demonstrate that our framework is capable of generating diverse novel gaze images, achieving superior image quality and gaze estimation accuracy compared to previous state-of-the-art methods. YoungChan Choi, HengFei Wang, YiHua Cheng, Boeun Kim, Hyung Jin Chang, Younggeun Choi 0001, Sang-Il Choi |
ACM Multimedia | 5 |
| 2025 | AuGQ: Augmented quantization granularity to overcome accuracy degradation for sub-byte quantized deep neural networks
Ahmed Mujtaba, Wai-Kong Lee, ByoungChul Ko, Hyung Jin Chang, Seong Oun Hwang |
Appl. Intell. | 4 |
| 2025 | A unified framework for unsupervised action learning via global-to-local motion transformer
Boeun Kim, Hyung Jin Chang, Tae-Hyun Oh |
Pattern Recognit. | 3 |
| 2025 | Bridging domain spaces for unsupervised domain adaptation
Jaemin Na, Heechul Jung, Hyung Jin Chang, Wonjun Hwang |
Pattern Recognit. | 3 |
| 2025 | Behavior-Aware Knowledge-Embedded Model for Driver Attention PredictionabstractAccurately predicting driver attention is crucial for enhancing advanced driving assistance systems and autonomous vehicles, attracting increasing research interest. Most existing approaches, rooted in general, task-free saliency detection, adopt data-driven paradigms to correlate bottom-up environmental situations with attention distributions. However, they often overlook the complex top-down task-driven aspects of driver attention that are fundamental for the safe navigation of driving tasks, leading to limitations in handling real-world scenarios. In this paper, we take an initial step to explore and introduce BKnet, a Behavior-aware Knowledge-embedded model that innovatively integrates driving behaviors and empirical knowledge. Specifically, inspired by the human long-term cognitive process, we introduce a novel knowledge memory mechanism. It dynamically associates varied traffic scenarios with consistent driving behaviors, fostering the generation of robust behavior-aware empirical knowledge representations. To this end, BKnet facilitates a nuanced and comprehensive simulation of drivers’ attention mechanisms, driven synergistically by both top-down and bottom-up processes. Additionally, we further contribute to the field by collecting a novel Behavior-Aware Driver Attention (BADA) dataset. To the best of our knowledge, BADA is the first attention dataset explicitly incorporated into real-world driving behavior tasks from multiple drivers. Lastly, comprehensive experiments underscore BKnet’s superiority over existing state-of-the-art approaches and validate the effectiveness and necessity of integrating behavior-aware knowledge into driver attention prediction. Yuchen Zhou 0002, Chao Gou, Zipeng Guo, Yihua Cheng, Hyung Jin Chang |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2024 | NCRF: Neural Contact Radiance Fields for Free-Viewpoint Rendering of Hand-Object InteractionabstractModeling hand-object interactions is a fundamentally challenging task in 3D computer vision. Despite remarkable progress that has been achieved in this field, existing methods still fail to synthesize the hand-object interaction photo-realistically, suffering from degraded rendering quality caused by the heavy mutual occlusions between the hand and the object, and inaccurate hand-object pose estimation. To tackle these challenges, we present a novel free-viewpoint rendering framework, Neural Contact Radiance Field (NCRF), to reconstruct hand-object interactions from a sparse set of videos. In particular, the proposed NCRF framework consists of two key components: (a) A contact optimization field that predicts an accurate contact field from 3D query points for achieving desirable contact between the hand and the object. (b) A hand-object neural radiance field to learn an implicit hand-object representation in a static canonical space, in concert with the specifically designed hand-object motion field to produce observation-to-canonical correspondences. We jointly learn these key components where they mutually help and regularize each other with visual and geometric constraints, producing a high-quality hand-object reconstruction that achieves photorealistic novel view synthesis. Extensive experiments on HO3D and DexYCB datasets show that our approach outperforms the current state-of-the-art in terms of both rendering quality and pose estimation accuracy. Zhongqun Zhang, Jifei Song, Eduardo Pérez-Pellitero, Yiren Zhou, Hyung Jin Chang, Ales Leonardis |
3DV | 5 |
| 2024 | Dual Prototype-Driven Objectness Decoupling for Cross-Domain Object Detection in Urban Scene
Jaemin Na, Joong-Won Hwang, Hyung Jin Chang, Wonjun Hwang |
ACCV (8) | 4 |
| 2024 | Few Exemplar-Based General Medical Image Segmentation via Domain-Aware Selective Adaptation
Qiming Huang, Hyung Jin Chang, Jianbo Jiao |
ACCV (2) | 6 |
| 2024 | What Do You See in Vehicle? Comprehensive Vision Solution for In-Vehicle Gaze EstimationabstractDriver's eye gaze holds a wealth of cognitive and intentional cues crucial for intelligent vehicles. Despite its sig-nificance, research on in-vehicle gaze estimation remains limited due to the scarcity of comprehensive and well-annotated datasets in real driving scenarios. In this pa-per, we present three novel elements to advance in-vehicle gaze research. Firstly, we introduce IVGaze, a pioneering dataset capturing in-vehicle gaze, collected from 125 sub-jects and covering a large range of gaze and head poses within vehicles. In this dataset, we propose a new vision-based solution for in-vehicle gaze collection, introducing a refined gaze target calibration method to tackle annotation challenges. Second, our research focuses on in-vehicle gaze estimation leveraging the IvGaze. In-vehicle face images often suffer from low resolution, prompting our in-troduction of a gaze pyramid transformer that leverages transformer-based multilevel features integration. Expanding upon this, we introduce the dual-stream gaze pyramid transformer (GazeDPTR). Employing perspective transfor-mation, we rotate virtual cameras to normalize images, uti-lizing camera pose to merge normalized and original images for accurate gaze estimation. GazeDPTR shows state-of-the-art performance on the IVGaze dataset. Thirdly, we explore a novel strategy for gaze zone classification by extending the GazeDPTR. A foundational tri-plane and project gaze onto these planes are newly defined. Leveraging both positional features from the projection points and visual attributes from images, we achieve superior performance compared to relying solely on visual features, sub-stantiating the advantage of gaze estimation. The project is available at https://yihua.zone/work/ivgaze. Yihua Cheng, Yaning Zhu, Zongji Wang, Hongquan Hao, Yongwei Liu, Shiqing Cheng, Xi Wang 0021, Hyung Jin Chang |
CVPR | 8 |
| 2024 | MoST: Motion Style Transformer Between Diverse Action ContentsabstractWhile existing motion style transfer methods are effective between two motions with identical content, their performance significantly diminishes when transferring style between motions with different contents. This challenge lies in the lack of clear separation between content and style of a motion. To tackle this challenge, we propose a novel motion style transformer that effectively disentangles style from content and generates a plausible motion with transferred style from a source motion. Our distinctive approach to achieving the goal of disentanglement is twofold: (1) a new architecture for motion style transformer with 'part-attentive style modulator across body parts' and ‘Siamese encoders that encode style and content features separately’; (2) style disentanglement loss. Our method outperforms existing methods and demonstrates exceptionally high quality, particularly in motion pairs with different contents, without the need for heuristic post-processing. Codes are available at https://github.com/Boeun-Kim/MoST. Boeun Kim, Hyung Jin Chang, Jin Young Choi 0002 |
CVPR | 3 |
| 2024 | GeoReF: Geometric Alignment Across Shape Variation for Category-level Object Pose RefinementabstractObject pose refinement is essential for robust object pose estimation. Previous work has made significant progress to-wards instance-level object pose refinement. Yet, category-level pose refinement is a more challenging problem due to large shape variations within a category and the discrep-ancies between the target object and the shape prior. To address these challenges, we introduce a novel architecture for category-level object pose refinement. Our approach in-tegrates an HS-Iayer and learnable affine transformations, which aims to enhance the extraction and alignment of Geometric information. Additionally, we introduce a cross-cloud transformation mechanism that efficiently merges di-verse data sources. Finally, we push the limits of our model by incorporating the shape prior information for translation and size error prediction. We conducted extensive ex-periments to demonstrate the effectiveness of the proposed framework. Through extensive quantitative experiments, we demonstrate significant improvement over the baseline method by a large margin across all metrics.11Project page: https://lynne-zheng-linfang.github.io/georef.github.io Linfang Zheng, Tze Ho Elden Tse, Chen Wang 0123, Yinghan Sun, Hua Chen 0007, Ales Leonardis, Wei Zhang 0013, Hyung Jin Chang |
CVPR | 8 |
| 2024 | Benchmarks and Challenges in Pose Estimation for Egocentric Hand Interactions with Objects
Zicong Fan, Takehiko Ohkawa, Linlin Yang 0001, Nie Lin, Zhishan Zhou, Jiajun Liang, Zhong Gao, Xuanyang Zhang, Feng Lu 0005, Karim Abou Zeid, Bastian Leibe, Jeongwan On, Seungryul Baek, Saurabh Gupta 0001, Yoichi Sato 0001, Otmar Hilliges, Hyung Jin Chang, Angela Yao |
ECCV (25) | 23 |
| 2024 | NL2Contact: Natural Language Guided 3D Hand-Object Contact Modeling with Diffusion Model
Zhongqun Zhang, Hengfei Wang, Ziwei Yu, Yihua Cheng, Angela Yao, Hyung Jin Chang |
ECCV (28) | 6 |
| 2024 | TextGaze: Gaze-Controllable Face Generation with Natural LanguageabstractGenerating face image with specific gaze information has attracted considerable attention in recent years. Existing approaches typically input gaze values directly for face generation, which is unnatural and requires annotated gaze datasets for training, thereby limiting its application. In this paper, we present a novel gaze-controllable face generation task that overcomes these limitations. Our approach inputs textual descriptions that describe human gaze and head behavior and generates corresponding face images. Our work first introduces a text-of-gaze dataset containing over 90k text descriptions spanning a dense distribution of gaze and head poses. We further propose a gaze-controllable text-to-face method. Our method contains a sketch-conditioned face diffusion module and a model-based sketch diffusion module. We define a face sketch based on facial landmarks and eye segmentation map. It provides a structured and detailed foundation for generating facial images. The face diffusion module generates face images from the face sketch, and the sketch diffusion module employs a 3D face model to generate face sketch from text description. Experiments on the FFHQ dataset show the effectiveness of our method. Our dataset is available at https://github.com/hengfei-wang/TextGaze. Hengfei Wang, Zhongqun Zhang, Yihua Cheng, Hyung Jin Chang |
ACM Multimedia | 4 |
| 2024 | Multi-Modal Gaze Following in Conversational ScenariosabstractGaze following estimates gaze targets of in-scene person by understanding human behavior and scene information. Existing methods usually analyze scene images for gaze following. However, compared with visual images, audio also provides crucial cues for determining human behavior. This suggests that we can further improve gaze following considering audio cues. In this paper, we explore gaze following tasks in conversational scenarios. We propose a novel multimodal gaze following framework based on our observation "audiences tend to focus on the speaker". We first leverage the correlation between audio and lips, and classify speakers and listeners in a scene. We then use the identity information to enhance scene images and propose a gaze candidate estimation network. The network estimates gaze candidates from enhanced scene images and we use MLP to match subjects with candidates as classification tasks. Existing gaze following datasets focus on visual images while ignore audios. To evaluate our method, we collect a conversational dataset, VideoGazeSpeech (VGS), which is the first gaze following dataset including images and audio. Our method significantly outperforms existing methods in VGS datasets. The visualization result also prove the advantage of audio cues in gaze following tasks. Our work will inspire more researches in multi-modal gaze following estimation. Zhongqun Zhang, Nora Horanyi, Jaewon Moon, Yihua Cheng, Hyung Jin Chang |
WACV | 6 |
| 2023 | BoIR: Box-Supervised Instance Representation for Multi-Person Pose Estimation
Uyoung Jeong, Seungryul Baek, Hyung Jin Chang, Kwang In Kim |
BMVC | 3 |
| 2023 | High-Fidelity Eye Animatable Neural Radiance Fields for Human Face
Hengfei Wang, Zhongqun Zhang, Yihua Cheng, Hyung Jin Chang |
BMVC | 4 |
| 2023 | Mutual Information-Based Temporal Difference Learning for Human Pose Estimation in VideoabstractTemporal modeling is crucial for multi-frame human pose estimation. Most existing methods directly employ optical flow or deformable convolution to predict full-spectrum motion fields, which might incur numerous irrelevant cues, such as a nearby person or background. Without further efforts to excavate meaningful motion priors, their results are suboptimal, especially in complicated spatio-temporal interactions. On the other hand, the temporal difference has the ability to encode representative motion information which can potentially be valuable for pose estimation but has not been fully exploited. In this paper, we present a novel multi-frame human pose estimation framework, which employs temporal differences across frames to model dynamic contexts and engages mutual information objectively to facilitate useful motion information disentanglement. To be specific, we design a multi-stage Temporal Difference Encoder that performs incremental cascaded learning conditioned on multi-stage feature difference sequences to derive informative motion representation. We further propose a Representation Disentanglement module from the mutual information perspective, which can grasp discriminative task-relevant motion signals by explicitly defining useful and noisy constituents of the raw motion features and minimizing their mutual information. These place us to rank No.1 in the Crowd Pose Estimation in Complex Events Challenge on benchmark dataset HiEve, and achieve state-of-the-art performance on three benchmarks PoseTrack2017, PoseTrack2018, and PoseTrack21. Runyang Feng, Yixing Gao 0001, Xueqing Ma, Tze Ho Elden Tse, Hyung Jin Chang |
CVPR | 5 |
| 2023 | GazeNeRF: 3D-Aware Gaze Redirection with Neural Radiance FieldsabstractWe propose GazeNeRF, a 3D-aware method for the task of gaze redirection. Existing gaze redirection methods operate on 2D images and struggle to generate 3D consistent results. Instead, we build on the intuition that the face region and eyeballs are separate 3D structures that move in a coordinated yet independent fashion. Our method leverages recent advancements in conditional image-based neural radiance fields and proposes a two-stream architecture that predicts volumetric features for the face and eye regions separately. Rigidly transforming the eye features via a 3D rotation matrix provides fine-grained control over the desired gaze angle. The final, redirected image is then attained via differentiable volume compositing. Our experiments show that this architecture outperforms naively conditioned NeRF baselines as well as previous state-of-the-art 2D gaze redirection methods in terms of redirection accuracy and identity preservation. Code and models will be released for research purposes. Alessandro Ruzzi, Xiangwei Shi, Xi Wang 0021, Gengyan Li 0001, Shalini De Mello, Hyung Jin Chang, Xucong Zhang, Otmar Hilliges |
CVPR | 6 |
| 2023 | HS-Pose: Hybrid Scope Feature Extraction for Category-level Object Pose EstimationabstractIn this paper, we focus on the problem of category-level object pose estimation, which is challenging due to the large intra-category shape variation. 3D graph convolution (3D-GC) based methods have been widely used to extract local geometric features, but they have limitations for complex shaped objects and are sensitive to noise. Moreover, the scale and translation invariant properties of 3D-GC restrict the perception of an object's size and translation information. In this paper, we propose a simple network structure, the HS-layer, which extends 3D-GC to extract hybrid scope latent features from point cloud data for category-level object pose estimation tasks. The proposed HS-layer: 1) is able to perceive local-global geometric structure and global information, 2) is robust to noise, and 3) can encode size and translation information. Our experiments show that the simple replacement of the 3D-GC layer with the proposed HS-layer on the baseline method (GPV-Pose) achieves a significant improvement, with the performance increased by 14.5% on 5°2cm metric and 10.3% on IoU75. Our method outperforms the state-of-the-art methods by a large margin (8.3% on 5°2cm, 6.9% on IoU75) on REAL275 dataset and runs in real-time (50 FPS)11Codeisavailable: https://github.com/Lynne-Zheng-Linfang/HS-Pose. Linfang Zheng, Chen Wang 0123, Yinghan Sun, Esha Dasgupta, Hua Chen 0007, Ales Leonardis, Wei Zhang 0013, Hyung Jin Chang |
CVPR | 8 |
| 2023 | DiffPose: SpatioTemporal Diffusion Model for Video-Based Human Pose EstimationabstractDenoising diffusion probabilistic models that were initially proposed for realistic image generation have recently shown success in various perception tasks (e.g., object detection and image segmentation) and are increasingly gaining attention in computer vision. However, extending such models to multi-frame human pose estimation is non-trivial due to the presence of the additional temporal dimension in videos. More importantly, learning representations that focus on keypoint regions is crucial for accurate localization of human joints. Nevertheless, the adaptation of the diffusion-based methods remains unclear on how to achieve such objective. In this paper, we present DiffPose, a novel diffusion architecture that formulates video-based human pose estimation as a conditional heatmap generation problem. First, to better leverage temporal information, we propose SpatioTemporal Representation Learner which aggregates visual evidences across frames and uses the resulting features in each denoising step as a condition. In addition, we present a mechanism called Lookup-based Multi-Scale Feature Interaction that determines the correlations between local joints and global contexts across multiple scales. This mechanism generates delicate representations that focus on keypoint regions. Altogether, by extending diffusion models, we show two unique characteristics from DiffPose on pose estimation task: (i) the ability to combine multiple sets of pose estimates to improve prediction accuracy, particularly for challenging joints, and (ii) the ability to adjust the number of iterative steps for feature refinement without retraining the model. DiffPose sets new state-of-the-art results on three benchmarks: PoseTrack2017, PoseTrack2018, and PoseTrack21. Runyang Feng, Yixing Gao 0001, Tze Ho Elden Tse, Xueqing Ma, Hyung Jin Chang |
ICCV | 5 |
| 2023 | Spectral Graphormer: Spectral Graph-based Transformer for Egocentric Two-Hand Reconstruction using Multi-View Color ImagesabstractWe propose a novel transformer-based framework that reconstructs two high fidelity hands from multi-view RGB images. Unlike existing hand pose estimation methods, where one typically trains a deep network to regress hand model parameters from single RGB image, we consider a more challenging problem setting where we directly regress the absolute root poses of two-hands with extended forearm at high resolution from egocentric view. As existing datasets are either infeasible for egocentric viewpoints or lack background variations, we create a large-scale synthetic dataset with diverse scenarios and collect a real dataset from multi-calibrated camera setup to verify our proposed multi-view image feature fusion strategy. To make the reconstruction physically plausible, we propose two strategies: (i) a coarse-to-fine spectral graph convolution decoder to smoothen the meshes during upsampling and (ii) an optimisation-based refinement stage at inference to prevent self-penetrations. Through extensive quantitative and qualitative evaluations, we show that our framework is able to produce realistic two-hand reconstructions and demonstrate the generalisation of synthetic-trained models to real data, as well as real-time AR/VR applications. Tze Ho Elden Tse, Franziska Mueller 0001, Zhengyang Shen, Danhang Tang, Thabo Beeler, Mingsong Dou, Yinda Zhang 0001, Sasa Petrovic, Hyung Jin Chang, Jonathan Taylor 0001, Bardia Doosti |
ICCV | 9 |
| 2023 | Clothes Grasping and Unfolding Based on RGB-D Semantic SegmentationabstractClothes grasping and unfolding is a core step in robotic-assisted dressing. Most existing works leverage depth images of clothes to train a deep learning-based model to recognize suitable grasping points. These methods often utilize physics engines to synthesize depth images to reduce the cost of real labeled data collection. However, the natural domain gap between synthetic and real images often leads to poor performance of these methods on real data. Furthermore, these approaches often struggle in scenarios where grasping points are occluded by the clothing item itself. To address the above challenges, we propose a novel Bi-directional Fractal Cross Fusion Network (BiFCNet) for semantic segmentation, enabling recognition of graspable regions in order to provide more possibilities for grasping. Instead of using depth images only, we also utilize RGB images with rich color features as input to our network in which the Fractal Cross Fusion (FCF) module fuses RGB and depth data by considering global complex features based on fractal geometry. To reduce the cost of real data collection, we further propose a data augmentation method based on an adversarial strategy, in which the color and geometric transformations simultaneously process RGB and depth data while maintaining the label correspondence. Finally, we present a pipeline for clothes grasping and unfolding from the perspective of semantic segmentation, through the addition of a strategy for grasp point selection from segmentation regions based on clothing flatness measures, while taking into account the grasping direction. We evaluate our BiFCNet on the public dataset NYUDv2 and obtained comparable performance to current state-of-the-art models. We also deploy our model on a Baxter robot, running extensive grasping and unfolding experiments as part of our ablation studies, achieving an 84% success rate. Xingyu Zhu 0014, Xin Wang 0035, Jonathan Freer, Hyung Jin Chang, Yixing Gao 0001 |
ICRA | 4 |
| 2023 | Switching Temporary Teachers for Semi-Supervised Semantic SegmentationabstractThe teacher-student framework, prevalent in semi-supervised semantic segmentation, mainly employs the exponential moving average (EMA) to update a single teacher's weights based on the student's. However, EMA updates raise a problem in that the weights of the teacher and student are getting coupled, causing a potential performance bottleneck. Furthermore, this problem may become more severe when training with more complicated labels such as segmentation masks but with few annotated data. This paper introduces Dual Teacher, a simple yet effective approach that employs dual temporary teachers aiming to alleviate the coupling problem for the student. The temporary teachers work in shifts and are progressively improved, so consistently prevent the teacher and student from becoming excessively close. Specifically, the temporary teachers periodically take turns generating pseudo-labels to train a student model and maintain the distinct characteristics of the student model for each epoch. Consequently, Dual Teacher achieves competitive performance on the PASCAL VOC, Cityscapes, and ADE20K benchmarks with remarkably shorter training times than state-of-the-art methods. Moreover, we demonstrate that our approach is model-agnostic and compatible with both CNN- and Transformer-based models. Code is available at https://github.com/naver-ai/dual-teacher. Jaemin Na, Jung-Woo Ha 0001, Hyung Jin Chang, Dongyoon Han, Wonjun Hwang |
NeurIPS | 3 |
| 2023 | G-DAIC: A Gaze Initialized Framework for Description and Aesthetic-Based Image CroppingabstractWe propose a new gaze-initialised optimisation framework to generate aesthetically pleasing image crops based on user description. We extended the existing description-based image cropping dataset by collecting user eye movements corresponding to the image captions. To best leverage the contextual information to initialise the optimisation framework using the collected gaze data, this work proposes two gaze-based initialisation strategies, Fixed Grid and Region Proposal. In addition, we propose the adaptive Mixed scaling method to find the optimal output despite the size of the generated initialisation region and the described part of the image. We address the runtime limitation of the state-of-the-art method by implementing the Early termination strategy to reduce the number of iterations required to produce the output. Our experiments show that G-DAIC reduced the runtime by 92.11%, and the quantitative and qualitative experiments demonstrated that the proposed framework produces higher quality and more accurate image crops w.r.t. user intention. Nora Horanyi, Ales Leonardis, Hyung Jin Chang |
Proc. ACM Hum. Comput. Interact. | 4 |
| 2022 | Collaborative Learning for Hand and Object Reconstruction with Attention-guided Graph ConvolutionabstractEstimating the pose and shape of hands and objects under interaction finds numerous applications including aug-mented and virtual reality. Existing approaches for hand and object reconstruction require explicitly defined physical constraints and known objects, which limits its application domains. Our algorithm is agnostic to object models, and it learns the physical rules governing hand-object interaction. This requires automatically inferring the shapes and physi-cal interaction of hands and (potentially unknown) objects. We seek to approach this challenging problem by proposing a collaborative learning strategy where two-branches of deep networks are learning from each other. Specifically, we transfer hand mesh information to the object branch and vice versa for the hand branch. The resulting optimi-sation (training) problem can be unstable, and we address this via two strategies: (i) attention-guided graph convo-lution which helps identify and focus on mutual occlusion and (ii) unsupervised associative loss which facilitates the transfer of information between the branches. Experiments using four widely-used benchmarks show that our frame-work achieves beyond state-of-the-art accuracy in 3D pose estimation, as well as recovers dense 3D hand and object shapes. Each technical component above contributes meaningfully in the ablation study. Tze Ho Elden Tse, Kwang In Kim, Ales Leonardis, Hyung Jin Chang |
CVPR | 4 |
| 2022 | Global-Local Motion Transformer for Unsupervised Skeleton-Based Action Learning
Boeun Kim, Hyung Jin Chang, Jin Young Choi 0002 |
ECCV (4) | 2 |
| 2022 | Contrastive Vicinal Space for Unsupervised Domain Adaptation
Jaemin Na, Dongyoon Han, Hyung Jin Chang, Wonjun Hwang |
ECCV (34) | 3 |
| 2022 | S2Contact: Graph-Based Network for 3D Hand-Object Contact Estimation with Semi-supervised Learning
Tze Ho Elden Tse, Zhongqun Zhang, Kwang In Kim, Ales Leonardis, Feng Zheng 0001, Hyung Jin Chang |
ECCV (1) | 6 |
| 2022 | Towards Generic 3D Tracking in RGBD Videos: Benchmark and Baseline
Zhongqun Zhang, Zhe Li 0008, Hyung Jin Chang, Ales Leonardis, Feng Zheng 0001 |
ECCV (22) | 4 |
| 2022 | TP-AE: Temporally Primed 6D Object Pose Tracking with Auto-EncodersabstractFast and accurate tracking of an object's motion is one of the key functionalities of a robotic system for achieving reliable interaction with the environment. This paper focuses on the instance-level six-dimensional (6D) pose tracking problem with a symmetric and textureless object under occlusion. We propose a Temporally Primed 6D pose tracking framework with Auto-Encoders (TP-AE) to tackle the pose tracking problem. The framework consists of a prediction step and a temporally primed pose estimation step. The prediction step aims to quickly and efficiently generate a guess on the object's real-time pose based on historical information about the target object's motion. Once the prior prediction is obtained, the temporally primed pose estimation step embeds the prior pose into the RGB-D input, and leverages auto-encoders to reconstruct the target object with higher quality under occlusion, thus improving the framework's performance. Extensive experiments show that the proposed 6D pose tracking method can accurately estimate the 6D pose of a symmetric and textureless object under occlusion, and significantly outperforms the state-of-the-art on T-LESS dataset while running in real-time at 26 FPS. Linfang Zheng, Ales Leonardis, Tze Ho Elden Tse, Nora Horanyi, Hua Chen 0007, Wei Zhang 0013, Hyung Jin Chang |
ICRA | 7 |
| 2022 | Novel-View Synthesis of Human Tourist PhotosabstractWe present a novel framework for performing novel-view synthesis on human tourist photos. Given a tourist photo from a known scene, we reconstruct the photo in 3D space through modeling the human and the background independently. We generate a deep buffer from a novel viewpoint of the reconstruction and utilize a deep network to translate the buffer into a photo-realistic rendering of the novel view. We additionally present a method to relight the renderings, allowing for relighting of both human and background to match either the provided input image or any other. The key contributions of our paper are: 1) a framework for performing novel view synthesis on human tourist photos, 2) an appearance transfer method for relighting of humans to match synthesized backgrounds, and 3) a method for estimating lighting properties from a single human photo. We demonstrate the proposed framework on photos from two different scenes of various tourists. Jonathan Freer, Kwang Moo Yi, Wei Jiang 0034, Jongwon Choi 0002, Hyung Jin Chang |
WACV | 5 |
| 2022 | Repurposing existing deep networks for caption and aesthetic-guided image cropping
Nora Horanyi, Kedi Xia, Kwang Moo Yi, Abhishake Kumar Bojja, Ales Leonardis, Hyung Jin Chang |
Pattern Recognit. | 6 |
| 2022 | Learning a Model-Driven Variational Network for Deformable Image RegistrationabstractData-driven deep learning approaches to image registration can be less accurate than conventional iterative approaches, especially when training data is limited. To address this issue and meanwhile retain the fast inference speed of deep learning, we propose VR-Net, a novel cascaded variational network for unsupervised deformable image registration. Using a variable splitting optimization scheme, we first convert the image registration problem, established in a generic variational framework, into two sub-problems, one with a point-wise, closed-form solution and the other one being a denoising problem. We then propose two neural layers (i.e. warping layer and intensity consistency layer) to model the analytical solution and a residual U-Net (termed generalized denoising layer) to formulate the denoising problem. Finally, we cascade the three neural layers multiple times to form our VR-Net. Extensive experiments on three (two 2D and one 3D) cardiac magnetic resonance imaging datasets show that VR-Net outperforms state-of-the-art deep learning methods on registration accuracy, whilst maintaining the fast inference speed of deep learning and the data-efficiency of variational models. Xi Jia, Alexander Thorley, Wei Chen 0092, Huaqi Qiu, LinLin Shen, Iain B. Styles, Hyung Jin Chang, Ales Leonardis, Antonio M. Simoes Monteiro de Marvao, Declan P. O'Regan, Daniel Rueckert, Jinming Duan 0001 |
IEEE Trans. Medical Imaging | 7 |
| 2021 | Apparently Irrational Choice as Optimal Sequential Decision MakingabstractIn this paper, we propose a normative approach to modeling apparently human irrational decision making (cognitive biases) that makes use of inherently rational computational mechanisms. We view preferential choice tasks as sequential decision making problems and formulate them as Partially Observable Markov Decision Processes (POMDPs). The resulting sequential decision model learns what information to gather about which options, whether to calculate option values or make comparisons between options and when to make a choice. We apply the model to choice problems where context is known to influence human choice, an effect that has been taken as evidence that human cognition is irrational. Our results show that the new model approximates a bounded optimal cognitive policy and makes quantitative predictions that correspond well to evidence about human choice. Furthermore, the model uses context to help infer which option has a maximum expected value while taking into account computational cost and cognitive limits. In addition, it predicts when, and explains why, people stop evidence accumulation and make a decision. We argue that the model provides evidence that apparent human irrationalities are emergent consequences of processes that prefer higher value (rational) policies. Haiyang Chen 0002, Hyung Jin Chang, Andrew Howes 0001 |
AAAI | 2 |
| 2021 | Class-Attentive Diffusion Network for Semi-Supervised ClassificationabstractRecently, graph neural networks for semi-supervised classification have been widely studied. However, existing methods only use the information of limited neighbors and do not deal with the inter-class connections in graphs. In this paper, we propose Adaptive aggregation with Class-Attentive Diffusion (AdaCAD), a new aggregation scheme that adaptively aggregates nodes probably of the same class among K-hop neighbors. To this end, we first propose a novel stochastic process, called Class-Attentive Diffusion (CAD), that strengthens attention to intra-class nodes and attenuates attention to inter-class nodes. In contrast to the existing diffusion methods with a transition matrix determined solely by the graph structure, CAD considers both the node features and the graph structure with the design of our class-attentive transition matrix that utilizes a classifier. Then, we further propose an adaptive update scheme that leverages different reflection ratios of the diffusion result for each node depending on the local class-context. As the main advantage, AdaCAD alleviates the problem of undesired mixing of inter-class features caused by discrepancies between node labels and the graph topology. Built on AdaCAD, we construct a simple model called Class-Attentive Diffusion Network (CAD-Net). Extensive experiments on seven benchmark datasets consistently demonstrate the efficacy of the proposed method and our CAD-Net significantly outperforms the state-of-the-art methods. Code is available at https://github.com/ljin0429/CAD-Net. Jongin Lim 0002, Daeho Um, Hyung Jin Chang, Jin Young Choi 0002 |
AAAI | 3 |
| 2021 | FS-Net: Fast Shape-Based Network for Category-Level 6D Object Pose Estimation With Decoupled Rotation MechanismabstractIn this paper, we focus on category-level 6D pose and size estimation from a monocular RGB-D image. Previous methods suffer from inefficient category-level pose feature extraction, which leads to low accuracy and inference speed. To tackle this problem, we propose a fast shape-based network (FS-Net) with efficient category-level feature extraction for 6D pose estimation. First, we design an orientation aware autoencoder with 3D graph convolution for latent feature extraction. Thanks to the shift and scale-invariance properties of 3D graph convolution, the learned latent feature is insensitive to point shift and object size. Then, to efficiently decode category-level rotation information from the latent feature, we propose a novel decoupled rotation mechanism that employs two decoders to complementarily access the rotation information. For translation and size, we estimate them by two residuals: the difference between the mean of object points and ground truth translation, and the difference between the mean size of the category and ground truth size, respectively. Finally, to increase the generalization ability of the FS-Net, we propose an on-line box-cage based 3D deformation mechanism to augment the training data. Extensive experiments on two benchmark datasets show that the proposed method achieves state-of-the-art performance in both category- and instance-level 6D object pose estimation. Especially in category-level pose estimation, without extra synthetic data, our method outperforms existing methods by 6.3% on the NOCS-REAL dataset1. Wei Chen 0092, Xi Jia, Hyung Jin Chang, Jinming Duan 0001, LinLin Shen, Ales Leonardis |
CVPR | 3 |
| 2021 | VaB-AL: Incorporating Class Imbalance and Difficulty With Variational Bayes for Active LearningabstractActive Learning for discriminative models has largely been studied with the focus on individual samples, with less emphasis on how classes are distributed or which classes are hard to deal with. In this work, we show that this is harmful. We propose a method based on the Bayes’ rule, that can naturally incorporate class imbalance into the Active Learning framework. We derive that three terms should be considered together when estimating the probability of a classifier making a mistake for a given sample; i) probability of mislabelling a class, ii) likelihood of the data given a predicted class, and iii) the prior probability on the abundance of a predicted class. Implementing these terms requires a generative model and an intractable likelihood estimation. Therefore, we train a Variational Auto Encoder (VAE) for this purpose. To further tie the VAE with the classifier and facilitate VAE training, we use the classifiers’ deep feature representations as input to the VAE. By considering all three probabilities, among them, especially the data imbalance, we can substantially improve the potential of existing methods under limited data budget. We show that our method can be applied to classification tasks on multiple different datasets – including one that is a real-world dataset with heavy data imbalance – significantly outperforming the state of the art. Jongwon Choi 0002, Kwang Moo Yi, Jinho Choo, Byoungjip Kim, Jin-Yeop Chang, Youngjune Gwon, Hyung Jin Chang |
CVPR | 8 |
| 2021 | FixBi: Bridging Domain Spaces for Unsupervised Domain AdaptationabstractUnsupervised domain adaptation (UDA) methods for learning domain invariant representations have achieved remarkable progress. However, most of the studies were based on direct adaptation from the source domain to the target domain and have suffered from large domain discrepancies. In this paper, we propose a UDA method that effectively handles such large domain discrepancies. We introduce a fixed ratio-based mixup to augment multiple intermediate domains between the source and target domain. From the augmented-domains, we train the source-dominant model and the target-dominant model that have complementary characteristics. Using our confidence-based learning methodologies, e.g., bidirectional matching with high-confidence predictions and self-penalization using low-confidence predictions, the models can learn from each other or from its own results. Through our proposed methods, the models gradually transfer domain knowledge from the source to the target domain. Extensive experiments demonstrate the superiority of our proposed method on three public benchmarks: Office-31, Office-Home, and VisDA-2017.1 Jaemin Na, Heechul Jung, Hyung Jin Chang, Wonjun Hwang |
CVPR | 3 |
| 2021 | Unsupervised Hyperbolic Representation Learning via Message Passing Auto-EncodersabstractMost of the existing literature regarding hyperbolic embedding concentrate upon supervised learning, whereas the use of unsupervised hyperbolic embedding is less well explored. In this paper, we analyze how unsupervised tasks can benefit from learned representations in hyperbolic space. To explore how well the hierarchical structure of un-labeled data can be represented in hyperbolic spaces, we design a novel hyperbolic message passing auto-encoder whose overall auto-encoding is performed in hyperbolic space. The proposed model conducts auto-encoding the networks via fully utilizing hyperbolic geometry in message passing. Through extensive quantitative and qualitative analyses, we validate the properties and benefits of the unsupervised hyperbolic representations. Codes are available at https://github.com/junhocho/HGCAE. Jiwoong Park, Hyung Jin Chang, Jin Young Choi 0002 |
CVPR | 3 |
| 2021 | Nesterov Accelerated ADMM for Fast Diffeomorphic Image Registration
Alexander Thorley, Xi Jia, Hyung Jin Chang, Karina Bunting, Victoria Stoll, Antonio M. Simoes Monteiro de Marvao, Declan P. O'Regan, Georgios V. Gkoutos, Dipak Kotecha, Jinming Duan 0001 |
MICCAI (4) | 3 |
| 2021 | Motion-aware ensemble of three-mode trackers for unmanned aerial vehicles
Kyuewang Lee, Hyung Jin Chang, Jongwon Choi 0002, Byeongho Heo, Ales Leonardis, Jin Young Choi 0002 |
Mach. Vis. Appl. | 2 |
| 2020 | G2L-Net: Global to Local Network for Real-Time 6D Pose Estimation With Embedding Vector FeaturesabstractIn this paper, we propose a novel real-time 6D object pose estimation framework, named G2L-Net. Our network operates on point clouds from RGB-D detection in a divide-and-conquer fashion. Specifically, our network consists of three steps. First, we extract the coarse object point cloud from the RGB-D image by 2D detection. Second, we feed the coarse object point cloud to a translation localization network to perform 3D segmentation and object translation prediction. Third, via the predicted segmentation and translation, we transfer the fine object point cloud into a local canonical coordinate, in which we train a rotation localization network to estimate initial object rotation. In the third step, we define point-wise embedding vector features to capture viewpoint-aware information. To calculate more accurate rotation, we adopt a rotation residual estimator to estimate the residual between initial rotation and ground truth, which can boost initial pose estimation performance. Our proposed G2L-Net is real-time despite the fact multiple steps are stacked via the proposed coarse-to-fine framework. Extensive experiments on two benchmark datasets show that G2L-Net achieves state-of-the-art performance in terms of both accuracy and speed. Wei Chen 0092, Xi Jia, Hyung Jin Chang, Jinming Duan 0001, Ales Leonardis |
CVPR | 3 |
| 2020 | Combining Task Predictors via Enhancing Joint Predictability
Kwang In Kim, Christian Richardt, Hyung Jin Chang |
ECCV (16) | 3 |
| 2020 | SeqHAND: RGB-Sequence-Based 3D Hand Pose and Shape Estimation
John Yang 0001, Hyung Jin Chang, Seungeui Lee, Nojun Kwak |
ECCV (12) | 2 |
| 2020 | PointPoseNet: Point Pose Network for Robust 6D Object Pose EstimationabstractIn this paper, we propose a novel pipeline to estimate 6D object pose from RGB-D images of known objects present in complex scenes. The pipeline directly operates on raw point clouds extracted from RGB-D scans. Specifically, our method takes the point cloud as input and regresses the point-wise unit vectors pointing to the 3D keypoints. We then use these vectors to generate keypoint hypotheses from which the 6D object pose hypotheses are computed. Finally, we select the best 6D object pose from the hypotheses based on a proposed scoring mechanism with geometry constraints. Extensive experiments show that the proposed method is robust against the variety in object shape and appearance as well as occlusions between objects, and that our method outperforms the state-of-the-art methods on the LINEMOD and Occlusion LINEMOD datasets. Wei Chen 0092, Jinming Duan 0001, Hector Basevi, Hyung Jin Chang, Ales Leonardis |
WACV | 4 |
| 2020 | Body Pose Sonification for a View-Independent Auditory Aid to Blind Rock ClimbersabstractRock climbing is a sport in which blind people have traditionally found it extremely difficult to excel due to the high degree of visual problem solving required, and also the requirement to climb with a sighted assistant. We present a system which automates the role of the sighted assistant in order to provide blind people with the freedom to climb and train on their own. We address climbing-specific limitations of a state-of-the-art skeleton tracking system, and discuss the ways in which we mitigated these limitations using post-processing techniques tuned specially for a climbing scenario. We also describe the auditory feedback system used to instruct the blind climber, and demonstrate that a user can learn to follow it in a relatively short time by showing a significant improvement in performance over just a few trials with the system. Joseph Ramsay, Hyung Jin Chang |
WACV | 2 |
| 2019 | PMnet: Learning of Disentangled Pose and Movement for Unsupervised Motion Retargeting
Jongin Lim 0002, Hyung Jin Chang, Jin Young Choi 0002 |
BMVC | 2 |
| 2019 | Joint Manifold Diffusion for Combining Predictions on Decoupled ObservationsabstractWe present a new predictor combination algorithm that improves a given task predictor based on potentially relevant reference predictors. Existing approaches are limited in that, to discover the underlying task dependence, they either require known parametric forms of all predictors or access to a single fixed dataset on which all predictors are jointly evaluated. To overcome these limitations, we design a new non-parametric task dependence estimation procedure that automatically aligns evaluations of heterogeneous predictors across disjoint feature sets. Our algorithm is instantiated as a robust manifold diffusion process that jointly refines the estimated predictor alignments and the corresponding task dependence. We apply this algorithm to the relative attributes ranking problem and demonstrate that it not only broadens the application range of predictor combination approaches but also outperforms existing methods even when applied to classical predictor combination settings. Kwang In Kim, Hyung Jin Chang |
CVPR | 2 |
| 2019 | Symmetric Graph Convolutional Autoencoder for Unsupervised Graph Representation LearningabstractWe propose a symmetric graph convolutional autoencoder which produces a low-dimensional latent representation from a graph. In contrast to the existing graph autoencoders with asymmetric decoder parts, the proposed autoencoder has a newly designed decoder which builds a completely symmetric autoencoder form. For the reconstruction of node features, the decoder is designed based on Laplacian sharpening as the counterpart of Laplacian smoothing of the encoder, which allows utilizing the graph structure in the whole processes of the proposed autoencoder architecture. In order to prevent the numerical instability of the network caused by the Laplacian sharpening introduction, we further propose a new numerically stable form of the Laplacian sharpening by incorporating the signed graphs. In addition, a new cost function which finds a latent representation and a latent affinity matrix simultaneously is devised to boost the performance of image clustering tasks. The experimental results on clustering, link prediction and visualization tasks strongly support that the proposed model is stable and outperforms various state-of-the-art algorithms. Jiwoong Park, Minsik Lee 0001, Hyung Jin Chang, Kyuewang Lee, Jin Young Choi 0002 |
ICCV | 3 |
| 2018 | Context-Aware Deep Feature Compression for High-Speed Visual TrackingabstractWe propose a new context-aware correlation filter based tracking framework to achieve both high computational speed and state-of-the-art performance among real-time trackers. The major contribution to the high computational speed lies in the proposed deep feature compression that is achieved by a context-aware scheme utilizing multiple expert auto-encoders; a context in our framework refers to the coarse category of the tracking target according to appearance patterns. In the pre-training phase, one expert auto-encoder is trained per category. In the tracking phase, the best expert auto-encoder is selected for a given target, and only this auto-encoder is used. To achieve high tracking performance with the compressed feature map, we introduce extrinsic denoising processes and a new orthogonality loss term for pre-training and fine-tuning of the expert autoencoders. We validate the proposed context-aware framework through a number of experiments, where our method achieves a comparable performance to state-of-the-art trackers which cannot run in real-time, while running at a significantly fast speed of over 100 fps. Jongwon Choi 0002, Hyung Jin Chang, Tobias Fischer 0001, Sangdoo Yun, Kyuewang Lee, Jiyeoup Jeong, Yiannis Demiris, Jin Young Choi 0002 |
CVPR | 2 |
| 2018 | RT-GENE: Real-Time Eye Gaze Estimation in Natural Environments
Tobias Fischer 0001, Hyung Jin Chang, Yiannis Demiris |
ECCV (10) | 2 |
| 2018 | Transferring Visuomotor Learning from Simulation to the Real World for Robotics Manipulation TasksabstractHand-eye coordination is a requirement for many manipulation tasks including grasping and reaching. However, accurate hand-eye coordination has shown to be especially difficult to achieve in complex robots like the iCub humanoid. In this work, we solve the hand-eye coordination task using a visuomotor deep neural network predictor that estimates the arm's joint configuration given a stereo image pair of the arm and the underlying head configuration. As there are various unavoidable sources of sensing error on the physical robot, we train the predictor on images obtained from simulation. The images from simulation were modified to look realistic using an image-to-image translation approach. In various experiments, we first show that the visuomotor predictor provides accurate joint estimates of the iCub's hand in simulation. We then show that the predictor can be used to obtain the systematic error of the robot's joint measurements on the physical iCub robot. We demonstrate that a calibrator can be designed to automatically compensate this error. Finally, we validate that this enables accurate reaching of objects while circumventing manual fine-calibration of the robot. Phuong D. H. Nguyen, Tobias Fischer 0001, Hyung Jin Chang, Ugo Pattacini, Giorgio Metta, Yiannis Demiris |
IROS | 3 |
| 2018 | Highly Articulated Kinematic Structure Estimation Combining Motion and Skeleton InformationabstractIn this paper, we present a novel framework for unsupervised kinematic structure learning of complex articulated objects from a single-view 2D image sequence. In contrast to prior motion-based methods, which estimate relatively simple articulations, our method can generate arbitrarily complex kinematic structures with skeletal topology via a successive iterative merging strategy. The iterative merge process is guided by a density weighted skeleton map which is generated from a novel object boundary generation method from sparse 2D feature points. Our main contributions can be summarised as follows: (i) An unsupervised complex articulated kinematic structure estimation method that combines motion segments with skeleton information. (ii) An iterative fine-to-coarse merging strategy for adaptive motion segmentation and structural topology embedding. (iii) A skeleton estimation method based on a novel silhouette boundary generation from sparse feature points using an adaptive model selection method. (iv) A new highly articulated object dataset with ground truth annotation. We have verified the effectiveness of our proposed method in terms of computational time and estimation accuracy through rigorous experiments with multiple datasets. Our experiments show that the proposed method outperforms state-of-the-art methods both quantitatively and qualitatively. Hyung Jin Chang, Yiannis Demiris |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2018 | Learning Kinematic Structure Correspondences Using Multi-Order SimilaritiesabstractIn this paper, we present a novel framework for finding the kinematic structure correspondences between two articulated objects in videos via hypergraph matching. In contrast to appearance and graph alignment based matching methods, which have been applied among two similar static images, the proposed method finds correspondences between two dynamic kinematic structures of heterogeneous objects in videos. Thus our method allows matching the structure of objects which have similar topologies or motions, or a combination of the two. Our main contributions can be summarised as follows: (i) casting the kinematic structure correspondence problem into a hypergraph matching problem by incorporating multi-order similarities with normalising weights, (ii) introducing a structural topology similarity measure by aggregating topology constrained subgraph isomorphisms, (iii) measuring kinematic correlations between pairwise nodes, and (iv) proposing a combinatorial local motion similarity measure using geodesic distance on the Riemannian manifold. We demonstrate the robustness and accuracy of our method through a number of experiments on synthetic and real data, outperforming various other state of the art methods. Our method is not limited to a specific application nor sensor, and can be used as building block in applications such as action recognition, human motion retargeting to robots, and articulated object manipulation amongst others. Hyung Jin Chang, Tobias Fischer 0001, Maxime Petit, Martina Zambelli, Yiannis Demiris |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2017 | Attentional Correlation Filter Network for Adaptive Visual TrackingabstractWe propose a new tracking framework with an attentional mechanism that chooses a subset of the associated correlation filters for increased robustness and computational efficiency. The subset of filters is adaptively selected by a deep attentional network according to the dynamic properties of the tracking target. Our contributions are manifold, and are summarised as follows: (i) Introducing the Attentional Correlation Filter Network which allows adaptive tracking of dynamic targets. (ii) Utilising an attentional network which shifts the attention to the best candidate modules, as well as predicting the estimated accuracy of currently inactive modules. (iii) Enlarging the variety of correlation filters which cover target drift, blurriness, occlusion, scale changes, and flexible aspect ratio. (iv) Validating the robustness and efficiency of the attentional mechanism for visual tracking through a number of experiments. Our method achieves similar performance to non real-time trackers, and state-of-the-art performance amongst real-time trackers. Jongwon Choi 0002, Hyung Jin Chang, Sangdoo Yun, Tobias Fischer 0001, Yiannis Demiris, Jin Young Choi 0002 |
CVPR | 2 |
| 2017 | Variational Autoencoded Regression: High Dimensional Regression of Visual Data on Complex ManifoldabstractThis paper proposes a new high dimensional regression method by merging Gaussian process regression into a variational autoencoder framework. In contrast to other regression methods, the proposed method focuses on the case where output responses are on a complex high dimensional manifold, such as images. Our contributions are summarized as follows: (i) A new regression method estimating high dimensional image responses, which is not handled by existing regression algorithms, is proposed. (ii) The proposed regression method introduces a strategy to learn the latent space as well as the encoder and decoder so that the result of the regressed response in the latent space coincide with the corresponding response in the data space. (iii) The proposed regression is embedded into a generative model, and the whole procedure is developed by the variational autoencoder framework. We demonstrate the robustness and effectiveness of our method through a number of experiments on various visual data regression problems. Young Joon Yoo, Sangdoo Yun, Hyung Jin Chang, Yiannis Demiris, Jin Young Choi 0002 |
CVPR | 3 |
| 2017 | Latent Regression Forest: Structured Estimation of 3D Hand PosesabstractIn this paper we present the latent regression forest (LRF), a novel framework for real-time, 3D hand pose estimation from a single depth image. Prior discriminative methods often fall into two categories: holistic and patch-based. Holistic methods are efficient but less flexible due to their nearest neighbour nature. Patch-based methods can generalise to unseen samples by consider local appearance only. However, they are complex because each pixel need to be classified or regressed during testing. In contrast to these two baselines, our method can be considered as a structured coarse-to-fine search, starting from the centre of mass of a point cloud until locating all the skeletal joints. The searching process is guided by a learnt latent tree model which reflects the hierarchical topology of the hand. Our main contributions can be summarised as follows: (i) Learning the topology of the hand in an unsupervised, data-driven manner. (ii) A new forest-based, discriminative framework for structured search in images, as well as an error regression step to avoid error accumulation. (iii) A new multi-view hand pose dataset containing 180 K annotated images from 10 different subjects. Our experiments on two datasets show that the LRF outperforms baselines and prior arts in both accuracy and efficiency. Danhang Tang, Hyung Jin Chang, Alykhan Tejani, Tae-Kyun Kim 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2016 | Kinematic Structure Correspondences via Hypergraph MatchingabstractIn this paper, we present a novel framework for finding the kinematic structure correspondence between two objects in videos via hypergraph matching. In contrast to prior appearance and graph alignment based matching methods which have been applied among two similar static images, the proposed method finds correspondences between two dynamic kinematic structures of heterogeneous objects in videos. Our main contributions can be summarised as follows: (i) casting the kinematic structure correspondence problem into a hypergraph matching problem, incorporating multi-order similarities with normalising weights, (ii) a structural topology similarity measure by a new topology constrained subgraph isomorphism aggregation, (iii) a kinematic correlation measure between pairwise nodes, and (iv) a combinatorial local motion similarity measure using geodesic distance on the Riemannian manifold. We demonstrate the robustness and accuracy of our method through a number of experiments on complex articulated synthetic and real data. Hyung Jin Chang, Tobias Fischer 0001, Maxime Petit, Martina Zambelli, Yiannis Demiris |
CVPR | 1 |
| 2016 | Visual Tracking Using Attention-Modulated Disintegration and IntegrationabstractIn this paper, we present a novel attention-modulated visual tracking algorithm that decomposes an object into multiple cognitive units, and trains multiple elementary trackers in order to modulate the distribution of attention according to various feature and kernel types. In the integration stage it recombines the units to memorize and recognize the target object effectively. With respect to the elementary trackers, we present a novel attentional feature-based correlation filter (AtCF) that focuses on distinctive attentional features. The effectiveness of the proposed algorithm is validated through experimental comparison with state-of-theart methods on widely-used tracking benchmark datasets. Jongwon Choi 0002, Hyung Jin Chang, Jiyeoup Jeong, Yiannis Demiris, Jin Young Choi 0002 |
CVPR | 2 |
| 2016 | Iterative path optimisation for personalised dressing assistance using vision and force informationabstractWe propose an online iterative path optimisation method to enable a Baxter humanoid robot to assist human users to dress. The robot searches for the optimal personalised dressing path using vision and force sensor information: vision information is used to recognise the human pose and model the movement space of upper-body joints; force sensor information is used for the robot to detect external force resistance and to locally adjust its motion. We propose a new stochastic path optimisation method based on adaptive moment estimation. We first compare the proposed method with other path optimisation algorithms on synthetic data. Experimental results show that the performance of the method achieves the smallest error with fewer iterations and less computation time. We also evaluate real-world data by enabling the Baxter robot to assist real human users with their dressing. Yixing Gao 0001, Hyung Jin Chang, Yiannis Demiris |
IROS | 2 |
| 2016 | Transition Hough forest for trajectory-based action recognitionabstractIn this paper, we propose a new discriminative framework based on Hough forests that enables us to efficiently recognize and localize sequential data in the form of spatio-temporal trajectories. Contrary to traditional decision forest-based methods where predictions are made independently of its output temporal context, we introduce the concept of "transition", which enforces the temporal coherence of estimations and further enhances the discrimination between action classes. We start applying our proposed framework to the problem of recognizing and localizing fingertip written trajectories in mid-air using an egocentric camera. To this purpose, we present a new challenging dataset that allows us to evaluate and compare our method with previous approaches. Finally, we apply our framework to general human action recognition using local spatio-temporal trajectories obtaining comparable to state-of-the-art performance on a public benchmark. Guillermo Garcia-Hernando, Hyung Jin Chang, Ismael Serrano, Oscar Déniz-Suárez, Tae-Kyun Kim 0001 |
WACV | 2 |
| 2016 | Spatio-Temporal Hough Forest for efficient detection-localisation-recognition of fingerwriting in egocentric camera
Hyung Jin Chang, Guillermo Garcia-Hernando, Danhang Tang, Tae-Kyun Kim 0001 |
Comput. Vis. Image Underst. | 1 |
| 2015 | Unsupervised learning of complex articulated kinematic structures combining motion and skeleton informationabstractIn this paper we present a novel framework for unsupervised kinematic structure learning of complex articulated objects from a single-view image sequence. In contrast to prior motion information based methods, which estimate relatively simple articulations, our method can generate arbitrarily complex kinematic structures with skeletal topology by a successive iterative merge process. The iterative merge process is guided by a skeleton distance function which is generated from a novel object boundary generation method from sparse points. Our main contributions can be summarised as follows: (i) Unsupervised complex articulated kinematic structure learning by combining motion and skeleton information. (ii) Iterative fine-to-coarse merging strategy for adaptive motion segmentation and structure smoothing. (iii) Skeleton estimation from sparse feature points. (iv) A new highly articulated object dataset containing multi-stage complexity with ground truth. Our experiments show that the proposed method out-performs state-of-the-art methods both quantitatively and qualitatively. Hyung Jin Chang, Yiannis Demiris |
CVPR | 1 |
| 2015 | User modelling for personalised dressing assistance by humanoid robotsabstractAssistive robots can improve the well-being of disabled or frail human users by reducing the burden that activities of daily living impose on them. To enable personalised assistance, such robots benefit from building a user-specific model, so that the assistance is customised to the particular set of user abilities. In this paper, we present an end-to-end approach for home-environment assistive humanoid robots to provide personalised assistance through a dressing application for users who have upper-body movement limitations. We use randomised decision forests to estimate the upper-body pose of users captured by a top-view depth camera, and model the movement space of upper-body joints using Gaussian mixture models. The movement space of each upper-body joint consists of regions with different reaching capabilities. We propose a method which is based on real-time upper-body pose and user models to plan robot motions for assistive dressing. We validate each part of our approach and test the whole system, allowing a Baxter humanoid robot to assist human to wear a sleeveless jacket. Yixing Gao 0001, Hyung Jin Chang, Yiannis Demiris |
IROS | 2 |
| 2015 | STARE: Spatio-Temporal Attention Relocation for Multiple Structured Activities DetectionabstractWe present a spatio-temporal attention relocation (STARE) method, an information-theoretic approach for efficient detection of simultaneously occurring structured activities. Given multiple human activities in a scene, our method dynamically focuses on the currently most informative activity. Each activity can be detected without complete observation, as the structure of sequential actions plays an important role on making the system robust to unattended observations. For such systems, the ability to decide where and when to focus is crucial to achieving high detection performances under resource bounded condition. Our main contributions can be summarized as follows: 1) information-theoretic dynamic attention relocation framework that allows the detection of multiple activities efficiently by exploiting the activity structure information and 2) a new high-resolution data set of temporally-structured concurrent activities. Our experiments on applications show that the STARE method performs efficiently while maintaining a reasonable level of accuracy. Kyuhwa Lee, Dimitri Ognibene, Hyung Jin Chang, Tae-Kyun Kim 0001, Yiannis Demiris |
IEEE Trans. Image Process. | 3 |
| 2015 | 3D Finger CAPE: Clicking Action and Position Estimation under Self-Occlusions in Egocentric ViewpointabstractIn this paper we present a novel framework for simultaneous detection of click action and estimation of occluded fingertip positions from egocentric viewed single-depth image sequences. For the detection and estimation, a novel probabilistic inference based on knowledge priors of clicking motion and clicked position is presented. Based on the detection and estimation results, we were able to achieve a fine resolution level of a bare hand-based interaction with virtual objects in egocentric viewpoint. Our contributions include: (i) a rotation and translation invariant finger clicking action and position estimation using the combination of 2D image-based fingertip detection with 3D hand posture estimation in egocentric viewpoint. (ii) a novel spatio-temporal random forest, which performs the detection and estimation efficiently in a single framework. We also present (iii) a selection process utilizing the proposed clicking action detection and position estimation in an arm reachable AR/VR space, which does not require any additional device. Experimental results show that the proposed method delivers promising performance under frequent self-occlusions in the process of selecting objects in AR/VR space whilst wearing an egocentric-depth camera-attached HMD. Youngkyoon Jang, Seungtak Noh, Hyung Jin Chang, Tae-Kyun Kim 0001, Woontack Woo |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2014 | Latent Regression Forest: Structured Estimation of 3D Articulated Hand PostureabstractIn this paper we present the Latent Regression Forest (LRF), a novel framework for real-time, 3D hand pose estimation from a single depth image. In contrast to prior forest-based methods, which take dense pixels as input, classify them independently and then estimate joint positions afterwards, our method can be considered as a structured coarse-to-fine search, starting from the centre of mass of a point cloud until locating all the skeletal joints. The searching process is guided by a learnt Latent Tree Model which reflects the hierarchical topology of the hand. Our main contributions can be summarised as follows: (i) Learning the topology of the hand in an unsupervised, data-driven manner. (ii) A new forest-based, discriminative framework for structured search in images, as well as an error regression step to avoid error accumulation. (iii) A new multi-view hand pose dataset containing 180K annotated images from 10 different subjects. Our experiments show that the LRF out-performs state-of-the-art methods in both accuracy and efficiency. Danhang Tang, Hyung Jin Chang, Alykhan Tejani, Tae-Kyun Kim 0001 |
CVPR | 2 |
| 2014 | Robust action recognition using local motion and group sparsity
Jungchan Cho, Minsik Lee 0001, Hyung Jin Chang, Songhwai Oh |
Pattern Recognit. | 3 |
| 2013 | Action Chart: A Representation for Efficient Recognition of Complex ActivityabstractIn this paper we propose an efficient method for the recognition of long and complex action streams. First, we design a new motion feature flow descriptor by composing low-level local features. Then a new data embedding method is developed in order to represent the motion flow as an one-dimensional sequence, whilst preserving useful motion information for recognition. Finally attentional motion spots (AMSs) are defined to automatically detect meaningful motion changes from the embedded one-dimensional sequence. An unsupervised learning strategy based on expectation maximization and a weighted Gaussian mixture model is then applied to the AMSs for each action class, resulting in an action representation which we refer to as Action Chart. The Action Chart is then used efficiently for recognizing each action class. Through comparison with the state-of-the-art methods, experimental results show that the Action Chart gives promising recognition performance with low computational load and can be used for abstracting long video sequences. Hyung Jin Chang, Jungchan Cho, Songhwai Oh, Kwang Moo Yi, Jin Young Choi 0002 |
BMVC | 1 |
| 2013 | Initialization-Insensitive Visual Tracking through Voting with Salient Local FeaturesabstractIn this paper we propose an object tracking method in case of inaccurate initializations. To track objects accurately in such situation, the proposed method uses "motion saliency" and "descriptor saliency" of local features and performs tracking based on generalized Hough transform (GHT). The proposed motion saliency of a local feature emphasizes features having distinctive motions, compared to the motions which are not from the target object. The descriptor saliency emphasizes features which are likely to be of the object in terms of its feature descriptors. Through these saliencies, the proposed method tries to "learn and find" the target object rather than looking for what was given at initialization, giving robust results even with inaccurate initializations. Also, our tracking result is obtained by combining the results of each local feature of the target and the surroundings with GHT voting, thus is robust against severe occlusions as well. The proposed method is compared against nine other methods, with nine image sequences, and hundred random initializations. The experimental results show that our method outperforms all other compared methods. Kwang Moo Yi, Hawook Jeong, Byeongho Heo, Hyung Jin Chang, Jin Young Choi 0002 |
ICCV | 4 |
| 2013 | Abstracted radon profiles for fingerprint recognitionabstractConventional minutiae-based fingerprint recognition approaches consider only local characteristics and their accuracy dramatically decreases as the number of available minutiae decreases. We propose new features based on Abstracted Radon Profile (ARP). Proposed method uses global properties of an image and it does not necessitate any heavy preprocessing as in classical methods. By using independent gradual patching via proposed multilayer architecture, local characteristics of an image are also preserved. ARP features have an advantage of being robust to zero mean additive noise. For sparse signal representation, dictionary is constructed from the ARP features of the training samples. Recognition is done by ℓ1-minimization with quadratic constraints, so this framework can handle dense noise by exploiting the fact that these errors are often sparse. Experimental results in assessing recognition performance demonstrate the proposed approach outperforms the conventional approaches in correlation and distance based comparisons. Computational time comparison result shows the proposed feature is more efficient than brute-force method of image alignment and promising for handling other pattern recognition problems as well. Tushar Sandhan, Hyung Jin Chang, Jin Young Choi 0002 |
ICIP | 2 |
| 2012 | Active attentional sampling for speed-up of background subtractionabstractIn this paper, we present an active sampling method to speed up conventional pixel-wise background subtraction algorithms. The proposed active sampling strategy is designed to focus on attentional region such as foreground regions. The attentional region is estimated by detection results of previous frame in a recursive probabilistic way. For the estimation of the attentional region, we propose a foreground probability map based on temporal, spatial, and frequency properties of foregrounds. By using this foreground probability map, active attentional sampling scheme is developed to make a minimal sampling mask covering almost foregrounds. The effectiveness of the proposed active sampling method is shown through various experiments. The proposed masking method successfully speeds up pixel-wise background subtraction methods approximately 6.6 times without deteriorating detection performance. Also realtime detection with Full HD video is successfully achieved through various conventional background subtraction algorithms. Hyung Jin Chang, Hawook Jeong, Jin Young Choi 0002 |
CVPR | 1 |
| 2012 | Robust moving object detection against fast illumination change
JinMin Choi, Hyung Jin Chang, Yung Jun Yoo, Jin Young Choi 0002 |
Comput. Vis. Image Underst. | 2 |
| 2011 | Modeling of moving object trajectory by spatio-temporal learning for abnormal behavior detectionabstractThis paper proposes a trajectory analysis method by handling the spatio-temporal property of trajectory. Not using similarity measures of two trajectories, our model analyzes overall path of a trajectory. Learning of spatio property is presented as semantic regions (e.g. go straight, turn left, turn right) that are clustered effectively using topic model. The temporal order of observations on a trajectory is taken into account using HMM for detecting global anomaly. Results of experiments show that modeling of semantic region and detecting of unusual trajectories are successful even in complex scenes. Hawook Jeong, Hyung Jin Chang, Jin Young Choi 0002 |
AVSS | 2 |
| 2011 | Tracking failure detection by imitating human visual perceptionabstractIn this paper, we present a tracking failure detection method by imitating human visual system. By adopting log-polar transformation, we could simulate properties of retina image, such as rotation and scaling invariance and foveal predominance. The rotation and scaling invariance helps to reduce false alarms caused by pose changes and intensify translational changes. Foveal predominant property helps to detect the tracking failing moment by amplifying the resolution around focus (tracking box center) and blurring the peripheries. Each ganglion cell corresponds to a pixel of log-polar image, and its adaptation is modeled as Gaussian mixture model. Its validity is shown through various experiments. Hyung Jin Chang, Myoung Soo Park, Hawook Jeong, Jin Young Choi 0002 |
ICIP | 1 |
| 2008 | Fast incremental learning for one-class support vector classifier using sample margin informationabstractIn this paper, we present a fast incremental one-class classifier algorithm for large scale problems. The proposed method reduces space and time complexities by reducing training set size during the training procedure using a criterion based on sample margin. After introducing the sample margin concept, we present the proposed algorithm and apply it to face detection database to show its efficiency and validity. Pyo Jae Kim, Hyung Jin Chang, Jin Young Choi 0002 |
ICPR | 2 |
| 2007 | Fast Support Vector Data Description Using K-Means Clustering
Pyo Jae Kim, Hyung Jin Chang, Dong Sung Song, Jin Young Choi 0002 |
ISNN (3) | 2 |