Hung-Kuo Chu

dblp:67/2946 · DBLP profile ↗
← Back
51ranked-venue papers
3as first author
17since 2021 · last 2025
0000-0001-7153-4411ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 46 · 3 first-author · 13 since 2021Artificial intelligence and machine learning · 6 · 3 since 2021Human-computer interaction and ubiquitous computing · 4 · 2 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Computer networks · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Bridging Coaching Knowledge and AI Feedback to Enhance Motor Learning in Basketball Shooting Mechanics Through a Knowledge-Based SOP Framework
Jian-Jia Weng, Calvin Ku, Jo-Chien Wang, Chih-Jen Cheng, Tica Lin, Yu-An Su, Tsung-Hsun Tsai, You-Yi Lin, Lun-Wei Ku, Hung-Kuo Chu, Min-Chun Hu 0001
CHI10
2025 Fine-grained Stroke Recognition in Broadcast Table Tennis Videos with ATDT
abstract
This study introduces an automated system for fine-grained stroke recognition in broadcast table tennis videos, designed to address challenges in manual annotation and tactical analysis during international competitions. The proposed framework integrates an Adaptive Temporal Difference Model with a Transformer Encoder (ATDT), leveraging a combination of Temporal Difference Networks (TDN) and Temporal Adaptive Modules (TAM) to enhance spatial and temporal feature extraction. To enhance feature discriminability, we employ supervised contrastive learning, which promotes better representation learning for fine-grained action recognition. The system is divided into two primary modules: the Action Segmentation Module (ASM) and the Action Recognition Module (ARM). ASM precisely identifies the start and end times of each stroke action by incorporating ball trajectory analysis to identify precise hit timings and placements. The precise segmentation facilitates the subsequent ARM to implement a three-stage recognition process: forehand and backhand classification, group-based classification, and intra-group action classification. This hierarchical approach improves the system’s ability to differentiate between subtle stroke variations, even under the constraints of low-resolution broadcast footage. To validate the framework, the MISTT dataset was collected, comprising 3,618 stroke action clips from 18 international matches, with professional player annotations. The proposed ATDT model outperformed existing methods, achieving a top-1 accuracy improvement of 18% for forehand strokes and 25.58% for backhand strokes compared to baseline models. Moreover, our automatic annotation system takes only 1/30 of the time compared to the manual annotation process, demonstrating its efficiency.
Tang-Chen Chang, Duen-Chian Jheng, Hsuan-Ya Liang, Bill Louis Harchan, Pu Ching, Tsung-Hsun Tsai, Chih-Yi Chang, Te-Cheng Wu, Yung-Hui Li, Tse-Yu Pan, Hung-Kuo Chu, Min-Chun Hu 0001
ACM Trans. Multim. Comput. Commun. Appl.11
2024 Panelformer: Sewing Pattern Reconstruction from 2D Garment Images
abstract
In this paper, we present a novel approach for reconstructing garment sewing patterns from 2D garment images. Our method addresses the challenge of handling occlusion in 2D images by leveraging the symmetric and correlated nature of garment panels. We introduce a transformer-based deep neural network called Panelformer that learns the parametric space of garment sewing patterns. The network comprises two components: the panel transformer and the stitch predictor. The panel transformer estimates the parametric panel shapes, including the occluded panels, by learning from the visible ones. The stitch predictor determines the stitching information among the predicted panels, enabling the reconstruction of the complete garment. To mitigate the overfitting problem caused by strong panel correlations, we propose two tailor-made data augmentation techniques: panel masking and garment mixing. These techniques generate a wider variety of panel combinations, enhancing the model’s robustness and generalization capability. We evaluate the effectiveness of Panelformer using a synthetic dataset with diverse garment types. The experimental results demonstrate that our method outperforms competing baselines and achieves comparable performance to NeuralTailor, which operates on 3D point cloud data. This validates the efficacy of our approach in the context of garment sewing pattern reconstruction. By utilizing 2D images as input, our method expands the potential applications of garment modeling and offers easy accessibility to end users. Our code is available online1.
Cheng-Hsiu Chen, Jheng-Wei Su, Min-Chun Hu 0001, Chih-Yuan Yao, Hung-Kuo Chu
WACV5
2024 VisionCoach: Design and Effectiveness Study on VR Vision Training for Basketball Passing
abstract
Vision Training is important for basketball players to effectively search for teammates who has wide-open opportunities to shoot, observe the defenders around the wide-open teammates and quickly choose a proper way to pass the ball to the most suitable one. We develop an immersive virtual reality (VR) system called VisionCoach to simulate the player's viewing perspective and generate three designed systematic vision training tasks to benefit the cultivating procedure. By recording the player's eye gazing and dribbling video sequence, the proposed system can analyze the vision-related behavior to understand the training effectiveness. To demonstrate the proposed VR training system can facilitate the cultivation of vision ability, we recruited 14 experienced players to participate in a 6-week between-subject study, and conducted a study by comparing the most frequently used 2D vision training method called Vision Performance Enhancement (VPE) program with the proposed system. Qualitative experiences and quantitative training results are reported to show that the proposed immersive VR training system can effectively improve player's vision ability in terms of gaze behavior and dribbling stability. Furthermore, training in the VR-VisionCoach Condition can transfer the learned abilities to real scenario more easily than training in the 2D-VPE Condition.
Pin-Xuan Liu, Tse-Yu Pan, Hsin-Shih Lin, Hung-Kuo Chu, Min-Chun Hu 0001
IEEE Trans. Vis. Comput. Graph.4
2023 SOFA: Style-based One-shot 3D Facial Animation Driven by 2D landmarks
abstract
We propose a 2D landmark-driven 3D facial animation framework trained without the need of 3D facial dataset. Our method decomposes the 3D facial avatar into geometry and texture. Given 2D landmarks as input, our models learn to estimate the parameters of FLAME and transfer the target texture into different facial expressions. The experiments show that our method achieves remarkable results. Using 2D landmarks as input data, our method has the potential to be deployed in a scenario that suffered from obtaining full RGB facial images (e.g., occluded by VR Head-mounted Display).
Pu Ching, Hung-Kuo Chu, Min-Chun Hu 0001
ICMR2
2023 Offensive Tactics Recognition in Broadcast Basketball Videos Based on 2D Camera View Player Heatmaps
abstract
It is essential for sports teams to review their offensive and defensive tactical execution performance as well as understand their opponents’ tactics in order to identify effective counterattack strategies. This study focuses on basketball offensive tactics recognition based on 2D camera view heatmaps. Most of the current tactics recognition methods learn the spatiotemporal correlation of players based on top-view trajectory information. To obtain correct top-view player trajectories, robust camera calibration and player tracking techniques are indispensable. However, for broadcast videos having large camera movement, serious player occlusions, and similar players’ jerseys, it is quite challenging to obtain accurate camera parameters and player tracking results, resulting in poor tactical analysis performance. Instead of applying camera calibration and player tracking, this study attempts to design a tactics recognition method that directly predicts the tactics class from 2D camera-view player heatmaps in the inference phase. Our proposed method uses a recurrent convolutional neural network with coordinate embedding to directly identify the tactics. Moreover, an auxiliary top-view player trajectory reconstruction module is added in the training phase to acquire better latent codes to represent the tactics. The experimental results show that for both supervised and unsupervised settings, our proposed method achieves comparable accuracy to the current tactics classification methods that rely on perfect top-view trajectory input.
subst Nico, Tse-Yu Pan, Herman Prawiro, Jain-Wei Peng, Wen-Cheng Chen, Hung-Kuo Chu, Min-Chun Hu 0001
ICMR6
2023 OmniScorer: Real-Time Shot Spot Analysis for Court View Basketball Videos
abstract
We propose a real-time shot spot analysis system specifically designed for basketball videos captured from a court view perspective, even in the presence of camera movements such as panning, zooming-in, and zooming-out. Our method consists of two stages: the first stage focuses on identifying the precise frame of the shot, while the second stage predicts the shot event category (i.e., 3-point shot, 2-point shot, or free-throw) and localizes the shot spot from a top-view perspective. Compared to existing end-to-end methods for shot event prediction, our method offers significant advantages. It effectively mitigates the overfitting problem and demonstrates superior performance in predicting 3-point shot and free-throw events. To the best of our knowledge, this work is the first real-time system capable of accurately localizing shot spots in basketball games captured by a moving camera with a court view.
Yen-Pin Cheng, Tsung-Hsun Tsai, Tai-Chen Tsai, Yi-Hsuan Chiu, Hung-Kuo Chu, Min-Chun Hu 0001
MMAsia5
2023 SLIBO-Net: Floorplan Reconstruction via Slicing Box Representation with Local Geometry Regularization
abstract
This paper focuses on improving the reconstruction of 2D floorplans from unstructured 3D point clouds. We identify opportunities for enhancement over the existing methods in three main areas: semantic quality, efficient representation, and local geometric details. To address these, we presents SLIBO-Net, an innovative approach to reconstructing 2D floorplans from unstructured 3D point clouds. We propose a novel transformer-based architecture that employs an efficient floorplan representation, providing improved room shape supervision and allowing for manageable token numbers. By incorporating geometric priors as a regularization mechanism and post-processing step, we enhance the capture of local geometric details. We also propose a scale-independent evaluation metric, correcting the discrepancy in error treatment between varying floorplan sizes. Our approach notably achieves a new state-of-the-art on the Structure3D dataset. The resultant floorplans exhibit enhanced semantic plausibility, substantially improving the overall quality and realism of the reconstructions. Our code and dataset are available online.
Jheng-Wei Su, Kuei-Yu Tung, Chihan Peng, Peter Wonka, Hung-Kuo Chu
NeurIPS5
2023 Img2Logo: Generating Golden Ratio Logos from Images
abstract
Abstract Logos are one of the most important graphic design forms that use an abstracted shape to clearly represent the spirit of a community. Among various styles of abstraction, a particular golden‐ratio design is frequently employed by designers to create a concise and regular logo. In this context, designers utilize a set of circular arcs with golden ratios (i.e., all arcs are taken from circles whose radii form a geometric series based on the golden ratio) as the design elements to manually approximate a target shape. This error‐prone process requires a large amount of time and effort, posing a significant challenge for design space exploration. In this work, we present a novel computational framework that can automatically generate golden ratio logo abstractions from an input image. Our framework is based on a set of carefully identified design principles and a constrained optimization formulation respecting these principles. We also propose a progressive approach that can efficiently solve the optimization problem, resulting in a sequence of abstractions that approximate the input at decreasing levels of detail. We evaluate our work by testing on images with different formats including real photos, clip arts, and line drawings. We also extensively validate the key components and compare our results with manual results by designers to demonstrate the effectiveness of our framework. Moreover, our framework can largely benefit design space exploration via easy specification of design parameters such as abstraction levels, golden circle sizes, etc.
Kai-Wen Hsiao, Yongliang Yang 0002, Yung-Chih Chiu, Min-Chun Hu 0001, Chih-Yuan Yao, Hung-Kuo Chu
Comput. Graph. Forum6
2023 Investigating Four Navigation Aids for Supporting Navigator Performance and Independence in Virtual Reality
abstract
Turn-by-turn navigation guidance is suggested to impair users’ independent wayfinding in the physical world. However, whether this impairment issue also exists in a virtual environment is underexplored. We compare map-based and live view-based turn-by-turn navigation aids with two additional navigation aids, reference-based and orientation-based, designed to provide directional knowledge. The results of our within-subjects experiment indicate that turn-by-turn navigation aids performed worse in building spatial awareness than reference-based guidance in virtual reality. In unaided wayfinding, reference-based guidance helped users navigate most efficiently. Unexpectedly, orientation-based guidance yielded poor navigation performance, similar to the turn-by-turn navigation aids. We found that the key to skill impairment is navigators’ tendency to rely on automated instructions. This suggests that turn-by-turn navigation aids need not be avoided, but rather that caution should be exercised to avoid the tendency of mindless instruction-following. Our qualitative findings also suggest the crucial roles of the navigation context and the suitability of the navigation aid for the context.
Ting-Yu Kuo, Yung-Ju Chang, Hung-Kuo Chu
Int. J. Hum. Comput. Interact.3
2023 Image-Based OA-Style Paper Pop-Up Design via Mixed-Integer Programming
abstract
Origami architecture (OA) is a fascinating papercraft that involves only a piece of paper with cuts and folds. Interesting geometric structures 'pop up' when the paper is opened. However, manually designing such a physically valid 2D paper pop-up plan is challenging since fold lines must jointly satisfy hard spatial constraints. Existing works on automatic OA-style paper pop-up design all focused on how to generate a pop-up structure that approximates a given target 3D model. This article presents the first OA-style paper pop-up design framework that takes 2D images instead of 3D models as input. Our work is inspired by the fact that artists often use 2D profiles to guide the design process, thus benefited from the high availability of 2D image resources. Due to the lack of 3D geometry information, we perform novel theoretic analysis to ensure the foldability and stability of the resultant design. Based on a novel graph representation of the paper pop-up plan, we further propose a practical optimization algorithm via mixed-integer programming that jointly optimizes the topology and geometry of the 2D plan. We also allow the user to interactively explore the design space by specifying constraints on fold lines. Finally, we evaluate our framework on various images with interesting 2D shapes. Experiments and comparisons exhibit both the efficacy and efficiency of our framework.
Chen Liu 0012, Kai-Wen Hsiao, Ying-Miao Kuo, Hung-Kuo Chu, Yongliang Yang 0002
IEEE Trans. Vis. Comput. Graph.5
2022 Layout-Guided Indoor Panorama Inpainting with Plane-Aware Normalization
Chao-Chen Gao, Cheng-Hsiu Chen, Jheng-Wei Su, Hung-Kuo Chu
ACCV (6)4
2022 Monocular 3D Human Pose Estimation with Domain Feature Alignment and Self Training
abstract
Despite great success in 3D monocular human pose estimation, the progress of accurate prediction for unseen poses or complex backgrounds is still limited due to the lack of labeled data. In this paper, we use synthetically generated images with 3D ground truth and unlabelled real data to address this domain gap challenge. Unlike recent works that apply the adversarial loss to their models, we propose a novel domain feature alignment method (DFA) that avoids the disadvantages of unstable training and wrong alignment. In addition, our method leverages self-training with data enhancement to create robust pseudo-labels for real data. The experimental results show the effectiveness of combining self-training with our DFA method on Human 3.6M testing data without using any 3D ground truth real data.
Yan-Hong Zhang, Calvin Ku, Min-Chun Hu 0001, Hung-Kuo Chu
ICME4
2022 ScoreActuary: Hoop-Centric Trajectory-Aware Network for Fine-Grained Basketball Shot Analysis
abstract
We propose a fine-grained basketball shot analysis system called ScoreActuary to analyze the players' shot events, which can be applied to game analysis, player training, and highlight generation. Given a basketball video as input, our system first detects/segments shot candidates and then analyzes "Shot Type", "Shot Result", and "Ball Status" of each shot candidate in real-time. Our approach is composed of a customized object detector and a trajectory-aware network to learn the information of ball trajectory. Compared to the existing methods that analyze basketball shots, our algorithm can better handle videos with arbitrary camera movements while improving the accuracy. To the best of our knowledge, this work is the first system that can analyze fine-grained shot events accurately in real basketball games with arbitrary camera movements.
Ting-Yang Kao, Tse-Yu Pan, Chen-Ni Chen, Tsung-Hsun Tsai, Hung-Kuo Chu, Min-Chun Hu 0001
ACM Multimedia5
2022 BetterSight: Immersive Vision Training for Basketball Players
abstract
Vision training is important for athletes and is key to win a sports game. Traditional vision training methods are suitable for sports that focus only on the ball. For basketball, however, players need to observe multiple moving objects (i.e., ball and players) concurrently on a large court. We propose BetterSight, an immersive vision training system for basketball that not only trains the vision of the player but also requires the player to dribble the ball stably, mimicking the situation in a real basketball game. BetterSight is composed of an Interaction Module (IM), a Training Content Generation Module (TCGM), and an Analysis Module (AM). IM allows the trainee to interact with the system more intuitively based on gesture and speech rather than the controller. TCGM simulates the training scenarios based on the training configurations selected by the trainee. AM collects the video sequences capturing the trainee and the trainee's eye movements during the training phase, and then analyzes the trainee's gaze, dribbling movements, and number of dribbles. The analyzed data can be used to evaluate the training effectiveness of using the proposed BetterSight.
Pin-Xuan Liu, Tse-Yu Pan, Hsin-Shih Lin, Hung-Kuo Chu, Min-Chun Hu 0001
ACM Multimedia4
2022 GetWild: A VR Editing System with AI-Generated 3D Object and Terrain
abstract
3D environment artists typically use 2D screens and 3D modeling software to achieve their creation. However, creating 3D content using 2D tools is counterintuitive. Moreover, the process would be inefficient for junior artists in the absence of a reference. We develop a system called GetWild, which employs artificial intelligence (AI) models to generate the prototype of 3D objects/terrain and allows users to further edit the generated content in the virtual space. With the aid of AI, the user can capture an image to obtain a rough 3D object model, or start with drawing simple sketches representing the river, the mountain peak and the mountain ridge to create a 3D terrain prototype. Further, the virtual reality (VR) technique is used to provide an immersive design environment and intuitive interaction (such as painting, sculpturing, coloring, and transformation) for users to edit the generated prototypes. Compared with the existing 3D modeling software and systems, the proposed VR editing system with AI-generated 3D objects/terrain provides a more efficient way for the user to create virtual artwork.
Shing Ming Wong, Chien-Wen Chen, Tse-Yu Pan, Hung-Kuo Chu, Min-Chun Hu 0001
ACM Multimedia4
2021 Manhattan Room Layout Reconstruction from a Single $360^{\circ }$ Image: A Comparative Study of State-of-the-Art Methods
Chuhang Zou, Jheng-Wei Su, Chihan Peng, Alex Colburn, Qi Shan, Peter Wonka, Hung-Kuo Chu, Derek Hoiem
Int. J. Comput. Vis.7
2020 Instance-Aware Image Colorization
abstract
Image colorization is inherently an ill-posed problem with multi-modal uncertainty. Previous methods leverage the deep neural network to map input grayscale images to plausible color outputs directly. Although these learning-based methods have shown impressive performance, they usually fail on the input images that contain multiple objects. The leading cause is that existing models perform learning and colorization on the entire image. In the absence of a clear figure-ground separation, these models cannot effectively locate and learn meaningful object-level semantics. In this paper, we propose a method for achieving instance-aware colorization. Our network architecture leverages an off-the-shelf object detector to obtain cropped object images and uses an instance colorization network to extract object-level features. We use a similar network to extract the full-image features and apply a fusion module to full object-level and image-level features to predict the final colors. Both colorization networks and fusion modules are learned from a large-scale dataset. Experimental results show that our work outperforms existing methods on different quality metrics and achieves state-of-the-art performance on image colorization.
Jheng-Wei Su, Hung-Kuo Chu, Jia-Bin Huang 0001
CVPR2
2020 Image Vectorization With Real-Time Thin-Plate Spline
abstract
The vector graphics with gradient mesh can be attributed to their compactness and scalability; however, they tend to fall short when it comes to real-time editing due to a lack of real-time rasterization and an efficient editing tool for image details. In this paper, we encode global manipulation geometries and local image details within a hybrid vector structure, using parametric patches and detailed features for localized and parallelized thin-plate spline interpolation in order to achieve good compressibility, interactive expressibility, and editability. The proposed system then automatically extracts an optimal set of detailed color features while considering the compression ratio of the image as well as reconstruction error and its characteristics applicable to the preservation of structural and irregular saliency of the image. The proposed real-time vector representation makes it possible to construct an interactive editing system for detail-maintained image magnification and color editing as well as material replacement in cross mapping, without maintaining spatial and temporal consistency while editing in a raster space. Experiments demonstrate that our representation method is superior to several state-of-the-art methods and as good as JPEG, while providing real-time editability and preserving structural and irregular saliency information.
Kuo-Wei Chen, Ying-Sheng Luo, Yu-Chi Lai, Yan-Lin Chen, Chih-Yuan Yao, Hung-Kuo Chu, Tong-Yee Lee
IEEE Trans. Multim.6
2020 Vid2Curve: simultaneous camera motion estimation and thin structure reconstruction from an RGB video
abstract
Thin structures, such as wire-frame sculptures, fences, cables, power lines, and tree branches, are common in the real world. It is extremely challenging to acquire their 3D digital models using traditional image-based or depth-based reconstruction methods, because thin structures often lack distinct point features and have severe self-occlusion. We propose the first approach that simultaneously estimates camera motion and reconstructs the geometry of complex 3D thin structures in high quality from a color video captured by a handheld camera. Specifically, we present a new curve-based approach to estimate accurate camera poses by establishing correspondences between featureless thin objects in the foreground in consecutive video frames, without requiring visual texture in the background scene to lock on. Enabled by this effective curve-based camera pose estimation strategy, we develop an iterative optimization method with tailored measures on geometry, topology as well as self-occlusion handling for reconstructing 3D thin structures. Extensive validations on a variety of thin structures show that our method achieves accurate camera pose estimation and faithful reconstruction of 3D thin structures with complex shape and topology at a level that has not been attained by other existing reconstruction methods.
Peng Wang 0099, Lingjie Liu, Nenglun Chen, Hung-Kuo Chu, Christian Theobalt, Wenping Wang 0001
ACM Trans. Graph.4
2020 Micrography QR Codes
abstract
This paper presents a novel algorithm to generate micrography QR codes, a novel machine-readable graphic generated by embedding a QR code within a micrography image. The unique structure of micrography makes it incompatible with existing methods used to combine QR codes with natural or halftone images. We exploited the high-frequency nature of micrography in the design of a novel deformation model that enables the skillful warping of individual letters and adjustment of font weights to enable the embedding of a QR code within a micrography. The entire process is supervised by a set of visual quality metrics tailored specifically for micrography, in conjunction with a novel QR code quality measure aimed at striking a balance between visual fidelity and decoding robustness. The proposed QR code quality measure is based on probabilistic models learned from decoding experiments using popular decoders with synthetic QR codes to capture the various forms of distortion that result from image embedding. Experiment results demonstrate the efficacy of the proposed method in generating micrography QR codes of high quality from a wide variety of inputs. The ability to embed QR codes with multiple scales makes it possible to produce a wide range of diverse designs. Experiments and user studies were conducted to evaluate the proposed method from a qualitative as well as quantitative perspective.
Shih-Hsuan Hung, Chih-Yuan Yao, Yu-Jen Fang, Ping Tan 0002, Ruen-Rone Lee, Alla Sheffer, Hung-Kuo Chu
IEEE Trans. Vis. Comput. Graph.7
2019 DuLa-Net: A Dual-Projection Network for Estimating Room Layouts From a Single RGB Panorama
abstract
We present a deep learning framework, called DuLa-Net, to predict Manhattan-world 3D room layouts from a single RGB panorama. To achieve better prediction accuracy, our method leverages two projections of the panorama at once, namely the equirectangular panorama-view and the perspective ceiling-view, that each contains different clues about the room layouts. Our network architecture consists of two encoder-decoder branches for analyzing each of the two views. In addition, a novel feature fusion structure is proposed to connect the two branches, which are then jointly trained to predict the 2D floor plans and layout heights. To learn more complex room layouts, we introduce the Realtor360 dataset that contains panoramas of Manhattan-world room layouts with different numbers of corners. Experimental results show that our work outperforms recent state-of-the-art in prediction accuracy and performance, especially in the rooms with non-cuboid layouts.
Shang-Ta Yang, Fu-En Wang, Chihan Peng, Peter Wonka, Min Sun 0001, Hung-Kuo Chu
CVPR6
2019 Generating Color Scribble Images using Multi-layered Monochromatic Strokes Dithering
abstract
Abstract Color scribbling is a unique form of illustration where artists use compact, overlapping, and monochromatic scribbles at microscopic scale to create astonishing colorful images at macroscopic scale. The creation process is skill‐demanded and time‐consuming, which typically involves drawing monochromatic scribbles layer‐by‐layer to depict true‐color subjects using a limited color palette delicately. In this work, we present a novel computational framework for automatic generation of color scribble images from arbitrary raster images. The core contribution of our work lies in a novel color dithering model tailor‐made for synthesizing a smooth color appearance using multiple layers of overlapped monochromatic strokes. Specifically, our system reconstructs the appearance of the input image by (i) generating layers of monochromatic scribbles based on a limited color palette derived from input image, and (ii) optimizing the drawing sequence among layers to minimize the visual color dissimilarity between dithered image and original image as well as the color banding artifacts. We demonstrate the effectiveness and robustness of our algorithm with various convincing results synthesized from a variety of input images with different stroke patterns. The experimental study further shows that our approach faithfully captures the scribble style and the color presentation at respectively microscopic and macroscopic scales, which is otherwise difficult for state‐of‐the‐art methods.
Yi-Hsiang Lo, Ruen-Rone Lee, Hung-Kuo Chu
Comput. Graph. Forum3
2018 Self-supervised Learning of Depth and Camera Motion from 360 ^\circ Videos
Fu-En Wang, Hou-Ning Hu, Hsien-Tzu Cheng, Juan-Ting Lin, Shang-Ta Yang, Meng-Li Shih, Hung-Kuo Chu, Min Sun 0001
ACCV (5)7
2018 A lightweight and efficient system for tracking handheld objects in virtual reality
abstract
While the content of virtual reality (VR) has grown explosively in recent years, the advance of designing user-friendly control interfaces in VR still remains a slow pace. The most commonly used device, such as gamepad or controller, has fixed shape and weight, and thus can not provide realistic haptic feedback when interacting with virtual objects in VR. In this work, we present a novel and lightweight tracking system in the context of manipulating handheld objects in VR. Specifically, our system can effortlessly synchronize the 3D pose of arbitrary handheld objects between the real world and VR in realtime performance. The tracking algorithm is simple, which delicately leverages the power of Leap Motion and IMU sensor to respectively track object's location and orientation. We demonstrate the effectiveness of our system with three VR applications use pencil, ping-pong paddle, and smartphone as control interfaces to provide users more immersive VR experience.
Ya-Kuei Chang, Jui-Wei Huang, Chien-Hua Chen, Chien-Wen Chen, Jain-Wei Peng, Min-Chun Hu 0001, Chih-Yuan Yao, Hung-Kuo Chu
VRST8
2018 Dual-MR: interaction with mixed reality using smartphones
abstract
Mixed reality (MR) has changed the perspective we see and interact with our world. While the current-generation of MR head-mounted devices (HMDs) are capable of generating high quality visual contents, interation in most MR applications typically relies on in-air hand gestures, gaze, or voice. These interfaces although are intuitive to learn, may easily lead to inaccurate operations due to fatigue or constrained by the environment. In this work, we present Dual-MR, a novel MR interation system that i) synchronizes the MR viewpoints of HMD and handheld smartphone, and ii) enables precise, tactile, immersive and user-friendly object-level manipulations throught the multi-touch input of smartphone. In addition, Dual-MR allows multiple users to join the same MR coordinate system to facilite the collaborate in the same physical space, which further broadens its usability. A preliminary user study shows that our system easily overwhelms the conventional interface, which combines in-air hand gesture and gaze, in the completion time for a series of 3D object manipulation tasks in MR.
Chi-Jung Lee, Hung-Kuo Chu
VRST2
2018 EZ-Manipulator: Designing a mobile, fast, and ambiguity-free 3D manipulation interface using smartphones
abstract
Interacting with digital contents in 3D is an essential task in various applications such as modeling packages, gaming, virtual reality, etc. Traditional interfaces using keyboard and mouse or trackball usually require a non-trivial amount of working space as well as a learning process. We present the design of EZ-Manipulator, a new 3D manipulation interface using smartphones that supports mobile, fast, and ambiguity-free interaction with 3D objects. Our system leverages the built-in multi-touch input and gyroscope sensor of smartphones to achieve 9 degrees-of-freedom axis-constrained manipulation and free-form rotation. Using EZ-Manipulator to manipulate objects in 3D is easy. The user merely has to perform intuitive singleor two-finger gestures and rotate the hand-held device to perform manipulations at fine-grained and coarse levels respectively.We further investigate the ambiguity in manipulation introduced by indirect manipulations using a multi-touch interface, and propose a dynamic virtual camera adjustment to effectively resolve the ambiguity. A preliminary study shows that our system has significant lower task completion time compared to conventional use of a keyboard–mouse interface, and provides a positive user experience to both novices and experts.
Po-Huan Tseng, Shih-Hsuan Hung, Pei-Ying Chiang, Chih-Yuan Yao, Hung-Kuo Chu
Comput. Vis. Media5
2018 Multi-view wire art
abstract
Wire art is the creation of three-dimensional sculptural art using wire strands. As the 2D projection of a 3D wire sculpture forms line drawing patterns, it is possible to craft multi-view wire sculpture art --- a static sculpture with multiple (potentially very different) interpretations when perceived at different viewpoints. Artists can effectively leverage this characteristic and produce compelling artistic effects. However, the creation of such multi-view wire sculpture is extremely time-consuming even by highly skilled artists. In this paper, we present a computational framework for automatic creation of multi-view 3D wire sculpture. Our system takes two or three user-specified line drawings and the associated viewpoints as inputs. We start with producing a sparse set of voxels via greedy selection approach such that their projections on the virtual cameras cover all the contour pixels of the input line drawings. The sparse set of voxels, however, do not necessary form one single connected component. We introduce a constrained 3D pathfinding algorithm to link isolated groups of voxels into a connected component while maintaining the similarity between the projected voxels and the line drawings. Using the reconstructed visual hull, we extract a curve skeleton and produce a collection of smooth 3D curves by fitting cubic splines and optimizing the curve deformation to best approximate the provided line drawings. We demonstrate the effectiveness of our system for creating compelling multi-view wire sculptures in both simulation and 3D physical printouts.
Kai-Wen Hsiao, Jia-Bin Huang 0001, Hung-Kuo Chu
ACM Trans. Graph.3
2018 Scale-aware black-and-white abstraction of 3D shapes
abstract
Flat design is a modern style of graphics design that minimizes the number of design attributes required to convey 3D shapes. This approach suits design contexts requiring simplicity and efficiency, such as mobile computing devices. This `less-is-more' design inspiration has posed significant challenges in practice since it selects from a restricted range of design elements (e.g., color and resolution) to represent complex shapes. In this work, we investigate a means of computationally generating a specialized 2D flat representation - image formed by black-and-white patches - from 3D shapes. We present a novel framework that automatically abstracts 3D man-made shapes into 2D binary images at multiple scales. Based on a set of identified design principles related to the inference of geometry and structure, our framework jointly analyzes the input 3D shape and its counterpart 2D representation, followed by executing a carefully devised layout optimization algorithm. The robustness and effectiveness of our method are demonstrated by testing it on a wide variety of man-made shapes and comparing the results with baseline methods via a pilot user study. We further present two practical applications that are likely to benefit from our work.
You-En Lin, Yongliang Yang 0002, Hung-Kuo Chu
ACM Trans. Graph.3
2017 AlphaRead: Support Unambiguous Referencing in Remote Collaboration with Readable Object Annotation
abstract
As experts and expertise are increasingly distributed across distance, remote collaboration on physical tasks also becomes popular. Physical collaboration requires collaborators to produce and resolve references to physical objects unambiguously. We present a novel annotation system called AlphaRead that enables users to add and see readable annotations of physical objects, such as labels in letter, in a dynamic video-mediated workspace. Explicit support for object readability can help people coordinate language and vision for collaboration, and allow them to directly read out object labels as a way to make unambiguous references. Object readability as a resource of linguistic references can reduce the ambiguity and complexity associated with traditional methods of referential expressions such as deictic pronouns ("this" or "that") or descriptions of object attributes. In a video-mediated collaboration study, by making objects referable with readable labels, we improved the communication efficiency over the alternative options of using raw video or video with non-readable annotations to collaborate. We also identified patterns of language behaviors that people exhibited with readable labels and discussed the implications to the design of collaboration support tools.
Yuan-Chia Chang, Hao-Chuan Wang, Hung-Kuo Chu, Shung-Ying Lin, Shuo-Ping Wang
CSCW3
2017 User-guided line abstraction using coherence and structure analysis
abstract
Line drawing is a style of image abstraction where the perceptual content of the image is conveyed using distinct straight or curved lines. However, extracting semantically salient lines is not trivial and mastered only by skilled artists. While many parametric filters have successfully extracted accurate and coherent lines, their results are sensitive to parameter choice and easily lead to either an excessive or insufficient number of lines. In this work, we present an interactive system to generate concise line abstractions of arbitrary images via a few user specified strokes. Specifically, the user simply has to provide a few intuitive strokes on the input images, including tracing roughly along edges and scribbling on the region of interest, through a sketching interface. The system then automatically extracts lines that are long, coherent and share similar textural structures to form a corresponding highly detailed line drawing. We have tested our system with a wide variety of images. Our experimental results show that our system outperforms state-of-the-art techniques in terms of quality and efficiency.
Hui-Chi Tsai, Ya-Hsuan Lee, Ruen-Rone Lee, Hung-Kuo Chu
Comput. Vis. Media4
2017 Generating Ambiguous Figure-Ground Images
abstract
Ambiguous figure-ground images, mostly represented as binary images, are fascinating as they present viewers a visual phenomena of perceiving multiple interpretations from a single image. In one possible interpretation, the white region is seen as a foreground figure while the black region is treated as shapeless background. Such perception can reverse instantly at any moment. In this paper, we investigate the theory behind this ambiguous perception and present an automatic algorithm to generate such images. We model the problem as a binary image composition using two object contours and approach it through a three-stage pipeline. The algorithm first performs a partial shape matching to find a good partial contour matching between objects. This matching is based on a content-aware shape matching metric, which captures features of ambiguous figure-ground images. Then we combine matched contours into a compound contour using an adaptive contour deformation, followed by computing an optimal cropping window and image binarization for the compound contour that maximize the completeness of object contours in the final composition. We have tested our system using a wide range of input objects and generated a large number of convincing examples with or without user guidance. The efficiency of our system and quality of results are verified through an extensive experimental study.
Ying-Miao Kuo, Hung-Kuo Chu, Ming-Te Chi, Ruen-Rone Lee, Tong-Yee Lee
IEEE Trans. Vis. Comput. Graph.2
2016 Synthesizing Emerging Images from Photographs
abstract
Emergence is the visual phenomenon by which humans recognize the objects in a seemingly noisy image through aggregating information from meaningless pieces and perceiving a whole that is meaningful. Such an unique mental skill renders emergence an effective scheme to tell humans and machines apart. Images that are detectable by human but difficult for an automatic algorithm to recognize are also referred as emerging images. A recent state-of-the-art work proposes to synthesize images of 3D objects that are detectable by human but difficult for an automatic algorithm to recognize. Their results are further verified to be easy for humans to recognize while posing a hard time for automatic machines. However, using 3D objects as inputs prevents their system from being practical and scalable for generating an infinite number of high quality images. For instance, the image quality may degrade quickly as the viewing and lighting conditions changing in 3D domain, and the available resources of 3D models are usually limited. However, using 3D objects as inputs brings drawbacks. For instance, the quality of results is sensitive to the viewing and lighting conditions in the 3D domain. The available resources of 3D models are usually limited, and thus restricts the scalability. This paper presents a novel synthesis technique to automatically generate emerging images from regular photographs, which are commonly taken with decent setting and widely accessible online. We adapt the previous system to the 2D setting of input photographs and develop a set of image-based operations. Our algorithm is also designed to support the difficulty level control of resultant images through a limited set of parameters. We conducted several experiments to validate the efficacy and efficiency of our system.
Cheng-Han Yang, Ying-Miao Kuo, Hung-Kuo Chu
ACM Multimedia3
2016 A simulation on grass swaying with dynamic wind force
abstract
Grass, lawn, and meadow are common features of outdoor scenes. However, animating the grass motion under the influence of wind could be a daunting task especially if the subtle variations between grass blades are considered. Collectively, those variations of individual grass blades may produce interesting phenomena, such as the eye catching wave-like motion of a meadow. In this work, we develop a framework that simulates the grass dynamics under a plausible wind field. We first test it with a small bunch of grass blades (Figure 1) and validate it with the motion of real-world grass. By utilizing the tessellation shaders and geometry shaders of modern GPUs, we make our grass model as efficient as possible so that it scales well to a large meadow, which is also exposed to a plausible wind field simulation. As a result, our grass simulation runs in real time for meadow scenes consisting of hundreds of thousands of grass blades, and produces convincing wave-like motions when the wind blows over the meadow.
Yi-Hsiang Lo, Hung-Kuo Chu, Ruen-Rone Lee, Chun-Fa Chang
I3D2
2016 Interactive Videos: Plausible Video Editing using Sparse Structure Points
abstract
Abstract Video remains the method of choice for capturing temporal events. However, without access to the underlying 3D scene models, it remains difficult to make object level edits in a single video or across multiple videos. While it may be possible to explicitly reconstruct the 3D geometries to facilitate these edits, such a workflow is cumbersome, expensive, and tedious. In this work, we present a much simpler workflow to create plausible editing and mixing of raw video footage using only sparse structure points (SSP) directly recovered from the raw sequences. First, we utilize user‐scribbles to structure the point representations obtained using structure‐from‐motion on the input videos. The resultant structure points, even when noisy and sparse, are then used to enable various video edits in 3D, including view perturbation, keyframe animation, object duplication and transfer across videos, etc. Specifically, we describe how to synthesize object images from new views adopting a novel image‐based rendering technique using the SSPs as proxy for the missing 3D scene information. We propose a structure‐preserving image warping on multiple input frames adaptively selected from object video, followed by a spatio‐temporally coherent image stitching to compose the final object image. Simple planar shadows and depth maps are synthesized for objects to generate plausible video sequence mimicking real‐world interactions. We demonstrate our system on a variety of input videos to produce complex edits, which are otherwise difficult to achieve.
Chia-Sheng Chang, Hung-Kuo Chu, Niloy J. Mitra
Comput. Graph. Forum2
2016 Feature-Aware Pixel Art Animation
abstract
Abstract Pixel art is a modern digital art in which high resolution images are abstracted into low resolution pixelated outputs using concise outlines and reduced color palettes. Creating pixel art is a labor intensive and skill‐demanding process due to the challenge of using limited pixels to represent complicated shapes. Not surprisingly, generating pixel art animation is even harder given the additional constraints imposed in the temporal domain. Although many powerful editors have been Designed to facilitate the creation of still pixel art images, the extension to pixel art animation remains an unexplored direction. Existing systems typically request users to craft individual pixels frame by frame, which is a tedious and error‐prone process. In this work, we present a novel animation framework tailored to pixel art images. Our system bases on conventional key‐frame animation framework and state‐of‐the‐art image warping techniques to generate an initial animation sequence. The system then jointly optimizes the prominent feature lines of individual frames respecting three metrics that capture the quality of the animation sequence in both spatial and temporal domains. We demonstrate our system by generating visually pleasing animations on a variety of pixel art images, which would otherwise be difficult by applying state‐of‐the‐art techniques due to severe artifacts.
Ming-Hsun Kuo, Yongliang Yang 0002, Hung-Kuo Chu
Comput. Graph. Forum3
2016 Court Reconstruction for Camera Calibration in Broadcast Basketball Videos
abstract
We introduce a technique of calibrating camera motions in basketball videos. Our method particularly transforms player positions to standard basketball court coordinates and enables applications such as tactical analysis and semantic basketball video retrieval. To achieve a robust calibration, we reconstruct the panoramic basketball court from a video, followed by warping the panoramic court to a standard one. As opposed to previous approaches, which individually detect the court lines and corners of each video frame, our technique considers all video frames simultaneously to achieve calibration; hence, it is robust to illumination changes and player occlusions. To demonstrate the feasibility of our technique, we present a stroke-based system that allows users to retrieve basketball videos. Our system tracks player trajectories from broadcast basketball videos. It then rectifies the trajectories to a standard basketball court by using our camera calibration method. Consequently, users can apply stroke queries to indicate how the players move in gameplay during retrieval. The main advantage of this interface is an explicit query of basketball videos so that unwanted outcomes can be prevented. We show the results in Figs. 1, 7, 9, 10 and our accompanying video to exhibit the feasibility of our technique.
Pei-Chih Wen, Wei-Chih Cheng, Yu-Shuen Wang, Hung-Kuo Chu, Nick C. Tang, Hong-Yuan Mark Liao
IEEE Trans. Vis. Comput. Graph.4
2016 A simulation on grass swaying with dynamic wind force
Ruen-Rone Lee, Yi Lo, Hung-Kuo Chu, Chun-Fa Chang
Vis. Comput.3
2015 Spatio-Temporal Learning of Basketball Offensive Strategies
abstract
Video-based group behavior analysis is drawing attention to its rich applications in sports, military, surveillance and biological observations. The recent advances in tracking techniques, based on either computer vision methodology or hardware sensors, further provide the opportunity of better solving this challenging task. Focusing specifically on the analysis of basketball offensive strategies, we introduce a systematic approach to establishing unsupervised modeling of group behaviors. In view that a possible group behavior (offensive strategy) could be of different duration and represented by dynamic player trajectories, the crux of our method is to automatically divide training data into meaningful clusters and learn their respective spatio-temporal model, which is established upon Gaussian mixture regression to account for intra-class spatio-temporal variations. The resulting strategy representation turns out to be flexible that can be used to not only establish the discriminant functions but also improve learning the models. We demonstrate the usefulness of our approach by exploring its effectiveness in analyzing a set of given basketball video clips.
Ching-Hang Chen, Tyng-Luh Liu, Yu-Shuen Wang, Hung-Kuo Chu, Nick C. Tang, Hong-Yuan Mark Liao
ACM Multimedia4
2015 Tone- and Feature-Aware Circular Scribble Art
abstract
Circular scribble art is a kind of line drawing where the seemingly random, noisy and shapeless circular scribbles at microscopic scale constitute astonishing grayscale images at macroscopic scale. Such a delicate skill has rendered the creation of circular scribble art a tedious and time-consuming task even for gifted artists. In this work, we present a novel method for automatic synthesis of circular scribble art. The synthesis problem is modeled as tracing along a virtual path using a parametric circular curve. To reproduce the tone and important edge structure of input grayscale images, the system adaptively adjusts the density and structure of virtual path, and dynamically controls the size, drawing speed and orientation of parametric circular curve during the synthesis. We demonstrate the potential of our system using several circular scribble images synthesized from a wide variety of grayscale images. A preliminary experimental studying is conducted to qualitatively and quantitatively evaluate our method. Results report that our method is efficient and generates convincing results comparable to artistic artworks.
Chun-Chia Chiu, Yi-Hsiang Lo, Ruen-Rone Lee, Hung-Kuo Chu
Comput. Graph. Forum4
2015 Pixel2Brick: Constructing Brick Sculptures from Pixel Art
abstract
LEGO®, a popular brick-based toy construction system, provides an affordable and convenient way of fabricating geometric shapes. However, building arbitrary shapes using LEGO bricks with restrictive colors and sizes is not trivial. It requires careful design process to produce appealing, stable and constructable brick sculptures. In this work, we investigate the novel problem of constructing brick sculptures from pixel art images. In contrast to previous efforts that focus on 3D models, pixel art contains rich visual contents for generating engaging LEGO designs. On the other hand, the characteristics of pixel art and corresponding brick sculpture pose new challenges to the design process. We present Pixel2Brick, a novel computational framework to automatically construct brick sculptures from pixel art. This is based on implementing a set of design guidelines concerning the visual quality as well as the structural stability of built sculptures. We demonstrate the effectiveness of our framework with various brick sculptures (both real and virtual) generated from a variety of pixel art images. Experimental results show that our framework is efficient and gains significant improvements over state-of-the-arts.
Ming-Hsun Kuo, You-En Lin, Hung-Kuo Chu, Ruen-Rone Lee, Yongliang Yang 0002
Comput. Graph. Forum3
2015 SmartAnnotator An Interactive Tool for Annotating Indoor RGBD Images
abstract
Abstract RGBD images with high quality annotations, both in the form of geometric (i.e., segmentation) and structural (i.e., how do the segments mutually relate in 3D) information, provide valuable priors for a diverse range of applications in scene understanding and image manipulation. While it is now simple to acquire RGBD images, annotating them, automatically or manually, remains challenging. We present SmartAnnotator, an interactive system to facilitate annotating raw RGBD images. The system performs the tedious tasks of grouping pixels, creating potential abstracted cuboids, inferring object interactions in 3D, and generates an ordered list of hypotheses. The user simply has to flip through the suggestions for segment labels, finalize a selection, and the system updates the remaining hypotheses. As annotations are finalized, the process becomes simpler with fewer ambiguities to resolve. Moreover, as more scenes are annotated, the system makes better suggestions based on the structural and geometric priors learned from previous annotation sessions. We test the system on a large number of indoor scenes across different users and experimental settings, validate the results on existing benchmark datasets, and report significant improvements over low‐level annotation alternatives. (Code and benchmark datasets are publicly available on the project page.)
Yu-Shiang Wong, Hung-Kuo Chu, Niloy J. Mitra
Comput. Graph. Forum2
2013 Halftone QR codes
abstract
QR code is a popular form of barcode pattern that is ubiquitously used to tag information to products or for linking advertisements. While, on one hand, it is essential to keep the patterns machine-readable; on the other hand, even small changes to the patterns can easily render them unreadable. Hence, in absence of any computational support, such QR codes appear as random collections of black/white modules, and are often visually unpleasant. We propose an approach to produce high quality visual QR codes, which we call halftone QR codes , that are still machine-readable. First, we build a pattern readability function wherein we learn a probability distribution of what modules can be replaced by which other modules. Then, given a text tag, we express the input image in terms of the learned dictionary to encode the source text. We demonstrate that our approach produces high quality results on a range of inputs and under different distortion effects.
Hung-Kuo Chu, Chia-Sheng Chang, Ruen-Rone Lee, Niloy J. Mitra
ACM Trans. Graph.1
2010 Camouflage images
abstract
Camouflage images contain one or more hidden figures that remain imperceptible or unnoticed for a while. In one possible explanation, the ability to delay the perception of the hidden figures is attributed to the theory that human perception works in two main phases: feature search and conjunction search. Effective camouflage images make feature based recognition difficult, and thus force the recognition process to employ conjunction search, which takes considerable effort and time. In this paper, we present a technique for creating camouflage images. To foil the feature search, we remove the original subtle texture details of the hidden figures and replace them by that of the surrounding apparent image. To leave an appropriate degree of clues for the conjunction search, we compute and assign new tones to regions in the embedded figures by performing an optimization between two conflicting terms, which we call immersion and standout , corresponding to hiding and leaving clues, respectively. We show a large number of camouflage images generated by our technique, with or without user guidance. We have tested the quality of the images in an extensive user study, showing a good control of the difficulty levels.
Hung-Kuo Chu, Wei-Hsin Hsu, Niloy J. Mitra, Daniel Cohen-Or, Tien-Tsin Wong, Tong-Yee Lee
ACM Trans. Graph.1
2009 Compatible quadrangulation by sketching
abstract
Abstract Mesh quadrangulation has received increasing attention in the past decade. While previous works have mostly focused on producing a high quality quad mesh of a single model, the connectivity of the quadrangulation is typically difficult to control and varies among models even with similar shapes. In this paper, we propose a novel interactive framework for quadrangulating a set of models collectively with compatible connectivity. Furthermore, we demonstrate its application to 3D mesh morphing. In our approach, the user interactively sketches a skeleton within each model, and our method automatically computes compatible base domains for all models from these skeletons, on which the models are parameterized. With this novel parameterization, it is very easy to generate a pleasing and smooth 3D morphing sequence among these compatible models. The method yields quadrangulation with comparable quality to existing approaches, but greatly simplifies compatible re‐meshing among a group of topologically equivalent models, in particular characters and animals models, with direct applications in shape blending and morphing. Copyright © 2009 John Wiley & Sons, Ltd.
Chih-Yuan Yao, Hung-Kuo Chu, Tong-Yee Lee
Comput. Animat. Virtual Worlds2
2009 Emerging images
abstract
Emergence refers to the unique human ability to aggregate information from seemingly meaningless pieces, and to perceive a whole that is meaningful. This special skill of humans can constitute an effective scheme to tell humans and machines apart. This paper presents a synthesis technique to generate images of 3D objects that are detectable by humans, but difficult for an automatic algorithm to recognize. The technique allows generating an infinite number of images with emerging figures. Our algorithm is designed so that locally the synthesized images divulge little useful information or cues to assist any segmentation or recognition procedure. Therefore, as we demonstrate, computer vision algorithms are incapable of effectively processing such images. However, when a human observer is presented with an emergence image, synthesized using an object she is familiar with, the figure emerges when observed as a whole. We can control the difficulty level of perceiving the emergence effect through a limited set of parameters. A procedure that synthesizes emergence images can be an effective tool for exploring and understanding the factors affecting computer vision techniques.
Niloy J. Mitra, Hung-Kuo Chu, Tong-Yee Lee, Lior Wolf, Yehezkel Yeshurun, Daniel Cohen-Or
ACM Trans. Graph.2
2009 Multiresolution Mean Shift Clustering Algorithm for Shape Interpolation
abstract
In this paper, we solve the problem of 3D shape interpolation with significant pose variation. For an ideal 3D shape interpolation, especially the articulated model, the shape should follow the movement of the underlying articulated structure and be transformed in a way that is as rigid as possible. Given input shapes with compatible connectivity, we propose a novel multiresolution mean shift (MMS) clustering algorithm to automatically extract their near-rigid components. Then, by building the hierarchical relationship among extracted components, we compute a common articulated structure for these input shapes. With the aid of this articulated structure, we solve the shape interpolation by combining 1) a global pose interpolation of near-rigid components from the source shape to the target shape with 2) a local gradient field interpolation for each pair of components, followed by solving a Poisson equation in order to reconstruct an interpolated shape. As a result, an aesthetically pleasing shape interpolation can be generated, with even the poses of shapes varying significantly. In contrast to a recent state-of-the-art work, the proposed approach can achieve comparable or even better results and have better computational efficiency as well.
Hung-Kuo Chu, Tong-Yee Lee
IEEE Trans. Vis. Comput. Graph.1
2008 Skeleton extraction by mesh contraction
abstract
Extraction of curve-skeletons is a fundamental problem with many applications in computer graphics and visualization. In this paper, we present a simple and robust skeleton extraction method based on mesh contraction. The method works directly on the mesh domain, without pre-sampling the mesh model into a volumetric representation. The method first contracts the mesh geometry into zero-volume skeletal shape by applying implicit Laplacian smoothing with global positional constraints. The contraction does not alter the mesh connectivity and retains the key features of the original mesh. The contracted mesh is then converted into a 1D curve-skeleton through a connectivity surgery process to remove all the collapsed faces while preserving the shape of the contracted mesh and the original topology. The centeredness of the skeleton is refined by exploiting the induced skeleton-mesh mapping. In addition to producing a curve skeleton, the method generates other valuable information about the object's geometry, in particular, the skeleton-vertex correspondence and the local thickness, which are useful for various applications. We demonstrate its effectiveness in mesh segmentation and skinning animation.
Oscar Kin-Chung Au, Chiew-Lan Tai, Hung-Kuo Chu, Daniel Cohen-Or, Tong-Yee Lee
ACM Trans. Graph.3
2007 Mesh pose-editing using examples
abstract
Abstract An easy‐to‐use mesh pose‐editing system is presented. We take advantage of both skeleton‐based and example‐based approaches in order to provide an intuitive way for artists to edit mesh poses. Our system automatically extracts the skeletons of the remaining example models once the skeleton of a reference mesh is constructed. In our editing system the desired skeleton can be easily and naturally posed using an inverse kinematics (IK) algorithm incorporated with searching the optimal weights in the defined skeleton space of examples meshes. Eventually, the desired shape with detailed deformation can be constructed by blending the example meshes. Experimental results show that the proposed system provides an easy and intuitive control on mesh pose‐editing. Copyright © 2007 John Wiley & Sons, Ltd.
Tong-Yee Lee, Chao-Hung Lin, Hung-Kuo Chu, Yu-Shuen Wang, Shao-Wei Yen, Chang-Rung Tsai
Comput. Animat. Virtual Worlds3
2006 Generating genus-n-to-m mesh morphing using spherical parameterization
abstract
Abstract Surface parameterization is a fundamental tool in computer graphics and benefits many applications such as texture mapping, morphing, and re‐meshing. Many spherical parameterization schemes with very nice properties have been proposed and widely used in the past. However, it is well known that the spherical parameterization is limited to genus‐0models. In this paper, we first propose a novel framework to extend spherical parameterization for handling a genus‐n surface. In this framework, we represent a surface S of arbitrary genus by a positive mesh O and several negative meshes Ni. Each negative surface is used to represent a hole. A positive surface O is obtained by removing all holes in the original surface S. Then, both positive and negative meshes are genus‐0 and can be spherically parameterized, respectively. To compute S, we can use a Boolean difference operation to subtract negative Nifrom a positive O. Next, we apply this novel framework to generate genus‐n‐to‐m mesh morphing application without restriction of n = m. Finally, there are many interesting non‐genus‐0 mesh morphing sequences generated. Copyright © 2006 John Wiley & Sons, Ltd.
Tong-Yee Lee, Chih-Yuan Yao, Hung-Kuo Chu, Ming-Jen Tai, Cheng-Chieh Chen
Comput. Animat. Virtual Worlds3
2005 Progressive mesh metamorphosis
abstract
Abstract This paper describes a new integrated scheme for metamorphosis between two closed manifold genus‐0 polyhedral models. Spherical parameterizations of the source and target models are created first. To control the morphing, any number of feature vertex pairs is specified and a fold‐over free warping method is used to align two spherical embeddings. Our method does not create a merged meta‐mesh or execute re‐meshing to construct a common connectivity for morphs. Alternatively, a scheme for the progressive connectivity transformation of two spherical parameterizations is employed to generate the intermediate meshes. A novel semi‐overlay with a geomorph scheme is proposed to reduce the popping effects caused by the connectivity transformation. We demonstrate several examples of aesthetically pleasing morphing sequences using the proposed scheme. Copyright © 2005 John Wiley & Sons, Ltd.
Chao-Hung Lin, Tong-Yee Lee, Hung-Kuo Chu, Chih-Yuan Yao
Comput. Animat. Virtual Worlds3