Tong-Yee Lee

dblp:18/4993 · DBLP profile ↗
← Back
142ranked-venue papers
26as first author
50since 2021 · last 2026
0000-0001-6699-2944ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 119 · 9 first-author · 48 since 2021Applied, interdisciplinary, general and emerging computing · 12 · 10 first-authorArtificial intelligence and machine learning · 7 · 2 first-author · 5 since 2021Systems, architecture and hardware · 6 · 5 first-authorComputer networks · 6 · 5 since 2021Human-computer interaction and ubiquitous computing · 3 · 3 since 2021
YearPublicationVenuePosition
2026 Bridging Cognitive Gap: Hierarchical Description Learning for Artistic Image Aesthetics Assessment
abstract
The aesthetic quality assessment task is crucial for developing a human-aligned quantitative evaluation system for AIGC. However, its inherently complex nature—spanning visual perception, cognition, and emotion—poses fundamental challenges. Although aesthetic descriptions offer a viable representation of this complexity, two critical challenges persist: (1) data scarcity and imbalance: existing dataset overly focuses on visual perception and neglects deeper dimensions due to the expensive manual annotation; and (2) model fragmentation: current visual networks isolate aesthetic attributes with multi-branch encoder, while multimodal methods represented by contrastive learning struggle to effectively process long-form textual descriptions. To resolve challenge (1), we first present the Refined Aesthetic Description (RAD) dataset, a large-scale (70k), multi-dimensional structured dataset, generated via an iterative pipeline without heavy annotation costs and easy to scale. To address challenge (2), we propose ArtQuant, an aesthetics assessment framework for artistic image which not only couple isolated aesthetic dimensions through joint description generation, but also better model long-text semantics with the help of LLM decoders. Besides, theoretical analysis confirms this symbiosis: RAD's semantic adequacy (data) and generation paradigm (model) collectively minimize prediction entropy, providing mathematical grounding for the framework. Our approach achieves state-of-the-art performance on several datasets while requiring only 33% of conventional training epochs, narrowing the cognitive gap between artistic image and aesthetic judgment. We will release both code and dataset to support future research.
Henglin Liu, Nisha Huang, Chang Liu 0071, Jiangpeng Yan, Huijuan Huang 0001, Jixuan Ying, Tong-Yee Lee, Pengfei Wan 0001, Xiangyang Ji
AAAI7
2026 HeadRouter: A Training-free Image Editing Framework for MM-DiTs by Adaptively Routing Attention Heads
abstract
Diffusion Transformers (DiTs) have exhibited robust capabilities in image generation tasks. However, accurate text-guided image editing for multimodal DiTs (MM-DiTs) still poses a significant challenge. Unlike UNet-based structures that could utilize self/cross-attention maps for semantic editing, MM-DiTs inherently lack support for explicit and consistent incorporated text guidance, resulting in semantic misalignment between the edited results and texts. In this study, we disclose the sensitivity of different attention heads to different image semantics within MM-DiTs and introduce HeadRouter , a training-free image editing framework that edits the source image by adaptively routing the text guidance to different attention heads in MM-DiTs. Furthermore, we propose a dual-token refinement module to refine text/image token representations for precise semantic guidance and accurate region expression. Experiments on multiple benchmarks demonstrate HeadRouter’s performance in terms of editing fidelity and image quality. The code is available at https://github.com/ICTMCG/HeadRouter .
Fan Tang, Juan Cao 0001, Xiaoyu Kong, Yuxin Zhang 0006, Jintao Li 0001, Oliver Deussen, Tong-Yee Lee
ACM Trans. Graph.8
2026 ArtCrafter: Text-Image Aligning Artistic Attribute Transfer via Embedding Reframing
abstract
Recent years have witnessed significant advancements in text-guided style transfer, primarily attributed to innovations in diffusion models. These models excel in conditional guidance, utilizing text or images to direct the sampling process. Traditional style transfer focuses on low-level visual features, such as brushstroke textures and color distributions, and appears more like applying an artistic filter to an image. Artistic attribute transfer, however, transcends the limitations of traditional style transfer by achieving the transfer of visual concepts from color and brushstrokes to high level aesthetic attributes such as composition, pose, and key semantic elements, resulting in more natural outcomes. Therefore, we propose an innovative text-to-image artistic attribute transfer framework named ArtCrafter. Specifically, we introduce an attention-based style extraction module, meticulously engineered to capture the subtle artistic attribute elements within an image. This module features a multi-layer architecture that leverages the capabilities of perceiver attention mechanisms to integrate fine-grained information. Additionally, we present a novel text-image aligning augmentation component that adeptly balances control over both modalities, enabling the model to efficiently map image and text embeddings into a shared feature space. We achieve this through attention operations that enable smooth information flow between modalities. Lastly, we incorporate an explicit modulation that seamlessly blends multimodal enhanced embeddings with original embeddings through an embedding reframing design, empowering the model to generate diverse outputs. Extensive experiments demonstrate that ArtCrafter yields impressive results in visual stylization, exhibiting exceptional levels of artistic attribute intensity, controllability, and diversity.
Nisha Huang, Kaer Huang, Yifan Pu, Jiangshan Wang, Yiqiang Yan, Xiu Li 0001, Tong-Yee Lee
IEEE Trans. Vis. Comput. Graph.8
2026 Make-Your-Anchor+: Temporal Consistent 2D Avatar Generation via Video Diffusion Prior
abstract
Despite the remarkable process of talking-head-based avatar-creating solutions, directly generating anchor-style videos with full-body motions remains challenging. In this study, we propose Make-Your-Anchor+, a novel system necessitating only a one-minute video clip of an individual for training, subsequently enabling the automatic generation of anchor-style videos with precise torso and hand movements. Specifically, we finetune a proposed structure-guided diffusion model on input video to render 3D mesh conditions into human appearances. We adopt a two-stage training strategy for the diffusion model, effectively mapping movements with specific appearances to create digital avatars for online streamers, live shopping hosts, and other applications. To produce arbitrary long temporal video, we extract human motion information from video diffusion prior by adapting the frame-wise diffusion model to pretrained video diffusion weights with lower cost, and a simple yet effective batch-overlapped temporal denoising module is proposed to bypass the constraints on video length during inference. Finally, a novel identity-specific face enhancement module is introduced to improve the visual quality of facial regions in the output videos. Comparative experiments demonstrate the system's effectiveness and superiority in visual quality, temporal coherence, and identity preservation, outperforming SOTA diffusion/non-diffusion methods.
Ziyao Huang 0002, Fan Tang, Juan Cao 0001, Yong Zhang 0034, Xiaodong Cun, Yihang Bo, Jintao Li 0001, Tong-Yee Lee
IEEE Trans. Vis. Comput. Graph.8
2026 GSDeformer: Direct, Real-Time and Extensible Cage-Based Deformation for 3D Gaussian Splatting
abstract
We present GSDeformer, a method that enables cage-based deformation on 3D Gaussian Splatting (3DGS). Our approach bridges cage-based deformation and 3DGS by using a proxy point-cloud representation. This point cloud is generated from 3D Gaussians, and deformations applied to the point cloud are translated into transformations on the 3D Gaussians. To handle potential bending caused by deformation, we incorporate a splitting process to approximate it. Our method does not modify or extend the core architecture of 3D Gaussian Splatting, making it compatible with any trained vanilla 3DGS or its variants. Additionally, we automate cage construction for 3DGS and its variants using a render-and-reconstruct approach. Experiments demonstrate that GSDeformer delivers superior deformation results compared to existing methods, is robust under extreme deformations, requires no retraining for editing, runs in real-time, and can be extended to other 3DGS variants.
Shuolin Xu, Hongchuan Yu, Tong-Yee Lee
IEEE Trans. Vis. Comput. Graph.4
2026 Interactive Visual Assessment for Text-to-Image Generation Models
abstract
Visual generation models have achieved remarkable progress in computer graphics applications but still face significant challenges in real-world deployment. Current assessment approaches for visual generation tasks typically follow an isolated three-phase framework: test input collection, model output generation, and user assessment. These fashions suffer from fixed coverage, evolving difficulty, and data leakage risks, limiting their effectiveness in comprehensively evaluating increasingly complex generation models. To address these limitations, we propose DyEval, an LLM-powered dynamic interactive visual assessment framework that facilitates collaborative evaluation between humans and generative models for text-to-image systems. DyEval features an intuitive visual interface that enables users to interactively explore and analyze model behaviors, while adaptively generating hierarchical, fine-grained, and diverse textual inputs to continuously probe the capability boundaries of the models based on their feedback. Additionally, to provide interpretable analysis for users to further improve tested models, we develop a contextual reflection module that mines failure triggers of test inputs and reflects model potential failure patterns, supporting in-depth analysis using the logical reasoning ability of LLM. Qualitative and quantitative experiments demonstrate that DyEval can effectively help users identify max up to 2.56 timesmore generation failures than conventional methods, and uncover complex and rare failure patterns, such as issues with pronoun generation and specific cultural context generation. Our framework provides valuable insights for improving generative models and has broad implications for advancing the reliability and capabilities of visual generation systems across various domains.
Xiaoyue Mi, Fan Tang, Juan Cao 0001, Qiang Sheng 0001, Ziyao Huang 0002, Peng Li 0030, Yang Liu 0005, Tong-Yee Lee
IEEE Trans. Vis. Comput. Graph.8
2026 Hiding in Plain Sight: Camouflaging Real-World Objects
abstract
Camouflage, a survival strategy perfected by nature, enables organisms to evade detection by blending seamlessly into their surroundings across multiple viewpoints. Replicating this remarkable ability in artificial settings remains a formidable challenge. While recent computational methods have advanced 2D camouflage, extending concealment into 3D is far more difficult: a single texture must reconcile drastically varying backgrounds, making effective camouflage across all views highly challenging. Prior approaches attempt to address this by designing multi-view textures, but their appearance representations are too simplistic to cope with strongly conflicting backgrounds and fail to account for physical light transport, resulting in breakdowns under realistic illumination. We introduce a new paradigm that formulates 3D camouflage as a multiview inverse rendering problem. Instead of treating concealment as texture synthesis, we directly optimize appearance within a physically based rendering framework, explicitly modeling reflections, shadows, refractions, and other complex light interactions. Our approach expands the solution space by combining BSDF-based material representation with anamorphic surface displacement, enabling camouflage that adapts naturally to both viewpoint and illumination changes. Through tailored loss functions and optimization strategies, our method produces results that are not only perceptually convincing but also physically consistent, marking a significant step toward practical, real-world 3D camouflage. Experiments demonstrate that our physically grounded formulation achieves robust concealment across diverse viewpoints and lighting scenarios, substantially outperforming prior image-based methods.
Dong-Yi Wu, Tong-Yee Lee
IEEE Trans. Vis. Comput. Graph.2
2025 Attend to Not Attended: Structure-then-Detail Token Merging for Post-training DiT Acceleration
abstract
Diffusion transformers have shown exceptional performance in visual generation but incur high computational costs. Token reduction techniques that compress models by sharing the denoising process among similar tokens have been introduced. However, existing approaches neglect the denoising priors of the diffusion models, leading to suboptimal acceleration and diminished image quality. This study proposes a novel concept: attend to prune feature redundancies in areas not attended by the diffusion process. We analyze the location and degree of feature redundancies based on the structure-then-detail denoising priors. Subsequently, we introduce SDTM, a structure-then-detail token merging approach that dynamically compresses feature redundancies. Specifically, we design dynamic visual token merging, compression ratio adjusting, and prompt reweighting for different stages. Served in a post-training way, the proposed method can be integrated seamlessly into any DiT architecture. Extensive experiments across various backbones, schedulers, and datasets showcase the superiority of our method, for example, it achieves 1.55× acceleration with negligible impact on image quality. Project page: https://github.com/ICTMCG/SDTM.
Haipeng Fang, Sheng Tang, Juan Cao 0001, Enshuo Zhang, Fan Tang, Tong-Yee Lee
CVPR6
2025 MaTe: Images are All You Need for Material Transfer via Diffusion Transformer
Nisha Huang, Henglin Liu, Yizhou Lin, Kaer Huang, Chubin Chen, Tong-Yee Lee, Xiu Li 0001
ICCV7
2025 ICE: Intercede Concept Erasure in Text-to-Image Diffusion Models
abstract
The success of diffusion models in text-to-image (T2I) generation has made it urgent to remove unwanted concepts, such as copyrighted, offensive, and unsafe ones, from pre-trained models in an accurate, timely, and cost-effective manner. However, limited by the inherent optimization perspective, existing methods have two major problems. Firstly, they overlook maintaining the global visual style during the erasure process, leading to significant style shifts. Secondly, excessive concept erasure causes relevant content to disappear or generates substitutes unrelated to the original object's attributes. Compared to other methods, our proposed ICE has unique advantages, as it can generate diverse visual features and achieve a balance between concept erasure and maintaining the semantic content of the target object. This is mainly achieved through our well-designed non-erasable features protector (NEFP) and augmented invariant constraints (AIC). Specifically, we enhance the protection of feature information by embedding an augmented orthogonal anchor concept matrix. Meanwhile, under controlled constraints, we introduce invariants into the embedding space to retain key semantics. This work specifically emphasizes the importance of focusing on feature expression and semantic protection in the concept erasure task for fully unleashing the performance of T2I models.
Yizhou Lin, Nisha Huang, Kaer Huang, Henglin Liu, Yiqiang Yan, Tong-Yee Lee, Xiu Li 0001
ACM Multimedia7
2025 In-Context Brush: Zero-shot Customized Subject Insertion with Context-Aware Latent Space Manipulation
abstract
Recent advances in diffusion models have enhanced multimodal-guided visual generation, enabling customized subject insertion that seamlessly “brushes” user-specified objects into a given image guided by textual prompts. However, existing methods often struggle to insert customized subjects with high fidelity and align results with the user’s intent through textual prompts. In this work, we propose In-Context Brush, a zero-shot framework for customized subject insertion by reformulating the task within the paradigm of in-context learning. Without loss of generality, we formulate the object image and the textual prompts as cross-modal demonstrations, and the target image with the masked region as the query. The goal is to inpaint the target image with the subject aligning textual prompts without model tuning. Building upon a pretrained MMDiT-based inpainting network, we perform test-time enhancement via dual-level latent space manipulation: intra-head latent feature shifting within each attention head that dynamically shifts attention outputs to reflect the desired subject semantics and inter-head attention reweighting across different heads that amplifies prompt controllability through differential attention prioritization. Extensive experiments and applications demonstrate that our approach achieves superior identity preservation, text alignment, and image quality compared to existing state-of-the-art methods, without requiring dedicated training or additional data collection. Project page: https://yuci-gpt.github.io/In-Context-Brush/.
Fan Tang, Lin Gao 0004, Oliver Deussen, Hongbin Yan, Jintao Li 0001, Juan Cao 0001, Tong-Yee Lee
SIGGRAPH Asia9
2025 View-Independent Wire Art Modeling via Manifold Fitting
abstract
Abstract This paper presents a novel fully automated method for generating view‐independent abstract wire art from 3D models. The main challenge in creating line art is to strike a balance among abstraction, structural clarity, 3D perception, and consistent aesthetics from different viewpoints. Many existing approaches have been proposed, including extracting wire art from mesh, reconstructing it from pictures, etc. But they all suffer from the fact that the wires are usually very unorganized and cumbersome and usually can only guarantee the observation effect of specific viewpoints. To overcome these problems, we propose a paradigm shift: instead of predicting the line segments directly, we consider the generation of wire art as an optimization‐driven manifold‐fitting problem. Thus we can abstract/generalize the 3D model while retaining the key properties necessary for appealing line art, including structural topology and connectivity, and maintain the three‐dimensionality of the line art with a multi‐perspective view. Experimental results show that our view‐independent method outperforms previous methods in terms of line simplicity, shape fidelity, and visual consistency.
Huiguang Huang, Dong-Yi Wu, Yu Cao 0019, Tong-Yee Lee
Comput. Graph. Forum5
2025 3DCMM: 3D Comprehensive Morphable Models With UV-UNet for Accurate Head Creation
abstract
In recent studies of 3D shape modelling and reconstruction, the focus has primarily been on the 3D face region. However, accurately creating the entire 3D head opens up a wide range of applications, including headwear design, cranial diagnosis, and avatar design. Therefore, we present our newly developed method of constructing 3D comprehensive morphable models (3DCMM) specifically tailored for human heads, along with a novel 3DCMM-based stepwise pipeline for creating accurate full 3D heads. Within our 3DCMM framework, we constructed a powerful 3D morphable face model with UV-UNet to generate the 3D face and predict the 3D scalp, resulting in a complete representation of the head. Additionally, our 3DCMM-based self-learning approach incorporates novel facial boundary-aware and structure-aware losses for highly accurate overall reconstructions of the entire facial region. Experimental evaluations demonstrate that our 3DCMM exhibits superior face representation power and achieves higher head prediction accuracy than existing models. Consequently, our 3DCMM-based 3D head creation method from a single image demonstrates outstanding performance capability on both face and head benchmarks.
Jie Zhang 0090, Kangneng Zhou, Yan Luximon, Tong-Yee Lee, Ping Li 0016
IEEE Trans. Multim.4
2025 SGG-Nets: Generic Rotation-Invariant Plugin Networks for Point Cloud Analysis
abstract
Rotation invariance is a crucial requirement for the analysis of 3D point clouds. However, current methods often achieve rotation invariance by employing specific network designs. These networks, though perform well on rotation-aware tasks, is inferior in general tasks such as classification and segmentation. On the other hand, many powerful point processing networks, such as PointNet++, DGCNN, etc., have general point processing abilities, but do not own the property of rotation invariance. In this paper, we propose a standalone rotation-invariant convolution operator called SGGConv (Spherical Geometric Graph-based Convolution) and two ways integrating it with common point-based networks. The networks equipped with SGGConvs are called SGG-Nets which promote the rotation-invariance ability of regular point networks without modifying their network architectures much. Our contributions are three-fold. First, we propose a rotation-invariant feature descriptor, namely Spherical Geometry Descriptor (SGD), which captures point-pair features in a Local Spherical Coordinate System (LSCS). Second, we propose the SGGConv based on SGD and LSCS with an efficient Graph-based Spherical Feature Passing (GSFP) mechanism. Thirdly, we define two modules S-SGGConvMdl and M-SGGConvMdl, which are used to integrate SGGConv into baseline point nets. We test SGG-Nets, such as SGG-PointNet++, SGG-DGCNN, SGG-RIConv++, on representative point cloud datasets. These models, equipped with our SGGConvs, not only enhance the rotation-invariance of the baseline network but also improve its performance on point cloud analysis tasks such as classification and part segmentation, without incurring too much computational overhead.
Jian Zhu 0001, Jianrong Yan, Jiebin Huang, Yongwei Nie, Bin Sheng 0001, Tong-Yee Lee
IEEE Trans. Multim.6
2025 B4M: Breaking Low-Rank Adapter for Making Content-Style Customization
abstract
Personalized generation paradigms empower designers to customize visual intellectual property with the help of textual descriptions by adapting pre-trained text-to-image models on a few images. Recent studies focus on simultaneously customizing content and detailed visual style in images but often struggle with entangling the two. In this study, we reconsider the customization of content and style concepts from the perspective of parameter space construction. Unlike existing methods that utilize a shared parameter space for content and style learning, we propose a novel framework that separates the parameter space to facilitate individual learning of content and style by introducing “partly learnable projection” (PLP) matrices to separate the original adapters into divided sub-parameter spaces. A “ break-for-make ” customization learning pipeline based on PLP is proposed: we first break the original adapters into “up projection” and “down projection” for content and style concept under orthogonal prior and then make the entity parameter space by reconstructing the content and style PLP matrices by using Riemannian preconditioning to adaptively balance content and style learning. Experiments on various styles, including textures, materials, and artistic style, show that our method outperforms state-of-the-art single/multiple concept learning pipelines regarding content-style-prompt alignment. Code is available at https://github.com/ICTMCG/Break-for-make .
Fan Tang, Juan Cao 0001, Yuxin Zhang 0006, Oliver Deussen, Weiming Dong, Jintao Li 0001, Tong-Yee Lee
ACM Trans. Graph.8
2025 Computer-Aided Colorization State-of-the-Science: A Survey
abstract
This article reviews published research in the field of computer-aided colorization technology. We argue that within this context, the colorization task can be considered to originate from computer graphics, advance by introducing computer vision, and progress towards the fusion of vision and graphics. Hence, we propose a specific taxonomy and organize the research work chronologically. We extend the existing reconstruction-based colorization evaluation techniques on the basis that aesthetic assessment should be introduced to ensure the computer-coloredimages closely satisfy human visual-related requirements. We then perform an aesthetic assessment using the proposed metric and existing evaluations, comparing the colorization performance of seven representative unconditional colorization models. Finally, we identify unresolved issues and propose fruitful areas for future research and development.
Yu Cao 0019, Xin Duan, Xiangqiao Meng, P. Y. Mok 0001, Ping Li 0016, Tong-Yee Lee
IEEE Trans. Vis. Comput. Graph.6
2025 MSEmbGAN: Multi-Stitch Embroidery Synthesis via Region-Aware Texture Generation
abstract
Convolutional neural networks (CNNs) are widely used for embroidery feature synthesis from images. However, they are still unable to predict diverse stitch types, which makes it difficult for the CNNs to effectively extract stitch features. In this paper, we propose a multi-stitch embroidery generative adversarial network (MSEmbGAN) that uses a region-aware texture generation sub-network to predict diverse embroidery features from images. To the best of our knowledge, our work is the first CNN-based generative adversarial network to succeed in this task. Our region-aware texture generation sub-network detects multiple regions in the input image using a stitch classifier and generates a stitch texture for each region based on its shape features. We also propose a colorization network with a color feature extractor, which helps achieve full image color consistency by requiring the color attributes of the output to closely resemble the input image. Because of the current lack of labeled embroidery image datasets, we provide a new multi-stitch embroidery dataset that is annotated with three single-stitch types and one multi-stitch type. Our dataset, which includes more than 30K high-quality multi-stitch embroidery images, more than 13K aligned content-embroidered images, and more than 17K unaligned images, is currently the largest embroidery dataset accessible, as far as we know. Quantitative and qualitative experimental results, including a qualitative user study, show that our MSEmbGAN outperforms current state-of-the-art embroidery synthesis and style-transfer methods on all evaluation indicators.
Xinrong Hu, Ping Li 0016, Bin Sheng 0001, Tong-Yee Lee
IEEE Trans. Vis. Comput. Graph.7
2025 CreativeSynth: Cross-Art-Attention for Artistic Image Synthesis With Multimodal Diffusion
abstract
Although remarkable progress has been made in image style transfer, style is just one of the components of artistic paintings. Directly transferring extracted style features to natural images often results in outputs with obvious synthetic traces. This is because key painting attributes including layout, perspective, shape, and semantics often cannot be conveyed and expressed through style transfer. Large-scale pretrained text-to-image generation models have demonstrated their capability to synthesize a vast amount of high-quality images. However, even with extensive textual descriptions, it is challenging to fully express the unique visual properties and details of paintings. Moreover, generic models often disrupt the overall artistic effect when modifying specific areas, making it more complicated to achieve a unified aesthetic in artworks. Our main novel idea is to integrate multimodal semantic information as a synthesis guide into artworks, rather than transferring style to the real world. We also aim to reduce the disruption to the harmony of artworks while simplifying the guidance conditions. Specifically, we propose an innovative multi-task unified framework called CreativeSynth, based on the diffusion model with the ability to coordinate multimodal inputs. CreativeSynth combines multimodal features with customized attention mechanisms to seamlessly integrate real-world semantic content into the art domain through Cross-Art-Attention for aesthetic maintenance and semantic fusion. We demonstrate the results of our method across a wide range of different art categories, proving that CreativeSynth bridges the gap between generative models and artistic expression.
Nisha Huang, Weiming Dong, Yuxin Zhang 0006, Fan Tang, Ronghui Li, Chongyang Ma, Xiu Li 0001, Tong-Yee Lee, Changsheng Xu
IEEE Trans. Vis. Comput. Graph.8
2025 Cartoon Animation Outpainting With Region-Guided Motion Inference
abstract
Cartoon animation video is a popular visual entertainment form worldwide, however many classic animations were produced in a 4:3 aspect ratio that is incompatible with modern widescreen displays. Existing methods like cropping lead to information loss while retargeting causes distortion. Animation companies still rely on manual labor to renovate classic cartoon animations, which is tedious and labor-intensive, but can yield higher-quality videos. Conventional extrapolation or inpainting methods tailored for natural videos struggle with cartoon animations due to the lack of textures in anime, which affects the motion estimation of the objects. In this article, we propose a novel framework designed to automatically outpaint 4:3 anime to 16:9 via region-guided motion inference. Our core concept is to identify the motion correspondences between frames within a sequence in order to reconstruct missing pixels. Initially, we estimate optical flow guided by region information to address challenges posed by exaggerated movements and solid-color regions in cartoon animations. Subsequently, frames are stitched to produce a pre-filled guide frame, offering structural clues for the extension of optical flow maps. Finally, a voting and fusion scheme utilizes learned fusion weights to blend the aligned neighboring reference frames, resulting in the final outpainting frame. Extensive experiments confirm the superiority of our approach over existing methods.
Huisi Wu, Chengze Li, Xueting Liu 0001, Zhenkun Wen, Tong-Yee Lee
IEEE Trans. Vis. Comput. Graph.6
2025 Shape Cloud Collage on Irregular Canvas
abstract
This paper addresses a challenging and novel problem in 2D shape cloud visualization: arranging irregular 2D shapes on an irregular canvas to minimize gaps and overlaps while emphasizing critical shapes by displaying them in larger sizes. The concept of a shape cloud is inspired by word clouds, which are widely used in visualization research to aesthetically summarize textual datasets by highlighting significant words with larger font sizes. We extend this concept to images, introducing shape clouds as a powerful and expressive visualization tool, guided by the principle that "a picture is worth a thousand words. Despite the potential of this approach, solutions in this domain remain largely unexplored." To bridge this gap, we develop a 2D shape cloud collage framework that compactly arranges 2D shapes, emphasizing important objects with larger sizes, analogous to the principles of word clouds. This task presents unique challenges, as existing 2D shape layout methods are not designed for scalable irregular packing. Applying these methods often results in suboptimal layouts, such as excessive empty spaces or inaccurate representations of the underlying data. To overcome these limitations, we propose a novel layout framework that leverages recent advances in differentiable optimization. Specifically, we formulate the irregular packing problem as an optimization task, modeling the object arrangement process as a differentiable pipeline. This approach enables fast and accurate end-to-end optimization, producing high-quality layouts. Experimental results show that our system efficiently creates visually appealing and high-quality shape clouds on arbitrary canvas shapes, outperforming existing methods.
Sheng-Yi Yao, Dong-Yi Wu, Thi Ngoc Hanh Le, Tong-Yee Lee
IEEE Trans. Vis. Comput. Graph.4
2025 MotionCrafter: Plug-and-Play Motion Guidance for Diffusion Models
abstract
The essence of a video lies in the dynamic motions. While text-to-video generative diffusion models have made significant strides in creating diverse content, effectively controlling specific motions through text prompts remains a challenge. By utilizing user-specified reference videos, the more precise guidance for character actions, object movements, and camera movements can be achieved. This gives rise to the task of motion customization, where the primary challenge lies in effectively decoupling the appearance and motion within a video clip. To address this challenge, we introduce MotionCrafter, a novel one-shot instance-guided motion customization method that is suitable for both pre-trained text-to-video and text-to-image diffusion models. MotionCrafter employs a parallel spatial-temporal architecture that integrates the reference motion into the temporal component of the base model, while independently adjusting the spatial module for character or style control. To enhance the disentanglement of motion and appearance, we propose an innovative dual-branch motion disentanglement approach, which includes a motion disentanglement loss and an appearance prior enhancement strategy. To facilitate more efficient learning of motions, we further propose a novel timestep-layered tuning strategy that directs the diffusion model to focus on motion-level information. Through comprehensive quantitative and qualitative experiments, along with user preference tests, we demonstrate that MotionCrafter can successfully integrate dynamic motions while maintaining the coherence and quality of the base model, providing a wide range of appearance generation capabilities. MotionCrafter can be applied to various personalized backbones in the community to generate videos with a variety of artistic styles.
Yuxin Zhang 0006, Weiming Dong, Fan Tang, Nisha Huang, Chongyang Ma, Pengfei Wan 0001, Tong-Yee Lee, Changsheng Xu
IEEE Trans. Vis. Comput. Graph.8
2025 A Comprehensive Evaluation of Arbitrary Image Style Transfer Methods
abstract
Despite the remarkable process in the field of arbitrary image style transfer (AST), inconsistent evaluation continues to plague style transfer research. Existing methods often suffer from limited objective evaluation and inconsistent subjective feedback, hindering reliable comparisons among AST variants. In this study, we propose a multi-granularity assessment system that combines standardized objective and subjective evaluations. We collect a fine-grained dataset considering a range of image contexts such as different scenes, object complexities, and rich parsing information from multiple sources. Objective and subjective studies are conducted using the collected dataset. Specifically, we innovate on traditional subjective studies by developing an online evaluation system utilizing a combination of point-wise, pair-wise, and group-wise questionnaires. Finally, we bridge the gap between objective and subjective evaluations by examining the consistency between the results from the two studies. We experimentally evaluate CNN-based, flow-based, transformer-based, and diffusion-based AST methods by the proposed multi-granularity assessment system, which lays the foundation for a reliable and robust evaluation. Providing standardized measures, objective data, and detailed subjective feedback empowers researchers to make informed comparisons and drive innovation in this rapidly evolving field.
Zijun Zhou, Fan Tang, Yuxin Zhang 0006, Oliver Deussen, Juan Cao 0001, Weiming Dong, Xiangtao Li, Tong-Yee Lee
IEEE Trans. Vis. Comput. Graph.8
2024 Make-Your-Anchor: A Diffusion-based 2D Avatar Generation Framework
abstract
Despite the remarkable process of talking-head-based avatar-creating solutions, directly generating anchor-style videos with full-body motions remains challenging. In this study, we propose Make-Your-Anchor, a novel system necessitating only a one-minute video clip of an individual for training, subsequently enabling the automatic generation of anchor-style videos with precise torso and hand movements. Specifically, we finetune a proposed structure-guided diffusion model on input video to render 3D mesh conditions into human appearances. We adopt a two-stage training strategy for the diffusion model, effectively binding movements with specific appearances. To produce arbitrary long temporal video, we extend the 2D U-Net in the frame-wise diffusion model to a 3D style without additional training cost, and a simple yet effective batch-overlapped temporal denoising module is proposed to bypass the constraints on video length during inference. Finally, a novel identity-specific face enhancement module is introduced to improve the visual quality of facial regions in the output videos. Comparative experiments demonstrate the effectiveness and superiority of the system in terms of visual quality, temporal coherence, and identity preservation, outperforming SOTA diffusion/non-diffusion methods. Project page: https://github.com/ICTMCG/Make-Your-Anchor.
Ziyao Huang 0002, Fan Tang, Yong Zhang 0034, Xiaodong Cun, Juan Cao 0001, Jintao Li 0001, Tong-Yee Lee
CVPR7
2024 Lighting Image/Video Style Transfer Methods by Iterative Channel Pruning
abstract
Deploying style transfer methods on resource-constrained devices is challenging, which limits their real-world applicability. To tackle this issue, we propose using pruning techniques to accelerate various visual style transfer methods. We argue that typical pruning methods may not be well-suited for style transfer methods and present an iterative correlation-based channel pruning (ICCP) strategy for encoder-transform-decoder-based image/video style transfer models. The correlation-based channel regularization preserves the feature distributions for content and style references, and the iterative pruning strategy prevents layer collapse when pruning on the encoder-decoder structure. Experiments demonstrate that the proposed ICCP can generate visual competitive results compared to SOTA style transfer methods and significantly reduces the number of parameters (at least 70K) and inference time. Model is available at https://github.com/wukx-wukx/ICCP.
Kexin Wu, Fan Tang, Oliver Deussen, Thi Ngoc Hanh Le, Weiming Dong, Tong-Yee Lee
ICASSP7
2024 Dance-to-Music Generation with Encoder-based Textual Inversion
abstract
The seamless integration of music with dance movements is essential for communicating the artistic intent of a dance piece. This alignment also significantly improves the immersive quality of gaming experiences and animation productions. Although there has been remarkable advancement in creating high-fidelity music from textual descriptions, current methodologies mainly focus on modulating overall characteristics such as genre and emotional tone. They often overlook the nuanced management of temporal rhythm, which is indispensable in crafting music for dance, since it intricately aligns the musical beats with the dancers’ movements. Recognizing this gap, we propose an encoder-based textual inversion technique to augment text-to-music models with visual control, facilitating personalized music generation. Specifically, we develop dual-path rhythm-genre inversion to effectively integrate the rhythm and genre of a dance motion sequence into the textual space of a text-to-music model. Contrary to traditional textual inversion methods, which directly update text embeddings to reconstruct a single target object, our approach utilizes separate rhythm and genre encoders to obtain text embeddings for two pseudo-words, adapting to the varying rhythms and genres. We collect a new dataset called In-the-wild Dance Videos (InDV) and demonstrate that our approach outperforms state-of-the-art methods across multiple evaluation metrics. Furthermore, our method is able to adapt to changes in tempo and effectively integrates with the inherent text-guided generation capability of the pre-trained model. Our source code and demo videos are available at https://github.com/lsfhuihuiff/Dance-to-music_Siggraph_Asia_2024.
Sifei Li, Weiming Dong, Yuxin Zhang 0006, Fan Tang, Chongyang Ma, Oliver Deussen, Tong-Yee Lee, Changsheng Xu
SIGGRAPH Asia7
2024 Deep learning-based importance map for content-aware media retargeting
Thi Ngoc Hanh Le, Tong-Yee Lee, Shih-Syun Lin, Weiming Dong
Multim. Tools Appl.2
2024 3DSN-Net: A 3-D Scale-Aware convNet With Nonlocal Context Guidance for Kidney and Tumor Segmentation From CT Volumes
abstract
Automatic kidney and tumor segmentation from CT volumes is a critical prerequisite/tool for diagnosis and surgical treatment (such as partial nephrectomy). However, it remains a particularly challenging issue as kidneys and tumors often exhibit large-scale variations, irregular shapes, and blurring boundaries. We propose a novel 3-D network to comprehensively tackle these problems; we call it 3DSN-Net. Compared with existing solutions, it has two compelling characteristics. First, with a new scale-aware feature extraction (SAFE) module, the proposed 3DSN-Net is capable of adaptively selecting appropriate receptive fields according to the sizes of targets instead of indiscriminately enlarging them, which is particularly essential for improving the segmentation accuracy of the tumor with large scale variation. Second, we propose a novel yet efficient nonlocal context guidance (NCG) mechanism to capture global dependencies to tackle irregular shapes and blurring boundaries of kidneys and tumors. Instead of directly harnessing a 3-D NCG mechanism, which makes the number of parameters exponentially increase and hence the network difficult to be trained under limited training data, we develop a 2.5D NCG mechanism based on projections of feature cubes, which achieves a tradeoff between segmentation accuracy and network complexity. We extensively evaluate the proposed 3DSN-Net on the famous KiTS dataset with many challenging kidney and tumor cases. Experimental results demonstrate our solution consistently outperforms state-of-the-art 3-D networks after being equipped with scale aware and NCG mechanisms, particularly for tumor segmentation.
Huisi Wu, Baiming Zhang, Zhuoying Li, Harry Qin, Tong-Yee Lee
IEEE Trans. Cybern.5
2024 Identity-Preserving Face Swapping via Dual Surrogate Generative Models
abstract
In this study, we revisit the fundamental setting of face-swapping models and reveal that only using implicit supervision for training leads to the difficulty of advanced methods to preserve the source identity. We propose a novel reverse pseudo-input generation approach to offer supplemental data for training face-swapping models, which addresses the aforementioned issue. Unlike the traditional pseudo-label-based training strategy, we assume that arbitrary real facial images could serve as the ground-truth outputs for the face-swapping network and try to generate corresponding input pair data. Specifically, we involve a source-creating surrogate that alters the attributes of the real image while keeping the identity, and a target-creating surrogate intends to synthesize attribute-preserved target images with different identities. Our framework, which utilizes proxy-paired data as explicit supervision to direct the face-swapping training process, partially fulfills a credible and effective optimization direction to boost the identity-preserving capability. We design explicit and implicit adaption strategies to better approximate the explicit supervision for face swapping. Quantitative and qualitative experiments on FF++, FFHQ, and wild images show that our framework could improve the performance of various face-swapping pipelines in terms of visual fidelity and ID preserving. Furthermore, we display applications with our method on re-aging, swappable attribute customization, cross-domain, and video face swapping. Code is available under https://github.com/ ICTMCG/CSCS.
Ziyao Huang 0002, Fan Tang, Yong Zhang 0034, Juan Cao 0001, Sheng Tang, Jintao Li 0001, Tong-Yee Lee
ACM Trans. Graph.8
2024 Suitable and Style-Consistent Multi-Texture Recommendation for Cartoon Illustrations
abstract
Texture plays an important role in cartoon illustrations to display object materials and enrich visual experiences. Unfortunately, manually designing and drawing an appropriate texture is not easy even for proficient artists, let alone novice or amateur people. While there exist tons of textures on the Internet, it is not easy to pick an appropriate one using traditional text-based search engines. Although several texture pickers have been proposed, they still require the users to browse the textures by themselves, which is still labor-intensive and time-consuming. In this article, an automatic texture recommendation system is proposed for recommending multiple textures to replace a set of user-specified regions in a cartoon illustration with visually pleasant look. Two measurements, the suitability measurement and the style-consistency measurement, are proposed to make sure that the recommended textures are suitable for cartoon illustration and at the same time mutually consistent in style. The suitability is measured based on the synthesizability, cartoonity, and region fitness of textures. The style-consistency is predicted using a learning-based solution since it is subjective to judge whether two textures are consistent in style. An optimization problem is formulated and solved via the genetic algorithm. Our method is validated on various cartoon illustrations, and convincing results are obtained.
Huisi Wu, Zhaoze Wang, Xueting Liu 0001, Tong-Yee Lee
ACM Trans. Multim. Comput. Commun. Appl.5
2024 AnimeDiffusion: Anime Diffusion Colorization
abstract
Being essential in animation creation, colorizing anime line drawings is usually a tedious and time-consuming manual task. Reference-based line drawing colorization provides an intuitive way to automatically colorize target line drawings using reference images. The prevailing approaches are based on generative adversarial networks (GANs), yet these methods still cannot generate high-quality results comparable to manually-colored ones. In this article, a new AnimeDiffusion approach is proposed via hybrid diffusions for the automatic colorization of anime face line drawings. This is the first attempt to utilize the diffusion model for reference-based colorization, which demands a high level of control over the image synthesis process. To do so, a hybrid end-to-end training strategy is designed, including phase 1 for training diffusion model with classifier-free guidance and phase 2 for efficiently updating color tone with a target reference colored image. The model learns denoising and structure-capturing ability in phase 1, and in phase 2, the model learns more accurate color information. Utilizing our hybrid training strategy, the network convergence speed is accelerated, and the colorization performance is improved. Our AnimeDiffusion generates colorization results with semantic correspondence and color consistency. In addition, the model has a certain generalization performance for line drawings of different line styles. To train and evaluate colorization methods, an anime face line drawing colorization benchmark dataset, containing 31,696 training data and 579 testing data, is introduced and shared. Extensive experiments and user studies have demonstrated that our proposed AnimeDiffusion outperforms state-of-the-art GAN-based methods and another diffusion-based model, both quantitatively and qualitatively.
Yu Cao 0019, Xiangqiao Meng, P. Y. Mok 0001, Tong-Yee Lee, Xueting Liu 0001, Ping Li 0016
IEEE Trans. Vis. Comput. Graph.4
2024 Retargeting Video With an End-to-End Framework
abstract
Video holds significance in computer graphics applications. Because of the heterogeneous of digital devices, retargeting videos becomes an essential function to enhance user viewing experience in such applications. In the research of video retargeting, preserving the relevant visual content in videos, avoiding flicking, and processing time are the vital challenges. Extending image retargeting techniques to the video domain is challenging due to the high running time. Prior work of video retargeting mainly utilizes time-consuming preprocessing to analyze frames. Plus, being tolerant of different video content, avoiding important objects from shrinking, and the ability to play with arbitrary ratios are the limitations that need to be resolved in these systems requiring investigation. In this paper, we present an end-to-end RETVI method to retarget videos to arbitrary aspect ratios. We eliminate the computational bottleneck in the conventional approaches by designing RETVI with two modules, content feature analyzer (CFA) and adaptive deforming estimator (ADE). The extensive experiments and evaluations show that our system outperforms previous work in quality and running time.
Thi Ngoc Hanh Le, Huiguang Huang, Tong-Yee Lee
IEEE Trans. Vis. Comput. Graph.4
2024 Regenerating Arbitrary Video Sequences With Distillation Path-Finding
abstract
If the video has long been mentioned as a widespread visualization form, the animation sequence in the video is mentioned as storytelling for people. Producing an animation requires intensive human labor from skilled professional artists to obtain plausible animation in both content and motion direction, incredibly for animations with complex content, multiple moving objects, and dense movement. This article presents an interactive framework to generate new sequences according to the users' preference on the starting frame. The critical contrast of our approach versus prior work and existing commercial applications is that novel sequences with arbitrary starting frame are produced by our system with a consistent degree in both content and motion direction. To achieve this effectively, we first learn the feature correlation on the frameset of the given video through a proposed network called RSFNet. Then, we develop a novel path-finding algorithm, SDPF, which formulates the knowledge of motion directions of the source video to estimate the smooth and plausible sequences. The extensive experiments show that our framework can produce new animations on the cartoon and natural scenes and advance prior works and commercial applications to enable users to obtain more predictable results.
Thi Ngoc Hanh Le, Sheng-Yi Yao, Chun-Te Wu, Tong-Yee Lee
IEEE Trans. Vis. Comput. Graph.4
2024 Image Collage on Arbitrary Shape via Shape-Aware Slicing and Optimization
abstract
Image collage is a very useful tool for visualizing an image collection. Most of the existing methods and commercial applications for generating image collages are designed on simple shapes, such as rectangular and circular layouts. This greatly limits the use of image collages in some artistic and creative settings. Although there are some methods that can generate irregularly-shaped image collages, they often suffer from severe image overlapping and excessive blank space. This prevents such methods from being effective information communication tools. In this article, we present a shape slicing algorithm and an optimization scheme that can create image collages of arbitrary shapes in an informative and visually pleasing manner given an input shape and an image collection. To overcome the challenge of irregular shapes, we propose a novel algorithm, called Shape-Aware Slicing, which partitions the input shape into cells based on medial axis and binary slicing tree. Shape-Aware Slicing,which is designed specifically for irregular shapes, takes human perception and shape structure into account to generate visually pleasing partitions. Then, the layout is optimized by analyzing input images with the goal of maximizing the total salient regions of the images. To evaluate our method, we conduct extensive experiments and compare our results against previous work. The evaluations show that our proposed algorithm can efficiently arrange image collections on irregular shapes and create visually superior results than prior work and existing commercial tools.
Dong-Yi Wu, Thi Ngoc Hanh Le, Sheng-Yi Yao, Yun-Chen Lin, Tong-Yee Lee
IEEE Trans. Vis. Comput. Graph.5
2024 MeshWGAN: Mesh-to-Mesh Wasserstein GAN With Multi-Task Gradient Penalty for 3D Facial Geometric Age Transformation
abstract
As the metaverse develops rapidly, 3D facial age transformation is attracting increasing attention, which may bring many potential benefits to a wide variety of users, e.g., 3D aging figures creation, 3D facial data augmentation and editing. Compared with 2D methods, 3D face aging is an underexplored problem. To fill this gap, we propose a new mesh-to-mesh Wasserstein generative adversarial network (MeshWGAN) with a multi-task gradient penalty to model a continuous bi-directional 3D facial geometric aging process. To the best of our knowledge, this is the first architecture to achieve 3D facial geometric age transformation via real 3D scans. As previous image-to-image translation methods cannot be directly applied to the 3D facial mesh, which is totally different from 2D images, we built a mesh encoder, decoder, and multi-task discriminator to facilitate mesh-to-mesh transformations. To mitigate the lack of 3D datasets containing children's faces, we collected scans from 765 subjects aged 5-17 in combination with existing 3D face databases, which provided a large training dataset. Experiments have shown that our architecture can predict 3D facial aging geometries with better identity preservation and age closeness compared to 3D trivial baselines. We also demonstrated the advantages of our approach via various 3D face-related graphics applications.
Jie Zhang 0090, Kangneng Zhou, Yan Luximon, Tong-Yee Lee, Ping Li 0016
IEEE Trans. Vis. Comput. Graph.4
2023 PhotoHelper: Portrait Photographing Guidance Via Deep Feature Retrieval and Fusion
abstract
We introduce a new photographing guidance (PhotoHelper) for amateur photographers to enhance their portrait photo quality using deep feature retrieval and fusion. In our model, we comprehensively integrate empirical aesthetic rules, traditional machine learning algorithms and deep neural networks to extract different kinds of features in both color and space aspects. With these features, we build a modified random forest with a structured photograph collection to identify types of photos. We also define the composition matching score to measure the similarity between the given photo and the reference photo. By combining all of the above processes, a one-stop deep portrait photographing guidance is constructed to provide users with professional reference photographs that are similar to the current scene and automatically generate spatial composition guidance according to the user-selected reference photo. Experiments and evaluations show that the aesthetic quality of portrait photos can be significantly improved via the composition guidance of our photographing guidance approach.
Bin Sheng 0001, Ping Li 0016, Tong-Yee Lee
IEEE Trans. Multim.4
2023 Creative and Progressive Interior Color Design with Eye-tracked User Preference
abstract
Interior scene colorization is vastly demanded in areas such as personalized architecture design. Existing works either require manual efforts to colorize individual objects or conform to fixed color patterns automatically learned from prior knowledge, whilst neglecting user preference. Quantitatively identifying user preferences is challenging, particularly at the early stage of the design process. The 3D setup also presents new challenges as the inhabitant can observe from any possible viewpoint. We propose a representative view selection method based on visual attention and a progressive preference inference model. We particularly focus on the progressive integration of eye-tracked user preference, which enables the assistance in creativity support and allows the possibility of convergent thinking. A series of user studies have been conducted to validate the effectiveness of the proposed view selection method, preference inference model and the creativity support mechanism.
Shihui Guo, Yubin Shi, Pintong Xiao, Yinan Fu, Juncong Lin, Wei Zeng 0004, Tong-Yee Lee
ACM Trans. Comput. Hum. Interact.7
2023 ProSpect: Prompt Spectrum for Attribute-Aware Personalization of Diffusion Models
abstract
Personalizing generative models offers a way to guide image generation with user-provided references. Current personalization methods can invert an object or concept into the textual conditioning space and compose new natural sentences for text-to-image diffusion models. However, representing and editing specific visual attributes such as material, style, and layout remains a challenge, leading to a lack of disentanglement and editability. To address this problem, we propose a novel approach that leverages the step-by-step generation process of diffusion models, which generate images from low to high frequency information, providing a new perspective on representing, generating, and editing images. We develop the Prompt Spectrum Space P*, an expanded textual conditioning space, and a new image representation method called ProSpect. ProSpect represents an image as a collection of inverted textual token embeddings encoded from per-stage prompts, where each prompt corresponds to a specific generation stage (i.e., a group of consecutive steps) of the diffusion model. Experimental results demonstrate that P* and ProSpect offer better disentanglement and controllability compared to existing methods. We apply ProSpect in various personalized attribute-aware image generation applications, such as image-guided or text-driven manipulations of materials, style, and layout, achieving previously unattainable results from a single image input without fine-tuning the diffusion models. Our source code is available at https://github.com/zyxElsa/ProSpect.
Yuxin Zhang 0006, Weiming Dong, Fan Tang, Nisha Huang, Chongyang Ma, Tong-Yee Lee, Oliver Deussen, Changsheng Xu
ACM Trans. Graph.7
2023 A Unified Arbitrary Style Transfer Framework via Adaptive Contrastive Learning
abstract
This work presents Unified Contrastive Arbitrary Style Transfer (UCAST), a novel style representation learning and transfer framework, that can fit in most existing arbitrary image style transfer models, such as CNN-based, ViT-based, and flow-based methods. As the key component in image style transfer tasks, a suitable style representation is essential to achieve satisfactory results. Existing approaches based on deep neural networks typically use second-order statistics to generate the output. However, these hand-crafted features computed from a single image cannot leverage style information sufficiently, which leads to artifacts such as local distortions and style inconsistency. To address these issues, we learn style representation directly from a large number of images based on contrastive learning by considering the relationships between specific styles and the holistic style distribution. Specifically, we present an adaptive contrastive learning scheme for style transfer by introducing an input-dependent temperature. Our framework consists of three key components: a parallel contrastive learning scheme for style representation and transfer, a domain enhancement (DE) module for effective learning of style distribution, and a generative network for style transfer. Qualitative and quantitative evaluations show the results of our approach are superior to those obtained via state-of-the-art methods. The code is available at https://github.com/zyxElsa/CAST_pytorch .
Yuxin Zhang 0006, Fan Tang, Weiming Dong, Chongyang Ma, Tong-Yee Lee, Changsheng Xu
ACM Trans. Graph.6
2023 Structure-aware Video Style Transfer with Map Art
abstract
Changing the style of an image/video while preserving its content is a crucial criterion to access a new neural style transfer algorithm. However, it is very challenging to transfer a new map art style to a certain video in which “content” comprises a map background and animation objects. In this article, we present a novel comprehensive system that solves the problems in transferring map art style in such video. Our system takes as input an arbitrary video, a map image, and an off-the-shelf map art image. It then generates an artistic video without damaging the functionality of the map and the consistency in details. To solve this challenge, we propose a novel network, Map Art Video Network (MAViNet), the tailored objective functions, and a rich training set with rich animation contents and different map structures. We have evaluated our method on various challenging cases and many comparisons with those of the related works. Our method substantially outperforms state-of-the-art methods in terms of visual quality and meets the mentioned criteria in this research domain.
Thi Ngoc Hanh Le, Ya-Hsuan Chen, Tong-Yee Lee
ACM Trans. Multim. Comput. Commun. Appl.3
2023 Animating Still Natural Images Using Warping
abstract
From a single still image, a looping video could be generated by imparting subtle motion to objects in the image. The results are a hybrid of photography and video. They contain gentle motion in some objects, while the rest of the image remains still. Existing techniques are successful in animating such images. However, there are still some drawbacks that need to be investigated, such as too-large computation time necessary to retrieve the matched videos or the challenges of controlling the desired motion not only in terms of a single region but also in terms of consistency in regions. In this work, we address these issues by proposing an interactive system with a novel warping method. The key idea of our approach is to utilize user’s annotations to impart motion to certain objects. With two proposed phases in terms of preserve-curve-warping and cycle warping, a looping video is generated. We demonstrate the effectiveness of our method via various experimental challenging results and evaluations. We show that with a simple and lightweight method, our system is able to deal with animating a still image’s problems and results in realistic motion and appealing videos. In addition, using our proposed system, it is easy to create plausible animation using simple user annotations without referencing the video database or machine learning models and allows ordinary users with minimal expertise to produce compelling results.
Thi Ngoc Hanh Le, Chih-Kuo Yeh, Ying-Chi Lin 0003, Tong-Yee Lee
ACM Trans. Multim. Comput. Commun. Appl.4
2023 Balance-Aware Grid Collage for Small Image Collections
abstract
Grid collages (GClg) of small image collections are popular and useful in many applications, such as personal album management, online photo posting, and graphic design. In this article, we focus on how visual effects influence individual preferences through various arrangements of multiple images under such scenarios. A novel balance-aware metric is proposed to bridge the gap between multi-image joint presentation and visual pleasure. The metric merges psychological achievements into the field of grid collage. To capture user preference, a bonus mechanism related to a user-specified special location in the grid and uniqueness values of the subimages is integrated into the metric. An end-to-end reinforcement learning mechanism empowers the model without tedious manual annotations. Experiments demonstrate that our metric can evaluate the GClg visual balance in line with human subjective perception, and the model can generate visually pleasant GClg results, which is comparable to manual designs.
Fan Tang, Weiming Dong, Feiyue Huang, Tong-Yee Lee, Changsheng Xu
IEEE Trans. Vis. Comput. Graph.5
2022 Learning a perceptual manifold with deep features for animation video resequencing
Charles C. Morace, Thi Ngoc Hanh Le, Sheng-Yi Yao, Shang-Wei Zhang, Tong-Yee Lee
Multim. Tools Appl.5
2022 Generating Virtual Wire Sculptural Art from 3D Models
abstract
Wire sculptures are objects sculpted by the use of wires. In this article, we propose practical methods to create 3D virtual wire sculptural art from a given 3D model. In contrast, most of the previous 3D wire art results are reconstructed from input 2D wire art images. Artists usually tend to design their wire art with a single wire if possible. If not possible, they try to create it with the least number of wires. To follow this general design trend, our proposed method generates 3D virtual wire art with the minimum number of continuous wire lines. To achieve this goal, we first adopt a greedy approach to extract important edges of a given 3D model. These extracted important edges become the basis for the subsequent lines to roughly represent the shape of the input model. Then, we connect them with the minimum number of continuous wire lines by the order obtained by optimally solving a traveling salesman problem with some constraints. Finally, we smooth the obtained 3D wires to simulate the real 3D wire results by artists. In addition, we also provide a user interface to control the winding of wires by their design preference. Finally, we experimentally show our 3D virtual wire results and evaluate these created results. As a result, the proposed method is computed effectively and interactively, and results are appealing and comparable to real 3D wire art work.
Chih-Kuo Yeh, Thi Ngoc Hanh Le, Zhi-Ying Hou, Tong-Yee Lee
ACM Trans. Multim. Comput. Commun. Appl.4
2022 C3 Assignment: Camera Cubemap Color Assignment for Creative Interior Design
abstract
Color design for 3D indoor scenes is a challenging problem due to many factors that need to be balanced. Although learning from images is a commonly adopted strategy, this strategy may be more suitable for natural scenes in which objects tend to have relatively fixed colors. For interior scenes consisting mostly of man-made objects, creative yet reasonable color assignments are expected. We propose$C^{3}$C3Assignment, a system providing diverse suggestions for interior color design while satisfying general global and local rules including color compatibility, color mood, contrast, and user preference. We extend these constraints from the image domain to$\mathbb {R}^3$, and formulate 3D interior color design as an optimization problem. The design is accomplished in an omnidirectional manner to ensure a comfortable experience when the inhabitant observes the interior scene from possible positions and directions. We design a surrogate-assisted evolutionary algorithm to efficiently solve the highly nonlinear optimization problem for interactive applications, and investigate the system performance concerning problem complexity, solver convergence, and suggestion diversity. Preliminary user studies have been conducted to validate the rule extension from 2D to 3D and to verify system usability.
Juncong Lin, Pintong Xiao, Yinan Fu, Yubin Shi, Hongran Wang, Shihui Guo, Ying He 0001, Tong-Yee Lee
IEEE Trans. Vis. Comput. Graph.8
2022 WYSIWYG Design of Hypnotic Line Art
abstract
Hypnotic line art is a modern form in which white narrow curved ribbons, with the width and direction varying along each path over a black background, provide a keen sense of 3D objects regarding surface shapes and topological contours. However, the procedure of manually creating such line art work can be quite tedious and time-consuming. In this article, we present an interactive system that offers a What-You-See-Is-What-You-Get (WYSIWYG) scheme for producing hypnotic line art images by integrating and placing evenly-spaced streamlines in tensor fields. With an input picture segmented, the user just needs to sketch a few illustrative strokes to guide the construction of a tensor field for each part of the objects therein. Specifically, we propose a new method which controls, with great precision, the aesthetic layout and artistic drawing of an array of streamlines in each tensor field to emulate the style of hypnotic line art. Given several parameters for streamlines such as density, thickness, and sharpness, our system is capable of generating professional-level hypnotic line art work. With great ease of use, it allows art designers to explore a wide variety of possibilities to obtain hypnotic line art results of their own preferences.
Chih-Kuo Yeh, Zhanping Liu, I-Hsuan Lin, Eugene Zhang, Tong-Yee Lee
IEEE Trans. Vis. Comput. Graph.5
2021 Content-and-disparity-aware stereoscopic video stabilization
Shih-Syun Lin, Thi Ngoc Hanh Le, Pang-Yu Wu, Tong-Yee Lee
Multim. Tools Appl.4
2021 Map art style transfer with multi-stage framework
Chiao-Yin Shih, Ya-Hsuan Chen, Tong-Yee Lee
Multim. Tools Appl.3
2021 Structure-Aware Motion Deblurring Using Multi-Adversarial Optimized CycleGAN
abstract
Recently, Convolutional Neural Networks (CNNs) have achieved great improvements in blind image motion deblurring. However, most existing image deblurring methods require a large amount of paired training data and fail to maintain satisfactory structural information, which greatly limits their application scope. In this paper, we present an unsupervised image deblurring method based on a multi-adversarial optimized cycle-consistent generative adversarial network (CycleGAN). Although original CycleGAN can handle unpaired training data well, the generated high-resolution images are probable to lose content and structure information. To solve this problem, we utilize a multi-adversarial mechanism based on CycleGAN for blind motion deblurring to generate high-resolution images iteratively. In this multi-adversarial manner, the hidden layers of the generator are gradually supervised, and the implicit refinement is carried out to generate high-resolution images continuously. Meanwhile, we also introduce the structure-aware mechanism to enhance the structure and detail retention ability of the multi-adversarial network for deblurring by taking the edge map as guidance information and adding multi-scale edge constraint functions. Our approach not only avoids the strict need for paired training data and the errors caused by blur kernel estimation, but also maintains the structural information better with multi-adversarial learning and structure-aware mechanism. Comprehensive experiments on several benchmarks have shown that our approach prevails the state-of-the-art methods for blind image motion deblurring.
Jie Chen 0097, Bin Sheng 0001, Ping Li 0016, Ping Tan 0002, Tong-Yee Lee
IEEE Trans. Image Process.7
2021 Fast Accurate and Automatic Brushstroke Extraction
abstract
Brushstrokes are viewed as the artist’s “handwriting” in a painting. In many applications such as style learning and transfer, mimicking painting, and painting authentication, it is highly desired to quantitatively and accurately identify brushstroke characteristics from old masters’ pieces using computer programs. However, due to the nature of hundreds or thousands of intermingling brushstrokes in the painting, it still remains challenging. This article proposes an efficient algorithm for brush Stroke extraction based on a Deep neural network, i.e., DStroke. Compared to the state-of-the-art research, the main merit of the proposed DStroke is to automatically and rapidly extract brushstrokes from a painting without manual annotation, while accurately approximating the real brushstrokes with high reliability. Herein, recovering the faithful soft transitions between brushstrokes is often ignored by the other methods. In fact, the details of brushstrokes in a master piece of painting (e.g., shapes, colors, texture, overlaps) are highly desired by artists since they hold promise to enhance and extend the artists’ powers, just like microscopes extend biologists’ powers. To demonstrate the high efficiency of the proposed DStroke, we perform it on a set of real scans of paintings and a set of synthetic paintings, respectively. Experiments show that the proposed DStroke is noticeably faster and more accurate at identifying and extracting brushstrokes, outperforming the other methods.
Yunfei Fu, Hongchuan Yu, Chih-Kuo Yeh, Tong-Yee Lee, Jian J. Zhang 0001
ACM Trans. Multim. Comput. Commun. Appl.4
2021 Content-Based Visual Summarization for Image Collections
abstract
With the surge of images in the information era, people demand an effective and accurate way to access meaningful visual information. Accordingly, effective and accurate communication of information has become indispensable. In this article, we propose a content-based approach that automatically generates a clear and informative visual summarization based on design principles and cognitive psychology to represent image collections. We first introduce a novel method to make representative and nonredundant summarizations of image collections, thereby ensuring data cleanliness and emphasizing important information. Then, we propose a tree-based algorithm with a two-step optimization strategy to generate the final layout that operates as follows: (1) an initial layout is created by constructing a tree randomly based on the grouping results of the input image set; (2) the layout is refined through a coarse adjustment in a greedy manner, followed by gradient back propagation drawing on the training procedure of neural networks. We demonstrate the usefulness and effectiveness of our method via extensive experimental results and user studies. Our visual summarization algorithm can precisely and efficiently capture the main content of image collections better than alternative methods or commercial tools.
Xingjia Pan, Fan Tang, Weiming Dong, Chongyang Ma, Yiping Meng, Feiyue Huang, Tong-Yee Lee, Changsheng Xu
IEEE Trans. Vis. Comput. Graph.7
2020 Disparity-preserving image rectangularization for stereoscopic panorama
I-Cheng Yeh 0001, Shih-Syun Lin, Shuo-Tse Hung, Tong-Yee Lee
Multim. Tools Appl.4
2020 Fast character modeling with sketch-based PDE surfaces
abstract
Abstract Virtual characters are 3D geometric models of characters. They have a lot of applications in multimedia. In this paper, we propose a new physics-based deformation method and efficient character modelling framework for creation of detailed 3D virtual character models. Our proposed physics-based deformation method uses PDE surfaces. Here PDE is the abbreviation of Partial Differential Equation, and PDE surfaces are defined as sculpting force-driven shape representations of interpolation surfaces. Interpolation surfaces are obtained by interpolating key cross-section profile curves and the sculpting force-driven shape representation uses an analytical solution to a vector-valued partial differential equation involving sculpting forces to quickly obtain deformed shapes. Our proposed character modelling framework consists of global modeling and local modeling. The global modeling is also called model building, which is a process of creating a whole character model quickly with sketch-guided and template-based modeling techniques. The local modeling produces local details efficiently to improve the realism of the created character model with four shape manipulation techniques. The sketch-guided global modeling generates a character model from three different levels of sketched profile curves called primary, secondary and key cross-section curves in three orthographic views. The template-based global modeling obtains a new character model by deforming a template model to match the three different levels of profile curves. Four shape manipulation techniques for local modeling are investigated and integrated into the new modelling framework. They include: partial differential equation-based shape manipulation, generalized elliptic curve-driven shape manipulation, sketch assisted shape manipulation, and template-based shape manipulation. These new local modeling techniques have both global and local shape control functions and are efficient in local shape manipulation. The final character models are represented with a collection of surfaces, which are modeled with two types of geometric entities: generalized elliptic curves (GECs) and partial differential equation-based surfaces. Our experiments indicate that the proposed modeling approach can build detailed and realistic character models easily and quickly.
Lihua You, Xiaosong Yang, JunJun Pan, Tong-Yee Lee, Shaojun Bian, Kun Qian 0009, Zulfiqar Habib, Allah Bux Sargano, Ismail Khalid Kazmi, Jian J. Zhang 0001
Multim. Tools Appl.4
2020 Image Vectorization With Real-Time Thin-Plate Spline
abstract
The vector graphics with gradient mesh can be attributed to their compactness and scalability; however, they tend to fall short when it comes to real-time editing due to a lack of real-time rasterization and an efficient editing tool for image details. In this paper, we encode global manipulation geometries and local image details within a hybrid vector structure, using parametric patches and detailed features for localized and parallelized thin-plate spline interpolation in order to achieve good compressibility, interactive expressibility, and editability. The proposed system then automatically extracts an optimal set of detailed color features while considering the compression ratio of the image as well as reconstruction error and its characteristics applicable to the preservation of structural and irregular saliency of the image. The proposed real-time vector representation makes it possible to construct an interactive editing system for detail-maintained image magnification and color editing as well as material replacement in cross mapping, without maintaining spatial and temporal consistency while editing in a raster space. Experiments demonstrate that our representation method is superior to several state-of-the-art methods and as good as JPEG, while providing real-time editability and preserving structural and irregular saliency information.
Kuo-Wei Chen, Ying-Sheng Luo, Yu-Chi Lai, Yan-Lin Chen, Chih-Yuan Yao, Hung-Kuo Chu, Tong-Yee Lee
IEEE Trans. Multim.7
2020 Image Retargetability
abstract
Real-world applications could benefit from the ability to automatically retarget an image to different aspect ratios and resolutions while preserving its visually and semantically important content. However, not all images can be equally processed. This study introduces the notion of image retargetability to describe how well a particular image can be handled by content-aware image retargeting. We propose to learn a deep convolutional neural network to rank photo retargetability, in which the relative ranking of photo retargetability is directly modeled in the loss function. Our model incorporates the joint learning of meaningful photographic attributes and image content information, which can facilitate the regularization of the complicated retargetability rating problem. To train and analyze this model, we collect a dataset that contains retargetability scores and meaningful image attributes assigned by six expert raters. The experiments demonstrate that our unified model can generate retargetability rankings that are highly consistent with human labels. To further validate our model, we show the applications of image retargetability in retargeting method selection, retargeting method assessment and generating a photo collage.
Fan Tang, Weiming Dong, Yiping Meng, Chongyang Ma, Fuzhang Wu, Tong-Yee Lee
IEEE Trans. Multim.7
2020 Intrinsic Image Decomposition with Step and Drift Shading Separation
abstract
Decomposing an image into the shading and reflectance layers remains challenging due to its severely under-constrained nature. We present an approach based on illumination decomposition that recovers the intrinsic images without additional information, e.g., depth or user interaction. Our approach is based on the rationale that the shading component contains the step and drift channels simultaneously. We decompose the illumination into two channels: the step shading, corresponding to the sharp shading changes due to cast shadow or abrupt shape changes; the drift shading, accounting for the smooth shading variations due to gradual illumination changes or slow shape changes. Due to such transformation of turning the conventional assumption that shading has smoothness as reasonable prior, our model has the advantages in handling real images, especially with the cast shadows or strong shape edges. We also apply a much stricter edge classifier along with a reinforcement process to enhance our method. We formulate the problem using a two-parameter energy function and split it into two energy functions corresponding to the reflectance and step shading. Experiments on the MIT dataset, the IIW dataset and the MPI Sintel dataset have shown the success of our approach over the state-of-the-art methods.
Bin Sheng 0001, Ping Li 0016, Yuxi Jin, Ping Tan 0002, Tong-Yee Lee
IEEE Trans. Vis. Comput. Graph.5
2020 Depth of Field Rendering Using Multilayer-Neighborhood Optimization
abstract
Depth of field (DOF) is utilized widely to deliver artistic effects in photography. However, existing post-processing techniques for rendering DOF effects introduce visual artifacts such as color leakage, blurring discontinuity, and the partial occlusion problems which limit the application of DOF. Traditionally, occluded pixels are ignored or not well estimated although they might make key contributions to images. In this paper, we propose a new filtering approach which takes approximated occluded pixels into account to synthesize the DOF effects for images. In our approach, images are separated into different layers based on depth. Besides, we utilize adaptive PatchMatch method to estimate the intensities of occluded pixels, especially in the background region. We again propose a new multilayer-neighborhood optimization to estimate occluded pixels contributions and render the images. Finally, we apply gathering filter to achieve the rendered images with elite DOF effects. Multiple experiments have shown that our approach can handle color leakage, blurring discontinuity and partial occlusion problem while providing high-quality DOF rendering effects.
Benxuan Zhang, Bin Sheng 0001, Ping Li 0016, Tong-Yee Lee
IEEE Trans. Vis. Comput. Graph.4
2019 High Relief from Brush Painting
abstract
Relief is an art form part way between 3D sculpture and 2D painting. We present a novel approach for generating a texture-mapped high-relief model from a single brush painting. Our aim is to extract the brushstrokes from a painting and generate the individual corresponding relief proxies rather than recovering the exact depth map from the painting, which is a tricky computer vision problem, requiring assumptions that are rarely satisfied. The relief proxies of brushstrokes are then combined together to form a 2.5D high-relief model. To extract brushstrokes from 2D paintings, we apply layer decomposition and stroke segmentation by imposing boundary constraints. The segmented brushstrokes preserve the style of the input painting. By inflation and a displacement map of each brushstroke, the features of brushstrokes are preserved by the resultant high-relief model of the painting. We demonstrate that our approach is able to produce convincing high-reliefs from a variety of paintings(with humans, animals, flowers, etc.). As a secondary application, we show how our brushstroke extraction algorithm could be used for image editing. As a result, our brushstroke extraction algorithm is specifically geared towards paintings with each brushstroke drawn very purposefully, such as Chinese paintings, Rosemailing paintings, etc.
Yunfei Fu, Hongchuan Yu, Chih-Kuo Yeh, Jian J. Zhang 0001, Tong-Yee Lee
IEEE Trans. Vis. Comput. Graph.5
2018 Photo Squarization by Deep Multi-Operator Retargeting
abstract
Squared forms of photos are widely used in social media as album covers or thumbnails of image streams. In this study, we realize photo squarization by modeling Retargeting Visual Perception Issues, which reflect human perception preference toward image ratargeting. General image retargeting techniques deal with three common issues, namely, salient content, object shape, and scene composition, to preserve the important information of original image. We propose a new way based on multi-operator techniques to investigate human behavior in balancing the three issues. We establish a new dataset and observe human behavior by inviting investigators to retarget images to square manually. We propose a data-driven approach composed of perception and distillation modules by using deep learning techniques to predict human perception preference. The perception part learns the relations among the three issues, and the distillation part transfers the learned relations to a simple but effective network. Our study contributes to deep learning literature by optimizing a network index and lightening its running burden. Experimental results show that photo squarization results generated by the proposed model are consistent with human visual perception results.
Fan Tang, Weiming Dong, Xiaopeng Zhang 0001, Oliver Deussen, Tong-Yee Lee
ACM Multimedia6
2018 Geometric and Textural Blending for 3D Model Stylization
abstract
Stylizing a 3D model with characteristic shapes or appearances is common in product design, particularly in the design of 3D model merchandise, such as souvenirs, toys, furniture, and stylized items. A model stylization approach is proposed in this study. The approach combines base and style models while preserving user-specified shape features of the base model and the attractive features of the style model with limited assistance from a user. The two models are first combined at the topological level. A tree-growing technique is utilized to search for all possible combinations of the two models. Second, the models are combined at textural and geometric levels by employing a morphing technique. Results show that the proposed approach generates various appealing models and allows users to control the diversity of the output models and adjust the blending degree between the base and style models. The results of this work are also experimentally compared with those of a recent work through a user study. The comparison indicates that our results are more appealing, feature-preserving, and reasonable than those of the compared previous study. The proposed system allows product designers to easily explore design possibilities and assists novice users in creating their own stylized models.
Yi-Jheng Huang, Wen-Chieh Lin, I-Cheng Yeh 0001, Tong-Yee Lee
IEEE Trans. Vis. Comput. Graph.4
2018 Generation of Escher Arts with Dual Perception
abstract
Escher transmutation is a graphic art that smoothly transforms one tile pattern into another tile pattern with dual perception. A classic example is the artwork called Sky and Water, in which a compelling figure-ground arrangement is applied to portray the transmutation of a bird in sky and a fish in water. The shape of a bird is progressively deformed and dissolves into the background while the background gradually reveals the shape of a fish. This paper introduces a system to create a variety of Escher-like transmutations, which includes the algorithms for initializing a tile pattern with dual figure-ground arrangement, for searching for the best matched shape of a user-specified motif from a database, and for transforming the content and shapes of tile patterns using a content-aware warping technique. The proposed system, integrating the graphic techniques of tile initialization, shape matching, and shape warping, allows users to create various Escher-like transmutations with minimal user interaction. Experimental results and conducted user studies demonstrate the feasibility and flexibility of the proposed system in Escher art generation.
Shih-Syun Lin, Charles C. Morace, Chao-Hung Lin, Li-Fong Hsu, Tong-Yee Lee
IEEE Trans. Vis. Comput. Graph.5
2017 Generating Ambiguous Figure-Ground Images
abstract
Ambiguous figure-ground images, mostly represented as binary images, are fascinating as they present viewers a visual phenomena of perceiving multiple interpretations from a single image. In one possible interpretation, the white region is seen as a foreground figure while the black region is treated as shapeless background. Such perception can reverse instantly at any moment. In this paper, we investigate the theory behind this ambiguous perception and present an automatic algorithm to generate such images. We model the problem as a binary image composition using two object contours and approach it through a three-stage pipeline. The algorithm first performs a partial shape matching to find a good partial contour matching between objects. This matching is based on a content-aware shape matching metric, which captures features of ambiguous figure-ground images. Then we combine matched contours into a compound contour using an adaptive contour deformation, followed by computing an optimal cropping window and image binarization for the compound contour that maximize the completeness of object contours in the final composition. We have tested our system using a wide range of input objects and generated a large number of convincing examples with or without user guidance. The efficiency of our system and quality of results are verified through an extensive experimental study.
Ying-Miao Kuo, Hung-Kuo Chu, Ming-Te Chi, Ruen-Rone Lee, Tong-Yee Lee
IEEE Trans. Vis. Comput. Graph.5
2017 Interactive High-Relief Reconstruction for Organic and Double-Sided Objects from a Photo
abstract
We introduce an interactive user-driven method to reconstruct high-relief 3D geometry from a single photo. Particularly, we consider two novel but challenging reconstruction issues: i) common non-rigid objects whose shapes are organic rather than polyhedral/symmetric, and ii) double-sided structures, where front and back sides of some curvy object parts are revealed simultaneously on image. To address these issues, we develop a three-stage computational pipeline. First, we construct a 2.5D model from the input image by user-driven segmentation, automatic layering, and region completion, handling three common types of occlusion. Second, users can interactively mark-up slope and curvature cues on the image to guide our constrained optimization model to inflate and lift up the image layers. We provide real-time preview of the inflated geometry to allow interactive editing. Third, we stitch and optimize the inflated layers to produce a high-relief 3D model. Compared to previous work, we can generate high-relief geometry with large viewing angles, handle complex organic objects with multiple occluded regions and varying shape profiles, and reconstruct objects with double-sided structures. Lastly, we demonstrate the applicability of our method on a wide variety of input images with human, animals, flowers, etc.
Chih-Kuo Yeh, Shi-Yang Huang, Pradeep Kumar Jayaraman, Chi-Wing Fu, Tong-Yee Lee
IEEE Trans. Vis. Comput. Graph.5
2016 Consistent Volumetric Warping Using Floating Boundaries for Stereoscopic Video Retargeting
abstract
The key to content-aware warping and cropping is adapting data to fit displays with various aspect ratios while preserving visually salient contents. Most previous studies achieve this objective by cropping insignificant contents near frame boundaries and consistently resizing frames through an optimization technique with various preservation constraints and fixed boundary conditions. These strategies significantly improve retargeting quality. However, warping under fixed boundary conditions may bound/limit the preservation of visually salient contents. Moreover, dynamic frame cropping and frame alignment may result in unnatural object/camera motions. In this paper, a floating boundary with volumetric warping and object-aware cropping is proposed to address these problems. In the proposed scheme, visually salient objects in the space-time domain are deformed as rigidly and as consistently as possible using information from matched objects and content-aware boundary constraints. The content-aware boundary constraints can retain visually salient contents in a fixed region with a desired resolution and aspect ratio, called critical region, during warping. Volumetric cropping with the fixed critical region is then performed to adjust stereoscopic videos to the desired aspect ratios. The strategies of warping and cropping using floating boundaries and spatiotemporal constraints enable our method to consistently preserve the temporal motions and spatial shapes of visually salient volumetric objects in the left and right videos as much as possible, thus leading to good content-aware retargeting. In addition, by considering shape, motion, and disparity preservation, the proposed scheme can be applied to various media, including images, stereoscopic images, videos, and stereoscopic videos. Qualitative and quantitative analyses of stereoscopic videos with diverse camera and considerable motions demonstrate a clear superiority of the proposed method over related methods in terms of retargeting quality.
Shih-Syun Lin, Chao-Hung Lin, Yu-Hsuan Kuo, Tong-Yee Lee
IEEE Trans. Circuits Syst. Video Technol.4
2016 Image Retargeting by Texture-Aware Synthesis
abstract
Real-world images usually contain vivid contents and rich textural details, which will complicate the manipulation on them. In this paper, we design a new framework based on exampled-based texture synthesis to enhance content-aware image retargeting. By detecting the textural regions in an image, the textural image content can be synthesized rather than simply distorted or cropped. This method enables the manipulation of textural & non-textural regions with different strategies since they have different natures. We propose to retarget the textural regions by example-based synthesis and non-textural regions by fast multi-operator. To achieve practical retargeting applications for general images, we develop an automatic and fast texture detection method that can detect multiple disjoint textural regions. We adjust the saliency of the image according to the features of the textural regions. To validate the proposed method, comparisons with state-of-the-art image retargeting techniques and a user study were conducted. Convincing visual results are shown to demonstrate the effectiveness of the proposed method.
Weiming Dong, Fuzhang Wu, Yan Kong, Xing Mei, Tong-Yee Lee, Xiaopeng Zhang 0001
IEEE Trans. Vis. Comput. Graph.5
2016 Measuring and Predicting Visual Importance of Similar Objects
abstract
Similar objects are ubiquitous and abundant in both natural and artificial scenes. Determining the visual importance of several similar objects in a complex photograph is a challenge for image understanding algorithms. This study aims to define the importance of similar objects in an image and to develop a method that can select the most important instances for an input image from multiple similar objects. This task is challenging because multiple objects must be compared without adequate semantic information. This challenge is addressed by building an image database and designing an interactive system to measure object importance from human observers. This ground truth is used to define a range of features related to the visual importance of similar objects. Then, these features are used in learning-to-rank and random forest to rank similar objects in an image. Importance predictions were validated on 5,922 objects. The most important objects can be identified automatically. The factors related to composition (e.g., size, location, and overlap) are particularly informative, although clarity and color contrast are also important. We demonstrate the usefulness of similar object importance on various applications, including image retargeting, image compression, image re-attentionizing, image admixture, and manipulation of blindness images.
Yan Kong, Weiming Dong, Xing Mei, Chongyang Ma, Tong-Yee Lee, Siwei Lyu, Feiyue Huang, Xiaopeng Zhang 0001
IEEE Trans. Vis. Comput. Graph.5
2015 Efficient QR Code Beautification With High Quality Visual Content
abstract
Quick response (QR) code is generally used for embedding messages such that people can conveniently use mobile devices to capture the QR code and acquire information through a QR code reader. In the past, the design of QR code generators only aimed to achieve high decodability and the produced QR codes usually look like random black-and-white patterns without visual semantics. In recent years, researchers have been tried to endow the QR code with aesthetic elements and QR code beautification has been formulated as an optimization problem that minimizes the visual perception distortion subject to acceptable decoding rate. However, the visual quality of the QR code generated by existing methods still leaves much to be desired. In this work, we propose a two-stage approach to generate QR code with high quality visual content. In the first stage, a baseline QR code with reliable decodability but poor visual quality is first synthesized based on the Gauss-Jordan elimination procedure. In the second stage, a rendering mechanism is designed to improve the visual quality while avoiding affecting the decodability of the QR code. The experimental results show that the proposed method substantially enhances the appearance of the QR code and the processing complexity is near real-time.
Shih-Syun Lin, Min-Chun Hu 0001, Chien-Han Lee, Tong-Yee Lee
IEEE Trans. Multim.4
2015 Morphable Word Clouds for Time-Varying Text Data Visualization
abstract
A word cloud is a visual representation of a collection of text documents that uses various font sizes, colors, and spaces to arrange and depict significant words. The majority of previous studies on time-varying word clouds focuses on layout optimization and temporal trend visualization. However, they do not fully consider the spatial shapes and temporal motions of word clouds, which are important factors for attracting people's attention and are also important cues for human visual systems in capturing information from time-varying text data. This paper presents a novel method that uses rigid body dynamics to arrange multi-temporal word-tags in a specific shape sequence under various constraints. Each word-tag is regarded as a rigid body in dynamics. With the aid of geometric, aesthetic, and temporal coherence constraints, the proposed method can generate a temporally morphable word cloud that not only arranges word-tags in their corresponding shapes but also smoothly transforms the shapes of word clouds over time, thus yielding a pleasing time-varying visualization. Using the proposed frame-by-frame and morphable word clouds, people can observe the overall story of a time-varying text data from the shape transition, and people can also observe the details from the word clouds in frames. Experimental results on various data demonstrate the feasibility and flexibility of the proposed method in morphable word cloud generation. In addition, an application that uses the proposed word clouds in a simulated exhibition demonstrates the usefulness of the proposed method.
Ming-Te Chi, Shih-Syun Lin, Shiang-Yi Chen, Chao-Hung Lin, Tong-Yee Lee
IEEE Trans. Vis. Comput. Graph.5
2015 2.5D Cartoon Hair Modeling and Manipulation
abstract
This paper addresses a challenging single-view modeling and animation problem with cartoon images. Our goal is to model the hairs in a given cartoon image with consistent layering and occlusion, so that we can produce various visual effects from just a single image. We propose a novel 2.5D modeling approach to deal with this problem. Given an input image, we first segment the hairs of the cartoon character into regions of hair strands. Then, we apply our novel layering metric, which is derived from the Gestalt psychology, to automatically optimize the depth ordering among the hair strands. After that, we employ our hair completion method to fill the occluded part of each hair strand, and create a 2.5D model of the cartoon hair. By using this model, we can produce various visual effects, e.g., we develop a simplified fluid simulation model to produce wind blowing animations with the 2.5D hairs. To further demonstrate the applicability and versatility of our method, we compare our results with real cartoon hair animations, and also apply our model to produce a wide variety of hair manipulation effects, including hair editing and hair braiding.
Chih-Kuo Yeh, Pradeep Kumar Jayaraman, Xiaopei Liu, Chi-Wing Fu, Tong-Yee Lee
IEEE Trans. Vis. Comput. Graph.5
2015 Erratum to: Optical illusion shape texturing using repeated asymmetric patterns
Ming-Te Chi, Chih-Yuan Yao, Eugene Zhang, Tong-Yee Lee
Vis. Comput.4
2014 Foldover-free shape deformation for biomedicine
Hongchuan Yu, Jian J. Zhang 0001, Tong-Yee Lee
J. Biomed. Informatics3
2014 Object-Coherence Warping for Stereoscopic Image Retargeting
abstract
This paper addresses the topic of content-aware stereoscopic image retargeting. The key to this topic is consistently adapting a stereoscopic image to fit displays with various aspect ratios and sizes while preserving visually salient content. Most methods focus on preserving the disparities and shapes of visually salient objects through nonlinear image warping, in which distortions caused by warping are propagated to homogenous and low-significance regions. However, disregarding the consistency of object deformation sometimes results in apparent distortions in both the disparities and shapes of objects. An object-coherence warping scheme is proposed to reduce this unwanted distortion. The basic idea is to utilize the information of matched objects rather than that of matched pixels in warping. Such information implies object correspondences in a stereoscopic image pair, which allows the generation of an object significance map and the consistent preservation of objects. This strategy enables our method to consistently preserve both the disparities and shapes of visually salient objects, leading to good content-aware retargeting. In the experiments, qualitative and quantitative analyses of various stereoscopic images show that our results are better than those generated by related methods in terms of consistency of object preservation.
Shih-Syun Lin, Chao-Hung Lin, Shu-Huai Chang, Tong-Yee Lee
IEEE Trans. Circuits Syst. Video Technol.4
2014 Summarization-Based Image Resizing by Intelligent Object Carving
abstract
Image resizing can be more effectively achieved with a better understanding of image semantics. In this paper, similar patterns that exist in many real-world images are analyzed. By interactively detecting similar objects in an image, the image content can be summarized rather than simply distorted or cropped. This method enables the manipulation of image pixels or patches as well as semantic objects in the scene during image resizing process. Given the special nature of similar objects in a general image, the integration of a novel object carving (OC) operator with the multi-operator framework is proposed for summarizing similar objects. The object removal sequence in the summarization strategy directly affects resizing quality. The method by which to evaluate the visual importance of the object as well as to optimally select the candidates for object carving is demonstrated. To achieve practical resizing applications for general images, a template matching-based method is developed. This method can detect similar objects even when they are of various colors, transformed in terms of perspective, or partially occluded. To validate the proposed method, comparisons with state-of-the-art resizing techniques and a user study were conducted. Convincing visual results are shown to demonstrate the effectiveness of the proposed method.
Weiming Dong, Tong-Yee Lee, Fuzhang Wu, Yan Kong, Xiaopeng Zhang 0001
IEEE Trans. Vis. Comput. Graph.3
2014 Drawing Road Networks with Mental Maps
abstract
Tourist and destination maps are thematic maps designed to represent specific themes in maps. The road network topologies in these maps are generally more important than the geometric accuracy of roads. A road network warping method is proposed to facilitate map generation and improve theme representation in maps. The basic idea is deforming a road network to meet a user-specified mental map while an optimization process is performed to propagate distortions originating from road network warping. To generate a map, the proposed method includes algorithms for estimating road significance and for deforming a road network according to various geometric and aesthetic constraints. The proposed method can produce an iconic mark of a theme from a road network and meet a user-specified mental map. Therefore, the resulting map can serve as a tourist or destination map that not only provides visual aids for route planning and navigation tasks, but also visually emphasizes the presentation of a theme in a map for the purpose of advertising. In the experiments, the demonstrations of map generations show that our method enables map generation systems to generate deformed tourist and destination maps efficiently.
Shih-Syun Lin, Chao-Hung Lin, Yan-Jhang Hu, Tong-Yee Lee
IEEE Trans. Vis. Comput. Graph.4
2014 Optical illusion shape texturing using repeated asymmetric patterns
Ming-Te Chi, Chih-Yuan Yao, Eugene Zhang, Tong-Yee Lee
Vis. Comput.4
2013 Illusory Motions on Surfaces
abstract
Illusory motions refer to the phenomena in which static images composed of certain colors and patterns lead to the illusion of motions. This paper presents a first approach to generating illusory motions on 3D surfaces which can be used for shape illustration as well as artistic visualization of line fields on surfaces. Our method extends previous work on generating illusory motions in the plane, which we adapt to 3D surfaces. In addition, we propose novel Repeated Asymmetric Patterns (RAPs) to visualize bidirectional flows, thus enabling the visualization of line fields in the plane and on surfaces. We demonstrate the effectiveness of our method with applications in shape illustration as well as line field visualization on surfaces.
Ming-Te Chi, Chih-Yuan Yao, Tong-Yee Lee, Eugene Zhang
CAD/Graphics3
2013 Patch-Based Image Warping for Content-Aware Retargeting
abstract
Image retargeting is the process of adapting images to fit displays with various aspect ratios and sizes. Most studies on image retargeting focus on shape preservation, but they do not fully consider the preservation of structure lines, which are sensitive to human visual system. In this paper, a patch-based retargeting scheme with an extended significance measurement is introduced to preserve shapes of both visually salient objects and structure lines while minimizing visual distortions. In the proposed scheme, a similarity transformation constraint is used to force visually salient content to undergo as-rigid-as-possible deformation, while an optimization process is performed to smoothly propagate distortions. These processes enable our approach to yield pleasing content-aware warping and retargeting. Experimental results and a user study show that our results are better than those generated by state-of-the-art approaches.
Shih-Syun Lin, I-Cheng Yeh 0001, Chao-Hung Lin, Tong-Yee Lee
IEEE Trans. Multim.4
2013 Content-Aware Video Retargeting Using Object-Preserving Warping
abstract
A novel content-aware warping approach is introduced for video retargeting. The key to this technique is adapting videos to fit displays with various aspect ratios and sizes while preserving both visually salient content and temporal coherence. Most previous studies solve this spatiotemporal problem by consistently resizing content in frames. This strategy significantly improves the retargeting results, but does not fully consider object preservation, sometimes causing apparent distortions on visually salient objects. We propose an object-preserving warping scheme with object-based significance estimation to reduce this unpleasant distortion. In the proposed scheme, visually salient objects in 3D space-time space are forced to undergo as-rigid-as-possible warping, while low-significance contents are warped as close as possible to linear rescaling. These strategies enable our method to consistently preserve both the spatial shapes and temporal motions of visually salient objects and avoid overdeformations on low-significance objects, yielding a pleasing motion-aware video retargeting. Qualitative and quantitative analyses, including a user study and experiments on complex videos containing diverse cameras and dynamic motions, show a clear superiority of our method over related video retargeting methods.
Shih-Syun Lin, Chao-Hung Lin, I-Cheng Yeh 0001, Shu-Huai Chang, Chih-Kuo Yeh, Tong-Yee Lee
IEEE Trans. Vis. Comput. Graph.6
2013 Spatially and Temporally Optimized Video Stabilization
abstract
Properly handling parallax is important for video stabilization. Existing methods that achieve the aim require either 3D reconstruction or long feature trajectories to enforce the subspace or epipolar geometry constraints. In this paper, we present a robust and efficient technique that works on general videos. It achieves high-quality camera motion on videos where 3D reconstruction is difficult or long feature trajectories are not available. We represent each trajectory as a Bézier curve and maintain the spatial relations between trajectories by preserving the original offsets of neighboring curves. Our technique formulates stabilization as a spatial-temporal optimization problem that finds smooth feature trajectories and avoids visual distortion. The Bézier representation enables strong smoothness of each feature trajectory and reduces the number of variables in the optimization problem. We also stabilize videos in a streaming fashion to achieve scalability. The experiments show that our technique achieves high-quality camera motion on a variety of challenging videos that are difficult for existing methods.
Yu-Shuen Wang, Feng Liu 0015, Pu-Sheng Hsu, Tong-Yee Lee
IEEE Trans. Vis. Comput. Graph.4
2013 Double-Sided 2.5D Graphics
abstract
This paper introduces double-sided 2.5D graphics, aiming at enriching the visual appearance when manipulating conventional 2D graphical objects in 2.5D worlds. By attaching a back texture image on a single-sided 2D graphical object, we can enrich the surface and texture detail on 2D graphical objects and improve our visual experience when manipulating and animating them. A family of novel operations on 2.5D graphics, including rolling, twisting, and folding, are proposed in this work, allowing users to efficiently create compelling 2.5D visual effects. Very little effort is needed from the user's side. In our experiment, various creative designs on double-sided graphics were worked out by the recruited participants including a professional artist, which show and demonstrate the feasibility and applicability of our proposed method.
Chih-Kuo Yeh, Peng Song 0001, Peng-Yen Lin, Chi-Wing Fu, Chao-Hung Lin, Tong-Yee Lee
IEEE Trans. Vis. Comput. Graph.6
2012 Human Motion Retrieval from Hand-Drawn Sketch
abstract
The rapid growth of motion capture data increases the importance of motion retrieval. The majority of the existing motion retrieval approaches are based on a labor-intensive step in which the user browses and selects a desired query motion clip from the large motion clip database. In this work, a novel sketching interface for defining the query is presented. This simple approach allows users to define the required motion by sketching several motion strokes over a drawn character, which requires less effort and extends the users’ expressiveness. To support the real-time interface, a specialized encoding of the motions and the hand-drawn query is required. Here, we introduce a novel hierarchical encoding scheme based on a set of orthonormal spherical harmonic (SH) basis functions, which provides a compact representation, and avoids the CPU/processing intensive stage of temporal alignment used by previous solutions. Experimental results show that the proposed approach can well retrieve the motions, and is capable of retrieve logically and numerically similar motions, which is superior to previous approaches. The user study shows that the proposed system can be a useful tool to input motion query if the users are familiar with it. Finally, an application of generating a 3D animation from a hand-drawn comics strip is demonstrated.
Min-Wen Chao, Chao-Hung Lin, Jackie Assa, Tong-Yee Lee
IEEE Trans. Vis. Comput. Graph.4
2012 Coherent Time-Varying Graph Drawing with Multifocus+Context Interaction
abstract
We present a new approach for time-varying graph drawing that achieves both spatiotemporal coherence and multifocus+context visualization in a single framework. Our approach utilizes existing graph layout algorithms to produce the initial graph layout, and formulates the problem of generating coherent time-varying graph visualization with the focus+context capability as a specially tailored deformation optimization problem. We adopt the concept of the super graph to maintain spatiotemporal coherence and further balance the needs for aesthetic quality and dynamic stability when interacting with time-varying graphs through focus+context visualization. Our method is particularly useful for multifocus+context visualization of time-varying graphs where we can preserve the mental map by preventing nodes in the focus from undergoing abrupt changes in size and location in the time sequence. Experiments demonstrate that our method strikes a good balance between maintaining spatiotemporal coherence and accentuating visual foci, thus providing a more engaging viewing experience for the users.
Kun-Chuan Feng, Chaoli Wang 0001, Han-Wei Shen, Tong-Yee Lee
IEEE Trans. Vis. Comput. Graph.4
2012 Region-Based Line Field Design Using Harmonic Functions
abstract
Field design has wide applications in graphics and visualization. One of the main challenges in field design has been how to provide users with both intuitive control over the directions in the field on one hand and robust management of its topology on the other hand. In this paper, we present a design paradigm for line fields that addresses this challenge. Rather than asking users to input all singularities as in most methods that offer topology control, we let the user provide a partitioning of the domain and specify simple flow patterns within the partitions. Represented by a selected set of harmonic functions, the elementary fields within the partitions are then combined to form continuous fields with rich appearances and well-determined topology. Our method allows a user to conveniently design the flow patterns while having precise and robust control over the topological structure. Based on the method, we developed an interactive tool for designing line fields from images, and demonstrated the utility of the fields in image stylization.
Chih-Yuan Yao, Ming-Te Chi, Tong-Yee Lee, Tao Ju 0001
IEEE Trans. Vis. Comput. Graph.3
2012 Social-Event-Driven Camera Control for Multicharacter Animations
abstract
In a virtual world, a group of virtual characters can interact with each other, and these characters may leave a group to join another. The interaction among individuals and groups often produces interesting events in a sequence of animation. The goal of this paper is to discover social events involving mutual interactions or group activities in multicharacter animations and automatically plan a smooth camera motion to view interesting events suggested by our system or relevant events specified by a user. Inspired by sociology studies, we borrow the knowledge in Proxemics, social force, and social network analysis to model the dynamic relation among social events and the relation among the participants within each event. By analyzing the variation of relation strength among participants and spatiotemporal correlation among events, we discover salient social events in a motion clip and generate an overview video of these events with smooth camera motion using a simulated annealing optimization method. We tested our approach on different motions performed by multiple characters. Our user study shows that our results are preferred in 66.19 percent of the comparisons with those by the camera control approach without event analysis and are comparable (51.79 percent) to professional results by an artist.
I-Cheng Yeh 0001, Wen-Chieh Lin, Tong-Yee Lee, Hsin-Ju Han, Jehee Lee, Manmyung Kim
IEEE Trans. Vis. Comput. Graph.3
2012 An RBF-Based Reparameterization Method for Constrained Texture Mapping
abstract
Texture mapping has long been used in computer graphics to enhance the realism of virtual scenes. However, to match the 3D model feature points with the corresponding pixels in a texture image, surface parameterization must satisfy specific positional constraints. However, despite numerous research efforts, the construction of a mathematically robust, foldover-free parameterization that is subject to positional constraints continues to be a challenge. In the present paper, this foldover problem is addressed by developing radial basis function (RBF)-based reparameterization. Given initial 2D embedding of a 3D surface, the proposed method can reparameterize 2D embedding into a foldover-free 2D mesh, satisfying a set of user-specified constraint points. In addition, this approach is mesh free. Therefore, generating smooth texture mapping results is possible without extra smoothing optimization.
Hongchuan Yu, Tong-Yee Lee, I-Cheng Yeh 0001, Xiaosong Yang, Wenxi Li, Jian J. Zhang 0001
IEEE Trans. Vis. Comput. Graph.2
2011 A graph-based shape matching scheme for 3D articulated objects
abstract
Abstract In this paper, a novel graph‐based shape matching scheme for three‐dimensional articulated objects is introduced. The underlying graph structure of a given 3D model is composed of its topological skeleton and local geometric features. Matching two graph structures is generally an NP‐hard combinatorial optimization problem. To reduce computation cost, two graphs are embedded on a high‐dimensional space, and then matched based on an extension of Earth Mover's Distance (EMD). Furthermore, the symmetric components of an articulated object are determined by a voting algorithm with a self‐matching strategy to refine the matching correspondences. Experimental results show that the proposed approach is robust, even when the models are under the surface disturbances of noise addition, smoothing, simplification, similarity transformation, and pose deformation. In addition, the proposed approach is capable of handling both global and partial shape matching. Copyright © 2011 John Wiley & Sons, Ltd.
Min-Wen Chao, Chao-Hung Lin, Chih-Chieh Chang, Tong-Yee Lee
Comput. Animat. Virtual Worlds4
2011 Efficient camera path planning algorithm for human motion overview
abstract
Abstract Camera path planning for character motions is a fundamental and important research topic, benefiting many animation applications. Existing optimal‐based approaches are generally computationally expensive and infeasible for interactive applications. In this paper, we propose an efficient approach that can take many constraints of finding the camera path into account and can potentially enable interactive camera control. Instead of solving a highly complicated camera optimization problem in a spatiotemporal four‐dimensional space, we heuristically determine the camera path based on an efficient greedy‐based tree traversal approach. The experimental results show that the proposed approach can efficiently generate a smooth, informative, and aesthetic camera path that can reveal the significant features of character motions. Moreover, the conducted user study also shows that the generated camera paths are comparable to those of a state‐of‐the‐art approach and those made by professional animators. Copyright © 2011 John Wiley & Sons, Ltd.
I-Cheng Yeh 0001, Chao-Hung Lin, Hung-Jen Chien, Tong-Yee Lee
Comput. Animat. Virtual Worlds4
2011 Scalable and coherent video resizing with per-frame optimization
abstract
The key to high-quality video resizing is preserving the shape and motion of visually salient objects while remaining temporally-coherent. These spatial and temporal requirements are difficult to reconcile, typically leading existing video retargeting methods to sacrifice one of them and causing distortion or waving artifacts. Recent work enforces temporal coherence of content-aware video warping by solving a global optimization problem over the entire video cube. This significantly improves the results but does not scale well with the resolution and length of the input video and quickly becomes intractable. We propose a new method that solves the scalability problem without compromising the resizing quality. Our method factors the problem into spatial and time/motion components: we first resize each frame independently to preserve the shape of salient regions, and then we optimize their motion using a reduced model for each pathline of the optical flow. This factorization decomposes the optimization of the video cube into sets of sub-problems whose size is proportional to a single frame's resolution and which can be solved in parallel. We also show how to incorporate cropping into our optimization, which is useful for scenes with numerous salient objects where warping alone would degenerate to linear scaling. Our results match the quality of state-of-the-art retargeting methods while dramatically reducing the computation time and memory consumption, making content-aware video resizing scalable and practical.
Yu-Shuen Wang, Jen-Hung Hsiao, Olga Sorkine-Hornung, Tong-Yee Lee
ACM Trans. Graph.4
2011 Feature-Preserving Volume Data Reduction and Focus+Context Visualization
abstract
The growing sizes of volumetric data sets pose a great challenge for interactive visualization. In this paper, we present a feature-preserving data reduction and focus+context visualization method based on transfer function driven, continuous voxel repositioning and resampling techniques. Rendering reduced data can enhance interactivity. Focus+context visualization can show details of selected features in context on display devices with limited resolution. Our method utilizes the input transfer function to assign importance values to regularly partitioned regions of the volume data. According to user interaction, it can then magnify regions corresponding to the features of interest while compressing the rest by deforming the 3D mesh. The level of data reduction achieved is significant enough to improve overall efficiency. By using continuous deformation, our method avoids the need to smooth the transition between low and high-resolution regions as often required by multiresolution methods. Furthermore, it is particularly attractive for focus+context visualization of multiple features. We demonstrate the effectiveness and efficiency of our method with several volume data sets from medical applications and scientific simulations.
Yu-Shuen Wang, Chaoli Wang 0001, Tong-Yee Lee, Kwan-Liu Ma
IEEE Trans. Vis. Comput. Graph.3
2011 Template-Based 3D Model Fitting Using Dual-Domain Relaxation
abstract
We introduce a template fitting method for 3D surface meshes. A given template mesh is deformed to closely approximate the input 3D geometry. The connectivity of the deformed template model is automatically adjusted to facilitate the geometric fitting and to ascertain high quality of the mesh elements. The template fitting process utilizes a specially tailored Laplacian processing framework, where in the first, coarse fitting stage we approximate the input geometry with a linearized biharmonic surface (a variant of LS-mesh), and then the fine geometric detail is fitted further using iterative Laplacian editing with reliable correspondence constraints and a local surface flattening mechanism to avoid foldovers. The latter step is performed in the dual mesh domain, which is shown to encourage near-equilateral mesh elements and significantly reduces the occurrence of triangle foldovers, a well-known problem in mesh fitting. To experimentally evaluate our approach, we compare our method with relevant state-of-the-art techniques and confirm significant improvements of results. In addition, we demonstrate the usefulness of our approach to the application of consistent surface parameterization (also known as cross-parameterization).
I-Cheng Yeh 0001, Chao-Hung Lin, Olga Sorkine-Hornung, Tong-Yee Lee
IEEE Trans. Vis. Comput. Graph.4
2010 Camouflage images
abstract
Camouflage images contain one or more hidden figures that remain imperceptible or unnoticed for a while. In one possible explanation, the ability to delay the perception of the hidden figures is attributed to the theory that human perception works in two main phases: feature search and conjunction search. Effective camouflage images make feature based recognition difficult, and thus force the recognition process to employ conjunction search, which takes considerable effort and time. In this paper, we present a technique for creating camouflage images. To foil the feature search, we remove the original subtle texture details of the hidden figures and replace them by that of the surrounding apparent image. To leave an appropriate degree of clues for the conjunction search, we compute and assign new tones to regions in the embedded figures by performing an optimization between two conflicting terms, which we call immersion and standout , corresponding to hiding and leaving clues, respectively. We show a large number of camouflage images generated by our technique, with or without user guidance. We have tested the quality of the images in an extensive user study, showing a good control of the difficulty levels.
Hung-Kuo Chu, Wei-Hsin Hsu, Niloy J. Mitra, Daniel Cohen-Or, Tien-Tsin Wong, Tong-Yee Lee
ACM Trans. Graph.6
2010 Motion-based video retargeting with optimized crop-and-warp
abstract
We introduce a video retargeting method that achieves high-quality resizing to arbitrary aspect ratios for complex videos containing diverse camera and dynamic motions. Previous content-aware retargeting methods mostly concentrated on spatial considerations, attempting to preserve the shape of salient objects in each frame by removing or distorting homogeneous background content. However, sacrificeable space is fundamentally limited in video, since object motion makes foreground and background regions correlated, causing waving and squeezing artifacts. We solve the retargeting problem by explicitly employing motion information and by distributing distortion in both spatial and temporal dimensions. We combine novel cropping and warping operators, where the cropping removes temporally-recurring contents and the warping utilizes available homogeneous regions to mask deformations while preserving motion. Variational optimization allows to find the best balance between the two operations, enabling retargeting of challenging videos with complex motions, numerous prominent objects and arbitrary depth variability. Our method compares favorably with state-of-the-art retargeting systems, as demonstrated in the examples and widely supported by the conducted user study.
Yu-Shuen Wang, Hui-Chih Lin, Olga Sorkine-Hornung, Tong-Yee Lee
ACM Trans. Graph.4
2010 Resizing by symmetry-summarization
abstract
Image resizing can be achieved more effectively if we have a better understanding of the image semantics. In this paper, we analyze the translational symmetry , which exists in many real-world images. By detecting the symmetric lattice in an image, we can summarize , instead of only distorting or cropping, the image content. This opens a new space for image resizing that allows us to manipulate, not only image pixels, but also the semantic cells in the lattice. As a general image contains both symmetry & non-symmetry regions and their natures are different, we propose to resize symmetry regions by summarization and non-symmetry region by warping. The difference in resizing strategy induces discontinuity at their shared boundary. We demonstrate how to reduce the artifact. To achieve practical resizing applications for general images, we developed a fast symmetry detection method that can detect multiple disjoint symmetry regions, even when the lattices are curved and perspectively viewed. Comparisons to state-of-the-art resizing techniques and a user study were conducted to validate the proposed method. Convincing visual results are shown to demonstrate its effectiveness.
Huisi Wu, Yu-Shuen Wang, Kun-Chuan Feng, Tien-Tsin Wong, Tong-Yee Lee, Pheng-Ann Heng
ACM Trans. Graph.5
2010 Real-Time Physics-Based 3D Biped Character Animation Using an Inverted Pendulum Model
abstract
We present a physics-based approach to generate 3D biped character animation that can react to dynamical environments in real time. Our approach utilizes an inverted pendulum model to online adjust the desired motion trajectory from the input motion capture data. This online adjustment produces a physically plausible motion trajectory adapted to dynamic environments, which is then used as the desired motion for the motion controllers to track in dynamics simulation. Rather than using Proportional-Derivative controllers whose parameters usually cannot be easily set, our motion tracking adopts a velocity-driven method which computes joint torques based on the desired joint angular velocities. Physically correct full-body motion of the 3D character is computed in dynamics simulation using the computed torques and dynamical model of the character. Our experiments demonstrate that tracking motion capture data with real-time response animation can be achieved easily. In addition, physically plausible motion style editing, automatic motion transition, and motion adaptation to different limb sizes can also be generated without difficulty.
Yao-Yang Tsai, Wen-Chieh Lin, Kuangyou B. Cheng, Jehee Lee, Tong-Yee Lee
IEEE Trans. Vis. Comput. Graph.5
2010 A novel semi-blind-and-semi-reversible robust watermarking scheme for 3D polygonal models
Chao-Hung Lin, Min-Wen Chao, Chan-Yu Liang, Tong-Yee Lee
Vis. Comput.4
2009 Compatible quadrangulation by sketching
abstract
Abstract Mesh quadrangulation has received increasing attention in the past decade. While previous works have mostly focused on producing a high quality quad mesh of a single model, the connectivity of the quadrangulation is typically difficult to control and varies among models even with similar shapes. In this paper, we propose a novel interactive framework for quadrangulating a set of models collectively with compatible connectivity. Furthermore, we demonstrate its application to 3D mesh morphing. In our approach, the user interactively sketches a skeleton within each model, and our method automatically computes compatible base domains for all models from these skeletons, on which the models are parameterized. With this novel parameterization, it is very easy to generate a pleasing and smooth 3D morphing sequence among these compatible models. The method yields quadrangulation with comparable quality to existing approaches, but greatly simplifies compatible re‐meshing among a group of topologically equivalent models, in particular characters and animals models, with direct applications in shape blending and morphing. Copyright © 2009 John Wiley & Sons, Ltd.
Chih-Yuan Yao, Hung-Kuo Chu, Tong-Yee Lee
Comput. Animat. Virtual Worlds4
2009 Emerging images
abstract
Emergence refers to the unique human ability to aggregate information from seemingly meaningless pieces, and to perceive a whole that is meaningful. This special skill of humans can constitute an effective scheme to tell humans and machines apart. This paper presents a synthesis technique to generate images of 3D objects that are detectable by humans, but difficult for an automatic algorithm to recognize. The technique allows generating an infinite number of images with emerging figures. Our algorithm is designed so that locally the synthesized images divulge little useful information or cues to assist any segmentation or recognition procedure. Therefore, as we demonstrate, computer vision algorithms are incapable of effectively processing such images. However, when a human observer is presented with an emergence image, synthesized using an object she is familiar with, the figure emerges when observed as a whole. We can control the difficulty level of perceiving the emergence effect through a limited set of parameters. A procedure that synthesizes emergence images can be an effective tool for exploring and understanding the factors affecting computer vision techniques.
Niloy J. Mitra, Hung-Kuo Chu, Tong-Yee Lee, Lior Wolf, Yehezkel Yeshurun, Daniel Cohen-Or
ACM Trans. Graph.3
2009 Motion-aware temporal coherence for video resizing
abstract
Temporal coherence is crucial in content-aware video retargeting. To date, this problem has been addressed by constraining temporally adjacent pixels to be transformed coherently. However, due to the motion-oblivious nature of this simple constraint, the retargeted videos often exhibit flickering or waving artifacts, especially when significant camera or object motions are involved. Since the feature correspondence across frames varies spatially with both camera and object motion, motion-aware treatment of features is required for video resizing. This motivated us to align consecutive frames by estimating interframe camera motion and to constrain relative positions in the aligned frames. To preserve object motion, we detect distinct moving areas of objects across multiple frames and constrain each of them to be resized consistently. We build a complete video resizing framework by incorporating our motion-aware constraints with an adaptation of the scale-and-stretch optimization recently proposed by Wang and colleagues. Our streaming implementation of the framework allows efficient resizing of long video sequences with low memory cost. Experiments demonstrate that our method produces spatiotemporally coherent retargeting results even for challenging examples with complex camera and object motion, which are difficult to handle with previous techniques.
Yu-Shuen Wang, Hongbo Fu 0001, Olga Sorkine-Hornung, Tong-Yee Lee, Hans-Peter Seidel
ACM Trans. Graph.4
2009 A High Capacity 3D Steganography Algorithm
abstract
In this paper, we present a very high-capacity and low-distortion 3D steganography scheme. Our steganography approach is based on a novel multilayered embedding scheme to hide secret messages in the vertices of 3D polygon models. Experimental results show that the cover model distortion is very small as the number of hiding layers ranges from 7 to 13 layers. To the best of our knowledge, this novel approach can provide much higher hiding capacity than other state-of-the-art approaches, while obeying the low distortion and security basic requirements for steganography on 3D models.
Min-Wen Chao, Chao-Hung Lin, Cheng-Wei Yu, Tong-Yee Lee
IEEE Trans. Vis. Comput. Graph.4
2009 Multiresolution Mean Shift Clustering Algorithm for Shape Interpolation
abstract
In this paper, we solve the problem of 3D shape interpolation with significant pose variation. For an ideal 3D shape interpolation, especially the articulated model, the shape should follow the movement of the underlying articulated structure and be transformed in a way that is as rigid as possible. Given input shapes with compatible connectivity, we propose a novel multiresolution mean shift (MMS) clustering algorithm to automatically extract their near-rigid components. Then, by building the hierarchical relationship among extracted components, we compute a common articulated structure for these input shapes. With the aid of this articulated structure, we solve the shape interpolation by combining 1) a global pose interpolation of near-rigid components from the source shape to the target shape with 2) a local gradient field interpolation for each pair of components, followed by solving a Poisson equation in order to reconstruct an interpolated shape. As a result, an aesthetically pleasing shape interpolation can be generated, with even the poses of shapes varying significantly. In contrast to a recent state-of-the-art work, the proposed approach can achieve comparable or even better results and have better computational efficiency as well.
Hung-Kuo Chu, Tong-Yee Lee
IEEE Trans. Vis. Comput. Graph.2
2008 Animation Key-Frame Extraction and Simplification Using Deformation Analysis
abstract
Three-dimensional animating meshes have been widely used in the computer graphics and video game industries. Reducing the animating mesh complexity is a common way of overcoming the rendering limitation or network bandwidth. Thus, we present a compact representation for animating meshes based on novel key-frames extraction and animating mesh simplification approaches. In contrast to the general simplification and key-frames extraction approaches which are driven by geometry metrics, the proposed methods are based on a deformation analysis of animating mesh to preserve both the geometric features and motion characteristics. These two approaches can produce a very compact animation representation in spatial and temporal domains, and therefore they can be beneficial in many applications such as progressive animation transmission and animation segmentation and transferring.
Tong-Yee Lee, Chao-Hung Lin, Yu-Shuen Wang, Tai-Guang Chen
IEEE Trans. Circuits Syst. Video Technol.1
2008 Motion overview of human actions
abstract
During the last decade, motion capture data has emerged and gained a leading role in animations, games and 3D environments. Many of these applications require the creation of expressive overview video clips capturing the human motion, however sufficient attention has not been given to this problem. In this paper, we present a technique that generates an overview video based on the analysis of motion capture data. Our method is targeted for applications of 3D character based animations, automating, for example, the action summary and gameplay overview in simulations and computer games. We base our method on quantum annealing optimization with an objective function that respects the analysis of the character motion and the camera movement constraints. It automatically generates a smooth camera control path, splitting it to several shots if required. To evaluate our method, we introduce a novel camera placement metric which is evaluated against previous work and conduct a user study comparing our results with the various systems.
Jackie Assa, Daniel Cohen-Or, I-Cheng Yeh 0001, Tong-Yee Lee
ACM Trans. Graph.4
2008 Skeleton extraction by mesh contraction
abstract
Extraction of curve-skeletons is a fundamental problem with many applications in computer graphics and visualization. In this paper, we present a simple and robust skeleton extraction method based on mesh contraction. The method works directly on the mesh domain, without pre-sampling the mesh model into a volumetric representation. The method first contracts the mesh geometry into zero-volume skeletal shape by applying implicit Laplacian smoothing with global positional constraints. The contraction does not alter the mesh connectivity and retains the key features of the original mesh. The contracted mesh is then converted into a 1D curve-skeleton through a connectivity surgery process to remove all the collapsed faces while preserving the shape of the contracted mesh and the original topology. The centeredness of the skeleton is refined by exploiting the induced skeleton-mesh mapping. In addition to producing a curve skeleton, the method generates other valuable information about the object's geometry, in particular, the skeleton-vertex correspondence and the local thickness, which are useful for various applications. We demonstrate its effectiveness in mesh segmentation and skinning animation.
Oscar Kin-Chung Au, Chiew-Lan Tai, Hung-Kuo Chu, Daniel Cohen-Or, Tong-Yee Lee
ACM Trans. Graph.5
2008 Self-animating images: illusory motion using repeated asymmetric patterns
abstract
Illusory motion in a still image is a fascinating research topic in the study of human motion perception. Physiologists and psychologists have attempted to understand this phenomenon by constructing simple, color repeated asymmetric patterns (RAP) and have found several useful rules to enhance the strength of illusory motion. Based on their knowledge, we propose a computational method to generate self-animating images. First, we present an optimized RAP placement on streamlines to generate illusory motion for a given static vector field. Next, a general coloring scheme for RAP is proposed to render streamlines. Furthermore, to enhance the strength of illusion and respect the shape of the region, a smooth vector field with opposite directional flow is automatically generated given an input image. Examples generated by our method are shown as evidence of the illusory effect and the potential applications for entertainment and design purposes.
Ming-Te Chi, Tong-Yee Lee, Yingge Qu, Tien-Tsin Wong
ACM Trans. Graph.2
2008 Optimized scale-and-stretch for image resizing
abstract
We present a "scale-and-stretch" warping method that allows resizing images into arbitrary aspect ratios while preserving visually prominent features. The method operates by iteratively computing optimal local scaling factors for each local region and updating a warped image that matches these scaling factors as closely as possible. The amount of deformation of the image content is guided by a significance map that characterizes the visual attractiveness of each pixel; this significance map is computed automatically using a novel combination of gradient and salience-based measures. Our technique allows diverting the distortion due to resizing to image regions with homogeneous content, such that the impact on perceptually important features is minimized. Unlike previous approaches, our method distributes the distortion in all spatial directions, even when the resizing operation is only applied horizontally or vertically, thus fully utilizing the available homogeneous regions to absorb the distortion. We develop an efficient formulation for the nonlinear optimization involved in the warping function computation, allowing interactive image resizing.
Yu-Shuen Wang, Chiew-Lan Tai, Olga Sorkine-Hornung, Tong-Yee Lee
ACM Trans. Graph.4
2008 Texture Mapping with Hard Constraints Using Warping Scheme
abstract
Texture mapping with positional constraints is an important and challenging problem in computer graphics. In this paper, we first present a theoretically robust, foldover-free 2D mesh warping algorithm. Then we apply this warping algorithm to handle mapping texture onto 3D meshes with hard constraints. The proposed algorithm is experimentally evaluated and compared with the state-of-the-art method for examples with more challenging constraints. These challenging constraints may lead to large distortions and foldovers. Experimental results show that the proposed scheme can generate more pleasing results and add fewer Steiner vertices on the 3D mesh embedding.
Tong-Yee Lee, Shao-Wei Yen, I-Cheng Yeh 0001
IEEE Trans. Vis. Comput. Graph.1
2008 Curve-Skeleton Extraction Using Iterative Least Squares Optimization
abstract
A curve skeleton is a compact representation of 3D objects and has numerous applications. It can be used to describe an object's geometry and topology. In this paper, we introduce a novel approach for computing curve skeletons for volumetric representations of the input models. Our algorithm consists of three major steps: 1) using iterative least squares optimization to shrink models and, at the same time, preserving their geometries and topologies, 2) extracting curve skeletons through the thinning algorithm, and 3) pruning unnecessary branches based on shrinking ratios. The proposed method is less sensitive to noise on the surface of models and can generate smoother skeletons. In addition, our shrinking algorithm requires little computation, since the optimization system can be factorized and stored in the pre-computational step. We demonstrate several extracted skeletons that help evaluate our algorithm. We also experimentally compare the proposed method with other well-known methods. Experimental results show advantages when using our method over other techniques.
Yu-Shuen Wang, Tong-Yee Lee
IEEE Trans. Vis. Comput. Graph.2
2008 Focus+Context Visualization with Distortion Minimization
abstract
The need to examine and manipulate large surface models is commonly found in many science, engineering, and medical applications. On a desktop monitor, however, seeing the whole model in detail is not possible. In this paper, we present a new, interactive Focus+Context method for visualizing large surface models. Our method, based on an energy optimization model, allows the user to magnify an area of interest to see it in detail while deforming the rest of the area without perceivable distortion. The rest of the surface area is essentially shrunk to use as little of the screen space as possible in order to keep the entire model displayed on screen. We demonstrate the efficacy and robustness of our method with a variety of models.
Yu-Shuen Wang, Tong-Yee Lee, Chiew-Lan Tai
IEEE Trans. Vis. Comput. Graph.2
2008 Stylized Rendering Using Samples of a Painted Image
abstract
We introduce a novel technique to generate painterly art map (PAM) for 3D non-photorealistic rendering. Our technique can automatically transfer brush stroke textures and color changes to 3D models from samples of a painted image. Therefore, the generation of stylized images/animation in the style of a given artwork can be achieved. This new approach works particularly well for a rich variety of brush strokes ranging from simple 1D and 2D line-art strokes to very complicated ones with significant variations in stroke characteristics. During the rendering/animation process, the coherence of brush stroke textures and color changes over 3D surfaces can be well maintained. With PAM, we can also easily generate the illusion of flow animation over a 3D surface to convey the shape of a model.
Chung-Ren Yan, Ming-Te Chi, Tong-Yee Lee, Wen-Chieh Lin
IEEE Trans. Vis. Comput. Graph.3
2008 Adaptive Geometry Image
abstract
We present a novel post-processing utility called adaptive geometry image (AGIM) for global parameterization techniques that can embed a 3D surface onto a rectangular1 domain. This utility first converts a single rectangular parameterization into many different tessellations of square geometry images(GIMs) and then efficiently packs these GIMs into an image called AGIM. Therefore, undersampled regions of the input parameterization can be up-sampled accordingly until the local reconstruction error bound is met. The connectivity of AGIM can be quickly computed and dynamically changed at rendering time. AGIM does not have T-vertices, and therefore no crack is generated between two neighboring GIMs at different tessellations. Experimental results show that AGIM can achieve significant PSNR gain over the input parameterization, AGIM retains the advantages of the original GIM and reduces the reconstruction error present in the original GIM technique. The AGIM is also for global parameterization techniques based on quadrilateral complexes. Using the approximate sampling rates, the PolyCube-based quadrilateral complexes with AGIM can outperform state-of-the-art multichart GIM technique in terms of PSNR.
Chih-Yuan Yao, Tong-Yee Lee
IEEE Trans. Vis. Comput. Graph.2
2008 Example-driven animation synthesis
Yu-Shuen Wang, Tong-Yee Lee
Vis. Comput.2
2007 Interactive Model Decomposition
abstract
In this paper, we propose an interactive model decomposition scheme. In preprocess, we automatically build a protrusive graph (PG) for any given 3D model. The purpose of the PG is to give the user a good clue to partition the model into visually-significant parts. Then, this scheme can interactively partition models according to the user-specified partitioning requirement. Finally, an iterative clustering is used to stabilize our partitions and a smoothing refinement is used to smooth the boundary between adjacent partitions. The experimental results show that the proposed scheme is a flexible and powerful method to decompose models into their significant components.
Yu-Shuen Wang, Tong-Yee Lee, Chao-Hung Lin
CAD/Graphics2
2007 Mesh pose-editing using examples
abstract
Abstract An easy‐to‐use mesh pose‐editing system is presented. We take advantage of both skeleton‐based and example‐based approaches in order to provide an intuitive way for artists to edit mesh poses. Our system automatically extracts the skeletons of the remaining example models once the skeleton of a reference mesh is constructed. In our editing system the desired skeleton can be easily and naturally posed using an inverse kinematics (IK) algorithm incorporated with searching the optimal weights in the defined skeleton space of examples meshes. Eventually, the desired shape with detailed deformation can be constructed by blending the example meshes. Experimental results show that the proposed system provides an easy and intuitive control on mesh pose‐editing. Copyright © 2007 John Wiley & Sons, Ltd.
Tong-Yee Lee, Chao-Hung Lin, Hung-Kuo Chu, Yu-Shuen Wang, Shao-Wei Yen, Chang-Rung Tsai
Comput. Animat. Virtual Worlds1
2006 Generating genus-n-to-m mesh morphing using spherical parameterization
abstract
Abstract Surface parameterization is a fundamental tool in computer graphics and benefits many applications such as texture mapping, morphing, and re‐meshing. Many spherical parameterization schemes with very nice properties have been proposed and widely used in the past. However, it is well known that the spherical parameterization is limited to genus‐0models. In this paper, we first propose a novel framework to extend spherical parameterization for handling a genus‐n surface. In this framework, we represent a surface S of arbitrary genus by a positive mesh O and several negative meshes Ni. Each negative surface is used to represent a hole. A positive surface O is obtained by removing all holes in the original surface S. Then, both positive and negative meshes are genus‐0 and can be spherically parameterized, respectively. To compute S, we can use a Boolean difference operation to subtract negative Nifrom a positive O. Next, we apply this novel framework to generate genus‐n‐to‐m mesh morphing application without restriction of n = m. Finally, there are many interesting non‐genus‐0 mesh morphing sequences generated. Copyright © 2006 John Wiley & Sons, Ltd.
Tong-Yee Lee, Chih-Yuan Yao, Hung-Kuo Chu, Ming-Jen Tai, Cheng-Chieh Chen
Comput. Animat. Virtual Worlds1
2006 Stylized and Abstract Painterly Rendering System Using a Multiscale Segmented Sphere Hierarchy
abstract
This paper presents a novel system framework for interactive, three-dimensional, stylized, abstract painterly rendering. In this framework, the input models are first represented using 3D point sets and then this point-based representation is used to build a multiresolution bounding sphere hierarchy. From the leaf to root nodes, spheres of various sizes are rendered into multiple-size strokes on the canvas. The proposed sphere hierarchy is developed using multiscale region segmentation. This segmentation task assembles spheres with similar attribute regularities into a meaningful region hierarchy. These attributes include colors, positions, and curvatures. This hierarchy is very useful in the following respects: 1) it ensures the screen-space stroke density, 2) controls different input model abstractions, 3) maintains region structures such as the edges/boundaries at different scales, and 4) renders models interactively. By choosing suitable abstractions, brush stroke, and lighting parameters, we can interactively generate various painterly styles. We also propose a novel scheme that reduces the popping effect in animation sequences. Many different stylized images can be generated using the proposed framework.
Ming-Te Chi, Tong-Yee Lee
IEEE Trans. Vis. Comput. Graph.2
2006 Segmenting a deforming mesh into near-rigid components
Tong-Yee Lee, Yu-Shuen Wang, Tai-Guang Chen
Vis. Comput.1
2005 Feature-Based Texture Synthesis
Tong-Yee Lee, Chung-Ren Yan
ICCSA (3)1
2005 A Fast 2D Shape Interpolation Technique
Ping-Hsien Lin, Tong-Yee Lee
ICCSA (3)2
2005 Feature-Constrained Texturing System for 3D Models
Tong-Yee Lee, Shaur-Uei Yan
KES (3)1
2005 Real-Time 3D Artistic Rendering System
Tong-Yee Lee, Shaur-Uei Yan, Yong-Nien Chen, Ming-Te Chi
KES (3)1
2005 Mesh decomposition using motion information from animation sequences
abstract
Abstract In computer graphics, mesh decomposition is a fundamental problem and it can benefit many applications. In this paper, we propose a novel mesh decomposition algorithm using motion information derived from a given animation sequence. The proposed algorithm first use principal component analysis (PCA) to construct a compact representation of a given animation sequence. Next, from this representation, we derive several motion parameters including motion complexity and similarity. Finally, we decompose a given mesh into sub‐meshes using derived motion information and subdivide the triangles along the cutting paths for the smoother borders between the mesh parts. Our experimental results show that this new decomposition scheme can bring the benefit of good compression ratios on animation sequences. Copyright © 2005 John Wiley & Sons, Ltd.
Tong-Yee Lee, Ping-Hsien Lin, Shaur-Uei Yan, Chun-Hao Lin
Comput. Animat. Virtual Worlds1
2005 Progressive mesh metamorphosis
abstract
Abstract This paper describes a new integrated scheme for metamorphosis between two closed manifold genus‐0 polyhedral models. Spherical parameterizations of the source and target models are created first. To control the morphing, any number of feature vertex pairs is specified and a fold‐over free warping method is used to align two spherical embeddings. Our method does not create a merged meta‐mesh or execute re‐meshing to construct a common connectivity for morphs. Alternatively, a scheme for the progressive connectivity transformation of two spherical parameterizations is employed to generate the intermediate meshes. A novel semi‐overlay with a geomorph scheme is proposed to reduce the popping effects caused by the connectivity transformation. We demonstrate several examples of aesthetically pleasing morphing sequences using the proposed scheme. Copyright © 2005 John Wiley & Sons, Ltd.
Chao-Hung Lin, Tong-Yee Lee, Hung-Kuo Chu, Chih-Yuan Yao
Comput. Animat. Virtual Worlds2
2005 Metamorphosis of 3D Polyhedral Models Using Progressive Connectivity Transformations
abstract
Three-dimensional metamorphosis is a powerful technique to produce a 3D shape transformation between two or more existing models. In this paper, we propose a novel 3D morphing technique that avoids creating a merged embedding that contains the faces, edges, and vertices of two given embeddings. This novel 3D morphing technique dynamically adds or removes vertices to gradually transform the connectivity of 3D polyhedrons from a source model into a target model and simultaneously creates the intermediate shapes. In addition, a priority control function provides the animators with control of arising or dissolving of input models' features in a morphing sequence. This is a useful tool to control a morphing sequence more easily and flexibly. Several examples of aesthetically pleasing morphs are demonstrated using the proposed method.
Chao-Hung Lin, Tong-Yee Lee
IEEE Trans. Vis. Comput. Graph.2
2004 Photo-realistic 3D Head Modeling Using Multi-view Images
Tong-Yee Lee, Ping-Hsien Lin, Tz-Hsien Yang
ICCSA (2)1
2004 Texture Mapping on Arbitrary 3D Surfaces
Tong-Yee Lee, Shaur-Uei Yan
ICCSA (2)1
2004 Camera-Sampling Field and Its Applications
abstract
In this paper, we propose a novel vector field, called a camera-sampling field, to represent the sampling density distribution of a pinhole camera. We give the derivation and discuss some essential properties of the camera-sampling field, including flux, divergence, curl, gradient, level surface, and sampling patterns. This vector field reveals camera-sampling concisely and facilitates camera sampling analysis. The usage for this vector field in several computer graphics applications is introduced, such as determining the splat kernel for image-based rendering, texture filtering, mipmap level selection, level transition criteria for LOD, and LDI-construction.
Ping-Hsien Lin, Tong-Yee Lee
IEEE Trans. Vis. Comput. Graph.2
2003 Morphology-Based 3D Volume Metamorphosis
Tong-Yee Lee, Chao-Hung Lin, Wen-Hsiu Wang
ICCSA (3)1
2003 A Hybrid Scheme for Interactive Rendering a Virtual Environment
Tong-Yee Lee, Ping-Hsien Lin, Tz-Hsien Yang
ICCSA (3)1
2003 Fast and Intuitive Metamorphosis of 3D Polyhedral Models Using SMCC Mesh Merging Scheme
abstract
A very fast and intuitive approach to generate the metamorphosis of two genus 0 3D polyhedral models is presented. There are two levels of correspondence specified by animators to control morphs. The higher level requires the animators to specify scatter features to decompose the input models into several corresponding patches. The lower level optionally allows the animators to specify extra features on each corresponding patch for finer correspondence control. Once these two levels of correspondence are established, the proposed schemes automatically and efficiently establish a complete one-to-one correspondence between two models. We propose a novel technique called SMCC (Structures of Minimal Contour Coverage) to efficiently and robustly merge corresponding embeddings. The SMCC scheme can compute merging in linear time. The performance of the proposed methods is comparable to or better than state-of-the-art 3D polyhedral metamorphosis. We demonstrate several examples of aesthetically pleasing morphs, which can be created very quickly and intuitively.
Tong-Yee Lee, Po-Hua Huang
IEEE Trans. Vis. Comput. Graph.1
2002 Feature-Guided Shape-Based Image Interpolation
abstract
A feature-guided image interpolation scheme is presented. It is an effective and improved, shape-based interpolation method used for interpolating image slices in medical applications. The proposed method integrates feature line-segments to guide the shape-based method for better shape interpolation. An automatic method for finding these line segments is given. The proposed feature-guided shape-based method can manage translation, rotation and scaling situations when the slices have similar shapes. It can also interpolate intermediate shapes when the successive slices do not have similar shapes. This method is experimentally evaluated using artificial and real two-dimensional and three-dimensional data. The proposed method generated satisfactory interpolated results in these experiments. We demonstrate the practicality, effectiveness and reproducibility of the proposed method for interpolating medical images.
Tong-Yee Lee, Chao-Hung Lin
IEEE Trans. Medical Imaging1
2001 Computer-aided prototype system for nose surgery
abstract
Rhinoplasty, or surgery to reshape the nose, is one of the most common of all plastic-surgery procedures. Rhinoplasty can enhance a patient's appearance and self-confidence, may also correct a birth defect or injury, or help relieve some breathing problem. In this paper, we present a three-dimensional (3-D) surgical simulation system, which can assist surgeons in planning rhinoplasty procedures. This system employs computer graphics and image-processing techniques for the simulation of a rhinoplasty. Although the presented algorithms themselves are not new, the proposed system exploits the new idea to apply 3-D morphing for rhinoplasty, and simulation results are useful for the physicians. According to patients' expectation of what they would like their noses to look like, our system simulates expected results. Our tools provide quantitative measurements of a nose structure. Using these quantitative results, surgeons can arrange appropriate preoperative plans for patients. Finally, experimental results and experiences are reported to evaluate the usefulness of the proposed system.
Tong-Yee Lee, Chao-Hung Lin, Han-Ying Lin
IEEE Trans. Inf. Technol. Biomed.1
2000 Morphology-based Three-dimensional Interpolation
abstract
In many medical applications, the number of available two-dimensional (2-D) images is always insufficient. Therefore, the three-dimensional (3-D) reconstruction must be accomplished by appropriate interpolation methods to fill gaps between available image slices. In this paper, we propose a morphology-based algorithm to interpolate the missing data. The proposed algorithm consists of several steps. First, the object or hole contours are extracted using conventional image-processing techniques. Second, the object or hole matching issue is evaluated. Prior to interpolation, the centroids of the objects are aligned. Next, we employ a dilation operator to transform digital images into distance maps and we correct the distance maps if required. Finally, we utilize an erosion operator to accomplish the interpolation. Furthermore, if multiple objects or holes are interpolated, we blend them together to complete the algorithm. We experimentally evaluate the proposed method against various synthesized cases reported in the literature. Experimental results show that the proposed method is able to handle general object interpolation effectively.
Tong-Yee Lee, Wen-Hsiu Wang
IEEE Trans. Medical Imaging1
1999 Interactive 3-D virtual colonoscopy system
abstract
We describe a low-cost three-dimensional (3-D) virtual colonoscopy system that is a noninvasive technique for examining the entire colon and can assist physicians in detecting polyps inside the colon. Using the helical CT data and proposed techniques, we can three-dimensionally reconstruct and visualize the inner surface of the colon. We generate high resolution of video views of the colon interior structures as if the viewer's eyes were inside the colon. The physicians can virtually navigate inside the colon in two different modes: interactive and automatic navigation, respectively. For automatic navigation, the flythrough path is determined a priori using the 3-D thinning and two-pass tracking schemes. The whole colon is spatially subdivided into several cells, and only potentially visible cells are taken into account during rendering. To further improve rendering efficiency, potentially visible cells are rendered at different levels of detail. Additionally, a chain of bounding volume in each cell is used to avoid penetrating through the colon during navigation. In comparison with previous work, the proposed system can efficiently accomplish required preprocessing tasks and afford adequate rendering speeds on a low-cost PC system.
Tong-Yee Lee, Ping-Hsien Lin, Chao-Hung Lin, Yung-Nien Sun, Xi-Zhang Lin
IEEE Trans. Inf. Technol. Biomed.1
1999 Three-dimensional facial model reconstruction and plastic surgery simulation
abstract
Facial model reconstruction and surgical simulation are essential to plastic surgery in today's medicine. Both can help surgeons to design appropriate repair plans and procedures prior to actual surgery. In this paper, we exploit a metamorphosis technique in our new design. First, using metamorphosis and vision techniques, we can establish three-dimensional facial models from a given photo. Second, we design several morphing operators, including augmentation, cutting, and lacerating. Experiments show that the proposed algorithms can successfully create acceptable facial models and generate realistically visual effects of surgical simulation.
Tong-Yee Lee, Yung-Nien Sun, Yung-Ching Lin, Leewen Lin, Chungnan Lee
IEEE Trans. Inf. Technol. Biomed.1
1998 Fast Feature-Based Metamorphosis and Operator Design
abstract
Metamorphosis is a powerful visual technique, for producing interesting transition between two images or volume data. Image or volume metamorphosis using simple features provides flexible and easy control of visual effect. The feature‐based image warping proposed by Beier and Neely is a brute‐force approach. In this paper, first, we propose optimization methods to reduce their warping time without noticeable loss of image quality. Second, we extend our methods to 3D volume data and propose several interesting warping operators allowing global and local metamorphosis of volume data.
Tong-Yee Lee, Young-Ching Lin, Leeween Lin, Yung-Nien Sun
Comput. Graph. Forum1
1997 A World-Wide Web based distributed animation environment
Chungnan Lee, Tong-Yee Lee, Tainchi Lu, Yao-Tsung Chen
Comput. Networks ISDN Syst.2
1997 Parallel Omplementation of a Ray Tracing Algorithm for Distributed Memory Parallel Computers
abstract
Ray tracing is a well known technique to generate life-like images. Unfortunately, ray tracing complex scenes can require large amounts of CPU time and memory storage. Distributed memory parallel computers with large memory capacities and high processing speeds are ideal candidates to perform ray tracing. However, the computational cost of rendering pixels and patterns of data access cannot be predicted until runtime. To parallelize such an application efficiently on distributed memory parallel computers, the issues of database distribution, dynamic data management and dynamic load balancing must be addressed. In this paper, we present a parallel implementation of a ray tracing algorithm on the Intel Delta parallel computer. In our database distribution, a small fraction of database is duplicated on each processor, while the remaining part is evenly distributed among groups of processors. In the system, there are multiple copies of the entire database in the memory of groups of processors. Dynamic data management is acheived by an ALRU cache scheme which can exploit image coherence to reduce data movements in ray tracing consecutive pixels. We balance load among processors by distributing subimages to processors in a global fashion based on previous workload requests. The success of our implementation depends crucially on a number of parameters which are experimentally evaluated. © 1997 John Wiley & Sons, Ltd.
Tong-Yee Lee, Cauligi S. Raghavendra, John B. Nicholas
Concurr. Pract. Exp.1
1997 A Web-based Distributed and Collaborative 3D Animation Environment
abstract
Many applications on the Web require active processing and co-ordination of services. In this paper we describe the design of a distributed 3D animation system built by integrating the Java language, parallel virtual machine (PVM) software, a collaborative mechanism, 3D computer graphics and the Web technologies. To achieve collaborative co-operation and functional independence, session control and system agents are devised in this system. In particular, we propose a simplified collaborative group definition, collaborative policies, and the state of participants to dynamically manage participants in this Web-based distributed environment. Based on our proposed mechanism, the server can efficiently determine the status of collaboration activities. © 1997 John Wiley & Sons, Ltd.
Tainchi Lu, Chungwen Chiang, Chungnan Lee, Tong-Yee Lee
Concurr. Pract. Exp.4
1997 Exploitation of Image Parallelism for Ray Tracing 3D Scenes on 2D Mesh Multicomputers
Tong-Yee Lee
Parallel Comput.1
1996 Image Composition Schemes for Sort-Last Polygon Rendering on 2D Mesh Multicomputers
abstract
In a sort-last polygon rendering system, the efficiency of image composition is very important for achieving fast rendering. In this paper, the implementation of a sort-last rendering system on a general purpose multicomputer system is described. A two-phase sort-last-full image composition scheme is described first, and then many variants of it are presented for 2D mesh message-passing multicomputers, such as the Intel Delta and Paragon. All the proposed schemes are analyzed and experimentally evaluated on Caltech's Intel Delta machine for our sort-last parallel polygon renderer. Experimental results show that sort-last-sparse strategies are better suited than sort-last-full schemes for software implementation on a general purpose multicomputer system. Further, interleaved composition regions perform better than coherent regions. In a large multicomputer system. Performance can be improved by carefully scheduling the tasks of rendering and communication. Using 512 processors to render our test scenes, the peak rendering rate achieved on a 282,144 triangle dataset is dose to 4.6 million triangles per second which is comparable to the speed of current state-of-the-art graphics workstations.
Tong-Yee Lee, Cauligi S. Raghavendra, John B. Nicholas
IEEE Trans. Vis. Comput. Graph.1
1995 AN Efficient Sort-Last Polygon Rendering Scheme on 2-D Mesh Parallel Computers
Tong-Yee Lee, Cauligi S. Raghavendra, John B. Nicholas
ICPP (3)1
1994 Experimental Evaluation of Load Balancing Strategies for Ray Tracing on Parallel Processors
abstract
Ray tracing is one of the computer graphics techniques used to render high quality images. Unfortunately, ray tracing complex scenes can require large amounts of CPU time, making the technique impractical for everyday use. Parallel ray tracing algorithms could potentially be used to reduce the high computational cost. However, pixel computation times can vary significantly, and naive attempts at parallelization give poor speedups due to load imbalance between the processors. In this paper, we evaluate the performance of three load balancing schemes for ray tracing on parallel processors, and propose two new load balancing strategies. To evaluate the performance, we implement all these strategies on the 512 processor Intel Touchstone Delta at Caltech.
Tong-Yee Lee, Cauligi S. Raghavendra, John B. Nicholas
ICPP (2)1
1993 A Fully Distributed Parallel Ray Tracing Scheme on the Delta Touchstone Machine
abstract
The authors describe a fully distributed, parallel algorithm for ray-tracing problem. Load balancing is achieved through the use of comb distribution to roughly assign the same amount of pixels to each processor first, and then dynamically redistribute excessive loads among processors to keep each processor busy. In this model, there is no need for a master node to be responsible for dynamic scheduling. When each node finishes its job, it just requests an extra job from one of its neighbors. The authors implement their algorithm on Intel Delta Touchstone machine with 2-D mesh network topology and provide simulation results. With their scheme, they can get good speedup and high efficiency without much communication overhead.>
Tong-Yee Lee, Cauligi S. Raghavendra, John B. Nicholas
HPDC1