I-Chao Shen

dblp:48/8319 · DBLP profile ↗
← Back
41ranked-venue papers
7as first author
27since 2021 · last 2026
0000-0003-4201-3793ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 33 · 7 first-author · 23 since 2021Human-computer interaction and ubiquitous computing · 9 · 7 since 2021Artificial intelligence and machine learning · 6 · 5 since 2021Computer networks · 1
YearPublicationVenuePosition
2026 Axis-Aligned Document Dewarping
abstract
Document dewarping is crucial for many applications. However, existing learning-based methods rely heavily on supervised regression with annotated data without fully leveraging the inherent geometric properties of physical documents. Our key insight is that a well-dewarped document is defined by its axis-aligned feature lines. This property aligns with the inherent axis-aligned nature of the discrete grid geometry in planar documents. Harnessing this property, we introduce three synergistic contributions: for the training phase, we propose an axis-aligned geometric constraint to enhance document dewarping; for the inference phase, we propose an axis alignment preprocessing strategy to reduce the dewarping difficulty; and for the evaluation phase, we introduce a new metric, Axis-Aligned Distortion (AAD), that not only incorporates geometric meaning and aligns with human visual perception but also demonstrates greater robustness. As a result, our method achieves state-of-the-art performance on multiple existing benchmarks, improving the AAD metric by 18.2% to 34.5%.
Chaoyun Wang, I-Chao Shen, Takeo Igarashi, Caigui Jiang
AAAI2
2026 Sketch-Guided Anime Hair Editing Using Multimodal Diffusion Transformer (Student Abstract)
abstract
Anime hair design is crucial but challenging, as it conveys personality and emotion through stylized geometry and layered structure. In this work, we propose a sketch-guided approach for intuitive control of multimodal diffusion transformers (MMDiT) to generate semantically consistent anime hairstyles. We adopt a wisp-level flowline input integrated with a fine-tuned MMDiT to transfer hairstyles while preserving character identity. We believe that this fine-grained sketch control within the MMDiT framework may offer a promising path for structured anime hair editing.
I-Chao Shen, Haoran Xie 0002
AAAI2
2025 FontCraft: Multimodal Font Design Using Interactive Bayesian Optimization
abstract
International audience
Yuki Tatsukawa, I-Chao Shen, Mustafa Doga Dogan, Anran Qi, Yuki Koyama 0001, Ariel Shamir, Takeo Igarashi
CHI2
2025 CompAct: Designing Interconnected Compliant Mechanisms with Targeted Actuation Transmissions
abstract
Compliant mechanisms enable the creation of compact and easy-to-fabricate devices for tangible interaction. This work explores interconnected compliant mechanisms consisting of multiple joints and rigid bodies to transmit and process displacements as signals that result from physical interactions. As these devices are difficult to design due to their vast and complex design space, we developed a graph-based design algorithm and computational tool to help users program and customize such computational functions and procedurally model physical designs. When combined with active materials with actuation and sensing capabilities, these devices can also render and detect haptic interaction. Our design examples demonstrate the tool's capability to respond to relevant HCI concepts, including building modular physical interface toolkits, encrypting tangible interactions, and customizing user augmentation for accessibility. We believe the tool will facilitate the generation of new interfaces with enriched affordance.
Humphrey Yang, I-Chao Shen, Nikolas Martelaro, Bo Zhu 0002, Haoran Xie 0002, Takeo Igarashi, Lining Yao
CHI2
2025 Interactive Multilayer Gaussian Garments for Low-Cost Try-On
abstract
Numerous recent works have utilized 3D Gaussian Splatting to represent high-fidelity digital avatars. However, none have enabled interactive multilayer Gaussian garments for virtual try-ons without relying on expensive hardware, such as a camera array and/or multiple GPUs. To enable affordable mix-and-match dressing—dressing 3D avatars with realistic and complex combinations of garments—it is crucial to handle the interactions between multiple layers of garments using consumer-level capturing hardware. To address this, we present a novel screenspace layer resolution method combined with physical simulation and Gaussian garments to enable realistic multilayer mix-and-match avatar dressing at interactive rates using low-cost hardware. As an offline process, we capture multiple static garments individually using only a single mobile camera on a static mannequin and then perform a dual reconstruction of Gaussians and simulation mesh. During runtime, these Gaussians are driven by a fast but simple physics simulator, whose output may contain inter-penetrations across garment layers. Our method fixes these in screenspace by rasterizing the simulation mesh from various camera views and culling the Gaussians that are skinned to unseen mesh triangles. We show the effectiveness of our approach by demonstrating mix-and-match dressing results at interactive rates using short-sleeves, long-sleeves, a fur vest, and a singlet. Additionally, we showcase a webcam-based interactive try-on application to further illustrate the capabilities of our system.
Ryan S. Zesch, I-Chao Shen, Haoran Xie 0002, Bo Zhu 0002, Shinjiro Sueda, Takeo Igarashi
Graphics Interface2
2025 NeRF is a Valuable Assistant for 3D Gaussian Splatting
abstract
We introduce NeRF-GS, a novel framework that jointly optimizes Neural Radiance Fields (NeRF) and 3D Gaussian Splatting (3DGS). This framework leverages the inherent continuous spatial representation of NeRF to mitigate several limitations of 3DGS, including sensitivity to Gaussian initialization, limited spatial awareness, and weak inter-Gaussian correlations, thereby enhancing its performance. In NeRF-GS, we revisit the design of 3DGS and progressively align its spatial features with NeRF, enabling both representations to be optimized within the same scene through shared 3D spatial information. We further address the formal distinctions between the two approaches by optimizing residual vectors for both implicit features and Gaussian positions to enhance the personalized capabilities of 3DGS. Experimental results on benchmark datasets show that NeRF-GS surpasses existing methods and achieves state-of-the-art performance. This outcome confirms that NeRF and 3DGS are complementary rather than competing, offering new insights into hybrid approaches that combine 3DGS and NeRF for efficient 3D scene representation.
Shuangkang Fang, I-Chao Shen, Takeo Igarashi, Yufeng Wang 0004, Zesheng Wang 0002, Yi Yang 0033, Wenrui Ding, Shuchang Zhou 0001
ICCV2
2025 MeshLLM: Empowering Large Language Models to Progressively Understand and Generate 3D Mesh
Shuangkang Fang, I-Chao Shen, Yufeng Wang 0004, Yi-Hsuan Tsai, Yi Yang 0033, Shuchang Zhou 0001, Wenrui Ding, Takeo Igarashi, Ming-Hsuan Yang 0001
ICCV2
2025 AutoSketch: VLM-assisted Style-Aware Vector Sketch Completion
abstract
Sketches are an important medium of expression and recently many works concentrate on automatic sketch creations. One such ability very useful for amateurs is text-based completion of a partial sketch to create a complex scene, while preserving the style of the partial sketch. Existing methods focus solely on generating sketch that match the content in the input prompt in a predefined style, ignoring the styles of the input partial sketches, e.g., the global abstraction level and local stroke styles. To address this challenge, we introduce AutoSketch, a style-aware vector sketch completion method that accommodates diverse sketch styles and supports iterative sketch completion. AutoSketch completes the input sketch in a style-consistent manner using a two-stage method. In the first stage, we initially optimize the strokes to match an input prompt augmented by style descriptions extracted from a vision-language model (VLM). Such style descriptions lead to non-photorealistic guidance images which enable more content to be depicted through new strokes. In the second stage, we utilize the VLM to adjust the strokes from the previous stage to adhere to the style present in the input partial sketch through an iterative style adjustment process. In each iteration, the VLM identifies a list of style differences between the input sketch and the strokes generated in the previous stage, translating these differences into adjustment codes to modify the strokes. We compare our method with existing methods using various sketch styles and prompts, perform extensive ablation studies and qualitative and quantitative evaluations, and demonstrate that AutoSketch can support diverse sketching scenarios.
Hsiao-Yuan Chin, I-Chao Shen, Yi-Ting Chiu, Ariel Shamir, Bing-Yu Chen 0004
SIGGRAPH Asia2
2025 Approximating Procedural Models of 3D Shapes with Neural Networks
abstract
Abstract Procedural modeling is a popular technique for 3D content creation and offers a number of advantages over alternative techniques for modeling 3D shapes. However, given a procedural model, predicting the procedural parameters of existing data provided in different modalities can be challenging. This is because the data may be in a different representation than the one generated by the procedural model, and procedural models are usually not invertible, nor are they differentiable. In this paper, we address these limitations and introduce an invertible and differentiable representation for procedural models. We approximate parameterized procedures with a neural network architecture NNProc that learns both the forward and inverse mapping of the procedural model by aligning the latent spaces of shape parameters and shapes. The network is trained in a manner that is agnostic to the inner workings of the procedural model, implying that models implemented in different languages or systems can be used. We demonstrate how the proposed representation can be used for both forward and inverse procedural modeling. Moreover, we show how NNProc can be used in conjunction with optimization for applications such as shape reconstruction from an image or a 3D Gaussian Splatting.
Ishtiaque Hossain, I-Chao Shen, Oliver van Kaick
Comput. Graph. Forum2
2025 LayoutRectifier: An Optimization-based Post-processing for Graphic Design Layout Generation
abstract
Abstract Recent deep learning methods can generate diverse graphic design layouts efficiently. However, these methods often create layouts with flaws, such as misalignment, unwanted overlaps, and unsatisfied containment. To tackle this issue, we propose an optimization‐based method called LayoutRectifier, which gracefully rectifies auto‐generated graphic design layouts to reduce these flaws while minimizing deviation from the generated layout. The core of our method is a two‐stage optimization. First, we utilize grid systems, which professional designers commonly use to organize elements, to mitigate misalignments through discrete search. Second, we introduce a novel box containment function designed to adjust the positions and sizes of the layout elements, preventing unwanted overlapping and promoting desired containment. We evaluate our method on content‐agnostic and content‐aware layout generation tasks and achieve better‐quality layouts that are more suitable for downstream graphic design tasks. Our method complements learning‐based layout generation methods and does not require additional training.
I-Chao Shen, Ariel Shamir, Takeo Igarashi
Comput. Graph. Forum1
2025 Real-Time Per-Garment Virtual Try-On with Temporal Consistency for Loose-Fitting Garments
abstract
Abstract Per‐garment virtual try‐on methods collect garment‐specific datasets and train networks tailored to each garment to achieve superior results. However, these approaches often struggle with loose‐fitting garments due to two key limitations: (1) They rely on human body semantic maps to align garments with the body, but these maps become unreliable when body contours are obscured by loose‐fitting garments, resulting in degraded outcomes; (2) They train garment synthesis networks on a per‐frame basis without utilizing temporal information, leading to noticeable jittering artifacts. To address the first limitation, we propose a two‐stage approach for robust semantic map estimation. First, we extract a garment‐invariant representation from the raw input image. This representation is then passed through an auxiliary network to estimate the semantic map. This enhances the robustness of semantic map estimation under loose‐fitting garments during garment‐specific dataset generation. To address the second limitation, we introduce a recurrent garment synthesis framework that incorporates temporal dependencies to improve frame‐to‐frame coherence while maintaining real‐time performance. We conducted qualitative and quantitative evaluations to demonstrate that our method outperforms existing approaches in both image quality and temporal coherence. Ablation studies further validate the effectiveness of the garment‐invariant representation and the recurrent synthesis framework.
Zaiqiang Wu, I-Chao Shen, Takeo Igarashi
Comput. Graph. Forum2
2025 The Mokume Dataset and Inverse Modeling of Solid Wood Textures
abstract
We present the Mokume dataset for solid wood texturing consisting of 190 cube-shaped samples of various hard and softwood species documented by high-resolution exterior photographs, annual ring annotations, and volumetric computed tomography (CT) scans. A subset of samples further includes photographs along slanted cuts through the cube for validation purposes. Using this dataset, we propose a three-stage inverse modeling pipeline to infer solid wood textures using only exterior photographs. Our method begins by evaluating a neural model to localize year rings on the cube face photographs. We then extend these exterior 2D observations into a globally consistent 3D representation by optimizing a procedural growth field using a novel iso-contour loss. Finally, we synthesize a detailed volumetric color texture from the growth field. For this last step, we propose two methods with different efficiency and quality characteristics: a fast inverse procedural texture method, and a neural cellular automaton (NCA). We demonstrate the synergy between the Mokume dataset and the proposed algorithms through comprehensive comparisons with unseen captured data. We also present experiments demonstrating the efficiency of our pipeline's components against ablations and baselines. Our code, the dataset, and reconstructions are available via https://mokumeproject.github.io/.
Maria Larsson, Hodaka Yamaguchi, Ehsan Pajouheshgar, I-Chao Shen, Kenji Tojo, Chia-Ming Chang 0003, Lars Hansson, Olof Broman, Takashi Ijiri, Ariel Shamir, Wenzel Jakob, Takeo Igarashi
ACM Trans. Graph.4
2025 StylePart: image-based shape part manipulation
abstract
Abstract Direct part-level manipulation of man-made shapes in an image is desired given its simplicity. However, it is not intuitive given the existing manually created cuboid and cylinder controllers. To tackle this problem, we present StylePart, a framework that enables direct shape manipulation of an image by leveraging generative models of both images and 3D shapes. Our key contribution is a shape-consistent latent mapping function that connects the image generative latent space and the 3D man-made shape attribute latent space. Our method “forwardly maps” the image content to its corresponding 3D shape attributes, where the shape part can be easily manipulated. The attribute codes of the manipulated 3D shape are then “backwardly mapped” to the image latent code to obtain the final manipulated image. By using both forward and backward mapping, an user can edit the image directly without resorting to any 3D workflow. We demonstrate our approach through various manipulation tasks, including part replacement, part resizing, and shape orientation manipulation, and evaluate its effectiveness through extensive ablation studies.
I-Chao Shen, Li-Wen Su, Yu-Ting Wu 0001, Bing-Yu Chen 0004
Vis. Comput.1
2024 Virtual Measurement Garment for Per-Garment Virtual Try-On
abstract
The popularity of virtual try-on methods has increased in recent years as they allow users to preview the appearance of garments on themselves without physically wearing them. However, existing image-based methods for general virtual try-on provide limited support to synthesize realistic and consistent garment images under different poses, due to two main difficulties: 1) the dataset used to train these methods contains a vast collection of garments, but they lack fine details of each garment; 2) they synthesize results by warping the front-view image of the target garment in a rest pose, which results in poor quality and detail for other viewpoints and poses. To overcome these drawbacks, per-garment virtual try-on methods train garment-specific networks that can produce high-quality results with fine-grained details for a particular target garment. However, existing per-garment virtual try-on methods require the use of a physical measurement garment, which limits their applicability. In this paper, we propose a novel per-garment virtual try-on method that leverages a virtual measurement garment, which eliminates the need for the physical measurement garment, to guide the synthesis of high-quality and temporally consistent garment images under various poses. Furthermore, we introduce a gap-filling module that effectively fills the gap between the synthesized garment and body parts. We conduct qualitative and quantitative evaluations against a state-of-the-art image-based virtual try-on method and ablation studies to demonstrate that our method achieves superior performance in terms of realism and consistency of the generated garment images.
Zaiqiang Wu, Toby Chong, I-Chao Shen, Takeo Igarashi
Graphics Interface4
2024 Learned Inference of Annual Ring Pattern of Solid Wood
abstract
Abstract We propose a method for inferring the internal anisotropic volumetric texture of a given wood block from annotated photographs of its external surfaces. The global structure of the annual ring pattern is represented using a continuous spatial scalar field referred to as the growth time field (GTF). First, we train a generic neural model that can represent various GTFs using procedurally generated training data. Next, we fit the generic model to the GTF of a given wood block based on surface annotations. Finally, we convert the GTF to an annual ring field (ARF) revealing the layered pattern and apply neural style transfer to render orientation‐dependent small‐scale features and colors on a cut surface. We show rendered results of various physically cut real wood samples. Our method has physical and virtual applications such as cut‐preview before subtractive fabricating solid wood artifacts and simulating object breaking.
Maria Larsson, Takashi Ijiri, I-Chao Shen, Hironori Yoshida, Ariel Shamir, Takeo Igarashi
Comput. Graph. Forum3
2024 FontCLIP: A Semantic Typography Visual-Language Model for Multilingual Font Applications
abstract
Abstract Acquiring the desired font for various design tasks can be challenging and requires professional typographic knowledge. While previous font retrieval or generation works have alleviated some of these difficulties, they often lack support for multiple languages and semantic attributes beyond the training data domains. To solve this problem, we present FontCLIP – a model that connects the semantic understanding of a large vision‐language model with typographical knowledge. We integrate typography‐specific knowledge into the comprehensive vision‐language knowledge of a pretrained CLIP model through a novel finetuning approach. We propose to use a compound descriptive prompt that encapsulates adaptively sampled attributes from a font attribute dataset focusing on Roman alphabet characters. FontCLIP's semantic typographic latent space demonstrates two unprecedented generalization abilities. First, FontCLIP generalizes to different languages including Chinese, Japanese, and Korean (CJK), capturing the typographical features of fonts across different languages, even though it was only finetuned using fonts of Roman characters. Second, FontCLIP can recognize the semantic attributes that are not presented in the training data. FontCLIP's dual‐modality and generalization abilities enable multilingual and cross‐lingual font retrieval and letter shape optimization, reducing the burden of obtaining desired fonts.
Yuki Tatsukawa, I-Chao Shen, Anran Qi, Yuki Koyama 0001, Takeo Igarashi, Ariel Shamir
Comput. Graph. Forum2
2024 Improving cache placement for efficient cache-based rendering
Yu-Ting Wu 0001, I-Chao Shen
Vis. Comput.2
2023 360MVSNet: Deep Multi-view Stereo Network with 360° Images for Indoor Scene Reconstruction
abstract
Recent multi-view stereo methods have achieved promising results with the advancement of deep learning techniques. Despite of the progress, due to the limited fields of view of regular images, reconstructing large indoor environments still requires collecting many images with sufficient visual overlap, which is quite labor-intensive. 360° images cover a much larger field of view than regular images and would facilitate the capture process. In this paper, we present 360MVSNet, the first deep learning network for multi-view stereo with 360° images. Our method combines uncertainty estimation with a spherical sweeping module for 360° images captured from multiple viewpoints in order to construct multi-scale cost volumes. By regressing volumes in a coarse-to-fine manner, high-resolution depth maps can be obtained. Furthermore, we have constructed EQMVS, a large-scale synthetic dataset that consists of over 50K pairs of RGB and depth maps in equirectangular projection. Experimental results demonstrate that our method can reconstruct large synthetic and real-world indoor scenes with significantly better completeness than previous traditional and learning-based methods while saving both time and effort in the data acquisition process.
Ching-Ya Chiu, Yu-Ting Wu 0001, I-Chao Shen, Yung-Yu Chuang
WACV3
2023 Data-guided Authoring of Procedural Models of Shapes
abstract
Abstract Procedural models enable the generation of a large amount of diverse shapes by varying the parameters of the model. However, writing a procedural model for replicating a collection of reference shapes is difficult, requiring much inspection of the original and replicated shapes during the development of the model. In this paper, we introduce a data‐guided method for aiding a programmer in creating a procedural model to replicate a collection of reference shapes. The user starts by writing an initial procedural model, and the system automatically predicts the model parameters for reference shapes, also grouping shapes by how well they are approximated by the current procedural model. The user can then update the procedural model based on the given feedback and iterate the process. Our system thus automates the tedious process of discovering the parameters that replicate reference shapes, allowing the programmer to focus on designing the high‐level rules that generate the shapes. We demonstrate through qualitative examples and a user study that our method is able to speed up the development time for creating procedural models of 2D and 3D man‐made shapes.
Ishtiaque Hossain, I-Chao Shen, Takeo Igarashi, Oliver van Kaick
Comput. Graph. Forum2
2023 Palette-Based and Harmony-Guided Colorization for Vector Icons
abstract
Abstract Colorizing icon is a challenging task, even for skillful artists, as it involves balancing aesthetics and practical considerations. Prior works have primarily focused on colorizing pixel‐based icons, which do not seamlessly integrate into the current vector‐based icon design workflow. In this paper, we propose a palette‐based colorization algorithm for vector icons without the need for rasterization. Our algorithm takes a vector icon and a five‐color palette as input and generates various colorized results for designers to choose from. Inspired by the common icon design workflow, we developed our algorithm to consist of two steps: generating a colorization template and performing the palette‐based color transfer. To generate the colorization templates, we introduce a novel vector icon colorization model that employs an MRF‐based loss and a color harmony loss. The color harmony loss encourages the alignment of the resulting color template with widely used harmony templates. We then map the predicted colorization template to chroma‐like palette colors to obtain diverse colorization results. We compare our results with those generated by previous pixel‐based icon colorization methods and validate the effectiveness of our algorithm by evaluations in both qualitative and quantitative measurements. Our method enables icon designers to explore diverse colorization results for a single icon using different color palettes while also efficiently evaluating the suitability of a color palette for a set of icons.
I-Chao Shen, Hsiao-Yuan Chin, Ruo-Xi Chen, Bing-Yu Chen 0004
Comput. Graph. Forum2
2023 EvIcon: Designing High-Usability Icon with Human-in-the-loop Exploration and IconCLIP
abstract
Abstract Interface icons are prevalent in various digital applications. Due to limited time and budgets, many designers rely on informal evaluation, which often results in poor usability icons. In this paper, we propose a unique human‐in‐the‐loop framework that allows our target users, that is novice and professional user interface (UI) designers, to improve the usability of interface icons efficiently. We formulate several usability criteria into a perceptual usability function and enable users to iteratively revise an icon set with an interactive design tool, EvIcon. We take a large‐scale pre‐trained joint image‐text embedding (CLIP) and fine‐tune it to embed icon visuals with icon tags in the same embedding space (IconCLIP). During the revision process, our design tool provides two types of instant perceptual usability feedback. First, we provide perceptual usability feedback modelled by deep learning models trained on IconCLIP embeddings and crowdsourced perceptual ratings. Second, we use the embedding space of IconCLIP to assist users in improving icons' visual distinguishability among icons within the user‐prepared icon set. To provide the perceptual prediction, we compiled IconCEPT10K, the first large‐scale dataset of perceptual usability ratings over 10,000 interface icons, by conducting a crowdsourcing study. We demonstrated that our framework could benefit UI designers' interface icon revision process with a wide range of professional experience. Moreover, the interface icons designed using our framework achieved better semantic distance and familiarity, verified by an additional online user study.
I-Chao Shen, Fu-Yin Cherng, Takeo Igarashi, Wen-Chieh Lin, Bing-Yu Chen 0004
Comput. Graph. Forum1
2022 StyleFaceUV: a 3D Face UV Map Generator for View-Consistent Face Image Synthesis
Wei-Chieh Chung, Jiankai Zhu, I-Chao Shen, Yu-Ting Wu 0001, Yung-Yu Chuang
BMVC3
2022 ODEN: Live Programming for Neural Network Architecture Editing
abstract
In deep learning application development, programmers tend to try different architectures and hyper-parameters until satisfied with the model performance. Nevertheless, program crashes due to tensor shape mismatch prohibit programmers, especially novice programmers, from smoothly going back and forth between neural network (NN) architecture editing and experimentation. We propose to leverage live programming techniques in NN architecture editing with an always-on visualization. When the user edits the program, the visualization can synchronously display tensor states and provide a warning message by continuously executing the program to prevent program crashes during experimentation. We implement the live visualization and integrate it into an IDE called ODEN that seamlessly supports the “edit→experiment→edit→···” repetition. With ODEN, the user can construct the neural network with the live visualization and transits into experimentation to instantly train and test the NN architecture. An exploratory user study is conducted to evaluate the usability, the limitations, and the potential of live visualization in ODEN.
Chunqi Zhao, I-Chao Shen, Tsukasa Fukusato, Jun Kato 0001, Takeo Igarashi
IUI2
2022 ClipGen: A Deep Generative Model for Clipart Vectorization and Synthesis
abstract
This article presents a novel deep learning-based approach for automatically vectorizing and synthesizing the clipart of man-made objects. Given a raster clipart image and its corresponding object category (e.g., airplanes), the proposed method sequentially generates new layers, each of which is composed of a new closed path filled with a single color. The final result is obtained by compositing all layers together into a vector clipart image that falls into the target category. The proposed approach is based on an iterative generative model that (i) decides whether to continue synthesizing a new layer and (ii) determines the geometry and appearance of the new layer. We formulated a joint loss function for training our generative model, including the shape similarity, symmetry, and local curve smoothness losses, as well as vector graphics rendering accuracy loss for synthesizing clipart recognizable by humans. We also introduced a collection of man-made object clipart, ClipNet, which is composed of closed-path layers, and two designed preprocessing tasks to clean up and enrich the original raw clipart. To validate the proposed approach, we conducted several experiments and demonstrated its ability to vectorize and synthesize various clipart categories. We envision that our generative model can facilitate efficient and intuitive clipart designs for novice users and graphic designers.
I-Chao Shen, Bing-Yu Chen 0004
IEEE Trans. Vis. Comput. Graph.1
2021 Per Garment Capture and Synthesis for Real-time Virtual Try-on
abstract
Virtual try-on is a promising application of computer graphics and human computer interaction that can have a profound real-world impact especially during this pandemic. Existing image-based works try to synthesize a try-on image from a single image of a target garment, but it inherently limits the ability to react to possible interactions. It is difficult to reproduce the change of wrinkles caused by pose and body size change, as well as pulling and stretching of the garment by hand. In this paper, we propose an alternative per garment capture and synthesis workflow to handle such rich interactions by training the model with many systematically captured images. Our workflow is composed of two parts: garment capturing and clothed person image synthesis. We designed an actuated mannequin and an efficient capturing process that collects the detailed deformations of the target garments under diverse body sizes and poses. Furthermore, we proposed to use a custom-designed measurement garment, and we captured paired images of the measurement garment and the target garments. We then learn a mapping between the measurement garment and the target garments using deep image-to-image translation. The customer can then try on the target garments interactively during online shopping. The proposed workflow requires certain manual labor, but we believe that the cost is acceptable given that the retailers are already paying significant costs for hiring professional photographers and models, stylists, and editors to take photographs for promotion. Our method can remove the need of hiring these costly professionals. We evaluated the effectiveness of the proposed system with ablation studies and quality comparison with previous virtual try-on methods. We perform a user study to show our promising virtual try-on performances. Moreover, we also demonstrate that we use our method for changing virtual costumes in video conferences. Finally, we provide the collected dataset as the cloth dataset parameterized by various viewing angles, body poses, and sizes.
Toby Chong, I-Chao Shen, Nobuyuki Umetani, Takeo Igarashi
UIST2
2021 Interactive Optimization of Generative Image Modelling using Sequential Subspace Search and Content-based Guidance
abstract
Abstract Generative image modeling techniques such as GAN demonstrate highly convincing image generation result. However, user interaction is often necessary to obtain desired results. Existing attempts add interactivity but require either tailored architectures or extra data. We present a human‐in‐the‐optimization method that allows users to directly explore and search the latent vector space of generative image modelling. Our system provides multiple candidates by sampling the latent vector space, and the user selects the best blending weights within the subspace using multiple sliders. In addition, the user can express their intention through image editing tools. The system samples latent vectors based on inputs and presents new candidates to the user iteratively. An advantage of our formulation is that one can apply our method to arbitrary pre‐trained model without developing specialized architecture or data. We demonstrate our method with various generative image modelling applications, and show superior performance in a comparative user study with prior art iGAN [ZKSE16].
Toby Chong, I-Chao Shen, Issei Sato, Takeo Igarashi
Comput. Graph. Forum2
2021 ClipFlip : Multi-view Clipart Design
abstract
Abstract We present an assistive system for clipart design by providing visual scaffolds from the unseen viewpoints. Inspired by the artists' creation process, our system constructs the visual scaffold by first synthesizing the reference 3D shape of the input clipart and rendering it from the desired viewpoint. The critical challenge of constructing this visual scaffold is to generate a reference 3D shape that matches the user's expectations in terms of object sizing and positioning while preserving the geometric style of the input clipart. To address this challenge, we propose a user‐assisted curve extrusion method to obtain the reference 3D shape. We render the synthesized reference 3D shape with a consistent style into the visual scaffold. By following the generated visual scaffold, the users can efficiently design clipart with their desired viewpoints. The user study conducted by an intuitive user interface and our generated visual scaffold suggests that our system is especially useful for estimating the ratio and scale between object parts and can save on average 57% of drawing time.
I-Chao Shen, Kuan-Hung Liu, Li-Wen Su, Yu-Ting Wu 0001, Bing-Yu Chen 0004
Comput. Graph. Forum1
2020 Director-360: Introducing Camera Handling to 360 Cameras
abstract
This work introduces the concept of camera handling for a 360 camera and proposes Director-360, a 360 camera enhanced with two novel handling techniques. Pointer and field-of-view (FoV) are designed to explicitly and implicitly capture the 360 photographer’s subject of interest within the 360 media at capture time. Pointer lets users specify a subject of interest about the scene by directly pointing the 360 camera as if using the camera as a flashlight, while FoV captures the user’s subject of interest within the 360 scene by mapping the user’s face direction to the 360 media. We described an implementation using deep-learning algorithms. We also presented the Director-360 Editor, which incorporates the handling data to streamline the post-editing process. To understand how Director-360 helps to compose 360 media, a pilot study was carried out to create video storytelling in three target scenarios. The results and user feedback from a pilot study were reported.
Hao-Juan Huang, I-Chao Shen, Li-Wei Chan 0001
MobileHCI2
2020 ZomeFab: Cost-Effective Hybrid Fabrication with Zometools
abstract
Abstract In recent years, personalized fabrication has received considerable attention because of the widespread use of consumer‐level three‐dimensional (3D) printers. However, such 3D printers have drawbacks, such as long production time and limited output size, which hinder large‐scale rapid‐prototyping. In this paper, for the time‐ and cost‐effective fabrication of large‐scale objects, we propose a hybrid 3D fabrication method that combines 3D printing and the Zometool construction set, which is a compact, sturdy and reusable structure for infill fabrication. The proposed method significantly reduces fabrication cost and time by printing only thin 3D outer shells. In addition, we design an optimization framework to generate both a Zometol structure and printed surface partitions by optimizing several criteria, including printability, material cost and Zometool structure complexity. Moreover, we demonstrate the effectiveness of the proposed method by fabricating various large‐scale 3D models.
I-Chao Shen, Ming-Shiuan Chen, Bing-Yu Chen 0004
Comput. Graph. Forum1
2018 Perception-driven semi-structured boundary vectorization
abstract
Artist-drawn images with distinctly colored, piecewise continuous boundaries, which we refer to as semi-structured imagery , are very common in online raster databases and typically allow for a perceptually unambiguous mental vector interpretation. Yet, perhaps surprisingly, existing vectorization algorithms frequently fail to generate these viewer-expected interpretations on such imagery. In particular, the vectorized region boundaries they produce frequently diverge from those anticipated by viewers. We propose a new approach to region boundary vectorization that targets semi-structured inputs and leverages observations about human perception of shapes to generate vector images consistent with viewer expectations. When viewing raster imagery observers expect the vector output to be an accurate representation of the raster input. However, perception studies suggest that viewers implicitly account for the lossy nature of the rasterization process and mentally smooth and simplify the observed boundaries. Our core algorithmic challenge is to balance these conflicting cues and obtain a piecewise continuous vectorization whose discontinuities, or corners, are aligned with human expectations. Our framework centers around a simultaneous spline fitting and corner detection method that combines a learned metric, that approximates human perception of boundary discontinuities on raster inputs, with perception-driven algorithmic discontinuity analysis. The resulting method balances local cues provided by the learned metric with global cues obtained by balancing simplicity and continuity expectations. Given the finalized set of corners, our framework connects those using simple, continuous curves that capture input regularities. We demonstrate our method on a range of inputs and validate its superiority over existing alternatives via an extensive comparative user study.
Shayan Hoshyari, Edoardo A. Dominici, Alla Sheffer, Nathan Carr 0001, Duygu Ceylan, I-Chao Shen
ACM Trans. Graph.7
2017 High-resolution 360 Video Foveated Stitching for Real-time VR
abstract
Abstract In virtual reality (VR) applications, the contents are usually generated by creating a 360° Video panorama of a real‐world scene. Although many capture devices are being released, getting high‐resolution panoramas and displaying a virtual world in real‐time remains challenging due to its computationally demanding nature. In this paper, we propose a real‐time 360° Video foveated stitching framework, that renders the entire scene in different level of detail, aiming to create a high‐resolution panoramic Video in real‐time that can be streamed directly to the client. Our foveated stitching algorithm takes Videos from multiple cameras as input, combined with measurements of human visual attention (i.e. the acuity map and the saliency map), can greatly reduce the number of pixels to be processed. We further parallelize the algorithm using GPU to achieve a responsive interface and validate our results via a user study. Our system accelerates graphics computation by a factor of 6 on a Google Cardboard display.
Wei-Tse Lee, Hsin-I Chen, Ming-Shiuan Chen, I-Chao Shen, Bing-Yu Chen 0004
Comput. Graph. Forum4
2016 Retargeting 3D Objects and Scenes with a General Framework
abstract
Abstract In this paper, we introduce an interactive method suitable for retargeting both 3D objects and scenes. Initially, the input object or scene is decomposed into a collection of constituent components enclosed by corresponding control bounding volumes which capture the intra‐structures of the object or semantic grouping of objects in the 3D scene. The overall retargeting is accomplished through a constrained optimization by manipulating the control bounding volumes. Without inferring the intricate dependencies between the components, we define a minimal set of constraints that maintain the spatial arrangement and connectivity between the components to regularize the valid retargeting results. The default retargeting behavior can then be easily altered by additional semantic constraints imposed by users. This strategy makes the proposed method highly flexible to process a wide variety of 3D objects and scenes under an unified framework. In addition, the proposed method achieved more general structure‐preserving pattern synthesis in both object and scene levels. We demonstrate the effectiveness of our method by applying it to several complicated 3D objects and scenes.
Yi-Ling Chen 0004, I-Chao Shen, Bing-Yu Chen 0004
Comput. Graph. Forum3
2016 Photo sundial: Estimating the time of capture in consumer photos
Tsung-Hung Tsai, Wei-Cih Jhou, Wen-Huang Cheng, Min-Chun Hu 0001, I-Chao Shen, Tekoing Lim, Kai-Lung Hua, Ahmed Ghoneim, M. Anwar Hossain 0001, Shintami Chusnul Hidayati
Neurocomputing5
2016 A scalable active framework for region annotation in 3D shape collections
abstract
Large repositories of 3D shapes provide valuable input for data-driven analysis and modeling tools. They are especially powerful once annotated with semantic information such as salient regions and functional parts. We propose a novel active learning method capable of enriching massive geometric datasets with accurate semantic region annotations. Given a shape collection and a user-specified region label our goal is to correctly demarcate the corresponding regions with minimal manual work. Our active framework achieves this goal by cycling between manually annotating the regions, automatically propagating these annotations across the rest of the shapes, manually verifying both human and automatic annotations, and learning from the verification results to improve the automatic propagation algorithm. We use a unified utility function that explicitly models the time cost of human input across all steps of our method. This allows us to jointly optimize for the set of models to annotate and for the set of models to verify based on the predicted impact of these actions on the human efficiency. We demonstrate that incorporating verification of all produced labelings within this unified objective improves both accuracy and efficiency of the active learning procedure. We automatically propagate human labels across a dynamic shape network using a conditional random field (CRF) framework, taking advantage of global shape-to-shape similarities, local feature similarities, and point-to-point correspondences. By combining these diverse cues we achieve higher accuracy than existing alternatives. We validate our framework on existing benchmarks demonstrating it to be significantly more efficient at using human input compared to previous techniques. We further validate its efficiency and robustness by annotating a massive shape dataset, labeling over 93,000 shape parts, across multiple model classes, and providing a labeled part collection more than one order of magnitude larger than existing ones.
Li Yi 0001, Vladimir G. Kim, Duygu Ceylan, I-Chao Shen, Mengyuan Yan, Hao Su 0001, Cewu Lu, Qixing Huang, Alla Sheffer, Leonidas J. Guibas
ACM Trans. Graph.4
2015 WonderLens: Optical Lenses and Mirrors for Tangible Interactions on Printed Paper
abstract
This work presents WonderLens, a system of optical lenses and mirrors for enabling tangible interactions on printed paper. When users perform spatial operations on the optical components, they deform the visual content that is printed on paper, and thereby provide dynamic visual feedback on user interactions without any display devices. The magnetic unit that is embedded in each lens and mirror allows the unit to be identified and tracked using an analog Hall-sensor grid that is placed behind the paper, so the system provides additional auditory and visual feedback through different levels of embodiment, further enhancing the interactivity with the printed content on the physical paper.
Rong-Hao Liang, I-Chao Shen, Yu-Chien Chan, Guan-Ting Chou, Li-Wei Chan 0001, De-Nian Yang, Mike Y. Chen, Bing-Yu Chen 0004
CHI2
2015 Data-driven Handwriting Synthesis in a Conjoined Manner
abstract
A person's handwriting appears differently within a typical range of variations, and the shapes of handwriting characters also show complex interaction with their nearby neighbors. This makes automatic synthesis of handwriting characters and paragraphs very challenging. In this paper, we propose a method for synthesizing handwriting texts according to a writer's handwriting style. The synthesis algorithm is composed by two phases. First, we create the multidimensional morphable models for different characters based on one writer's data. Then, we compute the cursive probability to decide whether each pair of neighboring characters are conjoined together or not. By jointly modeling the handwriting style and conjoined property through a novel trajectory optimization, final handwriting words can be synthesized from a set of collected samples. Furthermore, the paragraphs’ layouts are also automatically generated and adjusted according to the writer's style obtained from the same dataset. We demonstrate that our method can successfully synthesize an entire paragraph that mimic a writer's handwriting using his/her collected handwriting samples.
Hsin-I Chen, Tse-Ju Lin, Xiao-Feng Jian, I-Chao Shen, Bing-Yu Chen 0004
Comput. Graph. Forum4
2015 Gestalt Rule Feature Points
abstract
As the large online repositories of image and video data has emerged and continued to grow in number, the visual variations in such repositories has also increased dramatically. For example, the visual scene of a photograph can be changed into different colors by image editing tools or depicted by multiple representations, such as a painting and a hand-drawn sketch. The large visual variations tend to cause ambiguities for the existing computer vision algorithms to recognize the visual analogies of these images and often limit the potential of related applications. In this paper, therefore, we propose a new approach for detecting reliable visual features from images, with a particular focus on improving the repeatability of the local features in those images containing the same semantic contents (e.g., a landmark) but in different visual styles (e.g., a photo and a painting). We proposed a novel method for establishing visual correspondences between images based on the Gestalt theory, a psychological study of how human visions organize the visual perception. Experiments demonstrated the outperformance of our approach over the state-of-the-art local features in various computer vision tasks, such as cross domain image matching and retrieval.
I-Chao Shen, Wen-Huang Cheng
IEEE Trans. Multim.1
2015 Geometrically Consistent Stereoscopic Image Editing Using Patch-Based Synthesis
abstract
This paper presents a patch-based synthesis framework for stereoscopic image editing. The core of the proposed method builds upon a patch-based optimization framework with two key contributions: First, we introduce a depth-dependent patch-pair similarity measure for distinguishing and better utilizing image contents with different depth structures. Second, a joint patch-pair search is proposed for properly handling the correlation between two views. The proposed method successfully overcomes two main challenges of editing stereoscopic 3D media: (1) maintaining the depth interpretation, and (2) providing controllability of the scene depth. The method offers patch-based solutions to a wide variety of stereoscopic image editing problems, including depth-guided texture synthesis, stereoscopic NPR, paint by depth, content adaptation, and 2D to 3D conversion. Several challenging cases are demonstrated to show the effectiveness of the proposed method. The results of user studies also show that the proposed method produces stereoscopic images with good stereoscopics and visual quality.
Sheng-Jie Luo, Ying-Tse Sun, I-Chao Shen, Bing-Yu Chen 0004, Yung-Yu Chuang
IEEE Trans. Vis. Comput. Graph.3
2013 Artistic eye: recognizing key viewing points of popular sites
abstract
No abstract available.
Chih-Hsiang Hsu, I-Chao Shen, Wen-Huang Cheng, Shih-Wei Sun
MobiSys2
2013 Stroke-guided Image Synthesis for Skeletal Structure Editing
abstract
Abstract Creating variations of an image object is an important task, which usually requires manipulating the skeletal structure of the object. However, most existing methods (such as image deformation) only allow for stretching the skeletal structure of an object: modifying skeletal topology remains a challenge. This paper presents a technique for synthesizing image objects with different skeletal structures while respecting to an input image object. To apply this technique, a user firstly annotates the skeletal structure of the input object by specifying a number of strokes in the input image, and draws corresponding strokes in an output domain to generate new skeletal structures. Then, a number of the example texture pieces are sampled along the strokes in the input image and pasted along the strokes in the output domain with their orientations. The result is obtained by optimizing the texture sampling and seam computation. The proposed method is successfully used to synthesize challenging skeletal structures, such as skeletal branches, and a wide range of image objects with various skeletal structures, to demonstrate its effectiveness.
Sheng-Jie Luo, Chin-Yu Lin, I-Chao Shen, Bing-Yu Chen 0004
Comput. Graph. Forum3
2012 Perspective-aware warping for seamless stereoscopic image cloning
abstract
This paper presents a novel technique for seamless stereoscopic image cloning, which performs both shape adjustment and color blending such that the stereoscopic composite is seamless in both the perceived depth and color appearance. The core of the proposed method is an iterative disparity adaptation process which alternates between two steps: disparity estimation, which re-estimates the disparities in the gradient domain so that the disparities are continuous across the boundary of the cloned region; and perspective-aware warping, which locally re-adjusts the shape and size of the cloned region according to the estimated disparities. This process guarantees not only depth continuity across the boundary but also models local perspective projection in accordance with the disparities, leading to more natural stereoscopic composites. The proposed method allows for easy cloning of objects with intricate silhouettes and vague boundaries because it does not require precise segmentation of the objects. Several challenging cases are demonstrated to show that our method generates more compelling results compared to methods with only global shape adjustment.
Sheng-Jie Luo, I-Chao Shen, Bing-Yu Chen 0004, Wen-Huang Cheng, Yung-Yu Chuang
ACM Trans. Graph.2