VLDB 2026 Research / reviewers in the wild / expert
Xueting Liu 0001
dblp:83/6609
· DBLP profile ↗
43ranked-venue papers
6as first author
27since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 42 · 6 first-author · 26 since 2021Artificial intelligence and machine learning · 5 · 1 first-author · 5 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 since 2021Computer networks · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Line Drawing Abstraction Based on Line Importance Evaluation
Shilong Deng, Xueting Liu 0001, Chengze Li, Ping Li 0016, Zhenkun Wen, Huisi Wu |
CGI (1) | 2 |
| 2025 | Cartoon Animation Shading Removal
Zhenhua Ou, Chengze Li, Xueting Liu 0001, Zhenkun Wen, Huisi Wu |
CGI (1) | 3 |
| 2025 | Robust Character Stroke Segmentation For Diverse Fonts Via Contour Matching and Chain PropagationabstractStroke segmentation is a fundamental technique for various character analysis and synthesis applications. However, existing methods often face challenges such as over-segmentation, under-segmentation, low segmentation accuracy, and limited generalization capability when segmenting characters of diverse fonts. To address these issues, we propose a novel stroke segmentation method based on contour matching and utilize a similarity-based chain propagation strategy to tackle the challenges posed by fonts with significant structural and stylistic variations. Extensive visual and quantitative experiments on a newly created high-quality dataset demonstrate that our approach outperforms state-of-the-art methods and effectively handles a wide range of fonts. Xueting Liu 0001, Chengze Li, Zhenkun Wen, Huisi Wu |
ICIP | 2 |
| 2025 | Screentone-Preserved Manga RetargetingabstractAbstract As a popular comic style, manga offers a unique impression by utilizing a rich set ofbitonal patterns, or screentones, for illustration. However, screentones can easily be degraded when manga is resized in terms of aspect ratio and resolution for manga re‐layout and e‐manga migration applications. To tackle this problem, we propose the first automatic manga retargeting method that synthesizes a retargeted manga image while preserving the prominent structure and fine screentone intended by the manga artist. While modern natural photo retargeting methods can achieve prominent structure preservation, preserving screentones within arbitrarily shaped regions is very challenging due to two properties of manga: (i) pattern constancy under translation, and (ii) non‐compatibility with interpolation. To circumvent this barrier, we propose learning a quantized representation of screentones that is translation‐invariant and pointwisely representable through a tailored manga reconstruction network with a screentone‐anchored codebook. Thanks to these merits, we can perform the re‐synthesis operation using existing photo retargeting methods and achieve the desired manga retargeting results. We conducted extensive qualitative and quantitative experiments to validate the effectiveness of our method, and we achieved notably compelling results compared to alternative methods. Minshan Xie, Menghan Xia, Chengze Li, Xueting Liu 0001, Tien-Tsin Wong |
Comput. Graph. Forum | 4 |
| 2025 | Cartoon Animation Outpainting With Region-Guided Motion InferenceabstractCartoon animation video is a popular visual entertainment form worldwide, however many classic animations were produced in a 4:3 aspect ratio that is incompatible with modern widescreen displays. Existing methods like cropping lead to information loss while retargeting causes distortion. Animation companies still rely on manual labor to renovate classic cartoon animations, which is tedious and labor-intensive, but can yield higher-quality videos. Conventional extrapolation or inpainting methods tailored for natural videos struggle with cartoon animations due to the lack of textures in anime, which affects the motion estimation of the objects. In this article, we propose a novel framework designed to automatically outpaint 4:3 anime to 16:9 via region-guided motion inference. Our core concept is to identify the motion correspondences between frames within a sequence in order to reconstruct missing pixels. Initially, we estimate optical flow guided by region information to address challenges posed by exaggerated movements and solid-color regions in cartoon animations. Subsequently, frames are stitched to produce a pre-filled guide frame, offering structural clues for the extension of optical flow maps. Finally, a voting and fusion scheme utilizes learned fusion weights to blend the aligned neighboring reference frames, resulting in the final outpainting frame. Extensive experiments confirm the superiority of our approach over existing methods. Huisi Wu, Chengze Li, Xueting Liu 0001, Zhenkun Wen, Tong-Yee Lee |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2025 | Instance-guided anime editing with a curated large-scale dataset
Chengze Li, Xueting Liu 0001, Zhongping Ge |
Vis. Comput. | 3 |
| 2024 | SKETCH2MANGA: Shaded Manga Screening from Sketch with Diffusion ModelsabstractWhile manga is a popular entertainment form, creating manga is tedious, especially adding screentones to the created sketch, namely manga screening. Unfortunately, there is no existing method that tailors for automatic manga screening, probably due to the difficulty in generating shaded high-frequency screentones of high-quality. Classic manga screening approaches generally require user input to provide screentone exemplars or a reference manga image. Recent deep learning models enable automatic generation by learning from a large-scale dataset. However, the state-of-the-art models still fail to generate high-quality shaded screentones due to the lack of a tailored model and high-quality manga training data. In this paper, we propose a novel sketch-to-manga framework that first generates a color illustration from the sketch and then generates a screentoned manga based on the intensity guidance. Our method significantly outperforms existing methods in generating high-quality manga with shaded high-frequency screentones. Xueting Liu 0001, Chengze Li, Minshan Xie, Tien-Tsin Wong |
ICIP | 2 |
| 2024 | Suitable and Style-Consistent Multi-Texture Recommendation for Cartoon IllustrationsabstractTexture plays an important role in cartoon illustrations to display object materials and enrich visual experiences. Unfortunately, manually designing and drawing an appropriate texture is not easy even for proficient artists, let alone novice or amateur people. While there exist tons of textures on the Internet, it is not easy to pick an appropriate one using traditional text-based search engines. Although several texture pickers have been proposed, they still require the users to browse the textures by themselves, which is still labor-intensive and time-consuming. In this article, an automatic texture recommendation system is proposed for recommending multiple textures to replace a set of user-specified regions in a cartoon illustration with visually pleasant look. Two measurements, the suitability measurement and the style-consistency measurement, are proposed to make sure that the recommended textures are suitable for cartoon illustration and at the same time mutually consistent in style. The suitability is measured based on the synthesizability, cartoonity, and region fitness of textures. The style-consistency is predicted using a learning-based solution since it is subjective to judge whether two textures are consistent in style. An optimization problem is formulated and solved via the genetic algorithm. Our method is validated on various cartoon illustrations, and convincing results are obtained. Huisi Wu, Zhaoze Wang, Xueting Liu 0001, Tong-Yee Lee |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2024 | AnimeDiffusion: Anime Diffusion ColorizationabstractBeing essential in animation creation, colorizing anime line drawings is usually a tedious and time-consuming manual task. Reference-based line drawing colorization provides an intuitive way to automatically colorize target line drawings using reference images. The prevailing approaches are based on generative adversarial networks (GANs), yet these methods still cannot generate high-quality results comparable to manually-colored ones. In this article, a new AnimeDiffusion approach is proposed via hybrid diffusions for the automatic colorization of anime face line drawings. This is the first attempt to utilize the diffusion model for reference-based colorization, which demands a high level of control over the image synthesis process. To do so, a hybrid end-to-end training strategy is designed, including phase 1 for training diffusion model with classifier-free guidance and phase 2 for efficiently updating color tone with a target reference colored image. The model learns denoising and structure-capturing ability in phase 1, and in phase 2, the model learns more accurate color information. Utilizing our hybrid training strategy, the network convergence speed is accelerated, and the colorization performance is improved. Our AnimeDiffusion generates colorization results with semantic correspondence and color consistency. In addition, the model has a certain generalization performance for line drawings of different line styles. To train and evaluate colorization methods, an anime face line drawing colorization benchmark dataset, containing 31,696 training data and 579 testing data, is introduced and shared. Extensive experiments and user studies have demonstrated that our proposed AnimeDiffusion outperforms state-of-the-art GAN-based methods and another diffusion-based model, both quantitatively and qualitatively. Yu Cao 0019, Xiangqiao Meng, P. Y. Mok 0001, Tong-Yee Lee, Xueting Liu 0001, Ping Li 0016 |
IEEE Trans. Vis. Comput. Graph. | 5 |
| 2024 | Separating Shading and Reflectance From Cartoon IllustrationsabstractShading plays an important role in cartoon drawings to present the 3D lighting and depth information in a 2D image to improve the visual information and pleasantness. But it also introduces apparent challenges in analyzing and processing the cartoon drawings for different computer graphics and vision applications, such as segmentation, depth estimation, and relighting. Extensive research has been made in removing or separating the shading information to facilitate these applications. Unfortunately, the existing researches only focused on natural images, which are natively different from cartoons since the shading in natural images is physically correct and can be modeled based on physical priors. However, shading in cartoons is manually created by artists, which may be imprecise, abstract, and stylized. This makes it extremely difficult to model the shading in cartoon drawings. Without modeling the shading prior, in the paper, we propose a learning-based solution to separate the shading from the original colors using a two-branch system consisting of two subnetworks. To the best of our knowledge, our method is the first attempt in separating shading information from cartoon drawings. Our method significantly outperforms the methods tailored for natural images. Extensive evaluations have been performed with convincing results in all cases. Ziheng Ma, Chengze Li, Xueting Liu 0001, Huisi Wu, Zhenkun Wen |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2024 | Shading-Guided Manga Screening From ReferenceabstractManga screening is a critical process in manga production, which still requires intensive labor and cost. Existing manga screening methods either generate simple dotted screentones only or rely on color information and manual hints during screentone selection. Due to the large domain gap between line drawings and screened manga, and the difficulties in generating high-quality, properly selected and shaded screentones, even state-of-the-art deep learning methods cannot convert line drawings to screened manga well. Besides, ambiguity exists in the screening process since different artists may screen differently for the same line drawing. In this article, we propose to introduce shaded line drawing as the intermediate counterpart of the screened manga so that the manga screening task can be decomposed into two sub-tasks, generating shading from a line drawing and replacing shading with proper screentones. The reference image is adopted to resolve the ambiguity issue and provides options and controls on the generated screened manga. We proposed a reference-based shading generation network and a reference-based screentone generation module to achieve the two sub-tasks individually. We conduct extensive visual and quantitative experiments to verify the effectiveness of our system. Results and statistics show that our method outperforms existing methods on the manga screening task. Huisi Wu, Ziheng Ma, Wenliang Wu, Xueting Liu 0001, Chengze Li, Zhenkun Wen |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2024 | Appearance-Preserved Portrait-to-Anime Translation via Proxy-Guided Domain AdaptationabstractConverting a human portrait to anime style is a desirable but challenging problem. Existing methods fail to resolve this problem due to the large inherent gap between two domains that cannot be overcome by a simple direct mapping. For this reason, these methods struggle to preserve the appearance features in the original photo. In this article, we discover an intermediate domain, the coser portrait (portraits of humans costuming as anime characters), that helps bridge this gap. It alleviates the learning ambiguity and loosens the mapping difficulty in a progressive manner. Specifically, we start from learning the mapping between coser and anime portraits, and present a proxy-guided domain adaptation learning scheme with three progressive adaptation stages to shift the initial model to the human portrait domain. In this way, our model can generate visually pleasant anime portraits with well-preserved appearances given the human portrait. Our model adopts a disentangled design by breaking down the translation problem into two specific subtasks of face deformation and portrait stylization. This further elevates the generation quality. Extensive experimental results show that our model can achieve visually compelling translation with better appearance preservation and perform favorably against the existing methods both qualitatively and quantitatively. Our code and datasets are available at https://github.com/NeverGiveU/PDA-Translation. Wenpeng Xiao, Jiajie Mai, Xuemiao Xu, Chengze Li, Xueting Liu 0001, Shengfeng He |
IEEE Trans. Vis. Comput. Graph. | 7 |
| 2023 | Panel-Page-Aware Comic Genre UnderstandingabstractUsing a sequence of discrete still images to tell a story or introduce a process has become a tradition in the field of digital visual media. With the surge in these media and the requirements in downstream tasks, acquiring their main topics or genres in a very short time is urgently needed. As a representative form of the media, comic enjoys a huge boom as it has gone digital. However, different from natural images, comic images are divided by panels, and the images are not visually consistent from page to page. Therefore, existing works tailored for natural images perform poorly in analyzing comics. Considering the identification of comic genres is tied to the overall story plotting, a long-term understanding that makes full use of the semantic interactions between multi-level comic fragments needs to be fully exploited. In this paper, we propose [Formula: see text]Comic, a Panel-Page-aware Comic genre classification model, which takes page sequences of comics as the input and produces class-wise probabilities. [Formula: see text]Comic utilizes detected panel boxes to extract panel representations and deploys self-attention to construct panel-page understanding, assisted with interdependent classifiers to model label correlation. We develop the first comic dataset for the task of comic genre classification with multi-genre labels. Our approach is proved by experiments to outperform state-of-the-art methods on related tasks. We also validate the extensibility of our network to perform in the multi-modal scenario. Finally, we show the practicability of our approach by giving effective genre prediction results for whole comic books. Chenshu Xu, Xuemiao Xu, Nanxuan Zhao, Huaidong Zhang, Chengze Li, Xueting Liu 0001 |
IEEE Trans. Image Process. | 7 |
| 2023 | Multi-Scale Flow-Based Occluding Effect and Content Separation for Cartoon AnimationsabstractOccluding effects have been frequently used to present weather conditions and environments in cartoon animations, such as raining, snowing, moving leaves, and moving petals. While these effects greatly enrich the visual appeal of the cartoon animations, they may also cause undesired occlusions on the content area, which significantly complicate the analysis and processing of the cartoon animations. In this article, we make the first attempt to separate the occluding effects and content for cartoon animations. The major challenge of this problem is that, unlike natural effects that are realistic and small-sized, the effects of cartoons are usually stylistic and large-sized. Besides, effects in cartoons are manually drawn, so their motions are more unpredictable than realistic effects. To separate occluding effects and content for cartoon animations, we propose to leverage the difference in the motion patterns of the effects and the content, and capture the locations of the effects based on a multi-scale flow-based effect prediction (MFEP) module. A dual-task learning system is designed to extract the effect video and reconstruct the effect-removed content video at the same time. We apply our method on a large number of cartoon videos of different content and effects. Experiments show that our method significantly outperforms the existing methods. We further demonstrate how the separated effects and content facilitate the analysis and processing of cartoon videos through different applications, including segmentation, inpainting, and effect migration. Xuemiao Xu, Xueting Liu 0001 |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2023 | AddCR: a data-driven cartoon remastering
Yinghua Liu, Chengze Li, Xueting Liu 0001, Huisi Wu, Zhenkun Wen |
Vis. Comput. | 3 |
| 2022 | End-to-End Line Drawing VectorizationabstractVector graphics is broadly used in a variety of forms, such as illustrations, logos, posters, billboards, and printed ads. Despite its broad use, many artists still prefer to draw with pen and paper, which leads to a high demand of converting raster designs into the vector form. In particular, line drawing is a primary art and attracts many research efforts in automatically converting raster line drawings to vector form. However, the existing methods generally adopt a two-step approach, stroke segmentation and vectorization. Without vector guidance, the raster-based stroke segmentation frequently obtains unsatisfying segmentation results, such as over-grouped strokes and broken strokes. In this paper, we make an attempt in proposing an end-to-end vectorization method which directly generates vectorized stroke primitives from raster line drawing in one step. We propose a Transformer-based framework to perform stroke tracing like human does in an automatic stroke-by-stroke way with a novel stroke feature representation and multi-modal supervision to achieve vectorization with high quality and fidelity. Qualitative and quantitative evaluations show that our method achieves state of the art performance. Hanyuan Liu, Chengze Li, Xueting Liu 0001, Tien-Tsin Wong |
AAAI | 3 |
| 2022 | Left and Right Ventricular Segmentation Based on 3D Region-Aware U-NetabstractThe cardiac is one of the essential organs, and the segmentation of the left and right ventricular of cardiac is essential in diagnosing various heart diseases. The most popular method for the segmentation of 3D MRI images is the nnUNet. However, the 3D MRI volume of the ventricular contains other organs which interfere with the segmentation of the ventricular. Hence, we proposed a novel region-aware U-Net segmentation method RegUNet for ventricular segmentation. RegUNet improves the ventricular's segmentation performance by first capturing the region of interest (RoI) of the ventricular and then segmenting the ventricular with the captured RoI features, which reduces the segmentation module's difficulty by keeping the cardiac's features and leaving others such that RegUNet can focus on ventricular segmentation. Besides, since the model segments the ventricular with the captured RoI features, it saves the model's computing resources from identifying the background of the volume. Since 3D cardiac MRI volumes scanned by the different devices have diverse statistical characteristics, which causes the model's performance in processing the multi-source cardiac volumes to be unstable. We stabilize the model's performance with a multi-sources feature normalization strategy, which normalizes the feature from a different source with different parameters. We validated the proposed method on the M&MS dataset, a multi-sources 3D MRI cardiac segmentation dataset. Experiments showed that RegUNet's segmentation ability reached the state-of-the-art. Xueting Liu 0001, Huisi Wu, Zhenkun Wen, LinLin Shen |
CBMS | 3 |
| 2022 | Authenticity Identification of Qi Baishi's Shrimp Painting with Dynamic Token Enhanced Visual Transformer
Xueting Liu 0001, Huisi Wu, Fu Qi |
CGI | 3 |
| 2022 | Neural Recognition of Dashed Curves with Gestalt Law of ContinuityabstractDashed curve is a frequently used curve form and is widely used in various drawing and illustration applications. While humans can intuitively recognize dashed curves from disjoint curve segments based on the law of continuity in Gestalt psychology, it is extremely difficult for computers to model the Gestalt law of continuity and recognize the dashed curves since high-level semantic understanding is needed for this task. The various appear-ances and styles of the dashed curves posed on a potentially noisy background further complicate the task. In this paper, we propose an innovative Transformer-based framework to recognize dashed curves based on both high-level features and low-level clues. The framework manages to learn the computational analogy of the Gestalt Law in various do-mains to locate and extract instances of dashed curves in both raster and vector representations. Qualitative and quantitative evaluations demonstrate the efficiency and ro-bustness of our framework over all existing solutions. Hanyuan Liu, Chengze Li, Xueting Liu 0001, Tien-Tsin Wong |
CVPR | 3 |
| 2022 | Vectorizing Line Drawings of Arbitrary Thickness via Boundary-based Topology ReconstructionabstractAbstract Vectorization is a commonly used technique for converting raster images to vector format and has long been a research focus in computer graphics and vision. While a number of attempts have been made to extract the topology of line drawings and further convert them to vector representations, the existing methods commonly focused on resolving junctions composed of thin lines. They usually fail for line drawings composed of thick lines, especially at junctions. In this paper, we propose an automatic line drawing vectorization method that can reconstruct the topology of line drawings of arbitrary thickness. Our key observation is that no matter the lines are thin or thick, the boundaries of the lines always provide reliable hints for reconstructing the topology. For example, the boundaries of two continuous line segments at a junction are usually smoothly connected. By analyzing the continuity of boundaries, we can better analyze the topology at junctions. In particular, we first extract the skeleton of the input line drawing via thinning. Then we analyze the reliability of the skeleton points based on boundaries. Reliable skeleton points are preserved while unreliable skeleton points are reconstructed based on boundaries again. Finally, the skeleton after reconstruction is vectorized as the output. We apply our method on line drawings of various contents and styles. Satisfying results are obtained. Our method significantly outperforms existing methods for line drawings composed of thick lines. Xueting Liu 0001, Chengze Li, Huisi Wu, Zhenkun Wen |
Comput. Graph. Forum | 2 |
| 2022 | Reference-guided structure-aware deep sketch colorization for cartoonsabstractDigital cartoon production requires extensive manual labor to colorize sketches with visually pleasant color composition and color shading. During colorization, the artist usually takes an existing cartoon image as color guidance, particularly when colorizing related characters or an animation sequence. Reference-guided colorization is more intuitive than colorization with other hints, such as color points or scribbles, or textbased hints. Unfortunately, reference-guided colorization is challenging since the style of the colorized image should match the style of the reference image in terms of both global color composition and local color shading. In this paper, we propose a novel learning-based framework which colorizes a sketch based on a color style feature extracted from a reference color image. Our framework contains a color style extractor to extract the color feature from a color image, a colorization network to generate multi-scale output images by combining a sketch and a color feature, and a multi-scale discriminator to improve the reality of the output image. Extensive qualitative and quantitative evaluations show that our method outperforms existing methods, providing both superior visual quality and style reference consistency in the task of reference-based colorization. Xueting Liu 0001, Wenliang Wu, Chengze Li, Huisi Wu |
Comput. Vis. Media | 1 |
| 2021 | Deep Style Transfer for Line DrawingsabstractLine drawings are frequently used to illustrate ideas and concepts in digital documents and presentations. To compose a line drawing, it is common for users to retrieve multiple line drawings from the Internet and combine them as one image. However, different line drawings may have different line styles and are visually inconsistent when put together. In order that the line drawings can have consistent looks, in this paper, we make the first attempt to perform style transfer for line drawings. The key of our design lies in the fact that centerline plays a very important role in preserving line topology and extracting style features. With this finding, we propose to formulate the style transfer problem as a centerline stylization problem and solve it via a novel style-guided image-to-image translation network. Results and statistics show that our method significantly outperforms the existing methods both visually and quantitatively. Xueting Liu 0001, Wenliang Wu, Huisi Wu, Zhenkun Wen |
AAAI | 1 |
| 2021 | Deep Halftoning with Reversible Binary PatternabstractExisting halftoning algorithms usually drop colors and fine details when dithering color images with binary dot patterns, which makes it extremely difficult to recover the original information. To dispense the recovery trouble in future, we propose a novel halftoning technique that converts a color image into binary halftone with full restorability to the original version. The key idea is to implicitly embed those previously dropped information into the halftone patterns. So, the halftone pattern not only serves to reproduce the image tone, maintain the blue-noise randomness, but also represents the color information and fine details. To this end, we exploit two collaborative convolutional neural networks (CNNs) to learn the dithering scheme, under a nontrivial self-supervision formulation. To tackle the flatness degradation issue of CNNs, we propose a novel noise incentive block (NIB) that can serve as a generic CNN plug-in for performance promotion. At last, we tailor a guiding-aware training scheme that secures the convergence direction as regulated. We evaluate the invertible halftones in multiple aspects, which evidences the effectiveness of our method. Menghan Xia, Wenbo Hu 0002, Xueting Liu 0001, Tien-Tsin Wong |
ICCV | 3 |
| 2021 | Deep texture cartoonization via unsupervised appearance regularization
Huisi Wu, Xueting Liu 0001, Chengze Li, Wenliang Wu |
Comput. Graph. | 3 |
| 2021 | Deep boundary-aware semantic image segmentationabstractAbstract While extensive research efforts have been made in semantic image segmentation, the state‐of‐the‐art methods still suffer from blurry boundaries and mismatched objects due to the insufficient multiscale adaptability. In this paper, we propose a two‐branch convolutional neural network (CNN) approach to capture the multiscale context and the boundary information with the two branches, respectively. To capture the multiscale context, we propose to embed self‐attention mechanism to the atrous spatial pyramid pooling network. To capture the boundary information, we propose to fuse the low‐level features in boundary feature extraction for refining the extracted boundaries via a feature fusion layer (FFL). With FFL, our method can improve the segmentation result with clearer boundaries. A new loss function is proposed which contains a segmentation loss and a boundary loss. Experiments show that our method can predict the boundaries of objects more clearly and have better performance for small‐scale objects. Huisi Wu, Xueting Liu 0001, Ping Li 0016 |
Comput. Animat. Virtual Worlds | 4 |
| 2021 | Seamless manga inpainting with semantics awarenessabstractManga inpainting fills up the disoccluded pixels due to the removal of dialogue balloons or "sound effect" text. This process is long needed by the industry for the language localization and the conversion to animated manga. It is mostly done manually, as existing methods (mostly for natural image inpainting) cannot produce satisfying results. Manga inpainting is more tricky than natural image inpainting because its highly abstract illustration using structural lines and screentone patterns, which confuses the semantic interpretation and visual content synthesis. In this paper, we present the first manga inpainting method, a deep learning model, that generates high-quality results. Instead of direct inpainting, we propose to separate the complicated inpainting into two major phases, semantic inpainting and appearance synthesis. This separation eases both the feature understanding and hence the training of the learning model. A key idea is to disentangle the structural line and screentone, that helps the network to better distinguish the structural line and the screentone features for semantic interpretation. Both the visual comparison and the quantitative experiments evidence the effectiveness of our method and justify its superiority over existing state-of-the-art methods in the application of manga inpainting. Minshan Xie, Menghan Xia, Xueting Liu 0001, Chengze Li, Tien-Tsin Wong |
ACM Trans. Graph. | 3 |
| 2021 | Perceptual-Aware Sketch Simplification Based on Integrated VGG LayersabstractDeep learning has been recently demonstrated as an effective tool for raster-based sketch simplification. Nevertheless, it remains challenging to simplify extremely rough sketches. We found that a simplification network trained with a simple loss, such as pixel loss or discriminator loss, may fail to retain the semantically meaningful details when simplifying a very sketchy and complicated drawing. In this paper, we show that, with a well-designed multi-layer perceptual loss, we are able to obtain aesthetic and neat simplification results preserving semantically important global structures as well as fine details without blurriness and excessive emphasis on local structures. To do so, we design a multi-layer discriminator by fusing all VGG feature layers to differentiate sketches and clean lines. The weights used in layer fusing are automatically learned via an intelligent adjustment mechanism. Furthermore, to evaluate our method, we compare our method to state-of-the-art methods through multiple experiments, including visual comparison and intensive user study. Xuemiao Xu, Minshan Xie, Peiqi Miao, Wenpeng Xiao, Huaidong Zhang, Xueting Liu 0001, Tien-Tsin Wong |
IEEE Trans. Vis. Comput. Graph. | 7 |
| 2020 | Manga filling style conversion with screentone variational autoencoderabstractWestern color comics and Japanese-style screened manga are two popular comic styles. They mainly differ in the style of region-filling. However, the conversion between the two region-filling styles is very challenging, and manually done currently. In this paper, we identify that the major obstacle in the conversion between the two filling styles stems from the difference between the fundamental properties of screened region-filling and colored region-filling. To resolve this obstacle, we propose a screentone variational autoencoder, ScreenVAE, to map the screened manga to an intermediate domain. This intermediate domain can summarize local texture characteristics and is interpolative. With this domain, we effectively unify the properties of screening and color-filling, and ease the learning for bidirectional translation between screened manga and color comics. To carry out the bidirectional translation, we further propose a network to learn the translation between the intermediate domain and color comics. Our model can generate quality screened manga given a color comic, and generate color comic that retains the original screening intention by the bitonal manga artist. Several results are shown to demonstrate the effectiveness and convenience of the proposed method. We also demonstrate how the intermediate domain can assist other applications such as manga inpainting and photo-to-comic conversion. Minshan Xie, Chengze Li, Xueting Liu 0001, Tien-Tsin Wong |
ACM Trans. Graph. | 3 |
| 2019 | Colorblind-Shareable VideosabstractThe two distinctive visual experiences of binocular display, with and without stereoscopic glasses, have been recently utilized for visual sharing between the colorblind and the normal-vision audiences. However, all existing methods only work for still images, and lack of temporal consistency for video application. In this paper, we propose the first synthesis method for colorblind-sharable videos that possess the temporal consistency for both visual experiences of colorblind and normal-vision, and retains all other crucial characteristics for visual sharing with colorblind. We formulate this challenging multi-constraint problem as a global optimization and minimize an objective function consisting of temporal term, color preservation term, color distinguishability term, and binocular fusibility term. Qualitative and quantitative experiments are conducted to evaluate the effectiveness of the proposed method comparing to existing methods. Xinghong Hu, Xueting Liu 0001, Tien-Tsin Wong |
CW | 2 |
| 2019 | Colorblind-shareable videos by synthesizing temporal-coherent polynomial coefficientsabstractTo share the same visual content between color vision deficiencies (CVD) and normal-vision people, attempts have been made to allocate the two visual experiences of a binocular display (wearing and not wearing glasses) to CVD and normal-vision audiences. However, existing approaches only work for still images. Although state-of-the-art temporal filtering techniques can be applied to smooth the per-frame generated content, they may fail to maintain the multiple binocular constraints needed in our applications, and even worse, sometimes introduce color inconsistency (same color regions map to different colors). In this paper, we propose to train a neural network to predict the temporal coherent polynomial coefficients in the domain of global color decomposition. This indirect formulation solves the color inconsistency problem. Our key challenge is to design a neural network to predict the temporal coherent coefficients, while maintaining all required binocular constraints. Our method is evaluated on various videos and all metrics confirm that it outperforms all existing solutions. Xinghong Hu, Xueting Liu 0001, Zhuming Zhang, Menghan Xia, Chengze Li, Tien-Tsin Wong |
ACM Trans. Graph. | 2 |
| 2019 | Deep binocular tone mapping
Zhuming Zhang, Chu Han, Shengfeng He, Xueting Liu 0001, Xinghong Hu, Tien-Tsin Wong |
Vis. Comput. | 4 |
| 2018 | Binocular Tone Mapping with Improved Overall Contrast and Local DetailsabstractAbstract Tone mapping is a commonly used technique that maps the set of colors in high‐dynamic‐range (HDR) images to another set of colors in low‐dynamic‐range (LDR) images, to fit the need for print‐outs, LCD monitors and projectors. Unfortunately, during the compression of dynamic range, the overall contrast and local details generally cannot be preserved simultaneously. Recently, with the increased use of stereoscopic devices, the notion of binocular tone mapping has been proposed in the existing research study. However, the existing research lacks the binocular perception study and is unable to generate the optimal binocular pair that presents the most visual content. In this paper, we propose a novel perception‐based binocular tone mapping method, that can generate an optimal binocular image pair (generating left and right images simultaneously) from an HDR image that presents the most visual content by designing a binocular perception metric. Our method outperforms the existing method in terms of both visual and time performance. Zhuming Zhang, Xinghong Hu, Xueting Liu 0001, Tien-Tsin Wong |
Comput. Graph. Forum | 3 |
| 2018 | TransHist: Occlusion-robust shape detection in cluttered imagesabstractShape matching plays an important role in various computer vision and graphics applications such as shape retrieval, object detection, image editing, image retrieval, etc. However, detecting shapes in cluttered images is still quite challenging due to the incomplete edges and changing perspective. In this paper, we propose a novel approach that can efficiently identify a queried shape in a cluttered image. The core idea is to acquire the transformation from the queried shape to the cluttered image by summarising all point-to-point transformations between the queried shape and the image. To do so, we adopt a point-based shape descriptor, the pyramid of arc-length descriptor (PAD), to identify point pairs between the queried shape and the image having similar local shapes. We further calculate the transformations between the identified point pairs based on PAD. Finally, we summarise all transformations in a 4D transformation histogram and search for the main cluster. Our method can handle both closed shapes and open curves, and is resistant to partial occlusions. Experiments show that our method can robustly detect shapes in images in the presence of partial occlusions, fragile edges, and cluttered backgrounds. Chu Han, Xueting Liu 0001, Lok Tsun Sinn, Tien-Tsin Wong |
Comput. Vis. Media | 2 |
| 2018 | Invertible grayscaleabstractOnce a color image is converted to grayscale, it is a common belief that the original color cannot be fully restored, even with the state-of-the-art colorization methods. In this paper, we propose an innovative method to synthesize invertible grayscale. It is a grayscale image that can fully restore its original color. The key idea here is to encode the original color information into the synthesized grayscale, in a way that users cannot recognize any anomalies. We propose to learn and embed the color-encoding scheme via a convolutional neural network (CNN). It consists of an encoding network to convert a color image to grayscale, and a decoding network to invert the grayscale to color. We then design a loss function to ensure the trained network possesses three required properties: (a) color invertibility, (b) grayscale conformity, and (c) resistance to quantization error. We have conducted intensive quantitative experiments and user studies over a large amount of color images to validate the proposed method. Regardless of the genre and content of the color input, convincing results are obtained in all cases. Menghan Xia, Xueting Liu 0001, Tien-Tsin Wong |
ACM Trans. Graph. | 2 |
| 2018 | Globally Consistent Wrinkle-Aware Shading of Line DrawingsabstractShading is a tedious process for artists involved in 2D cartoon and manga production given the volume of contents that the artists have to prepare regularly over tight schedule. While we can automate shading production with the presence of geometry, it is impractical for artists to model the geometry for every single drawing. In this work, we aim to automate shading generation by analyzing the local shapes, connections, and spatial arrangement of wrinkle strokes in a clean line drawing. By this, artists can focus more on the design rather than the tedious manual editing work, and experiment with different shading effects under different conditions. To achieve this, we have made three key technical contributions. First, we model five perceptual cues by exploring relevant psychological principles to estimate the local depth profile around strokes. Second, we formulate stroke interpretation as a global optimization model that simultaneously balances different interpretations suggested by the perceptual cues and minimizes the interpretation discrepancy. Lastly, we develop a wrinkle-aware inflation method to generate a height field for the surface to support the shading region computation. In particular, we enable the generation of two commonly-used shading styles: 3D-like soft shading and manga-style flat shading. Pradeep Kumar Jayaraman, Chi-Wing Fu, Jianmin Zheng, Xueting Liu 0001, Tien-Tsin Wong |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2017 | Boundary-aware texture region segmentation from mangaabstractDue to the lack of color in manga (Japanese comics), black-and-white textures are often used to enrich visual experience. With the rising need to digitize manga, segmenting texture regions from manga has become an indispensable basis for almost all manga processing, from vectorization to colorization. Unfortunately, such texture segmentation is not easy since textures in manga are composed of lines and exhibit similar features to structural lines (contour lines). So currently, texture segmentation is still manually performed, which is labor-intensive and time-consuming. To extract a texture region, various texture features have been proposed for measuring texture similarity, but precise boundaries cannot be achieved since boundary pixels exhibit different features from inner pixels. In this paper, we propose a novel method which also adopts texture features to estimate texture regions. Unlike existing methods, the estimated texture region is only regarded an initial, imprecise texture region. We expand the initial texture region to the precise boundary based on local smoothness via a graph-cut formulation. This allows our method to extract texture regions with precise boundaries. We have applied our method to various manga images and satisfactory results were achieved in all cases. Xueting Liu 0001, Chengze Li, Tien-Tsin Wong |
Comput. Vis. Media | 1 |
| 2017 | Deep extraction of manga structural linesabstractExtraction of structural lines from pattern-rich manga is a crucial step for migrating legacy manga to digital domain. Unfortunately, it is very challenging to distinguish structural lines from arbitrary, highly-structured, and black-and-white screen patterns. In this paper, we present a novel data-driven approach to identify structural lines out of pattern-rich manga, with no assumption on the patterns. The method is based on convolutional neural networks. To suit our purpose, we propose a deep network model to handle the large variety of screen patterns and raise output accuracy. We also develop an efficient and effective way to generate a rich set of training data pairs. Our method suppresses arbitrary screen patterns no matter whether these patterns are regular, irregular, tone-varying, or even pictorial, and regardless of their scales. It outputs clear and smooth structural lines even if these lines are contaminated by and immersed in complex patterns. We have evaluated our method on a large number of mangas of various drawing styles. Our method substantially outperforms state-of-the-art methods in terms of visual quality. We also demonstrate its potential in various manga applications, including manga colorization, manga retargeting, and 2.5D manga generation. Chengze Li, Xueting Liu 0001, Tien-Tsin Wong |
ACM Trans. Graph. | 2 |
| 2017 | ASCII Art Synthesis from Natural PhotographsabstractWhile ASCII art is a worldwide popular art form, automatic generating structure-based ASCII art from natural photographs remains challenging. The major challenge lies on extracting the perception-sensitive structure from the natural photographs so that a more concise ASCII art reproduction can be produced based on the structure. However, due to excessive amount of texture in natural photos, extracting perception-sensitive structure is not easy, especially when the structure may be weak and within the texture region. Besides, to fit different target text resolutions, the amount of the extracted structure should also be controllable. To tackle these challenges, we introduce a visual perception mechanism of non-classical receptive field modulation (non-CRF modulation) from physiological findings to this ASCII art application, and propose a new model of non-CRF modulation which can better separate the weak structure from the crowded texture, and also better control the scale of texture suppression. Thanks to our non-CRF model, more sensible ASCII art reproduction can be obtained. In addition, to produce more visually appealing ASCII arts, we propose a novel optimization scheme to obtain the optimal placement of proportional-font characters. We apply our method on a rich variety of images, and visually appealing ASCII art can be obtained in all cases. Xuemiao Xu, Linyuan Zhong, Minshan Xie, Xueting Liu 0001, Harry Qin, Tien-Tsin Wong |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2016 | Globally optimal toon trackingabstractThe ability to identify objects or region correspondences between consecutive frames of a given hand-drawn animation sequence is an indispensable tool for automating animation modification tasks such as sequence-wide recoloring or shape-editing of a specific animated character. Existing correspondence identification methods heavily rely on appearance features, but these features alone are insufficient to reliably identify region correspondences when there exist occlusions or when two or more objects share similar appearances. To resolve the above problems, manual assistance is often required. In this paper, we propose a new correspondence identification method which considers both appearance features and motions of regions in a global manner. We formulate correspondence likelihoods between temporal region pairs as a network flow graph problem which can be solved by a well-established optimization algorithm. We have evaluated our method with various animation sequences and results show that our method consistently outperforms the state-of-the-art methods without any user guidance. Xueting Liu 0001, Tien-Tsin Wong, Pheng-Ann Heng |
ACM Trans. Graph. | 2 |
| 2016 | Text-aware balloon extraction from manga
Xueting Liu 0001, Chengze Li, Tien-Tsin Wong, Xuemiao Xu |
Vis. Comput. | 1 |
| 2015 | Region-based structure line detection for cartoonsabstractCartoons are a worldwide popular visual entertainment medium with a long history. Nowadays, with the boom of electronic devices, there is an increasing need to digitize old classic cartoons as a basis for further editing, including deformation, colorization, etc. To perform such editing, it is essential to extract the structure lines within cartoon images. Traditional edge detection methods are mainly based on gradients. These methods perform poorly in the face of compression artifacts and spatially-varying line colors, which cause gradient values to become unreliable. This paper presents the first approach to extract structure lines in cartoons based on regions. Our method starts by segmenting an image into regions, and then classifies them as edge regions and non-edge regions. Our second main contribution comprises three measures to estimate the likelihood of a region being a non-edge region. These measure darkness, local contrast, and shape. Since the likelihoods become unreliable as regions become smaller, we further classify regions using both likelihoods and the relationships to neighboring regions via a graph-cut formulation. Our method has been evaluated on a wide variety of cartoon images, and convincing results are obtained in all cases. Xueting Liu 0001, Tien-Tsin Wong, Xuemiao Xu |
Comput. Vis. Media | 2 |
| 2015 | Closure-aware sketch simplificationabstractIn this paper, we propose a novel approach to simplify sketch drawings. The core problem is how to group sketchy strokes meaningfully, and this depends on how humans understand the sketches. The existing methods mainly rely on thresholding low-level geometric properties among the strokes, such as proximity, continuity and parallelism. However, it is not uncommon to have strokes with equal geometric properties but different semantics. The lack of semantic analysis will lead to the inability in differentiating the above semantically different scenarios. In this paper, we point out that, due to the gestalt phenomenon of closure , the grouping of strokes is actually highly influenced by the interpretation of regions. On the other hand, the interpretation of regions is also influenced by the interpretation of strokes since regions are formed and depicted by strokes. This is actually a chicken-or-the-egg dilemma and we solve it by an iterative cyclic refinement approach. Once the formed stroke groups are stabilized, we can simplify the sketchy strokes by replacing each stroke group with a smooth curve. We evaluate our method on a wide range of different sketch styles and semantically meaningful simplification results can be obtained in all test cases. Xueting Liu 0001, Tien-Tsin Wong, Pheng-Ann Heng |
ACM Trans. Graph. | 1 |
| 2013 | Stereoscopizing cel animationsabstractWhile hand-drawn cel animation is a world-wide popular form of art and entertainment, introducing stereoscopic effect into it remains difficult and costly, due to the lack of physical clues. In this paper, we propose a method to synthesize convincing stereoscopic cel animations from ordinary 2D inputs, without labor-intensive manual depth assignment nor 3D geometry reconstruction. It is mainly automatic due to the need of producing lengthy animation sequences, but with the option of allowing users to adjust or constrain all intermediate results. The system fits nicely into the existing production flow of cel animation. By utilizing the T-junction cue available in cartoons, we first infer the initial, but not reliable, ordering of regions. One of our major contributions is to resolve the temporal inconsistency of ordering by formulating it as a graph-cut problem. However, the resultant ordering remains insufficient for generating convincing stereoscopic effect, as ordering cannot be directly used for depth assignment due to its discontinuous nature. We further propose to synthesize the depth through an optimization process with the ordering formulated as constraints. This is our second major contribution. The optimized result is the spatiotemporally smooth depth for synthesizing stereoscopic effect. Our method has been evaluated on a wide range of cel animations and convincing stereoscopic effect is obtained in all cases. Xueting Liu 0001, Xuan S. Yang, Linling Zhang, Tien-Tsin Wong |
ACM Trans. Graph. | 1 |