EDBT 2026 Demo / reviewers in the wild / expert
Zhouhui Lian
dblp:20/1780
· DBLP profile ↗
80ranked-venue papers
9as first author
42since 2021 · last 2026
0000-0002-2683-7170ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 59 · 5 first-author · 31 since 2021Artificial intelligence and machine learning · 44 · 5 first-author · 28 since 2021Databases, data management, data science and information retrieval · 4 · 2 since 2021Human-computer interaction and ubiquitous computing · 3 · 3 since 2021Systems, architecture and hardware · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | IndoorUAV: Benchmarking Vision-Language UAV Navigation in Continuous Indoor EnvironmentsabstractVision-Language Navigation (VLN) enables agents to navigate in complex environments by following natural language instructions grounded in visual observations. Although most existing work has focused on ground-based robots or outdoor Unmanned Aerial Vehicles (UAVs), indoor UAV-based VLN remains underexplored, despite its relevance to real-world applications such as inspection, delivery, and search-and-rescue in confined spaces. To bridge this gap, we introduce IndoorUAV, a novel benchmark and method specifically tailored for VLN with indoor UAVs. We begin by curating over 1,000 diverse and structurally rich 3D indoor scenes from the Habitat simulator. Within these environments, we simulate realistic UAV flight dynamics to collect diverse 3D navigation trajectories manually, further enriched through data augmentation techniques. Furthermore, we design an automated annotation pipeline to generate natural language instructions of varying granularity for each trajectory. This process yields over 16,000 high-quality trajectories, comprising the IndoorUAV-VLN subset, which focuses on long-horizon VLN. To support short-horizon planning, we segment long trajectories into sub-trajectories by selecting semantically salient keyframes and regenerating concise instructions, forming the IndoorUAV-VLA subset. Finally, we introduce IndoorUAV-Agent, a novel navigation model designed for our benchmark, leveraging task decomposition and multimodal reasoning. We hope IndoorUAV serves as a valuable resource to advance research on vision-language embodied AI in the indoor aerial navigation domain. Hanshuo Qiu, Qirong Yang, Zhouhui Lian |
AAAI | 5 |
| 2026 | TextFlux: An OCR-Free DiT Model for High-Fidelity Multilingual Scene Text SynthesisabstractAbstract Diffusion‐based scene text synthesis has progressed rapidly, yet existing methods commonly rely on additional visual conditioning modules and require large‐scale annotated data to support multilingual generation. In this work, we revisit the necessity of complex auxiliary modules and further explore an approach that simultaneously ensures glyph accuracy and achieves high‐fidelity scene integration, by leveraging diffusion models' inherent capabilities for contextual reasoning. To this end, we introduce TextFlux, a DiT‐based framework that enables multilingual scene text synthesis. The advantages of TextFlux can be summarized as follows: (1) OCR‐free model architecture. TextFlux eliminates the need for OCR encoders that are specifically used to extract visual text‐related features. (2) Strong multilingual scalability. TextFlux is effective in low‐resource multilingual settings, and achieves strong performance in newly added languages with fewer than 1,000 samples. (3) Streamlined training setup. TextFlux is trained with only 1% of the training data required by competing methods. (4) Controllable multi‐line text generation. TextFlux offers flexible multi‐line synthesis with precise line‐level control, outperforming methods restricted to single‐line or rigid layouts. Extensive experiments and visualizations demonstrate that TextFlux outperforms previous methods in both qualitative and quantitative evaluations. Our code is available at https://github.com/yyyyyxie/textflux . Jielei Zhang, Weihang Wang 0011, Longwen Gao, Zhouhui Lian |
Comput. Graph. Forum | 8 |
| 2026 | RSUniVLM: A unified vision-language model for remote sensing via Granularity-oriented MoE
Zhouhui Lian |
Pattern Recognit. | 2 |
| 2025 | Creating Your Editable 3D Photorealistic Avatar with Tetrahedron-constrained Gaussian SplattingabstractPersonalized 3D avatar editing holds significant promise due to its user-friendliness and availability to applications such as AR/VR and virtual try-ons. Previous studies have explored the feasibility of 3D editing, but often struggle to generate visually pleasing results, possibly due to the unstable representation learning under mixed optimization of geometry and texture in complicated reconstructed scenarios. In this paper, we aim to provide an accessible solution for ordinary users to create their editable 3D avatars with precise region localization, geometric adaptability, and photorealistic renderings. To tackle this challenge, we introduce a meticulously designed framework that decouples the editing process into local spatial adaptation and realistic appearance learning, utilizing a hybrid Tetrahedron-constrained Gaussian Splatting (TetGS) as the underlying representation. TetGS combines the controllable explicit structure of tetrahedral grids with the high-precision rendering capabilities of 3D Gaussian Splatting and is optimized in a progressive manner comprising three stages: 3D avatar instantiation from real-world monocular videos to provide accurate priors for TetGS initialization; localized spatial adaptation with explicitly partitioned tetrahedrons to guide the redistribution of Gaussian kernels; and geometry-based appearance generation with a coarse-to-fine activation strategy. Both qualitative and quantitative experiments demonstrate the effectiveness and superiority of our approach in generating photorealistic 3D editable avatars. Hanxi Liu, Yifang Men, Zhouhui Lian |
CVPR | 3 |
| 2025 | ART: Anonymous Region Transformer for Variable Multi-Layer Transparent Image GenerationabstractMulti-layer image generation is a fundamental task that enables users to isolate, select, and edit specific image layers, thereby revolutionizing interactions with generative models. In this paper, we introduce the Anonymous Region Transformer (ART), which facilitates the direct generation of variable multi-layer transparent images based on a global text prompt and an anonymous region layout. Inspired by Schema theory1, this anonymous region layout allows the generative model to autonomously determine which set of visual tokens should align with which text tokens, which is in contrast to the previously dominant semantic layout for the image generation task. In addition, the layer-wise region crop mechanism, which only selects the visual tokens belonging to each anonymous region, significantly reduces attention computation costs and enables the efficient generation of images with numerous distinct layers (e.g., 50+). When compared to the full attention approach, our method is over 12 times faster and exhibits fewer layer conflicts. Furthermore, we propose a high-quality multi-layer transparent image autoencoder that supports the direct encoding and decoding of the transparency of variable multi-layer images in a joint manner. By enabling precise control and scalable layer generation, ART establishes a new paradigm for interactive content creation. Yifan Pu, Zhicong Tang, Ruihong Yin, Haoxing Ye, Yuhui Yuan, Dong Chen 0003, Jianmin Bao, Sirui Zhang, Ji Li 0006, Xiu Li 0001, Zhouhui Lian, Gao Huang 0001, Baining Guo |
CVPR | 15 |
| 2025 | TexGaussian: Generating High-quality PBR Material via Octree-based 3D Gaussian SplattingabstractPhysically Based Rendering (PBR) materials play a crucial role in modern graphics, enabling photorealistic rendering across diverse environment maps. Developing an effective and efficient algorithm that is capable of automatically generating high-quality PBR materials rather than RGB texture for 3D meshes can significantly streamline the 3D content creation. Most existing methods leverage pre-trained 2D diffusion models for multi-view image synthesis, which often leads to severe inconsistency between the generated textures and input 3D meshes. This paper presents TexGaussian, a novel method that uses octant-aligned 3D Gaussian Splatting for rapid PBR material generation. Specifically, we place each 3D Gaussian on the finest leaf node of the octree built from the input 3D mesh to render the multi-view images not only for the albedo map but also for roughness and metallic. Moreover, our model is trained in a regression manner instead of diffusion denoising, capable of generating the PBR material for a 3D mesh in a single feed-forward process. Extensive experiments on publicly available benchmarks demonstrate that our method synthesizes more visually pleasing PBR materials and runs faster than previous methods in both unconditional and text-conditional scenarios, exhibiting better consistency with the given geometry. Our code and trained models are available at https://3d-aigc.github.io/TexGaussian. Bojun Xiong, Jialun Liu, Chenming Wu, Chen Zhao 0011, Errui Ding, Zhouhui Lian |
CVPR | 9 |
| 2025 | CalliReader: Contextualizing Chinese Calligraphy via an Embedding-Aligned Vision-Language ModelabstractChinese calligraphy, a UNESCO Heritage, remains computationally challenging due to visual ambiguity and cultural complexity. Existing AI systems fail to contextualize their intricate scripts, because of limited annotated data and poor visual-semantic alignment. We propose CalliReader, a vision-language model (VLM) that solves the Chinese Calligraphy Contextualization (CC$^2$) problem through three innovations: (1) character-wise slicing for precise character extraction and sorting, (2) CalliAlign for visual-text token compression and alignment, (3) embedding instruction tuning (e-IT) for improving alignment and addressing data scarcity. We also build CalliBench, the first benchmark for full-page calligraphic contextualization, addressing three critical issues in previous OCR and VQA approaches: fragmented context, shallow reasoning, and hallucination. Extensive experiments including user studies have been conducted to verify our CalliReader's \textbf{superiority to other state-of-the-art methods and even human professionals in page-level calligraphy recognition and interpretation}, achieving higher accuracy while reducing hallucination. Comparisons with reasoning models highlight the importance of accurate recognition as a prerequisite for reliable comprehension. Quantitative analyses validate CalliReader's efficiency; evaluations on document and real-world benchmarks confirm its robust generalization ability. Chenyi Huang, Feiyang Hao, Zhouhui Lian |
ICCV | 5 |
| 2025 | MMMG: A Massive, Multidisciplinary, Multi-Tier Generation Benchmark for Text-to-Image ReasoningabstractIn this paper, we introduce knowledge image generation as a new task, alongside the Massive Multi-Discipline Multi-Tier Knowledge-Image Generation Benchmark (MMMG) to probe the reasoning capability of image generation models.Knowledge images have been central to human civilization and to the mechanisms of human learning—a fact underscored by dual-coding theory and the picture-superiority effect.Generating such images is challenging, demanding multimodal reasoning that fuses world knowledge with pixel-level grounding into clear explanatory visuals.To enable comprehensive evaluation, MMMG offers $4,456$ expert-validated (knowledge) image-prompt pairs spanning $10$ disciplines, $6$ educational levels, and diverse knowledge formats such as charts, diagrams, and mind maps. To eliminate confounding complexity during evaluation, we adopt a unified Knowledge Graph (KG) representation. Each KG explicitly delineates a target image’s core entities and their dependencies.We further introduce MMMG-Score to evaluate generated knowledge images. This metric combines factual fidelity, measured by graph-edit distance between KGs, with visual clarity assessment.Comprehensive evaluations of $21$ state-of-the-art text-to-image generation models expose serious reasoning deficits—low entity fidelity, weak relations, and clutter—with GPT-4o achieving an MMMG-Score of only $50.20$, underscoring the benchmark’s difficulty.To spur further progress, we release FLUX-Reason (MMMG-Score of $34.45$), an effective and open baseline that combines a reasoning LLM with diffusion models and is trained on $16,000$ curated knowledge image–prompt pairs. Ryan Yuan, Haonan Cai, Ziyi Yue, Fatima Zohra Daha, Zhouhui Lian |
NeurIPS | 9 |
| 2025 | UTDesign: A Unified Framework for Stylized Text Editing and Generation in Graphic Design ImagesabstractAI-assisted graphic design has emerged as a powerful tool for automating the creation and editing of design elements such as posters, banners, and advertisements. While diffusion-based text-to-image models have demonstrated strong capabilities in visual content generation, their text rendering performance, particularly for small-scale typography and non-Latin scripts, remains limited. In this paper, we propose UTDesign, a unified framework for high-precision stylized text editing and conditional text generation in design images, supporting both English and Chinese scripts. Our framework introduces a novel DiT-based text style transfer model trained from scratch on a synthetic dataset, capable of generating transparent RGBA text foregrounds that preserve the style of reference glyphs. We further extend this model into a conditional text generation framework by training a multi-modal condition encoder on a curated dataset with detailed text annotations, enabling accurate, style-consistent text synthesis conditioned on background images, prompts, and layout specifications. Finally, we integrate our approach into a fully automated text-to-design (T2D) pipeline by incorporating pre-trained text-to-image (T2I) models and an MLLM-based layout planner. Extensive experiments demonstrate that UTDesign achieves state-of-the-art performance among open-source methods in terms of stylistic consistency and text accuracy, and also exhibits unique advantages compared to proprietary commercial approaches. Code and data for this paper are available at https://github.com/ZYM-PKU/UTDesign. Yuanpeng Gao, Jiwei Duan, Shisong Lin, Longfei Xiong, Zhouhui Lian |
SIGGRAPH Asia | 7 |
| 2025 | OctFusion: Octree-based Diffusion Models for 3D Shape GenerationabstractAbstract Diffusion models have emerged as a popular method for 3D generation. However, it is still challenging for diffusion models to efficiently generate diverse and high‐quality 3D shapes. In this paper, we introduce OctFusion, which can generate 3D shapes with arbitrary resolutions in 2.5 seconds on a single Nvidia 4090 GPU, and the extracted meshes are guaranteed to be continuous and manifold. The key components of OctFusion are the octree‐based latent representation and the accompanying diffusion models. The representation combines the benefits of both implicit neural representations and explicit spatial octrees and is learned with an octree‐based variational autoencoder. The proposed diffusion model is a unified multi‐scale U‐Net that enables weights and computation sharing across different octree levels and avoids the complexity of widely used cascaded diffusion schemes. We verify the effectiveness of OctFusion on the ShapeNet and Objaverse datasets and achieve state‐of‐the‐art performances on shape generation tasks. We demonstrate that OctFusion is extendable and flexible by generating high‐quality color fields for textured mesh generation and high‐quality 3D shapes conditioned on text prompts, sketches, or category labels. Our code and pre‐trained models are available at https://github.com/octree‐nn/octfusion . Bojun Xiong, Si-Tong Wei, Xin-Yang Zheng, Yan-Pei Cao 0001, Zhouhui Lian, Peng-Shuai Wang |
Comput. Graph. Forum | 5 |
| 2025 | Neural-Polyptych: Content Controllable Painting Recreation for Diverse GenresabstractTo bridge the gap between artists and non-specialists, we present a unified framework, Neural-Polyptych, to facilitate the creation of expansive, high-resolution paintings by seamlessly incorporating interactive hand-drawn sketches with fragments from original paintings. We have designed a multi-scale GAN-based architecture to decompose the generation process into two parts, each responsible for identifying global and local features. To enhance the fidelity of semantic details generated from users' sketched outlines, we introduce a Correspondence Attention module utilizing our Reference Bank strategy. This ensures the creation of high-quality, intricately detailed elements within the artwork. The final result is achieved by carefully blending these local elements while preserving coherent global consistency. Consequently, this methodology enables the production of digital paintings at megapixel scale, accommodating diverse artistic expressions and enabling users to recreate content in a controlled manner. We validate our approach to diverse genres of both Eastern and Western paintings. Applications such as large painting extension, texture shuffling, genre switching, mural art restoration, and recomposition can be successfully based on our framework. Dewen Guo, Zhouhui Lian, Jianhong Han, Jie Feng 0001, Bingfeng Zhou, Sheng Li 0008 |
Comput. Vis. Media | 3 |
| 2024 | DeepCalliFont: Few-Shot Chinese Calligraphy Font Synthesis by Integrating Dual-Modality Generative ModelsabstractFew-shot font generation, especially for Chinese calligraphy fonts, is a challenging and ongoing problem. With the help of prior knowledge that is mainly based on glyph consistency assumptions, some recently proposed methods can synthesize high-quality Chinese glyph images. However, glyphs in calligraphy font styles often do not meet these assumptions. To address this problem, we propose a novel model, DeepCalliFont, for few-shot Chinese calligraphy font synthesis by integrating dual-modality generative models. Specifically, the proposed model consists of image synthesis and sequence generation branches, generating consistent results via a dual-modality representation learning strategy. The two modalities (i.e., glyph images and writing sequences) are properly integrated using a feature recombination module and a rasterization loss function. Furthermore, a new pre-training strategy is adopted to improve the performance by exploiting large amounts of uni-modality data. Both qualitative and quantitative experiments have been conducted to demonstrate the superiority of our method to other state-of-the-art approaches in the task of few-shot Chinese calligraphy font synthesis. The source code can be found at https://github.com/lsflyt-pku/DeepCalliFont. Yitian Liu, Zhouhui Lian |
AAAI | 2 |
| 2024 | TextNeRF: A Novel Scene-Text Image Synthesis Method Based on Neural Radiance FieldsabstractAcquiring large-scale, well-annotated datasets is essential for training robust scene text detectors, yet the process is often resource-intensive and time-consuming. While some efforts have been made to explore the synthesis of scene text images, a notable gap remains between syn-thetic and authentic data. In this paper, we introduce a novel method that utilizes Neural Radiance Fields (NeRF) to model real-world scenes and emulate the data collection process by rendering images from diverse camera per-spectives, enriching the variability and realism of the synthesized data. A semi-supervised learning framework is proposed to categorize semantic regions within 3D scenes, ensuring consistent labeling of text regions across various viewpoints. Our method also models the pose, and view-dependent appearance of text regions, thereby offering precise control over camera poses and significantly improving the realism of text insertion and editing within scenes. Employing our technique on real-world scenes has led to the creation of a novel scene text image dataset (https://github.com/cuijl-ai/TextNeRF). Compared to other existing benchmarks, the proposed dataset is distinctive in providing not only standard annotations such as bounding boxes and transcriptions but also the information of 3D pose attributes for text regions, enabling a more detailed evaluation of the robustness of text detection algorithms. Through extensive experiments, we demonstrate the effectiveness of our proposed method in enhancing the performance of scene text detectors. Jialei Cui, Jianwei Du, Wenzhuo Liu, Zhouhui Lian |
CVPR | 4 |
| 2024 | En3D: An Enhanced Generative Model for Sculpting 3D Humans from 2D Synthetic DataabstractWe present En3D, an enhanced generative scheme for sculpting high-quality 3D human avatars. Unlike previous works that rely on scarce 3D datasets or limited 2D collections with imbalanced viewing angles and imprecise pose priors, our approach aims to develop a zero-shot 3D generative scheme capable of producing visually realistic, ge-ometrically accurate and content-wise diverse 3D humans without directly relying on pre-existing 3D or 2D assets. To address this challenge, we introduce a meticulously crafted workflow that implements accurate physical modeling to learn the enhanced 3D generative model from synthetic 2D data. During inference, we integrate optimization modules to bridge the gap between realistic appearances and coarse 3D shapes. Specifically, En3D comprises three modules: a 3D generator that accurately models generalizable 3D humans with realistic appearance from synthesized balanced, diverse, and structured human images; a geometry sculptor that enhances shape quality using multi-view normal constraints for intricate human structure; and a texturing module that disentangles explicit texture maps with fidelity and editability, leveraging semantical UV partitioning and a differentiable rasterizer: Experimental results show that our approach significantly outperforms prior works in terms of image quality, geometry accuracy and content diversity. We also showcase the applicability of our generated avatars for animation and editing, as well as the scalability of our approach for content-style free adaptation. Yifang Men, Biwen Lei, Yuan Yao 0013, Miaomiao Cui, Zhouhui Lian, Xuansong Xie |
CVPR | 5 |
| 2024 | 3DToonify: Creating Your High-Fidelity 3D Stylized Avatar Easily from 2D Portrait ImagesabstractVisual content creation has aroused a surge of interest given its applications in mobile photography and AR/VR. Portrait style transfer and 3D recovery from monocular images as two representative tasks have so far evolved independently. In this paper, we make a connection between the two, and tackle the challenging task of 3D portrait styl-ization - modeling high-fidelity 3D stylized avatars from captured 2D portrait images. However, naively combining the techniques from the two isolated areas may suf-fer from either inadequate stylization or absence of 3D as-sets. To this end, we propose 3DToonify, a new framework that introduces a progressive training scheme to achieve 3D style adaption on spatial neural representation (SNR). SNR is constructed with implicit fields and they are dynamically optimized by the progressive training scheme, which consists of three stages: guided prior learning, deformable geometry adaption and explicit texture adaption. In this way, stylized geometry and texture are learned in SNR in an explicit and structured way with only a single stylized exemplar needed. Moreover, our method obtains style-adaptive underlying structures (i.e., deformable geometry and exaggerated texture) and view-consistent styl-ized avatar rendering from arbitrary novel viewpoints. Both qualitative and quantitative experiments have been conducted to demonstrate the effectiveness and superiority of our method for automatically generating exemplar-guided 3D stylized avatars. Yifang Men, Hanxi Liu, Miaomiao Cui, Xuansong Xie, Zhouhui Lian |
CVPR | 6 |
| 2024 | UDiffText: A Unified Framework for High-Quality Text Synthesis in Arbitrary Images via Character-Aware Diffusion Models
Zhouhui Lian |
ECCV (31) | 2 |
| 2024 | CalliRewrite: Recovering Handwriting Behaviors from Calligraphy Images without SupervisionabstractHuman-like planning skills and dexterous manipulation have long posed challenges in the fields of robotics and artificial intelligence (AI). The task of reinterpreting calligraphy presents a formidable challenge, as it involves the decomposition of strokes and dexterous utensil control. Previous efforts have primarily focused on supervised learning of a single instrument, limiting the performance of robots in the realm of cross-domain text replication. To address these challenges, we propose CalliRewrite: a coarse-to-fine approach for robot arms to discover and recover plausible writing orders from diverse calligraphy images without requiring labeled demonstrations. Our model achieves fine-grained control of various writing utensils. Specifically, an unsupervised image-to-sequence model decomposes a given calligraphy glyph to obtain a coarse stroke sequence. Using an RL algorithm, a simulated brush is fine-tuned to generate stylized trajectories for robotic arm control. Evaluation in simulation and physical robot scenarios reveals that our method successfully replicates unseen fonts and styles while achieving integrity in unknown characters. To access our code and supplementary materials, please visit our project page: https://luoprojectpage.github.io/callirewrite/. Zekun Wu 0004, Zhouhui Lian |
ICRA | 3 |
| 2024 | Pano2Room: Novel View Synthesis from a Single Indoor PanoramaabstractRecent single-view 3D generative methods have made significant advancements by leveraging knowledge distilled from extensive 3D object datasets. However, challenges persist in the synthesis of 3D scenes from a single view, primarily due to the complexity of real-world environments and the limited availability of high-quality prior resources. In this paper, we introduce a novel approach called Pano2Room, designed to automatically reconstruct high-quality 3D indoor scenes from a single panoramic image. These panoramic images can be easily generated using a panoramic RGBD inpainter from captures at a single location with any camera. The key idea is to initially construct a preliminary mesh from the input panorama, and iteratively refine this mesh using a panoramic RGBD inpainter while collecting photo-realistic 3D-consistent pseudo novel views. Finally, the refined mesh is converted into a 3D Gaussian Splatting field and trained with the collected pseudo novel views. This pipeline enables the reconstruction of real-world 3D scenes, even in the presence of large occlusions, and facilitates the synthesis of photo-realistic novel views with detailed geometry. Extensive qualitative and quantitative experiments have been conducted to validate the superiority of our method in single-panorama indoor novel synthesis compared to the state-of-the-art. Our code and data are available at \url{https://github.com/TrickyGo/Pano2Room}. Guo Pu, Zhouhui Lian |
SIGGRAPH Asia | 3 |
| 2024 | CHWmaster: mastering Chinese handwriting via sliding-window recurrent neural networks
Shusen Tang, Zhouhui Lian |
Int. J. Document Anal. Recognit. | 3 |
| 2024 | HFH-Font: Few-shot Chinese Font Synthesis with Higher Quality, Faster Speed, and Higher ResolutionabstractThe challenge of automatically synthesizing high-quality vector fonts, particularly for writing systems (e.g., Chinese) consisting of huge amounts of complex glyphs, remains unsolved. Existing font synthesis techniques fall into two categories: 1) methods that directly generate vector glyphs, and 2) methods that initially synthesize glyph images and then vectorize them. However, the first category often fails to construct complete and correct shapes for complex glyphs, while the latter struggles to efficiently synthesize high-resolution (i.e., 1024 × 1024 or higher) glyph images while preserving local details. In this paper, we introduce HFH-Font, a few-shot font synthesis method capable of efficiently generating high-resolution glyph images that can be converted into high-quality vector glyphs. More specifically, our method employs a diffusion model-based generative framework with component-aware conditioning to learn different levels of style information adaptable to varying input reference sizes. We also design a distillation module based on Score Distillation Sampling for 1-step fast inference, and a style-guided super-resolution module to refine and upscale low-resolution synthesis results. Extensive experiments, including a user study with professional font designers, have been conducted to demonstrate that our method significantly outperforms existing font synthesis approaches. Experimental results show that our method produces high-fidelity, high-resolution raster images which can be vectorized into high-quality vector fonts. Using our method, for the first time, large-scale Chinese vector fonts of a quality comparable to those manually created by professional font designers can be automatically generated. Zhouhui Lian |
ACM Trans. Graph. | 2 |
| 2023 | DeepVecFont-v2: Exploiting Transformers to Synthesize Vector Fonts with Higher QualityabstractVector font synthesis is a challenging and ongoing problem in the fields of Computer Vision and Computer Graphics. The recently-proposed DeepVecFont [27] achieved state-of-the-art performance by exploiting information of both the image and sequence modalities of vector fonts. However, it has limited capability for handling long sequence data and heavily relies on an image-guided outline refinement post-processing. Thus, vector glyphs synthesized by DeepVecFont still often contain some distortions and artifacts and cannot rival human-designed results. To address the above problems, this paper proposes an enhanced version of DeepVecFont mainly by making the following three novel technical contributions. First, we adopt Transformers instead of RNNs to process sequential data and design a relaxation representation for vector outlines, markedly improving the model's capability and stability of synthesizing long and complex outlines. Second, we propose to sample auxiliary points in addition to control points to precisely align the generated and target Bézier curves or lines. Finally, to alleviate error accumulation in the sequential generation process, we develop a context-based self-refinement module based on another Transformer-based decoder to remove artifacts in the initially synthesized glyphs. Both qualitative and quantitative results demonstrate that the proposed method effectively resolves those intrinsic problems of the original DeepVecFont and outperforms existing approaches in generating English and Chinese vector fonts with complicated structures and diverse styles. Yuqing Wang 0006, Longhui Yu, Yuesheng Zhu, Zhouhui Lian |
CVPR | 5 |
| 2023 | VecFontSDF: Learning to Reconstruct and Synthesize High-Quality Vector Fonts via Signed Distance FunctionsabstractFont design is of vital importance in the digital content design and modern printing industry. Developing algorithms capable of automatically synthesizing vector fonts can significantly facilitate the font design process. However, existing methods mainly concentrate on raster image generation, and only a few approaches can directly synthesize vector fonts. This paper proposes an end-to-end trainable method, VecFontSDF, to reconstruct and synthesize high-quality vector fonts using signed distance functions (SDFs). Specifically, based on the proposed SDF-based implicit shape representation, VecFontSDF learns to model each glyph as shape primitives enclosed by several parabolic curves, which can be precisely converted to quadratic Bézier curves that are widely used in vector font products. In this manner, most image generation methods can be easily extended to synthesize vector fonts. Qualitative and quantitative experiments conducted on a publicly-available dataset demonstrate that our method obtains high-quality results on several tasks, including vector font reconstruction, interpolation, and few-shot vector font synthesis, markedly outperforming the state of the art. Zeqing Xia, Bojun Xiong, Zhouhui Lian |
CVPR | 3 |
| 2023 | CurveSDF: Binary Image Vectorization Using Signed Distance FieldsabstractBinary image vectorization is a classical and fundamental problem in the areas of Computer Graphics and Computer Vision. Existing image vectorization methods are mainly based on global optimization, typically failing to preserve important details on outlines due to the incapability of learning high-level knowledge from training data. To address this problem, we propose CurveSDF to facilitate the learning of vectorization for 2D outlines. Our method consists of the following three modules: convex separation, signed distance field (SDF) generation, and curve intersection calculation. Specifically, we first divide an input binary image into convex elements. Then, we use restrained curve-hyperplane divisions to generate their SDFs and precisely reconstruct the original image. Finally, we convert the generated SDFs to vector outlines composed of both Bezier curves and line segments. Moreover, our method is self-constrained, and thus there is no need to use any vector data for training. Experimental results demonstrate the effectiveness of our method and its superiority against other existing approaches for binary image vectorization. Zeqing Xia, Zhouhui Lian |
ICMR | 2 |
| 2023 | SinMPI: Novel View Synthesis from a Single Image with Expanded Multiplane ImagesabstractSingle-image novel view synthesis is a challenging and ongoing problem that aims to generate an infinite number of consistent views from a single input image. Although significant efforts have been made to advance the quality of generated novel views, less attention has been paid to the expansion of the underlying scene representation, which is crucial to the generation of realistic novel view images. This paper proposes SinMPI, a novel method that uses an expanded multiplane image (MPI) as the 3D scene representation to significantly expand the perspective range of MPI and generate high-quality novel views from a large multiplane space. The key idea of our method is to use Stable Diffusion [Rombach et al. 2021] to generate out-of-view contents, project all scene contents into an expanded multiplane image according to depths predicted by monocular depth estimators, and then optimize the multiplane image under the supervision of pseudo multi-view data generated by a depth-aware warping and inpainting module. Both qualitative and quantitative experiments have been conducted to validate the superiority of our method to the state of the art. Our code and data are available at https://github.com/TrickyGo/SinMPI. Guo Pu, Peng-Shuai Wang, Zhouhui Lian |
SIGGRAPH Asia | 3 |
| 2023 | Controllable Image Synthesis With Attribute-Decomposed GANabstractThis paper proposes Attribute-Decomposed GAN (ADGAN) and its enhanced version (ADGAN++) for controllable image synthesis, which can produce realistic images with desired attributes provided in various source inputs. The core ideas of the proposed ADGAN and ADGAN++ are both to embed component attributes into the latent space as independent codes and thus achieve flexible and continuous control of attributes via mixing and interpolation operations in explicit style representations. The major difference between them is that ADGAN processes all component attributes simultaneously while ADGAN++ utilizes a serial encoding strategy. More specifically, ADGAN consists of two encoding pathways with style block connections and is capable of decomposing the original hard mapping into multiple more accessible subtasks. In the source pathway, component layouts are extracted via a semantic parser and the segmented components are fed into a shared global texture encoder to obtain decomposed latent codes. This strategy allows for the synthesis of more realistic output images and the automatic separation of un-annotated component attributes. Although the original ADGAN works in a delicate and efficient manner, intrinsically it fails to handle the semantic image synthesizing task when the number of attribute categories is huge. To address this problem, ADGAN++ employs the serial encoding of different component attributes to synthesize each part of the target real-world image, and adopts several residual blocks with segmentation guided instance normalization to assemble the synthesized component images and refine the original synthesis result. The two-stage ADGAN++ is designed to alleviate the massive computational costs required when synthesizing real-world images with numerous attributes while maintaining the disentanglement of different attributes to enable flexible control of arbitrary component attributes of the synthesized images. Experimental results demonstrate the proposed methods' superiority over the state of the art in pose transfer, face style transfer, and semantic image synthesis, as well as their effectiveness in the task of component attribute transfer. Our code and data are publicly available at https://github.com/menyifang/ADGAN. Guo Pu, Yifang Men, Yiming Mao 0006, Yuning Jiang 0001, Wei-Ying Ma, Zhouhui Lian |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2023 | FontTransformer: Few-shot high-resolution Chinese glyph image synthesis via stacked transformers
Yitian Liu, Zhouhui Lian |
Pattern Recognit. | 2 |
| 2023 | Learning isometry-invariant representations for point cloud analysis
Xiao Sun 0011, Zhouhui Lian |
Pattern Recognit. | 3 |
| 2022 | Unpaired Cartoon Image Synthesis via Gated Cycle MappingabstractIn this paper, we present a general-purpose solution to cartoon image synthesis with unpaired training data. In contrast to previous works learning pre-defined cartoon styles for specified usage scenarios (portrait or scene), we aim to train a common cartoon translator which can not only simultaneously render exaggerated anime faces and realistic cartoon scenes, but also provide flexible user controls for desired cartoon styles. It is challenging due to the complexity of the task and the absence of paired data. The core idea of the proposed method is to introduce gated cycle mapping, that utilizes a novel gated mapping unit to produce the category-specific style code and embeds this code into cycle networks to control the translation process. For the concept of category, we classify images into different categories (e.g., 4 types: photo/cartoon portrait/scene) and learn finer-grained category translations rather than overall mappings between two domains (e.g., photo and cartoon). Furthermore, the proposed method can be easily extended to cartoon video generation with an auxiliary dataset and a new adaptive style loss. Experimental results demonstrate the superiority of the proposed method over the state of the art and validate its effectiveness in the brand-new task of general cartoon image synthesis. Yifang Men, Yuan Yao 0013, Miaomiao Cui, Zhouhui Lian, Xuansong Xie, Xian-Sheng Hua 0001 |
CVPR | 4 |
| 2022 | Aesthetic Text Logo Synthesis via Content-aware Layout InferringabstractText logo design heavily relies on the creativity and expertise of professional designers, in which arranging element layouts is one of the most important procedures. However, few attention has been paid to this task which needs to take many factors (e.g., fonts, linguistics, topics, etc.) into consideration. In this paper, we propose a content-aware layout generation network which takes glyph images and their corresponding text as input and synthesizes aesthetic layouts for them automatically. Specifically, we develop a dual-discriminator module, including a sequence discriminator and an image discriminator, to evaluate both the character placing trajectories and rendered shapes of synthesized text logos, respectively. Furthermore, we fuse the information of linguistics from texts and visual semantics from glyphs to guide layout prediction, which both play important roles in professional layout design. To train and evaluate our approach, we construct a dataset named as TextLogo3K, consisting of about 3,500 text logo images and their pixel-level annotations. Experimental studies on this dataset demonstrate the effectiveness of our approach for synthesizing visually-pleasing text logos and verify its superiority against the state of the art. Guo Pu, Wenhan Luo, Yexin Wang, Pengfei Xiong, Hongwen Kang, Zhouhui Lian |
CVPR | 7 |
| 2022 | Semi-supervised Semantic Segmentation via Prototypical Contrastive LearningabstractThe key idea of semi-supervised semantic segmentation is to leverage both labeled and unlabeled data. To achieve the goal, most existing methods resort to pseudo-labels for training. However, the dispersed feature distribution and biased category centroids could inevitably lead to the calculation deviation of feature distances and noisy pseudo labels. In this paper, we propose to denoise pseudo labels with representative prototypes. Specifically, to mitigate the effects of outliers, we first employ automatic clustering to model multiple prototypes with which the distribution of outliers can be better characterized. Then, a compact structure and clear decision boundary can be obtained by using contrastive learning. It is worth noting that our prototype-wise pseudo segmentation strategy can also be applied in most existing semantic segmentation networks. Experimental results show that our method outperforms other state-of-the-art approaches on both Cityscapes and Pascal VOC semantic segmentation datasets under various data partition protocols. Zenggui Chen, Zhouhui Lian |
ACM Multimedia | 2 |
| 2022 | CVFont: Synthesizing Chinese Vector Fonts via Deep Layout InferringabstractAbstract Creating a high‐quality Chinese vector font library, which can be directly used in real applications is time‐consuming and costly, since the font library typically consists of large amounts of vector glyphs. To address this problem, we propose a data‐driven system in which only a small number (about 10%) of Chinese glyphs need to be designed. Specifically, the system first automatically decomposes those input glyphs into vector components. Then, a layout prediction module based on deep neural networks is applied to learn the layout style of input characters. Finally, proper components are selected to assemble the glyph of each unseen character based on the predicted layout to build the font library that can be directly used in computers and smart mobile devices. Experimental results demonstrate that our system synthesizes high‐quality glyphs and significantly enhances the producing efficiency of Chinese vector fonts. Zhouhui Lian, Yichen Gao |
Comput. Graph. Forum | 1 |
| 2022 | DCT-net: domain-calibrated translation for portrait stylizationabstractThis paper introduces DCT-Net, a novel image translation architecture for few-shot portrait stylization. Given limited style exemplars (~100), the new architecture can produce high-quality style transfer results with advanced ability to synthesize high-fidelity contents and strong generality to handle complicated scenes (e.g., occlusions and accessories). Moreover, it enables full-body image translation via one elegant evaluation network trained by partial observations (i.e., stylized heads). Few-shot learning based style transfer is challenging since the learned model can easily become overfitted in the target domain, due to the biased distribution formed by only a few training examples. This paper aims to handle the challenge by adopting the key idea of "calibration first, translation later" and exploring the augmented global structure with locally-focused translation. Specifically, the proposed DCT-Net consists of three modules: a content adapter borrowing the powerful prior from source photos to calibrate the content distribution of target samples; a geometry expansion module using affine transformations to release spatially semantic constraints; and a texture translation module leveraging samples produced by the calibrated distribution to learn a fine-grained conversion. Experimental results demonstrate the proposed method's superiority over the state of the art in head stylization and its effectiveness on full image translation with adaptive deformations. Our code is publicly available at https://github.com/menyifang/DCT-Net. Yifang Men, Yuan Yao 0013, Miaomiao Cui, Zhouhui Lian, Xuansong Xie |
ACM Trans. Graph. | 4 |
| 2022 | DifferSketching: How Differently Do People Sketch 3D Objects?abstractMultiple sketch datasets have been proposed to understand how people draw 3D objects. However, such datasets are often of small scale and cover a small set of objects or categories. In addition, these datasets contain freehand sketches mostly from expert users, making it difficult to compare the drawings by expert and novice users, while such comparisons are critical in informing more effective sketch-based interfaces for either user groups. These observations motivate us to analyze how differently people with and without adequate drawing skills sketch 3D objects. We invited 70 novice users and 38 expert users to sketch 136 3D objects, which were presented as 362 images rendered from multiple views. This leads to a new dataset of 3,620 freehand multi-view sketches, which are registered with their corresponding 3D objects under certain views. Our dataset is an order of magnitude larger than the existing datasets. We analyze the collected data at three levels, i.e., sketch-level, stroke-level, and pixel-level, under both spatial and temporal characteristics, and within and across groups of creators. We found that the drawings by professionals and novices show significant differences at stroke-level, both intrinsically and extrinsically. We demonstrate the usefulness of our dataset in two applications: (i) freehand-style sketch synthesis, and (ii) posing it as a potential benchmark for sketch-based 3D reconstruction. Our dataset and code are available at https://chufengxiao.github.io/DifferSketching/. Chu-Feng Xiao 0001, Wanchao Su, Jing Liao 0001, Zhouhui Lian, Yi-Zhe Song, Hongbo Fu 0001 |
ACM Trans. Graph. | 4 |
| 2021 | FontRL: Chinese Font Synthesis via Deep Reinforcement LearningabstractAutomatic generation of Chinese fonts is a valuable but challenging task in areas of AI and Computer Graphics, mainly due to the huge amount of Chinese characters and their complex glyph structures. In this paper, we propose FontRL, a novel method for Chinese font synthesis by using deep reinforcement learning. Specifically, we first train a deep reinforcement learning model to obtain the Thin-Plate Spline (TPS) transformation that is able to modify the reference stroke skeleton in a mean font style into the skeleton of a required style for each stroke of every unseen Chinese character. Afterwards, we utilize a CNN model to predict the location and scale information of these strokes, and then assemble them to get the skeleton of the corresponding character. Finally, we convert each synthesized character skeleton into the glyph image via an image-to-image translation model. Both quantitative and qualitative experimental results demonstrate the superiority of the proposed FontRL compared to the state of the art. Our code is available at https://github.com/lsflyt-pku/FontRL. Yitian Liu, Zhouhui Lian |
AAAI | 2 |
| 2021 | BoW Pooling: A Plug-and-Play Unit for Feature Aggregation of Point CloudsabstractPoint cloud provides a compact and flexible representation for 3D shapes and recently attracts more and more attention due to the increasing demands in practical applications. The major challenge of handling such irregular data is how to achieve the permutation invariance of points in the input. Most of existing methods extract local descriptors that encode the geometry of local structure, followed by a symmetric function to form a global representation. The max pooling usually serves as the symmetric function and shows slight superiority compared to the average pooling. We argue that some discrimination information is inevitably missing when applying the max pooling across all local descriptors. In this paper, we propose the BoW pooling, a plug-and-play unit to substitute the max pooling. Our BoW pooling analyzes the set of local descriptors statistically and generates a histogram that reflects how the primitives in the dictionary constitute the overall geometry. Extensive experiments demonstrate that the proposed Bow pooling is efficient to improve the performance in point cloud classification, shape retrieval and segmentation tasks and outperforms other existing symmetric functions. Xiao Sun 0011, Zhouhui Lian |
AAAI | 3 |
| 2021 | High-Fidelity and Arbitrary Face EditingabstractCycle consistency is widely used for face editing. However, we observe that the generator tends to find a tricky way to hide information from the original image to satisfy the constraint of cycle consistency, making it impossible to maintain the rich details (e.g., wrinkles and moles) of non-editing areas. In this work, we propose a simple yet effective method named HifaFace to address the above-mentioned problem from two perspectives. First, we relieve the pressure of the generator to synthesize rich details by directly feeding the high-frequency information of the input image into the end of the generator. Second, we adopt an additional discriminator to encourage the generator to synthesize rich details. Specifically, we apply wavelet transformation to transform the image into multi-frequency domains, among which the high-frequency parts can be used to recover the rich details. We also notice that a fine-grained and wider-range control for the attribute is of great importance for face editing. To achieve this goal, we propose a novel attribute regression loss. Powered by the proposed framework, we achieve high-fidelity and arbitrary face editing, outperforming other state-of-the-art approaches. Yue Gao 0006, Fangyun Wei, Jianmin Bao, Shuyang Gu, Dong Chen 0003, Fang Wen 0001, Zhouhui Lian |
CVPR | 7 |
| 2021 | Bidirectional Regression for Arbitrary-Shaped Text Detection
Zhouhui Lian |
ICDAR (4) | 2 |
| 2021 | CentripetalText: An Efficient Text Instance Representation for Scene Text DetectionabstractScene text detection remains a grand challenge due to the variation in text curvatures, orientations, and aspect ratios. One of the hardest problems in this task is how to represent text instances of arbitrary shapes. Although many methods have been proposed to model irregular texts in a flexible manner, most of them lose simplicity and robustness. Their complicated post-processings and the regression under Dirac delta distribution undermine the detection performance and the generalization ability. In this paper, we propose an efficient text instance representation named CentripetalText (CT), which decomposes text instances into the combination of text kernels and centripetal shifts. Specifically, we utilize the centripetal shifts to implement pixel aggregation, guiding the external text pixels to the internal text kernels. The relaxation operation is integrated into the dense regression for centripetal shifts, allowing the correct prediction in a range instead of a specific value. The convenient reconstruction of text contours and the tolerance of prediction errors in our method guarantee the high detection accuracy and the fast inference speed, respectively. Besides, we shrink our text detector into a proposal generation module, namely CentripetalText Proposal Network (CPN), replacing Segmentation Proposal Network (SPN) in Mask TextSpotter v3 and producing more accurate proposals. To validate the effectiveness of our method, we conduct experiments on several commonly used scene text benchmarks, including both curved and multi-oriented text datasets. For the task of scene text detection, our approach achieves superior or competitive performance compared to other existing methods, e.g., F-measure of 86.3% at 40.0 FPS on Total-Text, F-measure of 86.1% at 34.8 FPS on MSRA-TD500, etc. For the task of end-to-end scene text recognition, our method outperforms Mask TextSpotter v3 by 1.1% in F-measure on Total-Text. Zhouhui Lian |
NeurIPS | 3 |
| 2021 | Write Like You: Synthesizing Your Cursive Online Chinese Handwriting via Metric-based Meta LearningabstractAbstract In this paper, we propose a novel Sequence‐to‐Sequence model based on metric‐based meta learning for the arbitrary style transfer of online Chinese handwritings. Unlike most existing methods that treat Chinese handwritings as images and are unable to reflect the human writing process, the proposed model directly handles sequential online Chinese handwritings. Generally, our model consists of three sub‐models: a content encoder, a style encoder and a decoder, which are all Recurrent Neural Networks. In order to adaptively obtain the style information, we introduce an attention‐based adaptive style block which has been experimentally proven to bring considerable improvement to our model. In addition, to disentangle the latent style information from characters written by any writers effectively, we adopt metric‐based meta learning and pre‐train the style encoder using a carefully‐designed discriminative loss function. Then, our entire model is trained in an end‐to‐end manner and the decoder adaptively receives the style information from the style encoder and the content information from the content encoder to synthesize the target output. Finally, by feeding the trained model with a content character and several characters written by a given user, our model can write that Chinese character in the user's handwriting style by drawing strokes one by one like humans. That is to say, as long as you write several Chinese character samples, our model can imitate your handwriting style when writing. In addition, after fine‐tuning the model with a few samples, it can generate more realistic handwritings that are difficult to be distinguished from the real ones. Both qualitative and quantitative experiments demonstrate the effectiveness and superiority of our method. Shusen Tang, Zhouhui Lian |
Comput. Graph. Forum | 2 |
| 2021 | TextPolar: irregular scene text detection using polar representation
Jie Chen 0091, Zhouhui Lian |
Int. J. Document Anal. Recognit. | 2 |
| 2021 | VSRNet: End-to-end video segment retrieval with text query
Xiang Long, Dongliang He, Shilei Wen, Zhouhui Lian |
Pattern Recognit. | 5 |
| 2021 | DeepVecFont: synthesizing high-quality vector fonts via dual-modality learningabstractAutomatic font generation based on deep learning has aroused a lot of interest in the last decade. However, only a few recently-reported approaches are capable of directly generating vector glyphs and their results are still far from satisfactory. In this paper, we propose a novel method, DeepVecFont, to effectively resolve this problem. Using our method, for the first time, visually-pleasing vector glyphs whose quality and compactness are both comparable to human-designed ones can be automatically generated. The key idea of our DeepVecFont is to adopt the techniques of image synthesis, sequence modeling and differentiable rasterization to exhaustively exploit the dual-modality information (i.e., raster images and vector outlines) of vector fonts. The highlights of this paper are threefold. First, we design a dual-modality learning strategy which utilizes both image-aspect and sequence-aspect features of fonts to synthesize vector glyphs. Second, we provide a new generative paradigm to handle unstructured data (e.g., vector glyphs) by randomly sampling plausible synthesis results to get the optimal one which is further refined under the guidance of generated structured data (e.g., glyph images). Finally, qualitative and quantitative experiments conducted on a publicly-available dataset demonstrate that our method obtains high-quality synthesis results in the applications of vector font generation and interpolation, significantly outperforming the state of the art. Yizhi Wang 0006, Zhouhui Lian |
ACM Trans. Graph. | 2 |
| 2020 | Controllable Person Image Synthesis With Attribute-Decomposed GANabstractThis paper introduces the Attribute-Decomposed GAN, a novel generative model for controllable person image synthesis, which can produce realistic person images with desired human attributes (e.g., pose, head, upper clothes and pants) provided in various source inputs. The core idea of the proposed model is to embed human attributes into the latent space as independent codes and thus achieve flexible and continuous control of attributes via mixing and interpolation operations in explicit style representations. Specifically, a new architecture consisting of two encoding pathways with style block connections is proposed to decompose the original hard mapping into multiple more accessible subtasks. In source pathway, we further extract component layouts with an off-the-shelf human parser and feed them into a shared global texture encoder for decomposed latent codes. This strategy allows for the synthesis of more realistic output images and automatic separation of un-annotated attributes. Experimental results demonstrate the proposed method's superiority over the state of the art in pose transfer and its effectiveness in the brand-new task of component attribute transfer. Yifang Men, Yiming Mao 0006, Yuning Jiang 0001, Wei-Ying Ma, Zhouhui Lian |
CVPR | 5 |
| 2020 | Exploring Font-independent Features for Scene Text RecognitionabstractScene text recognition (STR) has been extensively studied in last few years. Many recently-proposed methods are specially designed to accommodate the arbitrary shape, layout and orientation of scene texts, but ignoring that various font (or writing) styles also pose severe challenges to STR. These methods, where font features and content features of characters are tangled, perform poorly in text recognition on scene images with texts in novel font styles. To address this problem, we explore font-independent features of scene texts via attentional generation of glyphs in a large number of font styles. Specifically, we introduce trainable font embeddings to shape the font styles of generated glyphs, with the image feature of scene text only representing its essential patterns. The generation process is directed by the spatial attention mechanism, which effectively copes with irregular texts and generates higher-quality glyphs than existing image-to-image translation methods. Experiments conducted on several STR benchmarks demonstrate the superiority of our method compared to the state of the art. Zhouhui Lian |
ACM Multimedia | 2 |
| 2020 | DeepStroke: Understanding Glyph Structure with Semantic Segmentation and Tabu Search
Zhouhui Lian, Yingmin Tang, Jianguo Xiao |
MMM (1) | 2 |
| 2020 | EasyMesh: An efficient method to reconstruct 3D mesh from a single image
Xiao Sun 0011, Zhouhui Lian |
Comput. Aided Geom. Des. | 2 |
| 2020 | Sketch based modeling and editing via shape space exploration
Zhouhui Lian, Jianguo Xiao |
Multim. Tools Appl. | 2 |
| 2020 | Attribute2Font: creating fonts you want from attributesabstractFont design is now still considered as an exclusive privilege of professional designers, whose creativity is not possessed by existing software systems. Nevertheless, we also notice that most commercial font products are in fact manually designed by following specific requirements on some attributes of glyphs, such as italic, serif, cursive, width, angularity, etc. Inspired by this fact, we propose a novel model, Attribute2Font, to automatically create fonts by synthesizing visually pleasing glyph images according to user-specified attributes and their corresponding values. To the best of our knowledge, our model is the first one in the literature which is capable of generating glyph images in new font styles, instead of retrieving existing fonts, according to given values of specified font attributes. Specifically, Attribute2Font is trained to perform font style transfer between any two fonts conditioned on their attribute values. After training, our model can generate glyph images in accordance with an arbitrary set of font attribute values. Furthermore, a novel unit named Attribute Attention Module is designed to make those generated glyph images better embody the prominent font attributes. Considering that the annotations of font attribute values are extremely expensive to obtain, a semi-supervised learning scheme is also introduced to exploit a large number of unlabeled fonts. Experimental results demonstrate that our model achieves impressive performance on many tasks, such as creating glyph images in new font styles, editing existing fonts, interpolation among different fonts, etc. Yizhi Wang 0006, Yue Gao 0006, Zhouhui Lian |
ACM Trans. Graph. | 3 |
| 2019 | SCFont: Structure-Guided Chinese Font Generation via Deep Stacked NetworksabstractAutomatic generation of Chinese fonts that consist of large numbers of glyphs with complicated structures is now still a challenging and ongoing problem in areas of AI and Computer Graphics (CG). Traditional CG-based methods typically rely heavily on manual interventions, while recentlypopularized deep learning-based end-to-end approaches often obtain synthesis results with incorrect structures and/or serious artifacts. To address those problems, this paper proposes a structure-guided Chinese font generation system, SCFont, by using deep stacked networks. The key idea is to integrate the domain knowledge of Chinese characters with deep generative networks to ensure that high-quality glyphs with correct structures can be synthesized. More specifically, we first apply a CNN model to learn how to transfer the writing trajectories with separated strokes in the reference font style into those in the target style. Then, we train another CNN model learning how to recover shape details on the contour for synthesized writing trajectories. Experimental results validate the superiority of the proposed SCFont compared to the state of the art in both visual and quantitative assessments. Zhouhui Lian, Yingmin Tang, Jianguo Xiao |
AAAI | 2 |
| 2019 | DynTypo: Example-Based Dynamic Text Effects TransferabstractIn this paper, we present a novel approach for dynamic text effects transfer by using example-based texture synthesis. In contrast to previous works that require an input video of the target to provide motion guidance, we aim to animate a still image of the target text by transferring the desired dynamic effects from an observed exemplar. Due to the simplicity of target guidance and complexity of realistic effects, it is prone to producing temporal artifacts such as flickers and pulsations. To address the problem, our core idea is to find a common Nearest-neighbor Field (NNF) that would optimize the textural coherence across all keyframes simultaneously. With the static NNF for video sequences, we implicitly transfer motion properties from source to target. We also introduce a guided NNF search by employing the distance-based weight map and Simulated Annealing (SA) for deep direction-guided propagation to allow intense dynamic effects to be completely transferred with no semantic guidance provided. Experimental results demonstrate the effectiveness and superiority of our method in dynamic text effects transfer through extensive comparisons with state-of-the-art algorithms. We also show the potentiality of our method via multiple experiments for various application domains. Yifang Men, Zhouhui Lian, Yingmin Tang, Jianguo Xiao |
CVPR | 2 |
| 2019 | SRINet: Learning Strictly Rotation-Invariant Representations for Point Cloud Classification and SegmentationabstractPoint cloud analysis has drawn broader attentions due to its increasing demands in various fields. Despite the impressive performance has been achieved on several databases, researchers neglect the fact that the orientation of those point cloud data is aligned. Varying the orientation of point cloud may lead to the degradation of performance, restricting the capacity of generalizing to real applications where the prior of orientation is often unknown. In this paper, we propose the point projection feature, which is invariant to the rotation of the input point cloud. A novel architecture is designed to mine features of different levels. We adopt a PointNet-based backbone to extract global feature for point cloud, and the graph aggregation operation to perceive local shape structure. Besides, we introduce an efficient key point descriptor to assign each point with different response and help recognize the overall geometry. Mathematical analyses and experimental results demonstrate that the proposed method can extract strictly rotation-invariant representations for point cloud recognition and segmentation without data augmentation, and outperforms other state-of-the-art methods. Xiao Sun 0011, Zhouhui Lian, Jianguo Xiao |
ACM Multimedia | 2 |
| 2019 | FontRNN: Generating Large-scale Chinese Fonts via Recurrent Neural NetworkabstractAbstract Despite the recent impressive development of deep neural networks, using deep learning based methods to generate large‐scale Chinese fonts is still a rather challenging task due to the huge number of intricate Chinese glyphs, e.g., the official standard Chinese charset GB18030‐2000 consists of 27,533 Chinese characters. Until now, most existing models for this task adopt Convolutional Neural Networks (CNNs) to generate bitmap images of Chinese characters due to CNN based models' remarkable success in various applications. However, CNN based models focus more on image‐level features while usually ignore stroke order information when writing characters. Instead, we treat Chinese characters as sequences of points (i.e., writing trajectories) and propose to handle this task via an effective Recurrent Neural Network (RNN) model with monotonic attention mechanism, which can learn from as few as hundreds of training samples and then synthesize glyphs for remaining thousands of characters in the same style. Experimental results show that our proposed FontRNN can be used for synthesizing large‐scale Chinese fonts as well as generating realistic Chinese handwritings efficiently. Shusen Tang, Zeqing Xia, Zhouhui Lian, Yingmin Tang, Jianguo Xiao |
Comput. Graph. Forum | 3 |
| 2019 | Irregular scene text detection via attention guided border labeling
Jie Chen 0091, Zhouhui Lian, Yingmin Tang, Jianguo Xiao |
Sci. China Inf. Sci. | 2 |
| 2019 | Boosting scene character recognition by learning canonical forms of glyphs
Zhouhui Lian, Yingmin Tang, Jianguo Xiao |
Int. J. Document Anal. Recognit. | 2 |
| 2019 | Artistic glyph image synthesis via one-stage few-shot learningabstractAutomatic generation of artistic glyph images is a challenging task that attracts many research interests. Previous methods either are specifically designed for shape synthesis or focus on texture transfer. In this paper, we propose a novel model, AGIS-Net, to transfer both shape and texture styles in one-stage with only a few stylized samples. To achieve this goal, we first disentangle the representations for content and style by using two encoders, ensuring the multi-content and multi-style generation. Then we utilize two collaboratively working decoders to generate the glyph shape image and its texture image simultaneously. In addition, we introduce a local texture refinement loss to further improve the quality of the synthesized textures. In this manner, our one-stage model is much more efficient and effective than other multi-stage stacked methods. We also propose a large-scale dataset with Chinese glyph images in various shape and texture styles, rendered from 35 professional-designed artistic fonts with 7,326 characters and 2,460 synthetic artistic fonts with 639 characters, to validate the effectiveness and extendability of our method. Extensive experiments on both English and Chinese artistic glyph image datasets demonstrate the superiority of our model in generating high-quality stylized glyph images against other state-of-the-art methods. Yue Gao 0006, Zhouhui Lian, Yingmin Tang, Jianguo Xiao |
ACM Trans. Graph. | 3 |
| 2019 | EasyFont: A Style Learning-Based System to Easily Build Your Large-Scale Handwriting FontsabstractGenerating personal handwriting fonts with large amounts of characters is a boring and time-consuming task. For example, the official standard GB18030-2000 for commercial font products consists of 27,533 Chinese characters. Consistently and correctly writing out such huge amounts of characters is usually an impossible mission for ordinary people. To solve this problem, we propose a system, EasyFont , to automatically synthesize personal handwriting for all (e.g., Chinese) characters in the font library by learning style from a small number (as few as 1%) of carefully-selected samples written by an ordinary person. Major technical contributions of our system are twofold. First, we design an effective stroke extraction algorithm that constructs best-suited reference data from a trained font skeleton manifold and then establishes correspondence between target and reference characters via a non-rigid point set registration approach. Second, we develop a set of novel techniques to learn and recover users’ overall handwriting styles and detailed handwriting behaviors. Experiments including Turing tests with 97 participants demonstrate that the proposed system generates high-quality synthesis results, which are indistinguishable from original handwritings. Using our system, for the first time, the practical handwriting font library in a user’s personal style with arbitrarily large numbers of Chinese characters can be generated automatically. It can also be observed from our experiments that recently-popularized deep learning based end-to-end methods are not able to properly handle this task, which implies the necessity of expert knowledge and handcrafted rules for many applications. Zhouhui Lian, Jianguo Xiao |
ACM Trans. Graph. | 1 |
| 2019 | Image-driven unsupervised 3D model co-segmentation
Paul L. Rosin, Xianfang Sun, Jianguo Xiao, Zhouhui Lian |
Vis. Comput. | 5 |
| 2018 | A Common Framework for Interactive Texture TransferabstractIn this paper, we present a general-purpose solution to interactive texture transfer problems that better preserves both local structure and visual richness. It is challenging due to the diversity of tasks and the simplicity of required user guidance. The core idea of our common framework is to use multiple custom channels to dynamically guide the synthesis process. For interactivity, users can control the spatial distribution of stylized textures via semantic channels. The structure guidance, acquired by two stages of automatic extraction and propagation of structure information, provides a prior for initialization and preserves the salient structure by searching the nearest neighbor fields (NNF) with structure coherence. Meanwhile, texture coherence is also exploited to maintain similar style with the source image. In addition, we leverage an improved PatchMatch with extended NNF and matrix operations to obtain transformable source patches with richer geometric information at high speed. We demonstrate the effectiveness and superiority of our method on a variety of scenes through extensive comparisons with state-of-the-art algorithms. Yifang Men, Zhouhui Lian, Yingmin Tang, Jianguo Xiao |
CVPR | 2 |
| 2018 | Font Recognition in Natural Images via Transfer Learning
Zhouhui Lian, Yingmin Tang, Jianguo Xiao |
MMM (1) | 2 |
| 2018 | An evaluation of canonical forms for non-rigid 3D shape retrievalabstractCanonical forms attempt to factor out a non-rigid shape’s pose, giving a pose-neutral shape. This opens up the possibility of using methods originally designed for rigid shape retrieval for the task of non-rigid shape retrieval. We extend our recent benchmark for testing canonical form algorithms. Our new benchmark is used to evaluate a greater number of state-of-the-art canonical forms, on five recent non-rigid retrieval datasets, within two different retrieval frameworks. A total of fifteen different canonical form methods are compared. We find that the difference in retrieval accuracy between different canonical form methods is small, but varies significantly across different datasets. We also find that efficiency is the main difference between the methods. David Pickup, Xianfang Sun, Paul L. Rosin, Ralph R. Martin, Zhi-Quan Cheng, Zhouhui Lian, Sipin Nie, Longcun Jin, Gil Shamai, Yusuf Sahillioglu, Ladislav Kavan |
Graph. Model. | 7 |
| 2018 | Text effects transfer via distribution-aware texture synthesis
Shuai Yang 0001, Jiaying Liu 0001, Zhouhui Lian, Zongming Guo |
Comput. Vis. Image Underst. | 3 |
| 2017 | Incremental Kernel Null Space Discriminant Analysis for Novelty Detection
Zhouhui Lian, Jianguo Xiao |
CVPR | 2 |
| 2017 | Awesome Typography: Statistics-Based Text Effects TransferabstractIn this work, we explore the problem of generating fantastic special-effects for the typography. It is quite challenging due to the model diversities to illustrate varied text effects for different characters. To address this issue, our key idea is to exploit the analytics on the high regularity of the spatial distribution for text effects to guide the synthesis process. Specifically, we characterize the stylized patches by their normalized positions and the optimal scales to depict their style elements. Our method first estimates these two features and derives their correlation statistically. They are then converted into soft constraints for texture transfer to accomplish adaptive multi-scale texture synthesis and to make style element distribution uniform. It allows our algorithm to produce artistic typography that fits for both local texture patterns and the global spatial distribution in the example. Experimental results demonstrate the superiority of our method for various text effects over conventional style transfer methods. In addition, we validate the effectiveness of our algorithm with extensive artistic typography library generation. Shuai Yang 0001, Jiaying Liu 0001, Zhouhui Lian, Zongming Guo |
CVPR | 3 |
| 2017 | Structure-Aware Image Resizing for Chinese Characters
Chengdong Liu, Zhouhui Lian, Yingmin Tang, Jianguo Xiao |
MMM (1) | 2 |
| 2016 | Shape Retrieval of Non-rigid 3D Human Modelsabstract3D models of humans are commonly used within computer graphics and vision, and so the ability to distinguish between body shapes is an important shape retrieval problem. We extend our recent paper which provided a benchmark for testing non-rigid 3D shape retrieval algorithms on 3D human models. This benchmark provided a far stricter challenge than previous shape benchmarks. We have added 145 new models for use as a separate training set, in order to standardise the training data used and provide a fairer comparison. We have also included experiments with the FAUST dataset of human scans. All participants of the previous benchmark study have taken part in the new tests reported here, many providing updated results using the new data. In addition, further participants have also taken part, and we provide extra analysis of the retrieval results. A total of 25 different shape retrieval methods are compared. David Pickup, Xianfang Sun, Paul L. Rosin, Ralph R. Martin, Zhouhui Lian, Masaki Aono, A. Ben Hamza, Alexander M. Bronstein, Michael M. Bronstein, S. Bu, Umberto Castellani, S. Cheng, Valeria Garro, Andrea Giachetti 0001, Afzal Godil, Luca Isaia, Henry Johan, Long Lai, Bo Li 0013, Chenfeng Li, Hai-Sheng Li 0002, Roee Litman, Yijuan Lu, Li Sun 0004, Gary K. L. Tam, Atsushi Tatsuma, Jianbo Ye |
Int. J. Comput. Vis. | 6 |
| 2015 | Content-independent font recognition on a single Chinese character using sparse representationabstractFont recognition on a single Chinese character is a challenging task especially when the identity of the character is unknown and the number of possible font types is huge. In this paper, we propose a novel method using multi-scale sparse representation to solve the problem of large-scale font recognition on a single unknown Chinese character. Specifically, we first apply a saliency-based sampling approach, which exploits the saliency information of character contours, to segment local patches in multiple scales from salient regions. Then, corresponding local descriptors are extracted by implementing Sobel and Prewitt operators in 4 directions. After encoding the local descriptors into sparse codes, max pooling and spatial pyramid matching are employed to pool them into a sparse representation. Finally, a multi-scale sparse representation is obtained by concatenating three sparse representations which respectively correspond to three particular scales of local patches, and then the linear SVM classifier is utilized for font classification. Experiments performed on a large-scale database consisting of Chinese character images in 160 fonts show that our method achieves significantly better performance compared to the state of the art. Moreover, we also carry out experiments on a subset of the database to demonstrate the effectiveness of our saliency-based sampling approach and the proposed Sobel-Prewitt feature. Weikang Song, Zhouhui Lian, Yingmin Tang, Jianguo Xiao |
ICDAR | 2 |
| 2015 | Aesthetic Visual Quality Evaluation of Chinese Handwritings
Rongju Sun, Zhouhui Lian, Yingmin Tang, Jianguo Xiao |
IJCAI | 2 |
| 2015 | Text Detection in Natural Images Using Localized Stroke Width Transform
Wenyan Dong, Zhouhui Lian, Yingmin Tang, Jianguo Xiao |
MMM (1) | 2 |
| 2014 | FlexiFont: a flexible system to generate personal font librariesabstractThis paper proposes FlexiFont, a system designed to generate personal font libraries from the camera-captured character images. Compared with existing methods, our system is able to process most kinds of languages and the generated font libraries can be extended by adding new characters based on the user's requirement. Moreover, digital cameras instead of scanners are chosen as the input devices, so that it is more convenient for common people to use the system. First of all, the users should choose a default template or define their own templates, then write the characters on the printed templates according to the certain instructions. After the users upload the photos of the templates with written characters, the system will automatically correct the perspective and split the whole photo into a set of individual character images. As the final step, FlexiFont will denoise, vectorize, and normalize each character image before storing it into a TrueType file. Experimental results demonstrate the robustness and efficiency of our system. Wanqiong Pan, Zhouhui Lian, Rongju Sun, Yingmin Tang, Jianguo Xiao |
ACM Symposium on Document Engineering | 2 |
| 2014 | Non-rigid point set registration for Chinese characters using structure-guided coherent point driftabstractThis paper proposes a non-rigid point set registration method called Structure-Guided Coherent Point Drift (SGCPD). The key idea of our method is to utilize structural information and combine the global and local point registrations together to improve the original Coherent Point Drift (CPD) algorithm. Specifically, given two point sets, we first align them using the CPD method with Localized Operator (CPDLO). Then we divide the target point set into several subsets and apply CPDLO to each subset. Finally, we implement the above two procedures until convergence. In this manner, more detailed information can be well exploited and thus higher registration accuracy can be achieved. Experimental results demonstrate that our method outperforms the original CPD approach on both point registration accuracy and skeleton decomposition accuracy for Chinese characters. Zhouhui Lian, Yingmin Tang, Jianguo Xiao |
ICIP | 2 |
| 2014 | A Data-Driven Personalized Digital Ink for Chinese Characters
Tianyang Yi, Zhouhui Lian, Yingmin Tang, Jianguo Xiao |
MMM (1) | 2 |
| 2014 | Skeleton-guided vectorization of Chinese calligraphy imagesabstractHow to automatically generate compact and high-quality vectorization for Chinese calligraphy images is a challenging problem, since these images usually suffer from noisy contours and discontinuous strokes. In this paper, we propose a skeleton guided approach to vectorize Chinese calligraphy images. Since the skeleton reflects the writing trace and it is less influenced by contour noises, our method could extract the important writing style from the noisy contours. Specifically, in our method, the calligraphy image is first preprocessed by binarization and denoising. Then salient contour points are detected by a novel algorithm. Afterwards, under the guidance of skeleton information, the salient points are classified into corner points and joint points. Finally, a dynamic curve fitting procedure is applied to generate the vectorization result. Experimental results demonstrate that our skeleton-guided approach could automatically distinguish tiny features from contour noises and thus obtains more visually satisfactory performance compared to other existing methods. Wanqiong Pan, Zhouhui Lian, Yingmin Tang, Jianguo Xiao |
MMSP | 2 |
| 2013 | Automatic Correspondence Finding for Chinese Characters Using Graph MatchingabstractAutomatically establishing correspondence between Chinese characters is a challenging task. In this paper, we propose a novel method to solve this problem. Given two Chinese characters, we first extract and properly prune the skeleton of each character to get the key points and the connectivity relations of these points. Then, the similarity between each pair of key points is calculated via the comparison of their local features. Afterwards, a set of edges are constructed by considering both the connectivity relations and k nearest neighbors (k-nn) of each point. Finally, correspondence between two characters is established by applying a guided graph matching algorithm. Experimental results demonstrate the effectiveness of our method for the correspondence problem of Chinese characters in both printing and handwritten styles. Moreover, we also show that our method can be utilized to automatically extract strokes from Chinese characters. Zhouhui Lian, Yingmin Tang, Jianguo Xiao |
ICIG | 2 |
| 2013 | Feature-Preserved 3D Canonical Form
Zhouhui Lian, Afzal Godil, Jianguo Xiao |
Int. J. Comput. Vis. | 1 |
| 2013 | CM-BOF: visual similarity-based 3D shape retrieval using Clock Matching and Bag-of-Features
Zhouhui Lian, Afzal Godil, Xianfang Sun, Jianguo Xiao |
Mach. Vis. Appl. | 1 |
| 2013 | A comparison of methods for non-rigid 3D shape retrieval
Zhouhui Lian, Afzal Godil, Benjamin Bustos, Mohamed Daoudi, Jeroen Hermans, Shun Kawamura, Yukinori Kurita, Guillaume Lavoué, Hien Van Nguyen, Ryutarou Ohbuchi, Yuki Ohkita, Yuya Ohishi, Fatih Porikli, Martin Reuter 0001, Ivan Sipiran, Dirk Smeets, Paul Suetens, Hedi Tabia, Dirk Vandermeulen |
Pattern Recognit. | 1 |
| 2012 | A new convexity measurement for 3D meshesabstractThis paper presents a novel convexity measurement for 3D meshes. The new convexity measure is calculated by minimizing the ratio of the summed area of valid regions in a mesh's six views, which are projected on faces of the bounding box whose edges are parallel to the coordinate axes, to the sum of three orthogonal projected areas of the mesh. The complete definition, theoretical analysis, and a computing algorithm of our convexity measure are explicitly described. This paper also proposes a new 3D shape descriptor CD (i.e., Convexity Distribution) based on the distribution of above-mentioned ratios, which are computed by randomly rotating the mesh around its center, to better describe the object's convexity-related properties compared to existing convexity measurements. Our experiments not only show that the proposed convexity measure corresponds well with human intuition, but also demonstrate the effectiveness of the new convexity measure and the new shape descriptor by significantly improving the performance of other methods in the application of 3D shape retrieval. Zhouhui Lian, Afzal Godil, Paul L. Rosin, Xianfang Sun |
CVPR | 1 |
| 2010 | Non-rigid 3D shape retrieval using Multidimensional Scaling and Bag-of-FeaturesabstractMatching non-rigid shapes is a challenging research field in content-based 3D object retrieval. In this paper, we present an image-based method to effectively address this problem. Multidimensional Scaling (MDS) and Principal Component Analysis (PCA) are first applied to each object to calculate its canonical form, which is afterward represented by 66 depth-buffer images captured on the vertices of an unit geodesic sphere. Then, each image is described as a word histogram obtained by the vector quantization of the image's salient local features. Finally, a multi-view shape matching scheme is carried out to measure the dissimilarity between two models. Experimental results on the McGill Articulated Shape Benchmark database demonstrate that, our method obtains better retrieval performance compared to the state-of-the-art. Zhouhui Lian, Afzal Godil, Xianfang Sun |
ICIP | 1 |
| 2010 | Visual Similarity Based 3D Shape Retrieval Using Bag-of-FeaturesabstractThis paper presents a novel 3D shape retrieval method, which uses Bag-of-Features and an efficient multi-view shape matching scheme. In our approach, a properly normalized object is first described by a set of depth-buffer views captured on the surrounding vertices of a given unit geodesic sphere. We then represent each view as a word histogram generated by the vector quantization of the view's salient local features. The dissimilarity between two 3D models is measured by the minimum distance of their all (24) possible matching pairs. This paper also investigates several critical issues including the influence of the number of views, codebook, training data, and distance function. Experiments on four commonly-used benchmarks demonstrate that: 1) Our approach obtains superior performance in searching for rigid models. 2) The local feature and global feature based methods are somehow complementary. Moreover, a linear combination of them significantly outperforms the state-of-the-art in terms of retrieval accuracy. Zhouhui Lian, Afzal Godil, Xianfang Sun |
Shape Modeling International | 1 |
| 2010 | Rectilinearity of 3D Meshes
Zhouhui Lian, Paul L. Rosin, Xianfang Sun |
Int. J. Comput. Vis. | 1 |