VLDB 2026 Research / reviewers in the wild / expert
Xiaonan Fang 0001
dblp:169/7194-1 · also Xiao-Nan Fang 0001
· DBLP profile ↗
11ranked-venue papers
5as first author
8since 2021 · last 2025
0000-0002-4787-5977ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 9 · 4 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | $C^{2}D$C2D: Context-Aware Concept Decomposition for Personalized Text-to-Image SynthesisabstractConcept decomposition is a technique for personalized text-to-image synthesis which learns textual embeddings of subconcepts from images that depicting an original concept. The learned subconcepts can then be composed to create new images. However, existing methods fail to address the issue of contextual conflicts when subconcepts from different sources are combined because contextual information remains encapsulated within the subconcept embeddings. To tackle this problem, we propose a Context-aware Concept Decomposition ($C^{2}D$C2D) framework. Specifically, we introduce a Similarity-Guided Divergent Embedding (SGDE) method to obtain subconcept embeddings. Then, we eliminate the latent contextual dependence between the subconcept embeddings and reconstruct the contextual information using an independent contextual embedding. This independent context can be combined with various subconcepts, enabling more controllable text-to-image synthesis based on subconcept recombination. Extensive experimental results demonstrate that our method outperforms existing approaches in both image quality and contextual consistency. Jiang Xin, Xiaonan Fang 0001, Xueling Zhu, Ju Ren 0001, Yaoxue Zhang |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2025 | Video sketching using multi-domain guidance and implicit encoding
Xiaonan Fang 0001, Muhan Chang |
Vis. Comput. | 1 |
| 2025 | Privacy-aware Real-Time Target Person Matting in Multi-Person Scenes Using Dual Encoder-Decoder Networks
Jiang Xin, Xiaonan Fang 0001, Xueling Zhu, Ruyi Dai, Ju Ren 0001, Wenzhen Yue, Yaoxue Zhang |
Vis. Comput. | 2 |
| 2024 | Single-Video Temporal Consistency Enhancement with Rolling Guidance
Xiaonan Fang 0001, Song-Hai Zhang |
CVM (2) | 1 |
| 2023 | Learning Local Contrast for Crisp Edge Detection
Xiaonan Fang 0001, Song-Hai Zhang |
J. Comput. Sci. Technol. | 1 |
| 2022 | Local Homography Estimation on User-Specified Textureless Regions
Zheng Chen 0016, Xiaonan Fang 0001, Song-Hai Zhang |
J. Comput. Sci. Technol. | 2 |
| 2022 | User-Guided Deep Human Image Matting Using Arbitrary TrimapsabstractImage matting is widely studied for accurate foreground extraction. Most algorithms, including deep-learning based solutions, require a carefully edited trimap. Recent works attempt to combine the segmentation stage and matting stage in one CNN model, but errors occurring at the segmentation stage lead to unsatisfactory matte. We propose a user-guided approach for practical human matting. More precisely, we provide a good automatic initial matting and a natural way of interaction that reduces the workload of drawing trimaps and allows users to guide the matting in ambiguous situation. We also combine the segmentation and matting stage in an end-to-end CNN architecture and introduce a residual-learning module to support convenient stroke-based interaction. The proposed model learns to propagate the input trimap and modify the deep image features, which can efficiently correct the segmentation errors. Our model supports arbitrary forms of trimaps from carefully edited to totally unknown maps. Our model also allows users to choose from different foreground estimations according to their preference. We collected a large human matting dataset consisting of 12K real-world human images with complex background and human-object relations. The proposed model is trained on the new dataset with a novel trimap generation strategy that enables the model to tackle different test situations and highly improves the interaction efficiency. Our method outperforms other state-of-the-art automatic methods and achieve competitive accuracy when high-quality trimaps are provided. Experiments indicate that our interactive matting strategy is superior to separately estimating the trimap and alpha matte using two models. Xiaonan Fang 0001, Song-Hai Zhang, Tao Chen 0015, Xian Wu 0004, Ariel Shamir, Shi-Min Hu 0001 |
IEEE Trans. Image Process. | 1 |
| 2021 | Prominent Structures for Video Analysis and EditingabstractWe present prominent structures in video, a representation of visually strong, spatially sparse and temporally stable structural units, for use in video analysis and editing. With a novel quality measurement of prominent structures in video, we develop a general framework for prominent structure computation, and an efficient hierarchical structure alignment algorithm between a pair of videos. The prominent structural unit map is proposed to encode both binary prominence guidance and numerical strength and geometry details for each video frame. Even though the detailed appearance of videos could be visually different, the proposed alignment algorithm can find matched prominent structure sub-volumes. Prominent structures in video support a wide range of video analysis and editing applications including graphic match-cut between successive videos, instant cut editing, finding transition portals from a video collection, structure-aware video re-ranking, visualizing human action differences, etc. Miao Wang 0004, Xiaonan Fang 0001, Ariel Shamir, Shi-Min Hu 0001 |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2020 | JMNet: A joint matting network for automatic human mattingabstractWe propose a novel end-to-end deep learning framework, the Joint Matting Network (JMNet), to automatically generate alpha mattes for human images. We utilize the intrinsic structures of the human body as seen in images by introducing a pose estimation module, which can provide both global structural guidance and a local attention focus for the matting task. Our network model includes a pose network, a trimap network, a matting network, and a shared encoder to extract features for the above three networks. We also append a trimap refinement module and utilize gradient loss to provide a sharper alpha matte. Extensive experiments have shown that our method outperforms state-of-theart human matting techniques; the shared encoder leads to better performance and lower memory costs. Our model can process real images downloaded from the Internet for use in composition applications. Xian Wu 0004, Xiaonan Fang 0001, Tao Chen 0015 |
Comput. Vis. Media | 2 |
| 2019 | Learning Explicit Smoothing Kernels for Joint Image FilteringabstractAbstract Smoothing noises while preserving strong edges in images is an important problem in image processing. Image smoothing filters can be either explicit (based on local weighted average) or implicit (based on global optimization). Implicit methods are usually time‐consuming and cannot be applied to joint image filtering tasks, i.e., leveraging the structural information of a guidance image to filter a target image. Previous deep learning based image smoothing filters are all implicit and unavailable for joint filtering. In this paper, we propose to learn explicit guidance feature maps as well as offset maps from the guidance image and smoothing parameter that can be utilized to smooth the input itself or to filter images in other target domains. We design a deep convolutional neural network consisting of a fully‐convolution block for guidance and offset maps extraction together with a stacked spatially varying deformable convolution block for joint image filtering. Our models can approximate several representative image smoothing filters with high accuracy comparable to state‐of‐the‐art methods, and serve as general tools for other joint image filtering tasks, such as color interpolation, depth map upsampling, saliency map upsampling, flash/non‐flash image denoising and RGB/NIR image denoising. Xiaonan Fang 0001, Miao Wang 0004, Ariel Shamir, Shi-Min Hu 0001 |
Comput. Graph. Forum | 1 |
| 2017 | Practical automatic background substitution for live videoabstractIn this paper we present a novel automatic background substitution approach for live video. The objective of background substitution is to extract the foreground from the input video and then combine it with a new background. In this paper, we use a color line model to improve the Gaussian mixture model in the background cut method to obtain a binary foreground segmentation result that is less sensitive to brightness differences. Based on the high quality binary segmentation results, we can automatically create a reliable trimap for alpha matting to refine the segmentation boundary. To make the composition result more realistic, an automatic foreground color adjustment step is added to make the foreground look consistent with the new background. Compared to previous approaches, our method can produce higher quality binary segmentation results, and to the best of our knowledge, this is the first time such an automatic and integrated background substitution system has been proposed which can run in real time, which makes it practical for everyday applications. Hao-Zhi Huang 0001, Xiaonan Fang 0001, Yufei Ye 0001, Song-Hai Zhang, Paul L. Rosin |
Comput. Vis. Media | 2 |