EDBT 2026 Demo / reviewers in the wild / expert
Xueming Li 0002
dblp:67/2097-2
· DBLP profile ↗
33ranked-venue papers
0as first author
21since 2021 · last 2026
0000-0003-1058-2799ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 18 · 12 since 2021Artificial intelligence and machine learning · 14 · 8 since 2021Computer networks · 2 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | RUCLIP: Robust concept unlearning in CLIP via semantic anchors
Yue Zhang 0016, Qinghong Yin, Xianlin Zhang, Xueming Li 0002 |
Expert Syst. Appl. | 6 |
| 2026 | CRColor: Cycle reference learning for exemplar-based image colorization
Mingdao Wang, Xianlin Zhang, Xueming Li 0002, Yue Zhang 0016 |
Neurocomputing | 4 |
| 2026 | Efficient motion-centric CLIP for compressed video action recognition
Jiangwan Zhou, Xueming Li 0002 |
Pattern Recognit. | 4 |
| 2025 | Sketch-1-to-3: One Single Sketch to 3D Detailed Face Reconstructionabstract3D face reconstruction from a single sketch is a critical yet underexplored task with significant practical applications. The primary challenges stem from the substantial modality gap between 2D sketches and 3D facial structures, including: (1) accurately extracting facial keypoints from 2D sketches; (2) preserving diverse facial expressions and fine-grained texture details; and (3) training a high-performing model with limited data. In this paper, we propose Sketch-1-to-3, a novel framework for realistic 3D face reconstruction from a single sketch, to address these challenges. Specifically, we first introduce the Geometric Contour and Texture Detail (GCTD) module, which enhances the extraction of geometric contours and texture details from facial sketches. Additionally, we design a deep learning architecture with a domain adaptation module and a tailored loss function to align sketches with the 3D facial space, enabling high-fidelity expression and texture reconstruction. To facilitate evaluation and further research, we construct SketchFaces, a real hand-drawn facial sketch dataset, and Syn-SketchFaces, a synthetic facial sketch dataset. Extensive experiments demonstrate that Sketch-1-to-3 achieves state-of-the-art performance in sketch-based 3D face reconstruction. Liting Wen, Zimo Yang, Xianlin Zhang, Chi Ding, Mingdao Wang, Xueming Li 0002 |
MMAsia | 6 |
| 2025 | Spcolor: Semantic prior guided exemplar-based image colorization
Xianlin Zhang, Mingdao Wang, Xueming Li 0002, Yue Zhang 0016 |
Pattern Recognit. | 4 |
| 2024 | Modeling the skeleton-language uncertainty for 3D action recognition
Mingdao Wang, Xianlin Zhang, Xueming Li 0002, Yue Zhang 0016 |
Neurocomputing | 4 |
| 2024 | Exemplar-based video colorization with long-term spatiotemporal dependency
Xueming Li 0002, Xianlin Zhang, Mingdao Wang, Jiatong Han, Yue Zhang 0016 |
Knowl. Based Syst. | 2 |
| 2024 | Learning Representations by Contrastive Spatio-Temporal Clustering for Skeleton-Based Action RecognitionabstractSelf-supervised representation learning has proven constructive for skeleton-based action recognition. For better performance, existing methods mainly focus on 1) multi-modal data augmentations and 2) triplet contrastive samples construction. However, designing these strategies is always heuristics and hard. Instead of exploring more similar strategies, this paper addresses this issue with a different view and proposes a novel Contrastive Spatio-Temporal Clustering (CSTC) module. CSTC constructs a supervised signal (pseudo-label) of action sequences in an online clustering manner, and it is complementary to the recent data augmentations or triplet contrastive samples construction strategies. Specifically, CSTC can be formulated as an optimal transport problem. we introduce the spatio-temporal regularizations into the original optimal transport term to guide the pseudo-label generation, i.e., a semantic regularization learned by frame index is proposed to constrain the frame order, and a prior normal distribution regularization based on sampling characteristics of samples is proposed to maintain the dependability of spatial cluster assignments. Furthermore, to enhance the learning of latent features, we propose a Bidirectional Cross-modal Clustering Consistency Objective (B3CO) to enforce cluster assignments consistency for different modalities of the same sample. Last, since fusing spatial and temporal clustering losses directly during back-propagation will confuse the learned dimension-specific semantics, we propose a simple yet effective training strategy to fix it by training the model using these two losses alternately. By integrating the above designs into the MoCo framework, we propose a Contrastive Spatio-Temporal Clustering Network (CSTCN), which can excavate cross-modal discriminative spatio-temporal features in the clustering space. Experimental results on NTU RGB+D 60, NTU RGB+D 120, and PKU-MMD II datasets show that CSTCN achieves state-of-the-art performance in both single- and multi-modal models, especially in the KNN and semi-supervised evaluation protocols. Besides, the key module CSTC shows good generalization capability, and achieves consistent performance improvement on the basis of several state-of-the-art methods which focus on data augmentations and triplet contrastive samples construction. Mingdao Wang, Xueming Li 0002, Xianlin Zhang, Lei Ma 0003, Yue Zhang 0016 |
IEEE Trans. Multim. | 2 |
| 2024 | PlanNet: A Generative Model for Component-Based Plan SynthesisabstractWe propose a novel generative model named as PlanNet for component-based plan synthesis. The proposed model consists of three modules, a wave function collapse algorithm to create large-scale wireframe patterns as the embryonic forms of floor plans, and two deep neural networks to outline the plausible boundary from each squared pattern, and meanwhile estimate the potential semantic labels for the components. In this manner, we use PlanNet to generate a large-scale component-based plan dataset with 10 K examples. Given an input boundary, our method retrieves dataset plan examples with similar configurations to the input, and then transfers the space layout from a user-selected plan example to the input. Benefiting from our interactive workflow, users can recursively subdivide individual components of the plans to enrich the plan contents, thus designing more complex plans for larger scenes. Moreover, our method also adopts a random selection algorithm to make the variations on semantic labels of the plan components, aiming at enriching the 3D scenes that the output plans are suited for. To demonstrate the quality and versatility of our generative model, we conduct intensive experiments, including the analysis of plan examples and their evaluations, plan synthesis with both hard and soft boundary constraints, and 3D scenes designed with the plan subdivision on different scales. We also compare our results with the state-of-the-art floor plan synthesis methods to validate the feasibility and efficacy of the proposed generative model. Qiang Fu 0004, Shuhan He, Xueming Li 0002, Hongbo Fu 0001 |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2024 | Magic Furniture: Design Paradigm of Multi-Function AssemblyabstractAssembly-based furniture with movable parts enables shape and structure reconfiguration, thus supporting multiple functions. Although a few attempts have been made for facilitating the creation of multi-function objects, designing such a multi-function assembly with the existing solutions often requires high imagination of designers. We develop the Magic Furniture system for users to easily create such designs simply given multiple cross-category objects. Our system automatically leverages the given objects as references to generate a 3D model with movable boards driven by back-and-forth movement mechanisms. By controlling the states of these mechanisms, a designed multi-function furniture object can be reconfigured to approximate the shapes and functions of the given objects. To ensure the designed furniture easy to transform between different functions, we perform an optimization algorithm to choose a proper number of movable boards and determine their shapes and sizes, following a set of design guidelines. We demonstrate the effectiveness of our system through various multi-function furniture designed with different sets of reference inputs and various movement constraints. We also evaluate the design results through several experiments including comparative and user studies. Qiang Fu 0004, Fan Zhang 0063, Xueming Li 0002, Hongbo Fu 0001 |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2023 | SSHRF-GAN: Spatial-Spectral Joint High Receptive Field GAN for Old Photo Restoration
Duren Wen, Xueming Li 0002, Yue Zhang 0016 |
PRCV (9) | 2 |
| 2023 | Component-aware generative autoencoder for structure hybrid and shape completionabstractAssembling components of man-made objects to create new structures or complete 3D shapes is a popular approach in 3D modeling techniques. Recently, leveraging deep neural networks for assembly-based 3D modeling has been widely studied. However, exploring new component combinations even across different categories is still challenging for most of the deep-learning-based 3D modeling methods. In this paper, we propose a novel generative autoencoder that tackles the component combinations for 3D modeling of man-made objects. We use the segmented input objects to create component volumes that have redundant components and random configurations. By using the input objects and the associated component volumes to train the autoencoder, we can obtain an object volume consisting of components with proper quality and structure as the network output. Such a generative autoencoder can be applied to either multiple object categories for structure hybrid or a single object category for shape completion. We conduct a series of evaluations and experimental results to demonstrate the usability and practicability of our method. Fan Zhang 0063, Qiang Fu 0004, Yang Liu 0132, Xueming Li 0002 |
Graph. Model. | 4 |
| 2023 | Fuzzy-based indoor scene modeling with differentiated examplesabstractWell-designed indoor scenes incorporate interior design knowledge, which has been an essential prior for most indoor scene modeling methods. However, the layout qualities of indoor scene datasets are often uneven, and most existing data-driven methods do not differentiate indoor scene examples in terms of quality. In this work, we aim to explore an approach that leverages datasets with differentiated indoor scene examples for indoor scene modeling. Our solution conducts subjective evaluations on lightweight datasets having various room configurations and furniture layouts, via pairwise comparisons based on fuzzy set theory. We also develop a system to use such examples to guide indoor scene modeling using user-specified objects. Specifically, we focus on object groups associated with certain human activities, and define room features to encode the relations between the position and direction of an object group and the room configuration. To perform indoor scene modeling, given an empty room, our system first assesses it in terms of the user-specified object groups, and then places associated objects in the room guided by the assessment results. A series of experimental results and comparisons to state-of-the-art indoor scene synthesis methods are presented to validate the usefulness and effectiveness of our approach. Qiang Fu 0004, Shuhan He, Hongbo Fu 0001, Xueming Li 0002, Zhigang Deng 0001 |
Comput. Vis. Media | 4 |
| 2023 | AABLSTM: A Novel Multi-task Based CNN-RNN Deep Model for Fashion AnalysisabstractWith the rapid growth of online commerce and fashion-related applications, visual clothing analysis and recognition has become a hotspot in computer vision. In this paper, we propose a novel AABLSTM network, which is based on deep CNN-RNN, to solve the visual fashion analysis of clothing category classification, attribute detection, and landmark localization. The designed fashion model is leveraged with the multi-task driven mechanism as follows: firstly, a bidirectional LSTM (Bi-LSTM) branch is proposed for efficiently mining the semantic association between related attributes so as to improve the precision of clothing category classification and attribute detection; then, an imitated hourglass sub-network of “down-up sampling” is constructed for boosting the accuracy of fashion landmark localization; and finally, a specially designed multi-loss function is constructed to better optimize the network training. Extensive experimental results on large-scale fashion datasets demonstrate the superior performance of our approach. Xianlin Zhang, Mengling Shen, Xueming Li 0002, Xiaojie Wang 0006 |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2023 | Rule of thirds-aware reinforcement learning for image aesthetic cropping
Xuewei Li 0005, Gang Zhang 0008, Yuquan Wu, Xueming Li 0002 |
Vis. Comput. | 4 |
| 2022 | TIR: A Two-Stage Insect Recognition Method for Convolutional Neural Network
Yunqi Feng 0002, Yang Liu 0132, Xianlin Zhang, Xueming Li 0002 |
PRCV (2) | 4 |
| 2022 | Indoor layout programming via virtual navigation detectors
Qiang Fu 0004, Hongbo Fu 0001, Zhigang Deng 0001, Xueming Li 0002 |
Sci. China Inf. Sci. | 4 |
| 2022 | Reinforcement learning cropping method based on comprehensive feature and aesthetics assessmentabstractAbstract Automatic image cropping can change the composition to improve the aesthetic quality of the images. Most of the existing automatic image cropping methods based on specific features need to generate a large number of candidate cropping windows. It is very time‐consuming and can only produce a limited aspect ratio results. In the face of these situations, a reinforcement learning cropping method based on comprehensive feature and aesthetics assessment is proposed. It does not need to produce a large number of candidate windows. Its gradually cropping mode is more in line with the process of image cropping by human. What is more, the proposed method takes the image aesthetic assessment into consideration. Experimental results show that the proposed method improves the cropping efficiency and achieves excellent cropping effect on the open Flickr Cropping Dataset and CUHK Image Cropping Dataset. The proposed method can overcome the shortages of existing methods. Xueming Li 0002, Xuewei Li 0005 |
IET Image Process. | 2 |
| 2022 | Hierarchical graph attention network with pseudo-metapath for skeleton-based action recognition
Mingdao Wang, Xueming Li 0002, Xianlin Zhang, Yue Zhang 0016 |
Neurocomputing | 2 |
| 2022 | A deformable CNN-based triplet model for fine-grained sketch-based image retrieval
Xianlin Zhang, Mengling Shen, Xueming Li 0002, Fangxiang Feng |
Pattern Recognit. | 3 |
| 2021 | A method of inpainting moles and acne on the high-resolution face photosabstractAbstract With the rapid development of mobile phones, more and more high‐resolution photos are taken. The demand for high‐resolution image inpainting is becoming increasingly urgent. In order to repair high‐resolution face images automatically and quickly, this paper proposes an improved generative adversarial networks method. Firstly, we made a high‐resolution dataset for training and testing, and abandoned the traditional 256256 size data. Secondly, since the existing methods can only repair the mask with fixed size and shape on the image, when the global average pooling layer is used in the network, the improved network can repair the moles and acne with arbitrary sizes and shapes on the human face photos. Finally, in order to achieve optimal performance of the network, a mixed loss function is used in training. The experimental results prove that our method has not only achieved good results in qualitative results, but also achieved excellent results in quantitative results. Xuewei Li 0005, Xueming Li 0002, Xianlin Zhang, Yang Liu 0132, Jiayi Liang, Ziliang Guo, Keyu Zhai |
IET Image Process. | 2 |
| 2020 | Interactive Design and Preview of Colored Snapshots of Indoor ScenesabstractAbstract This paper presents an interactive system for quickly designing and previewing colored snapshots of indoor scenes. Different from high‐quality 3D indoor scene rendering, which often takes several minutes to render a moderately complicated scene under a specific color theme with high‐performance computing devices, our system aims at improving the effectiveness of color theme design of indoor scenes and employs an image colorization approach to efficiently obtain high‐resolution snapshots with editable colors. Given several pre‐rendered, multi‐layer, gray images of the same indoor scene snapshot, our system is designed to colorize and merge them into a single colored snapshot. Our system also assists users in assigning colors to certain objects/components and infers more harmonious colors for the unassigned objects based on pre‐collected priors to guide the colorization. The quickly generated snapshots of indoor scenes provide previews of interior design schemes with different color themes, making it easy to determine the personalized design of indoor scenes. To demonstrate the usability and effectiveness of this system, we present a series of experimental results on indoor scenes of different types, and compare our method with a state‐of‐the‐art method for indoor scene material and color suggestion and offline/online rendering software packages. Qiang Fu 0004, Hai Yan, Hongbo Fu 0001, Xueming Li 0002 |
Comput. Graph. Forum | 4 |
| 2020 | Human-centric metrics for indoor scene assessment and synthesis
Qiang Fu 0004, Hongbo Fu 0001, Hai Yan, Xiaowu Chen 0001, Xueming Li 0002 |
Graph. Model. | 6 |
| 2019 | A survey on freehand sketch recognition and retrieval
Xianlin Zhang, Xueming Li 0002, Yang Liu 0132, Fangxiang Feng |
Image Vis. Comput. | 2 |
| 2018 | Better freehand sketch synthesis for sketch-based image retrieval: Beyond image edges
Xianlin Zhang, Xueming Li 0002, Xuewei Li 0005, Mengling Shen |
Neurocomputing | 2 |
| 2016 | ForgetMeNot: Memory-Aware Forensic Facial Sketch MatchingabstractWe investigate whether it is possible to improve the performance of automated facial forensic sketch matching by learning from examples of facial forgetting over time. Forensic facial sketch recognition is a key capability for law enforcement, but remains an unsolved problem. It is extremely challenging because there are three distinct contributors to the domain gap between forensic sketches and photos: The well-studied sketch-photo modality gap, and the less studied gaps due to (i) the forgetting process of the eye-witness and (ii) their inability to elucidate their memory. In this paper, we address the memory problem head on by introducing a database of 400 forensic sketches created at different time-delays. Based on this database we build a model to reverse the forgetting process. Surprisingly, we show that it is possible to systematically "un-forget" facial details. Moreover, it is possible to apply this model to dramatically improve forensic sketch recognition in practice: we achieve the state of the art results when matching 195 benchmark forensic sketches against corresponding photos and a 10,030 mugshot database. Shuxin Ouyang, Timothy M. Hospedales, Yi-Zhe Song, Xueming Li 0002 |
CVPR | 4 |
| 2016 | A survey on heterogeneous face recognition: Sketch, infra-red, 3D and low-resolution
Shuxin Ouyang, Timothy M. Hospedales, Yi-Zhe Song, Xueming Li 0002, Chen Change Loy, Xiaogang Wang 0001 |
Image Vis. Comput. | 4 |
| 2016 | Structure-Constrained Low-Rank and Partial Sparse Representation with Sample Selection for image classification
Yang Liu 0132, Xueming Li 0002, Haixu Liu |
Pattern Recognit. | 2 |
| 2014 | Cross-Modal Face Matching: Beyond Viewed Sketches
Shuxin Ouyang, Timothy M. Hospedales, Yi-Zhe Song, Xueming Li 0002 |
ACCV (2) | 4 |
| 2014 | Structure-constrained low-rank and partial sparse representation for image classificationabstractIn this paper, a novel Structure-Constrained Low-Rank and Partial Sparse Representation algorithm for image classification is proposed. First, a Structure-Constrained Low-Rank dictionary learning algorithm is proposed, which imposes both structure and low-rank restriction on the coefficient matrix. Second, under the assumption that the representation of test sample is sparse and correlated with the learned representation of training samples, we concatenate training samples and test samples to form a data matrix and find a low-rank and sparse representation of the data matrix over learned dictionary by low-rank matrix recovery technique. Experimental results demonstrate the effectiveness of the proposed algorithm. Yang Liu 0132, Haixu Liu, Xueming Li 0002 |
ICIP | 4 |
| 2014 | Hyperspectral image classification using sparse representation-based classifierabstractBased on the idea of SRC(Sparse Representation based Classification), a novel approach HSIC-SRC is proposed in this paper. Unlike most existing algorithms for HSIC(HSI Classification) via sparse representation, our main contributions lie in two aspects, 1) Considering the performance of SRC depending on the quality of dictionary, we employ LC-KSVD(Label Consistent KSVD) algorithm which joints the regularization of discriminative sparse-code error and the regularization of classification error into the sparsity model, to learn a more `compact' and `discriminative' dictionary for sparse representation, hence we can achieve better accuracy on classification; 2) Considering the neighboring pixels usually composed by similar materials, their spectral characteristics are highly correlated. Therefore, we employ the contextual information(or spatial information) into the sparse model, to obtain more optimized sparse representation of pixels. The proposed HSIC-SRC is applied to the well known hyperspectral image `University of Pavia' for classification, and experimental results show that it outperforms some other state-of-the-art classification algorithms, such as SVM, SP, OMP and SPG-l1, with only one simple linear classifier. Yufang Tang, Xueming Li 0002, Yang Liu 0132, Jizhe Wang, Shuchang Liu 0005 |
IGARSS | 2 |
| 2014 | Sparse dimensionality reduction based on compressed sensingabstractIn this paper, we propose a novel approach SDR-CS (Sparse Dimensionality Reduction based on CS) based on compressed sensing to reduce dimensionality. With certain constraint of objective function, our semi-supervised learning method utilizes instance to construct the optimally sparse dictionary in the training dataset, employs K-SVD and OMP algorithms to improve the convergence rate of learning, and then reduces the dimensionality of sparse representation of original data by Gaussian random matrix as measurement matrix, to achieve the purpose of dimensionality reduction. Experimental results demonstrate that our overcomplete sparse dictionary can enhance the major underlying structure characteristics of sparse representation, which are mapped into the regions with continuous dimensionality, not the same dimensionality, and improve the discrimination among data which belong to different classes. Only with the constraint of l2-norm, the proposed SDR-CS method has better performance of dimensionality reduction in the MNIST dataset, and it is superior to other existing methods with constraints of l2/l1-norm, achieving the classification error rate of 0.03. Yufang Tang, Xueming Li 0002, Yang Liu 0132, Jizhe Wang |
WCNC | 2 |
| 2013 | A novel method for stereo matching using Gabor Feature Image and Confidence MaskabstractIn this paper, we present a novel local-based algorithm for stereo matching using Gabor-Feature-Image and Confidence-Mask. Various local-based schemes have been proposed in recent years, most of them mainly use color difference as evaluation criterion when constructing the initial cost volume, however, color channel is highly sensitive to noise, illumination changes, etc. Therefore, we develop a new cost function based on Gabor-Feature-Image for obtaining a more accurate matching cost volume. Furthermore, in order to eliminate the matching ambiguities brought by the winnertakes-all method, an effective disparity refinement strategy using Confidence-Mask is implemented to select and refine the less reliable pixels. The proposed algorithm ranks 23th out of over 150 (global-based and local-based) methods on Middlebury data sets, both quantitative and qualitative evaluation show that it is comparable to state-of-the-art local-based stereo matching algorithms. Haixu Liu, Yang Liu 0132, Shuxin Ouyang, Xueming Li 0002 |
VCIP | 5 |