VLDB 2026 Research / reviewers in the wild / expert
Qiang Li 0024
dblp:72/872-24
· DBLP profile ↗
23ranked-venue papers
3as first author
10since 2021 · last 2026
0000-0003-0958-9926ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 16 · 2 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 3 first-author · 4 since 2021Databases, data management, data science and information retrieval · 5 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
12 papers |
Representation and self-supervised learning · 20% Graph learning · 18% Generative modeling · 11% | |
| Computer graphics and multimedia
2 papers |
Visual content generation and editing · 60% Rendering · 40% | |
| Databases, data mining, and information retrieval
1 paper |
Data integration and cleaning · 100% |
Topics — the 30 heaviest of 39, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Representation and self-supervised learning
contrastive learning |
1.2 | 2 | 2023 | LiftedCL: Lifting Contrastive Learning for Human-Centric Perception · ICLR 2023 Exploring Set Similarity for Dense Self-supervised Representation Learning · CVPR 2022 |
Machine learning › Generative modeling
generative adversarial network |
1.2 | 2 | 2023 | FEditNet: Few-Shot Editing of Latent Semantics in GAN Spaces · AAAI 2023 BlendGAN: Implicitly GAN Blending for Arbitrary Stylized Face Generation · NeurIPS 2021 |
Visual content generation and editing
3d content creation |
1.0 | 1 | 2026 | Sketch2Avatar: Geometry-Guided 3D Full-Body Human Generation in 360° From Hand-Drawn Sketches · IEEE Trans. Vis. Comput. Graph. 2026 |
Rendering
neural rendering |
1.0 | 1 | 2026 | Sketch2Avatar: Geometry-Guided 3D Full-Body Human Generation in 360° From Hand-Drawn Sketches · IEEE Trans. Vis. Comput. Graph. 2026 |
Machine learning › Graph learning
network embedding |
0.7 | 2 | 2019 | Adversarial Training Methods for Network Embedding · WWW 2019 Adversarial Network Embedding · AAAI 2018 |
Machine learning › Representation and self-supervised learning › latent space
latent space manipulation |
0.7 | 1 | 2023 | FEditNet: Few-Shot Editing of Latent Semantics in GAN Spaces · AAAI 2023 |
Computer vision › Vision and language
cross-modal alignment |
0.6 | 1 | 2022 | CRIS: CLIP-Driven Referring Image Segmentation · CVPR 2022 |
Computer vision › Segmentation and scene understanding
referring image segmentation |
0.6 | 1 | 2022 | CRIS: CLIP-Driven Referring Image Segmentation · CVPR 2022 |
Machine learning › Learning paradigms
multi-label classification |
0.5 | 2 | 2016 | Correlated Logistic Model With Elastic Net Regularization for Multilabel Image Classification · IEEE Trans. Image Process. 2016 Conditional Graphical Lasso for Multi-label Image Classification · CVPR 2016 |
Computer vision › 3D vision › pose estimation
3d hand pose estimation |
0.4 | 1 | 2019 | End-to-End Hand Mesh Recovery From a Monocular RGB Image · ICCV 2019 |
Computer vision › 3D vision
3d human reconstruction |
0.4 | 1 | 2019 | End-to-End Hand Mesh Recovery From a Monocular RGB Image · ICCV 2019 |
Machine learning › Trustworthy machine learning › robustness
adversarial robustness |
0.4 | 1 | 2019 | Adversarial Training Methods for Network Embedding · WWW 2019 |
Machine learning › Trustworthy machine learning › robustness › adversarial robustness
adversarial training |
0.4 | 1 | 2019 | Adversarial Training Methods for Network Embedding · WWW 2019 |
Computer vision › 3D vision › 3d human reconstruction
hand mesh reconstruction |
0.4 | 1 | 2019 | End-to-End Hand Mesh Recovery From a Monocular RGB Image · ICCV 2019 |
Computer vision › Face, body and person analysis › human pose estimation › articulated pose estimation
hand pose estimation |
0.4 | 1 | 2019 | End-to-End Hand Mesh Recovery From a Monocular RGB Image · ICCV 2019 |
Machine learning › Graph learning › network embedding
adversarial network embedding |
0.3 | 1 | 2018 | Adversarial Network Embedding · AAAI 2018 |
Machine learning › Optimization for machine learning
combinatorial optimization |
0.3 | 1 | 2018 | Large-Scale Order Dispatch in On-Demand Ride-Hailing Platforms: A Learning and Planning Approach · KDD 2018 |
Machine learning › Graph learning
graph representation learning |
0.3 | 1 | 2018 | Adversarial Network Embedding · AAAI 2018 |
Knowledge, reasoning and agents › Planning, search and constraint satisfaction › planning
large-scale planning |
0.3 | 1 | 2018 | Large-Scale Order Dispatch in On-Demand Ride-Hailing Platforms: A Learning and Planning Approach · KDD 2018 |
Knowledge, reasoning and agents › Planning, search and constraint satisfaction › planning
order dispatching |
0.3 | 1 | 2018 | Large-Scale Order Dispatch in On-Demand Ride-Hailing Platforms: A Learning and Planning Approach · KDD 2018 |
Machine learning › Graph learning › graph representation learning › structural embedding
structure-preserving embedding |
0.3 | 1 | 2018 | Adversarial Network Embedding · AAAI 2018 |
Machine learning › Graph learning
graph clustering |
0.3 | 1 | 2017 | Improving Stochastic Block Models by Incorporating Power-Law Degree Characteristic · IJCAI 2017 |
Machine learning › Graph learning
stochastic block model |
0.3 | 1 | 2017 | Improving Stochastic Block Models by Incorporating Power-Law Degree Characteristic · IJCAI 2017 |
Computer vision › Image recognition and object detection
image classification |
0.2 | 1 | 2016 | Correlated Logistic Model With Elastic Net Regularization for Multilabel Image Classification · IEEE Trans. Image Process. 2016 |
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › regression › generalized linear model
logistic regression |
0.2 | 1 | 2016 | Correlated Logistic Model With Elastic Net Regularization for Multilabel Image Classification · IEEE Trans. Image Process. 2016 |
Data integration and cleaning › data preprocessing
data cleaning |
0.2 | 1 | 2016 | Random Mixed Field Model for Mixed-Attribute Data Restoration · AAAI 2016 |
Data integration and cleaning
data reconstruction |
0.2 | 1 | 2016 | Random Mixed Field Model for Mixed-Attribute Data Restoration · AAAI 2016 |
Data integration and cleaning › missing data
missing value imputation |
0.2 | 1 | 2016 | Random Mixed Field Model for Mixed-Attribute Data Restoration · AAAI 2016 |
Machine learning › Generative modeling
latent space interpretation |
0.2 | 1 | 2023 | FEditNet: Few-Shot Editing of Latent Semantics in GAN Spaces · AAAI 2023 |
Computer vision › Segmentation and scene understanding
instance segmentation |
0.2 | 1 | 2022 | Exploring Set Similarity for Dense Self-supervised Representation Learning · CVPR 2022 |
Methods — techniques the papers use, named apart from their topics
contrastive learning · 1.2weighted blending · 1.0transformer-based generator · 1.0self-supervised learning · 1.0neural rendering · 1.0geometry-guided generation · 1.0few-shot learning · 0.7vision-language decoder · 0.6nearest neighbor search · 0.6attention · 0.6CLIP · 0.6differentiable re-projection loss · 0.4deep learning · 0.4structured mean-field variational inference · 0.2random mixed field prior · 0.2probabilistic generative model · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Sketch2Avatar: Geometry-Guided 3D Full-Body Human Generation in 360° From Hand-Drawn SketchesabstractGenerating full-body humans in 360$^{\circ }$∘ has broad applications in digital entertainment, online education and art design. Existing works primarily rely on coarse conditions such as body pose to guide the generation, lacking detailed control over the synthesized results. Regarding this limitation, sketches offer a promising alternative as an expressive condition that enables more explicit and precise control. However, current sketch-based generation methods focus on faces or common objects, how to transfer sketches into 360 $^{\circ }$∘ full-body humans remains unexplored. To bridge this gap, we propose Sketch2Avatar, the first generative model to achieve 3D full-body human generation from hand-drawn sketches. Our model is capable of synthesizing sketch-aligned and 360$^{\circ }$∘-consistent full-body human images by leveraging the geometry information extracted from sketches to guide the 3D representation generation and neural rendering. Specifically, we propose sketchguided 3D representation generation to model the 3D human and maintain the alignment between input sketches and generated humans. Our transformer-based generator incorporates spatial feature guidance and latent modulation derived from sketches to produce high-quality 3D representations. Additionally, our designed bodyaware neural rendering utilizes 3D human body priors from sketches, simplifying the learning of articulated body poses and complex body shapes. To train and evaluate our model, we construct a large-scale dataset comprising approximately 19 K 2D full-body human images and their corresponding sketches in a hand-drawn style. Experimental results demonstrate that our Sketch2Avatar can transfer hand-drawn sketches into photo-realistic 360$^{\circ }$∘ full-body human images with precise sketch-human alignment. Ablation studies further validate the effectiveness of our design choices. Qiang Li 0024, Jie Zhang 0090, Anthony Kong, Ping Li 0016 |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2025 | CFinBench: A Comprehensive Chinese Financial Benchmark for Large Language ModelsabstractYing Nie, Binwei Yan, Tianyu Guo, Hao Liu, Haoyu Wang, Wei He, Binfan Zheng, Weihao Wang, Qiang Li, Weijian Sun, Yunhe Wang, Dacheng Tao. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Binwei Yan, Tianyu Guo 0001, Wei He 0001, Binfan Zheng, Qiang Li 0024, Weijian Sun, Yunhe Wang 0001, Dacheng Tao |
NAACL (Long Papers) | 9 |
| 2024 | SED: Searching Enhanced Decoder with switchable skip connection for semantic segmentation
Zhibin Quan, Qiang Li 0024, Dejun Zhu, Wankou Yang |
Pattern Recognit. | 3 |
| 2023 | FEditNet: Few-Shot Editing of Latent Semantics in GAN SpacesabstractGenerative Adversarial networks (GANs) have demonstrated their powerful capability of synthesizing high-resolution images, and great efforts have been made to interpret the semantics in the latent spaces of GANs. However, existing works still have the following limitations: (1) the majority of works rely on either pretrained attribute predictors or large-scale labeled datasets, which are difficult to collect in most cases, and (2) some other methods are only suitable for restricted cases, such as focusing on interpretation of human facial images using prior facial semantics. In this paper, we propose a GAN-based method called FEditNet, aiming to discover latent semantics using very few labeled data without any pretrained predictors or prior knowledge. Specifically, we reuse the knowledge from the pretrained GANs, and by doing so, avoid overfitting during the few-shot training of FEditNet. Moreover, our layer-wise objectives which take content consistency into account also ensure the disentanglement between attributes. Qualitative and quantitative results demonstrate that our method outperforms the state-of-the-art methods on various datasets. The code is available at https://github.com/THU-LYJ-Lab/FEditNet. Mengfei Xia, Yezhi Shu, Yuji Wang, Yukun Lai, Qiang Li 0024, Pengfei Wan 0001, Zhongyuan Wang 0006, Yong-Jin Liu 0001 |
AAAI | 5 |
| 2023 | LiftedCL: Lifting Contrastive Learning for Human-Centric Perception
Qiang Li 0024, Wankou Yang |
ICLR | 2 |
| 2023 | Pyramid Geometric Consistency Learning For Semantic Segmentation
Qiang Li 0024, Zhibin Quan, Wankou Yang |
Pattern Recognit. | 2 |
| 2022 | CRIS: CLIP-Driven Referring Image SegmentationabstractReferring image segmentation aims to segment a referent via a natural linguistic expression. Due to the distinct data properties between text and image, it is challenging for a network to well align text and pixel-level features. Existing approaches use pretrained models to facilitate learning, yet separately transfer the language/vision knowledge from pretrained models, ignoring the multi-modal corresponding information. Inspired by the recent advance in Contrastive Language-Image Pretraining (CLIP), in this paper, we propose an end-to-end CLIP-Driven Referring Image Segmen-tation framework (CRIS). To transfer the multi-modal knowledge effectively, CRIS resorts to vision-language decoding and contrastive learning for achieving the text-to-pixel alignment. More specifically, we design a vision-language decoder to propagate fine-grained semantic information from textual representations to each pixel-level activation, which promotes consistency between the two modalities. In addition, we present text-to-pixel contrastive learning to explicitly enforce the text feature similar to the related pixel-level features and dissimilar to the irrelevances. The experimental results on three benchmark datasets demonstrate that our proposed framework significantly outperforms the state-of-the-art performance without any post-processing. Zhaoqing Wang, Qiang Li 0024, Xunqiang Tao, Yandong Guo, Mingming Gong, Tongliang Liu |
CVPR | 3 |
| 2022 | Exploring Set Similarity for Dense Self-supervised Representation LearningabstractBy considering the spatial correspondence, dense self-supervised representation learning has achieved superior performance on various dense prediction tasks. However, the pixel-level correspondence tends to be noisy because of many similar misleading pixels, e.g., backgrounds. To address this issue, in this paper, we propose to explore set similarity (SetSim) for dense self-supervised representation learning. We generalize pixel-wise similarity learning to set-wise one to improve the robustness because sets contain more semantic and structure information. Specifically, by resorting to attentional features of views, we establish the corresponding set, thus filtering out noisy backgrounds that may cause incorrect correspondences. Meanwhile, these at-tentional features can keep the coherence of the same image across different views to alleviate semantic inconsistency. We further search the cross-view nearest neighbours of sets and employ the structured neighbourhood information to enhance the robustness. Empirical evaluations demonstrate that SetSim surpasses or is on par with state-of-the-art meth-ods on object detection, keypoint detection, instance segmen-tation, and semantic segmentation. Zhaoqing Wang, Qiang Li 0024, Pengfei Wan 0001, Nannan Wang 0001, Mingming Gong, Tongliang Liu |
CVPR | 2 |
| 2021 | BlendGAN: Implicitly GAN Blending for Arbitrary Stylized Face GenerationabstractGenerative Adversarial Networks (GANs) have made a dramatic leap in high-fidelity image synthesis and stylized face generation. Recently, a layer-swapping mechanism has been developed to improve the stylization performance. However, this method is incapable of fitting arbitrary styles in a single model and requires hundreds of style-consistent training images for each style. To address the above issues, we propose BlendGAN for arbitrary stylized face generation by leveraging a flexible blending strategy and a generic artistic dataset. Specifically, we first train a self-supervised style encoder on the generic artistic dataset to extract the representations of arbitrary styles. In addition, a weighted blending module (WBM) is proposed to blend face and style representations implicitly and control the arbitrary stylization effect. By doing so, BlendGAN can gracefully fit arbitrary styles in a unified model while avoiding case-by-case preparation of style-consistent training images. To this end, we also present a novel large-scale artistic face dataset AAHQ. Extensive experiments demonstrate that BlendGAN outperforms state-of-the-art methods in terms of visual quality and style diversity for both latent-guided and reference-guided stylized face synthesis. Mingcong Liu, Qiang Li 0024, Zekui Qin, Pengfei Wan 0001 |
NeurIPS | 2 |
| 2021 | Adversarial training regularization for negative sampling based network embedding
Quanyu Dai, Xiao Shen 0001, Zimu Zheng, Liang Zhang 0042, Qiang Li 0024, Dan Wang 0002 |
Inf. Sci. | 5 |
| 2019 | End-to-End Hand Mesh Recovery From a Monocular RGB ImageabstractIn this paper, we present a HAnd Mesh Recovery (HAMR) framework to tackle the problem of reconstructing the full 3D mesh of a human hand from a single RGB image. In contrast to existing research on 2D or 3D hand pose estimation from RGB or/and depth image data, HAMR can provide a more expressive and useful mesh representation for monocular hand image understanding. In particular, the mesh representation is achieved by parameterizing a generic 3D hand model with shape and relative 3D joint angles. By utilizing this mesh representation, we can easily compute the 3D joint locations via linear interpolations between the vertexes of the mesh, while obtain the 2D joint locations with a projection of the 3D joints. To this end, a differentiable re-projection loss can be defined in terms of the derived representations and the ground-truth labels, thus making our framework end-to-end trainable. Qualitative experiments show that our framework is capable of recovering appealing 3D hand mesh even in the presence of severe occlusions. Quantitatively, our approach also outperforms the state-of-the-art methods for both 2D and 3D hand pose estimation from a monocular RGB image on several benchmark datasets. Qiang Li 0024, Hong Mo |
ICCV | 2 |
| 2019 | Ranking Network Embedding via Adversarial Learning
Quanyu Dai, Qiang Li 0024, Liang Zhang 0042, Dan Wang 0002 |
PAKDD (3) | 2 |
| 2019 | Adversarial Training Methods for Network EmbeddingabstractNetwork Embedding is the task of learning continuous node representations for networks, which has been shown effective in a variety of tasks such as link prediction and node classification. Most of existing works aim to preserve different network structures and properties in low-dimensional embedding vectors, while neglecting the existence of noisy information in many real-world networks and the overfitting issue in the embedding learning process. Most recently, generative adversarial networks (GANs) based regularization methods are exploited to regularize embedding learning process, which can encourage a global smoothness of embedding vectors. These methods have very complicated architecture and suffer from the well-recognized non-convergence problem of GANs. In this paper, we aim to introduce a more succinct and effective local regularization method, namely adversarial training, to network embedding so as to achieve model robustness and better generalization performance. Firstly, the adversarial training method is applied by defining adversarial perturbations in the embedding space with an adaptive L2 norm constraint that depends on the connectivity pattern of node pairs. Though effective as a regularizer, it suffers from the interpretability issue which may hinder its application in certain real-world scenarios. To improve this strategy, we further propose an interpretable adversarial training method by enforcing the reconstruction of the adversarial examples in the discrete graph domain. These two regularization methods can be applied to many existing embedding models, and we take DeepWalk as the base model for illustration in the paper. Empirical evaluations in both link prediction and node classification demonstrate the effectiveness of the proposed methods. Quanyu Dai, Xiao Shen 0001, Liang Zhang 0042, Qiang Li 0024, Dan Wang 0002 |
WWW | 4 |
| 2019 | Adapting Stochastic Block Models to Power-Law Degree DistributionsabstractStochastic block models (SBMs) have been playing an important role in modeling clusters or community structures of network data. But, it is incapable of handling several complex features ubiquitously exhibited in real-world networks, one of which is the power-law degree characteristic. To this end, we propose a new variant of SBM, termed power-law degree SBM (PLD-SBM), by introducing degree decay variables to explicitly encode the varying degree distribution over all nodes. With an exponential prior, it is proved that PLD-SBM approximately preserves the scale-free feature in real networks. In addition, from the inference of variational E-Step, PLD-SBM is indeed to correct the bias inherited in SBM with the introduced degree decay factors. Furthermore, experiments conducted on both synthetic networks and two real-world datasets including Adolescent Health Data and the political blogs network verify the effectiveness of the proposed model in terms of cluster prediction accuracies. Maoying Qiao, Jun Yu 0002, Wei Bian 0003, Qiang Li 0024, Dacheng Tao |
IEEE Trans. Cybern. | 4 |
| 2018 | Adversarial Network EmbeddingabstractLearning low-dimensional representations of networks has proved effective in a variety of tasks such as node classification, link prediction and network visualization. Existing methods can effectively encode different structural properties into the representations, such as neighborhood connectivity patterns, global structural role similarities and other high-order proximities. However, except for objectives to capture network structural properties, most of them suffer from lack of additional constraints for enhancing the robustness of representations. In this paper, we aim to exploit the strengths of generative adversarial networks in capturing latent features, and investigate its contribution in learning stable and robust graph representations. Specifically, we propose an Adversarial Network Embedding (ANE) framework, which leverages the adversarial learning principle to regularize the representation learning. It consists of two components, i.e., a structure preserving component and an adversarial learning component. The former component aims to capture network structural properties, while the latter contributes to learning robust representations by matching the posterior distribution of the latent representations to given priors. As shown by the empirical results, our method is competitive with or superior to state-of-the-art approaches on benchmark network embedding tasks. Quanyu Dai, Qiang Li 0024, Dan Wang 0002 |
AAAI | 2 |
| 2018 | Large-Scale Order Dispatch in On-Demand Ride-Hailing Platforms: A Learning and Planning ApproachabstractWe present a novel order dispatch algorithm in large-scale on-demand ride-hailing platforms. While traditional order dispatch approaches usually focus on immediate customer satisfaction, the proposed algorithm is designed to provide a more efficient way to optimize resource utilization and user experience in a global and more farsighted view. In particular, we model order dispatch as a large-scale sequential decision-making problem, where the decision of assigning an order to a driver is determined by a centralized algorithm in a coordinated way. The problem is solved in a learning and planning manner: 1) based on historical data, we first summarize demand and supply patterns into a spatiotemporal quantization, each of which indicates the expected value of a driver being in a particular state; 2) a planning step is conducted in real-time, where each driver-order-pair is valued in consideration of both immediate rewards and future gains, and then dispatch is solved using a combinatorial optimizing algorithm. Through extensive offline experiments and online AB tests, the proposed approach delivers remarkable improvement on the platform's efficiency and has been successfully deployed in the production system of Didi Chuxing. Zhixin Li 0006, Qingwen Guan, Dingshui Zhang, Qiang Li 0024, Junxiao Nan, Wei Bian 0003, Jieping Ye |
KDD | 5 |
| 2017 | Improving Stochastic Block Models by Incorporating Power-Law Degree CharacteristicabstractStochastic block models (SBMs) provide a statistical way modeling network data, especially in representing clusters or community structures. However, most block models do not consider complex characteristics of networks such as scale-free feature, making them incapable of handling degree variation of vertices, which is ubiquitous in real networks. To address this issue, we introduce degree decay variables into SBM, termed power-law degree SBM (PLD-SBM), to model the varying probability of connections between node pairs. The scale-free feature is approximated by a power-law degree characteristic. Such a property allows PLD-SBM to correct the distortion of degree distribution in SBM, and thus improves the performance of cluster prediction. Experiments on both simulated networks and two real-world networks including the Adolescent Health Data and the political blogs network demonstrate the validity of the motivation of PLD-SBM, and its practical superiority. Maoying Qiao, Jun Yu 0002, Wei Bian 0003, Qiang Li 0024, Dacheng Tao |
IJCAI | 4 |
| 2016 | Random Mixed Field Model for Mixed-Attribute Data RestorationabstractNoisy and incomplete data restoration is a critical preprocessing step in developing effective learning algorithms, which targets to reduce the effect of noise and missing values in data. By utilizing attribute correlations and/or instance similarities, various techniques have been developed for data denoising and imputation tasks. However, current existing data restoration methods are either specifically designed for a particular task, or incapable of dealing with mixed-attribute data. In this paper, we develop a new probabilistic model to provide a general and principled method for restoring mixed-attribute data. The main contributions of this study are twofold: a) a unified generative model, utilizing a generic random mixed field (RMF) prior, is designed to exploit mixed-attribute correlations; and b) a structured mean-field variational approach is proposed to solve the challenging inference problem of simultaneous denoising and imputation. We evaluate our method by classification experiments on both synthetic data and real benchmark datasets. Experiments demonstrate, our approach can effectively improve the classification accuracy of noisy and incomplete data by comparing with other data restoration methods. Qiang Li 0024, Wei Bian 0003, Jane You, Dacheng Tao |
AAAI | 1 |
| 2016 | Conditional Graphical Lasso for Multi-label Image ClassificationabstractMulti-label image classification aims to predict multiple labels for a single image which contains diverse content. By utilizing label correlations, various techniques have been developed to improve classification performance. However, current existing methods either neglect image features when exploiting label correlations or lack the ability to learn image-dependent conditional label structures. In this paper, we develop conditional graphical Lasso (CGL) to handle these challenges. CGL provides a unified Bayesian framework for structure and parameter learning conditioned on image features. We formulate the multi-label prediction as CGL inference problem, which is solved by a mean field variational approach. Meanwhile, CGL learning is efficient due to a tailored proximal gradient procedure by applying the maximum a posterior (MAP) methodology. CGL performs competitively for multi-label image classification on benchmark datasets MULAN scene, PASCAL VOC 2007 and PASCAL VOC 2012, compared with the state-of-the-art multi-label classification algorithms. Qiang Li 0024, Maoying Qiao, Wei Bian 0003, Dacheng Tao |
CVPR | 1 |
| 2016 | Correlated Logistic Model With Elastic Net Regularization for Multilabel Image ClassificationabstractIn this paper, we present correlated logistic (CorrLog) model for multilabel image classification. CorrLog extends conventional logistic regression model into multilabel cases, via explicitly modeling the pairwise correlation between labels. In addition, we propose to learn the model parameters of CorrLog with elastic net regularization, which helps exploit the sparsity in feature selection and label correlations and thus further boost the performance of multilabel classification. CorrLog can be efficiently learned, though approximately, by regularized maximum pseudo likelihood estimation, and it enjoys a satisfying generalization bound that is independent of the number of labels. CorrLog performs competitively for multilabel image classification on benchmark data sets MULAN scene, MIT outdoor scene, PASCAL VOC 2007, and PASCAL VOC 2012, compared with the state-of-the-art multilabel classification algorithms. Qiang Li 0024, Bo Xie 0002, Jane You, Wei Bian 0003, Dacheng Tao |
IEEE Trans. Image Process. | 1 |
| 2016 | Mutual Component Analysis for Heterogeneous Face RecognitionabstractHeterogeneous face recognition, also known as cross-modality face recognition or intermodality face recognition, refers to matching two face images from alternative image modalities. Since face images from different image modalities of the same person are associated with the same face object, there should be mutual components that reflect those intrinsic face characteristics that are invariant to the image modalities. Motivated by this rationality, we propose a novel approach called Mutual Component Analysis (MCA) to infer the mutual components for robust heterogeneous face recognition. In the MCA approach, a generative model is first proposed to model the process of generating face images in different modalities, and then an Expectation Maximization (EM) algorithm is designed to iteratively learn the model parameters. The learned generative model is able to infer the mutual components (which we call the hidden factor , where hidden means the factor is unreachable and invisible, and can only be inferred from observations) that are associated with the person’s identity, thus enabling fast and effective matching for cross-modality face recognition. To enhance recognition performance, we propose an MCA-based multiclassifier framework using multiple local features. Experimental results show that our new approach significantly outperforms the state-of-the-art results on two typical application scenarios: sketch-to-photo and infrared-to-visible face recognition. Zhifeng Li 0001, Dihong Gong, Qiang Li 0024, Dacheng Tao, Xuelong Li 0001 |
ACM Trans. Intell. Syst. Technol. | 3 |
| 2015 | Robust Nonnegative Patch Alignment for Dimensionality ReductionabstractDimensionality reduction is an important method to analyze high-dimensional data and has many applications in pattern recognition and computer vision. In this paper, we propose a robust nonnegative patch alignment for dimensionality reduction, which includes a reconstruction error term and a whole alignment term. We use correntropy-induced metric to measure the reconstruction error, in which the weight is learned adaptively for each entry. For the whole alignment, we propose locality-preserving robust nonnegative patch alignment (LP-RNA) and sparsity-preserviing robust nonnegative patch alignment (SP-RNA), which are unsupervised and supervised, respectively. In the LP-RNA, we propose a locally sparse graph to encode the local geometric structure of the manifold embedded in high-dimensional space. In particular, we select large p -nearest neighbors for each sample, then obtain the sparse representation with respect to these neighbors. The sparse representation is used to build a graph, which simultaneously enjoys locality, sparseness, and robustness. In the SP-RNA, we simultaneously use local geometric structure and discriminative information, in which the sparse reconstruction coefficient is used to characterize the local geometric structure and weighted distance is used to measure the separability of different classes. For the induced nonconvex objective function, we formulate it into a weighted nonnegative matrix factorization based on half-quadratic optimization. We propose a multiplicative update rule to solve this function and show that the objective function converges to a local optimum. Several experimental results on synthetic and real data sets demonstrate that the learned representation is more discriminative and robust than most existing dimensionality reduction methods. Xinge You, Weihua Ou, C. L. Philip Chen, Qiang Li 0024, Yuan Yan Tang |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2014 | Local Metric Learning for Exemplar-Based Object DetectionabstractObject detection has been widely studied in the computer vision community and it has many real applications, despite its variations, such as scale, pose, lighting, and background. Most classical object detection methods heavily rely on category-based training to handle intra-class variations. In contrast to classical methods that use a rigid category-based representation, exemplar-based methods try to model variations among positives by learning from specific positive samples. However, current existing exemplar-based methods either fail to use any training information or suffer from a significant performance drop when few exemplars are available. In this paper, we design a novel local metric learning approach to well handle exemplar-based object detection task. The main works are two-fold: 1) a novel local metric learning algorithm called exemplar metric learning (EML) is designed and 2) an exemplar-based object detection algorithm based on EML is implemented. We evaluate our method on two generic object detection data sets: UIUC-Car and UMass FDDB. Experiments show that compared with other exemplar-based methods, our approach can effectively enhance object detection performance when few exemplars are available. Xinge You, Qiang Li 0024, Dacheng Tao, Weihua Ou, Mingming Gong |
IEEE Trans. Circuits Syst. Video Technol. | 2 |