EDBT 2026 Demo / reviewers in the wild / expert
Wen Wang 0019
dblp:29/4680-19
· DBLP profile ↗
31ranked-venue papers
5as first author
26since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 15 · 3 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 15 · 5 first-author · 10 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 6 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Towards plastic and stable incremental learning: A dual-learner framework with cumulative parameter averaging
Wenju Sun, Qingyong Li, Wen Wang 0019 |
Pattern Recognit. | 4 |
| 2025 | Speech Token Prediction via Compressed-to-fine Language Modeling for Speech GenerationabstractNeural audio codecs, used as speech tokenizers, have demonstrated remarkable potential in the field of speech generation. However, to ensure high-fidelity audio reconstruction, neural audio codecs typically encode audio into long sequences of speech tokens, posing a significant challenge for downstream language models in long-context modeling. We observe that speech token sequences exhibit short-range dependency: due to the monotonic alignment between text and speech in text-to-speech (TTS) tasks, the prediction of the current token primarily relies on its local context, while long-range tokens contribute less to the current token prediction and often contain redundant information. Inspired by this observation, we propose a compressed-to-fine language modeling approach to address the challenge of long sequence speech tokens within neural codec language models: (1) Fine-grained Initial and Short-range Information: Our approach retains the prompt and local tokens during prediction to ensure text alignment and the integrity of paralinguistic information; (2) Compressed Long-range Context: Our approach compresses long-range token spans into compact representations to reduce redundant information while preserving essential semantics. Extensive experiments on various neural audio codecs and downstream language models validate the effectiveness and generalizability of the proposed approach, highlighting the importance of token compression in improving speech generation within neural codec language models. The demo of audio samples will be available at https://anonymous.4open.science/r/SpeechTokenPredictionViaCompressedToFinedLM. Wenrui Liu 0003, Qian Chen 0003, Wen Wang 0019, Guanrou Yang, Minghui Fang 0002, Jialong Zuo, Xiaoda Yang, Tao Jin 0004, Jin Xu 0010, Yafeng Chen, Jionghao Bai, Zhifang Guo |
ACM Multimedia | 3 |
| 2025 | MelodyEdit: Zero-shot Music Editing with Disentangled Inversion ControlabstractText-guided diffusion models revolutionize audio generation by adapting source audio to specific text prompts. However, existing zero-shot audio editing methods such as DDIM inversion accumulate errors across diffusion steps, reducing the effectiveness. Moreover, existing editing methods struggle with conducting complex non-rigid music edits while maintaining content integrity and high fidelity. To address these challenges, we propose MelodyEdit, a novel zero-shot music editing system based on innovative Disentangled Inversion Control (DIC) technique, which comprises Harmonized Attention Control and Disentangled Inversion. Disentangled Inversion disentangles the diffusion process into triple branches to rectify the deviated path of the source branch caused by DDIM inversion. Harmonized Attention Control unifies the mutual self-attention control and the cross-attention control with an intermediate Harmonic Branch to progressively generate the desired harmonic and melodic information in the target music. We also introduce ZoME-Bench, a comprehensive music editing benchmark with 1,100 samples covering ten distinct editing categories. ZoME-Bench facilitates both zero-shot and instruction-based music editing tasks. Our method outperforms state-of-the-art inversion techniques in editing fidelity and content preservation. Huadai Liu, Xiangtai Li, Wen Wang 0019, Qian Chen 0003, Rongjie Huang 0001, Zhou Zhao 0001, Wei Xue 0002 |
ACM Multimedia | 4 |
| 2025 | Task Arithmetic in Trust Region: A Training-Free Model Merging Approach to Navigate Knowledge ConflictsabstractMulti-task model merging offers an efficient solution for integrating knowledge from multiple fine-tuned models, mitigating the significant computational and storage demands associated with multi-task training. As a key technique in this field, Task Arithmetic (TA) defines task vectors by subtracting the pre-trained model (0 pre) from the fine-tuned task models in parameter space, then adjusting the weight between these task vectors and 0 pre to balance task-generalized and task-specific knowledge. Despite the promising performance of TA, conflicts can arise among the task vectors, particularly when different tasks require distinct model adaptations. In this paper, we formally define this issue as knowledge conflicts, characterized by the performance degradation of one task after merging with a model fine-tuned for another task. Through in-depth analysis, we show that these conflicts stem primarily from the components of task vectors that align with the gradient of task-specific losses at 0 pre. To address this, we propose Task Arithmetic in Trust Region (TATR), which defines the trust region as dimensions in the model parameter space that cause only small changes (corresponding to the task vector components with gradient orthogonal direction) in the task-specific losses. Restricting parameter merging within this trust region, TATR can effectively alleviate knowledge conflicts. Moreover, TATR serves as a plug-and-play module compatible with a wide range of TA-based methods. Extensive empirical evaluations on visual and visual-language tasks robustly demonstrate that TATR improves the multi-task performance of several TA-based model merging methods. Wenju Sun, Qingyong Li, Wen Wang 0019, Boyang Li 0001 |
ACM Multimedia | 3 |
| 2025 | EmoVoice: LLM-based Emotional Text-To-Speech Model with Freestyle Text PromptingabstractHuman speech goes beyond the mere transfer of information; it is a profound exchange of emotions and a connection between individuals. While Text-to-Speech (TTS) models have made huge progress, they still face challenges in controlling the emotional expression in the generated speech. In this work, we propose EmoVoice, a novel emotion-controllable TTS model that exploits large language models (LLMs) to enable fine-grained freestyle natural language emotion control, and a phoneme boost variant design that makes the model output phoneme tokens and audio tokens in parallel to enhance content consistency, inspired by chain-of-thought (CoT) and modality-of-thought (CoM) techniques. Besides, we introduce EmoVoice-DB, a high-quality 40-hour English emotion dataset featuring expressive speech and fine-grained emotion labels with natural language descriptions. EmoVoice achieves state-of-the-art performance on the English EmoVoice-DB test set using only synthetic training data, and on the Chinese Secap test set using our in-house data. We further investigate the reliability of existing emotion evaluation metrics and their alignment with human perceptual preferences, and explore using SOTA multimodal LLMs GPT-4o-audio and Gemini to assess emotional speech. Dataset, code, checkpoints and demo samples are available at https://github.com/yanghaha0908/EmoVoice. Guanrou Yang, Qian Chen 0003, Ziyang Ma 0001, Wen Wang 0019, Tianrui Wang, Yifan Yang 0005, Zhikang Niu, Wenrui Liu 0003, Fan Yu 0002, Zhihao Du, Zhifu Gao, Shiliang Zhang, Xie Chen 0001 |
ACM Multimedia | 6 |
| 2025 | Towards Minimizing Feature Drift in Model Merging: Layer-wise Task Vector Fusion for Adaptive Knowledge IntegrationabstractMulti-task model merging aims to consolidate knowledge from multiple fine-tuned task-specific experts into a unified model while minimizing performance degradation. Existing methods primarily approach this by minimizing differences between task-specific experts and the unified model, either from a parameter-level or a task-loss perspective. However, parameter-level methods exhibit a significant performance gap compared to the upper bound, while task-loss approaches entail costly secondary training procedures. In contrast, we observe that performance degradation closely correlates with feature drift, i.e., differences in feature representations of the same sample caused by model merging. Motivated by this observation, we propose Layer-wise Optimal Task Vector Merging (LOT Merging), a technique that explicitly minimizes feature drift between task-specific experts and the unified model in a layer-by-layer manner. LOT Merging can be formulated as a convex quadratic optimization problem, enabling us to analytically derive closed-form solutions for the parameters of linear and normalization layers. Consequently, LOT Merging achieves efficient model consolidation through basic matrix operations. Extensive experiments across vision and vision-language benchmarks demonstrate that LOT Merging significantly outperforms baseline methods, achieving improvements of up to 4.4% (ViT-B/32) over state-of-the-art approaches. The source code is available at https://github.com/SunWenJu123/model-merging. Wenju Sun, Qingyong Li, Wen Wang 0019, Yang Liu 0352, Boyang Li 0001 |
NeurIPS | 3 |
| 2025 | Decoupled likelihood modeling: A scalable approach for incremental generalized category discovery
Wenju Sun, Qingyong Li, Wen Wang 0019 |
Neurocomputing | 5 |
| 2024 | Incremental Learning via Robust Parameter Posterior FusionabstractThe posterior estimation of parameters based on Bayesian theory is a crucial technique in Incremental Learning (IL). The estimated posterior is typically utilized to impose loss regularization, which aligns the current training model parameters with the previously learned posterior to mitigate catastrophic forgetting, a major challenge in IL. However, this additional loss regularization can also impose detriment to the model learning, preventing it from reaching the true global optimum. To overcome this limitation, this paper introduces a novel Bayesian IL framework, Robust Parameter Posterior Fusion (RP2F). Unlike traditional methods, RP2F directly estimates the parameter posterior for new data without introducing extra loss regularization, which allows the model to accommodate new knowledge more sufficiently. It then fuses this new posterior with the existing ones based on the Maximum A Posteriori (MAP) principle, ensuring effective knowledge sharing across tasks. Furthermore, RP2F incorporates a common parameter-robustness priori to facilitate a seamless integration during posterior fusion. Comprehensive experiments on CIFAR-10, CIFAR-100, and Tiny-ImageNet datasets show that RP2F not only effectively mitigates catastrophic forgetting but also achieves backward knowledge transfer. Wenju Sun, Qingyong Li, Wen Wang 0019 |
ACM Multimedia | 4 |
| 2024 | Spatial-temporal graph-guided global attention network for video-based person re-identification
Xiaobao Li, Wen Wang 0019, Qingyong Li |
Mach. Vis. Appl. | 2 |
| 2024 | CLDiff: Weakly Supervised Cloud Detection With Denoising Diffusion Probabilistic ModelsabstractCloud detection is an essential step in remote sensing (RS) image processing, contributing to various applications. However, existing fully supervised cloud detection methods rely on massive pixel-wise annotations, which are expensive and time-consuming. To alleviate the annotation burden, weakly supervised cloud detection (WSCD) has received extensive attention recently. One standard approach performs cloud detection within a classification paradigm, which inevitably faces category ambiguity when detecting semitransparent clouds. To tackle this problem, we propose a novel WSCD framework based on the diffusion model, termed CLDiff. Specifically, a multiscale feature rectification (MFR) module is introduced to extract multiscale semantic features in the encoder, enabling a definite identification of clouds and mitigating interference from bright objects in the background. Considering that clouds exhibit varying optical thicknesses, a diffusion decoder is developed to model the intraclass variations of clouds in a generative strategy, improving thin cloud detection. Initially, it devises a Gaussian modulation function to recalibrate ambiguous cloud activations and emphasize semitransparent clouds. Subsequently, these modulated activations serve as semantic guidance to optimize the diffusion process. This approach enables CLDiff to activate cloud contours under definite semantic conditions and avoids the additional branches for semantic learning as found in previous methods. Experimental results demonstrate that CLDiff achieves state-of-the-art performance in WSCD. A public reference implementation of this work in PyTorch is available athttps://github.com/YLiu-creator/CLDiff. Yang Liu 0352, Qingyong Li, Zhigang Yao, Tony Z. Qiu, Wen Wang 0019 |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2023 | Decoupling Learning and Remembering: a Bilevel Memory Framework with Knowledge Projection for Task-Incremental LearningabstractThe dilemma between plasticity and stability arises as a common challenge for incremental learning. In contrast, the human memory system is able to remedy this dilemma owing to its multilevel memory structure, which motivates us to propose a Bilevel Memory system with Knowledge Projection (BMKP) for incremental learning. BMKP decouples the functions of learning and remembering via a bilevel-memory design: a working memory responsible for adaptively model learning, to ensure plasticity; a long-term memory in charge of enduringly storing the knowledge incorporated within the learned model, to guarantee stability. However, an emerging issue is how to extract the learned knowledge from the working memory and assimilate it into the long-term memory. To approach this issue, we reveal that the parameters learned by the working memory are actually residing in a redundant high-dimensional space, and the knowledge incorporated in the model can have a quite compact representation under a group of pattern basis shared by all incremental learning tasks. Therefore, we propose a knowledge projection process to adaptively maintain the shared basis, with which the loosely organized model knowledge of working memory is projected into the compact representation to be remembered in the long-term memory. We evaluate BMKP on CIFAR-10, CIFAR-100, and Tiny-ImageNet. The experimental results show that BMKP achieves state-of-the-art performance with lower memory usage11The code is available at https://github.com/SunWenJu123/BMKP. Wenju Sun, Qingyong Li, Jing Zhang 0058, Wen Wang 0019 |
CVPR | 4 |
| 2023 | Confidence-adapted meta-interaction for unsupervised person re-identification
Xiaobao Li, Qingyong Li, Wenyuan Xue, Yang Liu 0352, Fengjiao Liang, Wen Wang 0019 |
Appl. Intell. | 6 |
| 2023 | Multi-granularity Pseudo-label Collaboration for unsupervised person re-identification
Xiaobao Li, Qingyong Li, Fengjiao Liang, Wen Wang 0019 |
Comput. Vis. Image Underst. | 4 |
| 2023 | Class Incremental Learning based on Identically Distributed Parallel One-Class Classifiers
Wenju Sun, Qingyong Li, Jing Zhang 0058, Wen Wang 0019 |
Neurocomputing | 4 |
| 2023 | A General Dual-Branch Framework for Land Cover Mapping Models With Multispectral DataabstractLand cover mapping based on multispectral images can, in principle, be considered an application of semantic segmentation, but land cover mapping inputs include near-infrared (NIR) data in addition to RGB data. It has been experimentally found that LULC mapping performance based on RGB data alone is better than that based on RGB and NIR data when using some established single-branch encoder–decoder models. To address this issue, we propose a dual-branch encoder–decoder (DBED) framework that can be applied to existing encoder–decoder models for semantic segmentation. First, the multispectral input data is divided into two parts: RGB and NIR and fed to the respective branch for encoding. The dual-branch structure facilitates cross-modal complementary information encoding without deteriorating the original RGB modality-specific feature extraction. Second, an attention module named the multispectral attention module (MSAM) is proposed to mine the contextual correlation between the multispectral feature maps, leading to further performance boosting. We apply this framework to three mainstream semantic segmentation models and validate it on the Gaofen Image Dataset (GID). Experimental results show that this structure brings performance improvements. The source code of DBED is publicly available athttps://github.com/MalignusCN/DBED. Qingyong Li, Yang Liu 0352, Wen Wang 0019 |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2023 | Exemplar-free class incremental learning via discriminative and comparable parallel one-class classifiers
Wenju Sun, Qingyong Li, Jing Zhang 0058, Danyu Wang, Wen Wang 0019 |
Pattern Recognit. | 5 |
| 2023 | Leveraging Physical Rules for Weakly Supervised Cloud Detection in Remote Sensing ImagesabstractCloud detection plays a significant role in remote sensing image applications. Existing deep learning-based cloud detection methods rely on massive precise pixel-wise annotations, which are time-consuming and expensive. To alleviate this problem, we propose a weakly supervised cloud detection framework that leverages physical rules to generate weak supervision for cloud detection in remote sensing images. Specifically, a rule-based adaptive pseudo labeling (RAPL) algorithm is devised to adaptively annotate potential cloud pixels based on cloud spectral properties without manual intervention. Unlike existing physical annotations using fixed thresholds, RAPL employs the bidirectional threshold segmentation and adaptive gating mechanism to annotate cloud and boundary masks with more explicit semantic categories and spatial structures separately. Subsequently, these pseudo masks are treated as weak supervision to optimize the heuristic cloud detection network for pixel-wise segmentation. Considering that clouds appear as complex geometric structures and nonuniform spectral reflectance, a deformable boundary refining module is designed to enhance the modeling ability of spatial transformation and activate sharp boundaries from translucent cloud regions. Moreover, a harmonic loss is employed to recognize clouds with nonuniform spectral reflectance and suppress the interference of bright backgrounds. Extensive experiments on the GF-1, L8 Biome, and WDCD datasets demonstrate that the proposed method achieves state-of-the-art results. A public reference implementation of this work in PyTorch is available at https://github.com/NiAn-creator/HeuristicCloudDetection. Yang Liu 0352, Qingyong Li, Xiaobao Li, Shuyi He, Fengjiao Liang, Zhigang Yao, Wen Wang 0019 |
IEEE Trans. Geosci. Remote. Sens. | 8 |
| 2022 | Distribution and gradient constrained embedding model for zero-shot learning with fewer seen samples
Jing Zhang 0058, Wen Wang 0019, Wenju Sun, Zhirong Yang, Qingyong Li |
Knowl. Based Syst. | 3 |
| 2022 | Semantic Segmentation of Remote Sensing Images With Self-Supervised Semantic-Aware InpaintingabstractSemantic segmentation of remote sensing imageries plays a crucial role in resource exploration, urban planning, weather forecasting, etc. For this task, deep learning-based methods have shown significant achievement, typically trained with large-scale labeled data. However, these methods often suffer the performance deterioration facing limited labeled data in real-world applications. To address this problem, a novel self-supervised semantic segmentation framework is proposed for remote sensing imageries with limited labeled data. Specifically, image inpainting is acted as pixel-level pretext task for learning dense feature representations suitable for semantic segmentation. Further, rather than trivially leveraging the conventional random inpainting strategy, a novel adversarial training scheme is proposed to drive the pretext task to adaptively mask and restore salient local regions. The adversarial training scheme consists of instructor network and inpainting network, the instructor network increasingly predicts meaningful salient regions as erased regions, and meanwhile the inpainting network seeks for restoring the corrupted image as pretext task to learn its intrinsic representation. Moreover, the structural similarity (SSIM) is applied as a patch-level loss function for semantic segmentation considering that remote sensing images are highly structured. The experimental results on the ISPRS Potsdam dataset demonstrate that our method outperforms state-of-the-art self-supervised methods and the ImageNet pre-training methods. The source code is available at https://github.com/JasmineBJTU/self-supervised_RSSS. Shuyi He, Qingyong Li, Yang Liu 0352, Wen Wang 0019 |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2022 | DCNet: A Deformable Convolutional Cloud Detection Network for Remote Sensing ImageryabstractRecently, deep convolutional neural networks (CNNs) have made important progress in cloud detection with powerful representation learning capability and yield significant performance. However, most existing CNN-based cloud detection methods still face serious challenges because of the variable geometry of clouds and the complexity of underlying surfaces. It is attributed that they only use the fixed grid to extract contextual information, which lacks internal mechanisms to handle the geometric transformations of clouds. To tackle this problem, we propose a deformable convolutional cloud detection network with an encoder-decoder architecture, named DCNet, which can enhance the adaptability of a model to cloud variations. Specifically, we introduce deformable convolution blocks at the encoder to capture saliency spatial contexts adaptively based on the morphological characteristics of clouds and generate high-level semantic representations. After this, we incorporate skip-connection mechanisms into the decoder that integrate low-level spatial contexts as guidance to recover high-level semantic pixel localization and export precise cloud-detection results. Extensive experiments on the GF-1 wide field-of-view (WFV) Satellite Imagery demonstrate that DCNet outperforms several state-of-the-art methods. A public reference implementation of our proposed model in PyTorch is available athttps://github.com/NiAn-creator/deformableCloudDetection.git. Yang Liu 0352, Wen Wang 0019, Qingyong Li, Min Min, Zhigang Yao |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2022 | A zero-shot learning framework via cluster-prototype matching
Jing Zhang 0058, Qingyong Li, Wen Wang 0019, Wenju Sun, Chuan Shi 0001, Zhengming Ding |
Pattern Recognit. | 4 |
| 2022 | An Unsupervised Multi-Shot Person Re-Identification Method via Mutual Normalized Sparse Representation and Stepwise LearningabstractDue to abundant prior information and widespread applications, multi-shot based person re-identification has drawn increasing attention in recent years. In this paper, the high labeling cost and huge unlabeled data motivate us to focus on the unsupervised scenario and a unified coarse-to-fine framework is proposed, named by Mutual Normalized Sparse Representation (MNSR). Our method is an iteration procedure and each iteration involves two key steps: label estimation and metric model learning. In the former, we present a MNSR model to infer the pairwise labels of cross-camera by endowing sparse representation coefficient with the probability property. MNSR explicitly takes the mutually correlation between cameras into consideration and thus produces more accurate results. Meanwhile, we propose a probability-guided positive pairwise label prediction method to mine hard positive samples. For the latter, we learn a metric model with the estimated pairwise labels as supervision. In this procedure, we select some reliable labels for training by configuring with a stepwise learning method, rather than use all the estimated pair samples. This procedure helps to prevent the noise samples damaging the learning of discriminative metric model, especially for the initial iterations. Extensive experiments are conducted on four publicly available datasets, including PRID 2011, iLIDS-VID, SAIVT-SoftBio and MARS, and the results demonstrate the superior performance of the MNSR method in comparison with state-of-the-art unsupervised multi-shot person re-identification methods. Xiaobao Li, Qingyong Li, Wen Wang 0019, Lijun Guo |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2021 | TGRNet: A Table Graph Reconstruction Network for Table Structure RecognitionabstractA table arranging data in rows and columns is a very effective data structure, which has been widely used in business and scientific research. Considering large-scale tabular data in online and offline documents, automatic table recognition has attracted increasing attention from the document analysis community. Though human can easily understand the structure of tables, it remains a challenge for machines to understand that, especially due to a variety of different table layouts and styles. Existing methods usually model a table as either the markup sequence or the adjacency matrix between different table cells, failing to address the importance of the logical location of table cells, e.g., a cell is located in the first row and the second column of the table. In this paper, we reformulate the problem of table structure recognition as the table graph reconstruction, and propose an end-to-end trainable table graph reconstruction network (TGRNet) for table structure recognition. Specifically, the proposed method has two main branches, a cell detection branch and a cell logical location branch, to jointly predict the spatial location and the logical location of different cells. Experimental results on three popular table recognition datasets and a new dataset with table graph annotations (TableGraph-350K) demonstrate the effectiveness of the proposed TGRNet for table structure recognition. Code and annotations will be made publicly available at https://github.com/xuewenyuan/TGRNet. Wenyuan Xue, Baosheng Yu, Wen Wang 0019, Dacheng Tao, Qingyong Li |
ICCV | 3 |
| 2021 | A Question Answering System for Unstructured Table ImagesabstractQuestion answering over tables is a very popular semantic parsing task in natural language processing (NLP). However, few existing methods focus on table images, even though there are usually large-scale unstructured tables in practice (e.g., table images). Table parsing from images is nontrivial since it is closely related to not only NLP but also computer vision (CV) to parse the tabular structure from an image. In this demo, we present a question answering system for unstructured table images. The proposed system mainly consists of 1) a table recognizer to recognize the tabular structure from an image and 2) a table parser to generate the answer to a natural language question over the table. In addition, to train the model, we further provide table images and structure annotations for two widely used semantic parsing datasets. Specifically, the test set is used for this demo, from where the users can either choose from default questions or enter a new custom question. Wenyuan Xue, Wen Wang 0019, Qingyong Li, Baosheng Yu, Yibing Zhan, Dacheng Tao |
ACM Multimedia | 3 |
| 2021 | Unsupervised Multi-shot Person Re-identification via Dynamic Bi-directional Normalized Sparse Representation
Xiaobao Li, Wen Wang 0019, Qingyong Li, Lijun Guo |
MMM (1) | 2 |
| 2021 | Fine-Grained Image-Text Retrieval via Discriminative Latent Space LearningabstractFine-grained image-text retrieval aims at searching relevant images among fine-grained classes given a text query or in a reverse way. The challenges are not only bridging the gap between two heterogeneous modalities but also dealing with large inter-class similarity and intra-class variance existed in fine-grained data. To deal with the above challenges, we propose a Discriminative Latent Space Learning (DLSL) method for fine-grained image-text retrieval. Concretely, image and text features are extracted for capturing the subtle difference in fine-grained data. Subsequently, based on the extracted features, we perform couple dictionary learning to align the heterogeneous data in a uniform latent space. To make such alignment discriminative enough for the fine-grained task, the learned latent space is endowed with discriminative property via learning a discriminative map. Comprehensive experiments on fine-grained datasets demonstrate the effectiveness of our approach. Wen Wang 0019, Qingyong Li |
IEEE Signal Process. Lett. | 2 |
| 2018 | Discriminant Analysis on Riemannian Manifold of Gaussian Distributions for Face Recognition With Image SetsabstractTo address the problem of face recognition with image sets, we aim to capture the underlying data distribution in each set and thus facilitate more robust classification. To this end, we represent image set as the Gaussian mixture model (GMM) comprising a number of Gaussian components with prior probabilities and seek to discriminate Gaussian components from different classes. Since in the light of information geometry, the Gaussians lie on a specific Riemannian manifold, this paper presents a method named discriminant analysis on Riemannian manifold of Gaussian distributions (DARG). We investigate several distance metrics between Gaussians and accordingly two discriminative learning frameworks are presented to meet the geometric and statistical characteristics of the specific manifold. The first framework derives a series of provably positive definite probabilistic kernels to embed the manifold to a high-dimensional Hilbert space, where conventional discriminant analysis methods developed in Euclidean space can be applied, and a weighted Kernel discriminant analysis is devised which learns discriminative representation of the Gaussian components in GMMs with their prior probabilities as sample weights. Alternatively, the other framework extends the classical graph embedding method to the manifold by utilizing the distance metrics between Gaussians to construct the adjacency graph, and hence the original manifold is embedded to a lower-dimensional and discriminative target manifold with the geometric structure preserved and the interclass separability maximized. The proposed method is evaluated by face identification and verification tasks on four most challenging and largest databases, YouTube Celebrities, COX, YouTube Face DB, and Point-and-Shoot Challenge, to demonstrate its superiority over the state-of-the-art.To address the problem of face recognition with image sets, we aim to capture the underlying data distribution in each set and thus facilitate more robust classification. To this end, we represent image set as the Gaussian mixture model (GMM) comprising a number of Gaussian components with prior probabilities and seek to discriminate Gaussian components from different classes. Since in the light of information geometry, the Gaussians lie on a specific Riemannian manifold, this paper presents a method named discriminant analysis on Riemannian manifold of Gaussian distributions (DARG). We investigate several distance metrics between Gaussians and accordingly two discriminative learning frameworks are presented to meet the geometric and statistical characteristics of the specific manifold. The first framework derives a series of provably positive definite probabilistic kernels to embed the manifold to a high-dimensional Hilbert space, where conventional discriminant analysis methods developed in Euclidean space can be applied, and a weighted Kernel discriminant analysis is devised which learns discriminative representation of the Gaussian components in GMMs with their prior probabilities as sample weights. Alternatively, the other framework extends the classical graph embedding method to the manifold by utilizing the distance metrics between Gaussians to construct the adjacency graph, and hence the original manifold is embedded to a lower-dimensional and discriminative target manifold with the geometric structure preserved and the interclass separability maximized. The proposed method is evaluated by face identification and verification tasks on four most challenging and largest databases, YouTube Celebrities, COX, YouTube Face DB, and Point-and-Shoot Challenge, to demonstrate its superiority over the state-of-the-art. Wen Wang 0019, Ruiping Wang 0001, Zhiwu Huang, Shiguang Shan, Xilin Chen 0001 |
IEEE Trans. Image Process. | 1 |
| 2017 | Discriminative Covariance Oriented Representation Learning for Face Recognition with Image SetsabstractFor face recognition with image sets, while most existing works mainly focus on building robust set models with hand-crafted feature, it remains a research gap to learn better image representations which can closely match the subsequent image set modeling and classification. Taking sample covariance matrix as set model in the light of its recent promising success, we present a Discriminative Covariance oriented Representation Learning (DCRL) framework to bridge the above gap. The framework constructs a feature learning network (e.g. a CNN) to project the face images into a target representation space, and the network is trained towards the goal that the set covariance matrix calculated in the target space has maximum discriminative ability. To encode the discriminative ability of set covariance matrices, we elaborately design two different loss functions, which respectively lead to two different representation learning schemes, i.e., the Graph Embedding scheme and the Softmax Regression scheme. Both schemes optimize the whole network containing both image representation mapping and set model classification in a joint learning manner. The proposed method is extensively validated on three challenging and large scale databases for the task of face recognition with image sets, i.e., YouTube Celebrities, YouTube Face DB and Point-and-Shoot Challenge. Wen Wang 0019, Ruiping Wang 0001, Shiguang Shan, Xilin Chen 0001 |
CVPR | 1 |
| 2017 | Prototype Discriminative Learning for Image Set ClassificationabstractThis letter presents a prototype discriminative learning (PDL) method for image set classification. We aim to simultaneously learn prototypes and a linear discriminative projection to drive that in the target subspace each image set can be discriminated with its nearest neighbor prototype. To reveal the unseen appearance variations implicitly in an image set, the prototypes are actually “virtual,” which do not certainly appear in the set but are searched in the corresponding affine hull. Moreover, to enhance the stability and robustness of the learned target subspace, an orthogonality constraint is imposed on the projection. Thus, to optimize the prototypes and the projection jointly, we design a specific gradient descent mechanism by updating the projection on Stiefel manifold and the prototypes in Euclidean space in an alternative optimization manner. Experimental results on four challenging databases demonstrate the superiority of the proposed PDL method. Wen Wang 0019, Ruiping Wang 0001, Shiguang Shan, Xilin Chen 0001 |
IEEE Signal Process. Lett. | 1 |
| 2016 | Prototype Discriminative Learning for Face Image Set Classification
Wen Wang 0019, Ruiping Wang 0001, Shiguang Shan, Xilin Chen 0001 |
ACCV (3) | 1 |
| 2015 | Discriminant analysis on Riemannian manifold of Gaussian distributions for face recognition with image setsabstractThis paper presents a method named Discriminant Analysis on Riemannian manifold of Gaussian distributions (DARG) to solve the problem of face recognition with image sets. Our goal is to capture the underlying data distribution in each set and thus facilitate more robust classification. To this end, we represent image set as Gaussian Mixture Model (GMM) comprising a number of Gaussian components with prior probabilities and seek to discriminate Gaussian components from different classes. In the light of information geometry, the Gaussians lie on a specific Riemannian manifold. To encode such Riemannian geometry properly, we investigate several distances between Gaussians and further derive a series of provably positive definite probabilistic kernels. Through these kernels, a weighted Kernel Discriminant Analysis is finally devised which treats the Gaussians in GMMs as samples and their prior probabilities as sample weights. The proposed method is evaluated by face identification and verification tasks on four most challenging and largest databases, YouTube Celebrities, COX, YouTube Face DB and Point-and-Shoot Challenge, to demonstrate its superiority over the state-of-the-art. Wen Wang 0019, Ruiping Wang 0001, Zhiwu Huang, Shiguang Shan, Xilin Chen 0001 |
CVPR | 1 |