EDBT 2026 Demo / reviewers in the wild / expert
Wei Jiang 0009
dblp:21/3839-9
· DBLP profile ↗
45ranked-venue papers
2as first author
24since 2021 · last 2026
0000-0002-9240-5851ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 24 · 2 first-author · 13 since 2021Graphics, computer vision, multimedia, augmented reality and games · 21 · 1 first-author · 11 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 2 since 2021Human-computer interaction and ubiquitous computing · 4 · 1 since 2021Security and privacy · 3 · 3 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Concord: Concept-Informed Diffusion for Dataset DistillationabstractDataset distillation (DD) has witnessed significant progress in creating small datasets that encapsulate rich information from large original ones. Particularly, methods based on generative priors show promising performance while maintaining computational efficiency and cross-architecture generalization. However, the generation process lacks explicit controllability for each sample. Previous distillation methods primarily match the real distribution from the perspective of the entire dataset, whereas overlooking concept completeness at the instance level. The missing or incorrectly represented object details cannot be efficiently compensated due to the constrained sample amount typical in DD settings. To this end, we propose incorporating the concept understanding of large language models (LLMs) to perform Concept-Informed Diffusion (Concord) for dataset distillation. Specifically, distinguishable and fine-grained concepts are retrieved based on category labels to inform the denoising process and refine essential object details. These concepts can be applied to any diffusion-based DD framework to enhance both the controllability and interpretability of the distilled image generation, without relying on pre-trained classifiers. We demonstrate the efficacy of Concord by achieving state-of-the-art performance on ImageNet-1K and subsets. Code is released at Concord. Jianyang Gu, Ruoxi Jia 0001, Saeed Vahidian, Vyacheslav Kungurtsev, Wei Jiang 0009, Yiran Chen 0001 |
WACV | 6 |
| 2026 | Stochastic style perturbation modelling for visible-Infrared person re-Identification with severely modality imbalance
Jianyang Gu, Mingyu Wang 0004, Q. M. Jonathan Wu, Wei Jiang 0009 |
Neural Networks | 6 |
| 2026 | Quantifying and Overcoming the Bias Nature of Modality for Visible-Infrared Person Re-Identification
Daoxun Xia, Q. M. Jonathan Wu, Wei Jiang 0009 |
IEEE Trans. Inf. Forensics Secur. | 6 |
| 2025 | Group Distributionally Robust Dataset Distillation with Risk MinimizationabstractDataset distillation (DD) has emerged as a widely adopted technique for crafting a synthetic dataset that captures the essential information of a training dataset, facilitating the training of accurate neural models. Its applications span various domains, including transfer learning, federated learning, and neural architecture search. The most popular methods for constructing the synthetic data rely on matching the convergence properties of training the model with the synthetic dataset and the training dataset. However, using the empirical loss as the criterion must be thought of as auxiliary in the same sense that the training set is an approximate substitute for the population distribution, and the latter is the data of interest. Yet despite its popularity, an aspect that remains unexplored is the relationship of DD to its generalization, particularly across uncommon subgroups. That is, how can we ensure that a model trained on the synthetic dataset performs well when faced with samples from regions with low population density? Here, the representativeness and coverage of the dataset become salient over the guaranteed training error at inference. Drawing inspiration from distributionally robust optimization, we introduce an algorithm that combines clustering with the minimization of a risk measure on the loss to conduct DD. We provide a theoretical rationale for our approach and demonstrate its effective generalization and robustness across subgroups through numerical experiments. Saeed Vahidian, Mingyu Wang 0004, Jianyang Gu, Vyacheslav Kungurtsev, Wei Jiang 0009, Yiran Chen 0001 |
ICLR | 5 |
| 2025 | CoMix: Collaborative Mixed Learning via Style Fuzzy Normalization for Visible-Infrared Person Re-IdentificationabstractVisible–infrared person re-identification (VI-ReID) focuses on accurately matching individuals across different imaging modalities. Existing studies focus on generating modality-consistent images at the pixel level through the use of generative adversarial networks (GANs) to mitigate the impact of modality discrepancies. However, these methods face significant challenges in overcoming the limitation that synthesized samples from different modalities may suffer from semantic distortion. In this work, we propose an online one-stage style fuzzy normalization (SFN) method to generate modality-fuzzy features in the latent space while regularizing the model’s predictions. Specifically, SFN adaptively mixes the feature statistics of two random modality instances of the same identity in a single forward pass during training. In this process, to enhance the richness of modality interaction information, we design a novel causality balance loss, which enforces the generated fuzzy features to be independent of their initial modality while simultaneously encouraging them to align more closely with the other modality. Furthermore, we introduce an identity-aware consistency loss to regularize the predictions between the original and SFN-generated features to ensure semantic consistency. In contrast to prior work, SFN is a plug-and-play module that does not rely on any generative-based models, making it highly adaptable to various network architectures. Extensive experiments were performed on three public cross-modality datasets to ensure fair and reliable comparisons. The empirical results demonstrate the clear superiority of our method over previous state-of-the-art methods. Jianyang Gu, Mingyu Wang 0004, Q. M. Jonathan Wu, Wei Jiang 0009 |
IEEE Trans. Syst. Man Cybern. Syst. | 6 |
| 2024 | Summarizing Stream Data for Memory-Constrained Online Continual LearningabstractReplay-based methods have proved their effectiveness on online continual learning by rehearsing past samples from an auxiliary memory. With many efforts made on improving training schemes based on the memory, however, the information carried by each sample in the memory remains under-investigated. Under circumstances with restricted storage space, the informativeness of the memory becomes critical for effective replay. Although some works design specific strategies to select representative samples, by only employing a small number of original images, the storage space is still not well utilized. To this end, we propose to Summarize the knowledge from the Stream Data (SSD) into more informative samples by distilling the training characteristics of real images. Through maintaining the consistency of training gradients and relationship to the past tasks, the summarized samples are more representative for the stream data compared to the original images. Extensive experiments are conducted on multiple online continual learning benchmarks to support that the proposed SSD method significantly enhances the replay effects. We demonstrate that with limited extra computational overhead, SSD provides more than 3% accuracy boost for sequential CIFAR-100 under extremely restricted memory buffer. Code in https://github.com/vimar-gu/SSD. Jianyang Gu, Kai Wang 0036, Wei Jiang 0009, Yang You 0001 |
AAAI | 3 |
| 2024 | Efficient Dataset Distillation via Minimax DiffusionabstractDataset distillation reduces the storage and computational consumption of training a network by generating a small surrogate dataset that encapsulates rich information of the original large-scale one. However, previous distillation methods heavily rely on the sample-wise iterative optimization scheme. As the images-per-class (IPC) setting or image resolution grows larger, the necessary computation will demand overwhelming time and resources. In this work, we intend to incorporate generative diffusion techniques for computing the surrogate dataset. Observing that key factors for constructing an effective surrogate dataset are representativeness and diversity, we design additional minimax criteria in the generative training to enhance these facets for the generated images of diffusion models. We present a theoretical model of the process as hierarchical diffusion control demonstrating the flexibility of the diffusion process to target these criteria without jeopardizing the faithfulness of the sample to the desired distribution. The proposed method achieves state-of-the-art validation performance while demanding much less computational resources. Under the 100-IPC setting on Image Woof, our method requires less than one-twentieth the distillation time of previous methods, yet yields even better performance. Source code and generated data are available in https://github.com/vimar-gu/MinimaxDiffusion. Jianyang Gu, Saeed Vahidian, Vyacheslav Kungurtsev, Wei Jiang 0009, Yang You 0001, Yiran Chen 0001 |
CVPR | 5 |
| 2024 | Dynamic gradient reactivation for backward compatible person re-identification
Xiao Pan 0001, Hao Luo 0004, Fan Wang 0019, Hao Li 0030, Wei Jiang 0009, Jianming Zhang 0005, Jianyang Gu, Peike Li |
Pattern Recognit. | 6 |
| 2024 | Modality Bias Calibration Network via Information Disentanglement for Visible-Infrared Person ReidentificationabstractVisible–infrared person reidentification (VI-ReID) in social surveillance systems involves analyzing social behavior using nonoverlapping cross-modality camera sets. It often has poor retrieval performance under modality gap. One way to alleviate such the modality discrepancy is to learn shared person features that are generalizable across different modalities. However, because of significant differences in color between the visible and infrared images, the learned share features are always inclined to specific information of corresponding modality. To this end, we propose a modality bias calibration network (MBCNet) that filters out identity-irrelevant interference and recalibrates the learned modality-shared features. Specifically, to emphasize the modality-shared cues, we employ a feature decomposition module in the feature-level to filter out style variations and extract identity-relevant discriminative cues from the residual feature. In order to achieve a better disentanglement, a dual ranking entropy constraint is further proposed to ensure that the learned features contain only identity-relevant information and discard style-relevant information. Simultaneously, we design a decorrelated orthogonality Loss to ensure the disentangled features are not correlated with each other. Through comprehensive experiments, we demonstrate that MBCNet significantly improves the cross-modality retrieval performance in social surveillance systems and effectively addresses the modality bias training issue. Hao Luo 0004, Xiantao Peng, Wei Jiang 0009 |
IEEE Trans. Comput. Soc. Syst. | 4 |
| 2024 | Region Generation and Assessment Network for Occluded Person Re-IdentificationabstractPerson Re-identification (ReID) plays a more and more crucial role in recent years with a wide range of applications. Existing ReID methods are suffering from the challenges of misalignment and occlusions, which degrade the performance dramatically. Most methods tackle such challenges by utilizing external tools to locate body parts or exploiting matching strategies. Nevertheless, the inevitable domain gap between the datasets utilized for external tools and the ReID datasets and the complicated matching process make these methods unreliable and sensitive to noises. In this paper, we propose a Region Generation and Assessment Network (RGANet) to effectively and efficiently detect the human body regions and highlight the important regions. In the proposed RGANet, we first devise a Region Generation Module (RGM) which utilizes the pre-trained CLIP to locate the human body regions using semantic prototypes extracted from text descriptions. Learnable prompt is designed to eliminate domain gap between CLIP datasets and ReID datasets. Then, to measure the importance of each generated region, we introduce a Region Assessment Module (RAM) that assigns confidence scores to different regions and reduces the negative impact of the occlusion regions by lower scores. The RAM consists of a discrimination-aware indicator and an invariance-aware indicator, where the former indicates the capability to distinguish from different identities and the latter represents consistency among the images of the same class of human body regions. Extensive experimental results for six widely-used benchmarks including three tasks (occluded, partial, and holistic) demonstrate the superiority of RGANet against state-of-the-art methods. Shuting He, Kai Wang 0036, Hao Luo 0004, Fan Wang 0019, Wei Jiang 0009, Henghui Ding |
IEEE Trans. Inf. Forensics Secur. | 6 |
| 2024 | VGSG: Vision-Guided Semantic-Group Network for Text-Based Person SearchabstractText-based Person Search (TBPS) aims to retrieve images of target pedestrian indicated by textual descriptions. It is essential for TBPS to extract fine-grained local features and align them crossing modality. Existing methods utilize external tools or heavy cross-modal interaction to achieve explicit alignment of cross-modal fine-grained features, which is inefficient and time-consuming. In this work, we propose a Vision-Guided Semantic-Group Network (VGSG) for text-based person search to extract well-aligned fine-grained visual and textual features. In the proposed VGSG, we develop a Semantic-Group Textual Learning (SGTL) module and a Vision-guided Knowledge Transfer (VGKT) module to extract textual local features under the guidance of visual local clues. In SGTL, in order to obtain the local textual representation, we group textual features from the channel dimension based on the semantic cues of language expression, which encourages similar semantic patterns to be grouped implicitly without external tools. In VGKT, a vision-guided attention is employed to extract visual-related textual features, which are inherently aligned with visual cues and termed vision-guided textual features. Furthermore, we design a relational knowledge transfer, including a vision-language similarity transfer and a class probability transfer, to adaptively propagate information of the vision-guided textual features to semantic-group textual features. With the help of relational knowledge transfer, VGKT is capable of aligning semantic-group textual features with corresponding visual features without external tools and complex pairwise interaction. Experimental results on two challenging benchmarks demonstrate its superiority over state-of-the-art methods. Shuting He, Hao Luo 0004, Wei Jiang 0009, Xudong Jiang 0001, Henghui Ding |
IEEE Trans. Image Process. | 3 |
| 2024 | Inter-Intra Modality Knowledge Learning and Clustering Noise Alleviation for Unsupervised Visible-Infrared Person Re-IdentificationabstractUnsupervised visible-infrared person re-identification (USVI-ReID) is a valuable yet under-explored task that addresses the challenge of person retrieval from different modalities without annotations. USVI-ReID is faced with large intra and inter modality discrepancy as well as pseudo label noises. In this paper, we propose a novel learning approach termed Spectrum-Cam Aware and Refined Cross-Pair learning strategy (SCA-RCP), crafted to concurrently tackle these issues. The Dual-Spectrum Augmentation (DSA) method is first presented to mine inter-modality knowledge between the visible and infrared modalities by expanding the diversity of spectra for both modalities. Concurrently, to learn intra-modality knowledge, we divide the clusters into camera-based proxies and introduce the Cross-modal Cam-Aware Proxy learning method (CCAP) to pull together camera-based proxies that correspond to the same identity across modalities. Finally, to mitigate clustering noise, the innovative Refined Cross-Pair learning strategy (RCP) is devised comprising the Intra-modality Label Refinement (ILR) and Bi-directional Cross Pairing (BCP) method. ILR calculates maximum-likely label for each instance by employing the intra-modal cluster memories as a classifier, while BCP establish dependable cross-paired labels in a bidirectional manner. Extensive experiments on the cross-modality datasets demonstrate the superior performance of our model over the state-of-art method. Xiantao Peng, Wei Jiang 0009 |
IEEE Trans. Knowl. Data Eng. | 4 |
| 2023 | MSINet: Twins Contrastive Search of Multi-Scale Interaction for Object ReIDabstractNeural Architecture Search (NAS) has been increasingly appealing to the society of object Re-Identification (ReID), for that task-specific architectures significantly improve the retrieval performance. Previous works explore new optimizing targets and search spaces for NAS ReID, yet they neglect the difference of training schemes between image classification and ReID. In this work, we propose a novel Twins Contrastive Mechanism (TCM) to provide more appropriate supervision for ReID architecture search. TCM reduces the category overlaps between the training and validation data, and assists NAS in simulating real-world ReID training schemes. We then design a Multi-Scale Interaction (MSI) search space to search for rational interaction operations between multi-scale features. In addition, we introduce a Spatial Alignment Module (SAM) to further enhance the attention consistency confronted with images from different sources. Under the proposed NAS scheme, a specific architecture is automatically searched, named as MSINet. Extensive experiments demonstrate that our method surpasses state-of-the-art ReID methods on both indomain and cross-domain scenarios. Source code available in https://github.com/vimar-gu/MSINet. Jianyang Gu, Kai Wang 0036, Hao Luo 0004, Chen Chen 0114, Wei Jiang 0009, Yuqiang Fang, Shanghang Zhang, Yang You 0001, Jian Zhao 0006 |
CVPR | 5 |
| 2023 | Primitive Generation and Semantic-Related Alignment for Universal Zero-Shot SegmentationabstractWe study universal zero-shot segmentation in this work to achieve panoptic, instance, and semantic segmentation for novel categories without any training samples. Such zero-shot segmentation ability relies on inter-class relationships in semantic space to transfer the visual knowledge learned from seen categories to unseen ones. Thus, it is desired to well bridge semantic-visual spaces and apply the semantic relationships to visual feature learning. We introduce a generative model to synthesize features for unseen categories, which links semantic and visual spaces as well as address the issue of lack of unseen training data. Furthermore, to mitigate the domain gap between semantic and visual spaces, firstly, we enhance the vanilla generator with learned primitives, each of which contains fine-grained attributes related to categories, and synthesize unseen features by selectively assembling these primitives. Secondly, we propose to disentangle the visual feature into the semantic-related part and the semantic-unrelated part that contains useful visual classification clues but is less relevant to semantic representation. The inter-class relationships of semantic-related visual features are then required to be aligned with those in semantic space, thereby transferring semantic knowledge to visual feature learning. The proposed approach achieves impressively state-of-the-art performance on zero-shot panoptic segmentation, instance segmentation, and semantic segmentation. Shuting He, Henghui Ding, Wei Jiang 0009 |
CVPR | 3 |
| 2023 | Semantic-Promoted Debiasing and Background Disambiguation for Zero-Shot Instance SegmentationabstractZero-shot instance segmentation aims to detect and precisely segment objects of unseen categories without any training samples. Since the model is trained on seen categories, there is a strong bias that the model tends to classify all the objects into seen categories. Besides, there is a natural confusion between background and novel objects that have never shown up in training. These two challenges make novel objects hard to be raised in the final instance segmentation results. It is desired to rescue novel objects from background and dominated seen categories. To this end, we propose D2Zero with Semantic-Promoted Debiasing and Background Disambiguation to enhance the performance of Zero-shot instance segmentation. Semantic-promoted debiasing utilizes inter-class semantic relationships to involve unseen categories in visual feature training and learns an input-conditional classifier to conduct dynamical classification based on the input image. Background disambiguation produces image-adaptive background representation to avoid mistaking novel objects for background. Extensive experiments show that we significantly outperform previous state-of-the-art methods by a large margin, e.g., 16.86% improvement on COCO. Shuting He, Henghui Ding, Wei Jiang 0009 |
CVPR | 3 |
| 2023 | DREAM: Efficient Dataset Distillation by Representative MatchingabstractDataset distillation aims to synthesize small datasets with little information loss from original large-scale ones for reducing storage and training costs. Recent state-of-the-art methods mainly constrain the sample synthesis process by matching synthetic images and the original ones regarding gradients, embedding distributions, or training trajectories. Although there are various matching objectives, currently the strategy for selecting original images is limited to naive random sampling. We argue that random sampling overlooks the evenness of the selected sample distribution, which may result in noisy or biased matching targets. Besides, the sample diversity is also not constrained by random sampling. These factors together lead to optimization instability in the distilling process and degrade the training efficiency. Accordingly, we propose a novel matching strategy named as Dataset distillation by REpresentAtive Matching (DREAM), where only representative original images are selected for matching. DREAM is able to be easily plugged into popular dataset distillation frameworks and reduce the distilling iterations by more than 8 times without performance drop. Given sufficient training time, DREAM further provides significant improvements and achieves state-of-the-art performances. Jianyang Gu, Kai Wang 0036, Wei Jiang 0009, Yang You 0001 |
ICCV | 5 |
| 2023 | Prototype Adaption and Projection for Few- and Zero-Shot 3D Point Cloud Semantic SegmentationabstractIn this work, we address the challenging task of few-shot and zero-shot 3D point cloud semantic segmentation. The success of few-shot semantic segmentation in 2D computer vision is mainly driven by the pre-training on large-scale datasets like imagenet. The feature extractor pre-trained on large-scale 2D datasets greatly helps the 2D few-shot learning. However, the development of 3D deep learning is hindered by the limited volume and instance modality of datasets due to the significant cost of 3D data collection and annotation. This results in less representative features and large intra-class feature variation for few-shot 3D point cloud segmentation. As a consequence, directly extending existing popular prototypical methods of 2D few-shot classification/segmentation into 3D point cloud segmentation won't work as well as in 2D domain. To address this issue, we propose a Query-Guided Prototype Adaption (QGPA) module to adapt the prototype from support point clouds feature space to query point clouds feature space. With such prototype adaption, we greatly alleviate the issue of large feature intra-class variation in point cloud and significantly improve the performance of few-shot 3D segmentation. Besides, to enhance the representation of prototypes, we introduce a Self-Reconstruction (SR) module that enables prototype to reconstruct the support mask as well as possible. Moreover, we further consider zero-shot 3D point cloud semantic segmentation where there is no support sample. To this end, we introduce category words as semantic information and propose a semantic-visual projection model to bridge the semantic and visual spaces. Our proposed method surpasses state-of-the-art algorithms by a considerable 7.90% and 14.82% under the 2-way 1-shot setting on S3DIS and ScanNet benchmarks, respectively. Shuting He, Xudong Jiang 0001, Wei Jiang 0009, Henghui Ding |
IEEE Trans. Image Process. | 3 |
| 2023 | Transformer-Based Domain-Specific Representation for Unsupervised Domain Adaptive Vehicle Re-IdentificationabstractFully-supervised vehicle re-identification (re-ID) methods are faced with performance degradation when applied to new image domains. Therefore, developing unsupervised domain adaptation (UDA) to transfer the knowledge from learned source domain to new unlabeled target domain becomes an indispensable task. It is challenging because different domains have various image appearances, such as different backgrounds, illuminations and resolutions, especially when cameras have different viewpoints. To tackle this domain gap issue, a novel Transformer-based Domain-Specific Representation learning network (TDSR) is proposed to dynamically focus on corresponding detailed hints for each domain. Specifically, with the source and target domain being trained simultaneously, a domain encoding module is proposed to introduce domain information into the network. The original features of source and target domains are enriched with these domain encodings first, and then sequentially processed by a Transformer encoder to model contextual information and a decoder to summarize the encoded features into the final domain-specific feature representations. Moreover, we propose a Contrastive Clustering Loss (CCL) to directly optimize the distribution of features at cluster level. Instances are overall pulled closer to the prototype of the same identity, and pushed farther from the prototypes of different identities. It helps compact the clusters in the latent space and improve the discriminative capability of the network, leading to more accurate pseudo-label assignment in TDSR. Our method outperforms the state-of-the-art UDA methods on vehicle re-ID benchmark datasets VeRi and VehicleID on both real-world to real-world and synthetic to real-world settings. Jianyang Gu, Shuting He, Wei Jiang 0009 |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2022 | SFGN: Representing the sequence with one super frame for video person re-identification
Xiao Pan 0001, Hao Luo 0004, Wei Jiang 0009, Jianming Zhang 0005, Jianyang Gu, Peike Li |
Knowl. Based Syst. | 3 |
| 2022 | Multi-View Evolutionary Training for Unsupervised Domain Adaptive Person Re-IdentificationabstractClustering-based approaches have been successfully applied to unsupervised domain adaptation (UDA) tasks for person re-identification (Re-ID), where no annotations are provided in target domain. However, the clustering process is sensitive to noises, leading to imperfect pseudo labels that could damage the training performance. In this work, we propose a Multi-view Evolutionary Training (MET) method to effectively reduce noises in clustering results from two dimensions. First, to improve the clustering accuracy at each time frame (i.e. snapshot quality), a Multi-view Diffusion (MvD) module is proposed. Through capturing data relationships from multiple viewpoints and aggregating their information, noises and bias from each individual viewpoint can be eliminated, and more reliable similarity matrix can be produced for clustering. Second, to improve the temporal consistency between clustering at different iterations, i.e. temporal consistency, we propose an Evolutionary Local Refinement (ELR) module, which utilizes the previous clustering results to guide and improve current results, and further make the training process more stable and robust. Extensive experiments demonstrate that our method can provide clustering results with high quality, and achieve state-of-the-art performance on UDA Re-ID. Jianyang Gu, Hao Luo 0004, Fan Wang 0019, Hao Li 0030, Wei Jiang 0009 |
IEEE Trans. Inf. Forensics Secur. | 6 |
| 2022 | Modality-transfer generative adversarial network and dual-level unified latent representation for visible thermal Person re-identification
Wei Jiang 0009, Hao Luo 0004 |
Vis. Comput. | 2 |
| 2021 | TransReID: Transformer-based Object Re-IdentificationabstractExtracting robust feature representation is one of the key challenges in object re-identification (ReID). Although convolution neural network (CNN)-based methods have achieved great success, they only process one local neighborhood at a time and suffer from information loss on details caused by convolution and downsampling operators (e.g. pooling and strided convolution). To overcome these limitations, we propose a pure transformer-based object ReID framework named TransReID. Specifically, we first encode an image as a sequence of patches and build a transformer-based strong baseline with a few critical improvements, which achieves competitive results on several ReID benchmarks with CNN-based methods. To further enhance the robust feature learning in the context of transformers, two novel modules are carefully designed. (i) The jigsaw patch module (JPM) is proposed to rearrange the patch embeddings via shift and patch shuffle operations which generates robust features with improved discrimination ability and more diversified coverage. (ii) The side information embeddings (SIE) is introduced to mitigate feature bias towards camera/view variations by plugging in learnable embeddings to incorporate these non-visual clues. To the best of our knowledge, this is the first work to adopt a pure transformer for ReID research. Experimental results of TransReID are superior promising, which achieve state-of-the-art performance on both person and vehicle ReID benchmarks. Code is available at https://github.com/heshuting555/TransReID. Shuting He, Hao Luo 0004, Pichao Wang, Fan Wang 0019, Hao Li 0030, Wei Jiang 0009 |
ICCV | 6 |
| 2021 | Learning user interest with improved triplet deep ranking and web-image priors for topic-related video summarization
Mengjuan Fei, Wei Jiang 0009 |
Expert Syst. Appl. | 2 |
| 2021 | An efficient global representation constrained by Angular Triplet loss for vehicle re-identification
Jianyang Gu, Wei Jiang 0009, Hao Luo 0004, Hongyan Yu |
Pattern Anal. Appl. | 2 |
| 2020 | SMAP: Single-Shot Multi-person Absolute 3D Pose Estimation
Jianan Zhen, Jiaming Sun 0002, Wentao Liu 0002, Wei Jiang 0009, Hujun Bao, Xiaowei Zhou 0001 |
ECCV (15) | 5 |
| 2020 | Analytical derivatives for differentiable renderer: 3D pose estimation by silhouette consistency
Zaiqiang Wu, Wei Jiang 0009, Hongyan Yu |
J. Vis. Commun. Image Represent. | 2 |
| 2020 | STNReID: Deep Convolutional Networks With Pairwise Spatial Transformer Networks for Partial Person Re-IdentificationabstractPartial person re-identification (ReID) is a challenging task because only partial information of person images is available for matching target persons. Few studies, especially on deep learning, have focused on matching partial person images with holistic person images. This study presents a novel deep partial ReID framework based on pairwise spatial transformer networks (STNReID), which can be trained on existing holistic person datasets. STNReID includes a spatial transformer network (STN) module and a ReID module. The STN module samples an affined image (a semantically corresponding patch) from the holistic image to match the partial image. The ReID module extracts the features of the holistic, partial, and affined images. Competition (or confrontation) is observed between the STN module and the ReID module, and two-stage training is applied to acquire a strong STNReID for partial ReID. Experimental results show that our STNReID obtains 66.7% and 54.6% rank-1 accuracies on Partial-ReID and Partial-iLIDS datasets, respectively. These values are at par with those obtained with state-of-the-art methods. Hao Luo 0004, Wei Jiang 0009, Chi Zhang 0026 |
IEEE Trans. Multim. | 2 |
| 2020 | A Strong Baseline and Batch Normalization Neck for Deep Person Re-IdentificationabstractThis study proposes a simple but strong baseline for deep person re-identification (ReID). Deep person ReID has achieved great progress and high performance in recent years. However, many state-of-the-art methods design complex network structures and concatenate multi-branch features. In the literature, some effective training tricks briefly appear in several papers or source codes. The present study collects and evaluates these effective training tricks in person ReID. By combining these tricks, the model achieves 94.5% rank-1 and 85.9% mean average precision on Market1501 with only using the global features of ResNet50. The performance surpasses all existing global- and part-based baselines in person ReID. We propose a novel neck structure named as batch normalization neck (BNNeck). BNNeck adds a batch normalization layer after global pooling layer to separate metric and classification losses into two different feature spaces because we observe they are inconsistent in one embedding space. Extended experiments show that BNNeck can boost the baseline, and our baseline can improve the performance of existing state-of-the-art methods. Our codes and models are available at: https://github.com/michuanhaohao/reid-strong-baseline. Hao Luo 0004, Wei Jiang 0009, Youzhi Gu, Fuxu Liu, Xingyu Liao, Shenqi Lai, Jianyang Gu |
IEEE Trans. Multim. | 2 |
| 2019 | An End-To-End Network for Panoptic SegmentationabstractPanoptic segmentation, which needs to assign a category label to each pixel and segment each object instance simultaneously, is a challenging topic. Traditionally, the existing approaches utilize two independent models without sharing features, which makes the pipeline inefficient to implement. In addition, a heuristic method is usually employed to merge the results. However, the overlapping relationship between object instances is difficult to determine without sufficient context information during the merging process. To address the problems, we propose a novel end-to-end Occlusion Aware Network (OANet) for panoptic segmentation, which can efficiently and effectively predict both the instance and stuff segmentation in a single network. Moreover, we introduce a novel spatial ranking module to deal with the occlusion problem between the predicted instances. Extensive experiments have been done to validate the performance of our proposed method and promising results have been achieved on the COCO Panoptic benchmark. Chao Peng 0001, Changqian Yu, Jingbo Wang 0003, Xu Liu 0017, Gang Yu 0002, Wei Jiang 0009 |
CVPR | 7 |
| 2019 | SphereReID: Deep hypersphere manifold embedding for person re-identification
Wei Jiang 0009, Hao Luo 0004, Mengjuan Fei |
J. Vis. Commun. Image Represent. | 2 |
| 2019 | AlignedReID++: Dynamically matching local information for person re-identification
Hao Luo 0004, Wei Jiang 0009, Jingjing Qian, Chi Zhang 0026 |
Pattern Recognit. | 2 |
| 2018 | SCPNet: Spatial-Channel Parallelism Network for Joint Holistic and Partial Person Re-identification
Hao Luo 0004, Lingxiao He, Chi Zhang 0026, Wei Jiang 0009 |
ACCV (2) | 6 |
| 2018 | Creating memorable video summaries that satisfy the user's intention for taking the videos
Mengjuan Fei, Wei Jiang 0009 |
Neurocomputing | 2 |
| 2018 | Optical plasma boundary reconstruction based on least squares for EAST TokamakabstractReconstructing the shape and position of plasma is an important issue in Tokamaks. Equilibrium and fitting (EFIT) code is generally used for plasma boundary reconstruction in some Tokamaks. However, this magnetic method still has some inevitable disadvantages. In this paper, we present an optical plasma boundary reconstruction algorithm. This method uses EFIT reconstruction results as the standard to create the optimally optical reconstruction. Traditional edge detection methods cannot extract a clear plasma boundary for reconstruction. Based on global contrast, we propose an edge detection algorithm to extract the plasma boundary in the image plane. Illumination in this method is robust. The extracted boundary and the boundary reconstructed by EFIT are fitted by same-order polynomials and the transformation matrix exists. To acquire this matrix without camera calibration, the extracted plasma boundary is transformed from the image plane to the Tokamak poloidal plane by a mathematical model, which is optimally resolved by using least squares to minimize the error between the optically reconstructed result and the EFIT result. Once the transform matrix is acquired, we can optically reconstruct the plasma boundary with only an arbitrary image captured. The error between the method and EFIT is presented and the experimental results of different polynomial orders are discussed. Hao Luo 0004, Zhengping Luo 0002, Chao Xu 0001, Wei Jiang 0009 |
Frontiers Inf. Technol. Electron. Eng. | 4 |
| 2018 | A novel compact yet rich key frame creation method for compressed video summarization
Mengjuan Fei, Wei Jiang 0009 |
Multim. Tools Appl. | 2 |
| 2017 | A late fusion approach for harnessing multi-cnn model high-level featuresabstractThe high-level feature representation of deep convo-lutional neural networks (ConvNets) has proven to be superior to hand-crafted low-level features. Thus, this study investigates the effect of fusing such high-level features from multi-deep ConvNets under an application of visual object/scene categorization. In which, three pre-trained ConvNets are exploited as feature extractors, a single hidden layer is adopted to transform the highlevel representations to a low dimensional feature space and that are fused to harness the rich cues of the individual features. The experimental outcomes on six benchmark databases demonstrate that regardless of variation in visual statistics and tasks the fusion of multi-ConvNets' high-level features can meliorate the classification accuracy compared with a single modality, different ConvNets contain complementary cues of visual contents, and the fusion is capable of producing a very competitive performance to the state-of-the-art methods. Besides that, this fusion approach can be an effective complement to the ConvNets. Akilan Thangarajah, Q. M. Jonathan Wu, Amin Safaei 0001, Wei Jiang 0009 |
SMC | 4 |
| 2017 | Memorable and rich video summarization
Mengjuan Fei, Wei Jiang 0009 |
J. Vis. Commun. Image Represent. | 2 |
| 2016 | New fusional framework combining sparse selection and clustering for key frame extractionabstractKey frame extraction can facilitate rapid browsing and efficient video indexing in many applications. However, to be effective, key frames must preserve sufficient video content while also being compact and representative. This study proposes a syncretic key frame extraction framework that combines sparse selection (SS) and mutual information‐based agglomerative hierarchical clustering (MIAHC) to generate effective video summaries. In the proposed framework, the SS algorithm is first applied to the original video sequences to obtain optimal key frames. Then, using content‐loss minimisation and representativeness ranking, several candidate key frames are efficiently selected and grouped as initial clusters. A post‐processor – an improved MIAHC – subsequently performs further processing to eliminate redundant images and generate the final key frames. The proposed framework overcomes issues such as information redundancy and computational complexity that afflict conventional SS methods by first obtaining candidate key frames instead of accurate key frames. Subsequently, application of the improved MIAHC to these candidate key frames rather than the original video not only results in the generation of accurate key frames, but also reduces the computation time for clustering large videos. The results of comparative experiments conducted on two benchmark datasets verify that the performance of the proposed SS–MIAHC framework is superior to that of conventional methods. Mengjuan Fei, Wei Jiang 0009, Zhendong Song |
IET Comput. Vis. | 2 |
| 2016 | Extreme learning machine based transfer learning for data classification
Wei Jiang 0009 |
Neurocomputing | 3 |
| 2016 | Multiple-kernel-learning-based extreme learning machine for classification design
Wei Jiang 0009 |
Neural Comput. Appl. | 3 |
| 2014 | Dual Gaussian mixture model with pixel history for background suppressionabstractBackground suppression has found several application areas in computer vision. Among others, abandoned object detection and moving object detection under dynamic backgrounds are two major application areas. Gaussian mixture model (GMM) is one of the most popular methods and has been applied in both of these application areas. However, the aforementioned areas present contradicting challenges for GMM. Abandoned objects need the GMM to slowly get adapted to prevent accidental assimilation of the foreground, while dynamic backgrounds need quick assimilation to prevent noisy detection. A novel dual GMM based on pixel history is proposed to provide practical solutions to both of these contradictory challenges. The method uses separate GMMs to represent foreground and background modes. A foreground mode is transferred to the GMM representing the background based on its history of occurrence. Thus, a quick assimilation is offered to dynamic backgrounds while abandoned objects face a very slow assimilation. In each case, old background is preserved. Extensive experiments are done to demonstrate the effectiveness of the proposed method and to compare it with a number of well known methods in literature. Dibyendu Mukherjee, Ashirbani Saha, Q. M. Jonathan Wu, Wei Jiang 0009 |
SMC | 4 |
| 2014 | Multimodal 3D histogram for moving object detectionabstractMoving object detection is a popular research area due to its vast application in computer vision. Based on application areas, a number of moving object detection methods have been discovered. This work proposes a real-time method based on 3D histogram and temporal multiple mode selection, suited towards a vast majority of dynamic and noisy backgrounds, congested backgrounds and slow foregrounds. It is a generalization of [1] demonstrating improved performance. The temporal distribution of a video sequence can provide a history of the foreground motion and relatively static background. However, a dynamic background is identified by multiple intensity levels in the temporal distribution, and hence, cannot be properly modeled by a single histogram mode. The proposed multiple mode selection process involves identifying the dominant modes of the temporal distribution and assigning them to a multimodal background. The work provides a detailed analysis of the proposition, a clear description of the proposed algorithm and an extensive number of tests with some of the well-known methods in literature. It also demonstrates the improvement over its predecessor in several types of scenarios covered by a large number of datasets. Both qualitative and quantitative comparisons are carried out to demonstrate the applicability and accuracy of the proposed method. Dibyendu Mukherjee, Ashirbani Saha, Q. M. Jonathan Wu, Wei Jiang 0009 |
SMC | 4 |
| 2014 | Fast sparse approximation of extreme learning machine
Wei Jiang 0009 |
Neurocomputing | 3 |
| 2010 | Panoramic 3D Reconstruction Using Stereo Multi-Perspective PanoramaabstractIn this paper, we present a novel approach to imaging a panoramic (360°) environment and computing its dense depth map. Our approach adopts a multi-baseline stereo strategy using a set of multi-perspective panoramas where large baseline lengths are available. We design two image acquisition rigs for capturing such multi-perspective panoramas. The first one is composed of two parallel stereo cameras. By rotating the rig about a vertical axis, we generate four multi-perspective panoramas by resampling the regular perspective images captured by the stereo cameras. Then a depth map is estimated from the four multi-perspective panoramas and an original perspective image using a multi-baseline matching technique with different types of epipolar constraints. The second one is composed of a single camera and two mirrors. By rotating the rig, we acquire a spatio-temporal volume that is made up of the sequential images captured by the camera. Then we estimate a depth map by extracting trajectories from the spatio-temporal volume by using a multi-baseline stereo technique by considering occlusions. We can consider both rotating rigs as a single rotating camera with a very large field of view (FOV), that offers a large baseline length in depth estimation. In addition, compared with a previous approach using two multi-perspective panoramas from a single rotating camera, our approach can reduce matching errors due to image noise, repeated patterns, and occlusions by multi-baseline stereo techniques. Experimental results using both synthetic and real images show that our approach produces high quality panoramic 3D reconstruction. Wei Jiang 0009, Shigeki Sugimoto, Masatoshi Okutomi |
Int. J. Pattern Recognit. Artif. Intell. | 1 |
| 2006 | Panoramic 3D Reconstruction Using Rotational Stereo Camera with Simple Epipolar ConstraintsabstractIn this paper, we propose a novel method for panoramic 3D scene recovery using rotational stereo cameras with simple epipolar constraints. By rotating two parallel stereo cameras about a vertical axis with a constant velocity, we acquire two sampled spacio-temporal volumes which are made of the sequential images captured by a uniform angular interval. The two spacio-temporal volumes can be resampled into a set of multi-perspective panoramas. We analyze the epipolar geometry among images (panoramas and original images) of two spacio-temporal volumes. The result shows that only three types of simple epipolar constraints (epipolar line is row or column of image) exist in the two spacio-temporal volumes. Then we compute a depth map from four image pairs using a multi-baseline algorithm with the three types of epipolar constraints; that is horizontal, vertical and combination of them. Experimental results using both synthetic and real images show that our approach produces high quality panoramic 3D reconstruction. Wei Jiang 0009, Masatoshi Okutomi, Shigeki Sugimoto |
CVPR (1) | 1 |