VLDB 2026 Research / reviewers in the wild / expert
Weiwei Shi 0003
dblp:44/8271-3
· DBLP profile ↗
22ranked-venue papers
7as first author
9since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 12 · 7 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 12 · 2 first-author · 6 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Deep semi-supervised learning method based on sample adaptive weights and discriminative feature learning
Weiwei Shi 0003, Xinhong Hei 0001 |
J. Vis. Commun. Image Represent. | 2 |
| 2024 | TransBoNet: Learning camera localization with Transformer Bottleneck and Attention
Xiaogang Song 0001, Hongjuan Li, Li Liang 0008, Weiwei Shi 0003, Guo Xie, Xinhong Hei 0001 |
Pattern Recognit. | 4 |
| 2024 | Unsupervised Monocular Estimation of Depth and Visual Odometry Using Attention and Depth-Pose Consistency LossabstractRecent studies have shown that joint depth and pose estimation using convolutional neural networks (CNNs) can learn unlabelled monocular frames. However, three problems remain: 1) CNNs can only extract local features due to the limited receptive field, 2) scale ambiguity is inherent in the monocular task, and 3) illness regions violate the photometric consistency assumption and produce large errors. We propose a novel framework, ADPDepth, with corresponding effective strategies to ameliorate the above problems. First, a PCAtt module is designed to capture the correlation between channels and efficiently extract multiscale spatial information using a multibranch parallel strategy. Second, depth-pose consistency loss is proposed based on the geometric consistency in depth and pose to constrain the scale between samples, eliminate scale ambiguity and obtain a globally consistent scale. To further improve performance, a cover mask is derived from depth-pose consistency for filtering dynamic objects and outliers to reduce the adverse effects of these illness regions. Extensive experiments are conducted on the KITTI, NYU-Depth and Make3D datasets. Based on public benchmarks, the experimental results confirm that the proposed ADPDepth framework achieves state-of-the-art performance. The effectiveness of each strategy is also verified in subsequent ablation experiments. Xiaogang Song 0001, Haoyue Hu, Li Liang 0008, Weiwei Shi 0003, Guo Xie, Xinhong Hei 0001 |
IEEE Trans. Multim. | 4 |
| 2023 | Non-Exemplar Class-Incremental Learning via Adaptive Old Class ReconstructionabstractIn the Class-Incremental Learning (CIL) task, rehearsal-based approaches have received a lot of attention recently. However, storing old class samples is often infeasible in application scenarios where device memory is insufficient or data privacy is important. Therefore, it is necessary to rethink Non-Exemplar Class-Incremental Learning (NECIL). In this paper, we propose a novel NECIL method named POLO with an adaPtive Old cLass recOnstruction mechanism, in which a density-based prototype reinforcement method (DBR), a topology-correction prototype adaptation method (TPA), and an adaptive prototype augmentation method (APA) are designed to reconstruct pseudo features of old classes in new incremental sessions. Specifically, the DBR focuses on the low-density features to maintain the model's discriminative ability for old classes. Afterward, the TPA is designed to adapt old class prototypes to new feature spaces in the incremental learning process. Finally, the APA is developed to further adapt pseudo feature spaces of old classes to new feature spaces. Experimental evaluations on four benchmark datasets demonstrate the effectiveness of our proposed method over the state-of-the-art NECIL methods. Shaokun Wang, Weiwei Shi 0003, Yuhang He 0001, Yihong Gong |
ACM Multimedia | 2 |
| 2023 | Image super-resolution with multi-scale fractal residual attention network
Xiaogang Song 0001, Wanbo Liu, Li Liang 0008, Weiwei Shi 0003, Guo Xie, Xinhong Hei 0001 |
Comput. Graph. | 4 |
| 2023 | Semantic Knowledge Guided Class-Incremental LearningabstractDriven by practical needs, research on Class-Incremental Learning (CIL) has received more and more attentions in recent years. A technical challenge to be conquered by CIL methods is the catastrophic forgetting problem, where the model’s performance improves rapidly on new classes while deteriorates drastically on old ones. The main causes behind catastrophic forgetting include network drifts, inter-class confusions, etc. In this paper, we propose a novel CIL method that solves the catastrophic forgetting problem from two aspects. First, to solve the inter-class confusion problem, we propose a novel Semantic knOwledge gUided ciL framework (SOUL) that consists of a CNN feature extractor and a Bi-GCN (Graph Convolutional Network) classifier. In each CIL session, we use the semantic knowledge extracted from the class labels to build two inter-class relation graphs among all the encountered old and new classes. Using these two relation graphs, we develop a Bi-GCN classifier to fuse two kinds of semantic relations in a balanced way, and then to transfer the inter-class relations from semantic modality to image classification weights. The entire SOUL framework is trained end-to-end by the standard BP algorithm, which optimizes the Bi-GCN classifier and the CNN feature extractor jointly. Second, to prevent the network drift, we develop the local topology preserving strategy that divides the global topological structure of the learned feature space into a set of local topological relations, and maintains these local relations at CIL session. Experimental evaluations demonstrate the state-of-the-art performance accuracies on benchmark image classification datasets. Shaokun Wang, Weiwei Shi 0003, Songlin Dong, Xinyuan Gao, Xiang Song 0005, Yihong Gong |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2023 | JAMSNet: A Remote Pulse Extraction Network Based on Joint Attention and Multi-Scale FusionabstractRemote photoplethysmography (rPPG) has been an active research topic in recent years. While most existing methods are focusing on eliminating motion artifacts in the raw traces obtained from single-scale region-of-interest (ROI), it is worth noting that there are some noise signals that cannot be effectively separated in single-scale space but can be separated more easily in multi-scale space. In this paper, we analyze the distribution of pulse signal and motion artifacts in different layers of a Gaussian pyramid. We propose a method that combines multi-scale analysis and neural network for pulse extraction in different scales, and a layer-wise attention mechanism to adaptively fuse the features according to signal strength. In addition, we propose spatial-temporal joint attention module and channel-temporal joint attention module to learn and exaggerate pulse features in the joint spaces, respectively. The proposed remote pulse extraction network is called Joint Attention and Multi-Scale fusion Network (JAMSNet). Extensive experiments have been conducted on two publicly available datasets and one self-collected dataset. The results show that the proposed JAMSNet shows better performance than state-of-the-art methods. Changchen Zhao, Hongsheng Wang, Huiling Chen 0001, Weiwei Shi 0003, Yuanjing Feng |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2022 | Transductive Semisupervised Deep HashingabstractDeep hashing methods have shown their superiority to traditional ones. However, they usually require a large amount of labeled training data for achieving high retrieval accuracies. We propose a novel transductive semisupervised deep hashing (TSSDH) method which is effective to train deep convolutional neural network (DCNN) models with both labeled and unlabeled training samples. TSSDH method consists of the following four main ingredients. First, we extend the traditional transductive learning (TL) principle to make it applicable to DCNN-based deep hashing. Second, we introduce confidence levels for unlabeled samples to reduce adverse effects from uncertain samples. Third, we employ a Gaussian likelihood loss for hash code learning to sufficiently penalize large Hamming distances for similar sample pairs. Fourth, we design the large-margin feature (LMF) regularization to make the learned features satisfy that the distances of similar sample pairs are minimized and the distances of dissimilar sample pairs are larger than a predefined margin. Comprehensive experiments show that the TSSDH method can produce superior image retrieval accuracies compared to the representative semisupervised deep hashing methods under the same number of labeled training samples. Weiwei Shi 0003, Yihong Gong, Badong Chen, Xinhong Hei 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2021 | Analogy-Detail Networks for Object RecognitionabstractThe human visual system can recognize object categories accurately and efficiently and is robust to complex textures and noises. To mimic the analogy-detail dual-pathway human visual cognitive mechanism revealed in recent cognitive science studies, in this article, we propose a novel convolutional neural network (CNN) architecture named analogy-detail networks (ADNets) for accurate object recognition. ADNets disentangle the visual information and process them separately using two pathways: the analogy pathway extracts coarse and global features representing the gist (i.e., shape and topology) of the object, while the detail pathway extracts fine and local features representing the details (i.e., texture and edges) for determining object categories. We modularize the architecture and encapsulate the two pathways into the analogy-detail block as the CNN building block to construct ADNets. For implementation, we propose a general principle that transmutes typical CNN structures into the ADNet architecture and applies the transmutation on representative baseline CNNs. Extensive experiments on CIFAR10, CIFAR100, street view house numbers, and ImageNet data sets demonstrate that ADNets significantly reduce the test error rates of the baseline CNNs by up to 5.76% and outperform other state-of-the-art architectures. Comprehensive analysis and visualizations further demonstrate that ADNets are interpretable and have a better shape-texture tradeoff for recognizing the objects with complex textures. Xiaopeng Hong, Weiwei Shi 0003, Xinyuan Chang, Yihong Gong |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2020 | Object detection with class aware region proposal network and focused attention objective
Yihong Gong, Weiwei Shi 0003, De Cheng |
Pattern Recognit. Lett. | 3 |
| 2020 | Multi-Target Multi-Camera Tracking by Tracklet-to-Target AssignmentabstractThis paper focuses on the Multi-Target Multi-Camera Tracking task (MTMCT), which aims at tracking multiple targets within a multi-camera network. As the trajectory of each target is inherently split into multiple sub-trajectories (namely local tracklets) in a multi-camera network, a major challenge of MTMCT is how to accurately match the local tracklets generated within each camera across different cameras and generate a complete global trajectory for each target, i.e., the cross-camera tracklet matching problem. We solve the cross-camera tracklet matching problem by TRACklet-to-Target Assignment (TRACTA), and propose the Restricted Non-negative Matrix Factorization (RNMF) algorithm to compute the optimal assignment solution that meets a set of constraints, which should be in force in practice. TRACTA can correct the tracking errors caused by occlusions and missed detections in local tracklets, and produce a complete global trajectory for each target across all the cameras. Moreover, we also develop an analytical way of estimating the total number of targets in the camera network, which plays an important role to compute the tracklet-to-target assignment. Experimental evaluations and ablation studies on four MTMCT benchmark datasets show the superiority of the proposed TRACTA method. Yuhang He 0001, Xing Wei 0001, Xiaopeng Hong, Weiwei Shi 0003, Yihong Gong |
IEEE Trans. Image Process. | 4 |
| 2019 | Consistency-Preserving deep hashing for fast person re-identification
Diangang Li, Yihong Gong, De Cheng, Weiwei Shi 0003, Xinyuan Chang |
Pattern Recognit. | 4 |
| 2019 | Normalized Non-Negative Sparse Encoder for Fast Image RepresentationabstractImage representation based on sparse coding generalizes the bag of words model. Although it reduces the reconstruction error for local features to achieve the state-of-the-art image classification performance, the large computational cost hinders the application of sparse coding-based image features. In this paper, we propose approximating a sparse code using the output of a simple neural network. The resulting parameter learning model for the neural network automatically incorporates non-negative and shift-invariant constraints, leading to an efficient normalized non-negative sparse coding (N3SC) sparse encoder. Without the use of the traditional iterative process to solve the sparse coding objective, the sparse encoder directly “converts” each local feature into a sparse code. We also introduce a method for training the encoder based on the auto-encoder method. In addition, we formally propose the corresponding sparse coding scheme called N3SC, which enforces both the non-negative constraint and the shift-invariant constraint in addition to the traditional sparse coding criteria. As demonstrated by several experiments, the obtained N3SC encoder requires only 3%-10% of the processing time for image feature extraction compared with the standard sparse coding scheme. At the same time, the features extracted using the exact solutions of the N3SC coding scheme and the N3SC encoder offer superior image classification accuracy compared to the accuracy of many existing sparse coding-based representations. Shizhou Zhang, Jinjun Wang, Weiwei Shi 0003, Yihong Gong, Yong Xia 0001, Yanning Zhang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2019 | Fine-Grained Image Classification Using Modified DCNNs Trained by Cascaded Softmax and Generalized Large-Margin LossesabstractWe develop a fine-grained image classifier using a general deep convolutional neural network (DCNN). We improve the fine-grained image classification accuracy of a DCNN model from the following two aspects. First, to better model the h -level hierarchical label structure of the fine-grained image classes contained in the given training data set, we introduce h fully connected (fc) layers to replace the top fc layer of a given DCNN model and train them with the cascaded softmax loss. Second, we propose a novel loss function, namely, generalized large-margin (GLM) loss, to make the given DCNN model explicitly explore the hierarchical label structure and the similarity regularities of the fine-grained image classes. The GLM loss explicitly not only reduces between-class similarity and within-class variance of the learned features by DCNN models but also makes the subclasses belonging to the same coarse class be more similar to each other than those belonging to different coarse classes in the feature space. Moreover, the proposed fine-grained image classification framework is independent and can be applied to any DCNN structures. Comprehensive experimental evaluations of several general DCNN models (AlexNet, GoogLeNet, and VGG) using three benchmark data sets (Stanford car, fine-grained visual classification-aircraft, and CUB-200-2011) for the fine-grained image classification task demonstrate the effectiveness of our method. Weiwei Shi 0003, Yihong Gong, De Cheng, Nanning Zheng 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2018 | Transductive Semi-Supervised Deep Learning Using Min-Max Features
Weiwei Shi 0003, Yihong Gong, Chris Ding, Zhiheng Ma, Nanning Zheng 0001 |
ECCV (5) | 1 |
| 2018 | Person re-identification by the asymmetric triplet and identification loss function
De Cheng, Yihong Gong, Weiwei Shi 0003, Shizhou Zhang |
Multim. Tools Appl. | 3 |
| 2018 | Correction to: Person re-identification by the symmetric triplet and identification loss function
De Cheng, Yihong Gong, Weiwei Shi 0003, Shizhou Zhang |
Multim. Tools Appl. | 3 |
| 2018 | Deep feature learning via structured graph Laplacian embedding for person re-identification
De Cheng, Yihong Gong, Xiaojun Chang, Weiwei Shi 0003, Alex Hauptmann 0001, Nanning Zheng 0001 |
Pattern Recognit. | 4 |
| 2018 | Entropy and orthogonality based deep discriminative feature learning for object recognition
Weiwei Shi 0003, Yihong Gong, De Cheng, Nanning Zheng 0001 |
Pattern Recognit. | 1 |
| 2018 | Improving CNN Performance Accuracies With Min-Max ObjectiveabstractWe propose a novel method for improving performance accuracies of convolutional neural network (CNN) without the need to increase the network complexity. We accomplish the goal by applying the proposed Min-Max objective to a layer below the output layer of a CNN model in the course of training. The Min-Max objective explicitly ensures that the feature maps learned by a CNN model have the minimum within-manifold distance for each object manifold and the maximum between-manifold distances among different object manifolds. The Min-Max objective is general and able to be applied to different CNNs with insignificant increases in computation cost. Moreover, an incremental minibatch training procedure is also proposed in conjunction with the Min-Max objective to enable the handling of large-scale training data. Comprehensive experimental evaluations on several benchmark data sets with both the image classification and face verification tasks reveal that employing the proposed Min-Max objective in the training process can remarkably improve performance accuracies of a CNN model in comparison with the same model trained without using this objective. Weiwei Shi 0003, Yihong Gong, Jinjun Wang, Nanning Zheng 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2018 | Training DCNN by Combining Max-Margin, Max-Correlation Objectives, and Correntropy Loss for Multilabel Image ClassificationabstractIn this paper, we build a multilabel image classifier using a general deep convolutional neural network (DCNN). We propose a novel objective function that consists of three parts, i.e., max-margin objective, max-correlation objective, and correntropy loss. The max-margin objective explicitly enforces that the minimum score of positive labels must be larger than the maximum score of negative labels by a predefined margin, which not only improves accuracies of the multilabel classifier, but also eases the threshold determination. The max-correlation objective can make the DCNN model learn a latent semantic space, which maximizes the correlations between the feature vectors of the training samples and their corresponding ground-truth label vectors projected into this space. Instead of using the traditional softmax loss, we adopt the correntropy loss from the information theory field to minimize the training errors of the DCNN model. The proposed framework can be end-to-end trained. Comprehensive experimental evaluations on Pascal VOC 2007 and MIR Flickr 25K multilabel benchmark data sets with four DCNN models, i.e., AlexNet, VGG-16, GoogLeNet, and ResNet demonstrate that the proposed objective function can remarkably improve the performance accuracies of a DCNN model for the task of multilabel image classification. Weiwei Shi 0003, Yihong Gong, Nanning Zheng 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2016 | Improving CNN Performance with Min-Max Objective
Weiwei Shi 0003, Yihong Gong, Jinjun Wang |
IJCAI | 1 |