EDBT 2026 Demo / reviewers in the wild / expert
Shenqi Lai
dblp:238/0159
· DBLP profile ↗
24ranked-venue papers
4as first author
22since 2021 · last 2026
0000-0001-8673-2218ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 16 · 4 first-author · 14 since 2021Artificial intelligence and machine learning · 12 · 2 first-author · 11 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Occluded person Re-Identification with noise injection
Can Yao, Xi Du, Deng Cai 0001, Shenqi Lai |
Pattern Recognit. | 6 |
| 2026 | 3 × 3 Kernel Is All You Need for VisionabstractMost modern Convolutional Neural Networks (CNNs) employ a multi-branch structure with various-sized convolutions to capture long- and short-range dependencies. However, these CNNs use large kernel convolutions (e.g., astonishingly 101 kernels) and specialized techniques (e.g., reparameterization and sparsity), increasing complexity in both training and inference stages. This paper focuses on designing an efficient CNN based on pure 3×3 convolutions without introducing complex operations and techniques. Specifically, we propose a Spatial Pyramid (SP) block, which consists of the Multi-branch Residual (MbR) module and the Gated-branch Residual (GbR) module. The MbR introduces multiscale pooling as the key component, thus capturing long-range visual cues through large down-sampling rates and shorter-range dependencies through low down-sampling rates while maintaining low computational complexity. Besides, the GbR uses one 3×3 convolution to refine dependencies along spatial and channel dimensions. Based on the SP block, we construct the Spatial Pyramid CNN (SPCNN), a model composed exclusively of Point-Wise Convolution and 3×3 Depth-Wise Convolution. Under comparable computational complexity, SPCNN significantly outperforms the state-of-the-art CNN PeLK (83.6% vs 82.6%) with only 3 × 3 kernels (compared to 101 × 101 kernels in PeLK). Besides, our SPCNN demonstrates comparability with state-of-the-art backbones in lightweight models, object detection, instance segmentation, and semantic segmentation. Moreover, evaluations of four image retrieval benchmarks also demonstrate the effectiveness. All codes are released at https://github.com/xiaolai-sqlai/SPCNN. Shenqi Lai, Mengjian Li, Haifeng Liu 0001, Xueming Qian, Deng Cai 0001, Yaxiong Wang |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2026 | LEViT: Locally Enhanced Vision Transformer for Efficient Object Re-IdentificationabstractVision Transformer (ViT) on object re-identification (ReID) has attracted significant attention recently. However, ViT-based ReID substantially increases computational complexity, imposing significant burdens during training and inference. This paper presents an efficient and effective ViT-based backbone for ReID tasks, called the Locally Enhanced Vision Transformer (LEViT). ViT models typically emphasize global relationship modeling, yet ReID tasks are more sensitive to local information. To address this gap, we propose a Locally Enhanced (LE) block to enhance local information by performing self-attention within local split windows. Since part-based models dominate ReID, calculating self-attention across all patches is computationally inefficient. We also replace the traditional Query-Key-Value projector with the Group Convolution (G-Conv) projector, enabling the model to capture local details more efficiently. Furthermore, G-Conv is integrated into the channel MLP to strengthen local feature sensitivity. Using these components, we develop two LEViT variants: LEViT-S and LEViT-L. To our knowledge, LEViT is the first highly adaptable ViT backbone for ReID tasks. Experimental evaluations demonstrate the effectiveness in five ReID datasets: Market1501, DukeMTMC, MSMT17, VeRi-776, and VehicleID. Notably, LEViT-S outperforms TransReID while requiring less than 10% computational complexity. Furthermore, LEViT obtains the state-of-the-art on three deep metric learning datasets: CUB-200-2011, Cars196, and University-1652. Our code will be available athttps://github.com/YuhuiWang99/LEViT. Shenqi Lai, Mingyuan Fan 0002, Junshi Huang, Haifeng Liu 0001, Deng Cai 0001, Xueming Qian, Yaxiong Wang |
IEEE Trans. Multim. | 1 |
| 2025 | A Pyramid Fusion MLP for Dense PredictionabstractRecently, MLP-based architectures have achieved competitive performance with convolutional neural networks (CNNs) and vision transformers (ViTs) across various vision tasks. However, most MLP-based methods introduce local feature interactions to facilitate direct adaptation to downstream tasks, thereby lacking the ability to capture global visual dependencies and multi-scale context, ultimately resulting in unsatisfactory performance on dense prediction. This paper proposes a competitive and effective MLP-based architecture called Pyramid Fusion MLP (PFMLP) to address the above limitation. Specifically, each block in PFMLP introduces multi-scale pooling and fully connected layers to generate feature pyramids, which are subsequently fused using up-sample layers and an additional fully connected layer. Employing different down-sample rates allows us to obtain diverse receptive fields, enabling the model to simultaneously capture long-range dependencies and fine-grained cues, thereby exploiting the potential of global context information and enhancing the spatial representation power of the model. Our PFMLP is the first lightweight MLP to obtain comparable results with state-of-the-art CNNs and ViTs on the ImageNet-1K benchmark.With larger FLOPs, it exceeds state-of-the-art CNNs, ViTs, and MLPs under similar computational complexity. Furthermore, experiments in object detection, instance segmentation, and semantic segmentation demonstrate that the visual representation acquired from PFMLP can be seamlessly transferred to downstream tasks, producing competitive results. All materials contain the training codes and logs are released at https://github.com/huangqiuyu/PFMLP. Qiuyu Huang, Zequn Jie, Lin Ma 0002, Li Shen 0008, Shenqi Lai |
IEEE Trans. Image Process. | 5 |
| 2024 | Open-Vocabulary Animal Keypoint Detection with Semantic-Feature Matching
Hao Zhang 0117, Lumin Xu, Shenqi Lai, Wenqi Shao, Nanning Zheng 0001, Ping Luo 0002, Yu Qiao 0001, Kaipeng Zhang |
Int. J. Comput. Vis. | 3 |
| 2024 | FMGNet: An efficient feature-multiplex group network for real-time vision task
Hao Zhang 0117, Kaipeng Zhang, Nanning Zheng 0001, Shenqi Lai |
Pattern Recognit. | 5 |
| 2024 | HF-HRNet: A Simple Hardware Friendly High-Resolution NetworkabstractHigh-resolution networks have made significant progress in dense prediction tasks such as human pose estimation and semantic segmentation. To better explore this high-resolution mechanism on mobile devices, Lite-HRNet incorporates shuffle operations to reduce computational complexity in the channel dimension, while Dite-HRNet employs dynamic convolution and pooling to capture long-range interactions with low computational complexity in the spatial dimension. The core idea behind both approaches is to efficiently capture information in either the channel or spatial dimension. However, shuffle operations and dynamic operations are not hardware-friendly. As a result, both Lite-HRNet and Dite-HRNet cannot achieve the desired inference speed on specialized devices, including Neural Processing Units (NPUs) and Graphics Processing Units (GPUs). To overcome these limitations, we present a simple Hardware-Friendly Lightweight High-resolution Network (HF-HRNet) based on our proposed Hardware-Friendly Uniform-sized Mug (HUM) block. HUM block mainly consists of the Cascaded Depthwise (CAD) block and Multi-Scale Context Embedding (MCE) block. The CAD block cascades depthwise convolutions to obtain a larger receptive field in the spatial dimension, while the MCE block aggregates multi-scale spatial feature information from different scales and adjusts channel features. Extensive experiments are conducted on human pose estimation (COCO, MPII) and semantic segmentation (Cityscapes), resulting in a better trade-off between inference speed and accuracy on both NPUs and GPUs. It is noteworthy that on the COCO test-dev set, HF-HRNet-30 outperforms Dite-HRNet-30 and Lite-HRNet-30 by 1.9 AP and 2.8 AP, respectively, while running about 13 times faster and 9 times faster on NPUs, respectively. Our code are publicly available for use: https://github.com/zhanghao5201/HF-HRNet. Hao Zhang 0117, Yujie Dun, Yixuan Pei, Shenqi Lai, Chengxu Liu 0001, Kaipeng Zhang, Xueming Qian |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2024 | Diversity-Learning Block: Conquer Feature Homogenization of MultibranchabstractVisual Geometry Group (VGG)-style ConvNet is an neural-network process units (NPU)-friendly network; however, the accuracy of this architecture cannot keep up with other well-designed network structures. Although some reparameterization methods are proposed to remedy this weakness, their performance suffers from the homogenization issue of parallel branches, and the preset shape of convolution kernels also influences spatial perception. To address this problem, we propose a diversity-learning (DL) block to build the DLNet, which could adaptively learn various features to enrich the feature space. To balance floating point of operations (FLOPs) and accuracy, groupwise operation is introduced and finally, a lightweight DL ConvNet DLGNet is obtained. Extensive evaluations have been conducted on different computer vision tasks, e.g., image classification [Canadian Institute For Advanced Research (CIFAR) and ImageNet], object detection [PASCAL visual object classes (VOC) and Microsoft Common Objects in Context (MS COCO)], and semantic segmentation (Cityscapes). The experimental results show that our proposed DLGNet can achieve comparable performance with the state-of-the-art networks while the speed is 183% faster than GhostNet and even over 600% faster than MobileNetV3 with similar accuracy when running on NPU. Junjie Yang 0008, Shenqi Lai, Xuan Wang 0018, Yaxiong Wang, Xueming Qian |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2023 | RaMLP: Vision MLP via Region-aware MixingabstractRecently, MLP-based architectures achieved impressive results in image classification against CNNs and ViTs. However, there is an obvious limitation in that their parameters are related to image sizes, allowing them to process only fixed image sizes. Therefore, they cannot directly adapt dense prediction tasks (e.g., object detection and semantic segmentation) where images are of various sizes. Recent methods tried to address it but brought two new problems, long-range dependencies or important visual cues are ignored. This paper presents a new MLP-based architecture, Region-aware MLP (RaMLP), to satisfy various vision tasks and address the above three problems. In particular, we propose a well-designed module, Region-aware Mixing (RaM). RaM captures important local information and further aggregates these important visual clues. Based on RaM, RaMLP achieves a global receptive field even in one block. It is worth noting that, unlike most existing MLP-based architectures that adopt the same spatial weights to all samples, RaM is region-aware and adaptively determines weights to extract region-level features better. Impressively, our RaMLP outperforms state-of-the-art ViTs, CNNs, and MLPs on both ImageNet-1K image classification and downstream dense prediction tasks, including MS-COCO object detection, MS-COCO instance segmentation, and ADE20K semantic segmentation. In particular, RaMLP outperforms MLPs by a large margin (around 1.5% Apb or 1.0% mIoU) on dense prediction tasks. The training code could be found at https://github.com/xiaolai-sqlai/RaMLP. Shenqi Lai, Xi Du, Kaipeng Zhang |
IJCAI | 1 |
| 2023 | Task-Adaptive Feature Disentanglement and Hallucination for Few-Shot ClassificationabstractFew-shot classification is a challenging task of computer vision and is critical to the data-sparse scenario like rare disease diagnosis. Feature augmentation is a straightforward way to alleviate the data-sparse issue in few-shot classification. However, mimicking the original feature distribution from a small amount of data is challenging. Existing augmentation-based methods are task-agnostic: the augmented feature is not with optimal intra-class diversity and inter-class discriminability concerning a certain task. To address this drawback, we propose a novel Task-adaptive Feature Disentanglement and Hallucination framework, dubbed TaFDH. Concretely, we first perceive the task information to disentangle the original feature into two components: class-irrelevant and class-specific features. Then more class-irrelevant features are decoded from a learned variational distribution, fused with the class-specific feature to get the augmented features. Finally, a generalized prior distribution over a quadratic classifier is meta-learned, which can be fast adapted to the class-specific posterior, thus further alleviating the inadequacy and uncertainty of feature hallucination via the nature of Bayesian inference. In this way, we construct a more discriminable embedding space with reasonable intra-class diversity instead of simply restoring the original embedding space, which can lead to a more precise decision boundary. We obtain the augmented features equipped with enhanced inter-class discriminability by highlighting the most discriminable part while boosting the intra-class diversity by fusing with the diverse generated class-irrelevant parts. Experiments on five multi-grained few-shot classification datasets demonstrate the superiority of our method. Li Shen 0008, Shenqi Lai, Chun Yuan 0003 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2023 | SCGNet: Shifting and Cascaded Group NetworkabstractMany lightweight networks have been proposed for resource-limited applications, however, they cannot be efficiently applied to neural-network processing units (NPUs) due to the limited operations supported by the NPUs, and few works focus on efficient network design on the NPUs. The basic blocks of networks such as MobileNetV2 and RegNet use smaller convolution kernels with relatively small receptive fields, which are not conducive to capturing large-scale spatial information. To address this weakness, we propose Shifting and Cascaded Group (SCG) block, where we cascade group convolutions with larger kernels to exploit multi-scale information and propose shifting group convolution to communicate channel information between different groups. Besides, we carefully devise our architecture guided by some principles and finally build a very efficient network called Shifting and Cascaded Group Network (SCGNet) on NPUs. To verify the superiority of our method, we conduct extensive experiments on various tasks including image classification, object detection, human pose estimation, person re-identification, and semantic segmentation to comprehensively evaluate the performance. Results on widely used datasets such as ImageNet, PASCAL VOC, COCO, MPII, Market-1501, DukeMTMC-ReID, CUHK03, and Cityscapes demonstrate that the proposed network is a more effective network on the corresponding vision tasks. Hao Zhang 0117, Shenqi Lai, Yaxiong Wang, Zongyang Da, Yujie Dun, Xueming Qian |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2023 | MPPM: A Mobile-Efficient Part Model for Object re-IDabstractObject re-identification (re-ID) is one of the core technologies in Multi-Object Tracking (MOT) that requires real-time decision-making. A Neural Processing Unit (NPU) is a low-power device that is dedicated to deploying neural network-based algorithms and has become one of the most important devices in today's mobile onboard systems. However, the current mainstream re-ID methods rarely consider the NPU characteristics, which makes it difficult for these methods to achieve both high onboard frame rates and high accuracies on an NPU. To address this problem, this paper focuses on designing a re-ID algorithm suitable for NPU deployment. The model of the object re-ID can be divided into two parts: theencoder(backbone) and thedecoder. In this article, a Mobile-efficient Pure Part Model (MPPM) is presented for re-ID task. First, for thebackboneof re-ID, we propose an efficient structure GogglesNet, which is composed of traditional convolutions. GogglesNet performs well on the re-ID task and can be comparable to lightweight networks on ImageNet with regard to accuracy and is faster on NPU. We then revisit the architectures of Pure Part Model (PPM) in person re-ID, including PCB and MGN, and propose a mobile-efficientdecoderDual Pattern Network (DPN) for re-ID. The proposed MPPM achieves comparable performance with MGN on five re-ID datasets Market-1501, DukeMTMC-reID, MSMT17, VeRi-776, and VehicleID, while the proposed parameter amount is only 10.2% of it, and the speed on NPU is more than eight times higher. Hanyang Jin, Shenqi Lai, Xueming Qian |
IEEE Trans. Multim. | 2 |
| 2022 | Scale adaption-guided human face detection
Cunying Ye, Xin Li 0134, Shenqi Lai, Yaxiong Wang, Xueming Qian |
Knowl. Based Syst. | 3 |
| 2022 | Dimension-aware attention for efficient mobile networks
Rongyun Mo, Shenqi Lai, Yan Yan 0001, Zhenhua Chai, Xiaolin Wei |
Pattern Recognit. | 2 |
| 2022 | Occlusion-Sensitive Person Re-Identification via Attribute-Based Shift AttentionabstractOccluded person re-identification is one of the most challenging tasks in security surveillance. Most existing methods focus on extracting human body features from occluded pedestrian images. This paper prioritizes a difference between occluded and non-occluded person re-ID: When computing the similarity between a holistic pedestrian image and an occluded pedestrian image, a certain part of the human body in this holistic image can be distractive for pedestrian retrieval. To solve this problem, we propose an occluded person re-ID framework named attribute-based shift attention network (ASAN). First, unlike other methods that use off-the-shelf tools to locate pedestrian body parts in the occluded images, we design an attribute-guided occlusion-sensitive pedestrian segmentation (AOPS) module. AOPS is a weakly supervised method that leverages the semantic-level attribute annotations in person re-ID datasets. Second, guided by the pedestrian masks provided by AOPS, a shift feature adaption (SFA) module extracts the visible part of the human body feature in a part-based manner. After that, a visible region matching (VRM) algorithm is proposed to filter out the interfer-ence information in the holistic person images during the retrieval phase and further purify the representation of pedestrian features. Extensive experiments with ablation analysis demonstrate our method’s effectiveness. And the state-of-the-art results are achieved on four occluded datasets Partial-REID, Partial-iLIDS, Occluded-DukeMTMC, and Occluded REID. Moreover, the experiments on two holistic person re-ID datasets Market-1501 and DukeMTMC-reID, and a vehicle re-ID dataset VeRi-776 show that ASAN also has a good generality. Hanyang Jin, Shenqi Lai, Xueming Qian |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2022 | DBCFace: Towards Pure Convolutional Neural Network Face DetectionabstractFace detection generally requires prior boxes and an extra non-maximum suppression(NMS) post-processing in modern deep learning methods. However, anchor design and anchor matching strategy significantly affect the performance of face detectors, so we have to spend a lot of time on anchor designing for different business scenarios. The other issue is that NMS cannot be easily parallelized and it may become a bottleneck of detection speed. In this paper, we propose a simple yet efficient pure convolutional neural network face detection method, named dual-branch center face detector(DBCFace for short), which solve face detection via a dual branch fully convolutional framework without extra anchor design and NMS. Extensive experiments are conducted on four popular face detection benchmarks, including AFW, PASCAL face, FDDB, and WIDER FACE, demonstrating that our method is comparable with state-of-the-art methods while the speed is faster. Xin Li 0134, Shenqi Lai, Xueming Qian |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2022 | SPGNet: Serial and Parallel Group NetworkabstractNeural-network Processing Units (NPU), which specializes in the acceleration of deep neural networks (DNN), is of great significance to latency-sensitive areas like robotics or edge computing. However, there are few works focusing on the network design for NPU in recent studies. Most of the popular lightweight structures (e.g. MobileNet) are designed with depthwise convolution, which has less computation in theory but is not friendly to existing hardwares, and the speed tested on NPU is not always satisfactory. Even under similar FLOPs (the number of multiply-accumulates), vanilla convolution operation is always faster than depthwise one. In this paper, we will propose a novel architecture named Serial and Parallel Group Network (SPGNet), which can capture discriminative multi-scale information and at the same time keep the structure compact. Extensive evaluations have been conducted on different computer vision tasks, e.g. image classification (CIFAR and ImageNet), object detection (PASCAL VOC and MS COCO) and person re-identification (Market-1501 and DukeMTMC-ReID). The experimental results show that our proposed SPGNet can achieve comparable performance with the state-of-the-art networks while the speed is 120% faster than MobileNetV2 under similar FLOPS and over 300% faster than GhostNet with similar accuracy on NPU. Xuan Wang 0018, Shenqi Lai, Zhenhua Chai, Xingjun Zhang, Xueming Qian |
IEEE Trans. Multim. | 2 |
| 2021 | Rethinking BiSeNet for Real-Time Semantic SegmentationabstractBiSeNet [28], [27] has been proved to be a popular two-stream network for real-time segmentation. However, its principle of adding an extra path to encode spatial information is time-consuming, and the backbones borrowed from pretrained tasks, e.g., image classification, may be inefficient for image segmentation due to the deficiency of task-specific design. To handle these problems, we propose a novel and efficient structure named Short-Term Dense Concatenate network (STDC network) by removing structure redundancy. Specifically, we gradually reduce the dimension of feature maps and use the aggregation of them for image representation, which forms the basic module of STDC network. In the decoder, we propose a Detail Aggregation module by integrating the learning of spatial information into low-level layers in single-stream manner. Finally, the low-level features and deep features are fused to predict the final segmentation results. Extensive experiments on Cityscapes and CamVid dataset demonstrate the effectiveness of our method by achieving promising trade-off between segmentation accuracy and inference speed. On Cityscapes, we achieve 71.9% mIoU on the test set with a speed of 250.4 FPS on NVIDIA GTX 1080Ti, which is 45.2% faster than the latest methods, and achieve 76.8% mIoU with 97.0 FPS while inferring on higher resolution images. Code is available at https://github.com/MichaelFan01/STDC-Seg. Mingyuan Fan 0002, Shenqi Lai, Junshi Huang, Xiaoming Wei, Zhenhua Chai, Junfeng Luo, Xiaolin Wei |
CVPR | 2 |
| 2021 | Feature Decomposition and Reconstruction Learning for Effective Facial Expression RecognitionabstractIn this paper, we propose a novel Feature Decomposition and Reconstruction Learning (FDRL) method for effective facial expression recognition. We view the expression information as the combination of the shared information (expression similarities) across different expressions and the unique information (expression-specific variations) for each expression. More specifically, FDRL mainly consists of two crucial networks: a Feature Decomposition Network (FDN) and a Feature Reconstruction Network (FRN). In particular, FDN first decomposes the basic features extracted from a backbone network into a set of facial action-aware latent features to model expression similarities. Then, FRN captures the intra-feature and inter-feature relationships for la-tent features to characterize expression-specific variations, and reconstructs the expression feature. To this end, two modules including an intra-feature relation modeling module and an inter-feature relation modeling module are developed in FRN. Experimental results on both the in-the-lab databases (including CK+, MMI, and Oulu-CASIA) and the in-the-wild databases (including RAF-DB and SFEW) show that the proposed FDRL method consistently achieves higher recognition accuracy than several state-of-the-art methods. This clearly highlights the benefit of feature decomposition and reconstruction for classifying expressions. Delian Ruan, Yan Yan 0001, Shenqi Lai, Zhenhua Chai, Chunhua Shen, Hanzi Wang |
CVPR | 3 |
| 2021 | Hashing person re-ID with self-distilling smooth relaxation
Hanyang Jin, Shenqi Lai, Guoshuai Zhao 0001, Xueming Qian |
Neurocomputing | 2 |
| 2021 | Preparing lessons: Improve knowledge distillation with better supervision
Tiancheng Wen, Shenqi Lai, Xueming Qian |
Neurocomputing | 2 |
| 2021 | Deep Transfer Hashing for Image RetrievalabstractDeep supervised hashing has emerged as an influential solution to large-scale semantic image retrieval problems in computer vision. In the light of recent progress, image label is the common way to define whether two images belong to the same category, but it contains little supervised information. The one-hot label can't accurately define the similarity of two images, which is important for image retrieval. In this paper, we propose an effective method, Deep Transfer Hashing(DTH) which uses the knowledge from teacher model as the supervised information. Inspired by knowledge distillation for model compression and deep hashing for fast image retrieval, we transfer the knowledge from a complex convolutional neural network(teacher) to a small neural network(student) which is used for fast image retrieval. The distance of the knowledge from teacher model can indicate the similarity of images. By minimizing the hashing codes distribution between the hashing layers of teacher model and student model, we can improve the retrieval performance. And we also evaluate the performance of the compressed model at inference stage. We test our method on widely used datasets CIFAR-10 and NUS-WIDE and we compare our method with other state-of-the-art methods in image retrieval domain. The experimental results show that our method can improve the image retrieval baseline by a large margin and better than other methods. Hongjia Zhai, Shenqi Lai, Hanyang Jin, Xueming Qian, Tao Mei 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2020 | A Strong Baseline and Batch Normalization Neck for Deep Person Re-IdentificationabstractThis study proposes a simple but strong baseline for deep person re-identification (ReID). Deep person ReID has achieved great progress and high performance in recent years. However, many state-of-the-art methods design complex network structures and concatenate multi-branch features. In the literature, some effective training tricks briefly appear in several papers or source codes. The present study collects and evaluates these effective training tricks in person ReID. By combining these tricks, the model achieves 94.5% rank-1 and 85.9% mean average precision on Market1501 with only using the global features of ResNet50. The performance surpasses all existing global- and part-based baselines in person ReID. We propose a novel neck structure named as batch normalization neck (BNNeck). BNNeck adds a batch normalization layer after global pooling layer to separate metric and classification losses into two different feature spaces because we observe they are inconsistent in one embedding space. Extended experiments show that BNNeck can boost the baseline, and our baseline can improve the performance of existing state-of-the-art methods. Our codes and models are available at: https://github.com/michuanhaohao/reid-strong-baseline. Hao Luo 0004, Wei Jiang 0009, Youzhi Gu, Fuxu Liu, Xingyu Liao, Shenqi Lai, Jianyang Gu |
IEEE Trans. Multim. | 6 |
| 2019 | Enhanced Normalized Mean Error loss for Robust Facial Landmark detection
Shenqi Lai, Zhenhua Chai, Shengxi Li, Huanhuan Meng, Mengzhao Yang, Xiaoming Wei |
BMVC | 1 |