VLDB 2026 Research / reviewers in the wild / expert
Fei Wang 0032
dblp:52/3194-32
· DBLP profile ↗
64ranked-venue papers
6as first author
51since 2021 · last 2026
0000-0002-1024-5867ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 54 · 4 first-author · 42 since 2021Graphics, computer vision, multimedia, augmented reality and games · 34 · 3 first-author · 25 since 2021Databases, data management, data science and information retrieval · 4 · 4 since 2021Human-computer interaction and ubiquitous computing · 3 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021Computer networks · 1 · 1 first-author · 1 since 2021Security and privacy · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | AutoAWG: Adverse Weather Generation with Adaptive Multi-Controls for Automotive VideosabstractPerception robustness under adverse weather remains a critical challenge for autonomous driving, with the core bottleneck being the scarcity of real-world video data in adverse weather. Existing weather generation approaches struggle to balance visual quality and annotation reusability. We present AutoAWG, a controllable Adverse Weather video Generation framework for Autonomous driving. Our method employs a semantics-guided adaptive fusion of multiple controls to balance strong weather stylization with high-fidelity preservation of safety-critical targets; leverages a vanishing point-anchored temporal synthesis strategy to construct training sequences from static images, thereby reducing reliance on synthetic data; and adopts masked training to enhance long-horizon generation stability. On the nuScenes validation set, AutoAWG significantly outperforms prior state-of-the-art methods: without first-frame conditioning, FID and FVD are relatively reduced by 50.0% and 16.1%; with first-frame conditioning, they are further reduced by 8.7% and 7.2%, respectively. Extensive qualitative and quantitative results demonstrate advantages in style fidelity, temporal consistency, and semantic–structural integrity, underscoring the practical value of AutoAWG for improving downstream perception in autonomous driving. Our code is available at: https://github.com/higherhu/AutoAWG Jiagao Hu, Daiguo Zhou, Danzhen Fu, Fuhao Li, Fei Wang 0032, Wenhua Liao, Jiayi Xie |
ICMR | 6 |
| 2026 | WAMNet: Wavelet-enhanced asymmetric mamba network for semantic segmentation of multimodal remote sensing images
Fei Wang 0032, Yanhong Yang, Haozheng Zhang, Chengkun Li, Yushan Xue, Shengyong Chen |
Neurocomputing | 1 |
| 2026 | Segmentation guided edge enhanced teacher-student for industrial anomaly detection
Yanhong Yang, Haozheng Zhang, Fei Wang 0032, Shengyong Chen |
Neurocomputing | 4 |
| 2026 | Corrigendum to "FaceChain-MMID: Generating highly identity-consistent realistic portraits via dividing & merging multi-modal representations" [Pattern Recognition 168 (2025) 111858]
Chao Xu 0023, Fei Wang 0032, Baigui Sun, Jian Zhao 0006 |
Pattern Recognit. | 2 |
| 2025 | Layer as Puzzle Pieces: Compressing Large Language Models through Layer ConcatenationabstractLarge Language Models (LLMs) excel at natural language processing tasks, but their massive size leads to high computational and storage demands.
Recent works have sought to reduce their model size through layer-wise structured pruning.
However, they tend to ignore retaining the capabilities in the pruned part.
In this work, we re-examine structured pruning paradigms and uncover several key limitations: 1) notable performance degradation due to direct layer removal, 2) incompetent linear weighted layer aggregation, and 3) the lack of effective post-training recovery mechanisms.
To address these limitations, we propose CoMe, including a progressive layer pruning framework with a Concatenation-based Merging technology and a hierarchical distillation post-training process.
Specifically, we introduce a channel sensitivity metric that utilizes activation intensity and weight norms for fine-grained channel selection.
Subsequently, we employ a concatenation-based layer merging method to fuse the most critical channels in the adjacent layers, enabling a progressive model size reduction.
Finally, we propose a hierarchical distillation protocol, which leverages the correspondences between the original and pruned model layers established during pruning, enabling efficient knowledge transfer.
Experiments on seven benchmarks show that CoMe achieves state-of-the-art performance; when pruning 30% of LLaMA-2-7b's parameters, the pruned model retains 83% of its original average accuracy. Fei Wang 0032, Li Shen 0008, Liang Ding 0006, Chao Xue 0003, Ye Liu 0014, Changxing Ding |
NeurIPS | 1 |
| 2025 | Unveiling the Power of Self-Supervision for Multi-View Multi-Human Association and TrackingabstractMulti-view multi-human association and tracking (MvMHAT), is an emerging yet important problem for multi-person scene video surveillance, aiming to track a group of people over time in each view, as well as to identify the same person across different views at the same time, which is different from previous MOT and multi-camera MOT tasks only considering the over-time human tracking. This way, the videos for MvMHAT require more complex annotations while containing more information for self-learning. In this work, we tackle this problem with an end-to-end neural network in a self-supervised learning manner. Specifically, we propose to take advantage of the spatial-temporal self-consistency rationale by considering three properties of reflexivity, symmetry, and transitivity. Besides the reflexivity property that naturally holds, we design the self-supervised learning losses based on the properties of symmetry and transitivity, for both appearance feature learning and assignment matrix optimization, to associate multiple humans over time and across views. Furthermore, to promote the research on MvMHAT, we build two new large-scale benchmarks for the network training and testing of different algorithms. Extensive experiments on the proposed benchmarks verify the effectiveness of our method. We have released the benchmark and code to the public. Wei Feng 0005, Fei Wang 0032, Rui-Ze Han, Yiyang Gan, Zekun Qian, Junhui Hou, Song Wang 0002 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2025 | DIST+: Knowledge Distillation From a Stronger Adaptive TeacherabstractThe paper introduces DIST, an innovative knowledge distillation method that excels in learning from a superior teacher model. DIST differentiates itself from conventional techniques by adeptly handling the often significant prediction discrepancies between the student and teacher models. It achieves this by focusing on maintaining the relationships between their predictions, implementing a correlation-based loss to explicitly capture the teacher's intrinsic inter-class relations. Moreover, DIST uniquely considers the semantic similarities between different instances and each class at the intra-class level. The method is further enhanced by two significant improvements: (1) A teacher acclimation strategy, which effectively reduces the discrepancy between teacher and student, thereby optimizing the distillation process. (2) An extension of the DIST loss from the logit level to the feature level, a modification that proves especially beneficial for dense prediction tasks. DIST stands out for its simplicity, practicality, and adaptability to various architectures, model sizes, and training strategies. It consistently delivers state-of-the-art results across a range of applications, including image classification, object detection, and semantic segmentation. Tao Huang 0020, Shan You, Fei Wang 0032, Chen Qian 0006, Chang Xu 0002 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2025 | FaceChain-MMID: Generating highly identity-consistent realistic portraits via dividing & merging multi-modal representations
Chao Xu 0023, Fei Wang 0032, Baigui Sun, Jian Zhao 0006 |
Pattern Recognit. | 2 |
| 2024 | Harnessing the Power of MLLMs for Transferable Text-to-Image Person ReIDabstractText-to-image person re-identification (ReID) retrieves pedestrian images according to textual descriptions. Manually annotating textual descriptions is time-consuming, restricting the scale of existing datasets and therefore the generalization ability of ReID models. As a result, we study the transferable text-to-image ReID problem, where we train a model on our proposed large-scale database and directly deploy it to various datasets for evaluation. We obtain substantial training data via Multi-modal Large Language Models (MLLMs). Moreover, we identify and address two key challenges in utilizing the obtained textual descriptions. First, an MLLM tends to generate descriptions with similar structures, causing the model to overfit specific sentence patterns. Thus, we propose a novel method that uses MLLMs to caption images according to various templates. These templates are obtained using a multi-turn dialogue with a Large Language Model (LLM). Therefore, we can build a large-scale dataset with diverse textual descriptions. Second, an MLLM may produce incorrect descriptions. Hence, we introduce a novel method that automatically identifies words in a description that do not correspond with the image. This method is based on the similarity between one text and all patch token embeddings in the image. Then, we mask these words with a larger probability in the subsequent training epoch, alleviating the impact of noisy textual descriptions. The experimental results demonstrate that our methods significantly boost the direct transfer text-to-image ReID performance. Benefiting from the pre-trained model weights, we also achieve state-of-the-art performance in the traditional evaluation settings.https://github.com/WentaoTan/MLLM4Text-ReID Wentao Tan, Changxing Ding, Jiayu Jiang, Fei Wang 0032, Yibing Zhan, Dapeng Tao |
CVPR | 4 |
| 2024 | PPDM++: Parallel Point Detection and Matching for Fast and Accurate HOI DetectionabstractHuman-Object Interaction (HOI) detection aims to understand human activities by detecting interaction triplets. Previous HOI detection methods adopt a two-stage instance-driven paradigm. Unfortunately, many non-interactive human-object pairs generated by the first stage are the main obstacle impeding HOI detectors from high efficiency and promising performance. To remedy this, we propose a novel top-down interaction-driven paradigm, detecting interactions first and bridging interactive human-object pairs through interactions. We formulate HOI as a point triplet human point, interaction point, object point and design a Parallel Point Detection and Matching (PPDM) framework. We further take advantage of two-stage methods and propose a novel framework, PPDM++, that detects the interactive human-object pairs by PPDM, then extracts region features for each pair to predict actions. The core of PPDM/PPDM++ is to convert the instance-driven bottom-up paradigm to an interaction-driven top-down paradigm, thus avoiding additional computation costs from traversing a tremendous number of non-interactive pairs. Benefiting from the advanced paradigm, PPDM/PPDM++ has achieved significant performance gains with high efficiency. PPDM-DLA-34 has achieved 19.94 mAP with 42 FPS as the first real-time HOI detector, and PPDM++-SwinB achieves 30.1 mAP with 17 FPS on HICO-DET dataset. We also built an application-oriented database named HOI-A, a supplement to the existing datasets. Yue Liao, Si Liu 0001, Yulu Gao, Aixi Zhang, Fei Wang 0032, Bo Li 0006 |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2024 | Weak Augmentation Guided Relational Self-Supervised LearningabstractSelf-supervised Learning (SSL) including the mainstream contrastive learning has achieved great success in learning visual representations without data annotations. However, most methods mainly focus on the instance level information (i.e., the different augmented images of the same instance should have the same feature or cluster into the same class), but there is a lack of attention on the relationships between different instances. In this paper, we introduce a novel SSL paradigm, which we term as relational self-supervised learning (ReSSL) framework that learns representations by modeling the relationship between different instances. Specifically, our proposed method employs sharpened distribution of pairwise similarities among different instances as relation metric, which is thus utilized to match the feature embeddings of different augmentations. To boost the performance, we argue that weak augmentations matter to represent a more reliable relation, and leverage momentum strategy for practical efficiency. The designed asymmetric predictor head and an InfoNCE warm-up strategy enhance the robustness to hyper-parameters and benefit the resulting performance. Experimental results show that our proposed ReSSL substantially outperforms the state-of-the-art methods across different network architectures, including various lightweight networks (e.g., EfficientNet and MobileNet). Mingkai Zheng, Shan You, Fei Wang 0032, Chen Qian 0006, Changshui Zhang, Xiaogang Wang 0005, Chang Xu 0002 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2024 | Can Linguistic Knowledge Improve Multimodal Alignment in Vision-Language Pretraining?abstractThe field of multimedia research has witnessed significant interest in leveraging multimodal pretrained neural network models to perceive and represent the physical world. Among these models, vision-language pretraining (VLP) has emerged as a captivating topic. Currently, the prevalent approach in VLP involves supervising the training process with paired image-text data. However, limited efforts have been dedicated to exploring the extraction of essential linguistic knowledge, such as semantics and syntax, during VLP and understanding its impact on multimodal alignment. In response, our study aims to shed light on the influence of comprehensive linguistic knowledge encompassing semantic expression and syntactic structure on multimodal alignment. To achieve this, we introduce SNARE , a large-scale multimodal alignment probing benchmark designed specifically for the detection of vital linguistic components, including lexical, semantic, and syntax knowledge. SNARE offers four distinct tasks: Semantic Structure, Negation Logic, Attribute Ownership, and Relationship Composition. Leveraging SNARE , we conduct holistic analyses of six advanced VLP models (BLIP, CLIP, Flava, X-VLM, BLIP2, and GPT-4), along with human performance, revealing key characteristics of the VLP model: (i) Insensitivity to complex syntax structures, relying primarily on content words for sentence comprehension. (ii) Limited comprehension of sentence combinations and negations. (iii) Challenges in determining actions or spatial relations within visual information, as well as difficulties in verifying the correctness of ternary relationships. Based on these findings, we propose the following strategies to enhance multimodal alignment in VLP: (1) Utilize a large generative language model as the language backbone in VLP to facilitate the understanding of complex sentences. (2) Establish high-quality datasets that emphasize content words and employ simple syntax, such as short-distance semantic composition, to improve multimodal alignment. (3) Incorporate more fine-grained visual knowledge, such as spatial relationships, into pretraining objectives. 1 Fei Wang 0032, Liang Ding 0006, Jun Rao, Ye Liu 0014, Li Shen 0008, Changxing Ding |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2023 | Boosting Novel Category Discovery Over Domains with Soft Contrastive Learning and All in One ClassifierabstractUnsupervised domain adaptation (UDA) has proven to be highly effective in transferring knowledge from a label-rich source domain to a label-scarce target domain. However, the presence of additional novel categories in the target domain has led to the development of open-set domain adaptation (ODA) and universal domain adaptation (UNDA). Existing ODA and UNDA methods treat all novel categories as a single, unified unknown class and attempt to detect it during training. However, we found that domain variance can lead to more significant view-noise in unsupervised data augmentation, which affects the effectiveness of contrastive learning (CL) and causes the model to be overconfident in novel category discovery. To address these issues, a framework named Soft-contrastive All-in-one Network (SAN) is proposed for ODA and UNDA tasks. SAN includes a novel data-augmentation-based soft contrastive learning (SCL) loss to fine-tune the backbone for feature transfer and a more human-intuitive classifier to improve new class discovery capability. The SCL loss weakens the adverse effects of the data augmentation view-noise problem which is amplified in domain transfer tasks. The All-in-One (AIO) classifier overcomes the overconfidence problem of current mainstream closed-set and open-set classifiers. Visualization and ablation experiments demonstrate the effectiveness of the proposed innovations. Furthermore, extensive experiment results on ODA and UNDA show that SAN outperforms existing state-of-the-art methods. Zelin Zang, Senqiao Yang, Fei Wang 0032, Baigui Sun, Xuansong Xie, Stan Z. Li |
ICCV | 4 |
| 2023 | SimMatchV2: Semi-Supervised Learning with Graph ConsistencyabstractSemi-Supervised image classification is one of the most fundamental problem in computer vision, which significantly reduces the need for human labor. In this paper, we introduce a new semi-supervised learning algorithm - Sim-MatchV2, which formulates various consistency regularizations between labeled and unlabeled data from the graph perspective. In SimMatchV2, we regard the augmented view of a sample as a node, which consists of a label and its corresponding representation. Different nodes are connected with the edges, which are measured by the similarity of the node representations. Inspired by the message passing and node classification in graph theory, we propose four types of consistencies, namely 1) node-node consistency, 2) node-edge consistency, 3) edge-edge consistency, and 4) edge-node consistency. We also uncover that a simple feature normalization can reduce the gaps of the feature norm between different augmented views, significantly improving the performance of SimMatchV2. Our SimMatchV2 has been validated on multiple semi-supervised learning benchmarks. Notably, with ResNet-50 as our backbone and 300 epochs of training, SimMatchV2 achieves 71.9% and 76.2% Top-1 Accuracy with 1% and 10% labeled examples on ImageNet, which significantly outperforms the previous methods and achieves state-of-the-art performance. Mingkai Zheng, Shan You, Lang Huang 0001, Fei Wang 0032, Chen Qian 0006, Chang Xu 0002 |
ICCV | 5 |
| 2023 | DiffNAS: Bootstrapping Diffusion Models by Prompting for Better ArchitecturesabstractDiffusion models have recently exhibited remarkable performance on synthetic data. After a diffusion path is selected, a base model, such as UNet, operates as a denoising autoencoder, primarily predicting noises that need to be eliminated step by step. Consequently, it is crucial to employ a model that aligns with the expected budgets to facilitate superior synthetic performance. In this paper, we meticulously analyze the diffusion model and engineer a base model search approach, denoted "DiffNAS". Specifically, we leverage GPT-4 as a supernet to expedite the search, supplemented with a search memory to enhance the results. Moreover, we employ RFID as a proxy to promptly rank the experimental outcomes produced by GPT-4. We also adopt a rapid-convergence training strategy to boost search efficiency. Rigorous experimentation corroborates that our algorithm can augment the search efficiency by $2 \times$ under GPT-based scenarios, while also attaining a performance of 2.82 with 0.37 improvement in FID on CIFAR10 relative to the benchmark IDDPM algorithm. Xiu Su, Shan You, Fei Wang 0032, Chen Qian 0006, Chang Xu 0002 |
ICDM | 4 |
| 2023 | Masked Distillation with Receptive Tokens
Tao Huang 0020, Yuan Zhang 0020, Shan You, Fei Wang 0032, Chen Qian 0006, Jian Cao 0002, Chang Xu 0002 |
ICLR | 4 |
| 2023 | DamoFD: Digging into Backbone Design on Face Detection
Yang Liu 0356, Jiankang Deng, Fei Wang 0032, Xuansong Xie, Baigui Sun |
ICLR | 3 |
| 2023 | Generating New Paintings by Semantic Guidance
Fei Wang 0032, Junzhou Xie, Weifeng Liu 0001 |
MMM (2) | 2 |
| 2023 | Knowledge Diffusion for DistillationabstractThe representation gap between teacher and student is an emerging topic in knowledge distillation (KD). To reduce the gap and improve the performance, current methods often resort to complicated training schemes, loss functions, and feature alignments, which are task-specific and feature-specific. In this paper, we state that the essence of these methods is to discard the noisy information and distill the valuable information in the feature, and propose a novel KD method dubbed DiffKD, to explicitly denoise and match features using diffusion models. Our approach is based on the observation that student features typically contain more noises than teacher features due to the smaller capacity of student model. To address this, we propose to denoise student features using a diffusion model trained by teacher features. This allows us to perform better distillation between the refined clean feature and teacher feature. Additionally, we introduce a light-weight diffusion model with a linear autoencoder to reduce the computation cost and an adaptive noise matching module to improve the denoising performance. Extensive experiments demonstrate that DiffKD is effective across various types of features and achieves state-of-the-art performance consistently on image classification, object detection, and semantic segmentation tasks. Code is available at https://github.com/hunto/DiffKD. Tao Huang 0020, Yuan Zhang 0020, Mingkai Zheng, Shan You, Fei Wang 0032, Chen Qian 0006, Chang Xu 0002 |
NeurIPS | 5 |
| 2023 | Searching for Network Width With Bilaterally Coupled NetworkabstractSearching for a more compact network width recently serves as an effective way of channel pruning for the deployment of convolutional neural networks (CNNs) under hardware constraints. To fulfil the searching, a one-shot supernet is usually leveraged to efficiently evaluate the performance w.r.t. different network widths. However, current methods mainly follow a unilaterally augmented (UA) principle for the evaluation of each width, which induces the training unfairness of channels in supernet. In this article, we introduce a new supernet called Bilaterally Coupled Network (BCNet) to address this issue. In BCNet, each channel is fairly trained and responsible for the same amount of network widths, thus each network width can be evaluated more accurately. Besides, we propose to reduce the redundant search space and present the BCNetV2 as the enhanced supernet to ensure rigorous training fairness over channels. Furthermore, we leverage a stochastic complementary strategy for training the BCNet, and propose a prior initial population sampling method to boost the performance of the evolutionary search. We also propose a new open-source width search benchmark on macro structures named Channel-Bench-Macro for the better comparisons of the width search algorithms with MobileNet- and ResNet-like architectures. Extensive experiments on the benchmark datasets demonstrate that our method can achieve state-of-the-art performance. Xiu Su, Shan You, Jiyang Xie 0001, Fei Wang 0032, Chen Qian 0006, Changshui Zhang, Chang Xu 0002 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2023 | Personality in Daily Life: Multi-Situational Physiological Signals Reflect Big-Five Personality TraitsabstractThe popularity of wearable physiological recording devices has opened up new possibilities for the assessment of personality traits in everyday life. Compared with traditional questionnaires or laboratory assessments, wearable device-based measurements can collect rich data about individual physiological activities in real-life situations without interfering with normal life, enabling a more comprehensive description of individual differences. The present study aimed to explore the assessment of individuals' Big-Five personality traits by physiological signals in daily life situations. A commercial bracelet was used to track the heart rate (HR) data from eighty college students (all male) enrolled in a special training program with a strictly-controlled daily schedule for ten consecutive working days. Their HR activities were divided into five daily situations (morning exercise, morning classes, afternoon classes, free time in the evening, and self-study situations) according to their daily schedule. Regression analyses with HR-based features in these five situations averaged across the ten days revealed significant cross-validated quantitative prediction correlations of 0.32 and 0.26 for the dimensions of Openness and Extraversion, with the prediction correlation trending significance for Conscientiousness and Neuroticism. Moreover, the multi-situation HR-based results were in general superior to those based on single-situation HR-based features, as well as those based on the multi-situation self-reported emotion ratings. Togetherour findings demonstrate the link between personality and daily HR measures using state-of-the-art commercial devices and could shed light on the development of Big-Five personality assessment based on daily multi-situation physiological measures. Xinyu Shui, Fei Wang 0032, Dan Zhang 0014 |
IEEE J. Biomed. Health Informatics | 4 |
| 2022 | GreedyNASv2: Greedier Search with a Greedy Path FilterabstractTraining a good supernet in one-shot NAS methods is difficult since the search space is usually considerably huge$(\mathrm{e}.\mathrm{g}.,\ 13^{21})$. In order to enhance the supernet's evaluation ability, one greedy strategy is to sample good paths, and let the supernet lean towards the good ones and ease its evaluation burden as a result. However, in practice the search can be still quite inefficient since the identification of good paths is not accurate enough and sampled paths still scatter around the whole search space. In this paper, we leverage an explicit path filter to capture the characteristics of paths and directly filter those weak ones, so that the search can be thus implemented on the shrunk space more greedily and efficiently. Concretely, based on the fact that good paths are much less than the weak ones in the space, we argue that the label of “weak paths” will be more confident and reliable than that of “good paths” in multi-path sampling. In this way, we thus cast the training of path filter in the positive and unlabeled (PU) learning paradigm, and also encourage a path embedding as better path/operation representation to enhance the identification capacity of the learned filter. By dint of this embedding, we can further shrink the search space by aggregating similar operations with similar embeddings, and the search can be more efficient and accurate. Extensive experiments validate the effectiveness of the proposed method GreedyNASv2. For example, our obtained GreedyNASv2-L achieves 81.1% Top-1 accuracy on ImageNet dataset, significantly outperforming the ResNet-50 strong baselines. Tao Huang 0020, Shan You, Fei Wang 0032, Chen Qian 0006, Changshui Zhang, Xiaogang Wang 0001, Chang Xu 0002 |
CVPR | 3 |
| 2022 | DyRep: Bootstrapping Training with Dynamic Re-parameterizationabstractStructural re-parameterization (Rep) methods achieve noticeable improvements on simple VGG-style networks. Despite the prevalence, current Rep methods simply re-parameterize all operations into an augmented network, including those that rarely contribute to the model's performance. As such, the price to pay is an expensive computational overhead to manipulate these unnecessary behaviors. To eliminate the above caveats, we aim to boot-strap the training with minimal cost by devising a dynamic re-parameterization (DyRep) method, which encodes Rep technique into the training process that dynamically evolves the network structures. Concretely, our proposal adaptively finds the operations which contribute most to the loss in the network, and applies Rep to enhance their representational capacity. Besides, to suppress the noisy and redundant operations introduced by Rep, we devise a de-parameterization technique for a more compact re-parameterization. With this regard, DyRep is more efficient than Rep since it smoothly evolves the given network instead of constructing an over-parameterized network. Experimental results demonstrate our effectiveness, e.g., DyRep improves the accuracy of ResNet-18 by 2.04% on ImageNet and reduces 22% runtime over the baseline. Code is avail-able at: https://github.com/hunto/DyRep. Tao Huang 0020, Shan You, Fei Wang 0032, Chen Qian 0006, Chang Xu 0002 |
CVPR | 5 |
| 2022 | Learning Where to Learn in Cross-View Self-Supervised LearningabstractSelf-supervised learning (SSL) has made enormous progress and largely narrowed the gap with the supervised ones, where the representation learning is mainly guided by a projection into an embedding space. During the projection, current methods simply adopt uniform aggregation of pixels for embedding; however, this risks involving object-irrelevant nuisances and spatial misalignment for different augmentations. In this paper, we present a new approach, Learning Where to Learn (LEWEL), to adaptively aggregate spatial information of features, so that the projected embeddings could be exactly aligned and thus guide the feature learning better. Concretely, we reinterpret the projection head in SSL as a per-pixel projection and predict a set of spatial alignment maps from the original features by this weight-sharing projection head. A spectrum of aligned embeddings is thus obtained by aggregating the features with spatial weighting according to these alignment maps. As a result of this adaptive alignment, we observe substantial improvements on both image-level prediction and dense prediction at the same time: LEWEL improves MoCov2 [15] by 1.6%/1.3%/0.5%/0.4% points, improves BYOL [14] by 1.3%/1.3%/0.7%/0.6% points, on ImageNet linear/semi-supervised classification, Pascal VOC semantic segmentation, and object detection, respectively.††Code: https://t.1y/ZI0A. Lang Huang 0001, Shan You, Mingkai Zheng, Fei Wang 0032, Chen Qian 0006, Toshihiko Yamasaki |
CVPR | 4 |
| 2022 | MogFace: Towards a Deeper Appreciation on Face DetectionabstractBenefiting from the pioneering design of generic object detectors, significant achievements have been made in the field of face detection. Typically, the architectures of the backbone, feature pyramid layer, and detection head module within the face detector all assimilate the excellent experience from general object detectors. However, several effective methods, including label assignment and scale-level data augmentation strategy, fail to maintain consistent superiority when applying on the face detector directly. Concretely, the former strategy involves a vast body of hyperparameters and the latter one suffers from the challenge of scale distribution bias between different detection tasks, which both limit their generalization abilities. Furthermore, in order to provide accurate face bounding boxes for facial down-stream tasks, the face detector imperatively requires the elimination of false alarms. As a result, practical solutions on label assignment, scale-level data augmentation, and reducing false alarms are necessary for advancing face detectors. In this paper, we focus on resolving three aforementioned challenges that exiting methods are difficult to finish off and present a novel face detector, termed MogFace. In our Mogface, three key components, Adaptive Online Incremental Anchor Mining Strategy, Selective Scale Enhancement Strategy and Hierarchical Context-Aware Module, are separately proposed to boost the performance of face detectors. Finally, to the best of our knowledge, our MogFace is the best face detector on the Wider Face leader-board, achieving all champions across different testing scenarios. The code is available at https://github.com/damo-cv/MogFace. Yang Liu 0069, Fei Wang 0032, Jiankang Deng, Baigui Sun, Hao Li 0030 |
CVPR | 2 |
| 2022 | A Keypoint-based Global Association Network for Lane DetectionabstractLane detection is a challenging task that requires predicting complex topology shapes of lane lines and distinguishing different types of lanes simultaneously. Earlier works follow a top-down roadmap to regress predefined anchors into various shapes of lane lines, which lacks enough flexibility to fit complex shapes of lanes due to the fixed anchor shapes. Lately, some works propose to formulate lane detection as a keypoint estimation problem to describe the shapes of lane lines more flexibly and gradually group adjacent keypoints belonging to the same lane line in a point-by-point manner, which is inefficient and time-consuming during postprocessing. In this paper, we propose a Global Association Network (GANet) to formulate the lane detection problem from a new perspective, where each keypoint is directly regressed to the starting point of the lane line instead of point-by-point extension. Concretely, the association of keypoints to their belonged lane line is conducted by predicting their offsets to the corresponding starting points of lanes globally without dependence on each other, which could be done in parallel to greatly improve efficiency. In addition, we further propose a Lane-aware Feature Aggregator (LFA), which adaptively captures the local correlations between adjacent keypoints to supplement local information to the global association. Extensive experiments on two popular lane detection benchmarks show that our method outperforms previous methods with F1 score of 79.63% on CULane and 97.71% on Tusimple dataset with high FPS. Jinsheng Wang, Yinchao Ma, Shaofei Huang 0001, Tianrui Hui, Fei Wang 0032, Tianzhu Zhang 0001 |
CVPR | 5 |
| 2022 | SimMatch: Semi-supervised Learning with Similarity MatchingabstractLearning with few labeled data has been a longstanding problem in the computer vision and machine learning research community. In this paper, we introduced a new semi-supervised learning framework, SimMatch, which simulta-neously considers semantic similarity and instance similarity. In SimMatch, the consistency regularization will be applied on both semantic-level and instance-level. The different augmented views of the same instance are encouraged to have the same class prediction and similar similarity re-lationship respected to other instances. Next, we instanti-ated a labeled memory buffer to fully leverage the ground truth labels on instance-level and bridge the gaps between the semantic and instance similarities. Finally, we proposed the unfolding and aggregation operation which allows these two similarities be isomorphically transformed with each other. In this way, the semantic and instance pseudo-labels can be mutually propagated to generate more high-quality and reliable matching targets. Extensive ex-perimental results demonstrate that SimMatch improves the performance of semi-supervised learning tasks across dif-ferent benchmark datasets and different settings. Notably, with 400 epochs of training, SimMatch achieves 67.2%, and 74.4% Top-1 Accuracy with 1% and 10% labeled examples on ImageNet, which significantly outperforms the baseline methods and is better than previous semi-supervised learning frameworks. Mingkai Zheng, Shan You, Lang Huang 0001, Fei Wang 0032, Chen Qian 0006, Chang Xu 0002 |
CVPR | 4 |
| 2022 | ViTAS: Vision Transformer Architecture Search
Xiu Su, Shan You, Jiyang Xie 0001, Mingkai Zheng, Fei Wang 0032, Chen Qian 0006, Changshui Zhang, Xiaogang Wang 0001, Chang Xu 0002 |
ECCV (21) | 5 |
| 2022 | HEAD: HEtero-Assists Distillation for Heterogeneous Object Detectors
Luting Wang 0001, Yue Liao, Zeren Jiang, Jianlong Wu, Fei Wang 0032, Chen Qian 0006, Si Liu 0001 |
ECCV (9) | 6 |
| 2022 | ScaleNet: Searching for the Model to Scale
Jiyang Xie 0001, Xiu Su, Shan You, Zhanyu Ma, Fei Wang 0032, Chen Qian 0006 |
ECCV (21) | 5 |
| 2022 | Data Agnostic Filter Gating For Efficient Deep NetworksabstractFilter pruning is essential for deploying a well-trained CNN model on edge computation devices with a target computation budget (e.g., FLOPs). Current filter pruning methods mainly focus on leveraging feature maps to analyze the importance of filters, and prune those with less impact on the value of the CNN’s loss function, thereby ignoring the variance of input batches to differences in sparse structure over the filters. In this paper, we propose a data-agnostic filter pruning method that uses an auxiliary network named Dagger module to induce pruning with the pre-trained weights as input. Besides, to help prune filters with a preset FLOPs constraint, we utilize an explicit FLOPs-aware regularisation mechanism to directly promote pruning filters toward the target FLOPs. Experimental results on CIFAR-10 and ImageNet datasets show that the proposed filter pruning method surpasses the state-of-the-art. Hongyan Xu 0002, Xiu Su, Shan You, Tao Huang 0020, Fei Wang 0032, Chen Qian 0006, Changshui Zhang, Chang Xu 0002, Dadong Wang, Arcot Sowmya |
ICASSP | 5 |
| 2022 | Relational Surrogate Loss Learning
Tao Huang 0020, Zekang Li, Hua Lu 0017, Yong Shan, Shusheng Yang, Fei Wang 0032, Shan You, Chang Xu 0002 |
ICLR | 7 |
| 2022 | Cross-Modality Domain Adaptation for Freespace Detection: A Simple yet Effective BaselineabstractAs one of the fundamental functions of autonomous driving system, freespace detection aims at classifying each pixel of the image captured by the camera as drivable or non-drivable. Current works of freespace detection heavily rely on large amount of densely labeled training data for accuracy and robustness, which is time-consuming and laborious to collect and annotate. To the best of our knowledge, we are the first work to explore unsupervised domain adaptation for freespace detection to alleviate the data limitation problem with synthetic data. We develop a cross-modality domain adaptation framework which exploits both RGB images and surface normal maps generated from depth images. A Collaborative Cross Guidance (CCG) module is proposed to leverage the context information of one modality to guide the other modality in a cross manner, thus realizing inter-modality intra-domain complement. To better bridge the domain gap between source domain (synthetic data) and target domain (real-world data), we also propose a Selective Feature Alignment (SFA) module which only aligns the features of consistent foreground area between the two domains, thus realizing inter-domain intra-modality adaptation. Extensive experiments are conducted by adapting three different synthetic datasets to one real-world dataset for freespace detection respectively. Our method performs closely to fully supervised freespace detection methods (93.08% v.s. 97.50% F1 score) and outperforms other general unsupervised domain adaptation methods for semantic segmentation with large margins, which shows the promising potential of domain adaptation for freespace detection. Leyan Zhu, Shaofei Huang 0001, Tianrui Hui, Fei Wang 0032, Si Liu 0001 |
ACM Multimedia | 6 |
| 2022 | Green Hierarchical Vision Transformer for Masked Image ModelingabstractWe present an efficient approach for Masked Image Modeling (MIM) with hierarchical Vision Transformers (ViTs), allowing the hierarchical ViTs to discard masked patches and operate only on the visible ones. Our approach consists of three key designs. First, for window attention, we propose a Group Window Attention scheme following the Divide-and-Conquer strategy. To mitigate the quadratic complexity of the self-attention w.r.t. the number of patches, group attention encourages a uniform partition that visible patches within each local window of arbitrary size can be grouped with equal size, where masked self-attention is then performed within each group. Second, we further improve the grouping strategy via the Dynamic Programming algorithm to minimize the overall computation cost of the attention on the grouped patches. Third, as for the convolution layers, we convert them to the Sparse Convolution that works seamlessly with the sparse data, i.e., the visible patches in MIM. As a result, MIM can now work on most, if not all, hierarchical ViTs in a green and efficient way. For example, we can train the hierarchical ViTs, e.g., Swin Transformer and Twins Transformer, about 2.7$\times$ faster and reduce the GPU memory usage by 70%, while still enjoying competitive performance on ImageNet classification and the superiority on downstream COCO object detection benchmarks. Lang Huang 0001, Shan You, Mingkai Zheng, Fei Wang 0032, Chen Qian 0006, Toshihiko Yamasaki |
NeurIPS | 4 |
| 2022 | Knowledge Distillation from A Stronger TeacherabstractUnlike existing knowledge distillation methods focus on the baseline settings, where the teacher models and training strategies are not that strong and competing as state-of-the-art approaches, this paper presents a method dubbed DIST to distill better from a stronger teacher. We empirically find that the discrepancy of predictions between the student and a stronger teacher may tend to be fairly severer. As a result, the exact match of predictions in KL divergence would disturb the training and make existing methods perform poorly. In this paper, we show that simply preserving the relations between the predictions of teacher and student would suffice, and propose a correlation-based loss to capture the intrinsic inter-class relations from the teacher explicitly. Besides, considering that different instances have different semantic similarities to each class, we also extend this relational match to the intra-class level. Our method is simple yet practical, and extensive experiments demonstrate that it adapts well to various architectures, model sizes and training strategies, and can achieve state-of-the-art performance consistently on image classification, object detection, and semantic segmentation tasks. Code is available at: https://github.com/hunto/DIST_KD. Tao Huang 0020, Shan You, Fei Wang 0032, Chen Qian 0006, Chang Xu 0002 |
NeurIPS | 3 |
| 2022 | Where Does the Performance Improvement Come From?: - A Reproducibility Concern about Image-Text RetrievalabstractThis article aims to provide the information retrieval community with some reflections on recent advances in retrieval learning by analyzing the reproducibility of image-text retrieval models. Due to the increase of multimodal data over the last decade, image-text retrieval has steadily become a major research direction in the field of information retrieval. Numerous researchers train and evaluate image-text retrieval algorithms using benchmark datasets such as MS-COCO and Flickr30k. Research in the past has mostly focused on performance, with multiple state-of-the-art methodologies being suggested in a variety of ways. According to their assertions, these techniques provide improved modality interactions and hence more precise multimodal representations. In contrast to previous works, we focus on the reproducibility of the approaches and the examination of the elements that lead to improved performance by pretrained and nonpretrained models in retrieving images and text. Jun Rao, Fei Wang 0032, Liang Ding 0006, Shuhan Qi, Yibing Zhan, Weifeng Liu 0001, Dacheng Tao |
SIGIR | 2 |
| 2022 | SafeOSL: Ensuring memory safety of C via ownership-based intermediate languageabstractAbstract The unsafe features of C make it a big challenge to ensure memory safety of C programs, and often lead to memory errors that can result in vulnerabilities. Various formal verification techniques for ensuring memory safety of C have been proposed. However, most of them either have a high overhead, such as state explosion problem in model checking, or have false positives, such as abstract interpretation. In this article, by innovatively borrowing ownership system from Rust, we propose a novel and sound static memory safety analysis approach, named SafeOSL. Its basic idea is an ownership‐based intermediate language, called ownership system language (OSL), which captures the features of the ownership system in Rust. Ownership system specifies the relations among variables and memory locations, and maintains invariants that can ensure memory safety. The semantics of OSL is formalized in K‐framework, which is a rewriting‐logic based tool. C programs to be checked are first transformed into OSL programs and then detected by OSL semantics. Experimental results have demonstrated that SafeOSL is effective in detecting memory errors of C. Moreover, the translations and experiments indicate that the intermediate language OSL could be reused by other programming languages to detect memory errors. Xiaohua Yin, Shuanglong Kan, Guohua Shen, Zhe Chen 0011, Yang Liu 0003, Fei Wang 0032 |
Softw. Pract. Exp. | 7 |
| 2022 | Quantitative Personality Predictions From a Brief EEG RecordingabstractThe assessment of personality is crucial not only for scientific inquiries but also for real-world applications such as personnel selection. In this article, we propose and validate a novel implicit measure to predict an individual's levels in the Big Five personality traits from 5 minutes of electroencephalography (EEG) recordings. Participants viewed Chinese words with positive, negative, and neutral emotions. The multi-channel event-related potentials elicited by these emotional words were used to train a sparse regression model for personality prediction. Results from a large test sample of 196 participants indicated that the personality scores derived from the proposed measure reached significant correlations with a commonly used questionnaire (r = .50, .60, .49, .55, and .49 for agreeableness, conscientiousness, neuroticism, openness, and extraversion, respectively). The EEG-based personality scores showed good external validity as well, capable of predicting behavioral indices and psychological adjustment similar to self-reported scores. Besides, the EEG-based scores were relatively stable across time, as reflected by the test-retest reliability of .5 ∼ .7 for the five personality traits within a cohort of 33 participants 19-78 days later. These evaluations suggest that the proposed measure can serve as a viable alternative to conventional personality questionnaires in practice. Chengpeng Wu, Shimin Fu, Fei Wang 0032, Dan Zhang 0014 |
IEEE Trans. Affect. Comput. | 6 |
| 2021 | Reformulating HOI Detection As Adaptive Set PredictionabstractDetermining which image regions to concentrate is critical for Human-Object Interaction (HOI) detection. Conventional HOI detectors focus on either detected human and object pairs or pre-defined interaction locations, which limits learning of the effective features. In this paper, we reformulate HOI detection as an adaptive set prediction problem, with this novel formulation, we propose an Adaptive Set-based one-stage framework (AS-Net) with parallel instance and interaction branches. To attain this, we map a trainable interaction query set to an interaction prediction set with transformer. Each query adaptively aggregates the interaction-relevant features from global contexts through multi-head co-attention. Besides, the training process is supervised adaptively by matching each ground-truth with the interaction prediction. Furthermore, we design an effective instance-aware attention module to introduce instructive features from the instance branch into the interaction branch. Our method outperforms previous state-of-the-art methods without any extra human pose and language features on three challenging HOI detection datasets. Especially, we achieve over 31% relative improvement on a large scale HICO-DET dataset. Code is available at https://github.com/yoyomimi/AS-Net. Mingfei Chen, Yue Liao, Si Liu 0001, Zhiyuan Chen 0008, Fei Wang 0032, Chen Qian 0006 |
CVPR | 5 |
| 2021 | Collaborative Spatial-Temporal Modeling for Language-Queried Video Actor SegmentationabstractLanguage-queried video actor segmentation aims to predict the pixel-level mask of the actor which performs the actions described by a natural language query in the target frames. Existing methods adopt 3D CNNs over the video clip as a general encoder to extract a mixed spatio-temporal feature for the target frame. Though 3D convolutions are amenable to recognizing which actor is performing the queried actions, it also inevitably introduces misaligned spatial information from adjacent frames, which confuses features of the target frame and yields inaccurate segmentation. Therefore, we propose a collaborative spatial-temporal encoder-decoder framework which contains a 3D temporal encoder over the video clip to recognize the queried actions, and a 2D spatial encoder over the target frame to accurately segment the queried actors. In the decoder, a Language-Guided Feature Selection (LGFS) module is proposed to flexibly integrate spatial and temporal features from the two encoders. We also propose a Cross-Modal Adaptive Modulation (CMAM) module to dynamically recombine spatial- and temporal-relevant linguistic features for multimodal feature interaction in each stage of the two encoders. Our method achieves new state-of-the-art performance on two popular benchmarks with less computational overhead than previous approaches. Tianrui Hui, Shaofei Huang 0001, Si Liu 0001, Guanbin Li, Wenguan Wang, Jizhong Han, Fei Wang 0032 |
CVPR | 8 |
| 2021 | Prioritized Architecture Sampling With Monto-Carlo Tree SearchabstractOne-shot neural architecture search (NAS) methods significantly reduce the search cost by considering the whole search space as one network, which only needs to be trained once. However, current methods select each operation independently without considering previous layers. Besides, the historical information obtained with huge computation costs is usually used only once and then discarded. In this paper, we introduce a sampling strategy based on Monte Carlo tree search (MCTS) with the search space modeled as a Monte Carlo tree (MCT), which captures the dependency among layers. Furthermore, intermediate results are stored in the MCT for future decisions and a better exploration-exploitation balance. Concretely, MCT is updated using the training loss as a reward to the architecture performance; for accurately evaluating the numerous nodes, we propose node communication and hierarchical node selection methods in the training and search stages, respectively, making better uses of the operation rewards and hierarchical information. Moreover, for a fair comparison of different NAS methods, we construct an open-source NAS benchmark of a macro search space evaluated on CIFAR-10, namely NAS-Bench-Macro. Extensive experiments on NAS-Bench-Macro and ImageNet demonstrate that our method significantly improves search efficiency and performance. For example, by only searching 20 architectures, our obtained architecture achieves 78.0% top-1 accuracy with 442M FLOPs on ImageNet. Code (Benchmark) is available at: https://github.com/xiusu/NAS-Bench-Macro. Xiu Su, Tao Huang 0020, Yanxi Li 0001, Shan You, Fei Wang 0032, Chen Qian 0006, Changshui Zhang, Chang Xu 0002 |
CVPR | 5 |
| 2021 | BCNet: Searching for Network Width With Bilaterally Coupled NetworkabstractSearching for a more compact network width recently serves as an effective way of channel pruning for the deployment of convolutional neural networks (CNNs) under hardware constraints. To fulfill the searching, a one-shot supernet is usually leveraged to efficiently evaluate the performance w.r.t. different network widths. However, current methods mainly follow a unilaterally augmented (UA) principle for the evaluation of each width, which induces the training unfairness of channels in supernet. In this paper, we introduce a new supernet called Bilaterally Coupled Network (BCNet) to address this issue. In BCNet, each channel is fairly trained and responsible for the same amount of network widths, thus each network width can be evaluated more accurately. Besides, we leverage a stochastic complementary strategy for training the BCNet, and propose a prior initial population sampling method to boost the performance of the evolutionary search. Extensive experiments on benchmark CIFAR-10 and ImageNet datasets indicate that our method can achieve state-of-the-art or competing performance over other baseline methods. Moreover, our method turns out to further boost the performance of NAS models by refining their network widths. For example, with the same FLOPs budget, our obtained EfficientNet-B0 achieves 77.36% Top-1 accuracy on ImageNet dataset, surpassing the performance of original setting by 0.48%. Xiu Su, Shan You, Fei Wang 0032, Chen Qian 0006, Changshui Zhang, Chang Xu 0002 |
CVPR | 3 |
| 2021 | Towards Improving the Consistency, Efficiency, and Flexibility of Differentiable Neural Architecture SearchabstractMost differentiable neural architecture search methods construct a super-net for search and derive a target-net as its sub-graph for evaluation. There exists a significant gap between the architectures in search and evaluation. As a result, current methods suffer from an inconsistent, inefficient, and inflexible search process. In this paper, we introduce EnTranNAS that is composed of Engine-cells and Transit-cells. The Engine-cell is differentiable for architecture search, while the Transit-cell only transits a sub-graph by architecture derivation. Consequently, the gap between the architectures in search and evaluation is significantly reduced. Our method also spares much memory and computation cost, which speeds up the search process. A feature sharing strategy is introduced for more balanced optimization and more efficient search. Furthermore, we develop an architecture derivation method to replace the traditional one that is based on a hand-crafted rule. Our method enables differentiable sparsification, and keeps the derived architecture equivalent to that of Engine-cell, which further improves the consistency between search and evaluation. More importantly, it supports the search for topology where a node can be connected to prior nodes with any number of connections, so that the searched architectures could be more flexible. Our search on CIFAR-10 has an error rate of 2.22% with only 0.07 GPU-day. We can also directly perform the search on ImageNet with topology learnable and achieve a top-1 error rate of 23.8% in 2.1 GPU-day. Shan You, Fei Wang 0032, Chen Qian 0006, Zhouchen Lin |
CVPR | 4 |
| 2021 | MFR 2021: Masked Face Recognition CompetitionabstractThis paper presents a summary of the Masked Face Recognition Competitions (MFR) held within the 2021 International Joint Conference on Biometrics (IJCB 2021). The competition attracted a total of 10 participating teams with valid submissions. The affiliations of these teams are diverse and associated with academia and industry in nine different countries. These teams successfully submitted 18 valid solutions. The competition is designed to motivate solutions aiming at enhancing the face recognition accuracy of masked faces. Moreover, the competition considered the deployability of the proposed solutions by taking the compactness of the face recognition models into account. A private dataset representing a collaborative, multisession, real masked, capture scenario is used to evaluate the submitted solutions. In comparison to one of the topperforming academic face recognition solutions, 10 out of the 18 submitted solutions did score higher masked face verification accuracy. Fadi Boutros, Naser Damer, Jan Niklas Kolf, Kiran B. Raja, Florian Kirchbuchner, Ramachandra Raghavendra, Arjan Kuijper, Pengcheng Fang, Fei Wang 0032, David Montero 0002, Naiara Aginako, Basilio Sierra, Marcos Nieto Doncel, Mustafa Ekrem Erakin, Ugur Demir, Hazim Kemal Ekenel, Asaki Kataoka, Kohei Ichikawa, Shizuma Kubo, Jie Zhang 0071, Shiguang Shan, Klemen Grm, Vitomir Struc, Sachith Seneviratne, Nuran Kasthuriarachchi, Sanka Rasnayaka, Pedro C. Neto, Ana Filipa Sequeira, João Ribeiro Pinto, Mohsen Saffari, Jaime S. Cardoso 0001 |
IJCB | 10 |
| 2021 | Learning with Privileged TasksabstractMulti-objective multi-task learning aims to boost the performance of all tasks by leveraging their correlation and conflict appropriately. Nevertheless, in real practice, users may have preference for certain tasks, and other tasks simply serve as privileged or auxiliary tasks to assist the training of target tasks. The privileged tasks thus possess less or even no priority in the final task assessment by users. Motivated by this, we propose a privileged multiple descent algorithm to arbitrate the learning of target tasks and privileged tasks. Concretely, we introduce a privileged parameter so that the optimization direction does not necessarily follow the gradient from the privileged tasks, but concentrates more on the target tasks. Besides, we also encourage a priority parameter for the target tasks to control the potential distraction of optimization direction from the privileged tasks. In this way, the optimization direction can be more aggressively determined by weighting the gradients among target and privileged tasks, and thus highlight more the performance of target tasks under the unified multi-task learning context. Extensive experiments on synthetic and real-world datasets indicate that our method can achieve versatile Pareto solutions under varying preference for the target tasks. Yuru Song, Zan Lou, Shan You, Erkun Yang, Fei Wang 0032, Chen Qian 0006, Changshui Zhang, Xiaogang Wang 0001 |
ICCV | 5 |
| 2021 | Weakly Supervised Contrastive LearningabstractUnsupervised visual representation learning has gained much attention from the computer vision community because of the recent achievement of contrastive learning. Most of the existing contrastive learning frameworks adopt the instance discrimination as the pretext task, which treating every single instance as a different class. However, such method will inevitably cause class collision problems, which hurts the quality of the learned representation. Motivated by this observation, we introduced a weakly supervised contrastive learning framework (WCL) to tackle this issue. Specifically, our proposed framework is based on two projection heads, one of which will perform the regular instance discrimination task. The other head will use a graph-based method to explore similar samples and generate a weak label, then perform a supervised contrastive learning task based on the weak label to pull the similar images closer. We further introduced a K-Nearest Neighbor based multi-crop strategy to expand the number of positive samples. Extensive experimental results demonstrate WCL improves the quality of self-supervised representations across different datasets. Notably, we get a new state-of-the-art result for semi-supervised learning. With only 1% and 10% labeled examples, WCL achieves 65% and 72% ImageNet Top-1 Accuracy using ResNet50, which is even higher than SimCLRv2 with ResNet101. Mingkai Zheng, Fei Wang 0032, Shan You, Chen Qian 0006, Changshui Zhang, Xiaogang Wang 0001, Chang Xu 0002 |
ICCV | 2 |
| 2021 | Locally Free Weight Sharing for Network Width Search
Xiu Su, Shan You, Tao Huang 0020, Fei Wang 0032, Chen Qian 0006, Changshui Zhang, Chang Xu 0002 |
ICLR | 4 |
| 2021 | K-shot NAS: Learnable Weight-Sharing for NAS with K-shot SupernetsabstractIn one-shot weight sharing for NAS, the weights of each operation (at each layer) are supposed to be identical for all architectures (paths) in the supernet. However, this rules out the possibility of adjusting operation weights to cater for different paths, which limits the reliability of the evaluation results. In this paper, instead of counting on a single supernet, we introduce $K$-shot supernets and take their weights for each operation as a dictionary. The operation weight for each path is represented as a convex combination of items in a dictionary with a simplex code. This enables a matrix approximation of the stand-alone weight matrix with a higher rank ($K>1$). A \textit{simplex-net} is introduced to produce architecture-customized code for each path. As a result, all paths can adaptively learn how to share weights in the $K$-shot supernets and acquire corresponding weights for better evaluation. $K$-shot supernets and simplex-net can be iteratively trained, and we further extend the search to the channel dimension. Extensive experiments on benchmark datasets validate that K-shot NAS significantly improves the evaluation accuracy of paths and thus brings in impressive performance improvements. Xiu Su, Shan You, Mingkai Zheng, Fei Wang 0032, Chen Qian 0006, Changshui Zhang, Chang Xu 0002 |
ICML | 4 |
| 2021 | Workshop on Model MiningabstractHow to mine the knowledge in the pretrained models is of significance in achieving more promising performance, since practitioners have access to many pretrained models easily. This Workshop on Model Mining aims to investigate more diverse and advanced manners in mining knowledge within models, which tends to leverage the pretrained models more wisely, elegantly and systematically. There are many topics related to this workshop, such as distilling a lightweight model from a well-trained heavy model via teacher-student paradigm, and boosting the performance of the model by carefully designing the predecessor tasks, e.g., pre-training, self-supervised and contrastive learning. Model mining as a special way of data mining is relevant to SIGKDD, and its audience including researchers and engineers will benefit a lot for designing more advanced algorithms for their tasks. Shan You, Chang Xu 0002, Fei Wang 0032, Changshui Zhang |
KDD | 3 |
| 2021 | ReSSL: Relational Self-Supervised Learning with Weak AugmentationabstractSelf-supervised Learning (SSL) including the mainstream contrastive learning has achieved great success in learning visual representations without data annotations. However, most of methods mainly focus on the instance level information (\ie, the different augmented images of the same instance should have the same feature or cluster into the same class), but there is a lack of attention on the relationships between different instances. In this paper, we introduced a novel SSL paradigm, which we term as relational self-supervised learning (ReSSL) framework that learns representations by modeling the relationship between different instances. Specifically, our proposed method employs sharpened distribution of pairwise similarities among different instances as \textit{relation} metric, which is thus utilized to match the feature embeddings of different augmentations. Moreover, to boost the performance, we argue that weak augmentations matter to represent a more reliable relation, and leverage momentum strategy for practical efficiency. Experimental results show that our proposed ReSSL significantly outperforms the previous state-of-the-art algorithms in terms of both performance and training efficiency. Mingkai Zheng, Shan You, Fei Wang 0032, Chen Qian 0006, Changshui Zhang, Xiaogang Wang 0001, Chang Xu 0002 |
NeurIPS | 3 |
| 2021 | From Objects to a Whole PaintingabstractStyle image painting is the process of using some stylized strokes to redraw a reference image purposefully and meaningfully. It is a kind of style image generation. In recent years, the application of GAN has greatly improved the quality of generated images for style image generation. However, those methods which use GAN are usually non-serialized. To solve this problem, reinforcement learning and RNN methods are applied to the generation of style images based on strokes. To speed up training, stroke-based style image painting using CNN framework is also proposed. But none of the existing style image painting methods takes into account the content distribution in the reference image, which makes the painting process lack a clear order. We extract the feature of the object contained in the image through the content acquisition module and use the optimal transmission theory to construct a compound loss function to integrate the content information into the painting process. As more advanced macro content information is added to the painting process, our painting method can draw images in a more orderly way. We have conducted a lot of experiments to prove that our method is superior to the current state-of-the-art style image painting method. Fei Wang 0032, Baodi Liu, Weifeng Liu 0001 |
SMC | 1 |
| 2020 | A Coarse-to-Fine Adaptive Network for Appearance-Based Gaze EstimationabstractHuman gaze is essential for various appealing applications. Aiming at more accurate gaze estimation, a series of recent works propose to utilize face and eye images simultaneously. Nevertheless, face and eye images only serve as independent or parallel feature sources in those works, the intrinsic correlation between their features is overlooked. In this paper we make the following contributions: 1) We propose a coarse-to-fine strategy which estimates a basic gaze direction from face image and refines it with corresponding residual predicted from eye images. 2) Guided by the proposed strategy, we design a framework which introduces a bi-gram model to bridge gaze residual and basic gaze direction, and an attention component to adaptively acquire suitable fine-grained feature. 3) Integrating the above innovations, we construct a coarse-to-fine adaptive network named CA-Net and achieve state-of-the-art performances on MPIIGaze and EyeDiap. Yihua Cheng, Shiyao Huang, Fei Wang 0032, Chen Qian 0006, Feng Lu 0005 |
AAAI | 3 |
| 2020 | CentripetalNet: Pursuing High-Quality Keypoint Pairs for Object DetectionabstractKeypoint-based detectors have achieved pretty-well performance. However, incorrect keypoint matching is still widespread and greatly affects the performance of the detector. In this paper, we propose CentripetalNet which uses centripetal shift to pair corner keypoints from the same instance. CentripetalNet predicts the position and the centripetal shift of the corner points and matches corners whose shifted results are aligned. Combining position information, our approach matches corner points more accurately than the conventional embedding approaches do. Corner pooling extracts information inside the bounding boxes onto the border. To make this information more aware at the corners, we design a cross-star deformable convolution network to conduct feature adaption. Furthermore, we explore instance segmentation on anchor-free detectors by equipping our CentripetalNet with a mask prediction module. On COCO test-dev, our CentripetalNet not only outperforms all existing anchor-free detectors with an AP of 48.0% but also achieves comparable performance to the state-of-the-art instance segmentation approaches with a 40.2% Mask AP. Code is available at https: //github.com/KiveeDong/CentripetalNet. Zhiwei Dong, Guoxuan Li, Yue Liao, Fei Wang 0032, Pengju Ren, Chen Qian 0006 |
CVPR | 4 |
| 2020 | A Real-Time Cross-Modality Correlation Filtering Method for Referring Expression ComprehensionabstractReferring expression comprehension aims to localize the object instance described by a natural language expression. Current referring expression methods have achieved good performance. However, none of them is able to achieve real-time inference without accuracy drop. The reason for the relatively slow inference speed is that these methods artificially split the referring expression comprehension into two sequential stages including proposal generation and proposal ranking. It does not exactly conform to the habit of human cognition. To this end, we propose a novel Realtime Cross-modality Correlation Filtering method (RCCF). RCCF reformulates the referring expression comprehension as a correlation filtering process. The expression is first mapped from the language domain to the visual domain and then treated as a template (kernel) to perform correlation filtering on the image feature map. The peak value in the correlation heatmap indicates the center points of the target box. In addition, RCCF also regresses a 2-D object size and 2-D offset. The center point coordinates, object size and center point offset together to form the target bounding box. Our method runs at 40 FPS while achieving leading performance in RefClef, RefCOCO, RefCOCO+ and RefCOCOg benchmarks. In the challenging RefClef dataset, our methods almost double the state-of-the-art performance (34.70% increased to 63.79%). We hope this work can arouse more attention and studies to the new cross-modality correlation filtering framework as well as the one-stage framework for referring expression comprehension. Yue Liao, Si Liu 0001, Guanbin Li, Fei Wang 0032, Chen Qian 0006, Bo Li 0006 |
CVPR | 4 |
| 2020 | PPDM: Parallel Point Detection and Matching for Real-Time Human-Object Interaction DetectionabstractWe propose a single-stage Human-Object Interaction (HOI) detection method that has outperformed all existing methods on HICO-DET dataset at 37 fps on a single Titan XP GPU. It is the first real-time HOI detection method. Conventional HOI detection methods are composed of two stages, i.e., human-object proposals generation, and proposals classification. Their effectiveness and efficiency are limited by the sequential and separate architecture. In this paper, we propose a Parallel Point Detection and Matching (PPDM) HOI detection framework. In PPDM, an HOI is defined as a point triplet. Human and object points are the center of the detection boxes, and the interaction point is the midpoint of the human and object points. PPDM contains two parallel branches, namely point detection branch and point matching branch. The point detection branch predicts three points. Simultaneously, the point matching branch predicts two displacements from the interaction point to its corresponding human and object points. The human point and the object point originated from the same interaction point are considered as matched pairs. In our novel parallel architecture, the interaction points implicitly provide context and regularization for human and object detection. The isolated detection boxes unlikely to form meaningful HOI triplets are suppressed, which increases the precision of HOI detection. Moreover, the matching between human and object detection boxes is only applied around limited numbers of filtered candidate interaction points, which saves much computational cost. Additionally, we build a new application-oriented database named as HOI-A, which serves as a good supplement to the existing datasets. Yue Liao, Si Liu 0001, Fei Wang 0032, Chen Qian 0006, Jiashi Feng |
CVPR | 3 |
| 2020 | GreedyNAS: Towards Fast One-Shot NAS With Greedy SupernetabstractTraining a supernet matters for one-shot neural architecture search (NAS) methods since it serves as a basic performance estimator for different architectures (paths). Current methods mainly hold the assumption that a supernet should give a reasonable ranking over all paths. They thus treat all paths equally, and spare much effort to train paths. However, it is harsh for a single supernet to evaluate accurately on such a huge-scale search space (e.g., 7^21). In this paper, instead of covering all paths, we ease the burden of supernet by encouraging it to focus more on evaluation of those potentially-good ones, which are identified using a surrogate portion of validation data. Concretely, during training, we propose a multi-path sampling strategy with rejection, and greedily filter the weak paths. The training efficiency is thus boosted since the training space has been greedily shrunk from all paths to those potentially-good ones. Moreover, we further adopt an exploration and exploitation policy by introducing an empirical candidate path pool. Our proposed method GreedyNAS is easy-to-follow, and experimental results on ImageNet dataset indicate that it can achieve better Top-1 accuracy under same search space and FLOPs or latency level, but with only ~60% of supernet training cost. By searching on a larger space, our GreedyNAS can also obtain new state-of-the-art architectures. Shan You, Tao Huang 0020, Mingmin Yang, Fei Wang 0032, Chen Qian 0006, Changshui Zhang |
CVPR | 4 |
| 2020 | Local Correlation Consistency for Knowledge Distillation
Jianlong Wu, Hongyu Fang, Yue Liao, Fei Wang 0032, Chen Qian 0006 |
ECCV (12) | 5 |
| 2020 | Agree to Disagree: Adaptive Ensemble Knowledge Distillation in Gradient SpaceabstractDistilling knowledge from an ensemble of teacher models is expected to have a more promising performance than that from a single one. Current methods mainly adopt a vanilla average rule, i.e., to simply take the average of all teacher losses for training the student network. However, this approach treats teachers equally and ignores the diversity among them. When conflicts or competitions exist among teachers, which is common, the inner compromise might hurt the distillation performance. In this paper, we examine the diversity of teacher models in the gradient space and regard the ensemble knowledge distillation as a multi-objective optimization problem so that we can determine a better optimization direction for the training of student network. Besides, we also introduce a tolerance parameter to accommodate disagreement among teachers. In this way, our method can be seen as a dynamic weighting method for each teacher in the ensemble. Extensive experiments validate the effectiveness of our method for both logits-based and feature-based cases. Shangchen Du, Shan You, Jianlong Wu, Fei Wang 0032, Chen Qian 0006, Changshui Zhang |
NeurIPS | 5 |
| 2020 | ISTA-NAS: Efficient and Consistent Neural Architecture Search by Sparse CodingabstractNeural architecture search (NAS) aims to produce the optimal sparse solution from a high-dimensional space spanned by all candidate connections. Current gradient-based NAS methods commonly ignore the constraint of sparsity in the search phase, but project the optimized solution onto a sparse one by post-processing. As a result, the dense super-net for search is inefficient to train and has a gap with the projected architecture for evaluation. In this paper, we formulate neural architecture search as a sparse coding problem. We perform the differentiable search on a compressed lower-dimensional space that has the same validation loss as the original sparse solution space, and recover an architecture by solving the sparse coding problem. The differentiable search and architecture recovery are optimized in an alternate manner. By doing so, our network for search at each update satisfies the sparsity constraint and is efficient to train. In order to also eliminate the depth and width gap between the network in search and the target-net in evaluation, we further propose a method to search and evaluate in one stage under the target-net settings. When training finishes, architecture variables are absorbed into network weights. Thus we get the searched architecture and optimized parameters in a single run. In experiments, our two-stage method on CIFAR-10 requires only 0.05 GPU-day for search. Our one-stage method produces state-of-the-art performances on both CIFAR-10 and ImageNet at the cost of only evaluation time. Shan You, Fei Wang 0032, Chen Qian 0006, Zhouchen Lin |
NeurIPS | 4 |
| 2020 | EEG responses to emotional videos can quantitatively predict big-five personality traits
Wenyu Liu 0001, Xuefei Long, Lilu Tang, Fei Wang 0032, Dan Zhang 0014 |
Neurocomputing | 6 |
| 2019 | Deep Comprehensive Correlation Mining for Image ClusteringabstractRecent developed deep unsupervised methods allow us to jointly learn representation and cluster unlabelled data. These deep clustering methods %like DAC start with mainly focus on the correlation among samples, e.g., selecting high precision pairs to gradually tune the feature representation, which neglects other useful correlations. In this paper, we propose a novel clustering framework, named deep comprehensive correlation mining~(DCCM), for exploring and taking full advantage of various kinds of correlations behind the unlabeled data from three aspects: 1) Instead of only using pair-wise information, pseudo-label supervision is proposed to investigate category information and learn discriminative features. 2) The features' robustness to image transformation of input space is fully explored, which benefits the network learning and significantly improves the performance. 3) The triplet mutual information among features is presented for clustering problem to lift the recently discovered instance-level deep mutual information to a triplet-level formation, which further helps to learn more discriminative features. Extensive experiments on several challenging datasets show that our method achieves good performance, e.g., attaining 62.3% clustering accuracy on CIFAR-10, which is 10.1% higher than the state-of-the-art results. Jianlong Wu, Keyu Long, Fei Wang 0032, Chen Qian 0006, Cheng Li 0009, Zhouchen Lin, Hongbin Zha |
ICCV | 3 |
| 2018 | The Devil of Face Recognition Is in the Noise
Fei Wang 0032, Liren Chen, Cheng Li 0009, Shiyao Huang, Chen Qian 0006, Chen Change Loy |
ECCV (9) | 1 |
| 2018 | What Makes a Champion: The Behavioral and Neural Correlates of Expertise in Multiplayer Online Battle Arena GamesabstractDespite the popularity of multiplayer online battle arena (MOBA) games, academic research on MOBA is still very limited. The current study aimed to fill this gap by exploring the behavioral and neural correlates of expertise for the most popular MOBA game, League of Legends (LOL). Three groups of LOL players with different expertise levels were recruited, including professional players, background-matched trainees, and age-matched students with no systematic LOL trainings. A series of behavioral tests and questionnaires was used to evaluate their general cognitive skills and their LOL-specific abilities were extracted from the neural activities (Electroencephalographs (EEG)s and Electrocardiographs (ECG)s) recorded during LOL matches. Using the behavioral features, both the students and the trainees could be significantly separated from the professional players (trainees vs. professional players, 61.11%; students vs. professional players, 66.67%), whereas the students and the trainees cannot be distinguished. Using the neural features, all three groups could be well separated with higher classification accuracies (students vs. trainees: 88.24%; trainees vs. professional players, 93.33%; students vs. professional players, 93.75%). The most contributing behavioral and neural indices were revealed as well, including multiple-object tracking capability, mental concentration, visuospatial attention ability, etc. The authors’ results for the first time showed the possibility of recognizing MOBA expertise using both behavioral and neural measurements and provided a framework for evaluation, selection, and training of professional MOBA players. Yue Ding 0005, Jingbo Ye, Fei Wang 0032, Dan Zhang 0014 |
Int. J. Hum. Comput. Interact. | 5 |
| 2017 | Residual Attention Network for Image ClassificationabstractIn this work, we propose Residual Attention Network, a convolutional neural network using attention mechanism which can incorporate with state-of-art feed forward network architecture in an end-to-end training fashion. Our Residual Attention Network is built by stacking Attention Modules which generate attention-aware features. The attention-aware features from different modules change adaptively as layers going deeper. Inside each Attention Module, bottom-up top-down feedforward structure is used to unfold the feedforward and feedback attention process into a single feedforward process. Importantly, we propose attention residual learning to train very deep Residual Attention Networks which can be easily scaled up to hundreds of layers. Extensive analyses are conducted on CIFAR-10 and CIFAR-100 datasets to verify the effectiveness of every module mentioned above. Our Residual Attention Network achieves state-of-the-art object recognition performance on three benchmark datasets including CIFAR-10 (3.90% error), CIFAR-100 (20.45% error) and ImageNet (4.8% single model and single crop, top-5 error). Note that, our method achieves 0.6% top-1 accuracy improvement with 46% trunk depth and 69% forward FLOPs comparing to ResNet-200. The experiment also demonstrates that our network is robust against noisy labels. Fei Wang 0032, Mengqing Jiang, Chen Qian 0006, Shuo Yang 0003, Cheng Li 0009, Xiaogang Wang 0001, Xiaoou Tang |
CVPR | 1 |