VLDB 2026 Research / reviewers in the wild / expert
Weifeng Ge
dblp:155/3277
· DBLP profile ↗
36ranked-venue papers
8as first author
28since 2021 · last 2025
0000-0002-6258-6225ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 30 · 7 first-author · 23 since 2021Artificial intelligence and machine learning · 26 · 7 first-author · 19 since 2021Systems, architecture and hardware · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | CocoER: Aligning Multi-Level Feature by Competition and Coordination for Emotion RecognitionabstractWith the explosion of human-machine interaction, emotion recognition has reignited attention. Previous works focus on improving visual feature fusion and reasoning from multiple image levels. Although it is non-trivial to deduce a person’s emotion by integrating multi-level feature (head, body and context), the emotion recognition results of each level is usually different from one another, which creates inconsistency in the prevailing feature alignment method and decrease recognition performance. In this work, we propose a multi-level image feature refinement method for emotion recognition (CocoER) to mitigate the impact caused by conflicting results from multi-level recognition. First, we leverage cross-level attention to improve visual feature consistency between hierarchically cropped head, body and context windows. Then, vocabulary informed alignment is incorporated into the recognition framework to produce pseudo label and guide hierarchical visual feature refinement. To effectively fuse multi-level feature, we elaborate on a competition process of eliminating irrelevant image level predictions and a coordination process to enhance the feature across all levels. Extensive experiments are executed on two popular datasets, and our method achieves state-of-the-art performance with multi-level interpretation results. Code is available at: https://github.com/bisno/CocoER. Xuli Shen, Hua Cai, Weilin Shen, Qing Xu 0017, Dingding Yu, Weifeng Ge, Xiangyang Xue 0001 |
CVPR | 6 |
| 2025 | Synthesizing Near-Boundary OOD Samples for Out-of-Distribution Detection
Kaixun Jiang, Zhaoyu Chen 0001, Bo Li 0115, Weifeng Ge |
ICCV | 6 |
| 2025 | DeFSS: Image-to-Mask Denoising Learning for Few-Shot Segmentation
Zishu Qin, Weifeng Ge |
ICCV | 3 |
| 2025 | GTAD: Global Temporal Aggregation Denoising Learning for 3D Semantic Occupancy PredictionabstractAccurately perceiving dynamic environments is a fundamental task for autonomous driving and robotic systems. Existing methods inadequately utilize temporal information, relying mainly on local temporal interactions between adjacent frames and failing to leverage global sequence information effectively. To address this limitation, we investigate how to effectively aggregate global temporal features from temporal sequences, aiming to achieve occupancy representations that efficiently utilize global temporal information from historical observations. For this purpose, we propose a global temporal aggregation denoising network named GTAD, introducing a global temporal information aggregation framework as a new paradigm for holistic 3D scene understanding. Our method employs an in-model latent denoising network to aggregate local temporal features from the current moment and global temporal features from historical sequences. This approach enables the effective perception of both fine-grained temporal information from adjacent frames and global temporal patterns from historical observations. As a result, it provides a more coherent and comprehensive understanding of the environment. Extensive experiments on the nuScenes and Occ3D-nuScenes benchmark and ablation studies demonstrate the superiority of our method. Yang Li 0041, Yisheng Deng, Weifeng Ge |
IROS | 5 |
| 2025 | StreamBridge: Turning Your Offline Video Large Language Model into a Proactive Streaming AssistantabstractWe present StreamBridge, a simple yet effective framework that seamlessly transforms offline Video-LLMs into streaming-capable models. It addresses two fundamental challenges in adapting existing models into online scenarios: (1) limited capability for multi-turn real-time understanding, and (2) lack of proactive response mechanisms. Specifically, StreamBridge incorporates (1) a memory buffer combined with a round-decayed compression strategy, supporting long-context multi-turn interactions, and (2) a decoupled, lightweight activation model that can be effortlessly integrated into existing Video-LLMs, enabling continuous proactive responses. To further support StreamBridge, we construct Stream-IT, a large-scale dataset tailored for streaming video understanding, featuring interleaved video-text sequences and diverse instruction formats. Extensive experiments show that StreamBridge significantly improves the streaming understanding capabilities of offline Video-LLMs across various tasks, outperforming even proprietary models such as GPT-4o and Gemini 1.5 Pro. Simultaneously, it achieves competitive or superior performance on standard video understanding benchmarks. Zhengfeng Lai, Weifeng Ge, Afshin Dehghan |
NeurIPS | 6 |
| 2025 | Enhancing Environmental Robustness in Few-Shot Learning via Conditional Representation LearningabstractFew-shot learning (FSL) has recently been extensively utilized to overcome the scarcity of training data in domain-specific visual recognition. In real-world scenarios, environmental factors such as complex backgrounds, varying lighting conditions, long-distance shooting, and moving targets often cause test images to exhibit numerous incomplete targets or noise disruptions. However, current research on evaluation datasets and methodologies has largely ignored the concept of "environmental robustness", which refers to maintaining consistent performance in complex and diverse physical environments. This neglect has led to a notable decline in the performance of FSL models during practical testing compared to their training performance. To bridge this gap, we introduce a new real-world multi-domain few-shot learning (RD-FSL) benchmark, which includes four domains and six evaluation datasets. The test images in this benchmark feature various challenging elements, such as camouflaged objects, small targets, and blurriness. Our evaluation experiments reveal that existing methods struggle to utilize training images effectively to generate accurate feature representations for challenging test images. To address this problem, we propose a novel conditional representation learning network (CRLNet) that integrates the interactions between training and testing images as conditional information in their respective representation processes. The main goal is to reduce intra-class variance or enhance inter-class variance at the feature representation level. Finally, comparative experiments reveal that CRLNet surpasses the current state-of-the-art methods, achieving performance improvements ranging from 6.83% to 16.98% across diverse settings and backbones. The source code and dataset are available at https://github.com/guoqianyu-alberta/Conditional-Representation-Learning. Jingrong Wu, Tianxing Wu 0001, Haofen Wang, Weifeng Ge |
IEEE Trans. Image Process. | 5 |
| 2025 | Cross-Modal Complementary Learning and Template-Based Reasoning Chains for Future Event Prediction in VideosabstractAlthough multi-modal large language models (MLLMs) have impressive cross-modal reasoning and prediction capabilities, a unified and rigorous evaluation standard is still lacking. In this paper, we propose a future event prediction task to evaluate their cross-modal temporal prediction capability. This task requires the model to generate descriptions of events that may occur in future based on the input premise video. We build a dataset on the existing datasets for model evaluation. This task faces many challenges, including the complexity of processing video data, such as understanding changes in objects, actions, and time dimensions within the video and the interference of redundant information. To address these challenges, we propose a novel cross-modal prediction framework that introduces cross-modal supplementary learning and template-based reasoning chains based on MLLMs. Cross-modal supplementary learning aims to promote visual and text information to supplement and mine their respective information, primarily to capture critical information in videos, relying on the adaptive temporal filter and casual Q-Former. The template-based reasoning chain drives GPT-4 to generate a series of template question pairs through design prompts, gradually guiding the model to perform hierarchical reasoning to support the final prediction. Through experimental evaluation, the performance of the current MLLMs may not meet the requirements, and our model outperforms all existing models in predicting future events. It shows that the capabilities of MLLMs can be further explored. Chenghang Lai, Weifeng Ge, Xiangyang Xue 0001 |
IEEE Trans. Multim. | 2 |
| 2025 | Adapting Multimodal Large Language Models for Video Question Answering by Capturing Question-Critical and Coherent MomentsabstractMultimodal Large Language Models (MLLMs) have demonstrated remarkable abilities in image-language reasoning. However, they deal with Video Question Answering (VideoQA) insufficiently, especially for questions demanding causal-temporal reasoning. Typically, they directly concatenate features of uniformly sampled frames as visual inputs for VideoQA. This gives rise to two challenges. For one thing, uniformly sampled frames are discrete and separately distributed across different timestamps, disrupting the coherence of question-critical events or actions. For another, it considers every scene within videos equally and introduces redundant frames that may distract the model from discovering the truth. Towards this, we highlight the importance of identifying continuous frames that are crucial for answering the questions, and propose a lightweight and differentiableCoherenceRecognizer (CoRe) to achieve this. Guided by the semantics of questions,CoRecomputes scores recording the relevance between each frame and the question, and selects a set of continuous frames with the highest scores for answer prediction. Additionally,CoReencodes the unselected frames into a short and coarse-grained representation as a completion of the general context. Equipped withCoRe, we can efficiently fine-tune the current MLLMs for VideoQA in an end-to-end manner, without suffering from the problems of incoherence or distraction. Extensive experiments demonstrate that our method achieves substantial improvements on several VideoQA benchmarks. Haibo Wang 0006, Chenghang Lai, Weifeng Ge |
IEEE Trans. Multim. | 3 |
| 2024 | Pixel-Level Semantic Correspondence Through Layout-Aware Representation Learning and Multi-Scale Matching IntegrationabstractEstablishing precise semantic correspondence across object instances in different images is a fundamental and challenging task in computer vision. In this task, difficulty arises often due to three challenges: confusing regions with similar appearance, inconsistent object scale, and indistinguishable nearby pixels. Recognizing these challenges, our paper proposes a novel semantic matching pipeline named LPMFlow toward extracting fine-grained semantics and geometry layouts for building pixel-level semantic correspondences. LPMFlow consists of three modules, each addressing one of the aforementioned challenges. The layout-aware representation learning module uniformly encodes source and target tokens to distinguish pixels or regions with similar appearances but different geometry semantics. The progressive feature superresolution module outputs four sets of 4D correlation tensors to generate accurate semantic flow between objects in different scales. Finally, the matching flow integration and refinement module is exploited to fuse matching flow in different scales to give the final flow predictions. The whole pipeline can be trained end-to-end, with a balance of computational cost and correspondence details. Extensive experiments based on benchmarks such as SPair-71K, PF-PASCAL, and PF-WILLOW have proved that the proposed method can well tackle the three challenges and outperform the previous methods, es-pecially in more stringent settings. Code is available at https://github.com/YXSUNMADMAX/LPMFlow. Yixuan Sun, Zhangyue Yin, Haibo Wang 0006, Yan Wang 0068, Xipeng Qiu, Weifeng Ge |
CVPR | 6 |
| 2024 | Q&A Prompts: Discovering Rich Visual Clues through Mining Question-Answer Prompts for VQA requiring Diverse World Knowledge
Haibo Wang 0006, Weifeng Ge |
ECCV (42) | 2 |
| 2024 | Visual-Language Collaborative Representation Network for Broad-Domain Few-Shot Image ClassificationabstractVisual-language models based on CLIP have shown remarkable abilities in general few-shot image classification. However, their performance drops in specialized fields such as healthcare or agriculture, because CLIP's pre-training does not cover all category data. Existing methods excessively depend on the multi-modal information representation and alignment capabilities acquired from CLIP pre-training, which hinders accurate generalization to unfamiliar domains. To address this issue, this paper introduces a novel visual-language collaborative representation network (MCRNet), aiming at acquiring a generalized capability for collaborative fusion and representation of multi-modal information. Specifically, MCRNet learns to generate relational matrices from an information fusion perspective to acquire aligned multi-modal features. This relationship generation strategy is category-agnostic, so it can be generalized to new domains. A class-adaptive fine-tuning inference technique is also introduced to help MCRNet efficiently learn alignment knowledge for new categories using limited data. Additionally, the paper establishes a new broad-domain few-shot image classification benchmark containing seven evaluation datasets from five domains. Comparative experiments demonstrate that MCRNet outperforms current state-of-the-art models, achieving an average improvement of 13.06% and 13.73% in the 1-shot and 5-shot settings, highlighting the superior performance and applicability of MCRNet across various domains. Jieji Ren, Haofen Wang, Tianxing Wu 0001, Weifeng Ge |
ACM Multimedia | 5 |
| 2024 | TagOOD: A Novel Approach to Out-of-Distribution Detection via Vision-Language Representations and Class Center Learning
Xinyu Zhou 0006, Kaixun Jiang, Lingyi Hong, Pinxue Guo, Zhaoyu Chen 0001, Weifeng Ge |
ACM Multimedia | 7 |
| 2024 | Weakly Supervised Gaussian Contrastive Grounding with Large Multimodal Models for Video Question AnsweringabstractVideo Question Answering (VideoQA) aims to answer natural language questions based on the information observed in videos. Despite the recent success of Large Multimodal Models (LMMs) in image-language understanding, they deal with VideoQA insufficiently, by simply taking uniformly sampled frames as visual inputs, which ignores question-relevant visual clues. Moreover, there are no human annotations for question-critical timestamps in existing VideoQA datasets. In light of this, we propose a weakly supervised framework to enforce the LMMs to reason out the answers with question-critical moments as visual inputs. Specifically, we fuse the question and answer pairs as event descriptions to find multiple keyframes as target moments and pseudo-labels. With these pseudo-labeled keyframes as additionally weak supervision, we devise a lightweight Gaussian-based Contrastive Grounding (GCG) module. GCG learns multiple Gaussian masks to characterize the temporal structure of the video, and sample question-critical frames as positive moments to be the visual inputs of LMMs. Extensive experiments on several benchmarks verify the effectiveness of our framework, and we achieve substantial improvements compared to previous state-of-the-art methods. Haibo Wang 0006, Chenghang Lai, Yixuan Sun, Weifeng Ge |
ACM Multimedia | 4 |
| 2024 | DeTrack: In-model Latent Denoising Learning for Visual Object TrackingabstractPrevious visual object tracking methods employ image-feature regression models or coordinate autoregression models for bounding box prediction. Image-feature regression methods heavily depend on matching results and do not utilize positional prior, while the autoregressive approach can only be trained using bounding boxes available in the training set, potentially resulting in suboptimal performance during testing with unseen data. Inspired by the diffusion model, denoising learning enhances the model’s robustness to unseen data. Therefore, We introduce noise to bounding boxes, generating noisy boxes for training, thus enhancing model robustness on testing data. We propose a new paradigm to formulate the visual object tracking problem as a denoising learning process. However, tracking algorithms are usually asked to run in real-time, directly applying the diffusion model to object tracking would severely impair tracking speed. Therefore, we decompose the denoising learning process into every denoising block within a model, not by running the model multiple times, and thus we summarize the proposed paradigm as an in-model latent denoising learning process. Specifically, we propose a denoising Vision Transformer (ViT), which is composed of multiple denoising blocks. In the denoising block, template and search embeddings are projected into every denoising block as conditions. A denoising block is responsible for removing the noise in a predicted bounding box, and multiple stacked denoising blocks cooperate to accomplish the whole denoising process. Subsequently, we
utilize image features and trajectory information to refine the denoised bounding box. Besides, we also utilize trajectory memory and visual memory to improve tracking stability. Experimental results validate the effectiveness of our approach, achieving competitive performance on several challenging datasets. The proposed in-model latent denoising tracker achieve real-time speed, rendering denoising learning applicable in the visual object tracking community. Xinyu Zhou 0006, Lingyi Hong, Kaixun Jiang, Pinxue Guo, Weifeng Ge |
NeurIPS | 6 |
| 2024 | Object-Centric Cross-Modal Knowledge Reasoning for Future Event Prediction in VideosabstractAlthough multi-modal large language models possess impressive cross-modal reasoning and prediction capabilities, they lack a unified and rigorous evaluation standard. In this paper, we introduce a future event prediction task to assess the cross-modal temporal prediction capabilities of these models. This task requires the model to generate descriptions of events that may occur in the future based on input video. To tackle this new task, we propose an object-centric cross-modal knowledge reasoning framework, which combines a basic information encoder, an adaptive multi-segment filter, a spatial-temporal relation encoder, a vision-text interaction module, and a pre-trained large language model decoder. The adaptive multi-segment filter captures selectively capture critical visual information in videos, enhancing the model’s focus on relevant features. The spatial-temporal relation encoder decomposes and associates the objects and scene information in the video. Additionally, the vision-text interaction module enhances the connection between visual sequences and their corresponding textual narratives, ensuring semantic coherence and consistency. To evaluate our framework, we constructed a dataset containing descriptions, dialogues of future events, and object-centric event reasoning chains. Experimental results indicate that the proposed framework outperforms all previous methods for future event prediction. Ablation studies further demonstrate the effectiveness of the designed modules. Chenghang Lai, Haibo Wang 0006, Weifeng Ge, Xiangyang Xue 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2023 | RankDNN: Learning to Rank for Few-Shot LearningabstractThis paper introduces a new few-shot learning pipeline that casts relevance ranking for image retrieval as binary ranking relation classification. In comparison to image classification, ranking relation classification is sample efficient and domain agnostic. Besides, it provides a new perspective on few-shot learning and is complementary to state-of-the-art methods. The core component of our deep neural network is a simple MLP, which takes as input an image triplet encoded as the difference between two vector-Kronecker products, and outputs a binary relevance ranking order. The proposed RankMLP can be built on top of any state-of-the-art feature extractors, and our entire deep neural network is called the ranking deep neural network, or RankDNN. Meanwhile, RankDNN can be flexibly fused with other post-processing methods. During the meta test, RankDNN ranks support images according to their similarity with the query samples, and each query sample is assigned the class label of its nearest neighbor. Experiments demonstrate that RankDNN can effectively improve the performance of its baselines based on a variety of backbones and it outperforms previous state-of-the-art algorithms on multiple few-shot learning benchmarks, including miniImageNet, tieredImageNet, Caltech-UCSD Birds, and CIFAR-FS. Furthermore, experiments on the cross-domain challenge demonstrate the superior transferability of RankDNN.The code is available at: https://github.com/guoqianyu-alberta/RankDNN. Haotong Gong, Xujun Wei, Yanwei Fu 0001, Yizhou Yu, Weifeng Ge |
AAAI | 7 |
| 2023 | MISC210K: A Large-Scale Dataset for Multi-Instance Semantic CorrespondenceabstractSemantic correspondence have built up a new way for object recognition. However current single-object matching schema can be hard for discovering commonalities for a category and far from the real-world recognition tasks. To fill this gap, we design the multi-instance semantic correspondence task which aims at constructing the correspondence between multiple objects in an image pair. To support this task, we build a multi-instance semantic correspondence (MISC) dataset from COCO Detection 2017 task called MISC210K. We construct our dataset as three steps: (1) category selection and data cleaning; (2) keypoint design based on 3D models and object description rules; (3) human-machine collaborative annotation. Following these steps, we select 34 classes of objects with 4,812 challenging images annotated via a well designed semi-automatic workflow, and finally acquire 218,179 image pairs with instance masks and instance-level keypoint pairs annotated. We design a dual-path collaborative learning pipeline to train instance-level co-segmentation task and fine-grained level correspondence task together. Benchmark evaluation and further ablation results with detailed analysis are provided with three future directions proposed. Our project is available on https://github.com/YXSUNMADMAX/MISC210K. Yixuan Sun, Haijing Guo, Yuzhou Zhao, Runmin Wu, Yizhou Yu, Weifeng Ge |
CVPR | 7 |
| 2023 | Correspondence Transformers with Asymmetric Feature Learning and Matching Flow Super-ResolutionabstractThis paper solves the problem of learning dense visual correspondences between different object instances of the same category with only sparse annotations. We decompose this pixel-level semantic matching problem into two easier ones: (i) First, local feature descriptors of source and target images need to be mapped into shared semantic spaces to get coarse matching flows. (ii) Second, matching flows in low resolution should be refined to generate accurate point-to-point matching results. We propose asymmetric feature learning and matching flow super-resolution based on vision transformers to solve the above problems. The asymmetric feature learning module exploits a biased cross-attention mechanism to encode token features of source images with their target counterparts. Then matching flow in low resolutions is enhanced by a super-resolution network to get accurate correspondences. Our pipeline is built upon vision transformers and can be trained in an end-to-end manner. Extensive experimental results on several popular benchmarks, such as PF-PASCAL, PF-WILLOW, and SPair-71 K, demonstrate that the proposed method can catch subtle semantic differences in pixels efficiently. Code is available on https://github.com/YXSUNMADMAX/ACTR. Yixuan Sun, Dongyang Zhao, Zhangyue Yin, Tao Gui, Weifeng Ge |
CVPR | 7 |
| 2023 | Weakly Supervised Learning of Semantic Correspondence through Cascaded Online Correspondence RefinementabstractIn this paper, we develop a weakly supervised learning algorithm to learn robust semantic correspondences from large-scale datasets with only image-level labels. Following the spirit of multiple instance learning (MIL), we decompose the weakly supervised correspondence learning problem into three stages: image-level matching, region-level matching, and pixel-level matching. We propose a novel cascaded online correspondence refinement algorithm to integrate MIL and the correspondence filtering and refinement procedure into a single deep network and train this network end-to-end with only image-level supervision, i.e., without point-to-point matching information. During the correspondence learning process, pixel-to-pixel matching pairs inferred from weak supervision are propagated, filtered, and enhanced through masked correspondence voting and calibration. Besides, we design a correspondence consistency check algorithm to select images with discriminative key points to generate pseudo-labels for classical matching algorithms. Finally, we filter out about 110,000 images from the ImageNet ILSVRC training set to formulate a new dataset, called SC-ImageNet. Experiments on several popular benchmarks indicate that pre-training on SC-ImageNet can improve the performance of state-ofthe-art algorithms efficiently. Our project is available on https://github.com/21210240056/SC-ImageNet. Yixuan Sun, Chenghang Lai, Qing Xu 0017, Xuli Shen, Weifeng Ge |
ICCV | 7 |
| 2023 | Hierarchical Visual Categories Modeling: A Joint Representation Learning and Density Estimation Framework for Out-of-Distribution DetectionabstractDetecting out-of-distribution inputs for visual recognition models has become critical in safe deep learning. This paper proposes a novel hierarchical visual category modeling scheme to separate out-of-distribution data from in-distribution data through joint representation learning and statistical modeling. We learn a mixture of Gaussian models for each in-distribution category. There are many Gaussian mixture models to model different visual categories. With these Gaussian models, we design an in-distribution score function by aggregating multiple Mahalanobis-based metrics. We don’t use any auxiliary outlier data as training samples, which may hurt the generalization ability of out-of-distribution detection algorithms. We split the ImageNet-1k dataset into ten folds randomly. We use one fold as the in-distribution dataset and the others as out-of-distribution datasets to evaluate the proposed method. We also conduct experiments on seven popular benchmarks, including CIFAR, iNaturalist, SUN, Places, Textures, ImageNet-O, and OpenImage-O. Extensive experiments indicate that the proposed method outperforms state-of-the-art algorithms clearly. Meanwhile, we find that our visual representation has a competitive performance when compared with features learned by classical methods. These results demonstrate that the proposed method hasn’t weakened the discriminative ability of visual recognition models and keeps high efficiency in detecting out-of-distribution samples. Xinyu Zhou 0006, Pinxue Guo, Yixuan Sun, Weifeng Ge |
ICCV | 6 |
| 2023 | Reading Relevant Feature from Global Representation Memory for Visual Object TrackingabstractReference features from a template or historical frames are crucial for visual object tracking. Prior works utilize all features from a fixed template or memory for visual object tracking. However, due to the dynamic nature of videos, the required reference historical information for different search regions at different time steps is also inconsistent. Therefore, using all features in the template and memory can lead to redundancy and impair tracking performance. To alleviate this issue, we propose a novel tracking paradigm, consisting of a relevance attention mechanism and a global representation memory, which can adaptively assist the search region in selecting the most relevant historical information from reference features. Specifically, the proposed relevance attention mechanism in this work differs from previous approaches in that it can dynamically choose and build the optimal global representation memory for the current frame by accessing cross-
frame information globally. Moreover, it can flexibly read the relevant historical information from the constructed memory to reduce redundancy and counteract the negative effects of harmful information. Extensive experiments validate the effectiveness of the proposed method, achieving competitive performance on five challenging datasets with 71 FPS. Xinyu Zhou 0006, Pinxue Guo, Lingyi Hong, Wei Zhang 0016, Weifeng Ge |
NeurIPS | 6 |
| 2023 | Multi-domain Information Fusion for Key-Points Guided GAN Inversion
Ruize Xu, Xiaowen Qiu, Boan He, Weifeng Ge |
PRCV (11) | 4 |
| 2022 | Attribute Surrogates Learning and Spectral Tokens Pooling in Transformers for Few-shot LearningabstractThis paper presents new hierarchically cascaded transformers that can improve data efficiency through attribute surrogates learning and spectral tokens pooling. Vision transformers have recently been thought of as a promising alternative to convolutional neural networks for visual recognition. But when there is no sufficient data, it gets stuck in overfitting and shows inferior performance. To improve data efficiency, we propose hierarchically cascaded transformers that exploit intrinsic image structures through spectral tokens pooling and optimize the learnable parameters through latent attribute surrogates. The intrinsic image structure is utilized to reduce the ambiguity between foreground content and background noise by spectral tokens pooling. And the attribute surrogate learning scheme is designed to benefit from the rich visual information in image-label pairs instead of simple visual concepts assigned by their labels. Our Hierarchically Cascaded Transformers, called HCTransformers, is built upon a self-supervised learning framework DINO and is tested on several popular few-shot learning benchmarks. In the inductive setting, HCTransformers surpass the DINO baseline by a large margin of 9.7% 5-way 1-shot accuracy and 9.17% 5-way 5-shot accuracy on miniImageNet, which demonstrates HCTransformers are efficient to extract discriminative features. Also, HCTransformers show clear advantages over SOTA few-shot classification methods in both 5-way 1-shot and 5-way 5-shot settings on four popular benchmark datasets, including miniImageNet, tieredImageNet, FC100, and CIFAR-FS. The trained weights and codes are available at https://github.com/StomachCold/HCTransformers. Yangji He, Weihan Liang, Dongyang Zhao, Weifeng Ge, Yizhou Yu |
CVPR | 5 |
| 2022 | FERV39k: A Large-Scale Multi-Scene Dataset for Facial Expression Recognition in VideosabstractCurrent benchmarks for facial expression recognition (FER) mainly focus on static images, while there are limited datasets for FER in videos. It is still ambiguous to evaluate whether performances of existing methods remain satisfactory in real-world application-oriented scenes. For example, the “Happy” expression with high intensity in Talk-Show is more discriminating than the same expression with low intensity in Official-Event. To fill this gap, we build a large-scale multi-scene dataset, coined as FERV39k. We analyze the important ingredients of constructing such a novel dataset in three aspects: (1) multi-scene hierarchy and expression class, (2) generation of candidate video clips, (3) trusted manual labelling process. Based on these guidelines, we select 4 scenarios subdivided into 22 scenes, annotate 86k samples automatically obtained from 4k videos based on the well-designed workflow, and finally build 38,935 video clips labeled with 7 classic expressions. Experiment benchmarks on four kinds of baseline frame-works were also provided and further analysis on their performance across different scenes and some challenges for future research were given. Besides, we systematically investigate key components of DFER by ablation studies. The baseline framework and our project are available on https://github.com/wangyanckxx/FERV39k. Yan Wang 0068, Yixuan Sun, Zhongying Liu, Shuyong Gao, Wei Zhang 0016, Weifeng Ge |
CVPR | 7 |
| 2022 | Towards Scalable and Fast Distributionally Robust Optimization for Data-Driven Deep LearningabstractWe introduce a scalable and fast method for solving distributionally robust optimization (DRO). Previous works have demonstrated that DRO outperforms empirical risk on a collection of inconsistent distribution of test data (the property of “uncertainty set”). However, DRO is hard to be applied for large-scale datasets and large parameterized model, due to the datapoint-level and non-differentiable objective function. In this paper, we formalize the DRO problem with the supremum of a family of subgroup-level loss functions. Subgroup loss is the cost function of partitioned uncertainty set. Then we implement the maximum of subgroup loss as the objective function and update model parameters by reweighting the descent direction, calculated from a differentiable objective function. Experimental results unveil that large parameterized models with the proposed method successfully adapt to uncertainty set whether the distribution contains out-of-domain or imbalanced property. Remarkably, with the explored reweighting strategy, the proposed algorithm effectively achieves competitive performance and robustness. Xuli Shen, Qing Xu 0017, Weifeng Ge, Xiangyang Xue 0001 |
ICDM | 4 |
| 2022 | DPCNet: Dual Path Multi-Excitation Collaborative Network for Facial Expression Representation Learning in VideosabstractCurrent works of facial expression learning in video consume significant computational resources to learn spatial channel feature representations and temporal relationships. To mitigate this issue, we propose a Dual Path multi-excitation Collaborative Network (DPCNet) to learn the critical information for facial expression representation from fewer keyframes in videos. Specifically, the DPCNet learns the important regions and keyframes from a tuple of four view-grouped frames by multi-excitation modules and produces dual-path representations of one video with consistency under two regularization strategies. A spatial-frame excitation module and a channel-temporal aggregation module are introduced consecutively to learn spatial-frame representation and generate complementary channel-temporal aggregation, respectively. Moreover, we design a multi-frame regularization loss to enforce the representation of multiple frames in the dual view to be semantically coherent. To obtain consistent prediction probabilities from the dual path, we further propose a dual path regularization loss, aiming to minimize the divergence between the distributions of two-path embeddings. Extensive experiments and ablation studies show that the DPCNet can significantly improve the performance of video-based FER and achieve state-of-the-art results on the large-scale DFEW dataset. Yan Wang 0068, Yixuan Sun, Wei Song 0007, Shuyong Gao, Zhaoyu Chen 0001, Weifeng Ge |
ACM Multimedia | 7 |
| 2021 | GraphFPN: Graph Feature Pyramid Network for Object DetectionabstractFeature pyramids have been proven powerful in image understanding tasks that require multi-scale features. State-of-the-art methods for multi-scale feature learning focus on performing feature interactions across space and scales using neural networks with a fixed topology. In this paper, we propose graph feature pyramid networks that are capable of adapting their topological structures to varying intrinsic image structures, and supporting simultaneous feature interactions across all scales. We first define an image specific superpixel hierarchy for each input image to represent its intrinsic image structures. The graph feature pyramid network inherits its structure from this superpixel hierarchy. Contextual and hierarchical layers are designed to achieve feature interactions within the same scale and across different scales. To make these layers more powerful, we introduce two types of local channel attention for graph neural networks by generalizing global channel attention for convolutional neural networks. The proposed graph feature pyramid network can enhance the multiscale features from a convolutional feature pyramid network.We evaluate our graph feature pyramid network in the object detection task by integrating it into the Faster R-CNN algorithm. The modified algorithm outperforms not only previous state-of-the-art feature pyramid based methods with a clear margin but also other popular detection methods on both MS-COCO 2017 validation and test datasets. Gangming Zhao, Weifeng Ge, Yizhou Yu |
ICCV | 2 |
| 2021 | Multi-scale Matching Networks for Semantic CorrespondenceabstractDeep features have been proven powerful in building accurate dense semantic correspondences in various previous works. However, the multi-scale and pyramidal hierarchy of convolutional neural networks has not been well studied to learn discriminative pixel-level features for semantic correspondence. In this paper, we propose a multi-scale matching network that is sensitive to tiny semantic differences between neighboring pixels. We follow the coarse-to-fine matching strategy and build a top-down feature and matching enhancement scheme that is coupled with the multi-scale hierarchy of deep convolutional neural networks. During feature enhancement, intra-scale enhancement fuses same-resolution feature maps from multiple layers together via local self-attention and cross-scale enhancement hallucinates higher-resolution feature maps along the top-down pathway. Besides, we learn complementary matching details at different scales thus the overall matching score is refined by features of different semantic levels gradually. Our multi-scale matching network can be trained end-to-end easily with few additional learnable parameters. Experimental results demonstrate that the proposed method achieves state-of-the-art performance on three popular benchmarks with high computational efficiency. The code has been released at https://github.com/wintersun661/MMNet. Dongyang Zhao, Zhenghao Ji, Gangming Zhao, Weifeng Ge, Yizhou Yu |
ICCV | 5 |
| 2019 | Weakly Supervised Complementary Parts Models for Fine-Grained Image Classification From the Bottom UpabstractGiven a training dataset composed of images and corresponding category labels, deep convolutional neural networks show a strong ability in mining discriminative parts for image classification. However, deep convolutional neural networks trained with image level labels only tend to focus on the most discriminative parts while missing other object parts, which could provide complementary information. In this paper, we approach this problem from a different perspective. We build complementary parts models in a weakly supervised manner to retrieve information suppressed by dominant object parts detected by convolutional neural networks. Given image level labels only, we first extract rough object instances by performing weakly supervised object detection and instance segmentation using Mask R-CNN and CRF-based segmentation. Then we estimate and search for the best parts model for each object instance under the principle of preserving as much diversity as possible. In the last stage, we build a bi-directional long short-term memory (LSTM) network to fuze and encode the partial information of these complementary parts into a comprehensive feature for image classification. Experimental results indicate that the proposed method not only achieves significant improvement over our baseline models, but also outperforms state-of-the-art algorithms by a large margin (6.7%, 2.8%, 5.2% respectively) on Stanford Dogs 120, Caltech-UCSD Birds 2011-200 and Caltech 256. Weifeng Ge, Xiangru Lin, Yizhou Yu |
CVPR | 1 |
| 2019 | Label-PEnet: Sequential Label Propagation and Enhancement Networks for Weakly Supervised Instance SegmentationabstractWeakly-supervised instance segmentation aims to detect and segment object instances precisely, given image-level labels only. Unlike previous methods which are composed of multiple offline stages, we propose Sequential Label Propagation and Enhancement Networks (referred as Label-PEnet) that progressively transforms image-level labels to pixel-wise labels in a coarse-to-fine manner. We design four cascaded modules including multi-label classification, object detection, instance refinement and instance segmentation, which are implemented sequentially by sharing the same backbone. The cascaded pipeline is trained alternatively with a curriculum learning strategy that generalizes labels from high level images to low-level pixels gradually with increasing accuracy. In addition, we design a proposal calibration module to explore the ability of classification networks to find key pixels that identify object parts, which serves as a post validation strategy running in the inverse order. We evaluate the efficiency of our Label-PEnet in mining instance masks on standard benchmarks: PASCAL VOC 2007 and 2012. Experimental results show that Label-PEnet outperforms the state-of-art algorithms by a clear margin, and obtains comparable performance even with fully supervised approaches. Weifeng Ge, Sheng Guo 0005, Matthew R. Scott |
ICCV | 1 |
| 2018 | Multi-Evidence Filtering and Fusion for Multi-Label Classification, Object Detection and Semantic Segmentation Based on Weakly Supervised LearningabstractSupervised object detection and semantic segmentation require object or even pixel level annotations. When there exist image level labels only, it is challenging for weakly supervised algorithms to achieve accurate predictions. The accuracy achieved by top weakly supervised algorithms is still significantly lower than their fully supervised counterparts. In this paper, we propose a novel weakly supervised curriculum learning pipeline for multi-label object recognition, detection and semantic segmentation. In this pipeline, we first obtain intermediate object localization and pixel labeling results for the training images, and then use such results to train task-specific deep networks in a fully supervised manner. The entire process consists of four stages, including object localization in the training images, filtering and fusing object instances, pixel labeling for the training images, and task-specific network training. To obtain clean object instances in the training images, we propose a novel algorithm for filtering, fusing and classifying object instances collected from multiple solution mechanisms. In this algorithm, we incorporate both metric learning and density-based clustering to filter detected object instances. Experiments show that our weakly supervised pipeline achieves state-of-the-art results in multi-label image classification as well as weakly supervised object detection and very competitive results in weakly supervised semantic segmentation on MS-COCO, PASCAL VOC 2007 and PASCAL VOC 2012. Weifeng Ge, Sibei Yang, Yizhou Yu |
CVPR | 1 |
| 2018 | Deep Metric Learning with Hierarchical Triplet Loss
Weifeng Ge, Dengke Dong, Matthew R. Scott |
ECCV (6) | 1 |
| 2018 | Image super-resolution via deterministic-stochastic synthesis and local statistical rectificationabstractSingle image superresolution has been a popular research topic in the last two decades and has recently received a new wave of interest due to deep neural networks. In this paper, we approach this problem from a different perspective. With respect to a downsampled low resolution image, we model a high resolution image as a combination of two components, a deterministic component and a stochastic component. The deterministic component can be recovered from the low-frequency signals in the downsampled image. The stochastic component, on the other hand, contains the signals that have little correlation with the low resolution image. We adopt two complementary methods for generating these two components. While generative adversarial networks are used for the stochastic component, deterministic component reconstruction is formulated as a regression problem solved using deep neural networks. Since the deterministic component exhibits clearer local orientations, we design novel loss functions tailored for such properties for training the deep regression network. These two methods are first applied to the entire input image to produce two distinct high-resolution images. Afterwards, these two images are fused together using another deep neural network that also performs local statistical rectification, which tries to make the local statistics of the fused image match the same local statistics of the groundtruth image. Quantitative results and a user study indicate that the proposed method outperforms existing state-of-the-art algorithms with a clear margin. Weifeng Ge, Bingchen Gong, Yizhou Yu |
ACM Trans. Graph. | 1 |
| 2017 | Borrowing Treasures from the Wealthy: Deep Transfer Learning through Selective Joint Fine-TuningabstractDeep neural networks require a large amount of labeled training data during supervised learning. However, collecting and labeling so much data might be infeasible in many cases. In this paper, we introduce a deep transfer learning scheme, called selective joint fine-tuning, for improving the performance of deep learning tasks with insufficient training data. In this scheme, a target learning task with insufficient training data is carried out simultaneously with another source learning task with abundant training data. However, the source learning task does not use all existing training data. Our core idea is to identify and use a subset of training images from the original source learning task whose low-level characteristics are similar to those from the target learning task, and jointly fine-tune shared convolutional layers for both tasks. Specifically, we compute descriptors from linear or nonlinear filter bank responses on training images from both tasks, and use such descriptors to search for a desired subset of training samples for the source learning task. Experiments demonstrate that our deep transfer learning scheme achieves state-of-the-art performance on multiple visual classification tasks with insufficient training data for deep learning. Such tasks include Caltech 256, MIT Indoor 67, and fine-grained classification problems (Oxford Flowers 102 and Stanford Dogs 120). In comparison to fine-tuning without a source domain, the proposed method can improve the classification accuracy by 2% - 10% using a single model. Codes and models are available at https://github.com/ZYYSzj/Selective-Joint-Fine-tuning. Weifeng Ge, Yizhou Yu |
CVPR | 1 |
| 2016 | Dynamic background estimation and complementary learning for pixel-wise foreground/background segmentation
Weifeng Ge, Zhenhua Guo 0001, Yuhan Dong, Youbin Chen |
Pattern Recognit. | 1 |
| 2014 | Background Subtraction with Dynamic Noise Sampling and Complementary LearningabstractBackground subtraction is a popular technique used in accurate foreground extraction with a stationary background. Since most outdoor surveillance videos are taken in complex environments, their "stationary" backgrounds change in some unknown patterns, which make the perfect foreground extraction very difficult. Based on visual background extractor (ViBe) scheme, in this paper we propose a new background subtraction algorithm which includes two innovative mechanisms and several other improved technique tricks. The paper inherits and develops background modeling based on pixel sample values, and use dynamic noise sampling and complementary learning to overcome the pixel-wise background model's intrinsic shortcomings. Besides, the algorithm works on the quantitative analysis without any estimation of the probability density function (pdf). Hence, it takes relatively low computational cost. Extensive experiments on a popular public dataset show that the proposed method has much better precision than ViBe, and could get the best precision and the highest average ranking compared with 27 state-of-the-art algorithms presented on the change detection website. Weifeng Ge, Yuhan Dong, Zhenhua Guo 0001, Youbin Chen |
ICPR | 1 |