Chunfeng Song

dblp:19/7485 · DBLP profile ↗
← Back
39ranked-venue papers
8as first author
22since 2021 · last 2025
0000-0003-1223-3242ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 33 · 8 first-author · 17 since 2021Graphics, computer vision, multimedia, augmented reality and games · 16 · 3 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Computer networks · 1 · 1 since 2021Security and privacy · 1
YearPublicationVenuePosition
2025 Multi-Modal Latent Variables for Cross-Individual Primary Visual Cortex Modeling and Analysis
abstract
Elucidating the functional mechanisms of the primary visual cortex (V1) remains a fundamental challenge in systems neuroscience. Current computational models face two critical limitations, namely the challenge of cross-modal integration between partial neural recordings and complex visual stimuli, and the inherent variability in neural characteristics across individuals, including differences in neuron populations and firing patterns. To address these challenges, we present a multi-modal identifiable variational autoencoder (miVAE) that employs a two-level disentanglement strategy to map neural activity and visual stimuli into a unified latent space. This framework enables robust identification of cross-modal correlations through refined latent space modeling. We complement this with a novel score-based attribution analysis that traces latent variables back to their origins in the source data space. Evaluation on a large-scale mouse V1 dataset demonstrates that our method achieves state-of-the-art performance in cross-individual latent representation and alignment, without requiring subject-specific fine-tuning, and exhibits improved performance with increasing data size. Significantly, our attribution algorithm successfully identifies distinct neuronal subpopulations characterized by unique temporal patterns and stimulus discrimination properties, while simultaneously revealing stimulus regions that show specific sensitivity to edge features and luminance variations. This scalable framework offers promising applications not only for advancing V1 research but also for broader investigations in neuroscience.
Yu Zhu 0008, Chunfeng Song, Wanli Ouyang, Tiejun Huang 0001
AAAI3
2025 Neuro-3D: Towards 3D Visual Decoding from EEG Signals
abstract
Human’s perception of the visual world is shaped by the stereo processing of 3D information. Understanding how the brain perceives and processes 3D visual stimuli in the real world has been a longstanding endeavor in neuroscience. Towards this goal, we introduce a new neuroscience task: decoding 3D visual perception from EEG signals, a neuroimaging technique that enables real-time monitoring of neural dynamics enriched with complex visual cues. To provide the essential benchmark, we first present EEG-3D, a pioneering dataset featuring multimodal analysis data and extensive EEG recordings from 12 subjects viewing 72 categories of 3D objects rendered in both videos and images. Furthermore, we propose Neuro-3D, a 3D visual decoding framework based on EEG signals. This framework adaptively integrates EEG features derived from static and dynamic stimuli to learn complementary and robust neural representations, which are subsequently utilized to recover both the shape and color of 3D objects through the proposed diffusion-based colored point cloud decoder. To the best of our knowledge, we are the first to explore EEG-based 3D visual decoding. Experiments indicate that Neuro-3D not only reconstructs colored 3D objects with high fidelity, but also learns effective neural representations that enable insightful brain region analysis. The code and dataset are available at https://github.com/gzq17/neuro-3D.
Zhanqiang Guo, Yonghao Song, Jiahui Bu, Weijian Mai, Qihao Zheng, Wanli Ouyang, Chunfeng Song
CVPR8
2025 MindAligner: Explicit Brain Functional Alignment for Cross-Subject Visual Decoding from Limited fMRI Data
abstract
Brain decoding aims to reconstruct visual perception of human subject from fMRI signals, which is crucial for understanding brain’s perception mechanisms. Existing methods are confined to the single-subject paradigm due to substantial brain variability, which leads to weak generalization across individuals and incurs high training costs, exacerbated by limited availability of fMRI data. To address these challenges, we propose MindAligner, an explicit functional alignment framework for cross-subject brain decoding from limited fMRI data. The proposed MindAligner enjoys several merits. First, we learn a Brain Transfer Matrix (BTM) that projects the brain signals of an arbitrary new subject to one of the known subjects, enabling seamless use of pre-trained decoding models. Second, to facilitate reliable BTM learning, a Brain Functional Alignment module is proposed to perform soft cross-subject brain alignment under different visual stimuli with a multi-level brain alignment loss, uncovering fine-grained functional correspondences with high interpretability. Experiments indicate that MindAligner not only outperforms existing methods in visual decoding under data-limited conditions, but also provides valuable neuroscience insights in cross-subject functional analysis. The code will be made publicly available.
Yuqin Dai, Zhouheng Yao, Chunfeng Song, Qihao Zheng, Weijian Mai, Kunyu Peng, Wanli Ouyang, Jian Yang 0003
ICML3
2025 Neural Representational Consistency Emerges from Probabilistic Neural-Behavioral Representation Alignment
abstract
Individual brains exhibit striking structural and physiological heterogeneity, yet neural circuits can generate remarkably consistent functional properties across individuals, an apparent paradox in neuroscience. While recent studies have observed preserved neural representations in motor cortex through manual alignment across subjects, the zero-shot validation of such preservation and its generalization to more cortices remain unexplored. Here we present PNBA (Probabilistic Neural-Behavioral Representation Alignment), a new framework that leverages probabilistic modeling to address hierarchical variability across trials, sessions, and subjects, with generative constraints preventing representation degeneration. By establishing reliable cross-modal representational alignment, PNBA reveals robust preserved neural representations in monkey primary motor cortex (M1) and dorsal premotor cortex (PMd) through zero-shot validation. We further establish similar representational preservation in mouse primary visual cortex (V1), reflecting a general neural basis. These findings resolve the paradox of neural heterogeneity by establishing zero-shot preserved neural representations across cortices and species, enriching neural coding insights and enabling zero-shot behavior decoding.
Yu Zhu 0008, Chunfeng Song, Wanli Ouyang, Tiejun Huang 0001
ICML2
2025 PaceLLM: Brain-Inspired Large Language Models for Long-Context Understanding
abstract
While Large Language Models (LLMs) demonstrate strong performance across domains, their long-context capabilities are limited by transient neural activations causing information decay and unstructured feed-forward network (FFN) weights leading to semantic fragmentation. Inspired by the brain’s working memory and cortical modularity, we propose PaceLLM, featuring two innovations: (1) a Persistent Activity (PA) Mechanism that mimics prefrontal cortex (PFC) neurons’ persistent firing by introducing an activation-level memory bank to dynamically retrieve, reuse, and update critical FFN states, addressing contextual decay; and (2) Cortical Expert (CE) Clustering that emulates task-adaptive neural specialization to reorganize FFN weights into semantic modules, establishing cross-token dependencies and mitigating fragmentation. Extensive evaluations show that PaceLLM achieves 6% improvement on LongBench’s Multi-document QA and 12.5–17.5% performance gains on $\infty$-Bench tasks, while extending measurable context length to 200K tokens in Needle-In-A-Haystack (NIAH) tests. This work pioneers brain-inspired LLM optimization and is complementary to other works. Besides, it can be generalized to any model and enhance their long-context performance and interpretability without structural overhauls.
Kangcong Li, Peng Ye 0006, Chongjun Tu, Lin Zhang 0055, Chunfeng Song, Qihao Zheng, Tao Chen 0003
NeurIPS5
2025 SynBrain: Enhancing Visual-to-fMRI Synthesis via Probabilistic Representation Learning
abstract
Deciphering how visual stimuli are transformed into cortical responses is a fundamental challenge in computational neuroscience. This visual-to-neural mapping is inherently a one-to-many relationship, as identical visual inputs reliably evoke variable hemodynamic responses across trials, contexts, and subjects. However, existing deterministic methods struggle to simultaneously model this biological variability while capturing the underlying functional consistency that encodes stimulus information. To address these limitations, we propose SynBrain, a generative framework that simulates the transformation from visual semantics to neural responses in a probabilistic and biologically interpretable manner. SynBrain introduces two key components: (i) BrainVAE models neural representations as continuous probability distributions via probabilistic learning while maintaining functional consistency through visual semantic constraints; (ii) A Semantic-to-Neural Mapper acts as a semantic transmission pathway, projecting visual semantics into the neural response manifold to facilitate high-fidelity fMRI synthesis. Experimental results demonstrate that SynBrain surpasses state-of-the-art methods in subject-specific visual-to-fMRI encoding performance. Furthermore, SynBrain adapts efficiently to new subjects with few-shot data and synthesizes high-quality fMRI signals that are effective in improving data-limited fMRI-to-image decoding performance. Beyond that, SynBrain reveals functional consistency across trials and subjects, with synthesized signals capturing interpretable patterns shaped by biological neural variability. Our code is available at https://github.com/MichaelMaiii/SynBrain.
Weijian Mai, Yu Zhu 0008, Zhouheng Yao, Dongzhan Zhou, Andrew Luo 0001, Qihao Zheng, Wanli Ouyang, Chunfeng Song
NeurIPS9
2025 CSBrain: A Cross-scale Spatiotemporal Brain Foundation Model for EEG Decoding
abstract
Understanding and decoding human brain activity from electroencephalography (EEG) signals is a fundamental problem in neuroscience and artificial intelligence, with applications ranging from cognition and emotion recognition to clinical diagnosis and brain–computer interfaces. While recent EEG foundation models have made progress in generalized brain decoding by leveraging unified architectures and large-scale pretraining, they inherit a scale-agnostic dense modeling paradigm from NLP and vision. This design overlooks an intrinsic property of neural activity—cross-scale spatiotemporal structure. Different EEG task patterns span a broad range of temporal and spatial scales, from brief neural activations to slow-varying rhythms, and from localized cortical activations to large-scale distributed interactions. Ignoring this diversity may lead to suboptimal representations and weakened generalization ability. To address these limitations, we propose CSBrain, a Cross-scale Spatiotemporal Brain foundation model for generalized EEG decoding. CSBrain introduces two key components: (i) Cross-scale Spatiotemporal Tokenization (CST), which aggregates multi-scale features within localized temporal windows and anatomical brain regions into compact scale-aware token representations; and (ii) Structured Sparse Attention (SSA), which models cross-window and cross-region dependencies for diverse decoding tasks, further enriching scale diversities while eliminating the spurious dependencies. CST and SSA are alternately stacked to progressively integrate cross-scale spatiotemporal dependencies. Extensive experiments across 11 representative EEG tasks and 16 datasets demonstrate that CSBrain consistently outperforms both task-specific models and strong foundation baselines. These results establish cross-scale modeling as a key inductive bias for generalized EEG decoding and highlight CSBrain as a robust backbone for future brain–AI research.
Yuchen Zhou 0002, Zichen Ren, Zhouheng Yao, Weiheng Lu, Kunyu Peng, Qihao Zheng, Chunfeng Song, Wanli Ouyang, Chao Gou
NeurIPS8
2024 GRAMO: geometric resampling augmentation for monocular 3D object detection
abstract
Abstract Data augmentation is widely recognized as an effective means of bolstering model robustness. However, when applied to monocular 3D object detection, non-geometric image augmentation neglects the critical link between the image and physical space, resulting in the semantic collapse of the extended scene. To address this issue, we propose two geometric-level data augmentation operators named Geometric-Copy-Paste (Geo-CP) and Geometric-Crop-Shrink (Geo-CS). Both operators introduce geometric consistency based on the principle of perspective projection, complementing the options available for data augmentation in monocular 3D. Specifically, Geo-CP replicates local patches by reordering object depths to mitigate perspective occlusion conflicts, and Geo-CS re-crops local patches for simultaneous scaling of distance and scale to unify appearance and annotation. These operations ameliorate the problem of class imbalance in the monocular paradigm by increasing the quantity and distribution of geometrically consistent samples. Experiments demonstrate that our geometric-level augmentation operators effectively improve robustness and performance in the KITTI and Waymo monocular 3D detection benchmarks.
He Guan, Chunfeng Song, Zhaoxiang Zhang 0001
Frontiers Comput. Sci.2
2023 Learning to Adapt Across Dual Discrepancy for Cross-Domain Person Re-Identification
abstract
Thanks to the advent of deep neural networks, recent years have witnessed rapid progress in person re-identification (re-ID). Deep-learning-based methods dominate the leadership of large-scale benchmarks, some of which even surpass the human-level performance. Despite their impressive performance under the single-domain setup, current fully-supervised re-ID models degrade significantly when transplanted to an unseen domain. According to the characteristics of the re-ID task, such degradation is mainly attributed to the dramatic variation within the target domain and the severe shift between the source and target domain, which we call dual discrepancy in this paper. To achieve a model that generalizes well to the target domain, it is desirable to take such dual discrepancy into account. In terms of the former issue, a prevailing solution is to enforce consistency between nearest-neighbors in the embedding space. However, we find that the search of neighbors is highly biased in our case due to the discrepancy across cameras. For this reason, we equip the vanilla neighborhood invariance approach with a camera-aware learning scheme. As for the latter issue, we propose a novel cross-domain mixup scheme. It works in conjunction with virtual prototypes which are employed to handle the disjoint label space between the two domains. In this way, we can realize the smooth transfer by introducing the interpolation between the two domains as a transition state. Extensive experiments on four public benchmarks demonstrate the superiority of our method. Without any auxiliary models and offline clustering procedure, it achieves competitive performance against existing state-of-the-art methods. The code is available at https://github.com/LuckyDC/generalizing-reid-improved.
Chuanchen Luo, Chunfeng Song, Zhaoxiang Zhang 0001
IEEE Trans. Pattern Anal. Mach. Intell.2
2023 CASIA-E: A Large Comprehensive Dataset for Gait Recognition
abstract
Gait recognition plays a special role in visual surveillance due to its unique advantage, e.g., long-distance, cross-view and non-cooperative recognition. However, it has not yet been widely applied. One reason for this awkwardness is the lack of a truly big dataset captured in practical outdoor scenarios. Here, the "big" at least means: (1) huge amount of gait videos; (2) sufficient subjects; (3) rich attributes; and (4) spatial and temporal variations. Moreover, most existing large-scale gait datasets are collected indoors, which have few challenges from real scenes, such as the dynamic and complex background clutters, illumination variations, vertical view variations, etc. In this article, we introduce a newly built big outdoor gait dataset, called CASIA-E. It contains more than one thousand people distributed over near one million videos. Each person involves 26 view angles and varied appearances caused by changes of bag carrying, dressing and walking styles. The videos are captured across five months and across three kinds of outdoor scenes. Soft biometric features are also recorded for all subjects including age, gender, height, weight, and nationality. Besides, we report an experimental benchmark and examine some meaningful problems that have not been well studied previously, e.g., the influence of million-level training videos, vertical view angles, walking styles, and the thermal infrared modality. We believe that such a big outdoor dataset and the experimental benchmark will promote the development of gait recognition in both academic research and industrial applications.
Chunfeng Song, Yongzhen Huang, Weining Wang 0001, Liang Wang 0001
IEEE Trans. Pattern Anal. Mach. Intell.1
2023 Weakly Supervised Semantic Segmentation via Box-Driven Masking and Filling Rate Shifting
abstract
Semantic segmentation has achieved huge progress via adopting deep Fully Convolutional Networks (FCN). However, the performance of FCN-based models severely rely on the amounts of pixel-level annotations which are expensive and time-consuming. Considering that bounding boxes also contain abundant semantic and objective information, an intuitive solution is to learn the segmentation with weak supervisions from the bounding boxes. How to make full use of the class-level and region-level supervisions from bounding boxes to estimate the uncertain regions is the critical challenge for the weakly supervised learning task. In this paper, we propose a mixture model to address this problem. First, we introduce a box-driven class-wise masking model (BCM) to remove irrelevant regions of each class. Moreover, based on the pixel-level segment proposal generated from the bounding box supervision, we calculate the mean filling rates of each class to serve as an important prior cue to guide the model ignoring the wrongly labeled pixels in proposals. To realize the more fine-grained supervision at instance-level, we further propose the anchor-based filling rate shifting module. Unlike previous methods that directly train models with the generated noisy proposals, our method can adjust the model learning dynamically with the adaptive segmentation loss. Thus it can help reduce the negative impacts from wrongly labeled proposals. Besides, based on the learned high-quality proposals with above pipeline, we explore to further boost the performance through two-stage learning. The proposed method is evaluated on the challenging PASCAL VOC 2012 benchmark and achieves 74.9 % and 76.4 % mean IoU accuracy under weakly and semi-supervised modes, respectively. Extensive experimental results show that the proposed method is effective and is on par with, or even better than current state-of-the-art methods.
Chunfeng Song, Wanli Ouyang, Zhaoxiang Zhang 0001
IEEE Trans. Pattern Anal. Mach. Intell.1
2022 The Devil Is in the Details: Window-based Attention for Image Compression
abstract
Learned image compression methods have exhibited superior rate-distortion performance than classical image compression standards. Most existing learned image compression models are based on Convolutional Neural Networks (CNNs). Despite great contributions, a main drawback of CNN based model is that its structure is not designed for capturing local redundancy, especially the nonrepetitive textures, which severely affects the reconstruction quality. Therefore, how to make full use of both global structure and local texture becomes the core problem for learning-based image compression. Inspired by recent progresses of Vision Transformer (ViT) and Swin Transformer, we found that combining the local-aware attention mechanism with the global-related feature learning could meet the expectation in image compression. In this paper, we first extensively study the effects of multiple kinds of attention mechanisms for local features learning, then introduce a more straightforward yet effective window-based local attention block. The proposed window-based attention is very flexible which could work as a plug-and-play component to enhance CNN and Transformer models. Moreover, we propose a novel Symmetrical TransFormer (STF) framework with absolute transformer blocks in the down-sampling encoder and up-sampling decoder. Extensive experimental evaluations have shown that the proposed method is effective and outperforms the state-of-the-art methods. The code is publicly available at https://github.com/Googolxx/STF.
Renjie Zou, Chunfeng Song, Zhaoxiang Zhang 0001
CVPR2
2022 Toward few-shot domain adaptation with perturbation-invariant representation and transferable prototypes
Junsong Fan, Yuxi Wang 0001, He Guan, Chunfeng Song, Zhaoxiang Zhang 0001
Frontiers Comput. Sci.4
2022 From Individual to Whole: Reducing Intra-class Variance by Feature Aggregation
Zhaoxiang Zhang 0001, Chuanchen Luo, Haiping Wu, Yuntao Chen, Naiyan Wang, Chunfeng Song
Int. J. Comput. Vis.6
2022 Dynamic video mix-up for cross-domain action recognition
Chunfeng Song, Shaolong Yue, Zhenyu Wang 0012, Jun Xiao 0005, Yanyang Liu
Neurocomputing2
2022 Identifying the key frames: An attention-aware sampling method for action recognition
Wenkai Dong, Zhaoxiang Zhang 0001, Chunfeng Song, Tieniu Tan
Pattern Recognit.3
2022 MonoPoly: A practical monocular 3D object detector
He Guan, Chunfeng Song, Zhaoxiang Zhang 0001, Tieniu Tan
Pattern Recognit.2
2021 Mask-guided contrastive attention and two-stream metric co-learning for person Re-identification
Chunfeng Song, Caifeng Shan, Yan Huang 0008, Liang Wang 0001
Neurocomputing1
2021 Video-Based Air Quality Measurement With Dual-Channel 3-D Convolutional Network
abstract
Air pollution detection and measurement is an important problem. Fast and effective explanation of the air quality is a necessary technique for monitoring the air pollution. However, the existing air quality measurement devices severely rely on many sensors, which are not only expensive but also inconvenient to carry. With the development of deep learning technology, computer vision-based task, such as the video processing and understanding have achieved great progress. Recently, image-based air quality measuring methods have been proposed and achieved satisfying accuracy in specific scenes. Whereas, the performance of those methods are not stable due to the noises in images and missed temporal relations between single images. With the rising of the short-video platform, the acquisition and dissemination of video data becomes more convenient. To address the problems of image-based air pollution measurement, we propose a video-based dual-channel 3-D convolution network for stable and accurate measuring. Besides the basic visual channel, we add a semantic channel to guide the network learn region-level features. The features from two channels are combined for stable prediction. Moreover, to evaluate the effectiveness of our method, we collect an outdoor video data set for air quality measurement through smart phones. Extensive experimental results show that our method achieves good results in the air quality measurement and outperforms the image-based methods.
Zhenyu Wang 0012, Shaolong Yue, Chunfeng Song
IEEE Internet Things J.3
2021 Adaptive super-resolution for person re-identification with low-resolution images
Yan Huang 0008, Chunfeng Song, Liang Wang 0001, Tieniu Tan
Pattern Recognit.3
2021 Multi-Domain Image-to-Image Translation via a Unified Circular Framework
abstract
The image-to-image translation aims to learn the corresponding information between the source and target domains. Several state-of-the-art works have made significant progress based on generative adversarial networks (GANs). However, most existing one-to-one translation methods ignore the correlations among different domain pairs. We argue that there is common information among different domain pairs and it is vital to multiple domain pairs translation. In this paper, we propose a unified circular framework for multiple domain pairs translation, leveraging a shared knowledge module across numerous domains. One selected translation pair can benefit from the complementary information from other pairs, and the sharing knowledge is conducive to mutual learning between domains. Moreover, absolute consistency loss is proposed and applied in the corresponding feature maps to ensure intra-domain consistency. Furthermore, our model can be trained in an end-to-end manner. Extensive experiments demonstrate the effectiveness of our approach on several complex translation scenarios, such as Thermal IR switching, weather changing, and semantic transfer tasks.
Yuxi Wang 0001, Zhaoxiang Zhang 0001, Chunfeng Song
IEEE Trans. Image Process.4
2021 Attention Guided Multiple Source and Target Domain Adaptation
abstract
Domain adaptation aims to alleviate the distribution discrepancy between source and target domains. Most conventional methods focus on one target domain setting adapted from one or multiple source domains while neglecting the multi-target domain setting. We argue that different target domains also have complementary information, which is very important for performance improvement. In this paper, we propose an Attention-guided Multiple source-and-target Domain Adaptation (AMDA) method to capture the context dependency information on transferable regions among multiple source and target domains. The innovation points of this paper are as follows: (1) We use numerous adversarial strategies to harvest sufficient information from multiple source and target domains, which extends the generalization and robustness of the feature pools. (2) We propose an intra-domain and inter-domain attention module to explore transferable context information. The proposed attention module can learn domain-invariant representations and reduce the negative transfer by focusing on transferable knowledge. Extensive experiments validate the effectiveness of our method with achieving state-of-the-art performance on several unsupervised domain adaptation datasets.
Yuxi Wang 0001, Zhaoxiang Zhang 0001, Chunfeng Song
IEEE Trans. Image Process.4
2020 CIAN: Cross-Image Affinity Net for Weakly Supervised Semantic Segmentation
abstract
Weakly supervised semantic segmentation with only image-level labels saves large human effort to annotate pixel-level labels. Cutting-edge approaches rely on various innovative constraints and heuristic rules to generate the masks for every single image. Although great progress has been achieved by these methods, they treat each image independently and do not take account of the relationships across different images. In this paper, however, we argue that the cross-image relationship is vital for weakly supervised segmentation. Because it connects related regions across images, where supplementary representations can be propagated to obtain more consistent and integral regions. To leverage this information, we propose an end-to-end cross-image affinity module, which exploits pixel-level cross-image relationships with only image-level labels. By means of this, our approach achieves 64.3% and 65.3% mIoU on Pascal VOC 2012 validation and test set respectively, which is a new state-of-the-art result by only using image-level labels for weakly supervised semantic segmentation, demonstrating the superiority of our approach.
Junsong Fan, Zhaoxiang Zhang 0001, Tieniu Tan, Chunfeng Song, Jun Xiao 0005
AAAI4
2020 Instance Guided Proposal Network for Person Search
abstract
Person detection networks have been widely used in person search. These detectors discriminate persons from the background and generate proposals of all the persons from a gallery of scene images for each query. However, such a large number of proposals have a negative influence on the following identity matching process because many distractors are involved. In this paper, we propose a new detection network for person search, named Instance Guided Proposal Network (IGPN), which can learn the similarity between query persons and proposals. Thus, we can decrease proposals according to the similarity scores. To incorporate information of the query into the detection network, we introduce the Siamese region proposal network to Faster-RCNN and we propose improved cross-correlation layers to alleviate the imbalance of parameters distribution. Furthermore, we design a local relation block and a global relation branch to leverage the proposal-proposal relations and query-scene relations, respectively. Extensive experiments show that our method improves the person search performance through decreasing proposals and achieves competitive performance on two large person search benchmark datasets, CUHK-SYSU and PRW.
Wenkai Dong, Zhaoxiang Zhang 0001, Chunfeng Song, Tieniu Tan
CVPR3
2020 Bi-Directional Interaction Network for Person Search
abstract
Existing works have designed end-to-end frameworks based on Faster-RCNN for person search. Due to the large receptive fields in deep networks, the feature maps of each proposal, cropped from the stem feature maps, involve redundant context information outside the bounding boxes. However, person search is a fine-grained task which needs accurate appearance information. Such context information can make the model fail to focus on persons, so the learned representations lack the capacity to discriminate various identities. To address this issue, we propose a Siamese network which owns an additional instance-aware branch, named Bi-directional Interaction Network (BINet). During the training phase, in addition to scene images, BINet also takes as inputs person patches which help the model discriminate identities based on human appearance. Moreover, two interaction losses are designed to achieve bi-directional interaction between branches at two levels. The interaction can help the model learn more discriminative features for persons in the scene. At the inference stage, only the major branch is applied, so BINet introduces no additional computation. Extensive experiments on two widely used person search benchmarks, CUHK-SYSU and PRW, have shown that our BINet achieves state-of-the-art results among end-to-end methods without loss of efficiency.
Wenkai Dong, Zhaoxiang Zhang 0001, Chunfeng Song, Tieniu Tan
CVPR3
2020 Learning Integral Objects With Intra-Class Discriminator for Weakly-Supervised Semantic Segmentation
abstract
Image-level weakly-supervised semantic segmentation (WSSS) aims at learning semantic segmentation by adopting only image class labels. Existing approaches generally rely on class activation maps (CAM) to generate pseudo-masks and then train segmentation models. The main difficulty is that the CAM estimate only covers partial foreground objects. In this paper, we argue that the critical factor preventing to obtain the full object mask is the classification boundary mismatch problem in applying the CAM to WSSS. Because the CAM is optimized by the classification task, it focuses on the discrimination across different image-level classes. However, the WSSS requires to distinguish pixels sharing the same image-level class to separate them into the foreground and the background. To alleviate this contradiction, we propose an efficient end-to-end Intra-Class Discriminator (ICD) framework, which learns intra-class boundaries to help separate the foreground and the background within each image-level class. Without bells and whistles, our approach achieves the state-of-the-art performance of image label based WSSS, with mIoU 68.0% on the VOC 2012 semantic segmentation benchmark, demonstrating the effectiveness of the proposed approach.
Junsong Fan, Zhaoxiang Zhang 0001, Chunfeng Song, Tieniu Tan
CVPR3
2020 Generalizing Person Re-Identification by Camera-Aware Invariance Learning and Cross-Domain Mixup
Chuanchen Luo, Chunfeng Song, Zhaoxiang Zhang 0001
ECCV (15)2
2020 Attentive Part-aware Networks for Partial Person Re- identification
abstract
Partial person re-identification (re-ID) refers to re-identify a person through occluded images. It suffers from two major challenges, i.e., insufficient training data and incomplete probe image. In this paper, we introduce a part-aware learning method for partial person re-identification. On the one hand, we adopt data augmentation operation to enrich the training data and improve the robustness of the model. On the other hand, we intuitively find that the partial person images usually have fixed percentages of parts, therefore, in partial person re-ID task, the probe image could be cropped from the pictures and divided into several different partial types following fixed ratios. Based on the cropped images, we propose the Cropping Type Consistency (CTC) loss to classify the cropping types of partial images. Moreover, in order to help the network better fit the generated and cropped data, we incorporate the Block Attention Mechanism (BAM) into the framework for attentive learning. To enhance the retrieval performance in the inference stage, we implement cropping on gallery images according to the predicted types of probe partial images. Through calculating feature distances between the partial image and the cropped holistic gallery images, the model can recognize the right person from the gallery. To validate the effectiveness of our approach, we conduct extensive experiments on the partial re- ID benchmarks and achieve state-of-the-art performance.
Lijuan Huo, Chunfeng Song, Zhengyi Liu, Zhaoxiang Zhang 0001
ICPR2
2020 Frame-GAN: Increasing the frame rate of gait videos with generative adversarial networks
Hong Ai, Chunfeng Song, Yan Huang 0008, Liang Wang 0001
Neurocomputing4
2019 Box-Driven Class-Wise Region Masking and Filling Rate Guided Loss for Weakly Supervised Semantic Segmentation
abstract
Semantic segmentation has achieved huge progress via adopting deep Fully Convolutional Networks (FCN). However, the performance of FCN based models severely rely on the amounts of pixel-level annotations which are expensive and time-consuming. To address this problem, it is a good choice to learn to segment with weak supervision from bounding boxes. How to make full use of the class-level and region-level supervisions from bounding boxes is the critical challenge for the weakly supervised learning task. In this paper, we first introduce a box-driven class-wise masking model (BCM) to remove irrelevant regions of each class. Moreover, based on the pixel-level segment proposal generated from the bounding box supervision, we could calculate the mean filling rates of each class to serve as an important prior cue, then we propose a filling rate guided adaptive loss (FR-Loss) to help the model ignore the wrongly labeled pixels in proposals. Unlike previous methods directly training models with the fixed individual segment proposals, our method can adjust the model learning with global statistical information. Thus it can help reduce the negative impacts from wrongly labeled proposals. We evaluate the proposed method on the challenging PASCAL VOC 2012 benchmark and compare with other methods. Extensive experimental results show that the proposed method is effective and achieves the state-of-the-art results.
Chunfeng Song, Yan Huang 0008, Wanli Ouyang, Liang Wang 0001
CVPR1
2019 Learning view invariant gait features with Two-Stream GAN
Yanyun Wang 0008, Chunfeng Song, Yan Huang 0008, Zhenyu Wang 0012, Liang Wang 0001
Neurocomputing2
2019 GaitNet: An end-to-end network for gait based human identification
Chunfeng Song, Yongzhen Huang, Yan Huang 0008, Liang Wang 0001
Pattern Recognit.1
2018 Learning Semantic Concepts and Order for Image and Sentence Matching
abstract
Image and sentence matching has made great progress recently, but it remains challenging due to the large visual-semantic discrepancy. This mainly arises from that the representation of pixel-level image usually lacks of high-level semantic information as in its matched sentence. In this work, we propose a semantic-enhanced image and sentence matching model, which can improve the image representation by learning semantic concepts and then organizing them in a correct semantic order. Given an image, we first use a multi-regional multi-label CNN to predict its semantic concepts, including objects, properties, actions, etc. Then, considering that different orders of semantic concepts lead to diverse semantic meanings, we use a context-gated sentence generation scheme for semantic order learning. It simultaneously uses the image global context containing concept relations as reference and the groundtruth semantic order in the matched sentence as supervision. After obtaining the improved image representation, we learn the sentence representation with a conventional LSTM, and then jointly perform image and sentence matching and sentence generation for model learning. Extensive experiments demonstrate the effectiveness of our learned semantic concepts and order, by achieving the state-of-the-art results on two public benchmark datasets.
Yan Huang 0008, Qi Wu 0001, Chunfeng Song, Liang Wang 0001
CVPR3
2018 Mask-Guided Contrastive Attention Model for Person Re-Identification
abstract
Person Re-identification (ReID) is an important yet challenging task in computer vision. Due to the diverse background clutters, variations on viewpoints and body poses, it is far from solved. How to extract discriminative and robust features invariant to background clutters is the core problem. In this paper, we first introduce the binary segmentation masks to construct synthetic RGB-Mask pairs as inputs, then we design a mask-guided contrastive attention model (MGCAM) to learn features separately from the body and background regions. Moreover, we propose a novel region-level triplet loss to restrain the features learnt from different regions, i.e., pulling the features from the full image and body region close, whereas pushing the features from backgrounds away. We may be the first one to successfully introduce the binary mask into person ReID task and the first one to propose region-level contrastive learning. We evaluate the proposed method on three public datasets, including MARS, Market-1501 and CUHK03. Extensive experimental results show that the proposed method is effective and achieves the state-of-the-art results. Mask and code will be released upon request.
Chunfeng Song, Yan Huang 0008, Wanli Ouyang, Liang Wang 0001
CVPR1
2015 Kinship Verification with Deep Convolutional Neural Networks
Kaihao Zhang, Yongzhen Huang, Chunfeng Song, Liang Wang 0001
BMVC3
2014 Deep auto-encoder based clustering
abstract
For unsupervised problems like clustering, linear or non-linear data transformations are widely used techniques. Generally, they are beneficial to data representation. However, if data have a complicated structure, these techniques would be unsatisfying for clustering. In this paper, we propose a new clustering method based on the deep auto-encoder network, which can learn a highly non-linear mapping function. Via simultaneously considering data reconstruction and compactness, our method can obtain stable and effective clustering. Experimental results on four databases demonstrate that the proposed model can achieve promising performance in terms of normalized mutual information, cluster purity and accuracy.
Chunfeng Song, Yongzhen Huang, Feng Liu 0036, Zhenyu Wang 0012, Liang Wang 0001
Intell. Data Anal.1
2014 Multiple spatial pooling for visual object recognition
Yongzhen Huang, Zifeng Wu, Liang Wang 0001, Chunfeng Song
Neurocomputing4
2013 Auto-encoder Based Data Clustering
Chunfeng Song, Feng Liu 0036, Yongzhen Huang, Liang Wang 0001, Tieniu Tan
CIARP (1)1
2009 The Sync Tracing Based on Improved Genetic Algorithm Neural Network
abstract
The elevator system is important in mine safety manufacture. Aiming the character of frequent startup and stop with nonlinearity, the sync tracing method based on improved genetic algorithm neural network is presented. Because the condition of the normal adaptation function is too free, the adaptation function is improved, which is the new function altering with input space, then, improved genetic algorithm neural network (IGANN) is established, the IGANN not only avoids getting into local extremum point, but also realizes sync tracing. It is proved by simulation of 400 kW assistant elevator in nine, that the sync tracing IGANN is effective for the character of frequent startup and stop with nonlinearity.
Yuanbin Hou, Chunfeng Song
IAS2