Zunlei Feng

dblp:191/2455 · DBLP profile ↗
← Back
134ranked-venue papers
14as first author
117since 2021 · last 2026
0000-0001-8640-8434ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 80 · 11 first-author · 71 since 2021Graphics, computer vision, multimedia, augmented reality and games · 73 · 8 first-author · 63 since 2021Databases, data management, data science and information retrieval · 10 · 1 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 9 · 8 since 2021Human-computer interaction and ubiquitous computing · 5 · 4 since 2021Computer networks · 2 · 1 first-author · 1 since 2021Security and privacy · 1 · 1 since 2021
YearPublicationVenuePosition
2026 D3-RSMDE: 40× Faster and High-Fidelity Remote Sensing Monocular Depth Estimation
abstract
Real-time, high-fidelity monocular depth estimation from remote sensing imagery is crucial for numerous applications, yet existing methods face a stark trade-off between accuracy and efficiency. Although using Vision Transformer (ViT) backbones for dense prediction is fast, they often exhibit poor perceptual quality. Conversely, diffusion models offer high fidelity but at a prohibitive computational cost. To overcome these limitations, we propose Depth Detail Diffusion for Remote Sensing Monocular Depth Estimation (D³-RSMDE), an efficient framework designed to achieve an optimal balance between speed and quality. Our framework first leverages a ViT-based module to rapidly generate a high-quality preliminary depth map construction, which serves as a structural prior, effectively replacing the time-consuming initial structure generation stage of diffusion models. Based on this prior, we propose a Progressive Linear Blending Refinement (PLBR) strategy, which uses a lightweight U-Net to refine the details in only a few iterations. The entire refinement step operates efficiently in a compact latent space supported by a Variational Autoencoder (VAE). Extensive experiments demonstrate that D³-RSMDE achieves a notable 11.85% reduction in the Learned Perceptual Image Patch Similarity (LPIPS) perceptual metric over leading models like Marigold, while also achieving over a 40× speedup in inference and maintaining VRAM usage comparable to lightweight ViT models.
Zunlei Feng, Haofei Zhang, Mingli Song, Jie Song 0011
AAAI3
2026 Multi-label Learning for Reliable Cervical Cytology Screening
Linyun Zhou, Jian Yang 0003, Xiuming Zhang, Zunlei Feng, Bingde Hu
ICIC (29)6
2026 Adaptive Location Hierarchy Learning for Long-Tailed Mobility Prediction
abstract
Human mobility prediction is crucial for applications ranging from location-based recommendations to urban planning, which aims to forecast users' next location visits based on historical trajectories. While existing mobility prediction models excel at capturing sequential patterns through diverse architectures for different scenarios, they are hindered by the long-tailed distribution of location visits, leading to biased predictions and limited applicability. This highlights the need for a solution that enhances the long-tailed prediction capabilities of these models with broad compatibility and efficiency across diverse architectures. To address this need, we propose the first architecture-agnostic plugin for long-tailed human mobility prediction, named \textbf{A}daptive \textbf{LO}cation \textbf{H}ier\textbf{A}rchy learning (ALOHA). Inspired by Maslow's theory of human motivation, we exploit and explore common mobility knowledge of head and tail locations derived from human mobility trajectories to effectively mitigate long-tailed bias. Specifically, we introduce an automatic pipeline to construct city-tailored location hierarchies based on Large Language Models (LLMs) and Chain-of-Thought (CoT) prompts, capturing high-level mobility semantics with minimal human verification. We further design an Adaptive Hierarchical Loss (AHL) that rebalances learning through Gumbel disturbance and node-wise adaptive weighting, enabling both exploitation of multi-level signals and exploration within semantically related groups. Extensive experiments across multiple state-of-the-art models demonstrate that ALOHA consistently improves long-tailed mobility prediction performance by up to 16.59\% while maintaining efficiency and robustness. Our code is at https://github.com/Star607/ALOHA.
Yu Wang 0176, Junshu Dai, Yuchen Ying, Hanyang Yuan, Zunlei Feng, Tongya Zheng, Mingli Song
WWW5
2026 Task-adaptive parameter optimization for medical image classification transfer learning
Xiangtong Du, Zhidong Liu, Weifan Xu, Zunlei Feng
Multim. Syst.4
2026 Self-improved holistic alignment for preference enhancement
Kejia Chen 0007, Jiawen Zhang 0005, Jiazhen Yang, Mingli Song, Zunlei Feng
Pattern Recognit.5
2026 Divide-and-conquer towards optimal adaptation of pre-trained model to medical tasks
Zhanghui Huang, Zunlei Feng, Xiaoyan Sun 0006, Shuifa Sun, Zhenming Yuan, Jun Yu 0002, Jian Zhang 0026
Pattern Recognit.2
2026 Morphology semantics-guided vision language alignment for cervical cell image classification
Jiaxin Lei, Zunlei Feng, Jingwen Ye, Zhenming Yuan, Jun Yu 0002
Pattern Recognit.2
2026 An Intelligent Interactive Visual Analytics System for Exploring Large and Multi-Scale Pathology Images
abstract
Pathology images are crucial for cancer diagnosis and treatment. Although artificial intelligence has driven rapid advancements in pathology image analysis, the interpretation of ultra-large and multi-scale pathology images in clinical practice still heavily relies on physicians' experience. Clinicians need to repeatedly zoom in and out on individual slides to compare and assess pathological details - a process that is both time-consuming and prone to visual fatigue. The system first employs a diffusion model to perform tissue segmentation on pathology images, then calculates pathological tissue proportions and morphological metrics. Finally, through multi-scale dynamic comparison and multi-level visual evaluation, the system facilitates comprehensive and precise analysis of pathology images. The system provides clinicians with an intelligent and interactive tool for pathology image interpretation, enabling efficient visualization and precise analysis of pathological details, thereby reducing the effort require for detailed analysis.
Chaoqing Xu, Xinyuan Fu, Liting Fang, Zunlei Feng, Xiuming Zhang, Can Wang 0001, Mingli Song, Wei Chen 0001
IEEE Trans. Vis. Comput. Graph.6
2025 Global Attribute-Association Pattern Aggregation for Graph Fraud Detection
abstract
Fraud is increasingly prevalent, and its patterns are frequently changing, posing challenges for fraud detection methods such as random forests and Graph Neural Networks (GNNs), which rely on bin-based and mixture features separately. The former may lose crucial graph-associated features, while the latter face incorrect feature fusion. To overcome these limitations, we propose an approach based on attribute-association pattern that leverages the distinct attribute and association patterns differentiating fraudulent from benign behaviors, to enhance fraud detection capabilities. Attribute features are adaptively split into separate bins to eliminate incorrect attribute fusion and combine association patterns through graph neighbor message passing, thereby deriving attribute-association pattern features. Using the learned attribute-association patterns, the fraud patterns between a single pattern and the patterns across the entire graph are globally aggregated. Extensive experiments comparing our approach with 24 methods on 7 datasets demonstrate that the proposed method achieves SOTA performance.
Mingjiang Duan, Da He, Tongya Zheng, Lingxiang Jia, Mingli Song, Xinyu Wang 0001, Zunlei Feng
AAAI7
2025 Association Pattern-enhanced Molecular Representation Learning
abstract
The applicability of drug molecules in various clinical scenarios is significantly influenced by a diverse range of molecular properties. By leveraging self-supervised conditions such as atom attributes and interatomic bonds, existing advanced molecular foundation models can generate expressive representations of these molecules. However, such models often overlook the fixed association patterns within molecules that influence physiological or chemical properties. In this paper, we introduce a novel association pattern-aware message passing method, which can serve as an effective yet general plug-and-play plugin, thereby enhancing the atom representations generated by molecular foundation models without requiring additional pretraining. Additionally, molecular property-specific pattern libraries are constructed to collect the generated interpretable common patterns that bind to these properties. Extensive experiments conducted on 11 benchmark molecular property prediction tasks across 8 advanced molecular foundation models demonstrate significant superiority of the proposed method, with performance improvements of up to approximately 20%. Furthermore, a property-specific pattern library is tailored for blood-brain barrier penetration, which has undergone corresponding mechanistic validation.
Lingxiang Jia, Yuchen Ying, Shaolun Yao, Jie Lei 0002, Jie Song 0011, Mingli Song, Zunlei Feng
AAAI9
2025 TED-DTMoA: Tri-Comparison Expertise Decision for Drug-Target Mechanism of Action
abstract
Machine-learned interactions between drugs and human protein targets play a crucial role in efficient and accurate drug discovery. However, the drug-target mechanism of action (DTMoA) prediction is actually a multi-class classification problem, which follows a long-tailed class distribution. Existing methods simply address whether the drugs and targets can interact and rarely consider these deep mechanisms. In this paper, we introduce TED-DTMoA, a novel DTMoA prediction framework that incorporates the divide-and-conquer strategy with tri-comparison options. Specifically, to reduce the learning difficulty of tail classes, we propose an expertise-based divide-and-conquer decision approach that combines the results of multiple independent expertise models for sub-tasks decomposed from the original prediction task. In addition, to enhance the discrimination of similar mechanism classes, we devise a tri-comparison learning strategy that defines the sub-task as the classification of triple options, such as expanding the classification task for classes A and B to include an extra “Neither of them” option. Extensive experiments conducted on various DTMoA datasets quantitatively demonstrate the proposed method achieves an approximately 13% performance improvement compared with advanced baselines. Moreover, our method exhibits an obvious superiority on the tail classes. Further analysis of the evolvability and generalization reveals the significant potential to be deployed in real-world scenes.
Lingxiang Jia, Zipeng Zhong, Shaolun Yao, Jie Song 0011, Mingli Song, Zunlei Feng
ECAI6
2025 Multimodal Sentiment Analysis with Parallel Attention and Correlation Fusion
Jie Lei 0002, Zunlei Feng, Ronghua Liang
ICANN (3)4
2025 Spatial-Temporal Reconstruction Error for AIGC-based Forgery Image Detection
abstract
The remarkable success of AI-Generated Content (AIGC), especially diffusion image generation models, brings about unprecedented creative applications, but also creates fertile ground for malicious counterfeiting and crime. A highly effective family of forgery image detection methods based on diffusion reconstruction error has emerged, as images generated by diffusion are more easily reconstructed by any diffusion model. However, we find that existing methods only use reconstruction error from a single time step, failing to fully leverage the entire reconstruction process. To this end, we propose to comprehensively consider every single time step to form the Temporal Reconstruction Error (TRE) that offers a richer feature representation. Furthermore, we design temporal aggregation and spatial focusing modules from two dimensions respectively to more effectively extract discriminative information from the TRE feature. Finally, we validate the proposed method on two popular datasets, and experimental results demonstrate that the proposed approach achieves state-of-the-art performance.
Chengji Shen, Zhenjiang Liu, Kai-Xuan Chen 0001, Jie Lei 0002, Mingli Song, Zunlei Feng
ICASSP6
2025 Dynamic Routing and Calibration for Few-Shot Object Detection
abstract
Few-shot object detection (FSOD), aiming to enhance the performance of novel object detection with limited labeled samples, has recently gained significant attention. Recent researches primarily focus on improving the generalization of novel classes and enhancing detector performance. However, the diversity of samples is often overlooked, and object proposals with inaccurate classifications or locations remain uncorrected. In this paper, we propose Dynamic Routing and Calibration for Few-Shot Object Detection (DRC-FSOD). Our approach includes a dynamic backbone routing that adapts to various samples by selecting appropriate backbones dynamically. Meanwhile, we construct a dynamic calibration module, which dynamically perform individual calibration for proposals based on their scores. Experimental results on MS COCO and Pascal VOC datasets show superiority over state-of-the-art methods.
Jie Lei 0002, Zunlei Feng, Ronghua Liang
ICASSP5
2025 Spatial-Temporal Forgery Trace Based Forgery Image Identification
Zunlei Feng, Jiachi Wang, Hengrui Lou, Binjia Zhou, Jie Lei 0002, Mingli Song, Yijun Bei
ICCV2
2025 Large Vision-Language Models are Generalist Solvers For Pathology Tasks
abstract
Leveraging the powerful capabilities of large language models (LLMs), large vision-language models (LVLMs) can perform a wide variety of tasks based on input images and user instructions. However, existing pathology-focused LVLMs are limited to relatively simple tasks, such as image captioning, visual question answering, and generating brief pathology reports, which restricts their clinical applicability. To enhance the practicality of pathology LVLMs and explore their performance boundaries across pathology tasks, we curated a multi-task pathology instruction-following dataset that better aligns with clinical needs. This dataset encompasses tasks such as cancer classification and grading, molecular subtype identification, and the detection of structures like nuclei, blood vessels, nerves, and lymph nodes. Extensive experiments were conducted on this dataset to identify key factors influencing the performance of LVLMs on these pathology tasks, and optimal solutions were proposed. Our findings provide valuable insights to advance the clinical application of large vision-language models in pathology.
Shengxuming Zhang, Hengrui Lou, Jing Zhang 0120, Xiuming Zhang, Mingli Song, Zunlei Feng
ICIP6
2025 Reinforced Model Merging
abstract
The success of large language models has garnered widespread attention for model merging techniques, especially training-free methods which combine model capabilities within the parameter space. However, two challenges remain: (1) uniform treatment of all parameters leads to performance degradation; (2) search-based algorithms are often inefficient. In this paper, we present an innovative framework termed Reinforced Model Merging (RMM), which encompasses an environment and agent tailored for merging tasks. These components interact to execute layer-wise merging actions, aiming to search the optimal merging architecture. Notably, RMM operates without any gradient computations on the original models, rendering it feasible for edge devices. Furthermore, by utilizing data subsets during the evaluation process, we addressed the bottleneck in the reward feedback phase, thereby accelerating RMM by up to 100 times. Extensive experiments demonstrate that RMM achieves state-of-the-art performance across various vision and NLP datasets and effectively overcomes the limitations of the existing baseline methods. Our code is available at https://github.com/WuDiHJQ/Reinforced-Model-Merging.
Jingwen Ye, Shunyu Liu 0001, Haofei Zhang, Jie Song 0011, Zunlei Feng, Mingli Song
ICME6
2025 Beyond the Label: Unveiling Fairness through Dynamic Attribute Projections in Classification
abstract
Image classification has been widely adopted in critical applications such as face recognition and medical imaging, but its prediction fairness has raised significant concerns. Existing fairness evaluation specifications and metrics have inherent limitations, which overlook certain correlations between target features and sensitive attributes. In this work, we introduce a novel evaluation specification for image classification models based on dynamic perturbations to address this challenge. Specifically, we propose an Attribute Projection Perturbation Strategy (APPS) and a projection-based fairness metric system to quantify the upper and lower bounds of fairness perturbations. By employing projection factors, sensitive attributes that may influence task-specific properties are mapped onto a unified dimension, enabling a multi-perspective examination and evaluation of the impact of these attributes on the fairness of prediction outcomes. Compared to existing metrics, the proposed evaluation specification demonstrates superior objectivity and interpretability across 24 image classification models, including CNN and ViT architectures.
Haoze Jiang, Zunlei Feng, Jiacong Hu, Binde Hu, Mingli Song, Yuanyu Wan
ICME2
2025 Assessing Safety Risks and Quantization-aware Safety Patching for Quantized Large Language Models
abstract
Quantized large language models (LLMs) have gained increasing attention and significance for enabling deployment in resource-constrained environments. However, emerging studies on a few calibration dataset-free quantization methods suggest that quantization may compromise the safety capabilities of LLMs, underscoring the urgent need for systematic safety evaluations and effective mitigation strategies. In this paper, we present comprehensive safety evaluations across various mainstream quantization techniques and diverse calibration datasets, utilizing widely accepted safety benchmarks. To address the identified safety vulnerabilities, we propose a quantization-aware safety patching framework, Q-resafe, to efficiently restore the safety capabilities of quantized LLMs while minimizing any adverse impact on utility. Extensive experiment results demonstrate that Q-resafe successfully re-aligns the safety of quantized LLMs with their pre-quantization counterparts, even under challenging evaluation scenarios. Project page: https://github.com/Thecommonirin/Qresafe.
Kejia Chen 0007, Jiawen Zhang 0005, Jiacong Hu, Yu Wang 0176, Jian Lou 0001, Zunlei Feng, Mingli Song
ICML6
2025 L-Diffusion: Laplace Diffusion for Efficient Pathology Image Segmentation
abstract
Pathology image segmentation plays a pivotal role in artificial digital pathology diagnosis and treatment. Existing approaches to pathology image segmentation are hindered by labor-intensive annotation processes and limited accuracy in tail-class identification, primarily due to the long-tail distribution inherent in gigapixel pathology images. In this work, we introduce the Laplace Diffusion Model, referred to as L-Diffusion, an innovative framework tailored for efficient pathology image segmentation. L-Diffusion utilizes multiple Laplace distributions, as opposed to Gaussian distributions, to model distinct components—a methodology supported by theoretical analysis that significantly enhances the decomposition of features within the feature space. A sequence of feature maps is initially generated through a series of diffusion steps. Following this, contrastive learning is employed to refine the pixel-wise vectors derived from the feature map sequence. By utilizing these highly discriminative pixel-wise vectors, the segmentation module achieves a harmonious balance of precision and robustness with remarkable efficiency. Extensive experimental evaluations demonstrate that L-Diffusion attains improvements of up to 7.16%, 26.74%, 16.52%, and 3.55% on tissue segmentation datasets, and 20.09%, 10.67%, 14.42%, and 10.41% on cell segmentation datasets, as quantified by DICE, MPA, mIoU, and FwIoU metrics. The source are available at https://github.com/Lweihan/LDiffusion.
Linyun Zhou, Yang Jian, Shengxuming Zhang, Xiangtong Du, Xiuming Zhang, Jing Zhang 0120, Chaoqing Xu, Mingli Song, Zunlei Feng
ICML10
2025 STD-FD: Spatio-Temporal Distribution Fitting Deviation for AIGC Forgery Identification
abstract
With the rise of AIGC technologies, particularly diffusion models, highly realistic fake images that can deceive human visual perception has become feasible. Consequently, various forgery detection methods have emerged. However, existing methods treat the generation process of fake images as either a black-box or an auxiliary tool, offering limited insights into its underlying mechanisms. In this paper, we propose Spatio-Temporal Distribution Fitting Deviation (STD-FD) for AIGC forgery detection, which explores the generative process in detail. By decomposing and reconstructing data within generative diffusion models, initial experiments reveal temporal distribution fitting deviations during the image reconstruction process. These deviations are captured through reconstruction noise maps for each spatial semantic unit, derived via a super-resolution algorithm. Critical discriminative patterns, termed DFactors, are identified through statistical modeling of these deviations. Extensive experiments show that STD-FD effectively captures distribution patterns in AIGC-generated data, demonstrating strong robustness and generalizability while outperforming state-of-the-art (SOTA) methods on major datasets. The source code is available at [this link](https://github.com/HengruiLou/STDFD).
Hengrui Lou, Zunlei Feng, Jinsong Geng, Erteng Liu, Jie Lei 0002, Lechao Cheng, Jie Song 0011, Mingli Song, Yijun Bei
ICML2
2025 Binning Encoder-Based Grouped Aggregation for Network Traffic Anomaly Detection
Lingyao Lu, Tongya Zheng, Haoye Wang, Zunlei Feng, Mingli Song
ICONIP (3)4
2025 Self-calibration Enhanced Whole Slide Pathology Image Analysis
abstract
Pathology images are considered the ``gold standard" for cancer diagnosis and treatment, with gigapixel images providing extensive tissue and cellular information. Existing methods fail to simultaneously extract global structural and local detail features for comprehensive pathology image analysis efficiently. To address these limitations, we propose a self-calibration enhanced framework for whole slide pathology image analysis, comprising three components: a global branch, a focus predictor, and a detailed branch. The global branch initially classifies using the pathological thumbnail, while the focus predictor identifies relevant regions for classification based on the last layer features of the global branch. The detailed extraction branch then assesses whether the magnified regions correspond to the lesion area. Finally, a feature consistency constraint between the global and detail branches ensures that the global branch focuses on the appropriate region and extracts sufficient discriminative features for final identification. These focused discriminative features can facilitate the discovery of novel prognostic tumor markers, from the perspective of feature uniqueness and tissue spatial distribution. Extensive experiment results demonstrate that the proposed framework can rapidly deliver accurate and explainable results for pathological grading and prognosis tasks.
Haoming Luo, Xiaotian Yu, Shengxuming Zhang, Jiabin Xia, Jian Yang 0003, Yuning Sun, Xiuming Zhang, Jing Zhang 0120, Zunlei Feng
IJCAI9
2025 DenseSAM: Semantic Enhance SAM for Efficient Dense Object Segmentation
abstract
Dense object segmentation is essential for various applications, particularly in pathology image and remote sensing image analysis. However, distinguishing numerous similar and densely packed objects in this task presents significant challenges. Several methods, including CNN- and ViT-based approaches, have been proposed to tackle these issues. Yet, models trained on limited datasets exhibit limited generalization ability. The Segment Anything Model (SAM) has recently achieved significant progress in zero-shot segmentation but relies heavily on precise positional guidance. However, providing numerous accurate location prompts in dense scenarios is time-consuming. To overcome this limitation, we conducted an in-depth exploration of the SAM mechanism and found that its strong generalization ability stems from the encoder’s edge detection capability, which is semantically independent, making location prompts essential for segmentation. This insight inspired the development of DenseSAM, which replaces location prompts with semantic guidance for automatic segmentation in dense scenarios. Specifically, it uses local details to weaken the edges of background objects, leverages global context to enhance intra-class feature similarity, while further increasing contrast with the background, and integrates a dual-head decoding process to enable lightweight automatic semantic segmentation. Extensive experiments on pathology images demonstrate that DenseSAM delivers remarkable performance with minimal training parameters, providing a cost-effective and efficient solution. Moreover, experiments on remote sensing images further validate its excellent scalability, making DenseSAM suitable for various dense object segmentation domains. The code is available at https://github.com/imAzhou/DenseSAM.
Linyun Zhou, Jiacong Hu, Shengxuming Zhang, Xiangtong Du, Mingli Song, Xiuming Zhang, Zunlei Feng
IJCAI7
2025 CorrDetail: Visual Detail Enhanced Self-Correction for Face Forgery Detection
abstract
With the swift progression of image generation technology, the widespread emergence of facial deepfakes poses significant challenges to the field of security, thus amplifying the urgent need for effective deepfake detection. Existing techniques for face forgery detection can broadly be categorized into two primary groups: visual-based methods and multimodal approaches. The former often lacks clear explanations for forgery details, while the latter, which merges visual and linguistic modalities, is more prone to the issue of hallucinations.To address these shortcomings, we introduce a visual detail enhanced self-correction framework, designated CorrDetail, for interpretable face forgery detection. CorrDetail is meticulously designed to rectify authentic forgery details when provided with error-guided questioning, with the aim of fostering the ability to uncover forgery details rather than yielding hallucinated responses. Additionally, to bolster the reliability of its findings, a visual fine-grained detail enhancement module is incorporated, supplying CorrDetail with more precise visual forgery details. Ultimately, a fusion decision strategy is devised to further augment the model's discriminative capacity in handling extreme samples, through the integration of visual information compensation and model bias reduction. Experimental results demonstrate that CorrDetail not only achieves state-of-the-art performance compared to the latest methodologies but also excels in accurately identifying forged details, all while exhibiting robust generalization capabilities.
Binjia Zhou, Hengrui Lou, Lizhe Chen, Dawei Luo, Jie Lei 0002, Zunlei Feng, Yijun Bei
IJCAI8
2025 A Large-scale Universal Evaluation Benchmark For Face Forgery Detection
Hengrui Lou, Zunlei Feng, Jinsong Geng, Erteng Liu, Lechao Cheng, Jie Lei 0002, Jie Song 0011, Mingli Song, Yijun Bei
ACM Multimedia2
2025 Association-Focused Path Aggregation for Graph Fraud Detection
abstract
Fraudulent activities have caused substantial negative social impacts and are exhibiting emerging characteristics such as intelligence and industrialization, posing challenges of high-order interactions, intricate dependencies, and the sparse yet concealed nature of fraudulent entities. Existing graph fraud detectors are limited by their narrow "receptive fields", as they focus only on the relations between an entity and its neighbors while neglecting longer-range structural associations hidden between entities. To address this issue, we propose a novel fraud detector based on Graph Path Aggregation (GPA). It operates through variable-length path sampling, semantic-associated path encoding, path interaction and aggregation, and aggregation-enhanced fraud detection. To further facilitate interpretable association analysis, we synthesize G-Internet, the first benchmark dataset in the field of internet fraud detection. Extensive experiments across datasets in multiple fraud scenarios demonstrate that the proposed GPA outperforms mainstream fraud detectors by up to +15% in Average Precision (AP). Additionally, GPA exhibits enhanced robustness to noisy labels and provides excellent interpretability by uncovering implicit fraudulent patterns across broader contexts. Code is available at https://github.com/horrible-dong/GPA.
Zunlei Feng, Jie Lei 0002, Mingli Song, Yang Gao 0001
NeurIPS3
2025 Signature Feature Sequence for Model Reuse Detection
Zunlei Feng, Jie Lei 0002
PRCV (2)5
2025 Activation Approximations Can Incur Safety Vulnerabilities in Aligned LLMs: Comprehensive Analysis and Defense
Jiawen Zhang 0005, Kejia Chen 0007, Lipeng He, Jian Lou 0001, Dan Li 0032, Zunlei Feng, Mingli Song, Jian Liu 0012, Kui Ren 0001, Xiaohu Yang 0001
USENIX Security Symposium6
2025 πFL: Private, atomic, incentive mechanism for federated learning based on blockchain
abstract
Federated learning (FL) is predicated on the provision of high-quality data by multiple clients, which is then used to train global models. A plethora of incentive mechanism studies have been conducted with the objective of promoting the provision of high-quality data by clients. These studies have focused on the distribution of benefits to clients. However, the incentives of federated learning are transactional in nature, and the issue of the atomicity of transactions has not been addressed. Furthermore, the data quality of individual clients participating in training varies, and they may participate negatively in training out of privacy leakage concerns.Consequently, we propose an inaugural atomistic incentive scheme with privacy preservation in the FL setting: πFL (privacy, atomic, incentive). This scheme establishes a more dependable training environment based on Shapley valuation, secure multi-party computation, and smart contracts. Consequently, it ensures that each client's contribution can be accurately measured and appropriately rewarded, improves the accuracy and efficiency of model training, and enhances the sustainability and reliability of the FL system. The efficacy of this mechanism has been demonstrated through comprehensive experimental analysis. It is evident that this mechanism not only protects the privacy of trainers and provides atomic training rewards but also improves the model performance of FL, with an accuracy improvement of at least 8%.
Kejia Chen 0007, Jiawen Zhang 0005, Xuanming Liu, Zunlei Feng, Xiaohu Yang 0001
Blockchain Res. Appl.4
2025 CoEF: Vehicular cooperative perception based on entropy theory and feature re-projection
Zunlei Feng, Gang Xiong 0001, Peijun Ye 0001, Guangmin Liu, Haina Tang, Fenghua Zhu
Expert Syst. Appl.2
2025 Deep feature response discriminative calibration
Linyun Zhou, Zunlei Feng, Mingli Song, Huiqiong Wang
Neurocomputing4
2025 Behavior capture guided engagement recognition
Yijun Bei, Songyuan Guo, Kewei Gao, Zunlei Feng
Pattern Recognit.4
2025 Target-Directed Progressive Gradient Adjusting for transfer learning
Yijun Bei, Kewei Gao, Zhuoyang Zhao, Erteng Liu, Zunlei Feng
Pattern Recognit.6
2025 Self-adaptive image-text fusion for medical image classification
Jian Zhang 0026, Kaihao He, Zunlei Feng, Shuifa Sun, Xiaoyan Sun 0006, Zhenming Yuan, Jun Yu 0002
Pattern Recognit.3
2024 ViT-Calibrator: Decision Stream Calibration for Vision Transformer
abstract
A surge of interest has emerged in utilizing Transformers in diverse vision tasks owing to its formidable performance. However, existing approaches primarily focus on optimizing internal model architecture designs that often entail significant trial and error with high burdens. In this work, we propose a new paradigm dubbed Decision Stream Calibration that boosts the performance of general Vision Transformers. To achieve this, we shed light on the information propagation mechanism in the learning procedure by exploring the correlation between different tokens and the relevance coefficient of multiple dimensions. Upon further analysis, it was discovered that 1) the final decision is associated with tokens of foreground targets, while token features of foreground target will be transmitted into the next layer as much as possible, and the useless token features of background area will be eliminated gradually in the forward propagation. 2) Each category is solely associated with specific sparse dimensions in the tokens. Based on the discoveries mentioned above, we designed a two-stage calibration scheme, namely ViT-Calibrator, including token propagation calibration stage and dimension propagation calibration stage. Extensive experiments on commonly used datasets show that the proposed approach can achieve promising results.
Zhijie Jia, Lechao Cheng, Yang Gao 0001, Jie Lei 0002, Yijun Bei, Zunlei Feng
AAAI7
2024 DGA-GNN: Dynamic Grouping Aggregation GNN for Fraud Detection
abstract
Fraud detection has increasingly become a prominent research field due to the dramatically increased incidents of fraud. The complex connections involving thousands, or even millions of nodes, present challenges for fraud detection tasks. Many researchers have developed various graph-based methods to detect fraud from these intricate graphs. However, those methods neglect two distinct characteristics of the fraud graph: the non-additivity of certain attributes and the distinguishability of grouped messages from neighbor nodes. This paper introduces the Dynamic Grouping Aggregation Graph Neural Network (DGA-GNN) for fraud detection, which addresses these two characteristics by dynamically grouping attribute value ranges and neighbor nodes. In DGA-GNN, we initially propose the decision tree binning encoding to transform non-additive node attributes into bin vectors. This approach aligns well with the GNN’s aggregation operation and avoids nonsensical feature generation. Furthermore, we devise a feedback dynamic grouping strategy to classify graph nodes into two distinct groups and then employ a hierarchical aggregation. This method extracts more discriminative features for fraud detection tasks. Extensive experiments on five datasets suggest that our proposed method achieves a 3% ~ 16% improvement over existing SOTA methods. Code is available at https://github.com/AtwoodDuan/DGA-GNN.
Mingjiang Duan, Tongya Zheng, Yang Gao 0001, Zunlei Feng, Xinyu Wang 0001
AAAI5
2024 Progressive Feature Self-Reinforcement for Weakly Supervised Semantic Segmentation
abstract
Compared to conventional semantic segmentation with pixel-level supervision, weakly supervised semantic segmentation (WSSS) with image-level labels poses the challenge that it commonly focuses on the most discriminative regions, resulting in a disparity between weakly and fully supervision scenarios. A typical manifestation is the diminished precision on object boundaries, leading to deteriorated accuracy of WSSS. To alleviate this issue, we propose to adaptively partition the image content into certain regions (e.g., confident foreground and background) and uncertain regions (e.g., object boundaries and misclassified categories) for separate processing. For uncertain cues, we propose an adaptive masking strategy and seek to recover the local information with self-distilled knowledge. We further assume that confident regions should be robust enough to preserve the global semantics, and introduce a complementary self-distillation method that constrains semantic consistency between confident regions and an augmented view with the same class labels. Extensive experiments conducted on PASCAL VOC 2012 and MS COCO 2014 demonstrate that our proposed single-stage approach for WSSS not only outperforms state-of-the-art counterparts but also surpasses multi-stage methods that trade complexity for accuracy.
Jingxuan He 0001, Lechao Cheng, Chaowei Fang, Zunlei Feng, Tingting Mu, Mingli Song
AAAI4
2024 Angle Robustness Unmanned Aerial Vehicle Navigation in GNSS-Denied Scenarios
abstract
Due to the inability to receive signals from the Global Navigation Satellite System (GNSS) in extreme conditions, achieving accurate and robust navigation for Unmanned Aerial Vehicles (UAVs) is a challenging task. Recently emerged, vision-based navigation has been a promising and feasible alternative to GNSS-based navigation. However, existing vision-based techniques are inadequate in addressing flight deviation caused by environmental disturbances and inaccurate position predictions in practical settings. In this paper, we present a novel angle robustness navigation paradigm to deal with flight deviation in point-to-point navigation tasks. Additionally, we propose a model that includes the Adaptive Feature Enhance Module, Cross-knowledge Attention-guided Module and Robust Task-oriented Head Module to accurately predict direction angles for high-precision navigation. To evaluate the vision-based navigation methods, we collect a new dataset termed as UAV_AR368. Furthermore, we design the Simulation Flight Testing Instrument (SFTI) using Google Earth to simulate different flight environments, thereby reducing the expenses associated with real flight testing. Experiment results demonstrate that the proposed model outperforms the state-of-the-art by achieving improvements of 26.0% and 45.6% in the success rate of arrival under ideal and disturbed circumstances, respectively.
Zunlei Feng, Haofei Zhang, Yang Gao 0001, Jie Lei 0002, Mingli Song
AAAI2
2024 Language Models-enhanced Semantic Topology Representation Learning For Temporal Knowledge Graph Extrapolation
abstract
Temporal Knowledge Graph (TKG) extrapolation aims to predict future missing facts based on historical information, which has exhibited both semantics and topology of events. The mainstream methods have advanced the prediction performance by exploring the potential of topology representations of TKGs based on dedicated temporal Graph Neural Networks (GNNs). Until recently, few Language Models (LM) based methods have attempted to model the semantic representations of TKGs, however, lacking specific designs for the topology information. Therefore, we propose a Semantic TOpology REpresentation learning (STORE) framework enhanced by LMs to bridge the gap between the semantics and topology of TKGs. Firstly, we tackle the challenge of long historical facts modeling by a time-aware sampling based on semantic priors to extract concise yet precise facts. Secondly, we handle the challenge of the interaction between topology and semantics by transforming graph representations into virtual tokens that are then integrated with generated prompts and fed into LMs. Finally, multi-head attention is adopted to obtain better semantic topology representations, thereby achieving joint optimization of both temporal GNNs and LMs. Extensive experiments on five datasets show that our STORE outperforms state-of-the-art GNNs- and LM-based methods.
Tianli Zhang, Tongya Zheng, Zhenbang Xiao, Zulong Chen, Liangyue Li, Zunlei Feng, Dongxiang Zhang, Mingli Song
CIKM6
2024 SecPE: Secure Prompt Ensembling for Private and Robust Large Language Models
abstract
With the growing popularity of LLMs among the general public users, privacy-preserving and adversarial robustness have become two pressing demands for LLM-based services, which have largely been pursued separately but rarely jointly. In this paper, to the best of our knowledge, we are among the first attempts towards robust and private LLM inference by tightly integrating two disconnected fields: private inference and prompt ensembling. The former protects users’ privacy by encrypting inference data transmitted and processed by LLMs, while the latter enhances adversarial robustness by yielding an aggregated output from multiple prompted LLM responses. Although widely recognized as effective individually, private inference for prompt ensembling together entails new challenges that render the naive combination of existing techniques inefficient. To overcome the hurdles, we propose SecPE, which designs efficient fully homomorphic encryption (FHE) counterparts for the core algorithmic building blocks of prompt ensembling. We conduct extensive experiments on 8 tasks to evaluate the accuracy, robustness, and efficiency of SecPE. The results show that SecPE maintains high clean accuracy and offers better robustness at the expense of merely 2.5% efficiency overhead compared to baseline private inference methods, indicating a satisfactory “accuracy-robustness-efficiency” tradeoff. For the efficiency of the encrypted Argmax operation that incurs major slowdown for prompt ensembling, SecPE is 35.4 times faster than the state-of-the-art peers, which can be of independent interest beyond this work.
Jiawen Zhang 0005, Kejia Chen 0007, Zunlei Feng, Jian Lou 0001, Mingli Song
ECAI3
2024 Improving Knowledge Distillation via Regularizing Feature Direction and Norm
Lechao Cheng, Manni Duan, Yongheng Wang, Zunlei Feng, Shu Kong
ECCV (24)5
2024 E3V-K5: An Authentic Benchmark for Redefining Video-Based Energy Expenditure Estimation
Shengxuming Zhang, Xinyu Wang 0001, Zunlei Feng, Mingli Song
ECCV (35)6
2024 Target Optimization Direction Guided Transfer Learning for Image Classification
abstract
At present, deep learning has made impressive achievements in various fields; however, effectively training deep neural networks on small data sets remains a significant challenge. Transfer learning, as a method of efficient training across multiple tasks, has been widely used to solve this problem. However, when the domain gap or the data volume difference between the two tasks is too large, the transfer learning may not perform well, and other optimization methods will be required to improve the performance. In this paper, we propose a new transfer learning method guided by the direction of objective optimization from the perspective of gradient. This method guides the gradient direction of the source task towards the gradient direction of the target task. In several similar and conflicting tasks, this method has achieved good results in efficiency and performance. In comparison with other transfer learning methods, the results shown by this method are generally better.
Kelvin Ting Zuo Han, Shengxuming Zhang, Gerard Marcos Freixas, Zunlei Feng, Cheng Jin 0001
ICASSP4
2024 Dynamic Neural Response Tuning
abstract
Artificial Neural Networks (ANNs) have gained widespread applications across various areas in recent years. The ANN design was initially inspired by principles of biology. The biological neural network's fundamental response process comprises information transmission and aggregation. The information transmission in biological neurons is often achieved by triggering action potentials that propagate through axons. ANNs utilize activation mechanisms to simulate such biological behavior. However, previous studies have only considered static response conditions, while the biological neuron's response conditions are typically dynamic, depending on multiple factors such as neuronal properties and the real-time environment. Therefore, the dynamic response conditions of biological neurons could help improve the static ones of existing activations in ANNs. Additionally, the biological neuron's aggregated response exhibits high specificity for different categories, allowing the nervous system to differentiate and identify objects. Inspired by these biological patterns, we propose a novel Dynamic Neural Response Tuning (DNRT) mechanism, which aligns the response patterns of ANNs with those of biological neurons. DNRT comprises Response-Adaptive Activation (RAA) and Aggregated Response Regularization (ARR), mimicking the biological neuron's information transmission and aggregation behaviors. RAA dynamically adjusts the response condition based on the characteristics and strength of the input signal. ARR is devised to enhance the network's ability to learn category specificity by imposing constraints on the network's response distribution. Extensive experimental studies indicate that the proposed DNRT is highly interpretable, applicable to various mainstream network architectures, and can achieve remarkable performance compared with existing neural response mechanisms in multiple tasks and domains. Code is available at https://github.com/horrible-dong/DNRT.
Linyun Zhou, Zunlei Feng, Mingli Song
ICLR5
2024 Visual-guided Query with Temporal Interaction for Video Object Segementation
abstract
The task of referring video object segmentation (RVOS) involves segmenting objects in video frames based on a given text description. However, most existing approaches treat the text directly as a query, neglecting the valuable visual and temporal information from the video. This limitation may cause the query unable to accurately perceive the target object. To address this issue, we introduce a visual-guided query with temporal interaction for referring video object segmentation (VQTI) approach. Our method capitalizes on frame-level features and video-level features to guide the query generation process, resulting in an enhanced perception of the target object. In addition, we introduce a spectral-guided segmentation optimizer module to enhance the fine-grained information, leading to more precise segmentation masks. Extensive experiments shows competitive performance against state-of-the-art approaches.
Jiaxin Qiu, Guoyu Yang, Jie Lei 0002, Zunlei Feng, Ronghua Liang
ICME4
2024 Critical Feature Sifting and Dynamic Aggregation for Anomalous Audio Sequence Detection
Erteng Liu, Kewei Gao, Jianhai Chen, Yijun Bei, Zunlei Feng
ICONIP (1)7
2024 MFCA: Multimodal Object Detection Based on Feature Calibration and Aggregation
Jie Lei 0002, Guoyu Yang, Zunlei Feng, Ronghua Liang
ICONIP (8)6
2024 Discriminative Feature Decoupling Enhancement for Speech Forgery Detection
Yijun Bei, Erteng Liu, Yang Gao 0001, Kewei Gao, Zunlei Feng
IJCAI7
2024 Improving Adversarial Robustness via Feature Pattern Consistency Constraint
Jiacong Hu, Jingwen Ye, Zunlei Feng, Jiazhen Yang, Shunyu Liu 0001, Xiaotian Yu, Lingxiang Jia, Mingli Song
IJCAI3
2024 Hundredfold Accelerating for Pathological Images Diagnosis and Prognosis through Self-reform Critical Region Focusing
Xiaotian Yu, Haoming Luo, Jiacong Hu, Xiuming Zhang, Yijun Bei, Mingli Song, Zunlei Feng
IJCAI9
2024 Deep Kernel Calibration
abstract
Korbinian Brodmann argued that the brain regions with different cytoarchitectures exhibited different cognitive functions, which are widely recognized as "Brodmann areas" in neuroscience. Inspired by this theory, we observe from experiments that the different well-trained convolutional kernels also hold their unique functionalities, indicated by their output activations that highly concentrate on the "different certain ranges". This discovery motivates us to devise a deep kernel-by-kernel feature calibration mechanism that "calibrates" the output distribution of a pre-trained convolutional kernel for a more concentrated activation interval, by eliminating the outlier activations. Towards this end, we develop two dedicated Brodmann calibration forms, termed as Hard Kernel Calibration (HKC) and Soft Kernel Calibration (SKC) that simply filters the overrange activations, or adaptively condense the activations in a weighted manner, respectively. As a flexible plug-and-play module, the proposed calibration demonstrates encouraging results on five benchmarks across sixteen network architectures and also triggers the functionality of significantly enhanced model robustness against adversarial attacks. The developed code is publicly available at https://github.com/tcmyxc/DKC.
Zunlei Feng, Jie Lei 0002, Huiqiong Wang, Zhongle Xie
IJCNN1
2024 GAN Doctor: Diagnosing and Treating Inherent Semantic Errors
abstract
Generative Adversarial Network (GAN), as a popular generative model in the field of Artificial Intelligence Generated Content (AIGC), has been intensively developed in previous research, with significant improvements in the quality and diversity of image generation. However, there are still many cases where the results are not satisfactory. A primary concern pertains to the chaotic and blurred local details within the generated images. In this work, through the diagnosis and analysis of high-quality and low-quality images produced by the GAN model, we identified that this issue stems from inherent semantic errors of the GAN, that is, convolutional kernels responsible for certain semantics are not properly involved in the generation process of corresponding image regions. To this end, we propose a straightforward yet effective treatment method, which constrains each image region to be generated by its corresponding semantic convolutional kernels. Experimental results demonstrate that our proposed optimization method can improve the issue of chaotic and blurred local regions in generated images and enhance the overall generation quality. Our work pioneers a novel paradigm for diagnosing and treating GANs, driving the research development and practical application of AIGC image generation technology.
Chengji Shen, Zunlei Feng, Zhongle Xie, Jie Lei 0002, Huiqiong Wang, Mingli Song
IJCNN2
2024 Loose Lesion Location Self-supervision Enhanced Colorectal Cancer Diagnosis
Tianhong Gao, Jie Song 0011, Xiaotian Yu, Shengxuming Zhang, Xiuming Zhang, Zipeng Zhong, Mingli Song, Zunlei Feng
MICCAI (11)12
2024 Fire and Smoke Detection with Burning Intensity Representation
Xiaoyi Han, Yanfei Wu, Nan Pu, Zunlei Feng, Qifei Zhang 0001, Yijun Bei, Lechao Cheng
MMAsia4
2024 Contextual Augmentation with Bias Adaptive for Few-Shot Video Object Segmentation
Shuaiwei Wang, Jie Lei 0002, Zunlei Feng, Ronghua Liang
MMM (1)4
2024 Dynamic-Static Graph Convolutional Network for Video-Based Facial Expression Recognition
Fahong Wang, Jie Lei 0002, Zeyu Zou, Zunlei Feng, Ronghua Liang
MMM (2)8
2024 Transformer Doctor: Diagnosing and Treating Vision Transformers
abstract
Due to its powerful representational capabilities, Transformers have gradually become the mainstream model in the field of machine vision. However, the vast and complex parameters of Transformers impede researchers from gaining a deep understanding of their internal mechanisms, especially error mechanisms. Existing methods for interpreting Transformers mainly focus on understanding them from the perspectives of the importance of input tokens or internal modules, as well as the formation and meaning of features. In contrast, inspired by research on information integration mechanisms and conjunctive errors in the biological visual system, this paper conducts an in-depth exploration of the internal error mechanisms of Transformers. We first propose an information integration hypothesis for Transformers in the machine vision domain and provide substantial experimental evidence to support this hypothesis. This includes the dynamic integration of information among tokens and the static integration of information within tokens in Transformers, as well as the presence of conjunctive errors therein. Addressing these errors, we further propose heuristic dynamic integration constraint methods and rule-based static integration constraint methods to rectify errors and ultimately improve model performance. The entire methodology framework is termed as Transformer Doctor, designed for diagnosing and treating internal errors within transformers. Through a plethora of quantitative and qualitative experiments, it has been demonstrated that Transformer Doctor can effectively address internal errors in transformers, thereby enhancing model performance.
Jiacong Hu, Hao Chen 0041, Kejia Chen 0007, Yang Gao 0001, Jingwen Ye, Xingen Wang, Mingli Song, Zunlei Feng
NeurIPS8
2024 Vision Mamba Mender
abstract
Mamba, a state-space model with selective mechanisms and hardware-aware architecture, has demonstrated outstanding performance in long sequence modeling tasks, particularly garnering widespread exploration and application in the field of computer vision. While existing works have mixed opinions of its application in visual tasks, the exploration of its internal workings and the optimization of its performance remain urgent and worthy research questions given its status as a novel model. Existing optimizations of the Mamba model, especially when applied in the visual domain, have primarily relied on predefined methods such as improving scanning mechanisms or integrating other architectures, often requiring strong priors and extensive trial and error. In contrast to these approaches, this paper proposes the Vision Mamba Mender, a systematic approach for understanding the workings of Mamba, identifying flaws within, and subsequently optimizing model performance. Specifically, we present methods for predictive correlation analysis of Mamba's hidden states from both internal and external perspectives, along with corresponding definitions of correlation scores, aimed at understanding the workings of Mamba in visual recognition tasks and identifying flaws therein. Additionally, tailored repair methods are proposed for identified external and internal state flaws to eliminate them and optimize model performance. Extensive experiments validate the efficacy of the proposed methods on prevalent Mamba architectures, significantly enhancing Mamba's performance.
Jiacong Hu, Anda Cao, Zunlei Feng, Shengxuming Zhang, Lingxiang Jia, Mingli Song
NeurIPS3
2024 Model LEGO: Creating Models Like Disassembling and Assembling Building Blocks
abstract
With the rapid development of deep learning, the increasing complexity and scale of parameters make training a new model increasingly resource-intensive. In this paper, we start from the classic convolutional neural network (CNN) and explore a paradigm that does not require training to obtain new models. Similar to the birth of CNN inspired by receptive fields in the biological visual system, we draw inspiration from the information subsystem pathways in the biological visual system and propose Model Disassembling and Assembling (MDA). During model disassembling, we introduce the concept of relative contribution and propose a component locating technique to extract task-aware components from trained CNN classifiers. For model assembling, we present the alignment padding strategy and parameter scaling strategy to construct a new model tailored for a specific task, utilizing the disassembled task-aware components. The entire process is akin to playing with LEGO bricks, enabling arbitrary assembly of new models, and providing a novel perspective for model creation and reuse. Extensive experiments showcase that task-aware components disassembled from CNN classifiers or new models assembled using these components closely match or even surpass the performance of the baseline, demonstrating its promising results for model reuse. Furthermore, MDA exhibits diverse potential applications, with comprehensive experiments exploring model decision route analysis, model compression, knowledge distillation, and more.
Jiacong Hu, Jingwen Ye, Yang Gao 0001, Xingen Wang, Zunlei Feng, Mingli Song
NeurIPS6
2024 Association Pattern-aware Fusion for Biological Entity Relationship Prediction
abstract
Deep learning-based methods significantly advance the exploration of associations among triple-wise biological entities (e.g., drug-target protein-adverse reaction), thereby facilitating drug discovery and safeguarding human health. However, existing researches only focus on entity-centric information mapping and aggregation, neglecting the crucial role of potential association patterns among different entities. To address the above limitation, we propose a novel association pattern-aware fusion method for biological entity relationship prediction, which effectively integrates the related association pattern information into entity representation learning. Additionally, to enhance the missing information of the low-order message passing, we devise a bind-relation module that considers the strong bind of low-order entity associations. Extensive experiments conducted on three biological datasets quantitatively demonstrate that the proposed method achieves about 4%-23% hit@1 improvements compared with state-of-the-art baselines. Furthermore, the interpretability of association patterns is elucidated in detail, thus revealing the intrinsic biological mechanisms and promoting it to be deployed in real-world scenarios. Our data and code are available at https://github.com/hry98kki/PatternBERP.
Lingxiang Jia, Yuchen Ying, Zunlei Feng, Zipeng Zhong, Shaolun Yao, Jiacong Hu, Mingjiang Duan, Xingen Wang, Jie Song 0011, Mingli Song
NeurIPS3
2024 Dual-Perspective Activation: Efficient Channel Denoising via Joint Forward-Backward Criterion for Artificial Neural Networks
abstract
The design of Artificial Neural Network (ANN) is inspired by the working patterns of the human brain. Connections in biological neural networks are sparse, as they only exist between few neurons. Meanwhile, the sparse representation in ANNs has been shown to possess significant advantages. Activation responses of ANNs are typically expected to promote sparse representations, where key signals get activated while irrelevant/redundant signals are suppressed. It can be observed that samples of each category are only correlated with sparse and specific channels in ANNs. However, existing activation mechanisms often struggle to suppress signals from other irrelevant channels entirely, and these signals have been verified to be detrimental to the network's final decision. To address the issue of channel noise interference in ANNs, a novel end-to-end trainable Dual-Perspective Activation (DPA) mechanism is proposed. DPA efficiently identifies irrelevant channels and applies channel denoising under the guidance of a joint criterion established online from both forward and backward propagation perspectives while preserving activation responses from relevant channels. Extensive experiments demonstrate that DPA successfully denoises channels and facilitates sparser neural representations. Moreover, DPA is parameter-free, fast, applicable to many mainstream ANN architectures, and achieves remarkable performance compared to other existing activation counterparts across multiple tasks and domains. Code is available at https://github.com/horrible-dong/DPA.
Chenchao Gao, Zunlei Feng, Jie Lei 0002, Bingde Hu, Xingen Wang, Mingli Song
NeurIPS3
2024 Behavior Capture Based Explainable Engagement Recognition
Yijun Bei, Songyuan Guo, Kewei Gao, Zunlei Feng, Yining Tong, Weimin Cai, Lechao Cheng
PRCV (10)4
2024 Benchmarking Multi-Scene Fire and Smoke Detection
Xiaoyi Han, Nan Pu, Zunlei Feng, Yijun Bei, Qifei Zhang 0001, Lechao Cheng
PRCV (11)3
2024 Multi-layer Tuning CLIP for Few-Shot Image Classification
Jinsong Geng, Cenyu Liu, Zunlei Feng, Yijun Bei
PRCV (5)5
2024 Graph Neural Networks-based hybrid framework for predicting particle crushing strength
Tongya Zheng, Tianli Zhang, Qingzheng Guan, Zunlei Feng, Mingli Song, Chun Chen 0001
Expert Syst. Appl.5
2024 Noise is the fatal poison: A Noise-aware Network for noisy dataset classification
Xiaotian Yu, Shengxuming Zhang, Lingxiang Jia, Mingli Song, Zunlei Feng
Neurocomputing6
2024 PatchDetector: Pluggable and non-intrusive patch for small object detection
Linyun Zhou, Shengxuming Zhang, Zunlei Feng, Mingli Song
Neurocomputing5
2024 NRD-Net: a noise-resistant distillation network for accurate diagnosis of prostate cancer with bi-parametric MRI images
Xiangtong Du, Ximing Wang, Zunlei Feng, Hai Deng
Multim. Tools Appl.4
2024 Life regression based patch slimming for vision transformers
Tianqi Shi, Lechao Cheng, Zunlei Feng, Mingli Song
Neural Networks6
2024 Interaction Pattern Disentangling for Multi-Agent Reinforcement Learning
abstract
Deep cooperative multi-agent reinforcement learning has demonstrated its remarkable success over a wide spectrum of complex control tasks. However, recent advances in multi-agent learning mainly focus on value decomposition while leaving entity interactions still intertwined, which easily leads to over-fitting on noisy interactions between entities. In this work, we introduce a novel interactiOn Pattern disenTangling (OPT) method, to disentangle the entity interactions into interaction prototypes, each of which represents an underlying interaction pattern within a subgroup of the entities. OPT facilitates filtering the noisy interactions between irrelevant entities and thus significantly improves generalizability as well as interpretability. Specifically, OPT introduces a sparse disagreement mechanism to encourage sparsity and diversity among discovered interaction prototypes. Then the model selectively restructures these prototypes into a compact interaction pattern by an aggregator with learnable weights. To alleviate the training instability issue caused by partial observability, we propose to maximize the mutual information between the aggregation weights and the history behaviors of each agent. Experiments on single-task, multi-task and zero-shot benchmarks demonstrate that the proposed method yields results superior to the state-of-the-art counterparts.
Shunyu Liu 0001, Jie Song 0011, Yihe Zhou, Na Yu 0001, Kai-Xuan Chen 0001, Zunlei Feng, Mingli Song
IEEE Trans. Pattern Anal. Mach. Intell.6
2024 DataMap: Dataset transferability map for medical image classification
Xiangtong Du, Zhidong Liu, Zunlei Feng, Hai Deng
Pattern Recognit.3
2024 Asymptotic Feature Pyramid Network for Labeling Pixels and Regions
abstract
Multi-scale features are crucial in encoding objects with varying scales in vision tasks. The classic top-down and bottom-up feature pyramid networks are a common strategy for multi-scale feature extraction. However, these approaches suffer from the loss or degradation of feature information, which impairs the fusion effect of non-adjacent levels. In this paper, we propose an Asymptotic Feature Pyramid Network (AFPN) that supports direct interaction between non-adjacent levels. AFPN starts by fusing two adjacent low-level features and asymptotic incorporates higher-level features into the fusion process. This fusion way avoids the significant semantic gap between non-adjacent levels. Adaptive spatial fusion operation is further used to mitigate potential multi-object information conflicts during feature fusion at each spatial location. To reduce parameters, computational requirements, and inference speed, we propose a Lightweight Asymptotic Feature Pyramid Network (LightAFPN) that uses the concept of reparametrization. We evaluate the proposed method on the MS-COCO 2017, PASCAL VOC and Cityscapes datasets in both object detection and semantic segmentation frameworks. Experimental evaluation shows that our method achieves more competitive results than other state-of-the-art feature pyramid networks. The code is available at https://github.com/gyyang23/AFPN.
Guoyu Yang, Jie Lei 0002, Zunlei Feng, Ronghua Liang
IEEE Trans. Circuits Syst. Video Technol.4
2024 Spatiotemporal-Augmented Graph Neural Networks for Human Mobility Simulation
abstract
Human mobility patterns have shown significant applications in policy-decision scenarios and economic behavior researches. The human mobility simulation task aims to generate human mobility trajectories given a small set of trajectory data, which have aroused much concern due to the scarcity and sparsity of human mobility data. Existing methods mostly rely on the static relationships of locations, while largely neglect the dynamic spatiotemporal effects of locations. On the one hand, spatiotemporal correspondences of visit distributions reveal the spatial proximity and the functionality similarity of locations. On the other hand, the varying durations in different locations hinder the iterative generation process of the mobility trajectory. Therefore, we propose a novel framework to model the dynamic spatiotemporal effects of locations, namelySpatioTemporal-Augmented gRaph neural networks (STAR). The STAR framework designs various spatiotemporal graphs to capture the spatiotemporal correspondences and builds a novel dwell branch to simulate the varying durations in locations, which is finally optimized in an adversarial manner. The comprehensive experiments over four real datasets for the human mobility simulation have verified the superiority of STAR tostate-of-the-artmethods. Our code is available athttps://github.com/Star607/STAR-TKDE.
Yu Wang 0176, Tongya Zheng, Shunyu Liu 0001, Zunlei Feng, Kai-Xuan Chen 0001, Yunzhi Hao, Mingli Song
IEEE Trans. Knowl. Data Eng.4
2024 Transition Propagation Graph Neural Networks for Temporal Networks
abstract
Researchers of temporal networks (e.g., social networks and transaction networks) have been interested in mining dynamic patterns of nodes from their diverse interactions. Inspired by recently powerful graph mining methods like skip-gram models and graph neural networks (GNNs), existing approaches focus on generating temporal node embeddings sequentially with nodes' sequential interactions. However, the sequential modeling of previous approaches cannot handles the transition structure between nodes' neighbors with limited memorization capacity. In detail, an effective method for the transition structures is required to both model nodes' personalized patterns adaptively and capture node dynamics accordingly. In this article, we propose a method, namely t ransition p ropagation g raph n eural n etworks (TIP-GNN), to tackle the challenges of encoding nodes' transition structures. The proposed TIP-GNN focuses on the bilevel graph structure in temporal networks: besides the explicit interaction graph, a node's sequential interactions can also be constructed as a transition graph. Based on the bilevel graph, TIP-GNN further encodes transition structures by multistep transition propagation and distills information from neighborhoods by a bilevel graph convolution. Experimental results over various temporal networks reveal the efficiency of our TIP-GNN, with at most 7.2% improvements of accuracy on temporal link prediction. Extensive ablation studies further verify the effectiveness and limitations of the transition propagation module. Our code is available at https://github.com/doujiang-zheng/TIP-GNN.
Tongya Zheng, Zunlei Feng, Tianli Zhang, Yunzhi Hao, Mingli Song, Xingen Wang, Xinyu Wang 0001, Ji Zhao 0016, Chun Chen 0001
IEEE Trans. Neural Networks Learn. Syst.2
2024 Self-Adaptive Clothing Mapping Based Virtual Try-on
abstract
VTON (Virtual Try-ON), as an innovative visual application in e-commerce scenarios with great commercial value, has been widely studied in recent years. Due to its better robustness and realistic effect, deformation-synthesize-based VTON has become the dominant approach in this field. Existing clothing deformation techniques optimize the mapping relations between the original clothing image and the ground truth (GT) image of the worn clothing. However, there are color differences between the original and GT clothing images caused by lighting, warping, and occlusion. The color differences may lead to misaligned clothing mapping by only minimizing the cost of pixel value difference. Another drawback is that taking the parsing prediction as GT will bring alignment remnant, rooting in the processing order of parsing and deformation. Aiming above two drawbacks, we put forward SAME-VTON (Self-Adaptive clothing Mapping basEd Virtual Try-ON) for achieving realistic virtual try-on results. The core of SAME-VTON is the self-adaptive clothing mapping technique, composed of two parts: a color-adaptive clothing mapping module and a parsing-adaptive prediction process. In the color-adaptive clothing mapping module, we map each pixel of the target clothing with a combination of multiple pixel values from the original clothing image, which considers both the position and color changes. Furthermore, different combination weights are learned to increase the diversity of color mapping. In the parsing-adaptive prediction process, the color-adaptive clothing mapping module is adopted to deform clothing first, then the human parsing result is predicted under the reference of the deformed clothing, which can avoid alignment remnant. Extensive experiments demonstrate that the proposed SAME-VTON with the self-adaptive clothing mapping technique can achieve optimal mapping in the case of large color differences and obtain superior results compared with existing deformation-synthesize-based VTON.
Chengji Shen, Zhenjiang Liu, Xin Gao 0032, Zunlei Feng, Mingli Song
ACM Trans. Multim. Comput. Commun. Appl.4
2024 HairStyle Editing via Parametric Controllable Strokes
abstract
In this work, we propose a stroke-based hairstyle editing network, dubbed HairstyleNet, allowing users to conveniently change the hairstyles of an image in an interactive fashion. Different from previous works, we simplify the hairstyle editing process where users can manipulate local or entire hairstyles by adjusting the parameterized hair regions. Our HairstyleNet consists of two stages: a stroke parameterization stage and a stroke-to-hair generation stage. In the stroke parameterization stage, we first introduce parametric strokes to approximate the hair wisps, where the stroke shape is controlled by a quadratic Bézier curve and a thickness parameter. Since rendering strokes with thickness to an image is not differentiable, we opt to leverage a neural renderer to construct the mapping from stroke parameters to a stroke image. Thus, the stroke parameters can be directly estimated from hair regions in a differentiable way, enabling us to flexibly edit the hairstyles of input images. In the stroke-to-hair generation stage, we design a hairstyle refinement network that first encodes coarsely composed images of hair strokes, face, and background into latent representations and then generates high-fidelity face images with desirable new hairstyles from the latent codes. Extensive experiments demonstrate that our HairstyleNet achieves state-of-the-art performance and allows flexible hairstyle manipulation.
Xinhui Song, Chen Liu 0028, Youyi Zheng, Zunlei Feng, Lincheng Li, Kun Zhou 0001, Xin Yu 0002
IEEE Trans. Vis. Comput. Graph.4
2023 Contrastive Identity-Aware Learning for Multi-Agent Value Decomposition
abstract
Value Decomposition (VD) aims to deduce the contributions of agents for decentralized policies in the presence of only global rewards, and has recently emerged as a powerful credit assignment paradigm for tackling cooperative Multi-Agent Reinforcement Learning (MARL) problems. One of the main challenges in VD is to promote diverse behaviors among agents, while existing methods directly encourage the diversity of learned agent networks with various strategies. However, we argue that these dedicated designs for agent networks are still limited by the indistinguishable VD network, leading to homogeneous agent behaviors and thus downgrading the cooperation capability. In this paper, we propose a novel Contrastive Identity-Aware learning (CIA) method, explicitly boosting the credit-level distinguishability of the VD network to break the bottleneck of multi-agent diversity. Specifically, our approach leverages contrastive learning to maximize the mutual information between the temporal credits and identity representations of different agents, encouraging the full expressiveness of credit assignment and further the emergence of individualities. The algorithm implementation of the proposed CIA module is simple yet effective that can be readily incorporated into various VD architectures. Experiments on the SMAC benchmarks and across different VD backbones demonstrate that the proposed method yields results superior to the state-of-the-art counterparts. Our code is available at https://github.com/liushunyu/CIA.
Shunyu Liu 0001, Yihe Zhou, Jie Song 0011, Tongya Zheng, Kai-Xuan Chen 0001, Tongtian Zhu, Zunlei Feng, Mingli Song
AAAI7
2023 Drift-aware Anomaly Detection for Non-stationary Time Series
abstract
Anomaly detection of time series is vital in various scenarios with explosively growing time series data. However, the non-stationary time series degrade the performance of current anomaly detection methods, where data drift causes unpredictable changes. This paper proposes a Drift-aware Anomaly Detection (DAD) method for detecting anomalies in non-stationary time series. DAD adopts a self-attention mechanism to learn an embedding, distinguishing the anomaly embeddings from the normal embeddings. Next, the KL divergence calculates the drift deviation between two data segments at adjacent periods. Then, the drift deviation module combined with the latent vector which is used to reconstruct the original vector. During the encoding stage of the time series, the latent code is modeled using different Gaussian mixture distributions and the data reconstruction error at each time tick is regarded as an anomaly metric. Furthermore, we propose a new metric to measure the degree of drift deviation for a dataset used for a fair experiment comparison. Experimental results on several public datasets and a newly collected sensor dataset demonstrate that for the non-stationary time series anomaly detection task, DAD outperforms state-of-the-art anomaly detection models up to 11.5% on the F1score.
Yang Gao 0001, Ying Li 0097, Zunlei Feng, Mingli Song, Chun Chen 0001
IEEE Big Data4
2023 How to Prevent the Continuous Damage of Noises to Model Training?
abstract
Deep learning with noisy labels is challenging and inevitable in many circumstances. Existing methods reduce the impact of mislabeled samples by reducing loss weights or screening, which highly rely on the model's superior discriminative power for identifying mislabeled samples. However, in the training stage, the trainee model is imperfect and will wrongly predict some mislabeled samples, which cause continuous damage to the model training. Consequently, there is a large performance gap between existing anti-noise models trained with noisy samples and models trained with clean samples. In this paper, we put forward a Gradient Switching Strategy (GSS) to prevent the continuous damage of mislabeled samples to the classifier. Theoretical analysis shows that the damage comes from the misleading gradient direction computed from the mislabeled samples. The trainee model will deviate from the correct optimization direction under the influence of the accumulated misleading gradient of mislabeled samples. To address this problem, the proposed GSS alleviates the damage by switching the gradient direction of each sample based on the gradient direction pool, which contains all-class gradient directions with different probabilities. During training, each gradient direction pool is updated iteratively, which assigns higher probabilities to potential principal directions for high-confidence samples. Conversely, uncertain samples are forced to explore in different directions rather than mislead model in a fixed direction. Extensive experiments show that GSS can achieve comparable performance with a model trained with clean data. Moreover, the proposed GSS is pluggable for existing frameworks. This idea of switching gradient directions provides a new perspective for future noisy-label learning.
Xiaotian Yu, Tianqi Shi, Zunlei Feng, Mingli Song
CVPR4
2023 A Loopback Network for Explainable Microvascular Invasion Classification
abstract
Microvascular invasion (MVI) is a critical factor for prognosis evaluation and cancer treatment. The current diagnosis of MVI relies on pathologists to manually find out cancerous cells from hundreds of blood vessels, which is time-consuming, tedious, and subjective. Recently, deep learning has achieved promising results in medical image analysis tasks. However, the unexplainability of black box models and the requirement of massive annotated samples limit the clinical application of deep learning based diagnostic methods. In this paper, aiming to develop an accurate, objective, and explainable diagnosis tool for MVI, we propose a Loopback Network (LoopNet) for classifying MVI efficiently. With the image-level category annotations of the collected Pathologic Vessel Image Dataset (PVID), LoopNet is devised to be composed binary classification branch and cell locating branch. The latter is devised to locate the area of cancerous cells, regular non-cancerous cells, and background. For healthy samples, the pseudo masks of cells supervise the cell locating branch to distinguish the area of regular non-cancerous cells and background. For each MVI sample, the cell locating branch predicts the mask of cancerous cells. Then the masked cancerous and non-cancerous areas of the same sample are input back to the binary classification branch separately. The loopback between two branches enables the category label to supervise the cell locating branch to learn the locating ability for cancerous areas. Experiment results show that the proposed LoopNet achieves 97.5% accuracy on MVI classification. Surprisingly, the proposed loopback mechanism not only enables LoopNet to predict the cancerous area but also facilitates the classification backbone to achieve better classification performance.
Shengxuming Zhang, Tianqi Shi, Xiuming Zhang, Jie Lei 0002, Zunlei Feng, Mingli Song
CVPR6
2023 LoSS: Local Structural Separation Hypergraph Convolutional Neural Network
abstract
Graph classification is a classic problem with practical applications in many real-life scenes. Existing graph neural networks, including GCN, GAT, and GIN, are proposed to extract useful features from complex graph structures. However, most existing methods’ feature extraction and aggregation inevitably mix the useful and redundant features, which will disturb the final classification performance. In this paper, to handle the above drawback, we put forward the Local Structural Separation Hypergraph Convolutional Neural Network (LoSS) based on two discoveries: most graph classification tasks only focus on a few groups of adjacent nodes, and different categories have their specific high response bits in graph embeddings. In LoSS, we first decouple the original graph into different hypergraphs and aggregate the features in each substructure, which aims to find useful features for the final classification. Next, the low-correlation feature suppression strategy is devised to suppress the irrelevant node-level and bit-level features in the forward inference process, effectively reducing the disturbance of redundant features. Experiments on five datasets show that the proposed LoSS can effectively locate and aggregate useful hypergraph features and achieve SOTA performance compared with existing methods.
Bingde Hu, Yang Gao 0001, Zunlei Feng, Mingli Song, Xinyu Wang 0001, Ying Li 0097
ECAI3
2023 Model Doctor for Diagnosing and Treating Segmentation Error
abstract
Despite the remarkable progress in semantic segmentation tasks with the advancement of deep neural networks, existing U-shaped hierarchical typical segmentation networks still suffer from local misclassification of categories and inaccurate target boundaries. In an effort to alleviate this issue, we propose a Model Doctor for semantic segmentation problems. The Model Doctor is designed to diagnose the aforementioned problems in existing pre-trained models and treat them without introducing additional data, with the goal of refining the parameters to achieve better performance. Extensive experiments on several benchmark datasets demonstrate the effectiveness of our method. Code is available at https://github.com/zhijiejia/DoctorForSeg.
Zhijie Jia, Kaiwen Hu, Lechao Cheng, Zunlei Feng, Mingli Song
ICIP5
2023 Team DETR: Guide Queries as a Professional Team in Detection Transformers
abstract
Recent proposed DETR variants have made tremendous progress in various scenarios due to their streamlined processes and remarkable performance. However, the learned queries usually explore the global context to generate the final set prediction, resulting in redundant burdens and unfaithful results. More specifically, a query is commonly responsible for objects of different scales and positions, which is a challenge for the query itself, and will cause spatial resource competition among queries. To alleviate this issue, we propose Team DETR, which leverages query collaboration and position constraints to embrace objects of interest more precisely. We also dynamically cater to each query member’s prediction preference, offering the query better scale and spatial priors. In addition, the proposed Team DETR is flexible enough to be adapted to other existing DETR variants without increasing parameters and calculations. Extensive experiments on the COCO dataset showcase that Team DETR achieves remarkable gains, especially for small and large objects. Code is available at https://github.com/horrible-dong/TeamDETR.
Linyun Zhou, Lechao Cheng, Zunlei Feng, Mingli Song
ICIP5
2023 CKR-Calibrator: Convolution Kernel Robustness Evaluation and Calibration
Yijun Bei, Jinsong Geng, Erteng Liu, Kewei Gao, Wenqi Huang 0002, Zunlei Feng
ICONIP (5)6
2023 Propheter: Prophetic Teacher Guided Long-Tailed Distribution Learning
Yongcheng Jing, Linyun Zhou, Wenqi Huang 0002, Lechao Cheng, Zunlei Feng, Mingli Song
ICONIP (4)6
2023 Improving Expressivity of GNNs with Subgraph-specific Factor Embedded Normalization
abstract
Graph Neural Networks~(GNNs) have emerged as a powerful category of learning architecture for handling graph-structured data. However, existing GNNs typically ignore crucial structural characteristics in node-induced subgraphs, which thus limits their expressiveness for various downstream tasks. In this paper, we strive to strengthen the representative capabilities of GNNs by devising a dedicated plug-and-play normalization scheme, termed as SUbgraph-sPEcific FactoR Embedded Normalization (SuperNorm), that explicitly considers the intra-connection information within each node-induced subgraph. To this end, we embed the subgraph-specific factor at the beginning and the end of the standard BatchNorm, as well as incorporate graph instance-specific statistics for improved distinguishable capabilities. In the meantime, we provide theoretical analysis to support that, with the elaborated SuperNorm, an arbitrary GNN is at least as powerful as the 1-WL test in distinguishing non-isomorphism graphs. Furthermore, the proposed SuperNorm scheme is also demonstrated to alleviate the over-smoothing phenomenon. Experimental results related to predictions of graph, node, and link properties on the eight popular datasets demonstrate the effectiveness of the proposed method. The code is available at https://github.com/chenchkx/SuperNorm.
Kai-Xuan Chen 0001, Shunyu Liu 0001, Tongtian Zhu, Ji Qiao, Yingjie Tian 0002, Tongya Zheng, Haofei Zhang, Zunlei Feng, Jingwen Ye, Mingli Song
KDD9
2023 Temporal Aggregation with Context Focusing for Few-Shot Video Object Detection
abstract
Few-shot video object detection focuses on finding all the objects in a given query video that belong to the same class, given only a few support images of the target object in an unseen class. Unfortunately, due to the object blur or occlusion in video frames, using single-frame object detection directly will greatly limit the accuracy. The issue is significantly worse in few-shot settings due to insufficient support and timedomain information. In this paper, we propose a temporal aggregation with context focusing framework (TACF) for few-shot video object detection, which aims to fully use the information between support images and adjacent video frames. The context focusing module effectively encodes the target object in adjacent frames according to the support images. Afterward, the temporal aggregation module implicitly extracts the most similar ROI features from these adjacent frames to obtain the target proposals. In the end, the matching network determines the category and bounding box by calculating the distance with the support images. Extensive experimental evaluations on FSVOD and FSYTV databases show that our method achieves more competitive results than image-based methods, naive video-based extensions, and the state-of-the-art few-shot video object detection method.
Jie Lei 0002, Fahong Wang, Zunlei Feng, Ronghua Liang
SMC4
2023 AFPN: Asymptotic Feature Pyramid Network for Object Detection
abstract
Multi-scale features are of great importance in encoding objects with scale variance in object detection tasks. A common strategy for multi-scale feature extraction is adopting the classic top-down and bottom-up feature pyramid networks. However, these approaches suffer from the loss or degradation of feature information, impairing the fusion effect of non-adjacent levels. This paper proposes an asymptotic feature pyramid network (AFPN) to support direct interaction at non-adjacent levels. AFPN is initiated by fusing two adjacent low-level features and asymptotically incorporates higher-level features into the fusion process. In this way, the larger semantic gap between non-adjacent levels can be avoided. Given the potential for multi-object information conflicts to arise during feature fusion at each spatial location, adaptive spatial fusion operation is further utilized to mitigate these inconsistencies. We incorporate the proposed AFPN into both two-stage and one-stage object detection frameworks and evaluate with the MS-COCO 2017 validation and test datasets. Experimental evaluation shows that our method achieves more competitive results than other state-of-the-art feature pyramid networks. The code is available at https://github.com/gyyang23/AFPN.
Guoyu Yang, Jie Lei 0002, Zhikuan Zhu, Siyu Cheng, Zunlei Feng, Ronghua Liang
SMC5
2023 Disassembling Convolutional Segmentation Network
Kaiwen Hu, Fangyuan Mao, Xinhui Song, Lechao Cheng, Zunlei Feng, Mingli Song
Int. J. Comput. Vis.6
2023 Reinforcement learning based web crawler detection for diversity and dynamics
Yang Gao 0001, Zunlei Feng, Mingli Song, Xingen Wang, Xinyu Wang 0001, Chun Chen 0001
Neurocomputing2
2023 DCAM: Disturbed class activation maps for weakly supervised semantic segmentation
Jie Lei 0002, Guoyu Yang, Shuaiwei Wang, Zunlei Feng, Ronghua Liang
J. Vis. Commun. Image Represent.4
2023 Conservative-Progressive Collaborative Learning for Semi-Supervised Semantic Segmentation
abstract
Pseudo supervision is regarded as the core idea in semi-supervised learning for semantic segmentation, and there is always a tradeoff between utilizing only the high-quality pseudo labels and leveraging all the pseudo labels. Addressing that, we propose a novel learning approach, called Conservative-Progressive Collaborative Learning (CPCL), among which two predictive networks are trained in parallel, and the pseudo supervision is implemented based on both the agreement and disagreement of the two predictions. One network seeks common ground via intersection supervision and is supervised by the high-quality labels to ensure a more reliable supervision, while the other network reserves differences via union supervision and is supervised by all the pseudo labels to keep exploring with curiosity. Thus, the collaboration of conservative evolution and progressive exploration can be achieved. To reduce the influences of the suspicious pseudo labels, the loss is dynamic re-weighted according to the prediction confidence. Extensive experiments demonstrate that CPCL achieves state-of-the-art performance for semi-supervised semantic segmentation.
Siqi Fan 0002, Fenghua Zhu, Zunlei Feng, Mingli Song, Fei-Yue Wang 0001
IEEE Trans. Image Process.3
2023 Knowledge Amalgamation for Object Detection With Transformers
abstract
Knowledge amalgamation (KA) is a novel deep model reusing task aiming to transfer knowledge from several well-trained teachers to a multi-talented and compact student. Currently, most of these approaches are tailored for convolutional neural networks (CNNs). However, there is a tendency that Transformers, with a completely different architecture, are starting to challenge the domination of CNNs in many computer vision tasks. Nevertheless, directly applying the previous KA methods to Transformers leads to severe performance degradation. In this work, we explore a more effective KA scheme for Transformer-based object detection models. Specifically, considering the architecture characteristics of Transformers, we propose to dissolve the KA into two aspects: sequence-level amalgamation (SA) and task-level amalgamation (TA). In particular, a hint is generated within the sequence-level amalgamation by concatenating teacher sequences instead of redundantly aggregating them to a fixed-size one as previous KA approaches. Besides, the student learns heterogeneous detection tasks through soft targets with efficiency in the task-level amalgamation. Extensive experiments on PASCAL VOC and COCO have unfolded that the sequence-level amalgamation significantly boosts the performance of students, while the previous methods impair the students. Moreover, the Transformer-based students excel in learning amalgamated knowledge, as they have mastered heterogeneous detection tasks rapidly and achieved superior or at least comparable performance to those of the teachers in their specializations.
Haofei Zhang, Feng Mao, Mengqi Xue, Gongfan Fang, Zunlei Feng, Jie Song 0011, Mingli Song
IEEE Trans. Image Process.5
2023 Distribution Knowledge Embedding for Graph Pooling
abstract
Graph-level representation learning is the pivotal step for downstream tasks that operate on the whole graph. The most common approach to this problem is graph pooling, where node features are typically averaged or summed to obtain the graph representations. However, pooling operations like averaging or summing inevitably cause severe information missing, which may severely downgrade the final performance. In this paper, we argue what is crucial to graph-level downstream tasks includes not only the topological structure but also thedistributionfrom which nodes are sampled. Therefore, powered by existing Graph Neural Networks (GNN), we propose a new plug-and-play pooling module, termed asDistribution Knowledge Embedding(DKEPool), where graphs are viewed as distributions on top of GNNs and the pooling goal is to summarize the entire distribution information instead of retaining a certain feature vector by simple predefined pooling operations. A DKEPool networkde factodisassembles representation learning into two stages,structure learninganddistribution learning. Structure learning follows a recursive neighborhood aggregation scheme to update node features where structure information is obtained. Distribution learning, on the other hand, omits node interconnections and focuses more on the distribution depicted by all the nodes. Extensive experiments on graph classification tasks demonstrate that the proposed DKEPool significantly and consistently outperforms the state-of-the-art methods. The code is avaliable athttps://github.com/chenchkx/dkepool
Kai-Xuan Chen 0001, Jie Song 0011, Shunyu Liu 0001, Na Yu 0001, Zunlei Feng, Gengshi Han, Mingli Song
IEEE Trans. Knowl. Data Eng.5
2023 HSDN: A High-Order Structural Semantic Disentangled Neural Network
abstract
Graph disentangling is a new promising direction that can help us to discover the latent patterns in the data and understand the behaviors of a graph learning model. Despite the many efforts in disentangling representation learning, few works focus on disentangling the latent factors behind a graph. Most current foci are mainly on studying node-level semantics in the graphs. Compared with node-level, the structure-level view can provide a new interpretable and in-depth insight into graph data. The study of structure-level relations enables us to reveal the high-order structural semantics in the data. To explore the complex high-order structural semantics in the data, we propose the High-order Structural Semantic Disentangled Neural Network (HSDN) to model the graph structure units and disentangle structural semantics. It's the first attempt to hypergraph disentangled networks. Unlike prior methods that disentangle factor graphs based on pair-wise relations only, we introduce hyperedges on pair-wise graphs to model structure units and disentangle the complex high-order structural semantics between different structures. Extensive experiments demonstrate that HSDN achieves state-of-the-art performances in terms of both disentangling and downstream tasks.
Bingde Hu, Xingen Wang, Zunlei Feng, Jie Song 0011, Ji Zhao 0016, Mingli Song, Xinyu Wang 0001
IEEE Trans. Knowl. Data Eng.3
2023 Temporal Aggregation and Propagation Graph Neural Networks for Dynamic Representation
abstract
Temporal graphs exhibit dynamic interactions between nodes over continuous time, whose topologies evolve with time elapsing. The whole temporal neighborhood of nodes reveals the varying preferences of nodes. However, previous works usually generate dynamic representation with limited neighbors for simplicity, which results in both inferior performance and high latency of online inference. Therefore, in this paper, we propose a novel method of temporal graph convolution with the whole neighborhood, namely Temporal Aggregation and Propagation Graph Neural Networks (TAP-GNN). Specifically, we first analyze the computational complexity of the dynamic representation problem by unfolding the temporal graph in a message-passing paradigm. The expensive complexity motivates us to design the AP (aggregation and propagation) block, which significantly reduces the repeated computation of historical neighbors. The final TAP-GNN supports online inference in the graph stream scenario, which incorporates the temporal information into node embeddings with a temporal activation function and a projection layer besides several AP blocks. Experimental results on various real-life temporal networks show that our proposed TAP-GNN outperforms existing temporal graph methods by a large margin in terms of both predictive performance and online inference latency.
Tongya Zheng, Xinchao Wang, Zunlei Feng, Jie Song 0011, Yunzhi Hao, Mingli Song, Xingen Wang, Xinyu Wang 0001, Chun Chen 0001
IEEE Trans. Knowl. Data Eng.3
2023 Ask-AC: An Initiative Advisor-in-the-Loop Actor-Critic Framework
abstract
Despite the promising results achieved, state-of-the-art interactive reinforcement learning schemes rely on passively receiving supervision signals from advisor experts, in the form of either continuous monitoring or predefined rules, which inevitably result in a cumbersome and expensive learning process. In this article, we introduce a novel initiative advisor-in-the-loop actor–critic (AC) framework, termed as Ask-AC, that replaces the unilateral advisor-guidance mechanism with a bidirectional learner-initiative one, and thereby enables a customized and efficacious message exchange between learner and advisor. At the heart of Ask-AC are two complementary components, namely, action requester and adaptive state selector, that can be readily incorporated into various discrete AC architectures. The former component allows the agent to initiatively seek advisor intervention in the presence of uncertain states, while the latter identifies the unstable states potentially missed by the former especially when environment changes, and then learns to promote the ask action on such states. Experimental results on both stationary and nonstationary environments and across different AC backbones demonstrate that the proposed framework significantly improves the learning efficiency of the agent, and achieves the performances on par with those obtained by continuous advisor monitoring.
Shunyu Liu 0001, Kai-Xuan Chen 0001, Na Yu 0001, Jie Song 0011, Zunlei Feng, Mingli Song
IEEE Trans. Syst. Man Cybern. Syst.5
2022 Model Doctor: A Simple Gradient Aggregation Strategy for Diagnosing and Treating CNN Classifiers
abstract
Recently, Convolutional Neural Network (CNN) has achieved excellent performance in the classification task. It is widely known that CNN is deemed as a 'blackbox', which is hard for understanding the prediction mechanism and debugging the wrong prediction. Some model debugging and explanation works are developed for solving the above drawbacks. However, those methods focus on explanation and diagnosing possible causes for model prediction, based on which the researchers handle the following optimization of models manually. In this paper, we propose the first completely automatic model diagnosing and treating tool, termed as Model Doctor. Based on two discoveries that 1) each category is only correlated with sparse and specific convolution kernels, and 2) adversarial samples are isolated while normal samples are successive in the feature space, a simple aggregate gradient constraint is devised for effectively diagnosing and optimizing CNN classifiers. The aggregate gradient strategy is a versatile module for mainstream CNN classifiers. Extensive experiments demonstrate that the proposed Model Doctor applies to all existing CNN classifiers, and improves the accuracy of 16 mainstream CNN classifiers by 1%~5%.
Zunlei Feng, Jiacong Hu, Sai Wu, Xiaotian Yu, Jie Song 0011, Mingli Song
AAAI1
2022 Enhanced Dual-Level Representations for Facial Expression Recognition
abstract
Facial expression is an essential factor in conveying human emotional states and intentions. A common strategy used for facial expression recognition (FER) is encoding expression representations from facial images. Although remarkable advancement has been made, challenges due to large variations of expression patterns and unavoidable hard samples still remain. In this paper, we propose dual-level representation enhancements (DLRE) addressing these issues. On one hand, mid-level representation enhancement (MRE) is introduced to avoid expression representation learning being dominated by a limited number of highly discriminative patterns. On the other hand, high-level representation enhancement (HRE) is introduced to alleviate the disturbance of misclassified representations especially for hard samples. The proposed method not only has stronger generalization capability to handle different variations of expression patterns but also greater discriminative power to capture the subtle distinctions of hard samples. Experimental evaluation on four popular databases, CK+, Oulu-CASIA, RAF-DB, and AffectNet, shows that our method achieves more competitive results than other state-of-the-art methods.
Jie Lei 0002, Zeyu Zou, Zunlei Feng, Ronghua Liang
ICIP5
2022 Comparison Knowledge Translation for Generalizable Image Classification
abstract
Deep learning has recently achieved remarkable performance in image classification tasks, which depends heavily on massive annotation. However, the classification mechanism of existing deep learning models seems to contrast to humans' recognition mechanism. With only a glance at an image of the object even unknown type, humans can quickly and precisely find other same category objects from massive images, which benefits from daily recognition of various objects. In this paper, we attempt to build a generalizable framework that emulates the humans' recognition mechanism in the image classification task, hoping to improve the classification performance on unseen categories with the support of annotations of other categories. Specifically, we investigate a new task termed Comparison Knowledge Translation (CKT). Given a set of fully labeled categories, CKT aims to translate the comparison knowledge learned from the labeled categories to a set of novel categories. To this end, we put forward a Comparison Classification Translation Network (CCT-Net), which comprises a comparison classifier and a matching discriminator. The comparison classifier is devised to classify whether two images belong to the same category or not, while the matching discriminator works together in an adversarial manner to ensure whether classified results match the truth. Exhaustive experiments show that CCT-Net achieves surprising generalization ability on unseen categories and SOTA performance on target categories.
Zunlei Feng, Sai Wu, Xiaotuan Jin, Zengliang He, Mingli Song, Huiqiong Wang
IJCAI1
2022 Distribution-Aware Graph Representation Learning for Transient Stability Assessment of Power System
abstract
The real-time transient stability assessment (TSA) plays a critical role in the secure operation of the power system. Although the classic numerical integration method, i.e. time-domain simulation (TDS), has been widely used in industry practice, it is inevitably trapped in a high computational complexity due to the high latitude sophistication of the power system. In this work, a data-driven power system estimation method is proposed to quickly predict the stability of the power system before TDS reaches the end of simulating time windows, which can reduce the average simulation time of stability assessment without loss of accuracy. As the topology of the power system is in the form of graph structure, graph neural network based representation learning is naturally suitable for learning the status of the power system. Motivated by observing the distribution information of crucial active power and reactive power on the power system's bus nodes, we thus propose a distribution-aware learning (DAL) module to explore an informative graph representation vector for describing the status of a power system. Then, TSA is re-defined as a binary classification task, and the stability of the system is determined directly from the resulting graph representation without numerical integration. Finally, we apply our method to the online TSA task. The case studies on the IEEE 39-bus system and Polish 2383-bus system demonstrate the effectiveness of our proposed method. The code is available at https://github.com/kxchern/dkepool-tsa
Kai-Xuan Chen 0001, Shunyu Liu 0001, Na Yu 0001, Jie Song 0011, Zunlei Feng, Mingli Song
IJCNN7
2022 Chemical Property Relation Guided Few-Shot Molecular Property Prediction
abstract
The ability of molecular property prediction is of great significance to drug discovery. However, molecular property prediction is essentially a few-shot problem, making it a challenge in deep learning applications in this scenario. Although some deep learning models, such as meta-learning, have been used in few-shot molecular property prediction, they neglect the relationships among different properties. In this paper, we introduce a chemical property relation modeling technique for generating the property relation map, which can guide the few-shot molecular property prediction with positively related properties. Some molecules sharing common properties are firstly picked out. Then, multiple property-aware graph neural networks are trained for extracting molecular representations from those molecules. Next, the Spearman's correlation is adopted to calculate the property relations with the property-aware matrix, which is built by calculating the pairwise similarity of the above molecular representations for each pair of properties. In the few-shot molecular property prediction task, the meta-learning strategy is adopted to learn common prediction knowledge from the meta-training categories, which are provided by the property relations. Extensive experiments on two public few-shot molecular property datasets demonstrate that the positively related properties are beneficial for the target property prediction, while negatively related properties have negative effects. With the guidance of positively related properties, the proposed method outperforms various state-of-the-art methods.
Shaolun Yao, Zunlei Feng, Jie Song 0011, Lingxiang Jia, Zipeng Zhong, Mingli Song
IJCNN2
2022 Flocking Birds of a Feather Together: Dual-step GAN Distillation via Realer-Fake Samples
abstract
Recent Generative Adversarial Networks (GAN) models deliver visually authentic images and find their applications in various applications. Due to their heavy size, however, it remains cumbersome to deploy state-of-the-art GAN on the edge side, such as mobile terminals. In this paper, we propose a practical and very effective approach for distilling GAN models, so as to produce competent student models that come with smaller sizes. Unlike other convolutional neural network models whose final output spans a small or reasonably-sized space, based on which knowledge distillation is carried out, GAN produces images as output that span an enormous space, making the distillation much more challenging. To this end, we propose an effective dual-step distillation approach tailored for GAN. Instead of directly enforcing the student output distribution to imitate that of the teacher, we partition the generated images into two sets, the “realer” one and the “fake” one, based on the image quality, and sequentially minimize the discrepancy first between fake images of the student and the teacher, and then between realer ones. This seemingly simple strategy turns out highly effective in terms of exploring output space structure and further the knowledge transfer. Experimental results have shown that the proposed method can be easily generalized and demonstrate the learned portable generator produces competitive or even better results both in quality and quantity.
Jingwen Ye, Zunlei Feng, Xinchao Wang
VCIP2
2022 Space and Level Cooperation Framework for Pathological Cancer Grading
abstract
Clinically, the pathological images are intuitive for cancer diagnosis and have been considered as the ‘gold standard’. There are two challenges for applying deep learning into the pathological images analysis: the ultra-large size and the noisy annotations. A pathological image usually contains billions of pixels, which is unsuitable for normal classification models. Furthermore, the ultra-large size and mixed cancerous cells compel the doctor to draw rough boundaries of lesion area according to the cancerous level, which brings two kinds of noisy labels: space noise (annotating inaccurate scope of cancerous area) and level noise (annotating inaccurate cancerous level). Based on the above findings, we propose the space and level cooperation framework, comprising a space-aware branch and a level-aware branch, for pathological cancer grading with noisy annotations. The space-aware branch first turns the ultra-large image into a Multilayer Superpixel (MS) graph, significantly reducing the size and preserving the global features. Then, a global-to-local rectifying strategy is adopted to solve the space noise. The level-aware branch adopts different grouped kernels and a novel grading loss function to handle level noise. Mean-while, two branches cooperate through complementing missing features of each other for handling the above two challenges. Extensive experiments demonstrate that with noisy annotations, the proposed framework achieves SOTA performance on our HCC dataset and two public datasets.
Xiaotian Yu, Zunlei Feng, Xiuming Zhang, Thomas Li
VCIP2
2022 CoEvo-Net: Coevolution Network for Video Highlight Detection
abstract
Video highlight detection (VHD) has emerged as a pressing task due to the unprecedentedly increasing amount of video data, such as those from e-commerce live-broadcasting platforms. Many approaches focus on exploiting text data, in the form of video description or time-sync comments, to facilitate the VHD task. Despite the promising results, they have largely overlooked the noises inherent in the text data and have mostly relied on isolating the feature of text and video. In this paper, we introduce a novel model to handle VHD, termed Coevolution Network (CoEvo-Net), that allows us to account for joint learning of the language and video features explicitly via a coevolution paradigm, in which features from the two data modalities progressively refine each other. This is achieved by a dedicated CoEvo-Cell that takes language and video together as inputs, extracts cross-modality, and filters the undesired parts of the input, such as words in a sentence. Furthermore, we release a large-scale dataset of e-commerce for VHD, in which each video is coupled with a sentence for description, to benchmark the sentence-based VHD approaches. Extensive experiments on the released dataset demonstrate that CoEvo-Net achieves state-of-the-art performance. Our dataset and code will be made publicly available.
Xinchao Wang, Xingen Wang, Zunlei Feng, Ruitao Liu, Mingli Song
IEEE Trans. Circuits Syst. Video Technol.5
2021 Visual Boundary Knowledge Translation for Foreground Segmentation
abstract
When confronted with objects of unknown types in an image, humans can effortlessly and precisely tell their visual boundaries. This recognition mechanism and underlying generalization capability seem to contrast to state-of-the-art image segmentation networks that rely on large-scale category-aware annotated training samples. In this paper, we make an attempt towards building models that explicitly account for visual boundary knowledge, in hope to reduce the training effort on segmenting unseen categories. Specifically, we investigate a new task termed as Boundary Knowledge Translation (BKT). Given a set of fully labeled categories, BKT aims to translate the visual boundary knowledge learned from the labeled categories, to a set of novel categories, each of which is provided only a few labeled samples. To this end, we propose a Translation Segmentation Network (Trans-Net), which comprises a segmentation network and two boundary discriminators. The segmentation network, combined with a boundary-aware self-supervised mechanism, is devised to conduct foreground segmentation, while the two discriminators work together in an adversarial manner to ensure an accurate segmentation of the novel categories under light supervision. Exhaustive experiments demonstrate that, with only tens of labeled samples as guidance, Trans-Net achieves close results on par with fully supervised methods.
Zunlei Feng, Lechao Cheng, Xinchao Wang, Xiang Wang 0010, Ya Jie Liu, Xiangtong Du, Mingli Song
AAAI1
2021 Edge-competing Pathological Liver Vessel Segmentation with Limited Labels
abstract
The microvascular invasion (MVI) is a major prognostic factor in hepatocellular carcinoma, which is one of the malignant tumors with the highest mortality rate. The diagnosis of MVI needs discovering the vessels that contain hepatocellular carcinoma cells and counting their number in each vessel, which depends heavily on experiences of the doctor, is largely subjective and time-consuming. However, there is no algorithm as yet tailored for the MVI detection from pathological images. This paper collects the first pathological liver image dataset containing $522$ whole slide images with labels of vessels, MVI, and hepatocellular carcinoma grades. The first and essential step for the automatic diagnosis of MVI is the accurate segmentation of vessels. The unique characteristics of pathological liver images, such as super-large size, multi-scale vessel, and blurred vessel edges, make the accurate vessel segmentation challenging. Based on the collected dataset, we propose an Edge-competing Vessel Segmentation Network (EVS-Net), which contains a segmentation network and two edge segmentation discriminators. The segmentation network, combined with an edge-aware self-supervision mechanism, is devised to conduct vessel segmentation with limited labeled patches. Meanwhile, two discriminators are introduced to distinguish whether the segmented vessel and background contain residual features in an adversarial manner. In the training stage, two discriminators are devised to compete for the predicted position of edges. Exhaustive experiments demonstrate that, with only limited labeled patches, EVS-Net achieves a close performance of fully supervised methods, which provides a convenient tool for the pathological liver vessel segmentation. Code is publicly available at https://github.com/wang97zh/EVS-Net.
Zunlei Feng, Xinchao Wang, Xiuming Zhang, Lechao Cheng, Jie Lei 0002, Mingli Song
AAAI1
2021 Tendentious Noise-rectifying Framework for Pathological HCC Grading
Xiaotian Yu, Zunlei Feng, Thomas Kwok To Li, Xiuming Zhang, Mingli Song
BMVC2
2021 Facial Expression Recognition by Expression-Specific Representation Swapping
Jie Lei 0002, Zeyu Zou, Zunlei Feng, Ronghua Liang
ICANN (2)6
2021 Mutual-Complementing Framework for Nuclei Detection and Segmentation in Pathology Image
abstract
Detection and segmentation of nuclei are fundamental analysis operations in pathology images, the assessments derived from which serve as the gold standard for cancer diagnosis. Manual segmenting nuclei is expensive and time-consuming. What’s more, accurate segmentation detection of nuclei can be challenging due to the large appearance variation, conjoined and overlapping nuclei, and serious degeneration of histological structures. Supervised methods highly rely on massive annotated samples. The existing two unsupervised methods are prone to failure on degenerated samples. This paper proposes a Mutual-Complementing Framework (MCF) for nuclei detection and segmentation in pathology images. Two branches of MCF are trained in the mutual-complementing manner, where the detection branch complements the pseudo mask of the segmentation branch, while the progressive trained segmentation branch complements the missing nucleus templates through calculating the mask residual between the predicted mask and detected result. In the detection branch, two response map fusion strategies and gradient direction based postprocessing are devised to obtain the optimal detection response. Furthermore, the confidence loss combined with the synthetic samples and self-finetuning is adopted to train the segmentation network with only high confidence areas. Extensive experiments demonstrate that MCF achieves comparable performance with only a few nucleus patches as supervision. Especially, MCF possesses good robustness (only dropping by about 6%) on degenerated samples, which are critical and common cases in clinical diagnosis.
Zunlei Feng, Xinchao Wang, Yining Mao, Thomas Li, Jie Lei 0002, Mingli Song
ICCV1
2021 Speech Guided Disentangled Visual Representation Learning for Lip Reading
abstract
Lip reading has achieved unparalleled development in recent years. However, existing methods have two main problems: 1) there is no explicit mechanism to ensure that the extracted visual features are only related to lip movements, resulting in degraded performance when video contains large variations, such as speakers’ poses; 2) quantities of labeled data are required to achieve good results, which are difficult to obtain in low-resource languages. In this paper, we propose a new visual representation learning method, SVLR, whose purpose is to extract disentangled, lip movements related visual features for lip reading task, by making use of quantities of unlabeled audio-visual data. This is achieved by explicitly disentangling the feature into lip movements related part and speaker identity related part. Then predicting speech from the disentangled features is used as the training objective to optimize model parameters. After this cross-modal training, the video encoder that extracts lip movements features is used as a feature extractor for the lip reading task. Various experiments on several word-level lip reading benchmarks have proved the effectiveness of the proposed method.
Zunlei Feng, Mingli Song
ICMI3
2021 Boundary Knowledge Translation based Reference Semantic Segmentation
abstract
Given a reference object of an unknown type in an image, human observers can effortlessly find the objects of the same category in another image and precisely tell their visual boundaries. Such visual cognition capability of humans seems absent from the current research spectrum of computer vision. Existing segmentation networks, for example, rely on a humongous amount of labeled data, which is laborious and costly to collect and annotate; besides, the performance of segmentation networks tend to downgrade as the number of the category increases. In this paper, we introduce a novel Reference semantic segmentation Network (Ref-Net) to conduct visual boundary knowledge translation. Ref-Net contains a Reference Segmentation Module (RSM) and a Boundary Knowledge Translation Module (BKTM). Inspired by the human recognition mechanism, RSM is devised only to segment the same category objects based on the features of the reference objects. BKTM, on the other hand, introduces two boundary discriminator branches to conduct inner and outer boundary segmentation of the target object in an adversarial manner, and translate the annotated boundary knowledge of open-source datasets into the segmentation network. Exhaustive experiments demonstrate that, with tens of finely-grained annotated samples as guidance, Ref-Net achieves results on par with fully supervised methods on six datasets. Our code can be found in the supplementary material.
Lechao Cheng, Zunlei Feng, Xinchao Wang, Ya Jie Liu, Jie Lei 0002, Mingli Song
IJCAI2
2021 A Location Constrained Dual-Branch Network for Reliable Diagnosis of Jaw Tumors and Cysts
Jiacong Hu, Zunlei Feng, Yining Mao, Jie Lei 0002, Mingli Song
MICCAI (7)2
2021 Shape Controllable Virtual Try-on for Underwear Models
abstract
Image virtual try-on task has abundant applications and has become a hot research topic recently. Existing 2D image-based virtual try-on methods aim to transfer a target clothing image onto a reference person, which has two main disadvantages: cannot control the size and length precisely; unable to accurately estimate the user's figure in the case of users wearing thick clothing, resulting in inaccurate dressing effect. In this paper, we put forward an akin task that aims to dress clothing for underwear models. To solve the above drawbacks, we propose a Shape Controllable Virtual Try-On Network (SC-VTON), where a graph attention network integrates the information of model and clothing to generate the warped clothing image. In addition, the control points are incorporated into SC-VTON for the desired clothing shape. Furthermore, by adding a Splitting Network and a Synthesis Network, we can use in-shop clothing/model pair data to help optimize the deformation module and generalize the task to the typical virtual try-on task. Extensive experiments show that the proposed method can achieve accurate shape control. Meanwhile, compared with other methods, our method can generate high-resolution results with detailed textures, which can be applied in real applications.
Xin Gao 0032, Zhenjiang Liu, Zunlei Feng, Chengji Shen, Kairi Ou, Haihong Tang, Mingli Song
ACM Multimedia3
2021 Towards End-to-End Embroidery Style Generation: A Paired Dataset and Benchmark
Jingwen Ye, Yixin Ji, Jie Song 0011, Zunlei Feng, Mingli Song
PRCV (4)4
2021 Disassembling object representations without labels
Zunlei Feng, Yongming He, Yike Yuan, Huiqiong Wang, Mingli Song
Neurocomputing1
2020 DEAL: Difficulty-Aware Active Learning for Semantic Segmentation
Shuai Xie, Zunlei Feng, Songtao Sun, Mingli Song
ACCV (1)2
2020 Disentangled Representation based Face Anti-Spoofing
abstract
Face anti-spoofing is an important problem for both academic research and industrial face recognition systems. Most of the existing face anti-spoofing methods take it as a classification task on individual static images, where motion pattern differences in consecutive real or fake face sequences are ignored. In this work, we propose a novel method to identify spoofing patterns using motion information. Different from previous methods, the proposed method makes the real or fake decision on the disentangled feature level, based on the observation that motion and spoofing pattern features could be disentangled from original image frames. We design a representation disentangling framework for this task, which is able to reconstruct both real and fake face sequences from the input. Meanwhile, the disentangled representations could be used to classify whether the input faces are real or fake. We perform several experiments on public face anti-spoofing datasets. The proposed method achieves SOTA results compared with existing methods.
Zunlei Feng, Zeyu Zou, Mingli Song, Jianping Shen
ICPR2
2020 Unsupervised Learning Facial Parameter Regressor for Action Unit Intensity Estimation via Differentiable Renderer
abstract
Facial action unit (AU) intensity is an index to describe all visually discernible facial movements. Most existing methods learn intensity estimator with limited AU data, while they lack of generalization ability out of the dataset. In this paper, we present a framework to predict the facial parameters (including identity parameters and AU parameters) based on a bone-driven face model (BDFM) under different views. The proposed framework consists of a feature extractor, a generator, and a facial parameter regressor. The regressor can fit the physical meaning parameters of the BDFM from a single face image with the help of the generator, which maps the facial parameters to the game-face images as a differentiable renderer. Besides, identity loss, loopback loss, and adversarial loss can improve the regressive results. Quantitative evaluations are performed on two public databases BP4D and DISFA, which demonstrates that the proposed method can achieve comparable or better performance than the state-of-the-art methods. What's more, the qualitative results also demonstrate the validity of our method in the wild.
Xinhui Song, Tianyang Shi, Zunlei Feng, Mingli Song, Jackie Lin, Chuanjie Lin, Changjie Fan, Yi Yuan 0002
ACM Multimedia3
2020 One-sample Guided Object Representation Disassembling
abstract
The ability to disassemble the features of objects and background is crucial for many machine learning tasks, including image classification, image editing, visual concepts learning, and so on. However, existing (semi-)supervised methods all need a large amount of annotated samples, while unsupervised methods can't handle real-world images with complicated backgrounds. In this paper, we introduce the One-sample Guided Object Representation Disassembling (One-GORD) method, which only requires one annotated sample for each object category to learn disassembled object representation from unannotated images. For the annotated one-sample, we first adopt some data augmentation strategies to generate some synthetic samples, which can guide the disassembling of the object features and background features. For the unannotated images, two self-supervised mechanisms: dual-swapping and fuzzy classification are introduced to disassemble object features from the background with the guidance of annotated one-sample. What's more, we devise two metrics to evaluate the disassembling performance from the perspective of representation and image, respectively. Experiments demonstrate that the One-GORD achieves competitive dissembling performance and can handle natural scenes with complicated backgrounds.
Zunlei Feng, Yongming He, Xinchao Wang, Xin Gao 0032, Jie Lei 0002, Cheng Jin 0001, Mingli Song
NeurIPS1
2020 Factorizable Graph Convolutional Networks
abstract
Graphs have been widely adopted to denote structural connections between entities. The relations are in many cases heterogeneous, but entangled together and denoted merely as a single edge between a pair of nodes. For example, in a social network graph, users in different latent relationships like friends and colleagues, are usually connected via a bare edge that conceals such intrinsic connections. In this paper, we introduce a novel graph convolutional network (GCN), termed as factorizable graph convolutional network (FactorGCN), that explicitly disentangles such intertwined relations encoded in a graph. FactorGCN takes a simple graph as input, and disentangles it into several factorized graphs, each of which represents a latent and disentangled relation among nodes. The features of the nodes are then aggregated separately in each factorized latent space to produce disentangled features, which further leads to better performances for downstream tasks. We evaluate the proposed FactorGCN both qualitatively and quantitatively on the synthetic and real-world datasets, and demonstrate that it yields truly encouraging results in terms of both disentangling and feature aggregation. Code is publicly available at https://github.com/ihollywhy/FactorGCN.PyTorch.
Yiding Yang, Zunlei Feng, Mingli Song, Xinchao Wang
NeurIPS2
2020 Adaptive Context Learning Network for Crowd Counting
abstract
The task of crowd counting is to estimate the accurate number of people in photos taken from unconstrained surveillance scenes. It is in general a challenging problem due to the input scale variations and perspective distortions. Previous methods make efforts to enhance the representation ability by using multi-scale features of the scene pictures. However, most of these methods directly add or fuse the features, in which the influences of different feature sizes are equally considered. In this paper, we propose a novel architecture called adaptive context learning network (ACLNet) to incorporate context of features in multiple levels. In this architecture, the original image features are enhanced by a multi-level feature generating module, and then the multi-level features are up-sampled to the same size and re-weighted for fusing. The ACLNet incorporates the context information existed in sub-regions of various scales adaptively, thus it is able to enhance the representative ability of multi-level features. We perform several experiments on public ShanghaiTech (A and B), UCF_CC_50 and NWPU-crowd datasets. Our proposed ACLNet achieves the state-of-the-art results compared with existing methods.
Guanqi Zeng, Zunlei Feng, Mingli Song, Jianping Shen
SMC3
2020 Neural Style Transfer: A Review
abstract
The seminal work of Gatys et al. demonstrated the power of Convolutional Neural Networks (CNNs) in creating artistic imagery by separating and recombining image content and style. This process of using CNNs to render a content image in different styles is referred to as Neural Style Transfer (NST). Since then, NST has become a trending topic both in academic literature and industrial applications. It is receiving increasing attention and a variety of approaches are proposed to either improve or extend the original NST algorithm. In this paper, we aim to provide a comprehensive overview of the current progress towards NST. We first propose a taxonomy of current algorithms in the field of NST. Then, we present several evaluation methods and compare different NST algorithms both qualitatively and quantitatively. The review concludes with a discussion of various applications of NST and open problems for future research. A list of papers discussed in this review, corresponding codes, pre-trained models and more comparison results are publicly available at: https://osf.io/f8tu4/.
Yongcheng Jing, Yezhou Yang, Zunlei Feng, Jingwen Ye, Yizhou Yu, Mingli Song
IEEE Trans. Vis. Comput. Graph.3
2019 CU-Net: Component Unmixing Network for Textile Fiber Identification
Zunlei Feng, Weixin Liang, Daocheng Tao, Anxiang Zeng, Mingli Song
Int. J. Comput. Vis.1
2019 Interpretable Partitioned Embedding for Intelligent Multi-item Fashion Outfit Composition
abstract
Intelligent fashion outfit composition has become more popular in recent years. Some deep-learning-based approaches reveal competitive composition. However, the uninterpretable characteristic makes such a deep-learning-based approach fail to meet the businesses’, designers’, and consumers’ urges to comprehend the importance of different attributes in an outfit composition. To realize interpretable and intelligent multi-item fashion outfit compositions, we propose a partitioned embedding network to learn interpretable embeddings from clothing items. The network contains two vital components: attribute partition module and partition adversarial module. In the attribute partition module, multiple attribute labels are adopted to ensure that different parts of the overall embedding correspond to different attributes. In the partition adversarial module, adversarial operations are adopted to achieve the independence of different parts. With the interpretable and partitioned embedding, we then construct an outfit-composition graph and an attribute matching map. Extensive experiments demonstrate that (1) the partitioned embedding have unmingled parts that correspond to different attributes and (2) outfits recommended by our model are more desirable in comparison with the existing methods.
Zunlei Feng, Zhenyun Yu, Yongcheng Jing, Sai Wu, Mingli Song, Yezhou Yang, Junxiao Jiang
ACM Trans. Multim. Comput. Commun. Appl.1
2018 Stroke Controllable Fast Style Transfer with Adaptive Receptive Fields
Yongcheng Jing, Yang Liu 0212, Yezhou Yang, Zunlei Feng, Yizhou Yu, Dacheng Tao, Mingli Song
ECCV (13)4
2018 Finer-Net: Cascaded Human Parsing with Hierarchical Granularity
abstract
Human parsing is a challenging and important task in various applications, such as dress collocation, clothing recommendation and action analysis. However, the existing methods are easily affected by pose variation and occlusion with requiring massive intensive annotations for fine-grained human segmentation. In this paper, we design a cascaded segmentation network with three stages to solve the above problems. Given a human image, we firstly predict the human joints as pose features. Secondly, these features along with the input image are fed into the first stage to obtain a primitive segmentation map to separate the human and the background. The primitive segmentation is then fed into the second stage with the original image to give a rough segmentation of human body. This procedure is repeated in the third stage to acquire a refined segmentation. Experimental results demonstrate the proposed method achieve superior performance than state-of-the-arts and show great generalization ability.
Jingwen Ye, Zunlei Feng, Yongcheng Jing, Mingli Song
ICME2
2018 Interpretable Partitioned Embedding for Customized Multi-item Fashion Outfit Composition
abstract
Intelligent fashion outfit composition becomes more and more popular in these years. Some deep learning based approaches reveal competitive composition recently. However, the uninterpretable characteristic makes such deep learning based approach cannot meet the designers, businesses and consumers' urge to comprehend the importance of different attributes in an outfit composition. To realize interpretable and customized multi-item fashion outfit compositions, we propose a partitioned embedding network to learn interpretable embeddings from clothing items. The network consists of two vital components: attribute partition module and partition adversarial module. In the attribute partition module, multiple attribute labels are adopted to ensure that different parts of the overall embedding correspond to different attributes. In the partition adversarial module, adversarial operations are adopted to achieve the independence of different parts. With the interpretable and partitioned embedding, we then construct an outfit composition graph and an attribute matching map. Extensive experiments demonstrate that 1) the partitioned embedding have unmingled parts which corresponding to different attributes and 2) outfits recommended by our model are more desirable in comparison with the existing methods.
Zunlei Feng, Zhenyun Yu, Yezhou Yang, Yongcheng Jing, Junxiao Jiang, Mingli Song
ICMR1
2018 Dual Swap Disentangling
abstract
Learning interpretable disentangled representations is a crucial yet challenging task. In this paper, we propose a weakly semi-supervised method, termed as Dual Swap Disentangling (DSD), for disentangling using both labeled and unlabeled data. Unlike conventional weakly supervised methods that rely on full annotations on the group of samples, we require only limited annotations on paired samples that indicate their shared attribute like the color. Our model takes the form of a dual autoencoder structure. To achieve disentangling using the labeled pairs, we follow a encoding-swap-decoding'' process, where we first swap the parts of their encodings corresponding to the shared attribute, and then decode the obtained hybrid codes to reconstruct the original input pairs. For unlabeled pairs, we follow theencoding-swap-decoding'' process twice on designated encoding parts and enforce the final outputs to approximate the input pairs. By isolating parts of the encoding and swapping them back and forth, we impose the dimension-wise modularity and portability of the encodings of the unlabeled samples, which implicitly encourages disentangling under the guidance of labeled pairs. This dual swap mechanism, tailored for semi-supervised setting, turns out to be very effective. Experiments on image datasets from a wide domain show that our model yields state-of-the-art disentangling performances.
Zunlei Feng, Xinchao Wang, Chenglong Ke, Anxiang Zeng, Dacheng Tao, Mingli Song
NeurIPS1
2018 Finding intrinsic color themes in images with human visual perception
Zunlei Feng, Wolong Yuan, Chunli Fu, Jie Lei 0002, Mingli Song
Neurocomputing1
2018 Scale insensitive and focus driven mobile screen defect detection in industry
Jie Lei 0002, Xin Gao 0032, Zunlei Feng, Huamou Qiu, Mingli Song
Neurocomputing3
2017 Graph-based color Gamut Mapping using neighbor metric
abstract
Colors are displayed in different ways on various devices, such as cameras, screens and printers. In order to achieve consistent appearance on these devices, color management is usually used, where the core part is Gamut Mapping Algorithm (GMA). However, the widely adopted Point-wise Gamut Mapping Algorithms (PGMAs) have been restricted to compromise between color accuracy and details. In this work, we firstly split color space into small cubes through sampling colors from it. Then, we built a 26-neighbored graph with the sample colors as vertexes and perceptual color differences between adjacent vertexes as weights. Based on the above graph, a new Multi-source Shortest Paths Algorithm (MSS-PA) is proposed to establish color mapping relationships between out-of-gamut colors and colors in gamut boundary. In the MSSPA, distance of shortest path between nonadjacent vertexes are used to replace those calculated directly using CIEDE2000, which is useful to measure large color difference. Experimental results show that our method achieves superior performance on the aspect of keeping accuracy and preserving details compared with HPMinDE and SGCK.
Zunlei Feng, Yongcheng Jing, Jie Lei 0002, Mingli Song
ICME1
2016 Which face is more attractive?
abstract
People are fond of sharing their photos of life experience in social networks, where the majority of photos are containing faces, especially in selfies. In fact, we are spontaneously evaluating the attractiveness of faces we see in daily life with personality and emotional traits within a single glance. Can we make a comparison on attractiveness between two face images automatically? In this paper, we capitalize on the fact that the difficulties in comparing the attractiveness of various face image pairs are not even and propose a hierarchical structure integrated with multiple comparative convolutional neural networks (CCNNs). First, five synthetic part faces are generated for each original face image to highlight the influences of parts in attractiveness. Both the original and synthetic face pairs are displayed to the subjects for attractiveness labeling. Multiple CCNNs are then trained on the labeled data with different kinds of pairs. Finally, all the CCNNs are combined to make a weighted comparison result for an input pair. The experimental results in two datasets show CCNNs can achieve significant performance improvement over other state-of-the-art methods. Quantifying the attractiveness of face images lends to many useful applications, such as choosing a better selfie and improving the results of related searching tasks.
Jie Lei 0002, Zunlei Feng, Mingli Song, Dacheng Tao
ICIP2