VLDB 2026 Research / reviewers in the wild / expert
Jie Lei 0002
dblp:61/5501-2
· DBLP profile ↗
45ranked-venue papers
10as first author
32since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 31 · 8 first-author · 22 since 2021Artificial intelligence and machine learning · 24 · 3 first-author · 17 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Association Pattern-enhanced Molecular Representation LearningabstractThe applicability of drug molecules in various clinical scenarios is significantly influenced by a diverse range of molecular properties. By leveraging self-supervised conditions such as atom attributes and interatomic bonds, existing advanced molecular foundation models can generate expressive representations of these molecules. However, such models often overlook the fixed association patterns within molecules that influence physiological or chemical properties. In this paper, we introduce a novel association pattern-aware message passing method, which can serve as an effective yet general plug-and-play plugin, thereby enhancing the atom representations generated by molecular foundation models without requiring additional pretraining. Additionally, molecular property-specific pattern libraries are constructed to collect the generated interpretable common patterns that bind to these properties. Extensive experiments conducted on 11 benchmark molecular property prediction tasks across 8 advanced molecular foundation models demonstrate significant superiority of the proposed method, with performance improvements of up to approximately 20%. Furthermore, a property-specific pattern library is tailored for blood-brain barrier penetration, which has undergone corresponding mechanistic validation. Lingxiang Jia, Yuchen Ying, Shaolun Yao, Jie Lei 0002, Jie Song 0011, Mingli Song, Zunlei Feng |
AAAI | 6 |
| 2025 | Multimodal Sentiment Analysis with Parallel Attention and Correlation Fusion
Jie Lei 0002, Zunlei Feng, Ronghua Liang |
ICANN (3) | 2 |
| 2025 | Spatial-Temporal Reconstruction Error for AIGC-based Forgery Image DetectionabstractThe remarkable success of AI-Generated Content (AIGC), especially diffusion image generation models, brings about unprecedented creative applications, but also creates fertile ground for malicious counterfeiting and crime. A highly effective family of forgery image detection methods based on diffusion reconstruction error has emerged, as images generated by diffusion are more easily reconstructed by any diffusion model. However, we find that existing methods only use reconstruction error from a single time step, failing to fully leverage the entire reconstruction process. To this end, we propose to comprehensively consider every single time step to form the Temporal Reconstruction Error (TRE) that offers a richer feature representation. Furthermore, we design temporal aggregation and spatial focusing modules from two dimensions respectively to more effectively extract discriminative information from the TRE feature. Finally, we validate the proposed method on two popular datasets, and experimental results demonstrate that the proposed approach achieves state-of-the-art performance. Chengji Shen, Zhenjiang Liu, Kai-Xuan Chen 0001, Jie Lei 0002, Mingli Song, Zunlei Feng |
ICASSP | 4 |
| 2025 | Dynamic Routing and Calibration for Few-Shot Object DetectionabstractFew-shot object detection (FSOD), aiming to enhance the performance of novel object detection with limited labeled samples, has recently gained significant attention. Recent researches primarily focus on improving the generalization of novel classes and enhancing detector performance. However, the diversity of samples is often overlooked, and object proposals with inaccurate classifications or locations remain uncorrected. In this paper, we propose Dynamic Routing and Calibration for Few-Shot Object Detection (DRC-FSOD). Our approach includes a dynamic backbone routing that adapts to various samples by selecting appropriate backbones dynamically. Meanwhile, we construct a dynamic calibration module, which dynamically perform individual calibration for proposals based on their scores. Experimental results on MS COCO and Pascal VOC datasets show superiority over state-of-the-art methods. Jie Lei 0002, Zunlei Feng, Ronghua Liang |
ICASSP | 2 |
| 2025 | Spatial-Temporal Forgery Trace Based Forgery Image Identification
Zunlei Feng, Jiachi Wang, Hengrui Lou, Binjia Zhou, Jie Lei 0002, Mingli Song, Yijun Bei |
ICCV | 6 |
| 2025 | STD-FD: Spatio-Temporal Distribution Fitting Deviation for AIGC Forgery IdentificationabstractWith the rise of AIGC technologies, particularly diffusion models, highly realistic fake images that can deceive human visual perception has become feasible. Consequently, various forgery detection methods have emerged. However, existing methods treat the generation process of fake images as either a black-box or an auxiliary tool, offering limited insights into its underlying mechanisms. In this paper, we propose Spatio-Temporal Distribution Fitting Deviation (STD-FD) for AIGC forgery detection, which explores the generative process in detail. By decomposing and reconstructing data within generative diffusion models, initial experiments reveal temporal distribution fitting deviations during the image reconstruction process. These deviations are captured through reconstruction noise maps for each spatial semantic unit, derived via a super-resolution algorithm. Critical discriminative patterns, termed DFactors, are identified through statistical modeling of these deviations. Extensive experiments show that STD-FD effectively captures distribution patterns in AIGC-generated data, demonstrating strong robustness and generalizability while outperforming state-of-the-art (SOTA) methods on major datasets. The source code is available at [this link](https://github.com/HengruiLou/STDFD). Hengrui Lou, Zunlei Feng, Jinsong Geng, Erteng Liu, Jie Lei 0002, Lechao Cheng, Jie Song 0011, Mingli Song, Yijun Bei |
ICML | 5 |
| 2025 | CorrDetail: Visual Detail Enhanced Self-Correction for Face Forgery DetectionabstractWith the swift progression of image generation technology, the widespread emergence of facial deepfakes poses significant challenges to the field of security, thus amplifying the urgent need for effective deepfake detection. Existing techniques for face forgery detection can broadly be categorized into two primary groups: visual-based methods and multimodal approaches. The former often lacks clear explanations for forgery details, while the latter, which merges visual and linguistic modalities, is more prone to the issue of hallucinations.To address these shortcomings, we introduce a visual detail enhanced self-correction framework, designated CorrDetail, for interpretable face forgery detection. CorrDetail is meticulously designed to rectify authentic forgery details when provided with error-guided questioning, with the aim of fostering the ability to uncover forgery details rather than yielding hallucinated responses. Additionally, to bolster the reliability of its findings, a visual fine-grained detail enhancement module is incorporated, supplying CorrDetail with more precise visual forgery details. Ultimately, a fusion decision strategy is devised to further augment the model's discriminative capacity in handling extreme samples, through the integration of visual information compensation and model bias reduction. Experimental results demonstrate that CorrDetail not only achieves state-of-the-art performance compared to the latest methodologies but also excels in accurately identifying forged details, all while exhibiting robust generalization capabilities. Binjia Zhou, Hengrui Lou, Lizhe Chen, Dawei Luo, Jie Lei 0002, Zunlei Feng, Yijun Bei |
IJCAI | 7 |
| 2025 | DOMR: Establishing Cross-View Segmentation via Dense Object MatchingabstractCross-view object correspondence involves matching objects between egocentric (first-person) and exocentric (third-person) views. It is a critical yet challenging task for visual understanding. In this work, we propose the Dense Object Matching and Refinement (DOMR) framework to establish dense object correspondences across views. The framework centers around the Dense Object Matcher (DOM) module, which jointly models multiple objects. Unlike methods that directly match individual object masks to image features, DOM leverages both positional and semantic relationships among objects to find correspondences. DOM integrates a proposal generation module with a dense matching module that jointly encodes visual, spatial, and semantic cues, explicitly constructing inter-object relationships to achieve dense matching among objects. Furthermore, we combine DOM with a mask refinement head designed to improve the completeness and accuracy of the predicted masks, forming the complete DOMR framework. Extensive evaluations on the Ego-Exo4D benchmark demonstrate that our approach achieves state-of-the-art performance with a mean IoU of 49.7% on Ego→Exo and 55.2% on Exo→Ego. These results outperform those of previous methods by 5.8% and 4.3%, respectively, validating the effectiveness of our integrated approach for cross-view understanding. Jitong Liao, Yulu Gao, Shaofei Huang 0001, Jialin Gao, Jie Lei 0002, Ronghua Liang, Si Liu 0001 |
ACM Multimedia | 5 |
| 2025 | A Large-scale Universal Evaluation Benchmark For Face Forgery Detection
Hengrui Lou, Zunlei Feng, Jinsong Geng, Erteng Liu, Lechao Cheng, Jie Lei 0002, Jie Song 0011, Mingli Song, Yijun Bei |
ACM Multimedia | 6 |
| 2025 | Association-Focused Path Aggregation for Graph Fraud DetectionabstractFraudulent activities have caused substantial negative social impacts and are exhibiting emerging characteristics such as intelligence and industrialization, posing challenges of high-order interactions, intricate dependencies, and the sparse yet concealed nature of fraudulent entities. Existing graph fraud detectors are limited by their narrow "receptive fields", as they focus only on the relations between an entity and its neighbors while neglecting longer-range structural associations hidden between entities. To address this issue, we propose a novel fraud detector based on Graph Path Aggregation (GPA). It operates through variable-length path sampling, semantic-associated path encoding, path interaction and aggregation, and aggregation-enhanced fraud detection. To further facilitate interpretable association analysis, we synthesize G-Internet, the first benchmark dataset in the field of internet fraud detection. Extensive experiments across datasets in multiple fraud scenarios demonstrate that the proposed GPA outperforms mainstream fraud detectors by up to +15% in Average Precision (AP). Additionally, GPA exhibits enhanced robustness to noisy labels and provides excellent interpretability by uncovering implicit fraudulent patterns across broader contexts. Code is available at https://github.com/horrible-dong/GPA. Zunlei Feng, Jie Lei 0002, Mingli Song, Yang Gao 0001 |
NeurIPS | 4 |
| 2025 | Signature Feature Sequence for Model Reuse Detection
Zunlei Feng, Jie Lei 0002 |
PRCV (2) | 6 |
| 2024 | ViT-Calibrator: Decision Stream Calibration for Vision TransformerabstractA surge of interest has emerged in utilizing Transformers in diverse vision tasks owing to its formidable performance. However, existing approaches primarily focus on optimizing internal model architecture designs that often entail significant trial and error with high burdens. In this work, we propose a new paradigm dubbed Decision Stream Calibration that boosts the performance of general Vision Transformers. To achieve this, we shed light on the information propagation mechanism in the learning procedure by exploring the correlation between different tokens and the relevance coefficient of multiple dimensions. Upon further analysis, it was discovered that 1) the final decision is associated with tokens of foreground targets, while token features of foreground target will be transmitted into the next layer as much as possible, and the useless token features of background area will be eliminated gradually in the forward propagation. 2) Each category is solely associated with specific sparse dimensions in the tokens. Based on the discoveries mentioned above, we designed a two-stage calibration scheme, namely ViT-Calibrator, including token propagation calibration stage and dimension propagation calibration stage. Extensive experiments on commonly used datasets show that the proposed approach can achieve promising results. Zhijie Jia, Lechao Cheng, Yang Gao 0001, Jie Lei 0002, Yijun Bei, Zunlei Feng |
AAAI | 5 |
| 2024 | Angle Robustness Unmanned Aerial Vehicle Navigation in GNSS-Denied ScenariosabstractDue to the inability to receive signals from the Global Navigation Satellite System (GNSS) in extreme conditions, achieving accurate and robust navigation for Unmanned Aerial Vehicles (UAVs) is a challenging task. Recently emerged, vision-based navigation has been a promising and feasible alternative to GNSS-based navigation. However, existing vision-based techniques are inadequate in addressing flight deviation caused by environmental disturbances and inaccurate position predictions in practical settings. In this paper, we present a novel angle robustness navigation paradigm to deal with flight deviation in point-to-point navigation tasks. Additionally, we propose a model that includes the Adaptive Feature Enhance Module, Cross-knowledge Attention-guided Module and Robust Task-oriented Head Module to accurately predict direction angles for high-precision navigation. To evaluate the vision-based navigation methods, we collect a new dataset termed as UAV_AR368. Furthermore, we design the Simulation Flight Testing Instrument (SFTI) using Google Earth to simulate different flight environments, thereby reducing the expenses associated with real flight testing. Experiment results demonstrate that the proposed model outperforms the state-of-the-art by achieving improvements of 26.0% and 45.6% in the success rate of arrival under ideal and disturbed circumstances, respectively. Zunlei Feng, Haofei Zhang, Yang Gao 0001, Jie Lei 0002, Mingli Song |
AAAI | 5 |
| 2024 | Visual-guided Query with Temporal Interaction for Video Object SegementationabstractThe task of referring video object segmentation (RVOS) involves segmenting objects in video frames based on a given text description. However, most existing approaches treat the text directly as a query, neglecting the valuable visual and temporal information from the video. This limitation may cause the query unable to accurately perceive the target object. To address this issue, we introduce a visual-guided query with temporal interaction for referring video object segmentation (VQTI) approach. Our method capitalizes on frame-level features and video-level features to guide the query generation process, resulting in an enhanced perception of the target object. In addition, we introduce a spectral-guided segmentation optimizer module to enhance the fine-grained information, leading to more precise segmentation masks. Extensive experiments shows competitive performance against state-of-the-art approaches. Jiaxin Qiu, Guoyu Yang, Jie Lei 0002, Zunlei Feng, Ronghua Liang |
ICME | 3 |
| 2024 | MFCA: Multimodal Object Detection Based on Feature Calibration and Aggregation
Jie Lei 0002, Guoyu Yang, Zunlei Feng, Ronghua Liang |
ICONIP (8) | 2 |
| 2024 | Deep Kernel CalibrationabstractKorbinian Brodmann argued that the brain regions with different cytoarchitectures exhibited different cognitive functions, which are widely recognized as "Brodmann areas" in neuroscience. Inspired by this theory, we observe from experiments that the different well-trained convolutional kernels also hold their unique functionalities, indicated by their output activations that highly concentrate on the "different certain ranges". This discovery motivates us to devise a deep kernel-by-kernel feature calibration mechanism that "calibrates" the output distribution of a pre-trained convolutional kernel for a more concentrated activation interval, by eliminating the outlier activations. Towards this end, we develop two dedicated Brodmann calibration forms, termed as Hard Kernel Calibration (HKC) and Soft Kernel Calibration (SKC) that simply filters the overrange activations, or adaptively condense the activations in a weighted manner, respectively. As a flexible plug-and-play module, the proposed calibration demonstrates encouraging results on five benchmarks across sixteen network architectures and also triggers the functionality of significantly enhanced model robustness against adversarial attacks. The developed code is publicly available at https://github.com/tcmyxc/DKC. Zunlei Feng, Jie Lei 0002, Huiqiong Wang, Zhongle Xie |
IJCNN | 4 |
| 2024 | GAN Doctor: Diagnosing and Treating Inherent Semantic ErrorsabstractGenerative Adversarial Network (GAN), as a popular generative model in the field of Artificial Intelligence Generated Content (AIGC), has been intensively developed in previous research, with significant improvements in the quality and diversity of image generation. However, there are still many cases where the results are not satisfactory. A primary concern pertains to the chaotic and blurred local details within the generated images. In this work, through the diagnosis and analysis of high-quality and low-quality images produced by the GAN model, we identified that this issue stems from inherent semantic errors of the GAN, that is, convolutional kernels responsible for certain semantics are not properly involved in the generation process of corresponding image regions. To this end, we propose a straightforward yet effective treatment method, which constrains each image region to be generated by its corresponding semantic convolutional kernels. Experimental results demonstrate that our proposed optimization method can improve the issue of chaotic and blurred local regions in generated images and enhance the overall generation quality. Our work pioneers a novel paradigm for diagnosing and treating GANs, driving the research development and practical application of AIGC image generation technology. Chengji Shen, Zunlei Feng, Zhongle Xie, Jie Lei 0002, Huiqiong Wang, Mingli Song |
IJCNN | 4 |
| 2024 | Contextual Augmentation with Bias Adaptive for Few-Shot Video Object Segmentation
Shuaiwei Wang, Jie Lei 0002, Zunlei Feng, Ronghua Liang |
MMM (1) | 3 |
| 2024 | Dynamic-Static Graph Convolutional Network for Video-Based Facial Expression Recognition
Fahong Wang, Jie Lei 0002, Zeyu Zou, Zunlei Feng, Ronghua Liang |
MMM (2) | 3 |
| 2024 | Dual-Perspective Activation: Efficient Channel Denoising via Joint Forward-Backward Criterion for Artificial Neural NetworksabstractThe design of Artificial Neural Network (ANN) is inspired by the working patterns of the human brain. Connections in biological neural networks are sparse, as they only exist between few neurons. Meanwhile, the sparse representation in ANNs has been shown to possess significant advantages. Activation responses of ANNs are typically expected to promote sparse representations, where key signals get activated while irrelevant/redundant signals are suppressed. It can be observed that samples of each category are only correlated with sparse and specific channels in ANNs. However, existing activation mechanisms often struggle to suppress signals from other irrelevant channels entirely, and these signals have been verified to be detrimental to the network's final decision. To address the issue of channel noise interference in ANNs, a novel end-to-end trainable Dual-Perspective Activation (DPA) mechanism is proposed. DPA efficiently identifies irrelevant channels and applies channel denoising under the guidance of a joint criterion established online from both forward and backward propagation perspectives while preserving activation responses from relevant channels. Extensive experiments demonstrate that DPA successfully denoises channels and facilitates sparser neural representations. Moreover, DPA is parameter-free, fast, applicable to many mainstream ANN architectures, and achieves remarkable performance compared to other existing activation counterparts across multiple tasks and domains. Code is available at https://github.com/horrible-dong/DPA. Chenchao Gao, Zunlei Feng, Jie Lei 0002, Bingde Hu, Xingen Wang, Mingli Song |
NeurIPS | 4 |
| 2024 | Asymptotic Feature Pyramid Network for Labeling Pixels and RegionsabstractMulti-scale features are crucial in encoding objects with varying scales in vision tasks. The classic top-down and bottom-up feature pyramid networks are a common strategy for multi-scale feature extraction. However, these approaches suffer from the loss or degradation of feature information, which impairs the fusion effect of non-adjacent levels. In this paper, we propose an Asymptotic Feature Pyramid Network (AFPN) that supports direct interaction between non-adjacent levels. AFPN starts by fusing two adjacent low-level features and asymptotic incorporates higher-level features into the fusion process. This fusion way avoids the significant semantic gap between non-adjacent levels. Adaptive spatial fusion operation is further used to mitigate potential multi-object information conflicts during feature fusion at each spatial location. To reduce parameters, computational requirements, and inference speed, we propose a Lightweight Asymptotic Feature Pyramid Network (LightAFPN) that uses the concept of reparametrization. We evaluate the proposed method on the MS-COCO 2017, PASCAL VOC and Cityscapes datasets in both object detection and semantic segmentation frameworks. Experimental evaluation shows that our method achieves more competitive results than other state-of-the-art feature pyramid networks. The code is available at https://github.com/gyyang23/AFPN. Guoyu Yang, Jie Lei 0002, Zunlei Feng, Ronghua Liang |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2023 | A Loopback Network for Explainable Microvascular Invasion ClassificationabstractMicrovascular invasion (MVI) is a critical factor for prognosis evaluation and cancer treatment. The current diagnosis of MVI relies on pathologists to manually find out cancerous cells from hundreds of blood vessels, which is time-consuming, tedious, and subjective. Recently, deep learning has achieved promising results in medical image analysis tasks. However, the unexplainability of black box models and the requirement of massive annotated samples limit the clinical application of deep learning based diagnostic methods. In this paper, aiming to develop an accurate, objective, and explainable diagnosis tool for MVI, we propose a Loopback Network (LoopNet) for classifying MVI efficiently. With the image-level category annotations of the collected Pathologic Vessel Image Dataset (PVID), LoopNet is devised to be composed binary classification branch and cell locating branch. The latter is devised to locate the area of cancerous cells, regular non-cancerous cells, and background. For healthy samples, the pseudo masks of cells supervise the cell locating branch to distinguish the area of regular non-cancerous cells and background. For each MVI sample, the cell locating branch predicts the mask of cancerous cells. Then the masked cancerous and non-cancerous areas of the same sample are input back to the binary classification branch separately. The loopback between two branches enables the category label to supervise the cell locating branch to learn the locating ability for cancerous areas. Experiment results show that the proposed LoopNet achieves 97.5% accuracy on MVI classification. Surprisingly, the proposed loopback mechanism not only enables LoopNet to predict the cancerous area but also facilitates the classification backbone to achieve better classification performance. Shengxuming Zhang, Tianqi Shi, Xiuming Zhang, Jie Lei 0002, Zunlei Feng, Mingli Song |
CVPR | 5 |
| 2023 | Temporal Aggregation with Context Focusing for Few-Shot Video Object DetectionabstractFew-shot video object detection focuses on finding all the objects in a given query video that belong to the same class, given only a few support images of the target object in an unseen class. Unfortunately, due to the object blur or occlusion in video frames, using single-frame object detection directly will greatly limit the accuracy. The issue is significantly worse in few-shot settings due to insufficient support and timedomain information. In this paper, we propose a temporal aggregation with context focusing framework (TACF) for few-shot video object detection, which aims to fully use the information between support images and adjacent video frames. The context focusing module effectively encodes the target object in adjacent frames according to the support images. Afterward, the temporal aggregation module implicitly extracts the most similar ROI features from these adjacent frames to obtain the target proposals. In the end, the matching network determines the category and bounding box by calculating the distance with the support images. Extensive experimental evaluations on FSVOD and FSYTV databases show that our method achieves more competitive results than image-based methods, naive video-based extensions, and the state-of-the-art few-shot video object detection method. Jie Lei 0002, Fahong Wang, Zunlei Feng, Ronghua Liang |
SMC | 2 |
| 2023 | AFPN: Asymptotic Feature Pyramid Network for Object DetectionabstractMulti-scale features are of great importance in encoding objects with scale variance in object detection tasks. A common strategy for multi-scale feature extraction is adopting the classic top-down and bottom-up feature pyramid networks. However, these approaches suffer from the loss or degradation of feature information, impairing the fusion effect of non-adjacent levels. This paper proposes an asymptotic feature pyramid network (AFPN) to support direct interaction at non-adjacent levels. AFPN is initiated by fusing two adjacent low-level features and asymptotically incorporates higher-level features into the fusion process. In this way, the larger semantic gap between non-adjacent levels can be avoided. Given the potential for multi-object information conflicts to arise during feature fusion at each spatial location, adaptive spatial fusion operation is further utilized to mitigate these inconsistencies. We incorporate the proposed AFPN into both two-stage and one-stage object detection frameworks and evaluate with the MS-COCO 2017 validation and test datasets. Experimental evaluation shows that our method achieves more competitive results than other state-of-the-art feature pyramid networks. The code is available at https://github.com/gyyang23/AFPN. Guoyu Yang, Jie Lei 0002, Zhikuan Zhu, Siyu Cheng, Zunlei Feng, Ronghua Liang |
SMC | 2 |
| 2023 | DCAM: Disturbed class activation maps for weakly supervised semantic segmentation
Jie Lei 0002, Guoyu Yang, Shuaiwei Wang, Zunlei Feng, Ronghua Liang |
J. Vis. Commun. Image Represent. | 1 |
| 2022 | Enhanced Dual-Level Representations for Facial Expression RecognitionabstractFacial expression is an essential factor in conveying human emotional states and intentions. A common strategy used for facial expression recognition (FER) is encoding expression representations from facial images. Although remarkable advancement has been made, challenges due to large variations of expression patterns and unavoidable hard samples still remain. In this paper, we propose dual-level representation enhancements (DLRE) addressing these issues. On one hand, mid-level representation enhancement (MRE) is introduced to avoid expression representation learning being dominated by a limited number of highly discriminative patterns. On the other hand, high-level representation enhancement (HRE) is introduced to alleviate the disturbance of misclassified representations especially for hard samples. The proposed method not only has stronger generalization capability to handle different variations of expression patterns but also greater discriminative power to capture the subtle distinctions of hard samples. Experimental evaluation on four popular databases, CK+, Oulu-CASIA, RAF-DB, and AffectNet, shows that our method achieves more competitive results than other state-of-the-art methods. Jie Lei 0002, Zeyu Zou, Zunlei Feng, Ronghua Liang |
ICIP | 1 |
| 2021 | Edge-competing Pathological Liver Vessel Segmentation with Limited LabelsabstractThe microvascular invasion (MVI) is a major prognostic factor in hepatocellular carcinoma, which is one of the malignant tumors with the highest mortality rate. The diagnosis of MVI needs discovering the vessels that contain hepatocellular carcinoma cells and counting their number in each vessel, which depends heavily on experiences of the doctor, is largely subjective and time-consuming. However, there is no algorithm as yet tailored for the MVI detection from pathological images. This paper collects the first pathological liver image dataset containing $522$ whole slide images with labels of vessels, MVI, and hepatocellular carcinoma grades. The first and essential step for the automatic diagnosis of MVI is the accurate segmentation of vessels. The unique characteristics of pathological liver images, such as super-large size, multi-scale vessel, and blurred vessel edges, make the accurate vessel segmentation challenging. Based on the collected dataset, we propose an Edge-competing Vessel Segmentation Network (EVS-Net), which contains a segmentation network and two edge segmentation discriminators. The segmentation network, combined with an edge-aware self-supervision mechanism, is devised to conduct vessel segmentation with limited labeled patches. Meanwhile, two discriminators are introduced to distinguish whether the segmented vessel and background contain residual features in an adversarial manner. In the training stage, two discriminators are devised to compete for the predicted position of edges. Exhaustive experiments demonstrate that, with only limited labeled patches, EVS-Net achieves a close performance of fully supervised methods, which provides a convenient tool for the pathological liver vessel segmentation. Code is publicly available at https://github.com/wang97zh/EVS-Net. Zunlei Feng, Xinchao Wang, Xiuming Zhang, Lechao Cheng, Jie Lei 0002, Mingli Song |
AAAI | 6 |
| 2021 | Facial Expression Recognition by Expression-Specific Representation Swapping
Jie Lei 0002, Zeyu Zou, Zunlei Feng, Ronghua Liang |
ICANN (2) | 1 |
| 2021 | Mutual-Complementing Framework for Nuclei Detection and Segmentation in Pathology ImageabstractDetection and segmentation of nuclei are fundamental analysis operations in pathology images, the assessments derived from which serve as the gold standard for cancer diagnosis. Manual segmenting nuclei is expensive and time-consuming. What’s more, accurate segmentation detection of nuclei can be challenging due to the large appearance variation, conjoined and overlapping nuclei, and serious degeneration of histological structures. Supervised methods highly rely on massive annotated samples. The existing two unsupervised methods are prone to failure on degenerated samples. This paper proposes a Mutual-Complementing Framework (MCF) for nuclei detection and segmentation in pathology images. Two branches of MCF are trained in the mutual-complementing manner, where the detection branch complements the pseudo mask of the segmentation branch, while the progressive trained segmentation branch complements the missing nucleus templates through calculating the mask residual between the predicted mask and detected result. In the detection branch, two response map fusion strategies and gradient direction based postprocessing are devised to obtain the optimal detection response. Furthermore, the confidence loss combined with the synthetic samples and self-finetuning is adopted to train the segmentation network with only high confidence areas. Extensive experiments demonstrate that MCF achieves comparable performance with only a few nucleus patches as supervision. Especially, MCF possesses good robustness (only dropping by about 6%) on degenerated samples, which are critical and common cases in clinical diagnosis. Zunlei Feng, Xinchao Wang, Yining Mao, Thomas Li, Jie Lei 0002, Mingli Song |
ICCV | 6 |
| 2021 | Flexible Knowledge Distillation with an Evolutional Network PopulationabstractDeep neural networks have continually surpassed traditional methods on a variety of computer vision tasks. Though deep neural networks are very powerful, the large number of parameters and complex structures consume considerable storage and calculation time, making it hard to deploy with limited resources. To tackle this issue, many recently proposed knowledge distillation approaches are aimed at obtaining a small student network to imitate a large teacher network. However, the student network structure is pre-defined and may be hard to train. In this paper, we propose to distill knowledge with an evolutional student network population. The population is initialized with several basic structures and each network is evaluated by the imitation ability (i.e., fitness) to the teacher network. By reusing the weights, we provide five enhancement options to strengthen the networks with high fitness and abandon the weak ones. By changing the fitness criterion, we can select networks to meet different requirements, such as balancing size and accuracy. This allows one to find a superior student network structure that better imitates the teacher model from various aspects with easier training. The experimental results demonstrate the proposed method can achieve superior performance of knowledge distillation with flexible student structures. Jie Lei 0002, Mingli Song, Jianping Shen, Ronghua Liang |
ICME | 1 |
| 2021 | Boundary Knowledge Translation based Reference Semantic SegmentationabstractGiven a reference object of an unknown type in an image, human observers can effortlessly find the objects of the same category in another image and precisely tell their visual boundaries. Such visual cognition capability of humans seems absent from the current research spectrum of computer vision. Existing segmentation networks, for example, rely on a humongous amount of labeled data, which is laborious and costly to collect and annotate; besides, the performance of segmentation networks tend to downgrade as the number of the category increases. In this paper, we introduce a novel Reference semantic segmentation Network (Ref-Net) to conduct visual boundary knowledge translation. Ref-Net contains a Reference Segmentation Module (RSM) and a Boundary Knowledge Translation Module (BKTM). Inspired by the human recognition mechanism, RSM is devised only to segment the same category objects based on the features of the reference objects. BKTM, on the other hand, introduces two boundary discriminator branches to conduct inner and outer boundary segmentation of the target object in an adversarial manner, and translate the annotated boundary knowledge of open-source datasets into the segmentation network. Exhaustive experiments demonstrate that, with tens of finely-grained annotated samples as guidance, Ref-Net achieves results on par with fully supervised methods on six datasets. Our code can be found in the supplementary material. Lechao Cheng, Zunlei Feng, Xinchao Wang, Ya Jie Liu, Jie Lei 0002, Mingli Song |
IJCAI | 5 |
| 2021 | A Location Constrained Dual-Branch Network for Reliable Diagnosis of Jaw Tumors and Cysts
Jiacong Hu, Zunlei Feng, Yining Mao, Jie Lei 0002, Mingli Song |
MICCAI (7) | 4 |
| 2020 | Faster Self-adaptive Deep Stereo
Xinchao Wang, Jie Song 0011, Jie Lei 0002, Mingli Song |
ACCV (1) | 4 |
| 2020 | One-sample Guided Object Representation DisassemblingabstractThe ability to disassemble the features of objects and background is crucial for many machine learning tasks, including image classification, image editing, visual concepts learning, and so on. However, existing (semi-)supervised methods all need a large amount of annotated samples, while unsupervised methods can't handle real-world images with complicated backgrounds. In this paper, we introduce the One-sample Guided Object Representation Disassembling (One-GORD) method, which only requires one annotated sample for each object category to learn disassembled object representation from unannotated images. For the annotated one-sample, we first adopt some data augmentation strategies to generate some synthetic samples, which can guide the disassembling of the object features and background features. For the unannotated images, two self-supervised mechanisms: dual-swapping and fuzzy classification are introduced to disassemble object features from the background with the guidance of annotated one-sample. What's more, we devise two metrics to evaluate the disassembling performance from the perspective of representation and image, respectively. Experiments demonstrate that the One-GORD achieves competitive dissembling performance and can handle natural scenes with complicated backgrounds. Zunlei Feng, Yongming He, Xinchao Wang, Xin Gao 0032, Jie Lei 0002, Cheng Jin 0001, Mingli Song |
NeurIPS | 5 |
| 2019 | Action Parsing-Driven Video Summarization Based on Reinforcement LearningabstractHow to manage, store, and index large numbers of videos is an urgent problem to be solved. Although there are many video summarization models achieving good results, models based on low-level features cannot summarize important semantic information and models based on semantic analysis need related text descriptions that do not exist for most videos. As a consequence, the mining semantic information contained in the video itself is a more feasible way. In this paper, we propose an action parsing-driven video summarization model based on reinforcement learning. The model is mainly divided into two parts, video cut by action parsing and video summarization based on reinforcement learning. In the first part, a sequential multiple instance learning model is trained with weakly annotated data to solve the problem of full annotation’s time consuming and weak annotation’s ambiguity. In the second part, we design a deep recurrent neural network-based video summarization model that selects the most distinguishable frames comparing with other actions. Meanwhile, the quality of the extracted key frames could be evaluated by the categorization accuracy. Experiments and comparison with state-of-the-art methods demonstrate the advantage of the proposed approach. Jie Lei 0002, Qiao Luan, Xinhui Song, Xiao Liu 0012, Dapeng Tao, Mingli Song |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2018 | Selective Zero-Shot Classification with Augmented Attributes
Jie Song 0011, Chengchao Shen, Jie Lei 0002, Anxiang Zeng, Kairi Ou, Dacheng Tao, Mingli Song |
ECCV (9) | 3 |
| 2018 | Finding intrinsic color themes in images with human visual perception
Zunlei Feng, Wolong Yuan, Chunli Fu, Jie Lei 0002, Mingli Song |
Neurocomputing | 4 |
| 2018 | Scale insensitive and focus driven mobile screen defect detection in industry
Jie Lei 0002, Xin Gao 0032, Zunlei Feng, Huamou Qiu, Mingli Song |
Neurocomputing | 1 |
| 2017 | Graph-based color Gamut Mapping using neighbor metricabstractColors are displayed in different ways on various devices, such as cameras, screens and printers. In order to achieve consistent appearance on these devices, color management is usually used, where the core part is Gamut Mapping Algorithm (GMA). However, the widely adopted Point-wise Gamut Mapping Algorithms (PGMAs) have been restricted to compromise between color accuracy and details. In this work, we firstly split color space into small cubes through sampling colors from it. Then, we built a 26-neighbored graph with the sample colors as vertexes and perceptual color differences between adjacent vertexes as weights. Based on the above graph, a new Multi-source Shortest Paths Algorithm (MSS-PA) is proposed to establish color mapping relationships between out-of-gamut colors and colors in gamut boundary. In the MSSPA, distance of shortest path between nonadjacent vertexes are used to replace those calculated directly using CIEDE2000, which is useful to measure large color difference. Experimental results show that our method achieves superior performance on the aspect of keeping accuracy and preserving details compared with HPMinDE and SGCK. Zunlei Feng, Yongcheng Jing, Jie Lei 0002, Mingli Song |
ICME | 5 |
| 2016 | Which face is more attractive?abstractPeople are fond of sharing their photos of life experience in social networks, where the majority of photos are containing faces, especially in selfies. In fact, we are spontaneously evaluating the attractiveness of faces we see in daily life with personality and emotional traits within a single glance. Can we make a comparison on attractiveness between two face images automatically? In this paper, we capitalize on the fact that the difficulties in comparing the attractiveness of various face image pairs are not even and propose a hierarchical structure integrated with multiple comparative convolutional neural networks (CCNNs). First, five synthetic part faces are generated for each original face image to highlight the influences of parts in attractiveness. Both the original and synthetic face pairs are displayed to the subjects for attractiveness labeling. Multiple CCNNs are then trained on the labeled data with different kinds of pairs. Finally, all the CCNNs are combined to make a weighted comparison result for an input pair. The experimental results in two datasets show CCNNs can achieve significant performance improvement over other state-of-the-art methods. Quantifying the attractiveness of face images lends to many useful applications, such as choosing a better selfie and improving the results of related searching tasks. Jie Lei 0002, Zunlei Feng, Mingli Song, Dacheng Tao |
ICIP | 1 |
| 2016 | Learning deep classifiers with deep featuresabstractVisual separability between different objects in various image classification tasks is highly uneven. As a consequence, humans need different levels of detailed descriptions to separate objects in multi-granularity similarities. Meanwhile, deep networks, such as convolutional neural networks (C-NNs) have demonstrated great ability in multilevel representations for an object. Unfortunately, existing methods with deep networks in classification typically use the output of the last layer as the only feature to train flat N-way classifiers, which fail to fit the multi-granularity character. In this paper, by regarding different CNN layers as multiple levels of abstraction, we propose a deep decision tree (DDT) to distinguish objects sharing great appearance similarities with utilizing features in all layers. First, deep features in multiple layers are extracted from deep networks as the input for building a DDT. Next, in the training phase, the features from earlier layers are selected for splitting on a deeper node. Finally, multiple DDTs are bagged to make the final prediction by taking the majority vote. The experimental results in two datasets show DDT can greatly improve the classification accuracy in multi-grained tasks than flat models. Jie Lei 0002, Xinhui Song, Mingli Song, Chun Chen 0001 |
ICME | 1 |
| 2016 | Event-based large scale surveillance video summarization
Xinhui Song, Jie Lei 0002, Dapeng Tao, Guanhong Yuan, Mingli Song |
Neurocomputing | 3 |
| 2016 | DeepChart: Combining deep convolutional networks and deep belief networks in chart classification
Binbin Tang, Xiao Liu 0012, Jie Lei 0002, Mingli Song, Dapeng Tao, Shuifa Sun, Fangmin Dong |
Signal Process. | 3 |
| 2015 | Whole-body humanoid robot imitation with pose similarity evaluation
Jie Lei 0002, Mingli Song, Ze-Nian Li, Chun Chen 0001 |
Signal Process. | 1 |
| 2014 | Humanoid Robot Imitation with Pose Similarity Metric LearningabstractImitation is considered to be a kind of social learning that allows the transfer of information, actions, behaviours, etc. Whereas current robots are unable to perform as many tasks as human, it is a natural way for them to learn by imitations, just as human does. With the humanoid robots being more intelligent, the field of robot imitation has getting noticeable advance. In this paper, we focus on the pose imitation between a human and a humanoid robot and learning a similarity metric between human pose and robot pose. In contrast to recent approaches that capture human data using expensive motion captures or only imitate the upper body movements, our framework adopts a Kinect instead and can deal with complex, whole body motions by keeping both single pose balance and pose sequence balance. Meanwhile, different from previous work that employs subjective evaluation, we propose a pose similarity metric based on the shared structure of the motion spaces of human and robot. The qualitative and quantitative experimental results demonstrate a satisfactory imitation performance and indicate that the proposed pose similarity metric is discriminative. Jie Lei 0002, Mingli Song, Ze-Nian Li, Chun Chen 0001, Xianghua Xu, Shiliang Pu |
ICPR | 1 |