EDBT 2026 Demo / reviewers in the wild / expert
Zhiming Luo
dblp:75/9709
· DBLP profile ↗
106ranked-venue papers
5as first author
80since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 56 · 4 first-author · 45 since 2021Artificial intelligence and machine learning · 45 · 2 first-author · 34 since 2021Applied, interdisciplinary, general and emerging computing · 14 · 12 since 2021Systems, architecture and hardware · 12 · 6 since 2021Security and privacy · 5 · 5 since 2021Computer networks · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | OneFont: A Unified Agent for End-to-End Font CreationabstractDespite recent advancements in font generation, practitioners still grapple with a laborious trial-and-error workflow. To streamline this, we propose OneFont, an end-to-end framework that interprets user intents via free-form dialogue, seamlessly integrating both glyph synthesis and refinement modules. We introduce the Font with Thought (FwT) paradigm, reframing font design as a reasoning task where the model plans actions and articulates design rationales. OneFont’s core planner is trained via a two-stage regimen to master this paradigm. First, we instill reasoning abilities via Supervised Fine-Tuning (SFT) on a new, comprehensive benchmark of 1,500 font families we built. Second, we refine the model's policy with a novel reinforcement learning algorithm, Group Relative Policy Optimization (GRPO), guided by a hybrid reward that assesses visual fidelity, rationale coherence, and transformation correctness. Extensive experiments show OneFont significantly surpasses existing methods in design quality and stroke precision across diverse scripts, validated on our new benchmark. We will release our dataset, code, and models. Yingxin Lai, Yufei Liu 0003, Jiaxing Chai, Zhiming Luo, Shaozi Li |
AAAI | 5 |
| 2026 | Disentangled self-supervised video camouflaged object detection and salient object detection
Haoke Xiao, Lv Tang, Bo Li 0115, Zhiming Luo, Shaozi Li |
Neural Networks | 4 |
| 2026 | Mitigating low-frequency bias: Feature recalibration and frequency attention regularization for adversarial robustness
Kejia Zhang 0003, Juanjuan Weng, Yuanzheng Cai, Shaozi Li, Zhiming Luo |
Neural Networks | 5 |
| 2026 | MTSCL-Net: Multi-level temporal spatial contrastive learning for robust breast tumor segmentation in DCE-MRI
Jiezhou He, Zhiming Luo, Songzhi Su, Shaozi Li |
Pattern Recognit. | 3 |
| 2026 | FADMB: Fully attention-based dual memory bank network for weakly supervised video anomaly detection
Zhiming Luo, Shuheng Huang, Jianzhe Gao, Shaozi Li |
Pattern Recognit. | 1 |
| 2026 | Exploring Frequencies via Feature Mixing and Meta-Learning for Improving Adversarial TransferabilityabstractRecent studies have shown that Deep Neural Networks (DNNs) are susceptible to adversarial attacks, with frequency-domain analysis underscoring the significance of high-frequency components in influencing model predictions. Conversely, targeting low-frequency components has been effective in enhancing attack transferability on black-box models. In this study, we introduce a frequency decomposition-based feature mixing method to exploit these frequency characteristics in both clean and adversarial samples. Our findings suggest that incorporating features of clean samples into adversarial features extracted from adversarial examples is more effective in attacking normally-trained models, while combining clean features with the adversarial features extracted from low-frequency parts decomposed from the adversarial samples yields better results in attacking defense models. However, a conflict issue arises when these two mixing approaches are employed simultaneously. To tackle the issue, we propose a cross-frequency meta-optimization approach comprising the meta-train step, meta-test step, and final update. In the meta-train step, we leverage the low-frequency components of adversarial samples to boost the transferability of attacks against defense models. Meanwhile, in the meta-test step, we utilize adversarial samples to stabilize gradients, thereby enhancing the attack's transferability against normally trained models. For the final update, we update the adversarial sample based on the gradients obtained from both meta-train and meta-test steps. Our proposed method is evaluated through extensive experiments on the ImageNet-Compatible dataset, affirming its effectiveness in improving the transferability of attacks on both normally-trained CNNs and defense models. The source code is available at https://github.com/WJJLL/MetaSSA. Juanjuan Weng, Zhiming Luo, Shaozi Li |
IEEE Trans. Image Process. | 2 |
| 2025 | Long-Tailed Out-of-Distribution Detection: Prioritizing Attention to TailabstractCurrent out-of-distribution (OOD) detection methods typically assume balanced in-distribution (ID) data, while most real-world data follow a long-tailed distribution. Previous approaches to long-tailed OOD detection often involve balancing the ID data by reducing the semantics of head classes. However, this reduction can severely affect the classification accuracy of ID data. The main challenge of this task lies in the severe lack of features for tail classes, leading to confusion with OOD data. To tackle this issue, we introduce a novel Prioritizing Attention to Tail (PATT) method using augmentation instead of reduction. Our main intuition involves using a mixture of von Mises-Fisher (vMF) distributions to model the ID data and a temperature scaling module to boost the confidence of ID data. This enables us to generate infinite contrastive pairs, implicitly enhancing the semantics of ID classes while promoting differentiation between ID and OOD data. To further strengthen the detection of OOD data without compromising the classification performance of ID data, we propose feature calibration during the inference phase. By extracting an attention weight from the training set that prioritizes the tail classes and reduces the confidence in OOD data, we improve the OOD detection capability. Extensive experiments verified that our method outperforms the current state-of-the-art methods on various benchmarks. Yina He, Yongcun Zhang, Juanjuan Weng, Shaozi Li, Zhiming Luo |
AAAI | 6 |
| 2025 | Font-Agent: Enhancing Font Understanding with Large Language ModelsabstractThe rapid development of generative models has significantly advanced font generation. However, limited exploration has been devoted to the evaluation and interpretability of graphical fonts. Existing quality assessment models can only provide basic visual analyses, such as recognizing clarity and brightness, without in-depth explanations. To address these limitations, we first constructed a large-scale multimodal dataset named the Diversity Font Dataset (DFD), comprising 135,000 font-text pairs. This dataset encompasses a wide range of generated font types and annotations, including language descriptions and quality assessments, thus providing a robust foundation for training and evaluating font analysis models. Based on this dataset, we developed a font agent built upon a Vision-Language Model (VLM) aiming to enhance font quality assessment and offer interpretable question-answering capabilities. Alongside the original visual encoder in VLM, we integrated an Edge-Aware Traces (EAT) module to capture detailed edge information of font strokes and components. Furthermore, we introduced a Dynamic Direct Preference Optimization (D-DPO) strategy to facilitate efficient model fine-tuning. Experimental results demonstrate that Font-Agent achieves state-of-the-art performance on the established dataset. To further evaluate the generalization ability of our algorithm, we conducted additional experiments on several public datasets. The results highlight the notable advantage of Font-Agent in both assessing the quality of generated fonts and comprehending their content. Yingxin Lai, Cuijie Xu, Haitian Shi, Zhiming Luo, Shaozi Li |
CVPR | 6 |
| 2025 | Towards Adversarial Robustness via Debiased High-Confidence Logit AlignmentabstractDespite the remarkable progress of deep neural networks (DNNs) in various visual tasks, their vulnerability to adversarial examples raises significant security concerns. Recent adversarial training methods leverage inverse adversarial attacks to generate high-confidence examples, aiming to align adversarial distributions with high-confidence class regions. However, our investigation reveals that under inverse adversarial attacks, high-confidence outputs are influenced by biased feature activations, causing models to rely on background features that lack a causal relationship with the labels. This spurious correlation bias leads to overfitting irrelevant background features during adversarial training, thereby degrading the model's robust performance and generalization capabilities. To address this issue, we propose Debiased High-Confidence Adversarial Training (DHAT), a novel approach that aligns adversarial logits with debiased high-confidence logits and restores proper attention by enhancing foreground logit orthogonality. Extensive experiments demonstrate that DHAT achieves state-of-the-art robustness on both CIFAR and ImageNet-1K benchmarks, while significantly improving generalization by mitigating the feature bias inherent in inverse adversarial training approaches. Code is available at https://github.com/KejiaZhang-Robust/DHAT. Kejia Zhang 0003, Juanjuan Weng, Shaozi Li, Zhiming Luo |
ICCV | 4 |
| 2025 | Forgery-Aware Adaptive CLIP for Generalizable Face Forgery Detection
Yongcun Zhang, Yingxin Lai, Guimin Shi, Zhiming Luo |
ICIC (3) | 6 |
| 2025 | SynPo: Boosting Training-Free Few-Shot Medical Segmentation via High-Quality Negative Prompts
Yufei Liu 0003, Haoke Xiao, Jiaxing Chai, Yongcun Zhang, Zijie Meng, Zhiming Luo |
MICCAI (5) | 7 |
| 2025 | Attentive Multi-Kernel Feature Aggregation Network for Cross-View Geo-LocalizationabstractCross-view geo-localization, which aims to match images of the same scene captured from diverse viewpoints from drone and satellite, presents a persistent challenge due to significant geometric distortions and appearance variations. Existing methods lack a comprehensive exploration and dynamic integration of spatial and channel attention mechanisms, while primarily focusing on extracting single-scale features. The former results in the model failing to focus on key regions of the feature map, while the latter may result in the inability to capture information at different scales. In this paper, we propose a novel Attentive Multi-kernel Feature Aggregation (AMFA) Network that incorporates a Synergistic Attention (SA) Module and a Multi-kernel Inception (MI) Module, which effectively addresses the challenges posed by significant variations in target building regions and contextual diversity in cross-view tasks. The SA Module adaptively fuses channel and spatial information to focus on the most discriminative regions within the feature maps. Building upon this, the MI Module extracts multi-scale features through parallel convolutional kernels, enabling a more comprehensive scene representation. Experiments on the University-1652 and SUES-200 datasets show that our method achieves state-of-the-art performance in cross-view geo-localization tasks. Shuheng Huang, Deyong Wu, Jinliang Lin, Zhiming Luo |
ICMR | 5 |
| 2025 | HCCM: Hierarchical Cross-Granularity Contrastive and Matching Learning for Natural Language-Guided Drones
Hao Ruan, Jinliang Lin, Yingxin Lai, Zhiming Luo, Shaozi Li |
ACM Multimedia | 4 |
| 2025 | Advancing Fine-Grained Spine Segmentation Through Visual-Language Model with Omni- and Pixel-Level Semantic Enhancements
Jianlong Cai, Sheng Lian, Dengfeng Pan, Guang-Yong Chen, Lei Li 0048, Zhiming Luo, Shuo Li 0001 |
PRCV (14) | 6 |
| 2025 | HRCUNet: Hierarchical Region Contrastive Learning for Segmentation of Breast Tumors in DCE-MRIabstractABSTRACT Segmenting breast tumors from dynamic contrast‐enhanced magnetic resonance images is a critical step in the early detection and diagnosis of breast cancer. However, this task becomes significantly more challenging due to the diverse shapes and sizes of tumors, which make it difficult to establish a unified perception field for modeling them. Moreover, tumor regions are often subtle or imperceptible during early detection, exacerbating the issue of extreme class imbalance. This imbalance can lead to biased training and challenge accurately segmenting tumor regions from the predominant normal tissues. To address these issues, we propose a hierarchical region contrastive learning approach for breast tumor segmentation. Our approach introduces a novel hierarchical region contrastive learning loss function that addresses the class imbalance problem. This loss function encourages the model to create a clear separation between feature embeddings by maximizing the inter‐class margin and minimizing the intra‐class distance across different levels of the feature space. In addition, we design a novel Attention‐based 3D Multi‐scale Feature Fusion Residual Module to explore more granular multi‐scale representations to improve the feature learning ability of tumors. Extensive experiments on two breast DCE‐MRI datasets demonstrate that the proposed algorithm is more competitive against several state‐of‐the‐art approaches under different segmentation metrics. Jiezhou He, Zhiming Luo, Wei Peng 0009, Songzhi Su, Shaozi Li |
Concurr. Comput. Pract. Exp. | 2 |
| 2025 | Niching-Based Two-Stage Differential Evolution for Feature SelectionabstractABSTRACT Feature selection is a vital preprocessing step aimed at identifying a subset of the most relevant features from high‐dimensional data to enhance model performance and reduce computational complexity. Differential Evolution (DE) algorithms have been extensively applied to this task by iteratively optimizing the selection probabilities or weights of features. However, many existing DE‐based approaches suffer from premature convergence and local optima entrapment due to an insufficient balance between global exploration and local exploitation. To address these challenges, we propose a niching‐based two‐stage mutation DE algorithm for feature selection. Firstly, the mutual information is utilized to initialize the population and reduce the number of features. Then, an improved two‐stage mutation operator is employed to balance the algorithm's exploitation and exploration. Additionally, duplicate individuals generated during the evolutionary process are replaced using a one‐bit evolutionary mutation to aid population evolution. Experimental results on 16 benchmark datasets demonstrate that the proposed method achieves superior classification accuracy compared to several state‐of‐the‐art approaches, validating its effectiveness and robustness in diverse feature selection scenarios. Deyong Wu, Zhiming Luo, Jiezhou He, Shaozi Li |
Concurr. Comput. Pract. Exp. | 2 |
| 2025 | Research on Transient Stability Model of Power System Based on Periodic Enhanced InformerabstractABSTRACT In view of the nonlinearity and complexity of power system connection, it is difficult to directly identify the key characteristics of power system stability, while there are characteristics unrelated to system stability. For this reason, this paper proposes a method to distinguish the transient stability margin of power grid based on long short memory recurrent network. This method uses unbalanced data clustering algorithm to perform clustering analysis on the historical data set, extract a variety of typical instability characteristics, and consider the periodic characteristics of the unstable sequence. The unstable value in the same period of the input sequence is combined with the output of the long and short term memory network (LSTM) network to output the instability prediction results. At the same time, the key value of the periodic instability in the input sequence is connected in series with the output of the Informer model to build a full connection layer, and combined with the convolutional neural network to build a power system transient stability model based on periodically enhanced Informer. Experimental results show that the model effectively solves the problem of periodic mode attenuation of non‐stationary load series and alleviates the redundancy problem of key features of power grid stability by building a multi‐dimensional periodic feature extraction channel. Zhiming Luo, Shaozi Li |
Concurr. Comput. Pract. Exp. | 2 |
| 2025 | Hierarchical vertical-aware and adaptive multi-scale network for three-dimensional object detection in maritime environmentsabstractAccurate three-dimensional (3D) object detection in maritime environments is critical for autonomous navigation. However, it remains challenging because of sparse point clouds, complex vertical structures, and extreme object scale variations. Existing 3D detectors are primarily designed for road scenes and often perform poorly in such conditions. Therefore, we propose a Hierarchical Vertical-aware and Adaptive Multi-scale Network (HVAM-Net), an anchor-free, single-stage deep learning framework tailored for maritime scenarios. HVAM-Net integrates three core modules: (1) a Hierarchical Pillar Encoding module that enhances vertical representation via exponential stratification and semantic-aware fusion; (2) an Adaptive Multi-scale Feature Extraction module that captures diverse spatial contexts via parallel atrous convolutions and attention-guided fusion; and (3) an Attention-Guided Dynamic Sampling module that refines upsampling by learning adaptive spatial offsets, enhancing semantic consistency in sparse regions. The effectiveness of HVAM-Net is validated through comprehensive comparisons with state-of-the-art 3D object detection methods. Experiments show that HVAM-Net achieves mean Average Precision scores of 86.7 %, 78 %, and 88 % on the self-collected, Thames River vessel, and simulated datasets, respectively, outperforming all baseline methods. Moreover, its resilience under adverse weather conditions and varying light detection and ranging configurations further confirms the strong generalization capability of this artificial intelligence-based approach in real-world maritime environments. Yutang Wang, Hangbin Wu, Yuanhang Kong, Zhiming Luo, Chun Liu 0003 |
Eng. Appl. Artif. Intell. | 6 |
| 2025 | Pick and mix reliable pseudo labels for scribble-supervised medical image segmentation
Jiawei Su, Zhiming Luo, Dazhen Lin, Lihui Lin, Shaozi Li |
Neurocomputing | 2 |
| 2025 | Cross-modality average precision optimization for visible thermal person re-identification
Yongguo Ling, Zhiming Luo, Dazhen Lin, Shaozi Li, Min Jiang 0005, Nicu Sebe, Zhun Zhong |
Pattern Recognit. | 2 |
| 2025 | Detect Changes Like Humans: Incorporating Semantic Priors for Improved Change DetectionabstractWhen given two similar images, humans identify their differences by comparing the appearance (e.g., color, texture) with the help of semantics (e.g., objects, relations). However, mainstream binary change detection models adopt a supervised training paradigm, where the annotated binary change map is the main constraint. Thus, such methods primarily emphasize difference-aware features between bi-temporal images, and the semantic understanding of changed landscapes is undermined, resulting in limited accuracy in the face of noise and illumination variations. To this end, this paper explores incorporating semantic priors from visual foundation models to improve the ability to detect changes. Firstly, we propose a Semantic-Aware Change Detection network (SA-CDNet), which transfers the knowledge of visual foundation models (i.e., FastSAM) to change detection. Inspired by the human visual paradigm, a novel dual-stream feature decoder is derived to distinguish changes by combining semantic-aware features and difference-aware features. Secondly, we explore a single-temporal pre-training strategy for better adaptation of visual foundation models. With pseudo-change data constructed from single-temporal segmentation datasets, we employ an extra branch of proxy semantic segmentation task for pre-training. We explore various settings like dataset combinations and landscape types, thus providing valuable insights. Experimental results on five challenging benchmarks demonstrate the superiority of our method over the existing state-of- the-art methods. The code is available at SA-CD. Yuhang Gan, Wenjie Xuan, Zhiming Luo, Zengmao Wang, Juhua Liu, Bo Du 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2025 | Improving Transferable Targeted Adversarial Attack via Normalized Logit Calibration and Truncated Feature MixingabstractThis paper aims to enhance the transferability of adversarial samples in targeted attacks, where attack success rates remain comparatively low. To achieve this objective, we propose two distinct techniques for improving the targeted transferability from the loss and feature aspects. First, in previous approaches, logit calibrations used in targeted attacks primarily focus on the logit margin between the targeted class and the untargeted classes among samples, neglecting the standard deviation of the logit. In contrast, we introduce a new normalized logit calibration method that jointly considers the logit margin and the standard deviation of logits. This approach effectively calibrates the logits, enhancing the targeted transferability. Second, previous studies have demonstrated that mixing the features of clean samples during optimization can significantly increase transferability. Building upon this, we further investigate a truncated feature mixing method to reduce the impact of the source training model, resulting in additional improvements. The truncated feature is determined by removing the Rank-1 feature associated with the largest singular value decomposed from the high-level convolutional layers of the clean sample. Extensive experiments conducted on the ImageNet-Compatible, CIFAR-10 and ImageNet-1k datasets demonstrate the individual and mutual benefits of our proposed two components, which outperform the state-of-the-art methods by a large margin in black-box targeted attacks. Juanjuan Weng, Zhiming Luo, Shaozi Li |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2025 | A Self-Adaptive Feature Extraction Method for Aerial-View Geo-LocalizationabstractCross-view geo-localization aims to match the same geographic location from different view images, e.g., drone-view images and geo-referenced satellite-view images. Due to UAV cameras' different shooting angles and heights, the scale of the same captured target building in the drone-view images varies greatly. Meanwhile, there is a difference in size and floor area for different geographic locations in the real world, such as towers and stadiums, which also leads to scale variants of geographic targets in the images. However, existing methods mainly focus on extracting the fine-grained information of the geographic targets or the contextual information of the surrounding area, which overlook the robust feature for scale changes and the importance of feature alignment. In this study, we argue that the key underpinning of this task is to train a network to mine a discriminative representation against scale variants. To this end, we design an effective and novel end-to-end network called Self-Adaptive Feature Extraction Network (Safe-Net) to extract powerful scale-invariant features in a self-adaptive manner. Safe-Net includes a global representation-guided feature alignment module and a saliency-guided feature partition module. The former applies an affine transformation guided by the global feature for adaptive feature alignment. Without extra region annotations, the latter computes saliency distribution for different regions of the image and adopts the saliency information to guide a self-adaptive feature partition on the feature map to learn a visual representation against scale variants. Experiments on two prevailing large-scale aerial-view geo-localization benchmarks, i.e., University-1652 and SUES-200, show that the proposed method achieves state-of-the-art results. In addition, our proposed Safe-Net has a significant scale adaptive capability and can extract robust feature representations for those query images with small target buildings. The source code of this study is available at: https://github.com/AggMan96/Safe-Net. Jinliang Lin, Zhiming Luo, Dazhen Lin, Shaozi Li, Zhun Zhong |
IEEE Trans. Image Process. | 2 |
| 2025 | Dual-Modality-Shared Learning and Label Refinement for Unsupervised Visible-Infrared Person ReIDabstractUnsupervised visible-infrared person re-identification (USVI-ReID) aims to match a person across two modalities without annotations. Current research primarily addresses the modality gap by establishing cross-modality correspondences through matching algorithms and utilizing memory banks for contrastive learning. However, the inherent noise in pseudo labels and neglect of hard samples often limit the efficacy of cross-modality learning. In this article, we propose a dual-modality-shared learning and label refinement (DLLR) algorithm for USVI-ReID. First, we leverage a cluster similarity matching (CSM) module and a cluster relationship-based label refinement (CRLR) algorithm to create and refine pseudo labels. Then, we adopt a weighted modality-shared memory (WMM) to construct memory banks by jointly considering sample distribution and feature differences, thereby enhancing the effectiveness of cross-modality learning. Extensive experiments on three publicly available datasets validate the effectiveness of our proposed method, which outperforms state-of-the-art methods. The code is available at https://github.com/CharRic/DLLR . Licun Dai, Zhiming Luo, Yongguo Ling, Jiaxing Chai, Shaozi Li |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2024 | Diversity-Authenticity Co-constrained Stylization for Federated Domain Generalization in Person Re-identificationabstractThis paper tackles the problem of federated domain generalization in person re-identification (FedDG re-ID), aiming to learn a model generalizable to unseen domains with decentralized source domains. Previous methods mainly focus on preventing local overfitting. However, the direction of diversifying local data through stylization for model training is largely overlooked. This direction is popular in domain generalization but will encounter two issues under federated scenario: (1) Most stylization methods require the centralization of multiple domains to generate novel styles but this is not applicable under decentralized constraint. (2) The authenticity of generated data cannot be ensured especially given limited local data, which may impair the model optimization. To solve these two problems, we propose the Diversity-Authenticity Co-constrained Stylization (DACS), which can generate diverse and authentic data for learning robust local model. Specifically, we deploy a style transformation model on each domain to generate novel data with two constraints: (1) A diversity constraint is designed to increase data diversity, which enlarges the Wasserstein distance between the original and transformed data; (2) An authenticity constraint is proposed to ensure data authenticity, which enforces the transformed data to be easily/hardly recognized by the local-side global/local model. Extensive experiments demonstrate the effectiveness of the proposed DACS and show that DACS achieves state-of-the-art performance for FedDG re-ID. Fengxiang Yang, Zhun Zhong, Zhiming Luo, Yifan He 0002, Shaozi Li, Nicu Sebe |
AAAI | 3 |
| 2024 | Learning to Distinguish Samples for Generalized Category Discovery
Fengxiang Yang, Nan Pu, Wenjing Li 0005, Zhiming Luo, Shaozi Li, Nicu Sebe, Zhun Zhong |
ECCV (65) | 4 |
| 2024 | CC-DA: Cross-Domain Consistency Data Augmentation for 3D Tumor SegmentationabstractDeep learning-based tumor segmentation in 3D medical images faces the challenges of limited annotated data and class imbalance. In this paper, we proposed a novel Cross-domain Consistency Data Augmentation (CC-DA) for 3D tumor segmentation. Specifically, we copy the tumor from source data and apply random transformations to enhance its diversity. Then, we paste the enhanced tumor into the organ area of target data to generate a new sample. This process can alleviate class imbalance by regulating the merged tumor pixel ratio. To further enhance the generated data credibility, we proposed a domain consistency constraint that aligns the source data distribution with the target data distribution. We conduct extensive experiments on KiTS19 and LiTS17 datasets. The promising results clearly show that our CC-DA method can effectively improve the existing state-of-the-art 3D tumor segmentation performance. Jiezhou He, Zhiming Luo, Wei Peng 0009, Songzhi Su, Shaozi Li |
ICASSP | 2 |
| 2024 | Selective Domain-Invariant Feature for Generalizable Deepfake DetectionabstractWith diverse presentation forgery methods emerging continually, detecting the authenticity of images has drawn growing attention. Although existing methods have achieved impressive accuracy in training dataset detection, they still perform poorly in the unseen domain and suffer from forgery of irrelevant information such as background and identity, affecting generalizability. To solve this problem, we proposed a novel framework Selective Domain-Invariant Feature (SDIF), which reduces the sensitivity to face forgery by fusing content features and styles. Specifically, we first use a Farthest-Point Sampling (FPS) training strategy to construct a task-relevant style sample representation space for fusing with content features. Then, we propose a dynamic feature extraction module to generate features with diverse styles to improve the performance and effectiveness of the feature extractor. Finally, a domain separation strategy is used to retain domain-related features to help distinguish between real and fake faces. Both qualitative and quantitative results in existing benchmarks and proposals demonstrate the effectiveness of our approach. Yingxin Lai, Yifan He 0002, Zhiming Luo, Shaozi Li |
ICASSP | 4 |
| 2024 | Modality-Dependent Sentiments Exploring for Multi-Modal Sentiment ClassificationabstractRecognizing human feelings from image and text is a core challenge of multi-modal data analysis, often applied in personalized advertising. Previous works aim at exploring the shared features, which are the matched contents between images and texts. However, the modality-dependent sentiment information (private features) in each modality is usually ignored by cross-modal interactions, the real sentiment is often reflected in one modality. In this paper, we propose a Modality-Dependent Sentiment Exploring framework (MDSE). First, to exploit the private features, we compare shared features with original image or text features, identifying previously overlooked unimodal features. Fusing the private and shared features can make the model more robust. Second, in order to obtain unified sentiment representations, we treat unimodal features and multi-modal fused features equally. We introduce a Modality-Agnostic Contrastive Loss (MACL) that performs contrastive learning between unimodal features and multi-modal fused features. The MACL can fully exploit sentiment information from multi-modal data and reduce the modality gap. Experiments on four public datasets demonstrate the effectiveness of our MDSE compared with existing methods. The full codes are available at https://github.com/royal-dargon/MDSE. Jingzhe Li, Chengji Wang, Zhiming Luo, Yuxian Wu, Xingpeng Jiang |
ICASSP | 3 |
| 2024 | DEEPOREDNET: Contrastive Learning-Based Attention-Weighted Dual Channel Residual Network for Ocular Redness AssessmentabstractOcular redness is highly prevalent worldwide and often accompanied by pain, discomfort, and vision problems, making it an essential signal for monitoring disease development and prognosis. Understanding the category of ocular redness is crucial for health. However, the intricate vascular structure of the ocular surface poses challenges in extracting meaningful features through conventional methods. Moreover, the subtle variations in the signs contribute to the difficulty in achieving accurate discrimination. In this paper, we propose a novel approach named contrastive learning-based attention-weighted dual channel residual network (DeepORedNet) to address the challenging problem. The effectiveness of the proposed network architecture has been validated through detailed experiments. Our proposed DeepORedNet achieves superior performance compared to the baseline models across all evaluation metrics. We hope the proposed framework can effectively facilitate the clinic assessment of ocular redness. Shaopan Wang, Jiezhou He, Jiaoyue Hu, Zuguo Liu, Zhiming Luo |
ICASSP | 6 |
| 2024 | Zero-Shot Co-Salient Object Detection FrameworkabstractCo-salient Object Detection (CoSOD) endeavors to replicate the human visual system’s capacity to recognize common and salient objects within a collection of images. Despite recent advancements in deep learning models, these models still rely on training with well-annotated CoSOD datasets. The exploration of training-free zero-shot CoSOD frameworks has been limited. In this paper, taking inspiration from the zero-shot transfer capabilities of foundational computer vision models, we introduce the first zero-shot CoSOD framework that harnesses these models without any training process. To achieve this, we introduce two novel components in our proposed framework: the group prompt generation (GPG) module and the co-saliency map generation (CMP) module. We evaluate the framework’s performance on widely-used datasets and observe impressive results. Our approach surpasses existing unsupervised methods and even outperforms fully supervised methods developed before 2020, while remaining competitive with some fully supervised methods developed before 2022. Haoke Xiao, Lv Tang, Bo Li 0115, Zhiming Luo, Shaozi Li |
ICASSP | 4 |
| 2024 | Omni-Granularity Embedding Network for Text-to-Image Person RetrievalabstractText-to-image person retrieval aims to identify the desired individual based on a textual description. As an instance-level retrieval problem, it has a large intra-class variance and a small inter-class variance. Although significant progress has been made, the omni-granularity matching issue remains unaddressed. Omni-granularity matching involves aligning words with multi-granularity image regions, challenging models to learn in an omni-granularity embedding space. In this paper, we introduce a novel Omni-Granularity Embedding Network (OGEN) for person representation learning. It addresses the omni-granularity matching issue by developing a Cross-Granularity Aggregation Module (CGAM). This module dynamically consolidates diverse granularity features for learning granularity-dependent and omni-granularity person representations. Additionally, a teacher-student knowledge transfer framework is introduced to minimize the inter-modality discrepancy, allowing CGAM to focus on modality-shared semantics. Due to the effectiveness of CGAM and the knowledge transfer framework, our OGEN enhances the Rank-1 accuracy of the Baseline by 8.54%, 9.89%, and 11.09% on three public datasets, respectively. Chengji Wang, Zhiming Luo, Shaozi Li |
ICME | 2 |
| 2024 | Mask Matching Network for Self-supervised Few-shot Medical Image SegmentationabstractExisting few-shot segmentation methods have achieved remarkable progress in medical image segmentation. However, many existing methods yield incomplete and discontinuous boundary predictions. In contrast, the Segment Anything Model (SAM) consistently produces clear, continuous, and comprehensive segmentation boundaries. Building on this observation, we propose a new two-step network called Mask Matching Network (MMNet) to introduce extra knowledge learned by SAM in natural images for few-shot medical image segmentation. Firstly, Q-Net has been utilized to locate some Regions of Interest (RoI) as prompts for SAM, allowing for the automatic generation of masks without relying on manual prompts. Secondly, we propose a novel Mask Matching Module (MMM), which considers both feature similarity and volume similarity as guidance to collaboratively mine the final segmentation from proposal masks. MMNet achieves state-of-the-art performance with remarkable improvements on two widely used datasets, abdominal MR (ABD) and cardiac MR (CMR), under two different settings. Zeyun Zhao, Jianzhe Gao, Zhiming Luo, Shaozi Li |
ICME | 4 |
| 2024 | TSESNet: Temporal-Spatial Enhanced Breast Tumor Segmentation in DCE-MRI Using Feature Perception and Separability
Jiezhou He, Zhiming Luo, Songzhi Su, Shaozi Li |
IJCAI | 3 |
| 2024 | A Collaborative Framework Using Multimodal Data and Adaptive Noise for Human Behavior Anomaly DetectionabstractHuman behavior anomaly detection in video aims to identify unusual behaviors that are crucial for public safety. Recently, there has been an increase in reconstruction or prediction-based methods that integrate diverse modal features to enhance anomaly detection. However, they use methods that independently or directly fusion multimodal features without fully considering the collaborative potential between multimodal features, which are susceptible to interference from semantic differences, thereby impacting detection performance. In contrast, we design a collaborative framework using multimodal data and adaptive noise for behavior anomaly detection. Our framework detects anomalies by analyzing the contrastive differences between two modalities alongside single-frame reconstruction errors. Specifically, we first learn the correlation between RGB and skeletal modalities for normal behavior through contrastive learning and use inter-modal contrast difference to detect motion anomalies. Additionally, we propose a single-frame reconstruction network that adaptively adds noise based on the importance of foreground features to detect appearance anomalies. Anomalies often occur in the motion foreground, and increasing noise in this area can make it more difficult to reconstruct anomalies. Extensive experiments validate the state-of-the-art performance of our method on three public datasets. Jianzhe Gao, Kejia Zhang 0003, Yifan He 0002, Zhiming Luo, Shaozi Li |
IJCNN | 5 |
| 2024 | QueryNet: A Unified Framework for Accurate Polyp Segmentation and Detection
Jiaxing Chai, Zhiming Luo, Jianzhe Gao, Licun Dai, Yingxin Lai, Shaozi Li |
MICCAI (8) | 2 |
| 2024 | VCLIPSeg: Voxel-Wise CLIP-Enhanced Model for Semi-supervised Medical Image Segmentation
Lei Li 0048, Sheng Lian, Zhiming Luo, Beizhan Wang, Shaozi Li |
MICCAI (9) | 3 |
| 2024 | A Multilevel Guidance-Exploration Network and Behavior-Scene Matching Method for Human Behavior Anomaly DetectionabstractHuman behavior anomaly detection aims to identify unusual human actions, playing a crucial role in intelligent surveillance and other areas. The current mainstream methods still adopt reconstruction or future frame prediction techniques. However, reconstructing or predicting low-level pixel features easily enables the network to achieve overly strong generalization ability, allowing anomalies to be reconstructed or predicted as effectively as normal data. Different from their methods, inspired by the Student-Teacher Network, we propose a novel framework called the Multilevel Guidance-Exploration Network (MGENet), which detects anomalies through the difference in high-level representation between the Guidance and Exploration network. Specifically, we first utilize the Normalizing Flow that takes skeletal keypoints as input to guide an RGB encoder, which takes unmasked RGB frames as input, to explore latent motion features. Then, the RGB encoder guides the mask encoder, which takes masked RGB frames as input, to explore the latent appearance feature. Additionally, we design a Behavior-Scene Matching Module to detect scene-related behavioral anomalies. Extensive experiments demonstrate that our proposed method achieves state-of-the-art performance on ShanghaiTech and UBnormal datasets, with AUC of 86.9% and 74.3%, respectively. The code is available at https://github.com/molu-ggg/GENet. Zhiming Luo, Jianzhe Gao, Yingxin Lai, Yifan He 0002, Shaozi Li |
ACM Multimedia | 2 |
| 2024 | Cross-Modality Perturbation Synergy Attack for Person Re-identificationabstractIn recent years, there has been significant research focusing on addressing security concerns in single-modal person re-identification (ReID) systems that are based on RGB images. However, the safety of cross-modality scenarios, which are more commonly encountered in practical applications involving images captured by infrared cameras, has not received adequate attention. The main challenge in cross-modality ReID lies in effectively dealing with visual differences between different modalities. For instance, infrared images are typically grayscale, unlike visible images that contain color information. Existing attack methods have primarily focused on the characteristics of the visible image modality, overlooking the features of other modalities and the variations in data distribution among different modalities. This oversight can potentially undermine the effectiveness of these methods in image retrieval across diverse modalities. This study represents the first exploration into the security of cross-modality ReID models and proposes a universal perturbation attack specifically designed for cross-modality ReID. This attack optimizes perturbations by leveraging gradients from diverse modality data, thereby disrupting the discriminator and reinforcing the differences between modalities. We conducted experiments on three widely used cross-modality datasets, namely RegDB, SYSU, and LLCM. The results not only demonstrate the effectiveness of our method but also provide insights for future improvements in the robustness of cross-modality ReID systems. Yunpeng Gong, Zhun Zhong, Yansong Qu, Zhiming Luo, Rongrong Ji, Min Jiang 0005 |
NeurIPS | 4 |
| 2024 | CPNet: Cross Prototype Network for Few-Shot Medical Image Segmentation
Zeyun Zhao, Jianzhe Gao, Zhiming Luo, Shaozi Li |
PRCV (15) | 3 |
| 2024 | Comparative evaluation of recent universal adversarial perturbations in image classification
Juanjuan Weng, Zhiming Luo, Dazhen Lin, Shaozi Li |
Comput. Secur. | 2 |
| 2024 | Learning transferable targeted universal adversarial perturbations by sequential meta-learning
Juanjuan Weng, Zhiming Luo, Dazhen Lin, Shaozi Li |
Comput. Secur. | 2 |
| 2024 | Learning multi-organ and tumor segmentation from partially labeled datasets by a conditional dynamic attention networkabstractSummary Multi‐organ segmentation is a critical prerequisite for many clinical applications. Deep learning‐based approaches have recently achieved promising results on this task. However, they heavily rely on massive data with multi‐organ annotated, which is labor‐ and expert‐intensive and thus difficult to obtain. In contrast, single‐organ datasets are easier to acquire, and many well‐annotated ones are publicly available. It leads to the partially labeled issue: How to learn a unified multi‐organ segmentation model from several single‐organ datasets? Pseudo‐label‐based methods and conditional information‐based methods make up the majority of existing solutions, where the former largely depends on the accuracy of pseudo‐labels, and the latter has a limited capacity for task‐related features. In this paper, we propose the Conditional Dynamic Attention Network (CDANet). Our approach is designed with two key components: (1) multisource parameter generator, fusing the conditional and multiscale information to better distinguish among different tasks, and (2) dynamic attention module, promoting more attention to task‐related features. We have conducted extensive experiments on seven partially labeled challenging datasets. The results show that our method achieved competitive results compared with the advanced approaches, with an average Dice score of 75.08%. Additionally, the Hausdorff Distance is 26.31, which is a competitive result. Lei Li 0048, Sheng Lian, Dazhen Lin, Zhiming Luo, Beizhan Wang, Shaozi Li |
Concurr. Comput. Pract. Exp. | 4 |
| 2024 | Mutual learning with reliable pseudo label for semi-supervised medical image segmentation
Jiawei Su, Zhiming Luo, Sheng Lian, Dazhen Lin, Shaozi Li |
Medical Image Anal. | 2 |
| 2024 | Reconstruct incomplete relation for incomplete modality brain tumor segmentation
Jiawei Su, Zhiming Luo, Chengji Wang, Sheng Lian, Xuejuan Lin, Shaozi Li |
Neural Networks | 2 |
| 2024 | Unsupervised multi-branch network with high-frequency enhancement for image dehazing
Zhiming Luo, Bo Du 0001, Laibin Chang, Jun Wan 0005 |
Pattern Recognit. | 2 |
| 2024 | Bridge Gap in Pixel and Feature Level for Cross-Modality Person Re-IdentificationabstractVisible thermal person re-identification (VT-ReID) plays a vital role in intelligent surveillance systems, particularly in weak lighting environments. VT-ReID faces substantial challenges, including the cross-modality gap and intra-class variations. Existing methods address these challenges through either pixel-level image translation techniques or feature-level metric learning techniques. However, the former approaches require additional computational costs and often generate noisy images, making model training challenging. The latter methods focus on constraining the relations between individual instances or class centers, while often ignoring joint consideration of the relationship between the two aspects. In addition, these works do not fully investigate the mutual benefits at both pixel-level and feature-level. To address these limitations, we propose a unified Dual-level Smooth Gap (DSG) learning framework that simultaneously smooths the cross-modality gap at the pixel and feature levels. Specifically, on the one hand, we develop a parameter-free Class-aware Modality Mix (CMM) to smooth the cross-modality gap at the pixel level. CMM can capture and explore internal information between the two modalities by mixing images from different modalities belonging to the same class. On the other hand, we devise an efficient Center-guided Metric Learning (CML) to reduce the inter-modality discrepancy and intra-class variations at the feature level. CML enhances model discrimination and generalization by enforcing constraints on both class centers and instances. Experiments on two benchmark datasets demonstrate the mutual benefits of our proposed and show the superior performance of our method over state-of-the-art methods. Yongguo Ling, Zhun Zhong, Zhiming Luo, Shaozi Li, Nicu Sebe |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2024 | Boosting Adversarial Transferability via Logits Mixup With Dominant Decomposed FeatureabstractRecent research has shown that adversarial samples are highly transferable and can be used to attack other unknown black-box Deep Neural Networks (DNNs). To improve the transferability of adversarial samples, several feature-based adversarial attack methods have been proposed to disrupt neuron activation in the middle layers. However, current state-of-the-art feature-based attack methods typically require additional computation costs for estimating the importance of neurons. To address this challenge, we propose a Singular Value Decomposition (SVD)-based feature-level attack method. Our approach is inspired by the discovery that eigenvectors associated with the larger singular values decomposed from the middle layer features exhibit superior generalization and attention properties. Specifically, we conduct the attack by retaining the dominant decomposed feature that corresponds to the largest singular value (i.e., Rank-1 decomposed feature) for computing the output logits before the final softmax. These logits are later integrated with the original logits to optimize adversarial examples. Our extensive experimental results verify the effectiveness of our proposed method, which can be easily integrated into various baselines to significantly enhance the transferability of adversarial samples for disturbing normally trained CNNs and advanced defense strategies. The source code is available at Link. Juanjuan Weng, Zhiming Luo, Shaozi Li, Dazhen Lin, Zhun Zhong |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2023 | Cross-Modality Earth Mover's Distance for Visible Thermal Person Re-identificationabstractVisible thermal person re-identification (VT-ReID) suffers from inter-modality discrepancy and intra-identity variations. Distribution alignment is a popular solution for VT-ReID, however, it is usually restricted to the influence of the intra-identity variations. In this paper, we propose the Cross-Modality Earth Mover's Distance (CM-EMD) that can alleviate the impact of the intra-identity variations during modality alignment. CM-EMD selects an optimal transport strategy and assigns high weights to pairs that have a smaller intra-identity variation. In this manner, the model will focus on reducing the inter-modality discrepancy while paying less attention to intra-identity variations, leading to a more effective modality alignment. Moreover, we introduce two techniques to improve the advantage of CM-EMD. First, Cross-Modality Discrimination Learning (CM-DL) is designed to overcome the discrimination degradation problem caused by modality alignment. By reducing the ratio between intra-identity and inter-identity variances, CM-DL leads the model to learn more discriminative representations. Second, we construct the Multi-Granularity Structure (MGS), enabling us to align modalities from both coarse- and fine-grained levels with the proposed CM-EMD. Extensive experiments show the benefits of the proposed CM-EMD and its auxiliary techniques (CM-DL and MGS). Our method achieves state-of-the-art performance on two VT-ReID benchmarks. Yongguo Ling, Zhun Zhong, Zhiming Luo, Fengxiang Yang, Donglin Cao, Yaojin Lin, Shaozi Li, Nicu Sebe |
AAAI | 3 |
| 2023 | Exploring Non-target Knowledge for Improving Ensemble Universal Adversarial AttacksabstractThe ensemble attack with average weights can be leveraged for increasing the transferability of universal adversarial perturbation (UAP) by training with multiple Convolutional Neural Networks (CNNs). However, after analyzing the Pearson Correlation Coefficients (PCCs) between the ensemble logits and individual logits of the crafted UAP trained by the ensemble attack, we find that one CNN plays a dominant role during the optimization. Consequently, this average weighted strategy will weaken the contributions of other CNNs and thus limit the transferability for other black-box CNNs. To deal with this bias issue, the primary attempt is to leverage the Kullback–Leibler (KL) divergence loss to encourage the joint contribution from different CNNs, which is still insufficient. After decoupling the KL loss into a target-class part and a non-target-class part, the main issue lies in that the non-target knowledge will be significantly suppressed due to the increasing logit of the target class. In this study, we simply adopt a KL loss that only considers the non-target classes for addressing the dominant bias issue. Besides, to further boost the transferability, we incorporate the min-max learning framework to self-adjust the ensemble weights for each CNN. Experiments results validate that considering the non-target KL loss can achieve superior transferability than the original KL loss by a large margin, and the min-max training can provide a mutual benefit in adversarial ensemble attacks. The source code is available at: https://github.com/WJJLL/ND-MM. Juanjuan Weng, Zhiming Luo, Zhun Zhong, Dazhen Lin, Shaozi Li |
AAAI | 2 |
| 2023 | Frequency-Aware Attentional Feature Fusion for Deepfake DetectionabstractVarious face manipulation techniques develop rapidly and can easily generate high-quality fake images or videos, posing significant ethical concerns when used for malicious purposes. Although recent works achieve significant performance in deepfake detection, they still suffer from overfitting issues. To deal with this problem, we propose a novel framework to aggregate diverse information for deepfake detection from both RGB and frequency. Specially, we first introduce a channel attention module to assemble local and global contexts to overcome the potential semantic inconsistency on local artifacts and global features. Then we design a spatial-frequency feature fusion module to fuse the RGB-frequency information comprehensively. Moreover, a variant attention module is further proposed to improve feature discrimination. Extensive experiments demonstrate that our method maintains comparable performance in intra-dataset and cross-dataset evaluation. Zhiming Luo, Guimin Shi, Shaozi Li |
ICASSP | 2 |
| 2023 | Boundary Difference over Union Loss for Medical Image Segmentation
Zhiming Luo, Shaozi Li |
MICCAI (4) | 2 |
| 2023 | TPNet: Enhancing Weakly Supervised Polyp Frame Detection with Temporal Encoder and Prototype-Based Memory Bank
Jianzhe Gao, Zhiming Luo, Shaozi Li |
PRCV (12) | 2 |
| 2023 | Towards Robust Person Re-Identification by Defending Against Universal AttackersabstractRecent studies show that deep person re-identification (re-ID) models are vulnerable to adversarial examples, so it is critical to improving the robustness of re-ID models against attacks. To achieve this goal, we explore the strengths and weaknesses of existing re-ID models, i.e., designing learning-based attacks and training robust models by defending against the learned attacks. The contributions of this paper are three-fold: First, we build a holistic attack-defense framework to study the relationship between the attack and defense for person re-ID. Second, we introduce a combinatorial adversarial attack that is adaptive to unseen domains and unseen model types. It consists of distortions in pixel and color space (i.e., mimicking camera shifts). Third, we propose a novel virtual-guided meta-learning algorithm for our attack-defense system. We leverage a virtual dataset to conduct experiments under our meta-learning framework, which can explore the cross-domain constraints for enhancing the generalization of the attack and the robustness of the re-ID model. Comprehensive experiments on three large-scale re-ID benchmarks demonstrate that: 1) Our combinatorial attack is effective and highly universal in cross-model and cross-dataset scenarios; 2) Our meta-learning algorithm can be readily applied to different attack and defense approaches, which can reach consistent improvement; 3) The defense model trained on the learning-to-learn framework is robust to recent SOTA attacks that are not even used during training. Fengxiang Yang, Juanjuan Weng, Zhun Zhong, Hong Liu 0009, Zheng Wang 0007, Zhiming Luo, Donglin Cao, Shaozi Li, Shin'ichi Satoh 0001, Nicu Sebe |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2023 | Dual-Stream Transformer With Distribution Alignment for Visible-Infrared Person Re-IdentificationabstractVisible-infrared person re-identification(VI-ReID) aims to match the person images captured by visible and infrared cameras and suffers from severe cross-modality discrepancy and intra-modality variations. Existing approaches mainly use convolution neural network (CNN)-based architectures to extract pedestrian features, which fail to capture the long-range dependencies within an image. In addition, previous works usually attempt to bridge the modality gap by using adversarial learning to generate style-consistent images or designing different feature-level metric learning constraints. However, few works consider the cross-modality disparity from the perspective of assessing overall distance distribution discrepancy. To address these problems, we design a pure Transformer-based Visible-Infrared (TransVI) network with a conventional two-stream structure, which can explicitly capture modality-specific representations and learn multi-modality sharable knowledge. TransVI can efficiently address the lack of global dependency in CNN-based architectures due to the multi-head self-attention modules in the transformer, which allows us to capture the long-range dependencies of pedestrian images. Furthermore, we introduce the Cross-Modality Dissimilarity-based Maximum Mean Discrepancy (CMD-MMD) constraint to handle the cross-modality discrepancy at the distance distribution level. Specifically, CMD-MMD leverages intra-modality distribution separability to guide inter-modality distribution separability learning, aligning pair-wise distance distributions of intra- and inter-modality for within-class and between-class, respectively. In this way, the distance distributions of intra- and inter-modality become more similar, significantly mitigating the cross-modality discrepancy and learning more modality invariant representations. Extensive experimental results on two public VI-ReID datasets confirm that our proposed framework can achieve state-of-the-art performance. Zehua Chai, Yongguo Ling, Zhiming Luo, Dazhen Lin, Min Jiang 0005, Shaozi Li |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2023 | Partial Siamese With Multiscale Bi-Codec Networks for Remote Sensing Image Haze RemovalabstractRecently, the U-Shaped networks has been widely explored in remote sensing image dehazing and obtained promising performance. However, most of the existing dehazing methods based on U-Shaped framework lack the reconstruction constraints of haze areas, which is particularly important to restore haze-free images. Moreover, their encoding and decoding layers cannot effectively fuse multi-scale features, resulting in deviations in the color and texture of the dehazing image. To address these issues, in this paper, we propose a Partial Siamese with Multiscale Bi-codec Dehazing Network (PSMB-Net) which is mainly composed of a Partial Siamese Framework (PSF) and a Multiscale Bi-codec Information Fusion (MBIF) module. Specifically, the PSF is proposed to create dehazing prior information to guide the network to build Siamese constraints and achieve improved dehazing results. Furthermore, we design a MBIF module which can enhance feature extraction, and the multi-scale information is used to improve the reconstruction ability of the network for the color and texture of the dehazing image. Experimental results on challenging benchmark datasets demonstrate the superiority of our PSMB-Net over state-of-the-art image dehazing methods. The source code is available at https://github.com/thislzm/PSMB-Net. Zhiming Luo, Bo Du 0001, Wen Yang 0001, Jun Wan 0005, Lefei Zhang |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2023 | Logit Margin Matters: Improving Transferable Targeted Adversarial Attack by Logit CalibrationabstractPrevious works have extensively studied the transferability of adversarial samples in untargeted black-box scenarios. However, it still remains challenging to craft targeted adversarial examples with higher transferability than non-targeted ones. Recent studies reveal that the traditional Cross-Entropy (CE) loss function is insufficient to learn transferable targeted adversarial examples due to the issue of vanishing gradient. In this work, we provide a comprehensive investigation of the CE loss function and find that the logit margin between the targeted and untargeted classes will quickly obtain saturation in CE, which largely limits the transferability. Therefore, in this paper, we devote to the goal of continually increasing the logit margin along the optimization to deal with the saturation issue and propose two simple and effective logit calibration methods, which are achieved by downscaling the logits with a temperature factor and an adaptive margin, respectively. Both of them can effectively encourage optimization to produce a larger logit margin and lead to higher transferability. Besides, we show that minimizing the cosine distance between the adversarial examples and the classifier weights of the target class can further improve the transferability, which is benefited from downscaling logits via L2-normalization. Experiments conducted on the ImageNet dataset validate the effectiveness of the proposed methods, which outperform the state-of-the-art methods in black-box targeted attacks. The source code is available at Link. Juanjuan Weng, Zhiming Luo, Shaozi Li, Nicu Sebe, Zhun Zhong |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2022 | Generalized Person Re-identification by Locating and Eliminating Domain-Sensitive Features
Fengxiang Yang, Zhiming Luo, Shaozi Li |
ACCV (6) | 3 |
| 2022 | Symmetrical Supervision with Transformer for Few-shot Medical Image SegmentationabstractFew-shot learning can potentially learn the target knowledge in extremely few data regimes. Existing few-shot medical image segmentation methods fail to consider the global anatomy correlation between the support and query sets. They generally adopt a weak one-way information transmission that can not fully explore the knowledge to segment query data. To address this problem, we propose a novel Symmetrical Supervision network based on traditional two-branch methods. We raise two main contributions: (1) The Symmetrical Supervision Mechanism is leveraged to strengthen the supervision of network training; (2) A transformer-based Global Feature Alignment module is introduced to increase the global consistency between the two branches. Experimental results on two challenging datasets (abdominal segmentation dataset CHAOS and cardiac segmentation dataset MS-CMRSeg) show a remarkable performance compared to other comparing methods. Yao Niu, Zhiming Luo, Sheng Lian, Lei Li 0048, Shaozi Li, Haixin Song |
BIBM | 2 |
| 2022 | Dynamic Selection Network For Rgb-D Salient Object DetectionabstractExisting RGB-D salient object detection (SOD) methods usually use elaborate fusion modules for exploring cross-modal information, which is computationally expensive and ignores the noise depth information. To deal with this issue, we propose a dynamic selection network (DSNet) for RGB-D salient object detection. Specifically, a cross-modal combination module (CCM) is proposed to fuse two modalities with a light computation. Then a dynamic selection module (DSM) adaptively learns the model parameter for the decoding based on the fused features. Furthermore, skip connection is used for hierarchical features combination between encoder and decoder. Experiments on four popular datasets demonstrate our model outperforms other state-of-the-art methods. Jinlin Zhou, Zhiming Luo, Shaozi Li |
ICIP | 2 |
| 2022 | Consistent response for automated multilabel thoracic disease classificationabstractSummary While recent studies on automated multilabel chest X‐ray (CXR) images classification have shown remarkable progress in leveraging complicated network and attention mechanisms, the automated detection on chest radiographs is still challenging because the pathological patterns are usually highly diverse in their sizes and locations. The CNN model will suffer from the complicated background and high diversity of diseases, which reduce the generalization and performance of the model. To solve these problems, we propose a dual‐distribution consistency (DDC) model, which increases the consistency from two aspects, that is, feature‐level and label‐level. This model integrates two novel loss functions: multilabel response consistency (MRC) loss and distribution consistency (DC) loss. Specifically, we use the original image and its transformed image as inputs to imitate different views of CXR images. The MRC loss encourages the multilabel‐wise attention maps to be consistent between the original CXR image and its transformed counterpart. And the DC loss can force their output probability distributions to be uniform. In this manner, we can make sure that the model can learn discriminative features by using a different view of CXR images. Experiments conducted on the ChestX‐ray14 dataset show the effectiveness of the proposed method. Jiawei Su, Zhiming Luo, Shaozi Li |
Concurr. Comput. Pract. Exp. | 2 |
| 2022 | Stock movement prediction via gated recurrent unit network based on reinforcement learning with incorporated attention mechanisms
Hongfeng Xu, Lei Chai, Zhiming Luo, Shaozi Li |
Neurocomputing | 3 |
| 2022 | Improving embedding learning by virtual attribute decoupling for text-based person search
Chengji Wang, Zhiming Luo, Yaojin Lin, Shaozi Li |
Neural Comput. Appl. | 2 |
| 2022 | Source-Free Open Compound Domain Adaptation in Semantic SegmentationabstractIn this work, we introduce a new concept, named source-free open compound domain adaptation (SF-OCDA), and study it in semantic segmentation. SF-OCDA is more challenging than the traditional domain adaptation but it is more practical. It jointly considers (1) the issues of data privacy and data storage and (2) the scenario of multiple target domains and unseen open domains. In SF-OCDA, only the source pre-trained model and the target data are available to learn the target model. The model is evaluated on the samples from the target and unseen open domains. To solve this problem, we present an effective framework by separating the training process into two stages: (1) pre-training a generalized source model and (2) adapting a target model with self-supervised learning. In our framework, we propose the Cross-Patch Style Swap (CPSS) to diversify samples with various patch styles in the feature-level, which can benefit the training of both stages. First, CPSS can significantly improve the generalization ability of the source model, providing more accurate pseudo-labels for the latter stage. Second, CPSS can reduce the influence of noisy pseudo-labels and also avoid the model overfitting to the target domain during self-supervised learning, consistently boosting the performance on the target and open domains. Experiments demonstrate that our method produces state-of-the-art results on the C-Driving dataset. Furthermore, our model also achieves the leading performance on CityScapes for domain generalization. Zhun Zhong, Zhiming Luo, Gim Hee Lee, Nicu Sebe |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2022 | Joint Representation Learning and Keypoint Detection for Cross-View Geo-LocalizationabstractIn this paper, we study the cross-view geo-localization problem to match images from different viewpoints. The key motivation underpinning this task is to learn a discriminative viewpoint-invariant visual representation. Inspired by the human visual system for mining local patterns, we propose a new framework called RK-Net to jointly learn the discriminative Representation and detect salient Keypoints with a single Network. Specifically, we introduce a Unit Subtraction Attention Module (USAM) that can automatically discover representative keypoints from feature maps and draw attention to the salient regions. USAM contains very few learning parameters but yields significant performance improvement and can be easily plugged into different networks. We demonstrate through extensive experiments that (1) by incorporating USAM, RK-Net facilitates end-to-end joint learning without the prerequisite of extra annotations. Representation learning and keypoint detection are two highly-related tasks. Representation learning aids keypoint detection. Keypoint detection, in turn, enriches the model capability against large appearance changes caused by viewpoint variants. (2) USAM is easy to implement and can be integrated with existing methods, further improving the state-of-the-art performance. We achieve competitive geo-localization accuracy on three challenging datasets, i. e., University-1652, CVUSA and CVACT. Our code is available at https://github.com/AggMan96/RK-Net. Jinliang Lin, Zhedong Zheng, Zhun Zhong, Zhiming Luo, Shaozi Li, Yi Yang 0001, Nicu Sebe |
IEEE Trans. Image Process. | 4 |
| 2021 | Learning to Attack Real-World Models for Person Re-identification via Virtual-Guided Meta-LearningabstractRecent advances in person re-identification (re-ID) have led to impressive retrieval accuracy. However, existing re-ID models are challenged by the adversarial examples crafted by adding quasi-imperceptible perturbations. Moreover, re-ID systems face the domain shift issue that training and testing domains are not consistent. In this study, we argue that learning powerful attackers with high universality that works well on unseen domains is an important step in promoting the robustness of re-ID systems. Therefore, we introduce a novel universal attack algorithm called ``MetaAttack'' for person re-ID. MetaAttack can mislead re-ID models on unseen domains by a universal adversarial perturbation. Specifically, to capture common patterns across different domains, we propose a meta-learning scheme to seek the universal perturbation via the gradient interaction between meta-train and meta-test formed by two datasets. We also take advantage of a virtual dataset (PersonX), instead of real ones, to conduct meta-test. This scheme not only enables us to learn with more comprehensive variation factors but also mitigates the negative effects caused by biased factors of real datasets. Experiments on three large-scale re-ID datasets demonstrate the effectiveness of our method in attacking re-ID models on unseen domains. Our final visualization results reveal some new properties of existing re-ID systems, which can guide us in designing a more robust re-ID model. Code and supplemental material are available at \url{https://github.com/FlyingRoastDuck/MetaAttack_AAAI21}. Fengxiang Yang, Zhun Zhong, Hong Liu 0009, Zheng Wang 0007, Zhiming Luo, Shaozi Li, Nicu Sebe, Shin'ichi Satoh 0001 |
AAAI | 5 |
| 2021 | Joint Noise-Tolerant Learning and Meta Camera Shift Adaptation for Unsupervised Person Re-IdentificationabstractThis paper considers the problem of unsupervised person re-identification (re-ID), which aims to learn discriminative models with unlabeled data. One popular method is to obtain pseudo-label by clustering and use them to optimize the model. Although this kind of approach has shown promising accuracy, it is hampered by 1) noisy labels produced by clustering and 2) feature variations caused by camera shift. The former will lead to incorrect optimization and thus hinders the model accuracy. The latter will result in assigning the intra-class samples of different cameras to different pseudo-label, making the model sensitive to camera variations. In this paper, we propose a unified framework to solve both problems. Concretely, we propose a Dynamic and Symmetric Cross-Entropy loss (DSCE) to deal with noisy samples and a camera-aware meta-learning algorithm (MetaCam) to adapt camera shift. DSCE can alleviate the negative effects of noisy samples and accommodate the change of clusters after each clustering step. MetaCam simulates cross-camera constraint by splitting the training data into meta-train and meta-test based on camera IDs. With the interacted gradient from meta-train and meta-test, the model is enforced to learn camera-invariant features. Extensive experiments on three re-ID benchmarks show the effectiveness and the complementary of the proposed DSCE and MetaCam. Our method outperforms the state-of-the-art methods on both fully unsupervised re-ID and unsupervised domain adaptive re-ID. Fengxiang Yang, Zhun Zhong, Zhiming Luo, Yuanzheng Cai, Yaojin Lin, Shaozi Li, Nicu Sebe |
CVPR | 3 |
| 2021 | Learning to Generalize Unseen Domains via Memory-based Multi-Source Meta-Learning for Person Re-IdentificationabstractRecent advances in person re-identification (ReID) obtain impressive accuracy in the supervised and unsupervised learning settings. However, most of the existing methods need to train a new model for a new domain by accessing data. Due to public privacy, the new domain data are not always accessible, leading to a limited applicability of these methods. In this paper, we study the problem of multi-source domain generalization in ReID, which aims to learn a model that can perform well on unseen domains with only several labeled source domains. To address this problem, we propose the Memory-based Multi-Source Meta-Learning (M3L) framework to train a generalizable model for unseen domains. Specifically, a meta-learning strategy is introduced to simulate the train-test process of domain generalization for learning more generalizable models. To overcome the unstable meta-optimization caused by the parametric classifier, we propose a memory-based identification loss that is non-parametric and harmonizes with meta-learning. We also present a meta batch normalization layer (MetaBN) to diversify meta-test features, further establishing the advantage of meta-learning. Experiments demonstrate that our M3L can effectively enhance the generalization ability of the model for unseen domains and can outperform the state-of-the-art methods on four large-scale ReID datasets. Zhun Zhong, Fengxiang Yang, Zhiming Luo, Yaojin Lin, Shaozi Li, Nicu Sebe |
CVPR | 4 |
| 2021 | Neighborhood Contrastive Learning for Novel Class DiscoveryabstractIn this paper, we address Novel Class Discovery (NCD), the task of unveiling new classes in a set of unlabeled samples given a labeled dataset with known classes. We exploit the peculiarities of NCD to build a new framework, named Neighborhood Contrastive Learning (NCL), to learn discriminative representations that are important to clustering performance. Our contribution is twofold. First, we find that a feature extractor trained on the labeled set generates representations in which a generic query sample and its neighbors are likely to share the same class. We exploit this observation to retrieve and aggregate pseudo-positive pairs with contrastive learning, thus encouraging the model to learn more discriminative representations. Second, we notice that most of the instances are easily discriminated by the network, contributing less to the contrastive loss. To overcome this issue, we propose to generate hard negatives by mixing labeled and unlabeled samples in the feature space. We experimentally demonstrate that these two ingredients significantly contribute to clustering performance and lead our model to outperform state-of-the-art methods by a large margin (e.g., clustering accuracy +13% on CIFAR-100 and +8% on ImageNet). Zhun Zhong, Enrico Fini, Subhankar Roy, Zhiming Luo, Elisa Ricci 0001, Nicu Sebe |
CVPR | 4 |
| 2021 | OpenMix: Reviving Known Knowledge for Discovering Novel Visual Categories in an Open WorldabstractIn this paper, we tackle the problem of discovering new classes in unlabeled visual data given labeled data from disjoint classes. Existing methods typically first pre-train a model with labeled data, and then identify new classes in unlabeled data via unsupervised clustering. However, the labeled data that provide essential knowledge are often underexplored in the second step. The challenge is that the labeled and unlabeled examples are from non-overlapping classes, which makes it difficult to build a learning relationship between them. In this work, we introduce Open-Mix to mix the unlabeled examples from an open set and the labeled examples from known classes, where their non-overlapping labels and pseudo-labels are simultaneously mixed into a joint label distribution. OpenMix dynamically compounds examples in two ways. First, we produce mixed training images by incorporating labeled examples with unlabeled examples. With the benefit of unique prior knowledge in novel class discovery, the generated pseudo-labels will be more credible than the original unlabeled predictions. As a result, OpenMix helps preventing the model from overfitting on unlabeled samples that may be assigned with wrong pseudo-labels. Second, the first way encourages the unlabeled examples with high class-probabilities to have considerable accuracy. We introduce these examples as reliable anchors and further integrate them with un-labeled samples. This enables us to generate more combinations in unlabeled examples and exploit finer object relations among the new classes. Experiments on three classification datasets demonstrate the effectiveness of the proposed OpenMix, which is superior to state-of-the-art methods in novel class discovery. Zhun Zhong, Linchao Zhu, Zhiming Luo, Shaozi Li, Yi Yang 0001, Nicu Sebe |
CVPR | 3 |
| 2021 | A Multi-Constraint Similarity Learning with Adaptive Weighting for Visible-Thermal Person Re-IdentificationabstractThe challenges of visible-thermal person re-identification (VT-ReID) lies in the inter-modality discrepancy and the intra-modality variations. An appropriate metric learning plays a crucial role in optimizing the feature similarity between the two modalities. However, most existing metric learning-based methods mainly constrain the similarity between individual instances or class centers, which are inadequate to explore the rich data relationships in the cross-modality data. Besides, most of these methods fail to consider the importance of different pairs, incurring an inefficiency and ineffectiveness of optimization. To address these issues, we propose a Multi-Constraint (MC) similarity learning method that jointly considers the cross-modality relationships from three different aspects, i.e., Instance-to-Instance (I2I), Center-to-Instance (C2I), and Center-to-Center (C2C). Moreover, we devise an Adaptive Weighting Loss (AWL) function to implement the MC efficiently. In the AWL, we first use an adaptive margin pair mining to select informative pairs and then adaptively adjust weights of mined pairs based on their similarity. Finally, the mined and weighted pairs are used for the metric learning. Extensive experiments on two benchmark datasets demonstrate the superior performance of the proposed over the state-of-the-art methods. Yongguo Ling, Zhiming Luo, Yaojin Lin, Shaozi Li |
IJCAI | 2 |
| 2021 | Text-based Person Search via Multi-Granularity Embedding LearningabstractMost existing text-based person search methods highly depend on exploring the corresponding relations between the regions of the image and the words in the sentence. However, these methods correlated image regions and words in the same semantic granularity. It 1) results in irrelevant corresponding relations between image and text, 2) causes an ambiguity embedding problem. In this study, we propose a novel multi-granularity embedding learning model for text-based person search. It generates multi-granularity embeddings of partial person bodies in a coarse-to-fine manner by revisiting the person image at different spatial scales. Specifically, we distill the partial knowledge from image scrips to guide the model to select the semantically relevant words from the text description. It can learn discriminative and modality-invariant visual-textual embeddings. In addition, we integrate the partial embeddings at each granularity and perform multi-granularity image-text matching. Extensive experiments validate the effectiveness of our method, which can achieve new state-of-the-art performance by the learned discriminative partial embeddings. Chengji Wang, Zhiming Luo, Yaojin Lin, Shaozi Li |
IJCAI | 2 |
| 2021 | Learning Consistency- and Discrepancy-Context for 2D Organ Segmentation
Lei Li 0048, Sheng Lian, Zhiming Luo, Shaozi Li, Beizhan Wang, Shuo Li 0001 |
MICCAI (1) | 3 |
| 2021 | A weakly supervised tooth-mark and crack detection method in tongue imageabstractAbstract Tongue diagnosis is one of the primary clinical diagnostic methods in Traditional Chinese Medicine. Recognizing the tooth‐marked tongue and the crackled tongue plays an essential role in evaluating the status of patients. Previous methods mainly focus on identifying whether a tongue image is a tooth‐marked tongue (cracked tongue) or not, while cannot provide more details. In this study, we propose a weakly supervised method for training the tooth‐mark and crack detection model by leveraging fully bounding‐box level annotated and coarse image‐level annotated tongue images. The proposed model is extended from the YOLO object detection model, and we add several classification branches for recognizing the tooth‐marked tongue and cracked tongue. The classification branch aims to predict the coarse label for both coarse‐labeled data and fully annotated data. The detection branch is used to locate the position of tooth marks and cracks from the fully annotated data. Finally, we utilize a multitask loss function for training the model. Experimental results on a challenging tongue image dataset demonstrate the effectiveness of our proposed weakly supervised method. Hui Weng, Lei Li 0048, Huangwei Lei, Zhiming Luo, Candong Li, Shaozi Li |
Concurr. Comput. Pract. Exp. | 4 |
| 2021 | High-Order-Interaction for weakly supervised Fine-Grained Visual Categorization
Nanyu Li, Zhiming Luo, Zhun Zhong, Shaozi Li |
Neurocomputing | 3 |
| 2021 | Divide-and-Merge the embedding space for cross-modality person search
Chengji Wang, Zhiming Luo, Zhun Zhong, Shaozi Li |
Neurocomputing | 2 |
| 2021 | APRIL: Anatomical prior-guided reinforcement learning for accurate carotid lumen diameter and intima-media thickness measurement
Sheng Lian, Zhiming Luo, Shaozi Li, Shuo Li 0001 |
Medical Image Anal. | 2 |
| 2021 | SAFD: single shot anchor free face detector
Chengji Wang, Zhiming Luo, Zhun Zhong, Shaozi Li |
Multim. Tools Appl. | 2 |
| 2021 | Learning to Adapt Invariance in Memory for Person Re-IdentificationabstractThis work considers the problem of unsupervised domain adaptation in person re-identification (re-ID), which aims to transfer knowledge from the source domain to the target domain. Existing methods are primary to reduce the inter-domain shift between the domains, which however usually overlook the relations among target samples. This paper investigates into the intra-domain variations of the target domain and proposes a novel adaptation framework w.r.t three types of underlying invariance, i.e., Exemplar-Invariance, Camera-Invariance, and Neighborhood-Invariance. Specifically, an exemplar memory is introduced to store features of samples, which can effectively and efficiently enforce the invariance constraints over the global dataset. We further present the Graph-based Positive Prediction (GPP) method to explore reliable neighbors for the target domain, which is built upon the memory and is trained on the source samples. Experiments demonstrate that 1) the three invariance properties are complementary and indispensable for effective domain adaptation, 2) the memory plays a key role in implementing invariance learning and improves the performance with limited extra computation cost, 3) GPP can facilitate the invariance learning and thus significantly improves the results, and 4) our approach produces new state-of-the-art adaptation accuracy on three re-ID large-scale benchmarks. Zhun Zhong, Liang Zheng 0001, Zhiming Luo, Shaozi Li, Yi Yang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2021 | A Global and Local Enhanced Residual U-Net for Accurate Retinal Vessel SegmentationabstractRetinal vessel segmentation is a critical procedure towards the accurate visualization, diagnosis, early treatment, and surgery planning of ocular diseases. Recent deep learning-based approaches have achieved impressive performance in retinal vessel segmentation. However, they usually apply global image pre-processing and take the whole retinal images as input during network training, which have two drawbacks for accurate retinal vessel segmentation. First, these methods lack the utilization of the local patch information. Second, they overlook the geometric constraint that retina only occurs in a specific area within the whole image or the extracted patch. As a consequence, these global-based methods suffer in handling details, such as recognizing the small thin vessels, discriminating the optic disk, etc. To address these drawbacks, this study proposes a Global and Local enhanced residual U-nEt (GLUE) for accurate retinal vessel segmentation, which benefits from both the globally and locally enhanced information inside the retinal region. Experimental results on two benchmark datasets demonstrate the effectiveness of the proposed method, which consistently improves the segmentation accuracy over a conventional U-Net and achieves competitive performance compared to the state-of-the-art. Sheng Lian, Lei Li 0048, Guiren Lian, Zhiming Luo, Shaozi Li |
IEEE ACM Trans. Comput. Biol. Bioinform. | 5 |
| 2020 | Asymmetric Co-Teaching for Unsupervised Cross-Domain Person Re-IdentificationabstractPerson re-identification (re-ID), is a challenging task due to the high variance within identity samples and imaging conditions. Although recent advances in deep learning have achieved remarkable accuracy in settled scenes, i.e., source domain, few works can generalize well on the unseen target domain. One popular solution is assigning unlabeled target images with pseudo labels by clustering, and then retraining the model. However, clustering methods tend to introduce noisy labels and discard low confidence samples as outliers, which may hinder the retraining process and thus limit the generalization ability. In this study, we argue that by explicitly adding a sample filtering procedure after the clustering, the mined examples can be much more efficiently used. To this end, we design an asymmetric co-teaching framework, which resists noisy labels by cooperating two models to select data with possibly clean labels for each other. Meanwhile, one of the models receives samples as pure as possible, while the other takes in samples as diverse as possible. This procedure encourages that the selected training samples can be both clean and miscellaneous, and that the two models can promote each other iteratively. Extensive experiments show that the proposed framework can consistently benefit most clustering based methods, and boost the state-of-the-art adaptation accuracy. Our code is available at https://github.com/FlyingRoastDuck/ACT_AAAI20. Fengxiang Yang, Ke Li 0015, Zhun Zhong, Zhiming Luo, Xing Sun 0001, Hao Cheng 0012, Feiyue Huang, Rongrong Ji, Shaozi Li |
AAAI | 4 |
| 2020 | Class-Aware Modality Mix and Center-Guided Metric Learning for Visible-Thermal Person Re-IdentificationabstractVisible thermal person re-identification (VT-REID) is an important and challenging task in that 1) weak lighting environments are inevitably encountered in real-world settings and 2) the inter-modality discrepancy is serious. Most existing methods either aim at reducing the cross-modality gap in pixel- and feature-level or optimizing cross-modality network by metric learning techniques. However, few works have jointly considered these two aspects and studied their mutual benefits. In this paper, we design a novel framework to jointly bridge the modality gap in pixel- and feature-level without additional parameters, as well as reduce the inter- and intra-modalities variations by a center-guided metric learning constraint. Specifically, we introduce the Class-aware Modality Mix (CMM) to generate internal information of the two modalities for reducing the modality gap in pixel-level. In addition, we exploit the KL-divergence to further align modality distributions on feature-level. On the other hand, we propose an efficient Center-guided Metric Learning (CML) method for decreasing the discrepancy within the inter- and intra-modalities, by enforcing constraints on class centers and instances. Extensive experiments on two datasets show the mutual advantage of the proposed components and demonstrate the superiority of our method over the state of the art. Yongguo Ling, Zhun Zhong, Zhiming Luo, Paolo Rota, Shaozi Li, Nicu Sebe |
ACM Multimedia | 3 |
| 2020 | A robust interclass and intraclass loss function for deep learning based tongue segmentationabstractSummary The fast and robust segmentation of tongue images is a prerequisite to achieve automatic tongue diagnosis in traditional Chinese medicine. In order to assist tongue diagnosis in real‐life scenarios, an ideal tongue segmentation method would need to obtain the entire tongue body as well as its precise contours. However, the similar appearance among the tongue body, the coating, and the lips hinders the performance for most unsupervised learning methods that primarily utilize low‐level visual features. On the other hand, although the supervised deep convolutional neural networks (DCNNs) that typically depend on the widely used crossentropy loss can achieve better accuracy, they are very prone to segment image to multiple trivially areas. To address both of the above issues, we make an attempt to boost the segmentation performance of DCNNs with a novel auxiliary loss function that seeks to exploit large margin learning for end‐to‐end tongue segmentation models. Specifically, we first propose a loss function that involves interclass and intraclass costs to directly measure the distance among pixels that belong to different connected regions. Then, we explore the potential ability of this loss function as a regularization for different segmentation networks such as those with attention modules and the deeply supervised network. Finally, a theoretical analysis for this learning scheme is presented. Through experiments on challenging datasets, we show that the proposed approach can be easily integrated into state‐of‐the‐art networks to boost their performance at the tongue segmentation task without bells and whistles. Yuanzheng Cai, Tao Wang 0047, Zhiming Luo |
Concurr. Comput. Pract. Exp. | 4 |
| 2020 | An iterative transfer learning framework for cross-domain tongue segmentationabstractSummary Tongue diagnosis is an important clinical examination in Traditional Chinese Medicine. As the first step of the diagnosis, the accuracy of tongue image segmentation directly affects the subsequent diagnosis. Recently, deep learning‐based methods have been applied for tongue image segmentation and achieve promising results. However, these methods usually work well on one dataset and degenerate significantly on different distributed datasets. To deal with this issue, we propose a framework named Iterative cross‐domain tongue segmentation in the study. First, we train a tongue image segmentation U‐Net model on the source dataset. Then, we propose a tongue assessment filter to select satisfying samples based on predictions of the U‐Net model from the target dataset. Following, we fine‐tune the model on the selected samples along with the source domain. Finally, we iterate between the filtering and the fine‐tuning steps until the model is converged. Experimental results on two tongue datasets show that our proposed method can improve the dice score on the target domain from 70.11% to 98.26%, as well as outperform state‐of‐the‐art comparing methods. Lei Li 0048, Zhiming Luo, Yuanzheng Cai, Candong Li, Shaozi Li |
Concurr. Comput. Pract. Exp. | 2 |
| 2020 | Pattern synthesis of thinned multi-input multi-output radar using difference set and differential evolutionabstractSummary To lower the side lobe level of MIMO (Multi‐input Multi‐output) radar, a new method, diffenence set and diffderential evolution (DSDE), was proposed. The newly proposed method can arrange the location of the linear MIMO radar antenna arrays by differential set (DS) and optimize the excitation amplitude of each antenna by Differential Evolution (DE). DS is a kind of analysis calculation method, which has faster calculation speed than intelligent optimization methods and has solid optimization ability in thinned antenna array optimization problems. Besides, DE is one of the best stochastic optimization methods which can keep the diversity and avoid premature of the population. Consequently, the proposed DSDE can achieve high performance in the arrangement of arrays and excitation amplitude optimization. Numerical experiments are conducted to test the performance of the proposed algorithm, and the results showed that the analysis optimization method and stochastic optimization method are suitable for solving the problem of pattern synthesis of MIMO radar and keep the diversity with better convergence. Guimin Shi, Zhiming Luo |
Concurr. Comput. Pract. Exp. | 4 |
| 2020 | SERU: A cascaded SE-ResNeXT U-Net for kidney and tumor segmentationabstractSummary According to statistics, kidney cancer is one of the most deadly cancer. An early and accurate diagnosis can significantly increase the cure rate. Accurate segmentation of kidney tumors in CT images plays an important role in kidney cancer diagnosis. However, it is a challenging task due to many different aspects, such as low contrast, irregular motion, diverse shapes, and sizes. For solving this issue, we proposed a SE‐R esNeXT U ‐Net (SERU) model in this study, which takes the advantages of SE‐Net, ResNeXT and U‐Net. Besides, we implement our model in a coarse‐to‐fine manner to utilize the information of context and key slices from the left and right kidney. We train and test our method on the KiTS19 Challenge. Experimental results demonstrate that our model can achieve promising results. Xiuzhen Xie, Lei Li 0048, Sheng Lian, Shaohao Chen, Zhiming Luo |
Concurr. Comput. Pract. Exp. | 5 |
| 2020 | Research on line overload identification of power system based on improved neural network algorithmabstractSummary Due to the continuous appearance of safety fault accidents in the practice process, operation safety has become the central task of various operation and management tasks of the power grid. Therefore, to establish a line overload identification and data control model for the power system, we first defined the vulnerability of complex power systems based on the analysis of each line and node. For finding the optimal parameters of this model, we proposed an improved optimization strategy by combining the genetic algorithm and BP neural network. To verified the effectiveness of our proposed method, we conducted experiments on a simulation on the IEEE 30‐node power system environment. Experimental results demonstrate that the proposed algorithms can establish an optimized overload identification model with better performance. This study can help to conduct reasonable adjustment when overload happens to the power system, and then reduce similar failure as well as enhance the operation safety. Zhiming Luo, Wangqing Lin, Shaozi Li |
Concurr. Comput. Pract. Exp. | 2 |
| 2020 | Hand gesture recognition based on attentive feature fusionabstractSummary Video‐based hand gesture recognition plays an important role in human‐computer interaction (HCI). Recent advanced methods usually add 3D convolutional neural networks to capture the information from both spatial and temporal dimensions. However, these methods suffer the issue of requiring large‐scale training data and high computational complexity. To address this issue, we proposed an attentive feature fusion framework for efficient hand‐gesture recognition. In our proposed model, we utilize a shallow two‐stream CNNs to capture the low‐level features from the original video frame and its corresponding optical flow. Following, we designed an attentive feature fusion module to selectively combine useful information from the previous two streams based on the attention mechanism. Finally, we obtain a compact embedding of a video by concatenating features from several short segments. To evaluate the effectiveness of our proposed framework, we train and test our method on a large‐scale video‐based hand gesture recognition dataset, Jester. Experimental results demonstrate that our approach obtains very competitive performance on the Jester dataset with a classification accuracy of 95.77%. Zhiming Luo, Huangbin Wu, Shaozi Li |
Concurr. Comput. Pract. Exp. | 2 |
| 2020 | A multi-source heterogeneous data analytic method for future price fluctuation prediction
Lei Chai, Hongfeng Xu, Zhiming Luo, Shaozi Li |
Neurocomputing | 3 |
| 2020 | Stock movement predictive network via incorporative attention mechanisms based on tweet and historical prices
Hongfeng Xu, Lei Chai, Zhiming Luo, Shaozi Li |
Neurocomputing | 3 |
| 2020 | Joint imbalanced classification and feature selection for hospital readmissions
Guodong Du 0002, Jia Zhang 0019, Zhiming Luo, Fenglong Ma, Lei Ma 0010, Shaozi Li |
Knowl. Based Syst. | 3 |
| 2020 | Leveraging Virtual and Real Person for Unsupervised Person Re-IdentificationabstractPerson re-identification (re-ID) is a challenging instance retrieval problem, especially when identity annotations are not available for training. Although modern deep re-ID approaches have achieved great improvement, it is still difficult to optimize the deep re-ID model and learn discriminative person representation without annotations in training data. To address this challenge, this study considers the problem of unsupervised person re-ID and introduces a novel approach to solve this problem by leveraging virtual and real data. Our approach includes two components: virtual person generation and training of the deep re-ID model. For virtual person generation, we learn a person generation model and a camera style transfer model using unlabeled real data to generate virtual persons with different poses and camera styles. The virtual data is formed as labeled training data, enabling subsequent training deep re-ID model in supervision. For training of the deep re-ID model, we divide it into three steps: 1) pre-training a coarse re-ID model by using virtual data; 2) collaborative filtering based positive pair mining from the real data; and 3) fine-tuning of the coarse re-ID model by leveraging the mined positive pairs and virtual data. The final re-ID model is achieved by iterating between step 2 and step 3 until convergence. Extensive experiments demonstrate the effectiveness of our method. Experimental results on two large-scale datasets, Market-1501 and DukeMTMC-reID, show the advantages of our method over state-of-the-art approaches in unsupervised person re-ID. Our code is now available online1. Fengxiang Yang, Zhun Zhong, Zhiming Luo, Sheng Lian, Shaozi Li |
IEEE Trans. Multim. | 3 |
| 2019 | Invariance Matters: Exemplar Memory for Domain Adaptive Person Re-IdentificationabstractThis paper considers the domain adaptive person re-identification (re-ID) problem: learning a re-ID model from a labeled source domain and an unlabeled target domain. Conventional methods are mainly to reduce feature distribution gap between the source and target domains. However, these studies largely neglect the intra-domain variations in the target domain, which contain critical factors influencing the testing performance on the target domain. In this work, we comprehensively investigate into the intra-domain variations of the target domain and propose to generalize the re-ID model w.r.t three types of the underlying invariance, i.e., exemplar-invariance, camera-invariance and neighborhood-invariance. To achieve this goal, an exemplar memory is introduced to store features of the target domain and accommodate the three invariance properties. The memory allows us to enforce the invariance constraints over global training batch without significantly increasing computation cost. Experiment demonstrates that the three invariance properties and the proposed memory are indispensable towards an effective domain adaptation system. Results on three re-ID domains show that our domain adaptation accuracy outperforms the state of the art by a large margin. Code is available at: https://github.com/zhunzhong07/ECN. Zhun Zhong, Liang Zheng 0001, Zhiming Luo, Shaozi Li, Yi Yang 0001 |
CVPR | 3 |
| 2019 | Manifold regularized discriminative feature selection for multi-label learning
Jia Zhang 0019, Zhiming Luo, Candong Li, Changen Zhou, Shaozi Li |
Pattern Recognit. | 2 |
| 2019 | Convolutional Neural Network With Shape Prior Applied to Cardiac MRI SegmentationabstractIn this paper, we present a novel convolutional neural network architecture to segment images from a series of short-axis cardiac magnetic resonance slices (CMRI). The proposed model is an extension of the U-net that embeds a cardiac shape prior and involves a loss function tailored to the cardiac anatomy. Since the shape prior is computed offline only once, the execution of our model is not limited by its calculation. Our system takes as input raw magnetic resonance images, requires no manual preprocessing or image cropping and is trained to segment the endocardium and epicardium of the left ventricle, the endocardium of the right ventricle, as well as the center of the left ventricle. With its multiresolution grid architecture, the network learns both high and low-level features useful to register the shape prior as well as accurately localize the borders of the cardiac regions. Experimental results obtained on the Automatic Cardiac Diagnostic Challenge - Medical Image Computing and Computer Assisted Intervention (ACDC-MICCAI) 2017 dataset show that our model segments multislices CMRI (left and right ventricle contours) in 0.18 s with an average Dice coefficient of [Formula: see text] and an average 3-D Hausdorff distance of [Formula: see text] mm. Clément Zotti, Zhiming Luo, Alain Lalande, Pierre-Marc Jodoin |
IEEE J. Biomed. Health Informatics | 2 |
| 2018 | Anchor Free Network for Multi-Scale Face DetectionabstractAnchor-based deep methods are the most widely used methods for face detection and have reached the state-of-the-art result. Compared with anchor-based methods that estimates the bounding-box rely on some pre-defined anchor boxes, anchor-free methods perform the localization by predicting the offsets of a pixel inside a face to its outside boundaries whose accuracies are much more precise. However, anchor-free methods suffer the drawback of low recall-rate mainly because 1) only using single scale features lead to miss detection of small faces, 2) the highly intra-class imbalance problem among different size faces. In this paper, to address these problems, we propose a unified anchor-free network for detecting multi-scale faces by leveraging the local and global contextual information of multi-layer features. We also utilize a scale aware sampling strategy to mitigate the intra-class imbalance issue which can adaptivity select the positive samples. Furthermore, a revised focal loss function is adopted to deal with the foreground/background imbalance issue. Experimental results on two benchmark datasets demonstrate the effective of our proposed method. Chengji Wang, Zhiming Luo, Sheng Lian, Shaozi Li |
ICPR | 2 |
| 2018 | Attention guided U-Net for accurate iris segmentation
Sheng Lian, Zhiming Luo, Zhun Zhong, Songzhi Su, Shaozi Li |
J. Vis. Commun. Image Represent. | 2 |
| 2018 | Improving deep ensemble vehicle classification by using selected adversarial samples
Wei Liu 0052, Zhiming Luo, Shaozi Li |
Knowl. Based Syst. | 2 |
| 2018 | Traffic Analytics With Low-Frame-Rate VideosabstractIn this paper, we investigate the possibility of monitoring highway traffic based on videos whose frame rate is too low to accurately estimate motion features. The goal of the proposed method is to recognize traffic conditions instead of measuring them, as is usually the case. The main advantage of our approach comes from its ability to process low-frame-rate videos for which motion features cannot be estimated. Our method takes advantage of the highly redundant nature of traffic scenes that are pictured from a top-down perspective showing vehicles on a predominant asphalted road surrounded by background objects. Due to the limited variety of objects pictured in traffic scenes, our method gets to learn features that are specific to such images. With these features, our method is able to segment traffic images, classify traffic scenes, and estimate traffic density without requiring motion features. Different convolutional neural network models are proposed to segment traffic images in three different classes (Road, Car, and Background), classify traffic images into different categories (Empty, Fluid, Heavy, and Jam), and predict traffic density. We also propose a procedure to perform transfer learning of any of these models to new traffic scenes. Zhiming Luo, Pierre-Marc Jodoin, Songzhi Su, Shaozi Li, Hugo Larochelle |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2018 | Spectral-Spatial Residual Network for Hyperspectral Image Classification: A 3-D Deep Learning FrameworkabstractIn this paper, we designed an end-to-end spectral-spatial residual network (SSRN) that takes raw 3-D cubes as input data without feature engineering for hyperspectral image classification. In this network, the spectral and spatial residual blocks consecutively learn discriminative features from abundant spectral signatures and spatial contexts in hyperspectral imagery (HSI). The proposed SSRN is a supervised deep learning framework that alleviates the declining-accuracy phenomenon of other deep learning models. Specifically, the residual blocks connect every other 3-D convolutional layer through identity mapping, which facilitates the backpropagation of gradients. Furthermore, we impose batch normalization on every convolutional layer to regularize the learning process and improve the classification performance of trained models. Quantitative and qualitative results demonstrate that the SSRN achieved the state-of-the-art HSI classification accuracy in agricultural, rural-urban, and urban data sets: Indian Pines, Kennedy Space Center, and University of Pavia. Zilong Zhong, Jonathan Li 0001, Zhiming Luo, Michael A. Chapman |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2018 | MIO-TCD: A New Benchmark Dataset for Vehicle Classification and LocalizationabstractThe ability to train on a large dataset of labeled samples is critical to the success of deep learning in many domains. In this paper, we focus on motor vehicle classification and localization from a single video frame and introduce the "MIOvision Traffic Camera Dataset" (MIO-TCD) in this context. MIO-TCD is the largest dataset for motorized traffic analysis to date. It includes 11 traffic object classes such as cars, trucks, buses, motorcycles, bicycles, pedestrians. It contains 786,702 annotated images acquired at different times of the day and different periods of the year by hundreds of traffic surveillance cameras deployed across Canada and the United States. The dataset consists of two parts: a "localization dataset", containing 137,743 full video frames with bounding boxes around traffic objects, and a "classification dataset", containing 648,959 crops of traffic objects from the 11 classes. We also report results from the 2017 CVPR MIO-TCD Challenge, that leveraged this dataset, and compare them with results for state-of-the-art deep learning architectures. These results demonstrate the viability of deep learning methods for vehicle localization and classification from a single video frame in real-life traffic scenarios. The topperforming methods achieve both accuracy and Kappa score above 96% on the classification dataset and mean-average precision of 77% on the localization dataset. We also identify scenarios in which state-of-the-art methods still fail and we suggest avenues to address these challenges. Both the dataset and detailed results are publicly available on-line [1]. Zhiming Luo, Frederic Branchaud-Charron, Carl Lemaire, Janusz Konrad, Shaozi Li, Akshaya Mishra, Andrew Achkar, Justin A. Eichel, Pierre-Marc Jodoin |
IEEE Trans. Image Process. | 1 |
| 2017 | Non-local Deep Features for Salient Object DetectionabstractSaliency detection aims to highlight the most relevant objects in an image. Methods using conventional models struggle whenever salient objects are pictured on top of a cluttered background while deep neural nets suffer from excess complexity and slow evaluation speeds. In this paper, we propose a simplified convolutional neural network which combines local and global information through a multi-resolution 4×5 grid structure. Instead of enforcing spacial coherence with a CRF or superpixels as is usually the case, we implemented a loss function inspired by the Mumford-Shah functional which penalizes errors on the boundary. We trained our model on the MSRA-B dataset, and tested it on six different saliency benchmark datasets. Results show that our method is on par with the state-of-the-art while reducing computation time by a factor of 18 to 100 times, enabling near real-time, high performance saliency detection. Zhiming Luo, Akshaya Kumar Mishra, Andrew Achkar, Justin A. Eichel, Shaozi Li, Pierre-Marc Jodoin |
CVPR | 1 |
| 2017 | A novel recurrent hybrid network for feature fusion in action recognition
Sheng Yu 0007, Zhiming Luo, Min Huang 0004, Shaozi Li |
J. Vis. Commun. Image Represent. | 4 |
| 2017 | Interactive deep learning method for segmenting moving objects
Yi Wang 0025, Zhiming Luo, Pierre-Marc Jodoin |
Pattern Recognit. Lett. | 2 |
| 2015 | Traffic analysis without motion featuresabstractIn this paper, we investigate the possibility of monitoring traffic without using any motion features. The goal of our system is to process videos with ultra-low frame rate, i.e. videos for which reliable motion features cannot be computed. In this work, we investigate how 2D spatial features combined with a machine learning method can assess traffic conditions such as fluid traffic, dense traffic, and traffic jam. The underlying hypothesis that we ought to validate is that traffic images are heavily characterized by their 2D spatial textures. In that perspective, we tested different 2D texture features and machine learning methods to see how accurate such an approach can be. We also performed a regression on the image descriptor in order to estimate traffic density. Experimental results obtained on the UCSD traffic dataset reveal that our approach generalizes well to various weather and lighting conditions. It even outperforms state-of-the-art traffic analysis methods relying on spatio-temporal features. Zhiming Luo, Pierre-Marc Jodoin, Shaozi Li, Songzhi Su |
ICIP | 1 |
| 2015 | Intelligent stock market instability index: Application to the Korean stock marketabstractIn order to monitor stock market instability in emerging markets, we propose a stock market instability index (SMII) with a corresponding p-value by using a model fitted to a stable period. More precisely, this study considers a random walk model and combines it with a nonparametric model by using Bayesian model averaging. The integrated stock market instability index (iSMII) and its p$-value are derived as a posterior expectation of the two models. In this study, an artificial neural network (ANN) is utilized as a nonparametric model. Young Min Kim 0002, Sung Kwon Han, Tae Yoon Kim, Kyong Joo Oh, Zhiming Luo, Chiho Kim |
Intell. Data Anal. | 5 |