VLDB 2026 Research / reviewers in the wild / expert
LinLin Shen
dblp:88/5607 · also Linlin Shen
· DBLP profile ↗
346ranked-venue papers
15as first author
240since 2021 · last 2027
0000-0003-1420-0815ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 204 · 8 first-author · 139 since 2021Graphics, computer vision, multimedia, augmented reality and games · 168 · 3 first-author · 125 since 2021Applied, interdisciplinary, general and emerging computing · 50 · 6 first-author · 29 since 2021Human-computer interaction and ubiquitous computing · 16 · 13 since 2021Security and privacy · 15 · 13 since 2021Databases, data management, data science and information retrieval · 5 · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2027 | TUS-DET: Open-vocabulary pretraining with standard planes for few-shot thyroid ultrasound lesion detection
Jiansong Zhang 0005, Shunlan Liu, Xiaoling Luo 0001, Guorong Lyu, LinLin Shen |
Expert Syst. Appl. | 6 |
| 2026 | FineXtrol: Controllable Motion Generation via Fine-Grained TextabstractRecent works have sought to enhance the controllability and precision of text-driven motion generation. Some approaches leverage large language models (LLMs) to produce more detailed texts, while others incorporate global 3D coordinate sequences as additional control signals. However, the former often introduces misaligned details and lacks explicit temporal cues, and the latter incurs significant computational cost when converting coordinates to standard motion representations. To address these issues, we propose FineXtrol, a novel control framework for efficient motion generation guided by temporally-aware, precise, user-friendly and fine-grained textual control signals that describe specific body part movements over time. In support of this framework, we design a hierarchical contrastive learning module that encourages the text encoder to produce more discriminative embeddings for our novel control signals, thereby improving motion controllability. Quantitative results show that FineXtrol achieves strong performance in controllable motion generation, while qualitative analysis demonstrates its flexibility in directing specific body part movements. Keming Shen, Bizhu Wu, Junliang Chen 0002, LinLin Shen |
AAAI | 5 |
| 2026 | Incomplete Multi-view Diabetic Retinopathy Grading via Self-Supervised Inter- and Intra-View RestorationabstractMulti-view diabetic retinopathy (DR) grading has achieved remarkable performance by capturing more comprehensive pathological features than single-view methods. However, complete multi-view fundus images are often difficult to obtain in clinical practice, and the performance degrades significantly when fewer views are available. To overcome this limitation, we propose the first incomplete multi-view DR grading framework, aiming to provide accurate diagnosis regardless of the number of available views. It introduces two novel modules. First, cross-view spatial correlation attention (CSCA) captures region correlations across views, automatically identifying and fusing diagnostically relevant spatial features to improve feature representation. Second, self-supervised mask consistency learning (SMCL) formulates a novel pretext task of missing-view information reconstruction by strategically masking inter- and intra-view regions, enabling the model to infer complete features from incomplete views. Benefiting from CSCA and SMCL, our method enhances structural feature consistency across views and effectively compensates for missing information during DR grading. Extensive experiments demonstrate that our method achieves state-of-the-art grading performance, particularly under realistic conditions where some views are unavailable. Zhihao Wu 0002, Jie Wen 0001, Wuzhen Shi, LinLin Shen |
AAAI | 5 |
| 2026 | KoCo: Conditioning Language Model Pre-training on Knowledge CoordinatesabstractStandard Large Language Model (LLM) pretraining typically treats corpora as flattened token sequences, often overlooking the realworld context that humans naturally rely on to contextualize information.To bridge this gap, we introduce Knowledge Coordinate Conditioning (KoCo), a simple method that maps every document into a three-dimensional semantic coordinate.By prepending these coordinates as textual prefixes for pre-training, we aim to equip the model with explicit contextual awareness to learn the documents within the real-world knowledge structure.Experiment results demonstrate that KoCo significantly enhances performance across 10 downstream tasks and accelerates pre-training convergence by approximately 30%.Furthermore, our analysis indicates that explicitly modeling knowledge coordinates helps the model distinguish stable facts from noise, effectively mitigating hallucination in generated outputs.# Role: Knowledge Taxonomist & Data Curator # Core Task: Your task is to analyze the provided [TEXT] and its source [URL], and generate a single, concise "Knowledge Context Meta-Tag."This tag serves solely to inform the training model of the "position", "sourcle", and "attribute" of the text it will read within the broader human knowledge system.Note that your goal is to **classify text**, not **summarize the content of the text**.Do not expand or modify existing information.-# Step Guide: 1. **Analysis [URL]:** * Check the domain name (e.g Yudong Li 0001, Jiawei Cai, LinLin Shen |
ACL (1) | 3 |
| 2026 | Benchmarking Egocentric Clinical Intent Understanding Capability for Medical Multimodal Large Language ModelsabstractShaonan Liu, Guo Yu, Xiaoling Luo, Shiyi Zheng, Jie Liu, Wenting Chen, Linlin Shen. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Shaonan Liu, Xiaoling Luo 0001, Shiyi Zheng, Jie Liu 0044, Wenting Chen, LinLin Shen |
ACL (1) | 7 |
| 2026 | Beyond the Leaderboard: Rethinking Medical Benchmarks for Large Language ModelsabstractWenxuan Wang, Zizhan Ma, Guo Yu, Yiu-Fai Cheung, Meidan Ding, Jie Liu, Wenting Chen, Linlin Shen. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Wenxuan Wang 0001, Zizhan Ma, Yiu-Fai Cheung, Meidan Ding, Jie Liu 0044, Wenting Chen, LinLin Shen |
ACL (1) | 8 |
| 2026 | Shape-aware and feature fused power line detection network
Shengdong Zhang, Xiaoqin Zhang 0002, Wenqi Ren, LinLin Shen, Jun Zhang 0011 |
Eng. Appl. Artif. Intell. | 4 |
| 2026 | Rewarding fine-grained image captioning with keyword group contrastive
Kailiang Ye, Zheng Lu 0002, LinLin Shen, Tianxiang Cui |
Expert Syst. Appl. | 3 |
| 2026 | Multi-source multi-task meta-learning with task-oriented distribution alignment for gastric cancer analysis in CT images
Ning Yuan, Yiyao Liu, Yingpeng Xie, Jixin Luan, Kuan Lv, Tianfu Wang 0001, Harry Qin, LinLin Shen, Guolin Ma, Bai Ying Lei |
Expert Syst. Appl. | 13 |
| 2026 | Text condition embedded regression network for automated dental implant abutment design
Mianjie Zheng, Xinquan Yang, Xiaoling Luo 0001, Xuefen Liu, He Meng 0008, LinLin Shen |
Expert Syst. Appl. | 8 |
| 2026 | Weakly Supervised Salient Object Detection with Text Supervision
Zhihao Wu 0002, Jie Wen 0001, LinLin Shen, Xiaopeng Fan 0001, Yong Xu 0001, Jian Yang 0003, David Zhang 0001 |
Int. J. Comput. Vis. | 3 |
| 2026 | MNSeg: Mamba-based 3D neuron segmentation integrated with bidirectional attention mechanism and topological loss
Xinle Dai, Qiufu Li, LinLin Shen, Wenting Chen, Weijia Fan |
Neurocomputing | 3 |
| 2026 | Fine-grained facial description generation with retrieval augmentation
Kailiang Ye, Zheng Lu 0002, LinLin Shen, Tianxiang Cui |
Neurocomputing | 3 |
| 2026 | Federated semi-supervised calibrated efficient fine-tuning of foundation models for medical image classification
Along He, Yanlin Wu, LinLin Shen, Ke Zou, Huazhu Fu |
Knowl. Based Syst. | 3 |
| 2026 | Compressed video-driven multimodal modeling and interaction for dynamic expression recognition
Weicheng Xie 0001, Junliang Zhang, Haijian Liang, LinLin Shen, Zhihui Lai 0001, Siyang Song, Zitong Yu |
Knowl. Based Syst. | 4 |
| 2026 | UDG-Prom: A unified dense-guided semantic prompting for cross-domain few-shot image segmentationabstract• MAF preserves low-level feature representations, while fusing global and local information to generate robust class-agnostic features. • TA2MP, as a unified feature transformation mechanism equipped with an automatic learnable prompt branch, reduces human reliance and disentangles domain- and class-specific information through contrastive learning. • UDG-Prom integrates the MAF and TA2MP modules to address the CD-FSS task with SAM. • Our model achieves competitive or superior performance compared to state-of-the-art methods on four CD-FSS benchmarks, and its strong generalization ability is comprehensively validated through evaluations on more difficult cross-domain datasets including CT-Lung (medical) and SUIM (underwater). Large Vision Models (LVMs), exemplified by SAM, contain powerful general knowledge from extensive pre-training, yet they often underperform in highly specialized domains. Building large models tailored for each domain is usually impractical due to the substantial cost of data collection and training. Therefore, a key challenge is how to tap into SAM’s strong knowledge base and transfer it effectively to new, domain-specific tasks, especially under Cross-Domain or Few-Shot constraints. Previous efforts have leveraged prior knowledge from foundation models for transfer learning; however, they typically target specific tasks and exhibit limited robustness in broader applications. To tackle this issue, we propose a Unified Dense-Guided Semantic Prompting framework (UDG-Prom), a new paradigm for Cross-Domain Few-Shot Segmentation (CD-FSS). First, a Multi-level Adaptation Framework (MAF) is used for integrated feature extraction as prior knowledge. Then, we incorporate a Task-Adaptive Auto Meta Prompt (TA 2 MP) module to enable the extraction of class-domain-agnostic features and generate high-quality, learnable visual prompts. By combining learnable prompts with a structured model and prototype disentanglement, this method retains SAM’s prior knowledge and effectively adapts to CD-FSS through category and domain cues. Extensive experiments on four benchmarks show that our model not only surpasses state-of-the-art CD-FSS approaches but also achieves a remarkable improvement in average accuracy. Xiangjian He, Xin Chen 0003, Jingxi Hu, LinLin Shen, Guoping Qiu |
Knowl. Based Syst. | 6 |
| 2026 | MFF-M3AD: A unified reconstruction method with multi-scale feature fusion for multi-category 3D anomaly detection
Hanzhe Liang, Yejin Tang, LinLin Shen, Jinbao Wang 0001, Can Gao |
Neural Networks | 4 |
| 2026 | AEPL: Adaptive empirical prototype learning with dynamic margins for deep face recognition
Weijia Fan, Zhixiang Cai, Chunsong Chen, Yanxi Liu 0004, Jiajun Wen 0001, Xi Jia, LinLin Shen, Jiancan Zhou, Qiufu Li |
Pattern Anal. Appl. | 8 |
| 2026 | BinaryAD: Efficient image anomaly detection via binarized representations
Bingyang Guo, Hanzhe Liang, LinLin Shen, Jinbao Wang 0001, Zhichao Lu |
Pattern Recognit. | 6 |
| 2026 | FProtoSeg: Fine-grained prototype alignment for Weakly Supervised Semantic Segmentation of histopathology images
Meidan Ding, Wenting Chen, Xiaoling Luo 0001, Haiqin Zhong, LinLin Shen |
Pattern Recognit. | 5 |
| 2026 | Learning discriminative features within forward-Forward algorithm using convolutional prototype
Qiufu Li, LinLin Shen |
Pattern Recognit. | 3 |
| 2026 | Zero-shot referring expression comprehension via guidance of Multimodal Large Language Models
Rouyi Li, Shiyi Zheng, Zhihao Wu 0002, LinLin Shen |
Pattern Recognit. | 5 |
| 2026 | A lightweight 3D anomaly detection method with rotationally invariant features
Hanzhe Liang, Jie Zhou 0009, Can Gao, Bingyang Guo, Jinbao Wang 0001, LinLin Shen |
Pattern Recognit. | 6 |
| 2026 | M-MambaS: Multimodal Mamba for small lesion segmentation
Gui Wang, Jianfeng Ren, LinLin Shen, Wooi Ping Cheah, Rong Qu |
Pattern Recognit. | 3 |
| 2026 | Open-world Weakly-Supervised Object Localization
Jinheng Xie, Zhaochuan Luo, Rouyi Li, Yawen Huang, Yuexiang Li, Yefeng Zheng 0001, Yang Zhang 0012, LinLin Shen, Zheng Shou 0001 |
Pattern Recognit. | 9 |
| 2026 | Enhancing 3D medical multi-modal large language models with integrated human body priors for computed tomography
Leilei Zeng, Jie Liu 0044, Wenting Chen, Chenyang Lyu, Wenxi Li, Shaonan Liu, Xiande Zhou, LinLin Shen |
Pattern Recognit. | 8 |
| 2026 | SDC-Net: Semi-supervised breast ultrasound lesion segmentation via semantic decoupling
Jiansong Zhang 0005, Zhuoqin Yang, Xiaoling Luo 0001, Shaozheng He, LinLin Shen |
Pattern Recognit. | 5 |
| 2026 | Wavelet-based physically guided normalization network for real-time traffic dehazing
Shengdong Zhang, Xiaoqin Zhang 0002, LinLin Shen, Shaohua Wan 0001, Wenqi Ren |
Pattern Recognit. | 3 |
| 2026 | S3DL: Sample-Aggregated Structured Supervised Dictionary Learning
Haiyan Yu 0003, Yucheng Peng, Jianfeng Ren, LinLin Shen, Xin Chen 0003, Ruibin Bai |
IEEE Signal Process. Lett. | 4 |
| 2026 | Hierarchical Multi-Modal Enhancement for Robust Transmission Line Detection
Shengdong Zhang, Xiaoqin Zhang 0002, Shaohua Wan 0001, Yujing M. Jiang, Wujie Zhou, LinLin Shen, Wenqi Ren |
IEEE Trans Autom. Sci. Eng. | 6 |
| 2026 | DeCenter: Density-Center Guided Perception Enhancement for UAV Object DetectionabstractUnmanned aerial vehicle (UAV) object detection is essential for applications such as surveillance, agriculture, and disaster response. However, UAV imagery often contains small, dense, and occluded objects, posing challenges for existing methods. To address these challenges, we propose DeCenter, a novel Density-Center Guided Perception Enhancement framework for UAV object detection. DeCenter is composed of two key modules that jointly enhance the perception of small and crowded objects. First, the Density-Guided Object Center Heatmap Generator (DOCHG) adaptively generates Gaussian kernel-based heatmaps according to local density information, guiding the model to emphasize central neighborhoods of objects in crowded regions. This mechanism reduces overlaps between adjacent instances and alleviates missed detections under occlusion. Second, the Density-Center Feature Enhancement module (DCFE) integrates complementary cues from density features and object centers, adaptively balancing region-level object distribution with fine-grained localization. By fusing these signals, DCFE enhances the quality of feature representations, making them more discriminative for dense small objects while suppressing background noise. Experimental results on VisDrone and UAVDT datasets show that DeCenter achieves competitive overall accuracy with clear improvements in detecting dense small objects, offering an effective solution for UAV object detection. The code will be available at https://github.com/bluuzzz/decenter. Zhiqing Shi, Zhihao Wu 0002, Jie Wen 0001, Mu Li 0005, Xiaopeng Fan 0001, Yaowei Wang 0001, LinLin Shen |
IEEE Trans. Circuits Syst. Video Technol. | 7 |
| 2026 | HemNet: Hemoglobin-Assistant Network for Video-Based Remote Photoplethysmography MeasurementabstractTraditional skin-contact physical sensors typically detect changes of blood volume to predict the periodicity of heartbeat by analyzing the absorption spectra of hemoglobin. However, the contact on human skin may cause uncomfortable feeling and induce difficulty for long-term monitoring. Recently, video-based remote photoplethysmography (rPPG) estimation approaches analyze the periodic facial color changes for matching cardiac cycle in a contactless manner. Nevertheless, the inherent relationship between the changes of facial color and blood volume is not fully exploited. Besides the influence of blood volume (i.e., hemoglobin), there are also other factors such as lighting and reflection that cause the change on facial color. We exploit the physical principles that cause skin color variations to separate the hemoglobin factor driven by blood volume. Based on the physical prior of the reflection of human skin, we introduce an rPPG estimation network assisted by decoupled hemoglobin sequence, named HemNet, which first explicitly leverages hemoglobin to assist rPPG signal estimation. To obtain meaningful hemoglobin from facial video, we design a human skin color disentangler that decouples the facial color variations into four significant features, i.e., hemoglobin, melanin, shading, and specular. We then present a multi-modality rPPG estimator that utilizes cross-covariance attention to extract fused feature from hemoglobin and RGB video inputs. Finally, an adaptive negative Pearson loss is proposed to effectively address phase misalignment between the blood volume in the finger and facial region during the training phase. We evaluate our HemNet on four widely used public benchmark datasets. The superiority of our method is demonstrated in both intra-dataset and cross-dataset test settings. The code is available at https://github.com/jingang-cv/hemnet. Ruize Wu, Jingang Shi, Xin Liu 0012, LinLin Shen, Yihong Gong, Guoying Zhao 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2026 | Dual-Stream Autoencoder With Spatial and High-Frequency Feature Interaction for Contactless Fingerprint Presentation Attack Detection
Feng Liu 0013, LinLin Shen |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2026 | GS2Physics: Semantic-Region-Aware Gaussian Splatting for Physical Property PredictionabstractPredicting the physical properties of reconstructed 3D assets is essential for virtual reality interactions. However, current systems often depend on manually assigning properties such as stiffness and density, which can be inefficient and prone to errors. To address this issue, we present GS2Physics, a novel framework based on 3D Gaussian Splatting. This framework is designed to predict physical properties accurately while maintaining improved consistency in semantic segmentation. Unlike existing approaches, which either struggle with region inconsistency or misalign semantic 3D features, GS2Physics embeds semantic-region-aware features directly into the Gaussian Splatting representation. This allows for region-consistent and accurate physical property prediction, achieving state-of-the-art performance on the ABO-500 mass prediction benchmark. To further evaluate our segmentation capabilities, we introduce PhysSeg-15, a subset dataset of ABO-500 featuring physical property segmentation masks for 15 different 3D objects captured from five viewpoints. Our method significantly outperforms existing approaches in segmentation accuracy. Qualitative results demonstrate more consistent material predictions across different object regions and improved accuracy in physical property prediction. In addition, we showcase the effectiveness of GS2Physics in 3D interaction tasks, where our predicted physical properties result in more realistic object motion. Our dataset and results are available at https://github.com/momaiyc/GS2Physics. Bin Huang 0016, Jiayi Lyu, Zehai Niu, LinLin Shen, Jinbao Wang 0001 |
IEEE Trans. Image Process. | 5 |
| 2026 | Multi-Granularity Facial Emotional Representation With Unlabeled Data and Textual SupervisionabstractFacial expressions (FEs) and action units (AUs) are facial emotional representations at different levels of granularity. In the past, recognizing them has often been treated as two separate tasks. There are also some methods that use the knowledge of one to aid in recognizing the other, but currently, unified models capable of recognizing both FEs and AUs simultaneously remain rare. In this paper, we construct a unified model with strong generalization capability to jointly perform facial expression recognition (FER) and action unit detection (AUD). Considering the extremely limited training samples annotated with both FEs and AUs, we introduce a large amount of unlabeled facial data from the wild. We carefully design category-specific confidence margins and leverage the correspondences between FEs and AUs to assign credible pseudo-labels to the unlabeled facial data. Furthermore, we incorporate semantically richer textual descriptions as supervision and refine them through visual perception, leveraging the inherent correlations between AUs and between FEs and AUs to enhance their precision. Extensive experiments demonstrate the superiority of the proposed method from various perspectives, including a unified zero-shot benchmark for exploring the model's comprehensive generalization capability to recognize facial emotional representations across multiple datasets, as well as within-domain and cross-domain evaluations after fine-tuning. The code for the proposed method is available at https://github.com/yuankaishen2001/MGFER. Kaishen Yuan, Zitong Yu, Xin Liu 0012, Bohao Xing, Yuting Zhang 0008, Weicheng Xie 0001, LinLin Shen, Björn W. Schuller |
IEEE Trans. Image Process. | 7 |
| 2026 | UINO-FSS: Unifying Representation Learning and Few-Shot Segmentation via Hierarchical Distillation and Mamba-HyperCorrelationabstractFew-shot semantic segmentation has attracted growing interest for its ability to generalize to novel object categories using only a few annotated samples. To address data scarcity, recent methods incorporate multiple foundation models to improve feature transferability and segmentation performance. However, they often rely on dual-branch architectures that combine pre-trained encoders to leverage complementary strengths, a design that limits flexibility and efficiency. This raises a fundamental question: "can we build a unified model that integrates knowledge from different foundation architectures?" Achieving this is, however, challenging due to the misalignment between class-agnostic segmentation capabilities and fine-grained discriminative representations. To this end, we present UINO-FSS (pronounced //), a novel framework built on the key observation that early-stage DINOv2 features exhibit distribution consistency with SAM's output embeddings. This consistency enables the integration of both models' knowledge into a single-encoder architecture via coarse-to-fine multimodal distillation. In particular, our segmenter consists of three core components: a bottleneck adapter for embedding alignment, a meta-visual prompt generator that leverages dense similarity volumes and semantic embeddings, and a mask decoder. Using hierarchical cross-model distillation, we effectively transfer SAM's knowledge into the segmenter, further enhanced by Mamba-based 4D correlation mining on support-query pairs. Extensive experiments show that UINO-FSS achieves new state-of-the-art results on COCO- $20^{i}$ under the 1-shot setting, with an mIoU of 64.5% (+2.2%), while also delivering competitive performance on PASCAL- $5^{i}$ . Zhiyue Tang, Wufeng Xue, Junkai Ji, LinLin Shen |
IEEE Trans. Image Process. | 6 |
| 2026 | ACGM: Attribute-Centric Graph Modeling Network for Concurrent Missing Tabular Data Imputation and COVID-19 PrognosisabstractCOVID-19 prognosis using clinical tabular data faces significant challenges due to missing values and class imbalance issues. Existing methods often overlook the complex high-order interrelationship among clinical attributes and struggle with training stability on imbalanced datasets. We propose ACGM, an attribute-centric graph modeling network that simultaneously addresses missing data imputation and COVID-19 prognosis. ACGM consists of three key modules: an attributes preprocessing module (APM) for coarse-grained imputation initialization, a graph-enhanced attributes imputation module (GEAIM) that models high-order inter-attribute relationships through graph structures, and a graph-enhanced disease prognosis module (GEDPM) that leverages these complex attribute interactions for final prediction. GEAIM and GEDPM employ a mean-teacher strategy with attributes graph matching to preserve high-order relationships, enhance training stability, and maintain structural integrity of attribute interactions. Extensive experiments are conducted on four public COVID-19 tabular datasets, demonstrating the superiority of our ACGM over existing methods. Through comprehensive interpretability analysis, we identify that attributes such as LDH, Difficulty In Breathing, and SaO2 significantly impact COVID-19 prognosis, aligning well with clinical insights and radiologist assessments. Zhuoru Wu, Wenting Chen, Xuechen Li 0001, Filippo Ruffini, Shaonan Liu, Lorenzo Tronchin, Domenico Albano, Eliodoro Faiella, Deborah Fazzini, Domiziana Santucci, Xiaoling Luo 0001, Valerio Guarrasi, Paolo Soda, LinLin Shen |
IEEE J. Biomed. Health Informatics | 14 |
| 2026 | Thyro-LMD: A Benchmark Dataset and Sample-Driven Data Loading, Attention, and Regularization for Long-Tailed Multi-Label Thyroid Ultrasound DiagnosisabstractDeveloping robust and effective computer-aided diagnostic (CAD) methods for thyroid ultrasound (TUS) remains a key challenge in medical imaging. Prior work has largely focused on binary or multi-class lesion classification, whereas real-world diagnosis follows standardized guidelines based on combinations of lexicon-level descriptors. These combinations naturally exhibit long-tailed distributions due to epidemiological patterns, limiting the robustness and generalizability of existing methods. Motivated by this, we introduce Thyro-LMD, the first long-tailed multi-label dataset for TUS. Using histopathology as the reference, Thyro-LMD provides retrospective, fine-grained annotations aligned with ACR TI-RADS lexicons and reveals a highly imbalanced label distribution. We benchmark representative methods, including end-to-end models, general-purpose multimodal large models (e.g., GPT-4o), and pretrained foundation models. While some methods show reasonable head-class performance, they struggle with body and tail classes. We therefore propose SynTUS-Net, a purpose-built baseline comprising collaborative modules addressing long-tailed multi-label challenges across data loading, feature encoding, and prediction regularization. SynTUS-Net achieves leading performance on Thyro-LMD, outperforming conventional traditional SOTA models by 5.3 Micro-F1 and 11.83 Macro-F1, and exceeding GPT-4o by 42.76 on Tail-F1. Extensive ablation studies confirm the contribution of each module. We believe Thyro-LMD and SynTUS-Net establish a clinically grounded benchmark and a new paradigm for interpretable and generalizable AI in ultrasound. Code and data will be released here. Jiansong Zhang 0005, Shunlan Liu, Xiaoling Luo 0001, Guorong Lyu, LinLin Shen |
IEEE Trans. Medical Imaging | 5 |
| 2026 | Ranking-Based Self-Supervised Representation Learning for Skeleton-Based Action RecognitionabstractRecently, researchers have achieved significant results in the skeleton-based action recognition. To better model the skeleton sequences, we drive the encoder to learn more discriminative representations in the self-supervised setting. We find that instead of clustering feature vectors to assign pseudo labels for samples as in DeepCluster, ranking them is a more reasonable, reliable, and efficient way to learn more effective feature representations. With this intuition, we propose a novel self-supervised learning framework,DeepRank. Specifically, we rank triplets of skeleton sequences with the ranking labels, obtained from the relative distances among them. Besides, to deeply mine complementary discriminative information that exists in different modalities of skeleton sequences, we further proposeMulti-ViewDeepRank(MV-DeepRank) to enable encoders to comprehensively learn complementary features from multiple modalities. Extensive experimental results on the NTU RGB+D, NTU RGB+D 120, PKU-MMD I, and PKU-MMD II datasets under various evaluation settings demonstrate the generality, transferability, and superiority of our proposed self-supervised learning frameworks. Notably, our frameworks surpass the previous methods that employ the same backbone networks as ours by at least 1.8% (ST-GCN) and 2.1% (STTFormer) under the finetuning setting. Additionally, DeepRank gains a significant advantage on computational complexities,$O(1)$, over the contrastive learning-based methods,$O(\rm{batch size})$, and the clustering-based methods,$O(\rm{number of clusters})$. Bizhu Wu, Junliang Chen 0002, Jinheng Xie, Qiufu Li, Jianfeng Ren, Ruibin Bai, Rong Qu, LinLin Shen |
IEEE Trans. Multim. | 8 |
| 2026 | LightQANet: Quantized and Adaptive Feature Learning for Low-Light Image EnhancementabstractLow-light image enhancement (LLIE) aims to improve illumination while preserving high-quality color and texture. However, existing methods often fail to extract reliable feature representations due to severely degraded pixel-level information under low-light conditions, resulting in poor texture restoration, color inconsistency, and artifact. To address these challenges, we propose LightQANet, a novel framework that introduces quantized and adaptive feature learning for low-light enhancement, aiming to achieve consistent and robust image quality across diverse lighting conditions. From the static modeling perspective, we design a Light Quantization Module (LQM) to explicitly extract and quantify illumination-related factors from image features. By enforcing structured light factor learning, LQM enhances the extraction of light-invariant representations and mitigates feature inconsistency across varying illumination levels. From the dynamic adaptation perspective, we introduce a Light-Aware Prompt Module (LAPM), which encodes illumination priors into learnable prompts to dynamically guide the feature learning process. LAPM enables the model to flexibly adapt to complex and continuously changing lighting conditions, further improving image enhancement. Extensive experiments on multiple low-light datasets demonstrate that our method achieves state-of-the-art performance, delivering superior qualitative and quantitative results across various challenging lighting scenarios. Xu Wu 0001, Zhihui Lai 0001, Xianxu Hou, Jie Zhou 0009, LinLin Shen |
IEEE Trans. Multim. | 6 |
| 2026 | HitBack: Transformer With Hierarchical-Semantic Cross Attention and Background Contrast for Weakly Supervised Wildlife Semantic SegmentationabstractMonitoring wildlife behavior and population changes is critical for conservation efforts. However, specialized analysis of large volumes of wildlife images is extremely chal lenging, necessitating the use of artificial intelligence techniques to automatically detect, segment, and classify species captured by trap cameras. Despite the increasing use of AI in wildlife monitoring, challenges with data quality and availability persist. The Snapshot Serengeti (SS) dataset only has image-level labels and very few bounding box labels, and there's no dataset with pixel-level labels due to the significant annotation costs. To this end, we create and release the large-scale Semantic Segmentation for Snapshots of the Serengeti (S4) dataset, consisting of 24K high-quality images across 47 species with precise masks, for both common and rare species. This dataset serves as a resource for developing semantic segmentation algorithms in wildlife studies. Additionally, we introduce HitBack, a novel method leveraging Hierarchical-Semantic Cross Attention (HCA) and Background Contrast (BC) for weakly supervised semantic segmentation (WSSS). The HCA module is used to capture both the shared and distinct features across species, and the BC module is designed to enhance foreground activation by ensuring consistency in the backgrounds. Extensive experiments on the newly proposed S4benchmark show that, our HitBack presents competitive performance when compared with the state-of-the-art models. The mIoU of HitBack is +10.4%, +14.7%, and +18.4% higher than that of ToCo, SIPE, and MCTformer, respectively. In addition, our HitBack even obtains performance that surpasses the fully supervised and semi-supervised methods when annotation data is limited. Code and datasets will be available at Github. Puxuan Xie, Xinshao Wang, Songhe Deng, Weizhao He, LinLin Shen |
IEEE Trans. Multim. | 6 |
| 2026 | FPGA Routing Congestion Prediction via Graph Learning-Aided Conditional GANabstractRouting congestion prediction expedites the closure of FPGA placement and routing (PnR). Current prediction methods employ convolutional models, taking advantage of their capacity of dealing with image-style inputs. However, these methods neglect the direct representation of circuit netlist and its information fusion with placement scheme. Moreover, the limited size of the convolutional kernel struggles to capture circuit connectivity in distant geometric regions. To address these issues, this article presents a graph-based routing congestion prediction framework that fuses the information contained in the circuit’s topological netlist and geometric placement scheme, and leverages a conditional generative adversarial network (cGAN) model to achieve optimized prediction performance compared to contemporary approaches. Our framework encompasses three key components: (1) the HeteroGraph, a heterogeneous graph that integrates a netlist subgraph and a layout subgraph by space mapping edges; (2) the HeteroGNN, a heterogeneous graph neural network that learns the latent features of both the circuit netlist and placement scheme through dual-space message-passing; and (3) the HeteroGNN-embedded cGAN, a model that combines the HeteroGNN with a cGAN for accurate FPGA routing congestion prediction. Compared to state-of-the-art approaches, our method reduces the routing congestion prediction’s root-mean-square error by 18.2% on the VTR7 benchmarks and by 15.0% on the large-scale Titan23 benchmarks. The code associated with this article can be found at https://github.com/AIPnR/FPGA_Hetero_Congestion_Prediction . Qingyu Yang 0004, Jingjin Li, Rui Li 0095, Yuting He 0002, Yajun Ha, LinLin Shen, Ruibin Bai, Heng Yu 0001 |
ACM Trans. Design Autom. Electr. Syst. | 6 |
| 2025 | DAMPER: A Dual-Stage Medical Report Generation Framework with Coarse-Grained MeSH Alignment and Fine-Grained Hypergraph MatchingabstractMedical report generation is crucial for clinical diagnosis and patient management, summarizing diagnoses and recommendations based on medical imaging. However, existing work often overlook the clinical pipeline involved in report writing, where physicians typically conduct an initial quick review followed by a detailed examination. Moreover, current alignment methods may lead to misaligned relationships. To address these issues, we propose DAMPER, a dual-stage framework for medical report generation that mimics the clinical pipeline of report writing in two stages. In the first stage, a MeSH-Guided Coarse-Grained Alignment (MCG) stage that aligns chest X-ray (CXR) image features with medical subject headings (MeSH) features to generate a rough keyphrase representation of the overall impression. In the second stage, a Hypergraph-Enhanced Fine-Grained Alignment (HFG) stage that constructs hypergraphs for image patches and report annotations, modeling high-order relationships within each modality and performing hypergraph matching to capture semantic correlations between image regions and textual phrases. Finally,the coarse-grained visual features, generated MeSH representations, and visual hypergraph features are fed into a report decoder to produce the final medical report. Extensive experiments on public datasets demonstrate the effectiveness of DAMPER in generating comprehensive and accurate medical reports, outperforming state-of-the-art methods across various evaluation metrics. Wenting Chen, Jie Liu 0044, Qisheng Lu, Xiaoling Luo 0001, LinLin Shen |
AAAI | 6 |
| 2025 | Like an Ophthalmologist: Dynamic Selection Driven Multi-View Learning for Diabetic Retinopathy GradingabstractDiabetic retinopathy (DR), with its large patient population, has become a formidable threat to human visual health. In the clinical diagnosis of DR, multi-view fundus images are considered to be more suitable for DR diagnosis because of the wide coverage of the field of view. Therefore, different from most of the previous single-view DR grading methods, we design a dynamic selection-driven multi-view DR grading method to fit clinical scenarios better. Since lesion information plays a key role in DR diagnosis, previous methods usually boost the model performance by enhancing the lesion feature. However, during the actual diagnosis, ophthalmologists not only focus on the crucial parts, but also exclude irrelevant features to ensure the accuracy of judgment. To this end, we introduce the idea of dynamic selection and design a series of selection mechanisms from fine granularity to coarse granularity. In this work, we first introduce an Ophthalmic Image Reader (OIR) agent to provide the model with pixel-level prompts of suspected lesion areas. Moreover, a Multi-View Token Selection Module (MVTSM) is designed to prune redundant feature tokens and realize dynamic selection of key information. In the final decision stage, we dynamically fuse multi-view features through the novel Multi-View Mixture of Experts Module (MVMoEM), to enhance key views and reduce the impact of conflicting views. Extensive experiments on a large multi-view fundus image dataset with 34,452 images demonstrate that our method performs favorably against state-of-the-art models. Xiaoling Luo 0001, Qihao Xu, Huisi Wu, Chengliang Liu 0003, Zhihui Lai 0001, LinLin Shen |
AAAI | 6 |
| 2025 | S³-Mamba: Small-Size-Sensitive Mamba for Lesion SegmentationabstractSmall lesions play a critical role in early disease diagnosis and intervention of severe infections. Popular models often face challenges in segmenting small lesions, as it occupies only a minor portion of an image, while down-sampling operations may inevitably lose focus on local features of small lesions. To tackle the challenges, we propose a Small-Size-Sensitive Mamba (S³-Mamba), which promotes the sensitivity to small lesions across three dimensions: channel, spatial, and training strategy. Specifically, an Enhanced Visual State Space block is designed to focus on small lesions through multiple residual connections to preserve local features, and selectively amplify important details while suppressing irrelevant ones through channel-wise attention. A Tensor-based Cross-feature Multi-scale Attention is designed to integrate input image features and intermediate-layer features with edge features and exploit the attentive support of features across multiple scales, thereby retaining spatial details of small lesions at various granularities. Finally, we introduce a novel regularized curriculum learning to automatically assess lesion size and sample difficulty, and gradually focus from easy samples to hard ones like small lesions. Extensive experiments on three medical image segmentation datasets show the superiority of our S³-Mamba, especially in segmenting small lesions. Gui Wang, Yuexiang Li, Wenting Chen, Meidan Ding, Wooi Ping Cheah, Rong Qu, Jianfeng Ren, LinLin Shen |
AAAI | 8 |
| 2025 | CA-Edit: Causality-Aware Condition Adapter for High-Fidelity Local Facial Attribute EditingabstractFor efficient and high-fidelity local facial attribute editing, most existing editing methods either require additional fine-tuning for different editing effects or tend to affect beyond the editing regions. Alternatively, inpainting methods can edit the target image region while preserving external areas. However, current inpainting methods still suffer from the generation misalignment with facial attributes description and the loss of facial skin details. To address these challenges, (i) a novel data utilization strategy is introduced to construct datasets consisting of attribute-text-image triples from a data-driven perspective, (ii) a Causality-Aware Condition Adapter is proposed to enhance the contextual causality modeling of specific details, which encodes the skin details from the original image while preventing conflicts between these cues and textual conditions. In addition, a Skin Transition Frequency Guidance technique is introduced for the local modeling of contextual causality via sampling guidance driven by low-frequency alignment. Extensive quantitative and qualitative experiments demonstrate the effectiveness of our method in boosting both fidelity and editability for localized attribute editing. Our codes will be made publicly available. Xiaole Xian, Xilin He, Zenghao Niu, Junliang Zhang, Weicheng Xie 0001, Siyang Song, Zitong Yu, LinLin Shen |
AAAI | 8 |
| 2025 | PerReactor: Offline Personalised Multiple Appropriate Facial Reaction GenerationabstractIn dyadic human-human interactions, individuals may express multiple different facial reactions in response to the same/similar behaviours expressed by their conversational partners depending on their personalised behaviour patterns. As a result, frequently-employed reconstruction loss-based strategies lead the training of previous automatic facial reaction generation (FRG) models to not only suffer from the 'one-to-many mapping' problem, but also fail to comprehensively consider the quality of the generated facial reactions. Besides, none of them considered such personalised behaviour patterns in generating facial reactions. In this paper, we propose the first adversarial FRG model training strategy which jointly learns appropriateness and realism discriminators to provide comprehensive task-specific supervision for training the target facial reaction generators, and reformulates the 'one-to-many (facial reactions) mapping' training problem as a 'one-to-one (distribution) mapping' training task, i.e., the FRG model is trained to output a distribution representing multiple appropriate/plausible facial reaction from each input human behaviour. In addition, our approach also serves as the first offline FRG approach that considers personalised behaviour patterns in generating of target individuals' facial reactions. Experiments show that our PerReactor not only largely outperformed all existing offline solutions for generating more appropriate, diverse and realistic facial reactions, but also is the first approach that can effectively generate personalised appropriate facial reactions. Hengde Zhu, Xiangyu Kong 0001, Weicheng Xie 0001, Xilin He, Lu Liu 0001, LinLin Shen, Wei Zhang 0243, Hatice Gunes, Siyang Song |
AAAI | 7 |
| 2025 | EAGLE: Expert-Guided Self-Enhancement for Preference Alignment in Pathology Large Vision-Language ModelabstractRecent advancements in Large Vision Language Models (LVLMs) show promise for pathological diagnosis, yet their application in clinical settings faces critical challenges of multimodal hallucination and biased responses. While preference alignment methods have proven effective in general domains, acquiring high-quality preference data for pathology remains challenging due to limited expert resources and domain complexity. In this paper, we propose EAGLE (Expert-guided self-enhancement for preference Alignment in patholoGy Large vision-languagE model), a novel framework that systematically integrates medical expertise into preference alignment. EAGLE consists of three key stages: initialization through supervised fine-tuning, self-preference creation leveraging expert prompting and medical entity recognition, and iterative preference following-tuning. The self-preference creation stage uniquely combines expert-verified chosen sampling with expert-guided rejected sampling to generate high-quality preference data, while the iterative tuning process continuously refines both data quality and model performance. Extensive experiments demonstrate that EAGLE significantly outperforms existing pathological LVLMs, effectively reducing hallucination and bias while maintaining pathological accuracy. The source code is available at https://github.com/meidandz/EAGLE. © 2025 Association for Computational Linguistics. Meidan Ding, Wenxuan Wang 0001, Haiqin Zhong, Xinheng Lyu, Wenting Chen, LinLin Shen |
ACL (1) | 8 |
| 2025 | Asclepius: A Spectrum Evaluation Benchmark for Medical Multi-Modal Large Language ModelsabstractThe significant breakthroughs of Medical Multi-Modal Large Language Models (Med-MLLMs) renovate modern healthcare with robust information synthesis and medical decision support. However, these models are often evaluated on benchmarks that are unsuitable for the Med-MLLMs due to the complexity of real-world diagnostics across diverse specialties. To address this gap, we introduce Asclepius, a novel Med-MLLM benchmark that comprehensively assesses Med-MLLMs in terms of: distinct medical specialties (cardiovascular, gas-troenterology, etc.) and different diagnostic capacities (perception, disease analysis, etc.). Grounded in 3 proposed core principles, Asclepius ensures a comprehensive evaluation by encompassing 15 medical specialties, stratifying into 3 main categories and 8 sub-categories of clinical tasks, and exempting overlap with existing VQA dataset. We further provide an in-depth analysis of 6 Med-MLLMs and compare them with 3 human specialists, providing insights into their competencies and limitations in various medical contexts. Our work not only advances the understanding of Med-MLLMs' capabilities but also sets a precedent for future evaluations and the safe deployment of these models in clinical environments. © 2025 Association for Computational Linguistics. Jie Liu 0044, Wenxuan Wang 0001, Yihang Su, Yudi Zhang 0005, Cheng-Yi Li, Wenting Chen, Xiaohan Xing, Kao-Jung Chang, LinLin Shen, Michael R. Lyu |
ACL (1) | 10 |
| 2025 | MedKAN: An Advanced Kolmogorov-Arnold Network for Medical Image ClassificationabstractRecent advancements in deep learning for image classification predominantly rely on convolutional neural networks (CNNs) or Transformer-based architectures. However, these models face notable challenges in medical imaging, particularly in capturing intricate texture details and contextual features. Kolmogorov-Arnold Networks (KANs) represent a novel class of architectures that enhance nonlinear transformation modeling, offering improved representation of complex features. In this work, we present MedKAN, a medical image classification framework built upon KAN and its convolutional extensions. MedKAN features two core modules: the Local Information KAN (LIK) module for fine-grained feature extraction and the Global Information KAN (GIK) module for broad contextual representation learning. By combining these modules, MedKAN achieves robust feature modeling and fusion. To address diverse computational needs, we introduce three scalable variants-MedKAN-S, MedKAN-B, and MedKAN-L. Experimental results on nine public medical imaging datasets demonstrate that MedKAN achieves superior performance compared to CNN- and Transformer-based models, highlighting its effectiveness and generalizability in medical image analysis. Code: https://github.com/SeriYann/MedKAN Zhuoqin Yang, Xu Wu 0001, LinLin Shen |
BIBM | 6 |
| 2025 | Can Foundation Models Really Segment Tumors? A Benchmarking Odyssey in Lung CT ImagingabstractAccurate lung tumor segmentation is crucial for improving diagnosis, treatment planning, and patient outcomes in oncology. However, the complexity of tumor morphology, size, and location poses significant challenges for automated segmentation. This study presents a comprehensive benchmarking analysis of deep learning-based segmentation models, comparing traditional architectures such as U-Net and DeepLabV3, selfconfiguring models like nnUNet, and foundation models like MedSAM, and MedSAM 2. Evaluating performance across two lung tumor segmentation datasets, we assess segmentation accuracy and computational efficiency under various learning paradigms, including few-shot learning and fine-tuning. The results reveal that while traditional models struggle with tumor delineation, foundation models, particularly MedSAM 2, outperform them in both accuracy and computational efficiency. These findings underscore the potential of foundation models for lung tumor segmentation, highlighting their applicability in improving clinical workflows and patient outcomes. Elena Mulero Ayllon, Massimiliano Mantegna, LinLin Shen, Paolo Soda, Valerio Guarrasi, Matteo Tortora |
CBMS | 3 |
| 2025 | MANet-CycleGAN: An Unsupervised LDCT Image Denoising Method Based on Channel Attention and Multi-scale Features
Jinglong Tian, Tianze Zhao, Zhijun Fan, LinLin Shen, Jieyao Wei, Qiumei Pu |
CVM (1) | 4 |
| 2025 | FaceBench: A Multi-View Multi-Level Facial Attribute VQA Dataset for Benchmarking Face Perception MLLMsabstractMultimodal large language models (MLLMs) have demonstrated remarkable capabilities in various tasks. However, effectively evaluating these MLLMs on face perception remains largely unexplored. To address this gap, we introduce FaceBench, a dataset featuring hierarchical multi-view and multi-level attributes specifically designed to assess the comprehensive face perception abilities of MLLMs. Initially, we construct a hierarchical facial attribute structure, which encompasses five views with up to three levels of attributes, totaling over 210 attributes and 700 attribute values. Based on the structure, the proposed FaceBench consists of 49,919 visual questionanswering (VQA) pairs for evaluation and 23,841 pairs for fine-tuning. Moreover, we further develop a robust face perception MLLM baseline, Face-LLaVA, by training with our proposed face VQA data. Extensive experiments on various mainstream MLLMs and Face-LLaVA are conducted to test their face perception ability, with results also compared against human performance. The results reveal that, the existing MLLMs are far from satisfactory in understanding the fine-grained facial attributes, while our Face-LLaVA significantly outperforms existing open-source models with a small amount of training data and is comparable to commercial ones like GPT-4o and Gemini. The dataset will be released at https://github.com/CVI-SZU/FaceBench Xusen Ma, Xianxu Hou, Meidan Ding, Yudong Li 0001, Junliang Chen 0002, Wenting Chen, Xiaoyang Peng, LinLin Shen |
CVPR | 9 |
| 2025 | MG-MotionLLM: A Unified Framework for Motion Comprehension and Generation across Multiple GranularitiesabstractRecent motion-aware large language models have demonstrated promising potential in unifying motion comprehension and generation. However, existing approaches primarily focus on coarse-grained motion-text modeling, where text describes the overall semantics of an entire motion sequence in just a few words. This limits their ability to handle fine-grained motion-relevant tasks, such as understanding and controlling the movements of specific body parts. To overcome this limitation, we pioneer MG-MotionLLM, a unified motion-language model for multi-granular motion comprehension and generation. We further introduce a comprehensive multi-granularity training scheme by incorporating a set of novel auxiliary tasks, such as localizing temporal boundaries of motion segments via detailed text as well as motion detailed captioning, to facilitate mutual reinforcement for motion-text modeling across various levels of granularity. Extensive experiments show that our MG-MotionLLM achieves superior performance on classical text-to-motion and motion-to-text tasks, and exhibits potential in novel fine-grained motion comprehension and editing tasks. Project page: CVI-SZU/MG-MotionLLM Bizhu Wu, Jinheng Xie, Keming Shen, Zhe Kong, Jianfeng Ren, Ruibin Bai, Rong Qu, LinLin Shen |
CVPR | 8 |
| 2025 | D3-Talker: Dual-Branch Decoupled Deformation Fields for Few-Shot 3D Talking Head SynthesisabstractA key challenge in 3D talking head synthesis lies in the reliance on a long-duration talking head video to train a new model for each target identity from scratch. Recent methods have attempted to address this issue by extracting general features from audio through pre-training models. However, since audio contains information irrelevant to lip motion, existing approaches typically struggle to map the given audio to realistic lip behaviors in the target face when trained on only a few frames, causing poor lip synchronization and talking head image quality. This paper proposes D3-Talker, a novel approach that constructs a static 3D Gaussian attribute field and employs audio and Facial Motion signals to independently control two distinct Gaussian attribute deformation fields, effectively decoupling the predictions of general and personalized deformations. We design a novel similarity contrastive loss function during pre-training to achieve more thorough decoupling. Furthermore, we integrate a Coarse-to-Fine module to refine the rendered images, alleviating blurriness caused by head movements and enhancing overall image quality. Extensive experiments demonstrate that D3-Talker outperforms state-of-the-art methods in both high-fidelity rendering and accurate audio-lip synchronization with limited training data. Kaijun Deng, Siyang Song, Jindong Xie, Wenhui Ma, LinLin Shen |
ECAI | 6 |
| 2025 | DEGSTalk: Decomposed Per-Embedding Gaussian Fields for Hair-Preserving Talking Face SynthesisabstractAccurately synthesizing talking face videos and capturing fine facial features for individuals with long hair presents a significant challenge. To tackle these challenges in existing methods, we propose a decomposed per-embedding Gaussian fields (DEGSTalk), a 3D Gaussian Splatting (3DGS)-based talking face synthesis method for generating realistic talking faces with long hairs. Our DEGSTalk employs Deformable Pre-Embedding Gaussian Fields, which dynamically adjust pre-embedding Gaussian primitives using implicit expression coefficients. This enables precise capture of dynamic facial regions and subtle expressions. Additionally, we propose a Dynamic Hair-Preserving Portrait Rendering technique to enhance the realism of long hair motions in the synthesized videos. Results show that DEGSTalk achieves improved realism and synthesis quality compared to existing approaches, particularly in handling complex facial dynamics and hair preservation. Our code is available at https://github.com/CVI-SZU/DEGSTalk. Kaijun Deng, Dezhi Zheng, Jindong Xie, Jinbao Wang 0001, Weicheng Xie 0001, LinLin Shen, Siyang Song |
ICASSP | 6 |
| 2025 | Big-Moe: Bypassing Isolated Gating For Generalized Multimodal Face Anti-SpoofingabstractIn the domain of facial recognition security, multimodal Face Anti-Spoofing (FAS) is essential for countering presentation attacks. However, existing technologies encounter challenges due to modality biases and imbalances, as well as domain shifts. Our research introduces a Mixture of Experts (MoE) model to address these issues effectively. We identified three limitations in traditional MoE approaches to multimodal FAS: (1) Coarse-grained experts’ inability to capture nuanced spoofing indicators; (2) Gated networks’ susceptibility to input noise affecting decision-making; (3) MoE’s sensitivity to prompt tokens leading to overfitting with conventional learning methods. To mitigate these, we propose the Bypass Isolated Gating MoE (BIG-MoE) framework, featuring: (1) Fine-grained experts for enhanced detection of subtle spoofing cues; (2) An isolation gating mechanism to counteract input noise; (3) A novel differential convolutional prompt bypass enriching the gating network with critical local features, thereby improving perceptual capabilities. Extensive experiments on four benchmark datasets demonstrate significant generalization performance improvement in multimodal FAS task. The code is released at https://github.com/murInJ/BIG-MoE. Zitong Yu, Xun Lin, Weicheng Xie 0001, LinLin Shen |
ICASSP | 5 |
| 2025 | High-Fidelity Editable Portrait Synthesis with 3D GAN InversionabstractThe 3D generative adversarial network (GAN) inversion converts an image into 3D representation to attain high-fidelity reconstruction and facilitate realistic image manipulation within the 3D latent space. However, previous approaches face challenges regarding the trade-off between the reconstruction ability and editability. That is, reversing a real-world image to a low-dimensional latent code would inevitably lead to information loss, and achieving a near-perfect reconstruction using high-rate triplane representation often limits the ability to manipulate the image freely in the latent space. To address these issues, we propose a novel latent conditioning encoder-based framework with the alignment between the low-dimensional latent and high-dimensional triplane. A non-semantic guided editing strategy bridges the intrinsic relation between the latent condition and triplane generation, making it possible to edit the high-dimensional representation by latent manipulation. As a result, our method can achieve high-fidelity reconstruction and editing simultaneously by directly controlling the latent code. Experimental results demonstrate that our approach excels in reconstruction and editing quality compared to previous 3D inversion methods. Furthermore, our method can also edit even real faces with large poses and out-of-domain cases. Jindong Xie, Yupei Lin, Jinbao Wang 0001, Xianxu Hou, LinLin Shen |
ICASSP | 6 |
| 2025 | Dual Encoders for Diffusion-based Image InpaintingabstractCurrent diffusion-based inpainting models struggle to preserve unmasked regions or generate highly coherent content. Additionally, it is hard for them to generate meaningful content for 3D inpainting. To tackle these challenges, we design a plug-and-play branch that runs through the entire generation process to enhance existing models. Specifically, we utilize dual encoders - a Convolutional Neural Network (CNN) encoder and the pre-trained Variational AutoEncoder (VAE) encoder, to encode masked images. The latent code and the feature map from the dual encoders are fed to diffusion models simultaneously. In addition, we apply Zero-padded initialization to solve the problem of mode collapse caused by this branch. Experiments on BrushBench and EditBench demonstrate that models with our plug-and-play branch can improve the coherence of inpainting, and our model achieves new state-of-the-art results. Dezhi Zheng, Kaijun Deng, Jinbao Wang 0001, LinLin Shen |
ICASSP | 4 |
| 2025 | DAP-MAE: Domain-Adaptive Point Cloud Masked Autoencoder for Effective Cross-Domain Learning
Qiufu Li, LinLin Shen |
ICCV | 3 |
| 2025 | SynFER: Towards Boosting Facial Expression Recognition With Synthetic DataabstractFacial expression datasets remain limited in scale due to the subjectivity of annotations and the labor-intensive nature of data collection. This limitation poses a significant challenge for developing modern deep learning-based facial expression analysis models, particularly foundation models, that rely on large-scale data for optimal performance. To tackle the overarching and complex challenge, instead of introducing a new large-scale dataset, we introduce SynFER (Synthesis of Facial Expressions with Refined Control), a novel synthetic framework for synthesizing facial expression image data based on high-level textual descriptions as well as more fine-grained and precise control through facial action units. To ensure the quality and reliability of the synthetic data, we propose a semantic guidance technique to steer the generation process and a pseudo-label generator to help rectify the facial expression labels for the synthetic images. To demonstrate the generation fidelity and the effectiveness of the synthetic data from SynFER, we conduct extensive experiments on representation learning using both synthetic data and real-world data. Results validate the efficacy of our approach and the synthetic data. Notably, our approach achieves a 67.23% classification accuracy on AffectNet when training solely with synthetic data equivalent to the AffectNet training set size, which increases to 69.84% when scaling up to five times the original size. Code is available here. Xilin He, Xiaole Xian, Bing Li 0024, Muhammad Haris Khan, ZongYuan Ge, Weicheng Xie 0001, Siyang Song, LinLin Shen, Bernard Ghanem, Xiangyu Yue 0001 |
ICCV | 9 |
| 2025 | WSI-LLaVA: A Multimodal Large Language Model for Whole Slide ImageabstractRecent advancements in computational pathology have produced patch-level Multi-modal Large Language Models (MLLMs), but these models are limited by their inability to analyze whole slide images (WSIs) comprehensively and their tendency to bypass crucial morphological features that pathologists rely on for diagnosis. To address these challenges, we first introduce WSI-Bench, a large-scale morphology-aware benchmark containing 180k VQA pairs from 9,850 WSIs across 30 cancer types, designed to evaluate MLLMs' understanding of morphological characteristics crucial for accurate diagnosis. Building upon this benchmark, we present WSI-LLaVA, a novel framework for gigapixel WSI understanding that employs a three-stage training approach: WSI-text alignment, feature space alignment, and task-specific instruction tuning. To better assess model performance in pathological contexts, we develop two specialized WSI metrics: WSI-Precision and WSI-Relevance. Experimental results demonstrate that WSI-LLaVA outperforms existing models across all capability dimensions, with a significant improvement in morphological analysis, establishing a clear correlation between morphological understanding and diagnostic accuracy. Yuci Liang, Xinheng Lyu, Wenting Chen, Meidan Ding, Xiangjian He, Xiaohan Xing, Sen Yang 0006, LinLin Shen |
ICCV | 11 |
| 2025 | Enhancing Adversarial Transferability by Balancing Exploration and Exploitation with Gradient-Guided SamplingabstractAdversarial attacks present a critical challenge to deep neural networks' robustness, particularly in transfer scenarios across different model architectures. However, the transferability of adversarial attacks faces a fundamental dilemma between Exploitation (maximizing attack potency) and Exploration (enhancing cross-model generalization). Traditional momentum-based methods over-prioritize Exploitation, i.e., higher loss maxima for attack potency but weakened generalization (narrow loss surface). Conversely, recent methods with inner-iteration sampling over-prioritize Exploration, i.e., flatter loss surfaces for cross-model generalization but weakened attack potency (suboptimal local maxima). To resolve this dilemma, we propose a simple yet effective Gradient-Guided Sampling (GGS), which harmonizes both objectives through guiding sampling along the gradient ascent direction to improve both sampling efficiency and stability. Specifically, based on MI-FGSM, GGS introduces inner-iteration random sampling and guides the sampling direction using the gradient from the previous inner-iteration (the sampling's magnitude is determined by a random distribution). This mechanism encourages adversarial examples to reside in balanced regions with both flatness for cross-model generalization and higher local maxima for strong attack potency. Comprehensive experiments across multiple DNN architectures and multimodal large language models (MLLMs) demonstrate the superiority of our method over state-of-the-art transfer attacks. Code is made available at https://github.com/anuin-cat/GGS. Zenghao Niu, Weicheng Xie 0001, Siyang Song, Zitong Yu, Feng Liu 0013, LinLin Shen |
ICCV | 6 |
| 2025 | FineMotion: A Dataset and Benchmark with Both Spatial and Temporal Annotation for Fine-Grained Motion Generation and Editing
Bizhu Wu, Jinheng Xie, Meidan Ding, Zhe Kong, Jianfeng Ren, Ruibin Bai, Rong Qu, LinLin Shen |
ICCV | 8 |
| 2025 | DeeperForward: Enhanced Forward-Forward Training for Deeper and Better PerformanceabstractWhile backpropagation effectively trains models, it presents challenges related to bio-plausibility, resulting in high memory demands and limited parallelism. Recently, Hinton (2022) proposed the Forward-Forward (FF) algorithm for high-parallel local updates. FF leverages squared sums as the local update target, termed goodness, and decouples goodness by normalizing the vector length to extract new features. However, this design encounters issues with feature scaling and deactivated neurons, limiting its application mainly to shallow networks. This paper proposes a novel goodness design utilizing **layer normalization** and **mean goodness** to overcome these challenges, demonstrating performance improvements even in 17-layer CNNs. Experiments on CIFAR-10, MNIST, and Fashion-MNIST show significant advantages over existing FF-based algorithms, highlighting the potential of FF in deep models. Furthermore, the model parallel strategy is proposed to achieve highly efficient training based on the property of local updates. Yang Zhang 0012, Weizhao He, Jiajun Wen 0001, LinLin Shen, Weicheng Xie 0001 |
ICLR | 5 |
| 2025 | BCE vs. CE in Deep Feature LearningabstractWhen training classification models, it expects that the learned features are compact within classes, and can well separate different classes. As the dominant loss function for training classification models, minimizing cross-entropy (CE) loss maximizes the compactness and distinctiveness, i.e., reaching neural collapse (NC). The recent works show that binary CE (BCE) performs also well in multi-class tasks. In this paper, we compare BCE and CE in deep feature learning. For the first time, we prove that BCE can also maximize the intra-class compactness and inter-class distinctiveness when reaching its minimum, i.e., leading to NC. We point out that CE measures the relative values of decision scores in the model training, implicitly enhancing the feature properties by classifying samples one-by-one. In contrast, BCE measures the absolute values of decision scores and adjust the positive/negative decision scores across all samples to uniformly high/low levels. Meanwhile, the classifier biases in BCE present a substantial constraint on the decision scores to explicitly enhance the feature properties in the training. The experimental results are aligned with above analysis, and show that BCE could improve the classification and leads to better compactness and distinctiveness among sample features. The codes have be released. Qiufu Li, Huibin Xiao, LinLin Shen |
ICML | 3 |
| 2025 | Any-to-Any Vision-Language Model for Multimodal X-ray Imaging and Radiological Report GenerationabstractGenerative models have revolutionized Artificial Intelligence (AI), particularly in multimodal applications. However, adapting these models to the medical domain poses unique challenges due to the complexity of medical data and the stringent need for clinical accuracy. In this work, we introduce a framework specifically designed for multimodal medical data generation. By enabling the generation of multi-view chest X-rays and their associated clinical report, it bridges the gap between general-purpose vision-language models and the specialized requirements of healthcare. Leveraging the MIMIC-CXR dataset, the proposed framework shows superior performance in generating high-fidelity images and semantically coherent reports. Our quantitative evaluation reveals significant results in terms of FID and BLEU scores, showcasing the quality of the generated data. Notably, our framework achieves comparable or even superior performance compared to real data on downstream disease classification tasks, underlining its potential as a tool for medical research and diagnostics. This study highlights the importance of domain-specific adaptations in enhancing the relevance and utility of generative models for clinical applications, paving the way for future advancements in synthetic multimodal medical data generation. Daniele Molino, Francesco Di Feola, LinLin Shen, Paolo Soda, Valerio Guarrasi |
IJCNN | 3 |
| 2025 | SoCDev: SoC Design Automation through TCL Code Generation via LLM-Based AgentsabstractThe design of System-on-Chip (SoC) is a complex task that requires considerable effort from designers. Recent advancements in Language Model (LLM)-based agents have brought opportunities to tackle this challenge. In this paper, we introduce SocDev, the first SoC design multi-agent system based on LLM, which can automatically generate SoC designs by creating TCL command scripts for the Vivado Design Suite. Additionally, we have developed SoCDevBench, a test bench to evaluate the capability of different models in generating SoC designs. We conducted experiments using four state-of-the-art LLMs with the test bench, assessing the pass rate, FPGA resource utilization, and logs throughout the process. The ablation study also verify the efficiency of each component. Furthermore, we successfully produced four functional SoCs with microprocessor and operating systems, and validated them using FPGAs. This study opens a new avenue for the generation of SoC designs leveraging LLMs. LinLin Shen, Qiuming Luo |
IJCNN | 3 |
| 2025 | 🤖 WSI-Agents: A Collaborative Multi-agent System for Multi-modal Whole Slide Image Analysis
Xinheng Lyu, Yuci Liang, Wenting Chen, Meidan Ding, Guolin Huang, Daokun Zhang, Xiangjian He, LinLin Shen |
MICCAI (5) | 9 |
| 2025 | TRRG: Towards Truthful Radiology Report Generation With Cross-Modal Disease Clue Enhanced Large Language Models
Yue Sun 0001, Tao Tan 0002, Chao Hao, Yawen Cui, Xinqi Su, Weicheng Xie 0001, LinLin Shen, Zitong Yu |
MICCAI (7) | 8 |
| 2025 | Smooth Online Multiple Appropriate Facial Reaction GenerationabstractIn dyadic interactions, facial reactions are crucial for conveying an individuals' responses to their conversational partners. Individuals may exhibit varied but appropriate facial reactions (AFRs) when perceiving the same behavioral expression. Although some recent methods can already respond multiple appropriate facial reactions to the given human speaker behaviors, the AFRs generated by these methods often fail to adequately preserve crucial head motions, leading to visual jitter and unnatural transitions between generated AFR segments. In this paper, we propose a novel and generic PFLPosNet framework which addresses the aforementioned problems at both pre-processing and post-processing stages, where a new pose-aware face behavior localization method PFL is introduced to retain the head pose displacement information from the source data. In addition, the framework proposes a real-time head pose adjustment method, PosNet, to ensure continuity and smoothness in the visual output of the model when using data with correct head pose displacement. Experimental results demonstrate that our approach not only generates more coherent and natural facial reaction sequences but also significantly outperforms existing online MAFRG methods in terms of continuity and smoothness. Our code is made available at https://github.com/rainforcetime/PFLPosNet. Weicheng Xie 0001, Chunlin Yan, Siyang Song, Zitong Yu, LinLin Shen, Laizhong Cui |
ACM Multimedia | 5 |
| 2025 | GM-DF: Generalized Multi-Scenario Deepfake DetectionabstractRecent advances in face forgery detection have shown strong in-domain performance but often fail to generalize to out-of-distribution data, especially when confronted with unseen manipulation techniques or domain shifts (e.g., lighting conditions, camera noise). We propose a novel Mixture-of-Experts framework, termed GM-DF, that decouples domain-specific and domain-invariant features to tackle cross-domain face forgery detection. Our method builds upon a foundation model (CLIP) and incorporates three key modules: (1) Dataset-Embedding Generator that leverages lightweight expert layers and database-aware feature normalization to adaptively modulate features at a per-domain level, capturing idiosyncratic cues without overfitting; (2) Multi-Dataset Representation mechanism that fuses these expert embeddings using scaled dot-product attention and integrates a mask image modeling (MIM) task to amplify local forgery artifacts; (3) Meta-Domain-Embedding Optimizer, inspired by MAML, which alternates between domain-specific (inner-loop) and domain-invariant (outer-loop) updates to facilitate rapid adaptation on new domains. Additionally, inspired by [13] (Yossi Gandelsman, Alexei A Efros, and Jacob Steinhardt. 2024. Interpreting the second-order effects of neurons in clip. arXiv preprint arXiv:2406.04341 (2024)) we introduce second-order feature propagation in the intermediate layers of CLIP to enhance fine-grained artifact cues and propose domain-class disentangled prompts to flexibly encode multi-domain text representations. Together, these strategies enable GM-DF to learn robust, shared forgery cues while preserving essential domain nuances. Our extensive experiments on multiple cross-domain benchmarks demonstrate that GM-DF significantly outperforms state-of-the-art approaches in both detection accuracy and domain transferability, reducing reliance on superficial artifacts and improving generalization to unseen forgeries. Importantly, our design requires minimal overhead beyond standard CLIP, making GM-DF both effective and computationally efficient for real-world face forgery detection. Yingxin Lai, Hongyang Wang 0001, Xiangui Kang, Bin Li 0011, LinLin Shen, Zitong Yu |
ACM Multimedia | 6 |
| 2025 | Taming Anomalies with Down-Up Sampling Networks: Group Center Preserving Reconstruction for 3D Anomaly DetectionabstractReconstruction-based methods have demonstrated very promising results for 3D anomaly detection. However, these methods face great challenges in handling high-precision point clouds due to the large scale and complex structure. In this study, a Down-Up Sampling Networks (DUS-Net) is proposed to reconstruct high-precision point clouds for 3D anomaly detection by preserving the group center geometric structure. The DUS-Net first introduces a Noise Generation module to generate noisy patches, which facilitates the diversity of training data and strengthens the feature representation for reconstruction. Then, a Down-sampling Network (Down-Net) is developed to learn an anomaly-free center point cloud from patches with noise injection. Subsequently, an Up-sampling Network (Up-Net) is designed to reconstruct high-precision point clouds by fusing multi-scale up-sampling features. Our method leverages group centers for construction, enabling the preservation of geometric structure and providing a more precise point cloud. Extensive experiments demonstrate the effectiveness of our proposed method, achieving state-of-the-art (SOTA) performance, with an Object-level AUROC of 79.9% and 79.5% and a Point-level AUROC of 71.2% and 84.7% on the Real3D-AD and Anomaly-ShapeNet datasets, respectively. Hanzhe Liang, Jie Zhang 0090, Tao Dai 0001, LinLin Shen, Jinbao Wang 0001, Can Gao |
ACM Multimedia | 4 |
| 2025 | DisFaceRep: Representation Disentanglement for Co-occurring Facial Components in Weakly Supervised Face ParsingabstractFace parsing aims to segment facial images into key components such as eyes, lips, and eyebrows. While existing methods rely on dense pixel-level annotations, such annotations are expensive and labor-intensive to obtain. To reduce annotation cost, we introduce Weakly Supervised Face Parsing (WSFP), a new task setting that performs dense facial component segmentation using only weak supervision, such as image-level labels and natural language descriptions. WSFP introduces unique challenges due to the high co-occurrence and visual similarity of facial components, which lead to ambiguous activations and degraded parsing performance. To address this, we propose DisFaceRep, a representation disentanglement framework designed to separate co-occurring facial components through both explicit and implicit mechanisms. Specifically, we introduce a co-occurring component disentanglement strategy to explicitly reduce dataset-level bias, and a text-guided component disentanglement loss to guide component separation using language supervision implicitly. Extensive experiments on CelebAMask-HQ, LaPa, and Helen demonstrate the difficulty of WSFP and the effectiveness of DisFaceRep, which significantly outperforms existing weakly supervised semantic segmentation methods. The code will be released at https://github.com/CVI-SZU/DisFaceRep. Xianxu Hou, Meidan Ding, Junliang Chen 0002, Kaijun Deng, Jinheng Xie, LinLin Shen |
ACM Multimedia | 7 |
| 2025 | Unknown Pixel Mask Based Fine-tuning of 2D Inpainting Models for Unbounded 3D Scene Generation from a Single ImageabstractConventional 2D inpainting models are trained using masks confined to 2D scenarios, resulting in meaningless content when applied to 3D-specific masks. These 3D-specific masks, termed Unknown Pixels (UP) masks, represent unseen pixels from novel viewpoints that remain obscured in the original input image. Existing methods attempt to mitigate this issue by employing post-processing techniques to transform UP masks into 2D equivalents, frequently suffering from unnatural distortions. To address these issues, we investigate the efficacy of directly training 2D inpainting models with UP masks to circumvent such distortions. In this paper, we introduce a novel framework designed to generate unbounded 3D scenes from a single image, guided by textual descriptions. Our approach leverages fine-tuned inpainting models that iteratively reconstruct incomplete images originating from pure projection. The generated points are then seamlessly integrated into the original point cloud via pixel-wise depth alignment. Extensive evaluations demonstrate that our framework outperforms existing methods in scene quality, processing speed, and memory efficiency. Dezhi Zheng, Kaijun Deng, Xianxu Hou, Jinbao Wang 0001, LinLin Shen |
ACM Multimedia | 6 |
| 2025 | MedChain: Bridging the Gap Between LLM Agents and Clinical Practice with Interactive SequenceabstractClinical decision making (CDM) is a complex, dynamic process crucial to healthcare delivery, yet it remains a significant challenge for artificial intelligence systems. While Large Language Model (LLM)-based agents have been tested on general medical knowledge using licensing exams and knowledge question-answering tasks, their performance in the CDM in real-world scenarios is limited due to the lack of comprehensive benchmark that mirror actual medical practice. To address this gap, we present MedChain, a dataset of 12,163 clinical cases that covers five key stages of clinical workflow. MedChain distinguishes itself from existing benchmarks with three key features of real-world clinical practice: personalization, interactivity, and sequentiality. Further, to tackle real-world CDM challenges, we also propose MedChain-Agent, an AI system that integrates a feedback mechanism and a MedCase-RAG module to learn from previous cases and adapt its responses. MedChain-Agent demonstrates remarkable adaptability in gathering information dynamically and handling sequential clinical tasks, significantly outperforming existing approaches. The relevant dataset and code will be released upon acceptance of this paper. Jie Liu 0044, Wenxuan Wang 0001, Zizhan Ma, Guolin Huang, Yihang Su, Kao-Jung Chang, Haoliang Li, LinLin Shen, Michael R. Lyu, Wenting Chen |
NeurIPS | 8 |
| 2025 | OTMamba: Ophthalmology Image Translation Using Guidance-Controllable Mamba Diffusion Model
Huijie Deng, Yamei Lu, Zhuoru Wu, Chunhui Zou, Bowei Yuan, Xiaoling Luo 0001, LinLin Shen |
PRCV (13) | 8 |
| 2025 | Token-Level Contrastive Learning for Open-World Weakly-Supervised Object Localization
Rouyi Li, Zhaochuan Luo, LinLin Shen |
PRCV (16) | 4 |
| 2025 | SplatID: Real-Time Lossless 3D Gaussian Splatting with Feature ID Generation and Frame Filtering
Wenhui Ma, LinLin Shen, Jinbao Wang 0001 |
PRCV (10) | 3 |
| 2025 | LabelGS: Label-Aware 3D Gaussian Splatting for 3D Scene Segmentation
Dezhi Zheng, Lei Wang 0018, Liping xiang, Kaijun Deng, Xiaowen Fu, LinLin Shen, Jinbao Wang 0001 |
PRCV (10) | 10 |
| 2025 | Towards Robust Training via Gradient-Diversified BackpropagationabstractNeural networks are prone to be vulnerable to adversarial attacks and domain shifts. Adversarial-driven methods including adversarial training and adversarial augmentation, have been frequently proposed to improve the model's robustness against adversarial attacks and distribution-shifted samples. Nonetheless, recent research on adversarial attacks has cast a spotlight on the robustness lacuna against attacks targeted at deep semantic layers. Our analysis reveals that previous adversarial-driven methods tend to generate overpowering perturbations in deep semantic layers, leading to distortion of the training for these layers. This can be primarily attributed to the exclusive utilization of loss functions on the output layer for adversarial gradient generation. This inherent practice projects an excessive adversarial impact on the deep semantic layers, elevating the difficulty of training such layers. Therefore, from the standing point of relaxing the excessive perturbations in the deep semantic layer and diversifying the adversarial gradients to ensure robust training for deep semantic layers, this paper proposes a novel Stochastic Loss Integration Method (SLIM), which can be instantiated into the existing adversarial-driven methods in a plug-and-play manner. Experimental results across diverse tasks, including classification and segmentation, as well as various areas such as adversarial robustness and domain generalization, validate the effectiveness of our proposed method. Furthermore, we provide an in-depth analysis to offer a comprehensive understanding of layer-wise training involving various loss terms. Xilin He, Qinliang Lin, Weicheng Xie 0001, Muhammad Haris Khan, Siyang Song, LinLin Shen |
WACV | 7 |
| 2025 | Enhancing anomaly detection with few-shot fine-tuned long text-to-image models
Jiajia An, Junbin Lu, Zhuoqin Yang, Jinbao Wang 0001, LinLin Shen |
Eng. Appl. Artif. Intell. | 8 |
| 2025 | A codebook-driven approach for low-light image enhancementabstractLow-light image enhancement (LLIE) aims to improve low-illumination images. However, existing methods face two challenges: (1) uncertainty in restoration from diverse brightness degradations; (2) loss of texture and color information caused by noise suppression and light enhancement. In this paper, we propose a novel enhancement approach, CodeEnhance, by leveraging discrete codebook priors and image refinement to address these challenges. In particular, we reframe LLIE as learning an image-to-code mapping from low-light images to discrete codebook, which has been learned from high-quality images. To enhance this process, a Semantic Embedding Module (SEM) is introduced to integrate semantic information with low-level features, and a Codebook Shift (CS) mechanism, designed to adapt the pre-learned codebook to better suit the distinct characteristics of our low-light dataset. Additionally, we present an Interactive Feature Transformation (IFT) module to refine texture and color information during image reconstruction, allowing for interactive enhancement based on user preferences. Extensive experiments on both real-world and synthetic benchmarks demonstrate that the incorporation of prior knowledge and controllable information transfer significantly enhances LLIE performance in terms of quality and fidelity. The proposed CodeEnhance exhibits superior robustness to various degradations, including uneven illumination, noise, and color distortion. The code can be obtained from https://github.com/csxuwu/CodeEnhance or https://www.scholat.com/laizhihui.cn . Xu Wu 0001, Xianxu Hou, Zhihui Lai 0001, Jie Zhou 0009, Witold Pedrycz, LinLin Shen |
Eng. Appl. Artif. Intell. | 7 |
| 2025 | Weakly supervised bounding-box generation for camera-trap image based animal detectionabstractAbstract In ecology, deep learning is improving the performance of camera‐trap image based wild animal analysis. However, high labelling cost becomes a big challenge, as it requires involvement of huge human annotation. For example, the Snapshot Serengeti (SS) dataset contains over 900,000 images, while only 322,653 contains valid animals, 68,000 volunteers were recruited to provide image level labels such as species, the no. of animals and five behaviour attributes such as standing, resting and moving etc. In contrast, the Gold Standard SS Bounding‐Box Coordinates (GSBBC for short) contains only 4011 images for training of object detection algorithms, as the annotation of bounding‐box for animals in the image, is much more costive. Such a no. of training images, is obviously insufficient. To address this, the authors propose a method to generate bounding‐boxes for a larger dataset using limited manually labelled images. To achieve this, the authors first train a wild animal detector using a small dataset (e.g. GSBBC) that is manually labelled to locate animals in images; then apply this detector to a bigger dataset (e.g. SS) for bounding‐box generation; finally, we remove false detections according to the existing label information of the images. Experiments show that detector trained with images whose bounding‐boxes are generated using the proposal, outperformed the existing camera‐trap image based animal detection, in terms of mean average precision (mAP). Compared with the traditional data augmentation method, our method improved the mAP by 21.3% and 44.9% for rare species, also alleviating the long‐tail issue in data distribution. In addition, detectors trained with the proposed method also achieve promising results when applied to classification and counting tasks, which are commonly required in wildlife research. Puxuan Xie, Renwu Gao, Weizeng Lu, LinLin Shen |
IET Comput. Vis. | 4 |
| 2025 | CLIMS++: Cross Language Image Matching with Automatic Context Discovery for Weakly Supervised Semantic Segmentation
Jinheng Xie, Songhe Deng, Xianxu Hou, Zhaochuan Luo, LinLin Shen, Yawen Huang, Yefeng Zheng 0001, Zheng Shou 0001 |
Int. J. Comput. Vis. | 5 |
| 2025 | YOLOCS: Object detection based on dense channel compression for feature spatial solidification
Weisheng Li 0001, Yujuan Tan, LinLin Shen, Jing Yu 0026, Haojie Fu |
Knowl. Based Syst. | 4 |
| 2025 | MLPFormer: MLP-integrated transformer for colorectal histopathology whole slide image segmentation
Xuechen Li 0001, Yanfei Zuo, LinLin Shen |
Neural Comput. Appl. | 6 |
| 2025 | An efficient and effective pore matching method using ResCNN descriptor and local outliers
Feng Liu 0013, Qiuheng Wang, Yanfeng Xiao, LinLin Shen |
Pattern Recognit. | 4 |
| 2025 | SymGraphAU: Prior knowledge based symbolic graph for action unit recognition
Weicheng Xie 0001, Junliang Zhang, Siyang Song, LinLin Shen, Zitong Yu |
Pattern Recognit. | 5 |
| 2025 | Distilled transformers with locally enhanced global representations for face forgery detection
Qiufu Li, Zitong Yu, LinLin Shen |
Pattern Recognit. | 4 |
| 2025 | Binarized Internal Fingerprint Reconstruction From Optical Coherence Tomography Based on Image Region RegressionabstractInternal fingerprint reconstruction is critical for bridging traditional fingerprint recognition with Optical Coherence Tomography (OCT)-based techniques. However, current reconstructed internal fingerprints often suffer from low ridge-valley contrast, noise interference, and ridge adherence issues. Traditional fingerprint enhancement techniques address these challenges but involve reconstructing 3D OCT fingerprints into 2D internal fingerprints, followed by enhancement. This two-step approach leads to module inconsistencies and difficulties in parameter setting during the enhancement process. To overcome these limitations, we for the first time propose a novel method that directly reconstructs binarized internal fingerprints. The proposed method employs an image region regression module that directly treats ridge blocks within B-scan images as regional units for regression, yielding 1D feature vectors representing ridges and valleys. Additionally, leveraging the continuity of information between adjacent B-scan images, a window adjustment function is introduced to refine the regression values, ensuring more stable binarized internal fingerprints. Experiments were conducted on publicly available OCT fingerprint benchmark datasets to compare the minutiae extraction and matching performance. The binarized internal fingerprints obtained by the proposed method achieved the highest mean NFIQ2 score. Based on the NBIS software compared to existing OCT internal fingerprint reconstruction methods, the proposed method achieved the lowest Equal Error Rate (EER) of 0.78%. In addition, compared to traditional fingerprint enhancement methods, the proposed method attained the highest F1-score for minutiae extraction at 72.39%. It also achieved the lowest EER and represented a 37.1% reduction compared to the best existing result. Feng Liu 0013, Wenfeng Zeng, LinLin Shen |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2025 | MFCLIP: Multi-Modal Fine-Grained CLIP for Generalizable Diffusion Face Forgery DetectionabstractThe rapid development of photo-realistic face generation methods has raised significant concerns in society and academia, highlighting the urgent need for robust and generalizable face forgery detection (FFD) techniques. Although existing approaches mainly capture face forgery patterns using image modality, other modalities like fine-grained noises and texts are not fully explored, which limits the generalization capability of the model. In addition, most FFD methods tend to identify facial images generated by GAN, but struggle to detect unseen diffusion-synthesized ones. To address the limitations, we aim to leverage the cutting-edge foundation model, contrastive language-image pre-training (CLIP), to achieve generalizable diffusion face forgery detection (DFFD). In this paper, we propose a novel multi-modal fine-grained CLIP (MFCLIP) model, which mines comprehensive and fine-grained forgery traces across image-noise modalities via language-guided face forgery representation learning, to facilitate the advancement of DFFD. Specifically, we devise a fine-grained language encoder (FLE) that extracts fine global language features from hierarchical text prompts. We design a multi-modal vision encoder (MVE) to capture global image forgery embeddings as well as fine-grained noise forgery patterns extracted from the richest patch, and integrate them to mine general visual forgery traces. Moreover, we build an innovative plug-and-play sample pair attention (SPA) method to emphasize relevant negative pairs and suppress irrelevant ones, allowing cross-modality sample pairs to conduct more flexible alignment. Extensive experiments and visualizations show that our model outperforms the state of the arts on different settings like cross-generator, cross-forgery, and cross-dataset evaluations. Our code will be available at https://github.com/Jenine-321/MFCLIP. Tianyi Wang 0006, Zitong Yu, Zan Gao 0001, LinLin Shen, Shengyong Chen |
IEEE Trans. Inf. Forensics Secur. | 5 |
| 2025 | READ3D-Net: Residual Autoencoder and GAN-Based 3-D Convolutional Network for Anomaly DetectionabstractVideo anomaly detection (VAD) is of great importance for a variety of real-time applications in video surveillance. Most deep learning-based anomaly detection algorithms adopt a one-class learning scheme to train a classifier using only normal data to distinguish between normal and abnormal events during the test phase. However, these methods, whether they are reconstruction or prediction models, commonly face the challenge of the model’s overly strong representation capability, which leads to excessive fitting of abnormal events and thus limits the performance of the model in diverse scenarios. To address these challenges, this work develops a novel residual autoencoder and generative adversarial network-based 3-D convolutional network, called READ3D-net, for anomaly detection. An adaptive multimodal pseudoanomaly generator is developed to simulate and generate diverse pseudoanomalies, aiming to enhance the model’s ability to extract the features with regard to “abnormal” behaviors while reducing the interference of background on the detection performance. In addition, residual structures are incorporated into the design of a reconstruction-based autoencoder model to enhance its feature extraction and discriminative capabilities. To further improve the model’s reconstruction ability on normal data, a dynamic generative adversarial strategy is proposed for effective feature learning. Extensive experiments conducted on public benchmark datasets demonstrate that the proposed model is more competitive than the state-of-the-art methods in video anomaly detection tasks, fully validating the effectiveness and practicality of the proposed approach. Yuwu Lu, Yinsheng Liu, Jiajun Wen 0001, Yang Zhang 0012, Yingyi Liang, Zhihui Lai 0001, LinLin Shen |
IEEE Trans. Ind. Informatics | 7 |
| 2025 | A Wavelet-Guided Deep Unfolding Network for Single Image Reflection RemovalabstractRemoving unwanted reflections from images is a fundamental yet challenging problem in low-level computer vision. Recent deep learning-based Single Image Reflection Removal (SIRR) methods have made significant progress. However, separating reflections from transmission content remains difficult, particularly in complex scenes where the two exhibit high visual similarity. Upon careful analysis, we find that reflections predominantly reside in the high-frequency components of an image. These reflections tend to distort fine details in the high-frequency range, while the low-frequency information remains relatively less affected. This observation motivates us to explore a frequency-aware approach for SIRR by leveraging the Discrete Wavelet Transform (DWT). The wavelet decomposition enables us to distinguish and isolate reflective artifacts in the frequency domain while preserving the transmission information. Building on this insight, we propose a novel Wavelet-guided Deep Unfolding Network (WDUNet) that leverages the strengths of wavelet decomposition and deep unfolding techniques to improve interpretability and generalization in SIRR. Specifically, we formulate an optimization-based reflection removal model using DWT and convolutional dictionaries. The proposed model is optimized via a proximal gradient algorithm and then unfolded into a neural network architecture, where all parameters are learned end-to-end during training. By combining wavelet domain analysis with deep unfolding, WDUNet enhances both the interpretability and generalization of SIRR methods. Additionally, we design and integrate the Low-frequency Parameter Estimation Module (LPEM) and High-frequency Parameter Estimation Module (HPEM) modules into WDUNet, allowing the network to automatically learn and optimize the models' hyperparameters. Extensive experiments conducted on four benchmark datasets demonstrate that WDUNet consistently outperforms existing state-of-the-art methods in both objective evaluation metrics and subjective visual quality. Qiufu Li, Xu Wu 0001, Nan Mu, LinLin Shen |
IEEE Trans. Image Process. | 6 |
| 2025 | Heterogeneous Domain Adaptation via Correlative and Discriminative Feature LearningabstractHeterogeneous domain adaptation seeks to learn an effective classifier or regression model for unlabeled target samples by using the well-labeled source samples but residing in different feature spaces and lying different distributions. Most recent works have concentrated on learning domain-invariant feature representations to minimize the distribution divergence via target pseudo-labels. However, two critical issues need to be further explored: 1) new feature representations should be not only domain-invariant but also category-correlative and discriminative and 2) alleviating the negative transfer caused by the incorrect pseudo-labeling target samples could boost the adaptation performance during the iterative learning process. To address these issues, in this paper, we put forward a novel heterogeneous domain adaptation method to learn category-correlative and discriminative representations, referred to as correlative and discriminative feature learning (CDFL). Specifically, CDFL aims to learn a feature space where class-specific feature correlations between the source and target domains are maximized, the divergences of marginal and conditional distribution between the source and target domains are minimized, and the distances of inter-class distribution are forced to be maximized to ensure the discriminative ability. Meanwhile, a selective pseudo-labeling procedure based on the correlation coefficient and classifier prediction is introduced to boost class-specific feature correlation and discriminative distribution alignment in an iteration way. Extensive experiments certify that CDFL outperforms the State-of-the-Art algorithms on five standard benchmarks. Yuwu Lu, Dewei Lin, LinLin Shen, Yicong Zhou, Jiahui Pan 0003 |
IEEE Trans. Multim. | 3 |
| 2025 | Progressive Pseudo Labeling for Multi-Dataset Detection Over Unified Label SpaceabstractExisting multi-dataset detection works mainly focus on the performance of detector on each of the datasets, with different label spaces. However, in real-world applications, a unified label space across multiple datasets is usually required. To address such a gap, we propose a progressive pseudo labeling (PPL) approach to detect objects across different datasets, over a unified label space. Specifically, we employ the widely used architecture of teacher-student model pair to jointly refine pseudo labels and train the unified object detector. The student model learns from both annotated labels and pseudo labels from the teacher model, which is updated by the exponential moving average (EMA) of the student. Three modules, i.e. Entropy-guided Adaptive Threshold (EAT), Global Classification Module (GCM) and Scene-Aware Fusion (SAF) strategy, are proposed to handle the noise of pseudo labels and fit the overall distribution. Extensive experiments are conducted on different multi-dataset benchmarks. The results demonstrate that our proposed method significantly outperforms the State-of-the-Art and is even comparable with supervised methods trained using annotations of all labels. Kai Ye 0004, Zepeng Huang, Yilei Xiong, Jinheng Xie, LinLin Shen |
IEEE Trans. Multim. | 6 |
| 2025 | Frequency Restoration and Modality Enforcement towards Resisting-corruption Multimodal Sentiment AnalysisabstractFor Multimodal Sentiment Analysis (MSA), previous methods concentrate on designing sophisticated fusion strategies and performing representation learning across heterogeneous modalities, aiming to leverage multimodal signals to detect human sentiment. However, these approaches fail to address the long-standing issue of corrupted modal details in videos, which may be caused by the challenge of the excessive loss of emotionally relevant semantics resulted from the degradation of detailed information. In this work, we aim to improve the robustness capacity of resisting corruption in MSA, by introducing a Hierarchical Frequency Restoration and Adaptive Modality Enforcement (HFR-AME) approach. The HFR-AME progressively recovers blurred detailed cues in each modality while enhancing the discriminative power of modal representations. Specifically, to reconstruct distinct frequency band features, we propose to equip the HFR module with a key component called the Frequency Multimodal UNet (FM-UNet), so as to utilize complementary modal features as conditions. This meticulous restoration process, performed from low to high frequency, facilitates the comprehensive recovery of intricate details. Meanwhile, to adaptively integrate these diverse frequency features, we introduce the AME module to enhance the beneficial modal frequencies while suppressing irrelevant ones, with the goal of strengthening the restored modal representations. Extensive experiments show our HFR-AME outperforms state-of-the-art methods on the CMU-MOSI and CMU-MOSEI datasets, improving 7-class accuracy by 0.5% and 0.6%, respectively. Further analysis also confirms its cross-lingual generalization and competitive computational efficiency. Our code is made available at https://github.com/nianhua20/HFR-AME . Weicheng Xie 0001, Haijian Liang, Zenghao Niu, Xianxu Hou, Siyang Song, Zitong Yu, LinLin Shen |
ACM Trans. Multim. Comput. Commun. Appl. | 7 |
| 2025 | ReactFace: Online Multiple Appropriate Facial Reaction Generation in Dyadic InteractionsabstractIn dyadic interaction, predicting the listener's facial reactions is challenging as different reactions could be appropriate in response to the same speaker's behaviour. Previous approaches predominantly treated this task as an interpolation or fitting problem, emphasizing deterministic outcomes but ignoring the diversity and uncertainty of human facial reactions. Furthermore, these methods often failed to model short-range and long-range dependencies within the interaction context, leading to issues in the synchrony and appropriateness of the generated facial reactions. To address these limitations, this paper reformulates the task as an extrapolation or prediction problem, and proposes an novel framework (called ReactFace) to generate multiple different but appropriate facial reactions from a speaker behaviour rather than merely replicating the corresponding listener facial behaviours. Our ReactFace generates multiple different but appropriate photo-realistic human facial reactions by: (i) learning an appropriate facial reaction distribution representing multiple different but appropriate facial reactions; and (ii) synchronizing the generated facial reactions with the speaker verbal and non-verbal behaviours at each time stamp, resulting in realistic 2D facial reaction sequences. Experimental results demonstrate the effectiveness of our approach in generating multiple diverse, synchronized, and appropriate facial reactions from each speaker's behaviour. The quality of the generated facial reactions is intimately tied to the speaker's speech and facial expressions, achieved through our novel speaker-listener interaction modules. Siyang Song, Weicheng Xie 0001, Micol Spitale, ZongYuan Ge, LinLin Shen, Hatice Gunes |
IEEE Trans. Vis. Comput. Graph. | 6 |
| 2024 | Boosting Adversarial Transferability across Model Genus by Deformation-Constrained WarpingabstractAdversarial examples generated by a surrogate model typically exhibit limited transferability to unknown target systems. To address this problem, many transferability enhancement approaches (e.g., input transformation and model augmentation) have been proposed. However, they show poor performances in attacking systems having different model genera from the surrogate model. In this paper, we propose a novel and generic attacking strategy, called Deformation-Constrained Warping Attack (DeCoWA), that can be effectively applied to cross model genus attack. Specifically, DeCoWA firstly augments input examples via an elastic deformation, namely Deformation-Constrained Warping (DeCoW), to obtain rich local details of the augmented input. To avoid severe distortion of global semantics led by random deformation, DeCoW further constrains the strength and direction of the warping transformation by a novel adaptive control strategy. Extensive experiments demonstrate that the transferable examples crafted by our DeCoWA on CNN surrogates can significantly hinder the performance of Transformers (and vice versa) on various tasks, including image classification, video action recognition, and audio recognition. Code is made available at https://github.com/LinQinLiang/DeCoWA. Qinliang Lin, Zenghao Niu, Xilin He, Weicheng Xie 0001, Yuanbo Hou, LinLin Shen, Siyang Song |
AAAI | 7 |
| 2024 | Fine-Grained Image-Text Alignment in Medical Imaging Enables Explainable Cyclic Image-Report GenerationabstractFine-grained vision-language models (VLM) have been widely used for inter-modality local alignment between the predefined fixed patches and textual words. However, in medical analysis, lesions exhibit varying sizes and positions, and using fixed patches may cause incomplete representations of lesions. Moreover, these methods provide explainability by using heatmaps to show the general image areas potentially associated with texts rather than specific regions, making their explanations not explicit and specific enough. To address these issues, we propose a novel Adaptive patch-word Matching (AdaMatch) model to correlate chest X-ray (CXR) image regions with words in medical reports and apply it to CXR-report generation to provide explainability for the generation process. AdaMatch exploits the fine-grained relation between adaptive patches and words to provide explanations of specific image regions with corresponding words. To capture the abnormal regions of varying sizes and positions, we introduce an Adaptive Patch extraction (AdaPatch) module to acquire adaptive patches for these regions adaptively. Aiming to provide explicit explainability for the CXR-report generation task, we propose an AdaMatch-based bidirectional LLM for Cyclic CXR-report generation (AdaMatch-Cyclic). It employs AdaMatch to obtain the keywords for CXR images and 'keypatches' for medical reports as hints to guide CXR-report generation. Extensive experiments on two publicly available CXR datasets validate the effectiveness of our method and its superior performance over existing methods. © 2024 Association for Computational Linguistics. Wenting Chen, LinLin Shen, Jiebo Luo 0001, Xiang Li 0001, Yixuan Yuan |
ACL (1) | 2 |
| 2024 | Multi-Scale Dynamic and Hierarchical Relationship Modeling for Facial Action Units RecognitionabstractHuman facial action units (AUs) are mutually related in a hierarchical manner, as not only they are associated with each other in both spatial and temporal domains but also AUs located in the same/close facial regions show stronger relationships than those of different facial regions. While none of existing approach thoroughly model such hi-erarchical inter-dependencies among AUs, this paper proposes to comprehensively model multi-scale AU-related dynamic and hierarchical spatiotemporal relationship among AUs for their occurrences recognition. Specifically, we first propose a novel multi-scale temporal differencing network with an adaptive weighting block to explicitly capture facial dynamics across frames at different spatial scales, which specifically considers the heterogeneity of range and mag-nitude in different AUs' activation. Then, a two-stage strategy is introduced to hierarchically model the relationship among AUs based on their spatial distribution (i.e., local and cross-region AU relationship modelling). Experimental results achieved on BP4D and DISFA show that our approach is the new state-of-the-art in the field of AU occurrence recognition. Our code is publicly available at https://github.com/CVI-SZU/MDHR. Zihan Wang 0005, Siyang Song, Songhe Deng, Weicheng Xie 0001, LinLin Shen |
CVPR | 6 |
| 2024 | APSeg: Auto-Prompt Network for Cross-Domain Few-Shot Semantic SegmentationabstractFew-shot semantic segmentation (FSS) endeavors to segment unseen classes with only a few labeled samples. Current FSS methods are commonly built on the assumption that their training and application scenarios share similar domains, and their performances degrade significantly while applied to a distinct domain. To this end, we propose to leverage the cutting-edge foundation model, the segment Anything Model (SAM), for generalization enhancement. The SAM however performs unsatisfactorily on domains that are distinct from its training data, which primarily comprise natural scene images, and it does not support automatic segmentation of specific semantics due to its interactive prompting mechanism. In our work, we introduce APSeg, a novel auto-prompt network for cross-domain few-shot semantic segmentation (CD-FSS), which is designed to be auto-prompted for guiding cross-domain segmentation. Specifically, we propose a Dual Prototype Anchor Transformation (DPAT) module that fuses pseudo query prototypes extracted based on cycle-consistency with support prototypes, allowing features to be transformed into a more stable domain-agnostic space. Additionally, a Meta Prompt (MPG) module is introduced to automatically generate prompt embeddings, eliminating the need for manual visual prompts. We build an efficient model which can be applied directly to target domains without fine-tuning. Extensive experiments on four cross-domain datasets show that our model outperforms the state-of-the-art CD-FSS method by 5.24% and 3.10% in average accuracy on 1-shot and 5-shot settings, respectively. Weizhao He, Yang Zhang 0012, LinLin Shen, Songhe Deng |
CVPR | 4 |
| 2024 | Tune-an-Ellipse: CLIP Has Potential to Find what you WantabstractVisual prompting of large vision language models such as CLIP exhibits intriguing zero-shot capabilities. A manually drawn red circle, commonly used for highlighting, can guide CLIP's attention to the surrounding region, to identify specific objects within an image. Without precise object proposals, however, it is insufficient for localization. Our novel, simple yet effective approach, i.e., Differentiable Visual Prompting, enables CLIP to zero-shot localize: given an image and a text prompt describing an object, we first pick a rendered ellipse from uniformly distributed anchor ellipses on the image grid via visual prompting, then use three loss functions to tune the ellipse coefficients to encap-sulate the target region gradually. This yields promising ex-perimental results for referring expression comprehension without precisely specified object proposals. In addition, we systematically present the limitations of visual prompting inherent in CLIP and discuss potential solutions. Jinheng Xie, Songhe Deng, Bing Li 0024, Yawen Huang, Yefeng Zheng 0001, Jürgen Schmidhuber, Bernard Ghanem, LinLin Shen, Zheng Shou 0001 |
CVPR | 9 |
| 2024 | MTaDCS: Moving Trace and Feature Density-Based Confidence Sample Selection Under Label Noise
Qingzheng Huang, Xilin He, Xiaole Xian, Qinliang Lin, Weicheng Xie 0001, Siyang Song, LinLin Shen, Zitong Yu |
ECCV (71) | 7 |
| 2024 | PointFaceFormer: Local and Global Attention Based Transformer for 3D Point Cloud Face RecognitionabstractExisting 3D point cloud-based facial recognition struggles to fully leverage both global and local information inherent in the 3D point cloud data. In this paper, we introduce the PointFaceFormer, the first Transformer model designed for 3D point cloud face recognition. It incorporates an attention mechanism based on dot product and cosine functions to construct a similarity Transformer architecture, which effectively extracts both local and global features from the point cloud data. Experimental results demonstrate that PointFaceFormer achieves a recognition accuracy of 89.08% and a verification accuracy of 76.93% on the large-scale facial point cloud dataset Lock3DFace, which is a new state-of-the-art in 3D face recognition. Furthermore, PointFaceFormer exhibits excellent generalization performance on cross-quality datasets. Additionally, we validate the effectiveness of the attention mechanism through ablation experiments, which justify the effectiveness of the proposed modules. Qiufu Li, Gui Wang, LinLin Shen |
FG | 4 |
| 2024 | Expression-Aware Masking and Progressive Decoupling for Cross-Database Facial Expression RecognitionabstractCross-database facial expression recognition (CD-FER) has been widely studied due to its promising applicability in real-life situations, while the generalization performance is the main concern in this task. For improving cross-database generalization, current works frequently resort to masked auto encoder (MAE) to learn the expression representation in an unsupervised manner, and disentanglement of expression and domain features. (i) For MAE, current algorithms mainly employ random masking, and leverage the reconstruction of these masked regions to enable networks to learn the expression representation. However, these masked regions are expression-irrelevant, can not well reflect the characteristics of expression, thus are not efficient enough in representation learning. To this end, we propose an expression-aware masking in MAE to improve the learning efficiency of expression representation, by guiding MAE to mask out expression-aware regions during training. (ii) For disentanglement of expression and domain features, current algorithms realize it mainly in the deep layers. However, the coupling of these features in the shallow layers are rarely concerned, which may largely affect the disentanglement performance in deep layers. Thus, we propose a progressive decoupler to disentangle these features block by block, to use the feature disentanglement in shallow layers to facilitate that in deep layers. Extensive quantitative and qualitative results on multiple expression datasets show that our method can largely outperform the state of the arts in terms of cross-database generalization performance. Xiaole Xian, Zihan Wang 0005, Weicheng Xie 0001, LinLin Shen |
FG | 5 |
| 2024 | Scale-Free And Task-Generic Attack: Generating Photo-Realistic Adversarial Patterns With Patch Quilting GeneratorabstractRecent CNN generator-based attack approaches can synthe-size unrestricted and semantically meaningful entities to the image, which are able to improve the transferability and robustness. However, such methods attack images by either synthesizing local adversarial entities, which are only suitable for attacking specific contents, or performing global attacks, which are only applicable to a specific image scale. In this paper, we propose a novel Patch Quilting Generative Adversarial Networks (PQ-GAN) to learn the first scale-free CNN generator that can be applied to attack images with arbitrary scales for various computer vision tasks. The principal investigation on transferability of the generated adversarial examples, robustness to defense frameworks, and visual quality assessment show that the proposed PQG-based attack framework outperforms the other nine state-of-the-art adversarial attack approaches when attacking the neural networks trained on two standard evaluation datasets (i.e., ImageNet and CityScapes). Our code is made available at https://github.com/XiangboGaoBarry/PQAttack. Xiangbo Gao, Qinliang Lin, Weicheng Xie 0001, LinLin Shen, Keerthy Kusumam, Siyang Song |
ICASSP | 5 |
| 2024 | Dynamic Data Sampler for Cross-Language Transfer Learning in Large Language ModelsabstractLarge Language Models (LLMs) have gained significant attention in the field of natural language processing (NLP) due to their wide range of applications. However, training LLMs for languages other than English poses significant challenges, due to the difficulty in acquiring large-scale corpus and the requisite computing resources. In this paper, we propose ChatFlow, a cross-language transfer-based LLM, to address these challenges and train large Chinese language models in a cost-effective manner. We employ a mix of Chinese, English, and parallel corpus to continuously train the LLaMA2 model, aiming to align cross-language representations and facilitate the knowledge transfer specifically to the Chinese language model. In addition, we use a dynamic data sampler to progressively transition the model from unsupervised pre-training to supervised fine-tuning. Experimental results demonstrate that our approach accelerates model convergence and achieves superior performance. We evaluate ChatFlow on popular Chinese and English benchmarks, the results indicate that it outperforms other Chinese models post-trained on LLaMA-2-7B. Yudong Li 0001, Zhe Zhao 0006, LinLin Shen, Cheng Hou, Xianxu Hou |
ICASSP | 5 |
| 2024 | Circular Decomposition and Cross-Modal Recombination for Multimodal Sentiment AnalysisabstractMultimodal Sentiment Analysis is a burgeoning research area, leveraging various modalities to predict the sentiment score. Nevertheless, previous studies have disregarded the impact of noise interference on specific modal sentiments during video recording, thereby compromising the accuracy of sentiment prediction. In this paper, we propose the Guided Circular Decomposition and Cross-Modal Recombination (GCD-CMR) model, which aims to eliminate contaminated sentiment features in a fine-grained way. To achieve this, we utilize tailored global information specific to each modality to guide the circular decomposing process in the GCD module, to produce a set of sentiment prototypes. Subsequently, in the CMR module, we align cross-modal sentiment prototypes and remove the contaminated prototypes for recombination. Experimental results on two publicly available datasets demonstrate that our model surpasses state-of-the-art models, confirming the effectiveness of our proposed method. We release the code at: https://github.com/nianhua20/GCD-CMR. Haijian Liang, Weicheng Xie 0001, Xilin He, Siyang Song, LinLin Shen |
ICASSP | 5 |
| 2024 | MERG: Multi-Dimensional Edge Representation Generation Layer for Graph Neural NetworksabstractEdges are essential in describing relationships among nodes. While existing graphs frequently use a single-value edge to describe association between each pair of node vectors, crucial relationships may be disregarded if they are not linearly correlated, which may limit graph analysis performance. Although some recent Graph Neural Networks (GNNs) can process graphs containing multi-dimensional edge features, they cannot convert single-value edge graphs to multi-dimensional edge graphs during propagation. This paper proposes a generic Multi-dimensional Edge Representation Generation (MERG) layer that can be inserted into any GNNs for heterogeneous graph analysis. It assigns multi-dimensional edge features for the input single-value edge graph, describing multiple task-specific and global context-aware relationship cues between each connected node pair. Results on eight graph benchmark datasets demonstrate that inserting the MERG layer into widely-used GNNs (e.g., GatedGCN and GAT) leads to major performance improvements, resulting in state-of-the-art (SOTA) results on seven out of eight evaluated datasets. Our code is publicly available at1. YuXin Song 0001, Aaron S. Jackson, Xi Jia, Weicheng Xie 0001, LinLin Shen, Hatice Gunes, Siyang Song |
ICASSP | 6 |
| 2024 | CLIP-Guided Bidirectional Prompt and Semantic Supervision for Dynamic Facial Expression RecognitionabstractDue to the insufficient semantic information supervision in existing works for dynamic facial expression recognition (DFER), videos with similar facial changes but different expressions may be easily confused. Thanks to the potential textual information for semantic supervision, contrastive language-image pretraining (CLIP) model provides a new direction for DFER. However, pre-trained CLIP based on image-text pairs has difficulty in capturing temporal features in the video domain. Therefore, we propose a novel visual language model that captures and aggregates dynamic features of expressions in semantic supervision via Inter-Frame Interaction Transformer (Inter-FIT) and Multi-Scale Temporal Aggregation (MSTA). Furthermore, though prompt learning is often used in CLIP to enhance semantic supervision, previous studies have only focused on the role of textual prompts, ignoring the importance of visual prompts in facilitating the relationality between the two. Therefore, we designed a Bidirectional Enhanced Prompt (BiEhPro) to facilitate the learning of this relationality between text and visual cues in enhancing semantic supervision. Extensive experiments and ablation studies on three benchmark datasets, i.e., DFEW, FERV39K, and MAFW, validate the effectiveness of our modules and algorithm. Code is publicly available at https://github.com/JunLiangZ/CLIP-Guided-DFER. Junliang Zhang, Xiaole Xian, Weicheng Xie 0001, LinLin Shen, Siyang Song |
IJCB | 6 |
| 2024 | WiNet: Wavelet-Based Incremental Learning for Efficient Medical Image Registration
Xinxing Cheng, Xi Jia, Wenqi Lu 0001, Qiufu Li, LinLin Shen, Alexander Krull, Jinming Duan 0001 |
MICCAI (2) | 5 |
| 2024 | 💎 GEM: Context-Aware Gaze EstiMation with Visual Search Behavior Matching for Chest Radiograph
Shaonan Liu, Wenting Chen, Jie Liu 0044, Xiaoling Luo 0001, LinLin Shen |
MICCAI (1) | 5 |
| 2024 | Multi-Dataset Multi-Task Learning for COVID-19 Prognosis
Filippo Ruffini, Lorenzo Tronchin, Zhuoru Wu, Wenting Chen, Paolo Soda, LinLin Shen, Valerio Guarrasi |
MICCAI (12) | 6 |
| 2024 | Simplify Implant Depth Prediction as Video Grounding: A Texture Perceive Implant Depth Prediction Network
Xinquan Yang, Xiaoling Luo 0001, Leilei Zeng, Yudi Zhang 0005, LinLin Shen, Yongqiang Deng |
MICCAI (5) | 6 |
| 2024 | A New Dataset and Baseline Model for Rectal Cancer Risk Assessment in Endoscopic Ultrasound Videos
Jiansong Zhang 0005, Peizhong Liu, LinLin Shen |
MICCAI (3) | 4 |
| 2024 | FLIP-80M: 80 Million Visual-Linguistic Pairs for Facial Language-Image Pre-TrainingabstractWhile significant progress has been made in multi-modal learning driven by large-scale image-text datasets, there is still a noticeable gap in the availability of such datasets within the facial domain. To facilitate and advance the field of facial representation learning, we present FLIP-80M, a large-scale visual-linguistic dataset comprising over 80 million face images paired with text descriptions. FLIP-80M is constructed by leveraging the large openly available image-text-pair dataset LAION-5B and a mixed-method approach to filter face-related pairs from both visual and linguistic perspectives. Our curation process involves face detection, face caption classification, text de-noising, and synthesis-based image augmentation. As a result, FLIP-80M stands as the largest face-text dataset to date. To evaluate the potential of our dataset, we fine-tune the CLIP model using the proposed FLIP-80M, to create FLIP (Facial Language-Image Pretraining) and assess its representation capabilities across various downstream tasks. Our experiments demonstrate that our FLIP model achieves state-of-the-art results in a range of face analysis tasks, including face parsing, face alignment, and face attribute classification. The dataset and models are available at https://github.com/ydli-ai/FLIP. Yudong Li 0001, Xianxu Hou, Dezhi Zheng, LinLin Shen, Zhe Zhao 0006 |
ACM Multimedia | 4 |
| 2024 | PerFRDiff: Personalised Weight Editing for Multiple Appropriate Facial Reaction GenerationabstractHuman facial reactions play crucial roles in dyadic human-human interactions, where individuals (i.e., listeners) with varying cognitive process styles may display different but appropriate facial reactions in response to an identical behaviour expressed by their conversational partners. While several existing facial reaction generation approaches are capable of generating multiple appropriate facial reactions (AFRs) in response to each given human behaviour, they fail to take human's personalised cognitive process in AFRs generation. In this paper, we propose the first online personalised multiple appropriate facial reaction generation (MAFRG) approach which learns a unique personalised cognitive style from the target human listener's previous facial behaviours and represents it as a set of network weight shifts. These personalised weight shifts are then applied to edit the weights of a pre-trained generic MAFRG model, allowing the obtained personalised model to naturally mimic the target human listener's cognitive process in its reasoning for multiple AFRs generations. Experimental results show that our approach not only largely outperformed all existing approaches in generating more appropriate and diverse generic AFRs, but also serves as the first reliable personalised MAFRG solution. Our code is made available at https://github.com/xk0720/PerFRDiff. Hengde Zhu, Xiangyu Kong 0001, Weicheng Xie 0001, LinLin Shen, Lu Liu 0001, Hatice Gunes, Siyang Song |
ACM Multimedia | 5 |
| 2024 | Towards High-resolution 3D Anomaly Detection via Group-Level Feature Contrastive LearningabstractHigh-resolution point clouds (HRPCD) anomaly detection (AD) plays a critical role in precision machining and high-end equipment manufacturing. Despite considerable 3D-AD methods that have been proposed recently, they still cannot meet the requirements of the HRPCD-AD task. There are several challenges: i) It is difficult to directly capture HRPCD information due to large amounts of points at the sample level; ii) The advanced transformer-based methods usually obtain anisotropic features, leading to degradation of the representation; iii) The proportion of abnormal areas is very small, which makes it difficult to characterize. To address these challenges, we propose a novel group-level feature-based network, called Group3AD, which has a significantly efficient representation ability. First, we design an Intercluster Uniformity Network (IUN) to present the mapping of different groups in the feature space as several clusters, and obtain a more uniform distribution between clusters representing different parts of the point clouds in the feature space. Then, an Intracluster Alignment Network (IAN) is designed to encourage groups within the cluster to be distributed tightly in the feature space. In addition, we propose an Adaptive Group-Center Selection (AGCS) based on geometric information to improve the pixel density of potential anomalous regions during inference. The experimental results verify the effectiveness of our proposed Group3AD, which surpasses Reg3D-AD by the margin of 5% in terms of object-level AUROC on Real3D-AD. We provide the code and supplementary information on our website: https://github.com/M-3LAB/Group3AD. Hongze Zhu, Guoyang Xie, Chengbin Hou, Tao Dai 0001, Can Gao, Jinbao Wang 0001, LinLin Shen |
ACM Multimedia | 7 |
| 2024 | Towards Combating Frequency Simplicity-biased Learning for Domain GeneralizationabstractDomain generalization methods aim to learn transferable knowledge from source domains that can generalize well to unseen target domains.
Recent studies show that neural networks frequently suffer from a simplicity-biased learning behavior which leads to over-reliance on specific frequency sets, namely as frequency shortcuts, instead of semantic information, resulting in poor generalization performance.
Despite previous data augmentation techniques successfully enhancing generalization performances, they intend to apply more frequency shortcuts, thereby causing hallucinations of generalization improvement.
In this paper, we aim to prevent such learning behavior of applying frequency shortcuts from a data-driven perspective. Given the theoretical justification of models' biased learning behavior on different spatial frequency components, which is based on the dataset frequency properties, we argue that the learning behavior on various frequency components could be manipulated by changing the dataset statistical structure in the Fourier domain.
Intuitively, as frequency shortcuts are hidden in the dominant and highly dependent frequencies of dataset structure, dynamically perturbating the over-reliance frequency components could prevent the application of frequency shortcuts.
To this end, we propose two effective data augmentation modules designed to collaboratively and adaptively adjust the frequency characteristic of the dataset, aiming to dynamically influence the learning behavior of the model and ultimately serving as a strategy to mitigate shortcut learning. Our code will be made publicly available. Xilin He, Qinliang Lin, Weicheng Xie 0001, Siyang Song, Muhammad Haris Khan, LinLin Shen |
NeurIPS | 8 |
| 2024 | HairDiffusion: Vivid Multi-Colored Hair Editing via Latent DiffusionabstractHair editing is a critical image synthesis task that aims to edit hair color and hairstyle using text descriptions or reference images, while preserving irrelevant attributes (e.g., identity, background, cloth). Many existing methods are based on StyleGAN to address this task. However, due to the limited spatial distribution of StyleGAN, it struggles with multiple hair color editing and facial preservation. Considering the advancements in diffusion models, we utilize Latent Diffusion Models (LDMs) for hairstyle editing. Our approach introduces Multi-stage Hairstyle Blend (MHB), effectively separating control of hair color and hairstyle in diffusion latent space. Additionally, we train a warping module to align the hair color with the target region. To further enhance multi-color hairstyle editing, we fine-tuned a CLIP model using a multi-color hairstyle dataset. Our method not only tackles the complexity of multi-color hairstyles but also addresses the challenge of preserving original colors during diffusion editing. Extensive experiments showcase the superiority of our method in editing multi-color hairstyles while preserving facial attributes given textual descriptions and reference images. Yang Zhang 0012, LinLin Shen, Kaijun Deng, Weizhao He, Jinbao Wang 0001 |
NeurIPS | 4 |
| 2024 | 3DFaceMAE: Pre-training of Masked Autoencoder Using Patch-Based Random Masking Reconstruction and Super-resolution for 3D Face Recognition
Qiufu Li, LinLin Shen, Junpeng Yang |
PRCV (11) | 3 |
| 2024 | Introduction to the special issue on IEEE CBMS 2022 mining healthcare: AI and machine learning for biomedicine
Rosa Sicilia, LinLin Shen, Alejandro Rodríguez González, KC Santosh, Peter J. F. Lucas |
Artif. Intell. Medicine | 2 |
| 2024 | Photo realistic synthetic dataset and multi-scale attention dehazing network
Shengdong Zhang, Xiaoqin Zhang 0002, Wenqi Ren, LinLin Shen, Li Zhao 0005, Jun Zhang 0011 |
Eng. Appl. Artif. Intell. | 4 |
| 2024 | Dual-branch interactive cross-frequency attention network for deep feature learning
Qiufu Li, LinLin Shen |
Expert Syst. Appl. | 2 |
| 2024 | Two-stream regression network for dental implant position prediction
Xinquan Yang, Xuechen Li 0001, Wenting Chen, LinLin Shen, Xin Li 0196, Yongqiang Deng |
Expert Syst. Appl. | 5 |
| 2024 | SCPMan: Shape context and prior constrained multi-scale attention network for pancreatic segmentation
Leilei Zeng, Xuechen Li 0001, Xinquan Yang, Wenting Chen, Jingxin Liu 0005, LinLin Shen |
Expert Syst. Appl. | 6 |
| 2024 | Unsupervised fabric defect detection with high-frequency feature mapping
Da Wan, Can Gao, Jie Zhou 0009, Xinrui Shen, LinLin Shen |
Multim. Tools Appl. | 5 |
| 2024 | GAN-based dehazing network with knowledge transferring
Shengdong Zhang, Xiaoqin Zhang 0002, LinLin Shen, En Fan |
Multim. Tools Appl. | 3 |
| 2024 | ImplantFormer: vision transformer-based implant position regression using dental CBCT data
Xinquan Yang, Xuechen Li 0001, Peixi Wu, LinLin Shen, Yongqiang Deng |
Neural Comput. Appl. | 5 |
| 2024 | Heterogeneous domain adaptation via incremental discriminative knowledge consistency
Yuwu Lu, Dewei Lin, Jiajun Wen 0001, LinLin Shen, Xuelong Li 0001, Zhenkun Wen |
Pattern Recognit. | 4 |
| 2024 | Network characteristics adaption and hierarchical feature exploration for robust object recognition
Weicheng Xie 0001, Gui Wang, LinLin Shen, Zhihui Lai 0001, Siyang Song |
Pattern Recognit. | 4 |
| 2024 | PointSurFace: Discriminative point cloud surface feature extraction for 3D face recognition
Junpeng Yang, Qiufu Li, LinLin Shen |
Pattern Recognit. | 3 |
| 2024 | Fine-Grained Temporal-Enhanced Transformer for Dynamic Facial Expression RecognitionabstractDynamic facial expression recognition (DFER) plays a vital role in understanding human emotions and behaviors. Existing efforts tend to fall into a single modality self-supervised pretraining learning paradigm, which limits the representation ability of models. Besides, coarse-grained temporal modeling struggles to capture subtle facial expression representations from various inputs. In this letter, we propose a novel method for DFER, termed fine-grained temporal-enhanced transformer (FTET-DFER), which consists of two stages. First, we employ the inherent correlation between visual and auditory modalities in real videos, to capture temporally dense representations such as facial movements and expressions, in a self-supervised audio-visual learning manner. Second, we utilize the learned embeddings as targets, to achieve the DFER. In addition, we design the FTET block to study fine-grained temporal-enhanced facial expression features based on intra-clip locally-enhanced relations as well as inter-clip locally-enhanced global relationships in videos. Extensive experiments show that FTET-DFER outperforms the state-of-the-arts through within-dataset and cross-dataset evaluation. LinLin Shen, Zitong Yu, Zan Gao 0001 |
IEEE Signal Process. Lett. | 3 |
| 2024 | A Lightweight and Noise-Robust Method for Internal OCT Fingerprint ReconstructionabstractOptical coherence tomography (OCT), as a non-invasive and high-resolution three-dimensional imaging technology, can capture biological tissue structure information under the skin of fingertips. This structure information facilitates stronger anti-spoofing capability of automatic fingerprint recognition systems (AFRSs), and the reconstructed internal fingerprint images based on the structural information are more robust against poor skin conditions. Various internal fingerprint reconstruction methods have been proposed, but these approaches often ignore the continuity of spatial structure information, have a large number of model parameters and are sensitive to noise. Specific to these problems, this paper proposes a lightweight and noise-robust point detection network (LNPDN) to reconstruct internal fingerprints. At first, by combining the ShuffleNet with the temporal shift module and self-attention, the continuity of spatial information is considered. Meanwhile, the previous refined tissue structural region segmentation task, which is highly affected by noise, is transformed into an easy noise-robust feature point detection mission. Then, these detected points are synthesized into a curve to represent the upper envelope of the viable epidermis by linear interpolation. Finally, internal fingerprint image is reconstructed by averaging those pixel values at a certain depth range below the envelope. The experimental results show the proposed feature point extraction model for the central vertex of ridge blocks reaches the F1-score value of 93.911%, and the average minimum point-segment distance between the proposed curve and the target curve is 1.475. It demonstrates that the proposed model can well extract the central vertex of the ridge blocks and the curve can reflect the location of the viable epidermis. We also compared the recognition capabilities of internal fingerprints extracted from 2138 OCT fingerprint volume data on the public OCT fingerprint benchmark dataset. Our method achieves the lowest equal error rate of 0.167%, with a relative reduction of 60.91% compared with state-of-the-art reconstruction methods. Feng Liu 0013, Wenfeng Zeng, LinLin Shen |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2024 | Generative Imperceptible Attack With Feature Learning Bias Reduction and Multi-Scale Variance RegularizationabstractExisting studies have shown that malicious and imperceptible adversarial samples may significantly weaken the reliability and validity of deep learning systems. Since gradient-based attack algorithms may result in higher generation latency or demand large computation overhead, generative attack methods are frequently considered. However, the effectiveness and imperceptibility are still the main concerns for these generative attacks, 1) biased feature learning may occur, i.e., these algorithms may generate undesirable feature perturbations for samples that are less likely to be successfully attacked; 2) the produced perturbation noises may be easily perceived by human eyes. To this end, we propose a novel generative attack by manipulating the feature update. The proposed algorithm has two main merits, 1) our Bias-reduced Feature Manipulation (BrFM) that differentiates the hard-to-attack (Hard2Attack) and easy-to-attack (Easy2Attack) features, can avoid the possible learning shortcut for different difficulties of features in attack process, by customizing perturbations for Hard2Attack features to make them behave oppositely to those of benign features; 2) our Multi-scale Variance Regularization (MsVR) can reduce the unnatural transitions of perturbations in mask edges and flat areas with low contrast, while simultaneously trading off a reasonable attack capacity. Extensive experiments on the datasets of Caltech-101 and Imagenette in terms of the attack success rate and four imperceptibility metrics, show the effectiveness of our attack paradigm over the related state-of-the-art generative attack methods. Our codes will be made publicly available. Weicheng Xie 0001, Zenghao Niu, Qinliang Lin, Siyang Song, LinLin Shen |
IEEE Trans. Inf. Forensics Secur. | 5 |
| 2024 | GenFace: A Large-Scale Fine-Grained Face Forgery Benchmark and Cross Appearance-Edge LearningabstractThe rapid advancement of photorealistic generators has reached a critical juncture where the discrepancy between authentic and manipulated images is increasingly indistinguishable. Thus, benchmarking and advancing techniques detecting digital manipulation become an urgent issue. Although there have been a number of publicly available face forgery datasets, the forgery faces are mostly generated using GAN-based synthesis technology, which does not involve the most recent technologies like diffusion. The diversity and quality of images generated by diffusion models have been significantly improved and thus a much more challenging face forgery dataset shall be used to evaluate SOTA forgery detection literature. In this paper, we propose a large-scale, diverse, and fine-grained high-fidelity dataset, namely GenFace, to facilitate the advancement of deepfake detection, which contains a large number of forgery faces generated by advanced generators such as the diffusion-based model and more detailed labels about the manipulation approaches and adopted generators. In addition to evaluating SOTA approaches on our benchmark, we design an innovative Cross Appearance-Edge Learning (CAEL) detector to capture multi-grained appearance and edge global representations, and detect discriminative and general forgery traces. Moreover, we devise an Appearance-Edge Cross-Attention (AECA) module to explore the various integrations across two domains. Extensive experiment results and visualizations show that our detection model outperforms the state of the arts on different settings like cross-generator, cross-forgery, and cross-dataset evaluations. Code and datasets will be available athttps://github.com/Jenine-321/GenFace. Zitong Yu, Tianyi Wang 0006, Xiaobin Huang, LinLin Shen, Zan Gao 0001, Jianfeng Ren |
IEEE Trans. Inf. Forensics Secur. | 5 |
| 2024 | Generative Adversarial and Self-Supervised Dehazing NetworkabstractOwing to the fast developments of economics, a lot of devices and objects have been connected and have formed the Internet of Things (IoT). Visual sensors have been applied in vehicle navigation, traffic situational awareness, and traffic safety management. However, the particles in the air degrade the imaging quality, which affects the performance of vehicle navigation, traffic situational awareness, and traffic safety management. Deep-learning-based dehazing methods were proposed to address this issue. However, these methods are trained with simulated hazy images and cannot generalize to natural haze images well. To address the domain shift problem, some methods resort to zero-shot learning or domain adaption to boost the generalization of the model on natural haze images. However, the relevance between dehazed results and clean images is ignored by zero-shot dehazing methods. Domain-adaption-based dehazing methods ignore the relationship between the dehazed results and the hazy images. To overcome these issues, a generative adversarial and self-supervised dehazing network is introduced to boost the dehazing performance on real haze images. First, generative adversarial is employed to construct the relevance between dehazed results and haze-free images, which can boost the natural appearance of dehazed results. Second, self-supervised learning is employed to construct the relevance between the dehazed results and hazy images, which can restrict the solution space of dehazing. To show the effectiveness of the proposed model, we conduct extensive experiments on real and simulated haze images. Compared with state-of-the-art methods, the proposed model achieves state-of-the-art dehazing performance. Shengdong Zhang, Xiaoqin Zhang 0002, Shaohua Wan 0001, Wenqi Ren, Liping Zhao 0005, LinLin Shen |
IEEE Trans. Ind. Informatics | 6 |
| 2024 | Cross-Layer Contrastive Learning of Latent Semantics for Facial Expression RecognitionabstractConvolutional neural networks (CNNs) have achieved significant improvement for the task of facial expression recognition. However, current training still suffers from the inconsistent learning intensities among different layers, i.e., the feature representations in the shallow layers are not sufficiently learned compared with those in deep layers. To this end, this work proposes a contrastive learning framework to align the feature semantics of shallow and deep layers, followed by an attention module for representing the multi-scale features in the weight-adaptive manner. The proposed algorithm has three main merits. First, the learning intensity, defined as the magnitude of the backpropagation gradient, of the features on the shallow layer is enhanced by cross-layer contrastive learning. Second, the latent semantics in the shallow-layer and deep-layer features are explored and aligned in the contrastive learning, and thus the fine-grained characteristics of expressions can be taken into account for the feature representation learning. Third, by integrating the multi-scale features from multiple layers with an attention module, our algorithm achieved the state-of-the-art performances, i.e. 92.21%, 89.50%, 62.82%, on three in-the-wild expression databases, i.e. RAF-DB, FERPlus, SFEW, and the second best performance, i.e. 65.29% on AffectNet dataset. Our codes will be made publicly available. Weicheng Xie 0001, Zhibin Peng, LinLin Shen, Wenya Lu, Yang Zhang 0012, Siyang Song |
IEEE Trans. Image Process. | 3 |
| 2024 | Enhancing Unsupervised Semantic Segmentation Through Context-Aware ClusteringabstractDespite the great progress of semantic segmentation with supervised learning, annotating large amounts of pixel-wise labels is, however, very expensive and time-consuming. To this end, Unsupervised Semantic Segmentation(USS) has been proposed to learn semantic segmentation, without any form of annotations. This approach involves dense prediction of semantics which is however challenging due to the unreliable nature of local representations. To solve this problem, we propose a newly context-aware unsupervised semantic segmentation framework, which aims to enhance the unsupervised semantic segmentation by leveraging contextual knowledge within and across images. In particular, we introduce a training strategy based on our Pyramid Semantic Guidance (PSG), which utilizes holistic semantics on pyramid views to guide pixel clustering with a siamese network-based framework. Additionally, we introduce a Context-Aware Embedding (CAE) module to fuse global features with low-level geometrical and appearance representations. We evaluate our method on the COCO-Stuff dataset and achieved competitive results compared to both the convolutional and ViT-based USS methods. Specifically, we attain significant improvements of +4.5% and +5% mIoU for Stuff and all class segmentation respectively, compared to previous approaches that employ unsupervised convolutional backbones. Yuan Wang 0083, Junliang Chen 0002, Songhe Deng, Zhi Wang 0001, LinLin Shen, Wenwu Zhu 0001 |
IEEE Trans. Multim. | 6 |
| 2024 | Taming Self-Supervised Learning for Presentation Attack Detection: De-Folding and De-MixingabstractBiometric systems are vulnerable to presentation attacks (PAs) performed using various PA instruments (PAIs). Even though there are numerous PA detection (PAD) techniques based on both deep learning and hand-crafted features, the generalization of PAD for unknown PAI is still a challenging problem. In this work, we empirically prove that the initialization of the PAD model is a crucial factor for generalization, which is rarely discussed in the community. Based on such observation, we proposed a self-supervised learning-based method, denoted as DF-DM. Specifically, DF-DM is based on a global-local view coupled with de-folding and de-mixing to derive the task-specific representation for PAD. During de-folding, the proposed technique will learn region-specific features to represent samples in a local pattern by explicitly minimizing the generative loss. While de-mixing drives detectors to obtain the instance-specific features with global information for more comprehensive representation by minimizing the interpolation-based consistency. Extensive experimental results show that the proposed method can achieve significant improvements in terms of both face and fingerprint PAD in more complicated and hybrid datasets when compared with the state-of-the-art methods. When training in CASIA-FASD and Idiap Replay-Attack, the proposed method can achieve an 18.60% equal error rate (EER) in OULU-NPU and MSU-MFSD, exceeding the baseline performance by 9.54%. The source code of the proposed technique is available at https://github.com/kongzhecn/dfdm. Zhe Kong, Wentian Zhang, Feng Liu 0013, Wenhan Luo, LinLin Shen, Ramachandra Raghavendra |
IEEE Trans. Neural Networks Learn. Syst. | 6 |
| 2024 | Automatic Learning Rate Adaption for Memristive Deep Learning SystemsabstractAs a possible device to further enhance the performance of the hybrid complementary metal oxide semiconductor (CMOS) technology in the hardware, the memristor has attracted widespread attention in implementing efficient and compact deep learning (DL) systems. In this study, an automatic learning rate tuning method for memristive DL systems is presented. Memristive devices are utilized to adjust the adaptive learning rate in deep neural networks (DNNs). The speed of the learning rate adaptation process is fast at first and then becomes slow, which consist of the memristance or conductance adjustment process of the memristors. As a result, no manual tuning of learning rates is required in the adaptive back propagation (BP) algorithm. While cycle-to-cycle and device-to-device variations could be a significant issue in memristive DL systems, the proposed method appears robust to noisy gradients, various architectures, and different datasets. Moreover, fuzzy control methods for adaptive learning are presented for pattern recognition, such that the over-fitting issue can be well addressed. To our best knowledge, this is the first memristive DL system using an adaptive learning rate for image recognition. Another highlight of the presented memristive adaptive DL system is that quantized neural network architecture is utilized, and there is therefore a significant increase in the training efficiency, without the loss of testing accuracy. Yang Zhang 0012, LinLin Shen |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2023 | TCSloT: Text Guided 3D Context and Slope Aware Triple Network for Dental Implant Position PredictionabstractIn implant prosthesis treatment, the surgical guide of implant is used to ensure accurate implantation. However, such design heavily relies on the manual location of the implant position. When deep neural network has been proposed to assist the dentist in locating the implant position, most of them take a single slice as input, which do not fully explore 3D contextual information and ignores the influence of implant slope. In this paper, we design a Text Guided 3D Context and Slope Aware Triple Network (TCSloT) to integrate the perception of contextual information from multiple adjacent slices and awareness of variation of implant slopes. A Texture Variation Perception (TVP) module is correspondingly design to process the multiple slices and capture the texture variation among slices and a Slope-Aware Loss (SAL) is proposed to dynamically assign adaptive weights for the regression head. Additionally, we design a conditional text guidance (CTG) module to integrate the text condition (i.e., left, middle and right) from the CLIP to assist the implant position prediction. Extensive experiments on a dental implant dataset through five-fold cross-validation, demonstrated that the proposed TCSloT achieves superior performance than existing methods. Xinquan Yang, Jinheng Xie, Xuechen Li 0001, LinLin Shen, Yongqiang Deng |
BIBM | 5 |
| 2023 | Multi-scale Contrastive Learning for Gastroenteroscopy ClassificationabstractIn gastroenteroscopy image analysis, numerous CADs demonstrate that deep learning aids doctors' diagnosis. The shapes and sizes of the lesions are varied. And in the clinic, the dataset appears to be data imbalanced. However, existing methods directly classify by texture and ignore lesions with various shapes and sizes. To address the issue above, we propose a deep neural network, which consists of multi-scale feature extraction, contrastive feature learning and a multi-scale feature fusion module. We train the contrastive feature learning module and multi-scale feature fusion module simultaneously to alleviate the issue of data distribution differences. Thus, the proposed network can better identify various categories. Extensive experiments on the Hyper Kvasir dataset show that the proposed Hybrid-M2CL outperforms the benchmark proposed by the dataset with 5.0% Macro Precision, 3.3% Macro Recall, 3.4% Macro F1-score, 3.3% Micro Precision, 3.6% MCC. In addition, it outperforms the SOTA by 1.1% Macro F1-score, 2.6% MCC, and 2.0% B-ACC. Xuechen Li 0001, Zhibin Peng, Wenting Chen, LinLin Shen, Guangyao Wu |
CBMS | 5 |
| 2023 | StyleGene: Crossover and Mutation of Region-level Facial Genes for Kinship Face SynthesisabstractHigh-fidelity kinship face synthesis has many potential applications, such as kinship verification, missing child identification, and social media analysis. However, it is challenging to synthesize high-quality descendant faces with genetic relations due to the lack of large-scale, high-quality annotated kinship data. This paper proposes RFG (Region-level Facial Gene) extraction framework to address this issue. We propose to use IGE (Image-based Gene Encoder), LGE (Latent-based Gene Encoder) and Gene Decoder to learn the RFGs of a given face image, and the relationships between RFGs and the latent space of Style-GAN2. As cycle-like losses are designed to measure the$\mathcal{L}_{2}$distances between the output of Gene Decoder and image encoder, and that between the output of LGE and IGE, only face images are required to train our framework, i.e. no paired kinship face data is required. Based upon the proposed RFGs, a crossover and mutation module is further designed to inherit the facial parts of parents. A Gene Pool has also been used to introduce the variations into the mutation of RFGs. The diversity of the faces of descendants can thus be significantly increased. Qualitative, quantitative, and subjective experiments on FIW, TSKinFace, and FF-Databases clearly show that the quality and diversity of kinship faces generated by our approach are much better than the existing state-of-the-art methods. Xianxu Hou, Zepeng Huang, LinLin Shen |
CVPR | 4 |
| 2023 | Activation Template Matching Loss for Explainable Face RecognitionabstractCan we construct an explainable face recognition network able to learn a facial part-based feature like eyes, nose, mouth and so forth, without any manual annotation or additionalsion datasets? In this paper, we propose a generic Explainable Channel Loss (ECLoss) to construct an explainable face recognition network. The explainable network trained with ECLoss can easily learn the facial part-based representation on the target convolutional layer, where an individual channel can detect a certain face part. Our experiments on dozens of datasets show that ECLoss achieves superior explainability metrics, and at the same time improves the performance of face verification without face alignment. In addition, our visualization results also illustrate the effectiveness of the proposed ECLoss. Huawei Lin 0001, Qiufu Li, LinLin Shen |
FG | 4 |
| 2023 | StyleAU: StyleGAN based Facial Action Unit Manipulation for Expression EditingabstractFacial expression editing has a wide range of applications, such as emotion detection, human-computer interaction, and social entertainment. However, existing expression editing methods either fail to allow for fine-grained editing, resulting in unnatural and unrealistic facial expressions, or generate artifacts and blurs, leading to poor image quality. In this paper, we propose a novel framework called StyleAU, which is based on StyleGAN and facial action units, to address these problems. Our framework leverages the pre-trained StyleGAN prior knowledge to enable action unit editing of the face in the StyleGAN latent space, allowing precise expression editing. In addition, we use an encoder to extract multi-scale content features to achieve high-fidelity image reconstruction. Our approach qualitatively and quantitatively outperforms competing methods for action unit manipulation and expression editing. Yanliang Guo, Xianxu Hou, Feng Liu 0013, LinLin Shen, Lei Wang 0018, Zhen Wang 0009, Peng Liu 0039 |
IJCB | 4 |
| 2023 | Shift from Texture-bias to Shape-bias: Edge Deformation-based Augmentation for Robust Object RecognitionabstractRecent studies have shown the vulnerability of CNNs under perturbation noises, which is partially caused by the reason that the well-trained CNNs are too biased toward the object texture, i.e., they make predictions mainly based on texture cues. To reduce this texture-bias, current studies resort to learning augmented samples with heavily perturbed texture to make networks be more biased toward relatively stable shape cues. However, such methods usually fail to achieve real shape-biased networks due to the insufficient diversity of the shape cues. In this paper, we propose to augment the training dataset by generating semantically meaningful shapes and samples, via a shape deformation-based online augmentation, namely as SDbOA. The samples generated by our SDbOA have two main merits. First, the augmented samples with more diverse shape variations enable networks to learn the shape cues more elaborately, which encourages the network to be shape-biased. Second, semantic-meaningful shape-augmentation samples could be produced by jointly regularizing the generator with object texture and edge-guidance soft constraint, where the edges are represented more robustly with a self information guided map to better against the noises on them. Extensive experiments under various perturbation noises demonstrate the obvious superiority of our shape-bias-motivated model over the state of the arts in terms of robustness performance. Code is available at https://github.com/C0notSilly/-ICCV-23-Edge-Deformation-based-Online-Augmentation. Xilin He, Qinliang Lin, Weicheng Xie 0001, Siyang Song, Feng Liu 0013, LinLin Shen |
ICCV | 7 |
| 2023 | UniFace: Unified Cross-Entropy Loss for Deep Face RecognitionabstractAs a widely used loss function in deep face recognition, the softmax loss cannot guarantee that the minimum positive sample-to-class similarity is larger than the maximum negative sample-to-class similarity. As a result, no unified threshold is available to separate positive sample-to-class pairs from negative sample-to-class pairs. To bridge this gap, we design a UCE (Unified Cross-Entropy) loss for face recognition model training, which is built on the vital constraint that all the positive sample-to-class similarities shall be larger than the negative ones. Our UCE loss can be integrated with margins for a further performance boost. The face recognition model trained with the proposed UCE loss, UniFace, was intensively evaluated using a number of popular public datasets like MFR, IJB-C, LFW, CFP-FP, AgeDB, and MegaFace. Experimental results show that our approach outperforms SOTA methods like SphereFace, CosFace, ArcFace, Partial FC, etc. Especially, till the submission of this work (Mar. 8, 2023), the proposed UniFace achieves the highest TAR@MR-All on the academic track of the MFR-ongoing challenge. $\color{Blue}{\mathbf{Code}}$ is publicly available. Jiancan Zhou, Xi Jia, Qiufu Li, LinLin Shen, Jinming Duan 0001 |
ICCV | 4 |
| 2023 | Multi Task-Based Facial Expression Synthesis with Supervision Learning and Feature Disentanglement of Image StyleabstractImage-to-Image synthesis paradigms have been widely used for facial expression synthesis. However, current generators are apt to either produce artifacts for largely posed and non-aligned faces or unduly change the identity information like AdaIN-based generator. In this work, we suggest to use image style feature to surrogate the expression cues in the generator, and propose a multi-task learning paradigm to explore this style information via the supervision learning and feature disentanglement. While the supervision learning can make the encoded style specifically represent the expression cues and enable the generator to produce correct expression, the feature disentanglement of content and style cues enables the generator to better preserve the identity information in expression synthesis. Experimental results show that the proposed algorithm can well reduce the artifacts for the synthesis of posed and non-aligned expressions, and achieves competitive performances in terms of FID, PNSR and classification accuracy, compared with four publicly available GANs. The code and pre-trained models are available at https://github.com/lumanxi236/MTSS. Wenya Lu, Zhibin Peng, Weicheng Xie 0001, Jiajun Wen 0001, Zhihui Lai 0001, LinLin Shen |
ICIP | 7 |
| 2023 | ADATS: Adaptive RoI-Align based Transformer for End-to-End Text SpottingabstractScene text spotting has attracted great attention in recent years. Compared with two-stage approaches that locate scene texts in the first stage and recognize them in the second stage, the advantages of joint location and recognition training are not fully explored. In this paper, we present an ADaptive RoI-Align based transformer for end-to-end Text Spotting (ADATS), which simultaneously locates and recognizes text with a single forward pass. By employing an Adaptive RoI-Align, the text features are extracted from the feature extraction network with the original aspect ratio, such that less information is lost during the alignment of arbitrarily-shaped scene text. Attention-based segmentation and recognition heads allow us to simultaneously optimize detection and recognition. Experiments on ICDAR 2015, MSRA-TD500, Total-Text, and CTW1500 demonstrate the effectiveness of our method. Zepeng Huang, Qi Wan, Junliang Chen 0002, Kai Ye 0004, LinLin Shen |
ICME | 6 |
| 2023 | Optimal Low-Rank QR Decomposition with an Application on RP-TSOD
Haiyan Yu 0003, Jianfeng Ren, Ruibin Bai, LinLin Shen |
ICONIP (14) | 4 |
| 2023 | CDNet: Cross-frequency Dual-branch Network for Face Anti-SpoofingabstractFace anti-spoofing (FAS) defends the facial image recognition systems against the spoof attacks. While the imperceptible spoof cues in the facial images are usually represented in the images' high-frequency components, existing methods do not fully explore them. In this paper, we introduce wavelet into face anti-spoofing and propose a Cross-frequency Dual-branch network (CDNet), which mainly contains two frequency branches to explore spoof cues from the input facial images' high- and low-frequency components generated by wavelet transforms. In CDNet, we design Frequency Attention Module (FAM) to fuse different internal frequency features learned by two frequency branches, and propose a Complementary Learning Module (CLM) to aggregate the two final frequency features. In addition, we present a resolution-aware Binary Cross-Entropy Loss to balance the training samples with different resolutions. We conduct comprehensive experiments on four datasets, and the results shows that our CDNet performs better than the previous state-of-the-art methods on both intra- and inter-dataset testing. Xiaobin Huang, Qiufu Li, LinLin Shen |
IJCNN | 3 |
| 2023 | TCEIP: Text Condition Embedded Regression Network for Dental Implant Position Prediction
Xinquan Yang, Jinheng Xie, Xuechen Li 0001, Xin Li 0196, LinLin Shen, Yongqiang Deng |
MICCAI (6) | 6 |
| 2023 | QA-CLIMS: Question-Answer Cross Language Image Matching for Weakly Supervised Semantic SegmentationabstractClass Activation Map (CAM) has emerged as a popular tool for weakly supervised semantic segmentation (WSSS), allowing the localization of object regions in an image using only image-level labels. However, existing CAM methods suffer from under-activation of target object regions and false-activation of background regions due to the fact that a lack of detailed supervision can hinder the model's ability to understand the image as a whole. In this paper, we propose a novel Question-Answer Cross-Language-Image Matching framework for WSSS (QA-CLIMS), leveraging the vision-language foundation model to maximize the text-based understanding of images and guide the generation of activation maps. First, a series of carefully designed questions are posed to the VQA (Visual Question Answering) model with Question-Answer Prompt Engineering (QAPE) to generate a corpus of both foreground target objects and backgrounds that are adaptive to query images. We then employ contrastive learning in a Region Image Text Contrastive (RITC) network to compare the obtained foreground and background regions with the generated corpus. Our approach exploits the rich textual information from the open vocabulary as additional supervision, enabling the model to generate high-quality CAMs with a more complete object region and reduce false-activation of background regions. We conduct extensive analysis to validate the proposed method and show that our approach performs state-of-the-art on both PASCAL VOC 2012 and MS COCO datasets. Songhe Deng, Jinheng Xie, LinLin Shen |
ACM Multimedia | 4 |
| 2023 | CCF-Net: A Cascade Center-Based Framework Towards Efficient Human Parts Detection
Kai Ye 0004, Haoqin Ji, Lei Wang 0018, Peng Liu 0039, LinLin Shen |
MMM (2) | 6 |
| 2023 | UniTSFace: Unified Threshold Integrated Sample-to-Sample Loss for Face RecognitionabstractSample-to-class-based face recognition models can not fully explore the cross-sample relationship among large amounts of facial images, while sample-to-sample-based models require sophisticated pairing processes for training. Furthermore, neither method satisfies the requirements of real-world face verification applications, which expect a unified threshold separating positive from negative facial pairs. In this paper, we propose a unified threshold integrated sample-to-sample based loss (USS loss), which features an explicit unified threshold for distinguishing positive from negative pairs. Inspired by our USS loss, we also derive the sample-to-sample based softmax and BCE losses, and discuss their relationship. Extensive evaluation on multiple benchmark datasets, including MFR, IJB-C, LFW, CFP-FP, AgeDB, and MegaFace, demonstrates that the proposed USS loss is highly efficient and can work seamlessly with sample-to-class-based losses. The embedded loss (USS and sample-to-class Softmax loss) overcomes the pitfalls of previous approaches and the trained facial model UniTSFace exhibits exceptional performance, outperforming state-of-the-art methods, such as CosFace, ArcFace, VPL, AnchorFace, and UNPG. Our code is available at https://github.com/CVI-SZU/UniTSFace. Qiufu Li, Xi Jia, Jiancan Zhou, LinLin Shen, Jinming Duan 0001 |
NeurIPS | 4 |
| 2023 | Learning Visual Prior via Generative Pre-TrainingabstractVarious stuff and things in visual data possess specific traits, which can be learned by deep neural networks and are implicitly represented as the visual prior, e.g., object location and shape, in the model. Such prior potentially impacts many vision tasks. For example, in conditional image synthesis, spatial conditions failing to adhere to the prior can result in visually inaccurate synthetic results. This work aims to explicitly learn the visual prior and enable the customization of sampling. Inspired by advances in language modeling, we propose to learn Visual prior via Generative Pre-Training, dubbed VisorGPT. By discretizing visual locations, e.g., bounding boxes, human pose, and instance masks, into sequences, VisorGPT can model visual prior through likelihood maximization. Besides, prompt engineering is investigated to unify various visual locations and enable customized sampling of sequential outputs from the learned prior. Experimental results demonstrate the effectiveness of VisorGPT in modeling visual prior and extrapolating to novel scenes, potentially motivating that discrete visual locations can be integrated into the learning paradigm of current language models to further perceive visual world. Code is available at https://sierkinhane.github.io/visor-gpt. Jinheng Xie, Kai Ye 0004, Yudong Li 0001, Yuexiang Li, Qinghong Lin, Yefeng Zheng 0001, LinLin Shen, Zheng Shou 0001 |
NeurIPS | 7 |
| 2023 | MaskDiffuse: Text-Guided Face Mask Removal Based on Diffusion Models
Jingxia Lu, Xianxu Hou, Zhibin Peng, LinLin Shen, Lixin Fan |
PRCV (6) | 5 |
| 2023 | Learning Adapters for Text-Guided Portrait Stylization with Pretrained Diffusion Models
Mintu Yang, Xianxu Hou, LinLin Shen, Lixin Fan |
PRCV (1) | 4 |
| 2023 | Adversarial Keyword Extraction and Semantic-Spatial Feature Aggregation for Clinical Report Guided Thyroid Nodule Segmentation
Yudi Zhang 0005, Wenting Chen, Xuechen Li 0001, LinLin Shen, Zhihui Lai 0001, Heng Kong |
PRCV (13) | 4 |
| 2023 | JLInst: Boundary-Mask Joint Learning for Instance Segmentation
Junliang Chen 0002, Zepeng Huang, LinLin Shen |
PRCV (12) | 4 |
| 2023 | Adaptive weighted rain streaks model-driven deep network for single image deraining
LinLin Shen, Zhihui Lai 0001 |
Expert Syst. Appl. | 2 |
| 2023 | Texture and semantic convolutional auto-encoder for anomaly detection and segmentationabstractAbstract Anomaly detection is a challenging task, especially detecting and segmenting tiny defect regions in images without anomaly priors. Although deep encoder‐decoder‐based convolutional neural networks have achieved good anomaly detection results, existing methods operate uniformly on all extracted image features without considering disentangling these features. To fully explore the texture and semantic information of images, A novel unsupervised anomaly detection method is proposed. Specifically, discriminative features are extracted from images by using a deep pre‐trained network, where shallow and deep features are aggregated into texture and semantic modules, respectively. Then, a feature fusion module is developed to interactively enable feature information in two different modules. The texture and semantic segmentation results are obtained by comparing the texture features and semantic features before and after reconstruction, respectively. Finally, an anomaly segmentation module is designed to generate anomaly detection results by integrating the results of the texture and semantic modules by setting a threshold. Experimental results on benchmark datasets for anomaly detection demonstrate that our proposed method can efficiently and effectively detect anomalies, outperforming some state‐of‐the‐art methods by 2.7% and 0.6% in classification and segmentation. Jintao Luo, Can Gao, Da Wan, LinLin Shen |
IET Comput. Vis. | 4 |
| 2023 | CT-Based Automatic Spine Segmentation Using Patch-Based Deep LearningabstractCT vertebral segmentation plays an essential role in various clinical applications, such as computer‐assisted surgical interventions, assessment of spinal abnormalities, and vertebral compression fractures. Automatic CT vertebral segmentation is challenging due to the overlapping shadows of thoracoabdominal structures such as the lungs, bony structures such as the ribs, and other issues such as ambiguous object borders, complicated spine architecture, patient variability, and fluctuations in image contrast. Deep learning is an emerging technique for disease diagnosis in the medical field. This study proposes a patch‐based deep learning approach to extract the discriminative features from unlabeled data using a stacked sparse autoencoder (SSAE). 2D slices from a CT volume are divided into overlapping patches fed into the model for training. A random under sampling (RUS)‐module is applied to balance the training data by selecting a subset of the majority class. SSAE uses pixel intensities alone to learn high‐level features to recognize distinctive features from image patches. Each image is subjected to a sliding window operation to express image patches using autoencoder high‐level features, which are then fed into a sigmoid layer to classify whether each patch is a vertebra or not. We validate our approach on three diverse publicly available datasets: VerSe, CSI‐Seg, and the Lumbar CT dataset. Our proposed method outperformed other models after configuration optimization by achieving 89.9% in precision, 90.2% in recall, 98.9% in accuracy, 90.4% in F‐score, 82.6% in intersection over union (IoU), and 90.2% in Dice coefficient (DC). The results of this study demonstrate that our model’s performance consistency using a variety of validation strategies is flexible, fast, and generalizable, making it suited for clinical application. Syed Furqan Qadri, Hongxiang Lin, LinLin Shen, Mubashir Ahmad, Salman Qadri, Salabat Khan, Maqbool Khan, Syeda Shamaila Zareen, Muhammad Azeem Akbar, Md Belal Bin Heyat, Saqib Qamar |
Int. J. Intell. Syst. | 3 |
| 2023 | Tongue size and shape classification fusing segmentation features for traditional Chinese medicine diagnosis
Xuechen Li 0001, Siting Zheng, LinLin Shen, Changen Zhou, Zhihui Lai 0001 |
Neural Comput. Appl. | 6 |
| 2023 | Learning from pseudo-lesion: a self-supervised framework for COVID-19 diagnosis
Xuechen Li 0001, Zhihao Jin, LinLin Shen |
Neural Comput. Appl. | 4 |
| 2023 | Enhanced robust spatial feature selection and correlation filter learning for UAV tracking
Jiajun Wen 0001, Hong-Lin Chu, Zhihui Lai 0001, Tianyang Xu 0001, LinLin Shen |
Neural Networks | 5 |
| 2023 | Deep generative image priors for semantic face manipulationabstractPrevious works on generative adversarial networks (GANs) mainly focus on how to synthesize high-fidelity images. In this paper, we present a framework to leverage the knowledge learned by GANs for semantic face manipulation. In particular, we propose to control the semantics of synthesized faces by adapting the latent codes with an attribute prediction model. Moreover, in order to achieve a more accurate estimation of different facial attributes, we propose to pretrain the attribute prediction model by inverting the synthesized face images back to the GAN latent space. As a result, our method explicitly considers the semantics encoded in the latent space of a pretrained GAN and is able to faithfully edit various attributes like eyeglasses, smiling, bald, age, mustache and gender for high-resolution face images. Extensive experiments show that our method has superior performance compared to state of the art for both face attribute prediction and semantic face manipulation. Xianxu Hou, LinLin Shen, Zhong Ming 0001, Guoping Qiu |
Pattern Recognit. | 2 |
| 2023 | Self-Supervised Learning of Person-Specific Facial Dynamics for Automatic Personality RecognitionabstractThis article aims to solve two important issues that frequently occur in existing automatic personality analysis systems: 1. Attempting to use very short video segments or even single frames, rather than long-term behaviour, to infer personality traits; 2. Lack of methods to encode person-specific facial dynamics for personality recognition. To deal with these issues, this paper first proposes a novel Rank Loss which utilizes the natural temporal evolution of facial actions, rather than personality labels, for self-supervised learning of facial dynamics. Our approach first trains a generic U-net style model that can infer general facial dynamics learned from a set of unlabelled face videos. Then, the generic model is frozen, and a set of intermediate filters are incorporated into this architecture. The self-supervised learning is then resumed with only person-specific videos. This way, the learned filters’ weights are person-specific, making them a valuable source for modeling person-specific facial dynamics. We then propose to concatenate the weights of the learned filters as a person-specific representation, which can be directly used to predict the personality traits without needing other parts of the network. We evaluate the proposed approach on both self-reported personality and apparent personality datasets. In addition to achieving promising results in the estimation of personality trait scores from videos, we show that the tasks conducted by the subject in the video matters, that fusion of a combination of tasks reaches highest accuracy, and that multi-scale dynamics are more informative than single-scale dynamics. Siyang Song, Shashank Jaiswal, Enrique Sánchez-Lozano, Georgios Tzimiropoulos, LinLin Shen, Michel F. Valstar |
IEEE Trans. Affect. Comput. | 5 |
| 2023 | Learning Person-Specific Cognition From Facial Reactions for Automatic Personality RecognitionabstractThis article proposes to recognise the true (self-reported) personality traits from the target subject's cognition simulated from facial reactions. This approach builds on the following two findings in cognitive science: (i) human cognition partially determines expressed behaviour and is directly linked to true personality traits; and (ii) in dyadic interactions, individuals’ nonverbal behaviours are influenced by their conversational partner's behaviours. In this context, we hypothesise that during a dyadic interaction, a target subject's facial reactions are driven by two main factors: their internal (person-specific) cognitive process, and the externalised nonverbal behaviours of their conversational partner. Consequently, we propose to represent the target subject's (defined as the listener) person-specific cognition in the form of a person-specific CNN architecture that has unique architectural parameters and depth, which takes audio-visual non-verbal cues displayed by the conversational partner (defined as the speaker) as input, and is able to reproduce the target subject's facial reactions. Each person-specific CNN is explored by the Neural Architecture Search (NAS) and a novel adaptive loss function, which is then represented as a graph representation for recognising the target subject's true personality. Experimental results not only show that the produced graph representations are well associated with target subjects’ personality traits in both human-human and human-machine interaction scenarios, and outperform the existing approaches with significant advantages, but also demonstrate that the proposed novel strategies help in learning more reliable personality representations. Siyang Song, Zilong Shao, Shashank Jaiswal, LinLin Shen, Michel F. Valstar, Hatice Gunes |
IEEE Trans. Affect. Comput. | 4 |
| 2023 | Adversarial Learning of Object-Aware Activation Map for Weakly-Supervised Semantic SegmentationabstractRecent years have witnessed impressive advances in the area of weakly-supervised semantic segmentation (WSSS). However, most of existing approaches are based on class activation maps (CAMs), which suffer from the under-segmentation problem (i.e., objects of interest are segmented partially). Although a number of literature works have been proposed to tackle this under-segmentation problem, we argue that these solutions built on CAMs may not be optimal for the WSSS task. Instead, in this paper we propose a network based on the object-aware activation map (OAM). The proposed network, termed OAM-Net, consists of four loss functions (foreground loss, background loss, average pixel and consistency loss) which ensure exactness, completeness, compactness and consistency of segmented objects via adversarial training. Compared to conventional CAM-based methods, our OAM-Net overcomes the under-segmentation drawback and significantly improves segmentation accuracy with negligible computational cost. A thorough comparison between OAM-Net and CAM-based approaches is carried out on the PASCAL VOC2012 dataset, and experimental results show that our network outperforms state-of-the-art approaches by a large margin. The code will be available soon. Junliang Chen 0002, Weizeng Lu, Yuexiang Li, LinLin Shen, Jinming Duan 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2023 | Weakly Supervised Pedestrian Segmentation for Person Re-IdentificationabstractPerson re-identification (RelD) is an important problem in intelligent surveillance and public security. Among all the solutions to this problem, existing mask-based methods first use a well-pretrained segmentation model to generate a foreground mask, in order to exclude the background from ReID. Then they perform the RelD task directly on the segmented pedestrian image. However, such a process requires extra datasets with pixel-level semantic labels. In this paper, we propose a Weakly Supervised Pedestrian Segmentation (WSPS) framework to produce the foreground mask directly from the RelD datasets. In contrast, our WSPS only requires image-level subject ID labels. To better utilize the pedestrian mask, we also propose the Image Synthesis Augmentation (ISA) technique to further augment the dataset. Experiments show that the features learned from our proposed framework are robust and discriminative. Compared with the baseline, the mAP of our framework is about 4.4%, 11.7%, and 4.0% higher on three widely used datasets including Market-1501, CUHK03, and MSMT17. The code will be available soon. Ziqi Jin, Jinheng Xie, Bizhu Wu, LinLin Shen |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2023 | Robust Twin Bounded Support Vector Classifier With Manifold RegularizationabstractSupport vector machine (SVM), as a supervised learning method, has different kinds of varieties with significant performance. In recent years, more research focused on nonparallel SVM, where twin SVM (TWSVM) is the typical one. In order to reduce the influence of outliers, more robust distance measurements are considered in these methods, but the discriminability of the models is neglected. In this article, we propose robust manifold twin bounded SVM (RMTBSVM), which considers both robustness and discriminability. Specifically, a novel norm, that is, capped$L_{1}$-norm, is used as the distance metric for robustness, and a robust manifold regularization is added to further improve the robustness and classification performance. In addition, we also use the kernel method to extend the proposed RMTBSVM for nonlinear classification. We introduce the optimization problems of the proposed model. Subsequently, effective algorithms for both linear and nonlinear cases are proposed and proved to be convergent. Moreover, the experiments are conducted to verify the effectiveness of our model. Compared with other methods under the SVM framework, the proposed RMTBSVM shows better classification accuracy and robustness. Junhong Zhang, Zhihui Lai 0001, Heng Kong, LinLin Shen |
IEEE Trans. Cybern. | 4 |
| 2023 | Precise Facial Landmark Detection by Reference Heatmap TransformerabstractMost facial landmark detection methods predict landmarks by mapping the input facial appearance features to landmark heatmaps and have achieved promising results. However, when the face image is suffering from large poses, heavy occlusions and complicated illuminations, they cannot learn discriminative feature representations and effective facial shape constraints, nor can they accurately predict the value of each element in the landmark heatmap, limiting their detection accuracy. To address this problem, we propose a novel Reference Heatmap Transformer (RHT) by introducing reference heatmap information for more precise facial landmark detection. The proposed RHT consists of a Soft Transformation Module (STM) and a Hard Transformation Module (HTM), which can cooperate with each other to encourage the accurate transformation of the reference heatmap information and facial shape constraints. Then, a Multi-Scale Feature Fusion Module (MSFFM) is proposed to fuse the transformed heatmap features and the semantic features learned from the original face images to enhance feature representations for producing more accurate target heatmaps. To the best of our knowledge, this is the first study to explore how to enhance facial landmark detection by transforming the reference heatmap information. The experimental results from challenging benchmark datasets demonstrate that our proposed method outperforms the state-of-the-art methods in the literature. Jun Wan 0005, Jun Liu 0036, Jie Zhou 0009, Zhihui Lai 0001, LinLin Shen, Ping Xiong 0001, Wenwen Min |
IEEE Trans. Image Process. | 5 |
| 2023 | TextFace: Text-to-Style Mapping Based Face Generation and ManipulationabstractAs a subtopic of text-to-image synthesis, text-to-face generation has great potential in face-related applications. In this paper, we propose a generic text-to-face framework, namely, TextFace, to achieve diverse and high-quality face image generation from text descriptions. We introduce text-to-style mapping, a novel method where the text description can be directly encoded into the latent space of a pretrained StyleGAN. Guided by our text-image similarity matching and face captioning-based text alignment, the textual latent code can be fed into the generator of a well-trained StyleGAN to produce diverse face images with high resolution (1024×1024). Furthermore, our model inherently supports semantic face editing using text descriptions. Finally, experimental results quantitatively and qualitatively demonstrate the superior performance of our model. Xianxu Hou, Yudong Li 0001, LinLin Shen |
IEEE Trans. Multim. | 4 |
| 2023 | Lifelong Age Transformation With a Deep Generative PriorabstractIn this paper, we consider the lifelong age progression and regression task, which requires to synthesize a persons appearance across a wide range of ages. We propose a simple yet effective learning framework to achieve this by exploiting the prior knowledge of faces captured by well-trained generative adversarial networks (GANs). Specifically, we first utilize a pretrained GAN to synthesize face images with different ages, with which we then learn to model the conditional aging process in the GAN latent space. Moreover, we also introduce a cycle consistency loss in the GAN latent space to preserve a persons identity. As a result, our model can reliably predict a person's appearance for different ages by modifying both shape and texture of the head. Both qualitative and quantitative experimental results demonstrate the superiority of our method over concurrent works. Furthermore, we demonstrate that our approach can also achieve high-quality age transformation for painting portraits and cartoon characters without additional age annotations. Xianxu Hou, Hanbang Liang, LinLin Shen, Zhong Ming 0001 |
IEEE Trans. Multim. | 4 |
| 2023 | Consistency Preservation and Feature Entropy Regularization for GAN Based Face EditingabstractGenerative Adversarial Network (GAN) has been widely used for image-to-image translation-based facial attribute editing. Existing GAN networks are likely to generate samples with anomalies, which may be caused by the lack of consistency preservation and feature entanglement. For preserving image consistency, many studies resorted to the design of the network framework and loss functions, e.g. cycle-consistency loss. However, the generator with the cycle-consistency loss could not well preserve the attribute-irrelevant features, and its feature-level noises may possibly cause synthesis abnormalities. For feature disentanglement, previous works were devoted to mining the implicit semantics of feature spaces, while these semantics are not stable and intuitive enough. For consistency preservation, we propose a target consistency loss to complement the cycle-consistency loss, and enable the network to learn to preserve features of the image more directly. Meanwhile, we filter out outlier feature maps to reduce the synthesis abnormalities and propose a dynamic dropout to better preserve the attribute-irrelevant features. For feature disentanglement, we encode the image semantics more stably and intuitively and propose an entropy regularization to decouple these semantics to allow independent editing of different attributes. The proposed modules are general and can be easily integrated with available image-to-image-based GAN models like StarGAN, AttGAN, and STGAN. Extensive experiments on CelebA dataset show that the our strategy can largely reduce the artifacts and better preserve the subtle facial features, and thus significantly improve the facial editing performance of these mainstream GAN models, in terms of FID, PSNR and SSIM. Additional experiments on realistic expression editing show that our method outperforms StarGAN on RaFD, and achieves much better generalization performances than the three baselines on datasets of FFHQ, RaFD and LFW. Weicheng Xie 0001, Wenya Lu, Zhibin Peng, LinLin Shen |
IEEE Trans. Multim. | 4 |
| 2022 | Left and Right Ventricular Segmentation Based on 3D Region-Aware U-NetabstractThe cardiac is one of the essential organs, and the segmentation of the left and right ventricular of cardiac is essential in diagnosing various heart diseases. The most popular method for the segmentation of 3D MRI images is the nnUNet. However, the 3D MRI volume of the ventricular contains other organs which interfere with the segmentation of the ventricular. Hence, we proposed a novel region-aware U-Net segmentation method RegUNet for ventricular segmentation. RegUNet improves the ventricular's segmentation performance by first capturing the region of interest (RoI) of the ventricular and then segmenting the ventricular with the captured RoI features, which reduces the segmentation module's difficulty by keeping the cardiac's features and leaving others such that RegUNet can focus on ventricular segmentation. Besides, since the model segments the ventricular with the captured RoI features, it saves the model's computing resources from identifying the background of the volume. Since 3D cardiac MRI volumes scanned by the different devices have diverse statistical characteristics, which causes the model's performance in processing the multi-source cardiac volumes to be unstable. We stabilize the model's performance with a multi-sources feature normalization strategy, which normalizes the feature from a different source with different parameters. We validated the proposed method on the M&MS dataset, a multi-sources 3D MRI cardiac segmentation dataset. Experiments showed that RegUNet's segmentation ability reached the state-of-the-art. Xueting Liu 0001, Huisi Wu, Zhenkun Wen, LinLin Shen |
CBMS | 6 |
| 2022 | Dual Fusion Mass Detector for Mammogram Mass DetectionabstractMammogram mass detection is a difficult task due to the mass character of the tiny area, fuzzy boundary, and occlusion. To address these problems, this paper proposes a novel detection network for mammogram mass detection. Firstly, we propose a novel feature fusion structure and Small Target Attention Module (STAM) to improve the model's ability to detect small masses. Secondly, Results-oriented Loss (ROL) is adopted to obtain better model performance. Finally, Incremental Positive Selection (IPS) is used to divide positive and negative anchors. The scarcity of breast mammogram images for training aggravates the difficulty of mass detection. Thus, we open our collected dataset, which contains 1456 mammogram images from 400 patients. Since the model includes a double feature fusion structure, the proposed network is named Dual Fusion Mass Detector (DFMD). Experiment results show that DFMD is robust to various variations on scale, blurry and occlusion. Zhihui Lai 0001, Heng Kong, LinLin Shen |
CBMS | 4 |
| 2022 | Breast Lesions Segmentation using Dual-level UNet (DL-UNet)abstractBreast disease is one of the primary diseases endangering women's health. Accurate segmentation of breast lesions can help doctors diagnose breast diseases. However, the size and morphology of breast lesions are different, and the intensity of breast tissue is uneven. Thus, it is challenging to segment the lesion area accurately. In this paper, we propose Dual-scale Feature Fusion (DSFF) module and Edgeloss to segment breast lesions. The DSFF module aims to integrate two-scale features and design another effective skip connection scheme to reduce false positive regions. To solve the problem of unclear segmentation boundary, we design Edgeloss for additional supervision on the boundary region to obtain a finer segmentation boundary. The experiment results show that the proposed DL-UNet with the DSFF module and new Edgeloss performs best in several classic networks. Yanjiao Zhao, Zhihui Lai 0001, LinLin Shen, Heng Kong |
CBMS | 3 |
| 2022 | Contrastive learning-based Adenoid Hypertrophy Grading Network Using Nasoendoscopic ImageabstractAdenoid hypertrophy is a common disease in children with otolaryngology diseases. Otolaryngologists usually use nasoendoscopy for adenoid hypertrophy screening, which is however tedious and time-consuming for the grading. So far, artificial intelligence technology has not been applied to the grading of nasoendoscopic adenoid. In this work, we firstly propose a novel multi-scale grading network, MIB-ANet, for adenoid hypertrophy classification. And we further propose a contrastive learning-based network to alleviate the overfitting problem of the model caused by lacking of nasoendoscopic adenoid images with high-quality annotations. The experimental results show that MIB-ANet shows the best grading performance compared to four classic CNNs, i.e., AlexNet, VGG16, ResNet50 and GoogleNet. Take$F_{1}$score as an example, MIB-ANet achieves 1.38% higher$F_{1}$score than the best baseline CNN - AlexNet. Due to the capability of the contrastive learning-based pre-training strategy in exploring unannotated data, the pre-training using SimCLR pretext task can consistently improve the performance of MIB-ANet when different ratios of the labeled training data are employed. The MIB-ANet pre-trained by SimCLR pretext task achieves 4.41%, 2.64%, 3.10%, and 1.71% higher$F_{1}$score when 25%, 50%, 75% and 100% of the training data are labeled, respectively. Siting Zheng, Xuechen Li 0001, Mingmin Bi, Xiaoshan Feng, Yunping Fan, LinLin Shen |
CBMS | 8 |
| 2022 | CSL: A Large-scale Chinese Scientific Literature DatasetabstractScientific literature serves as a high-quality corpus, supporting a lot of Natural Language Processing (NLP) research. However, existing datasets are centered around the English language, which restricts the development of Chinese scientific NLP. In this work, we present CSL, a large-scale Chinese Scientific Literature dataset, which contains the titles, abstracts, keywords and academic fields of 396k papers. To our knowledge, CSL is the first scientific document dataset in Chinese. The CSL can serve as a Chinese corpus. Also, this semi-structured data is a natural annotation that can constitute many supervised NLP tasks. Based on CSL, we present a benchmark to evaluate the performance of models across scientific domain tasks, i.e., summarization, keyword generation and text classification. We analyze the behavior of existing text-to-text models on the evaluation tasks and reveal the challenges for Chinese scientific NLP tasks, which provides a valuable reference for future research. Data and code will be publicly available. Yudong Li 0001, Zhe Zhao 0006, LinLin Shen, Weijie Liu 0002, Weiquan Mao |
COLING | 4 |
| 2022 | Frequency-driven Imperceptible Adversarial Attack on Semantic SimilarityabstractCurrent adversarial attack research reveals the vulnerability of learning-based classifiers against carefully crafted perturbations. However, most existing attack methods have inherent limitations in cross-dataset generalization as they rely on a classification layer with a closed set of categories. Furthermore, the perturbations generated by these methods may appear in regions easily perceptible to the human visual system (HVS). To circumvent the former problem, we propose a novel algorithm that attacks semantic similarity on feature representations. In this way, we are able to fool classifiers without limiting attacks to a specific dataset. For imperceptibility, we introduce the low-frequency constraint to limit perturbations within high-frequency components, ensuring perceptual similarity between adversarial examples and originals. Extensive experiments on three datasets (CIFAR-10, CIFAR-100, and ImageNet-1K) and three public online platforms indicate that our attack can yield misleading and transferable adversarial examples across architectures and datasets. Additionally, visualization results and quantitative performance (in terms of four different metrics) show that the proposed algorithm generates more imperceptible perturbations than the state-of-the-art methods. Code is made available at https://github.com/LinQinLiang/SSAH-adversarial-attack. Qinliang Lin, Weicheng Xie 0001, Bizhu Wu, Jinheng Xie, LinLin Shen |
CVPR | 6 |
| 2022 | Scene Consistency Representation Learning for Video Scene SegmentationabstractA long-term video, such as a movie or TV show, is composed of various scenes, each of which represents a series of shots sharing the same semantic story. Spotting the correct scene boundary from the long-term video is a challenging task, since a model must understand the storyline of the video to figure out where a scene starts and ends. To this end, we propose an effective Self-Supervised Learning (SSL) framework to learn better shot representations from unlabeled long-term videos. More specifically, we present an SSL scheme to achieve scene consistency, while exploring considerable data augmentation and shuffling methods to boost the model generalizability. Instead of explicitly learning the scene boundary features as in the previous methods, we introduce a vanilla temporal model with less inductive bias to verify the quality of the shot features. Our method achieves the state-of-the-art performance on the task of Video Scene Segmentation. Additionally, we suggest a more fair and reasonable benchmark to evaluate the performance of Video Scene Segmentation methods. The code is made available.11https://github.com/TencentYoutuResearch/SceneSegmentation-SCRL. Haoqian Wu, Yanan Luo, Ruizhi Qiao, Bo Ren 0002, Weicheng Xie 0001, LinLin Shen |
CVPR | 8 |
| 2022 | CLIMS: Cross Language Image Matching for Weakly Supervised Semantic SegmentationabstractIt has been widely known that CAM (Class Activation Map) usually only activates discriminative object regions and falsely includes lots of object-related backgrounds. As only a fixed set of image-level object labels are available to the WSSS (weakly supervised semantic segmentation) model, it could be very difficult to suppress those diverse background regions consisting of open set objects. In this paper, we propose a novel Cross Language Image Matching (CLIMS) framework, based on the recently introduced Contrastive Language-Image Pre-training (CLIP) model, for WSSS. The core idea of our framework is to introduce natural language supervision to activate more complete object regions and suppress closely-related open background regions. In particular, we design object, background region and text label matching losses to guide the model to excite more reasonable object regions for CAM of each category. In addition, we design a co-occurring background suppression loss to prevent the model from activating closely-related background regions, with a predefined set of class-related background text descriptions. These designs enable the proposed CLIMS to generate a more complete and compact activation map for the target objects. Extensive experiments on PASCAL VOC2012 dataset show that our CLIMS significantly outperforms the previous state-of-the-art methods. Code will be available at https://github.com/CVI-SZU/CLIMS. Jinheng Xie, Xianxu Hou, Kai Ye 0004, LinLin Shen |
CVPR | 4 |
| 2022 | C2 AM: Contrastive learning of Class-agnostic Activation Map for Weakly Supervised Object Localization and Semantic SegmentationabstractWhile class activation map (CAM) generated by image classification network has been widely used for weakly su-pervised object localization (WSOL) and semantic segmentation (WSSS), such classifiers usually focus on discriminative object regions. In this paper, we propose Contrastive learning for Class-agnostic Activation Map (C2AM) generation only using unlabeled image data, without the involvement of image-level supervision. The core idea comes from the observation that i) semantic information of fore-ground objects usually differs from their backgrounds; ii) foreground objects with similar appearance or background with similar color/texture have similar representations in the feature space. We form the positive and negative pairs based on the above relations and force the network to disentangle foreground and background with a class-agnostic activation map using a novel contrastive loss. As the network is guided to discriminate cross-image foreground-background, the class-agnostic activation maps learned by our approach generate more complete object regions. We successfully extracted from C2AM class-agnostic object bounding boxes for object localization and background cues to refine CAM generated by classification network for semantic segmentation. Extensive experiments on CUB-200-2011, ImageNet-1K, and PASCAL VOC2012 datasets show that both WSOL and WSSS can benefit from the proposed C2AM. Code will be available at https://github.com/CVI-SZUICCAM. Jinheng Xie, Jianfeng Xiang, Junliang Chen 0002, Xianxu Hou, LinLin Shen |
CVPR | 6 |
| 2022 | RamGAN: Region Attentive Morphing GAN for Region-Level Makeup Transfer
Jianfeng Xiang, Junliang Chen 0002, Wenshuang Liu, Xianxu Hou, LinLin Shen |
ECCV (22) | 5 |
| 2022 | Which One is Better? Self-supervised Temporal Coherence Learning for Skeleton Based Action RecognitionabstractRecently, researchers have achieved significant results in the skeleton based action recognition task. To better model the skeleton sequences, existing methods learned the feature representations in the self-supervised setting by solving pretext tasks, such as predicting the order of a shuffled skeleton sequence or verifying whether a given skeleton sequence is shuffled or not. However, these pretext tasks are either too challenging or too easy for the encoder to obtain a proper skeleton representation for action recognition. Therefore, we propose a novel self-pretraining pretext task, Which One Is Better (WOIB), to identify which one is more temporally coherent, given two shuffled skeleton sequences. Experiments on the NTU RGB+D, NTU RGB+D 120, and Kinetics-Skeleton datasets with different network architectures show significant improvements in recognition accuracy, demonstrating that such a well-designed pretext task is general and able to drive the encoder to learn more discriminative representations. Bizhu Wu, Mingyan Wu, Haoqin Ji, LinLin Shen |
IJCB | 4 |
| 2022 | Learning Multi-dimensional Edge Feature-based AU Relation Graph for Facial Action Unit RecognitionabstractThe activations of Facial Action Units (AUs) mutually influence one another. While the relationship between a pair of AUs can be complex and unique, existing approaches fail to specifically and explicitly represent such cues for each pair of AUs in each facial display. This paper proposes an AU relationship modelling approach that deep learns a unique graph to explicitly describe the relationship between each pair of AUs of the target facial display. Our approach first encodes each AU's activation status and its association with other AUs into a node feature. Then, it learns a pair of multi-dimensional edge features to describe multiple task-specific relationship cues between each pair of AUs. During both node and edge feature learning, our approach also considers the influence of the unique facial display on AUs' relationship by taking the full face representation as an input. Experimental results on BP4D and DISFA datasets show that both node and edge feature learning modules provide large performance improvements for CNN and transformer-based backbones, with our best systems achieving the state-of-the-art AU recognition results. Our approach not only has a strong capability in modelling relationship cues for AU recognition but also can be easily incorporated into various backbones. Our PyTorch code is made available at https://github.com/CVI-SZU/ME-GraphAU. Siyang Song, Weicheng Xie 0001, LinLin Shen, Hatice Gunes |
IJCAI | 4 |
| 2022 | Point Beyond Class: A Benchmark for Weakly Semi-supervised Abnormality Localization in Chest X-Rays
Haoqin Ji, Yuexiang Li, Jinheng Xie, Nanjun He, Yawen Huang, Dong Wei 0004, Xinrong Chen, LinLin Shen, Yefeng Zheng 0001 |
MICCAI (3) | 9 |
| 2022 | Sample Hardness Based Gradient Loss for Long-Tailed Cervical Cell Detection
Minmin Liu, Xuechen Li 0001, Xiangbo Gao, Junliang Chen 0002, LinLin Shen, Huisi Wu |
MICCAI (2) | 5 |
| 2022 | Talk2Face: A Unified Sequence-based Framework for Diverse Face Generation and Analysis TasksabstractFacial analysis is an important domain in computer vision and has received extensive research attention. For numerous downstream tasks with different input/output formats and modalities, existing methods usually design task-specific architectures and train them using face datasets collected in the particular task domain. In this work, we proposed a single model, Talk2Face, to simultaneously tackle a large number of face generation and analysis tasks, e.g. text guided face synthesis, face captioning and age estimation. Specifically, we cast different tasks into a sequence-to-sequence format with the same architecture, parameters and objectives. While text and facial images are tokenized to sequences, the annotation labels of faces for different tasks are also converted to natural languages for unified representation. We collect a set of 2.3M face-text pairs from available datasets across different tasks, to train the proposed model. Uniform templates are then designed to enable the model to perform different downstream tasks, according to the task context and target. Experiments on different tasks show that our model achieves better face generation and caption performances than SOTA approaches. On age estimation and multi-attribute classification, our model reaches competitive performance with those models specially designed and trained for these particular tasks. In practice, our model is much easier to be deployed to different facial analysis related tasks. Code and dataset will be available at https://github.com/ydli-ai/Talk2Face. Yudong Li 0001, Xianxu Hou, Zhe Zhao 0006, LinLin Shen, Xuefeng Yang, Kimmo Yan |
ACM Multimedia | 4 |
| 2022 | Content and Gradient Model-driven Deep Network for Single Image Reflection RemovalabstractSingle image reflection removal (SIRR) is an extremely challenging, ill-posed problem with many application scenarios. In recent years, massive deep learning-based methods have been proposed to remove undesirable reflections from a single input image. However, these methods lack interpretability and do not fully utilize the intrinsic physical structure of reflection images. In this paper, we propose a content and gradient-guided deep network (CGDNet) for single image reflection removal, which is a full-interpretable and model-driven network. Firstly, using the multi-scale convolutional dictionary, we design a novel single image reflection removal model, which combines the image content prior and gradient prior information. Then, the model is optimized using an optimization algorithm based on the proximal gradient technique and unfolded into a neural network, i.e., CGDNet. All the parameters of CGDNet can be automatically learned by end-to-end training. Besides, we introduce a reflection detection module into CGDNet to obtain a probabilistic confidence map and ensure that the network pays attention to reflection regions. Extensive experiments on four benchmark datasets demonstrate that CGDNet is more efficient than state-of-the-art methods in terms of both subjective and objective evaluations. Code is available at https://github.com/zynwl/CGDNet. LinLin Shen, Qiufu Li |
ACM Multimedia | 2 |
| 2022 | WaveSNet: Wavelet Integrated Deep Networks for Image Segmentation
Qiufu Li, LinLin Shen |
PRCV (4) | 2 |
| 2022 | Neuron segmentation using 3D wavelet integrated encoder-decoder networkabstractMOTIVATION: 3D neuron segmentation is a key step for the neuron digital reconstruction, which is essential for exploring brain circuits and understanding brain functions. However, the fine line-shaped nerve fibers of neuron could spread in a large region, which brings great computational cost to the neuron segmentation. Meanwhile, the strong noises and disconnected nerve fibers bring great challenges to the task. RESULTS: In this article, we propose a 3D wavelet and deep learning-based 3D neuron segmentation method. The neuronal image is first partitioned into neuronal cubes to simplify the segmentation task. Then, we design 3D WaveUNet, the first 3D wavelet integrated encoder-decoder network, to segment the nerve fibers in the cubes; the wavelets could assist the deep networks in suppressing data noises and connecting the broken fibers. We also produce a Neuronal Cube Dataset (NeuCuDa) using the biggest available annotated neuronal image dataset, BigNeuron, to train 3D WaveUNet. Finally, the nerve fibers segmented in cubes are assembled to generate the complete neuron, which is digitally reconstructed using an available automatic tracing algorithm. The experimental results show that our neuron segmentation method could completely extract the target neuron in noisy neuronal images. The integrated 3D wavelets can efficiently improve the performance of 3D neuron segmentation and reconstruction. AVAILABILITYAND IMPLEMENTATION: The data and codes for this work are available at https://github.com/LiQiufu/3D-WaveUNet. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Qiufu Li, LinLin Shen |
Bioinform. | 2 |
| 2022 | Reasonable object detection guided by knowledge of global context and category relationship
Haoqin Ji, Kai Ye 0004, Qi Wan, LinLin Shen |
Expert Syst. Appl. | 4 |
| 2022 | TW-GAN: Topology and width aware GAN for retinal artery/vein classification
Wenting Chen, Kai Ma 0002, Wei Ji 0011, Cheng Bian, Chunyan Chu, LinLin Shen, Yefeng Zheng 0001 |
Medical Image Anal. | 7 |
| 2022 | DigestPath: A benchmark dataset with challenge review for the pathological detection and segmentation of digestive-system
Qian Da, Zhongyu Li 0002, Yanfei Zuo, Chenbin Zhang, Jingxin Liu 0005, Wen Chen 0001, Jiahui Li 0005, Dou Xu, Hongmei Yi, Zhe Wang 0043, Li Zhang 0040, Xianying He, Xiaofan Zhang 0002, Ke Mei, Chuang Zhu, Weizeng Lu, LinLin Shen, Jun Shi 0006, Jun Li 0106, Sreehari S, Ganapathy Krishnamurthi, Jiangcheng Yang, Tiancheng Lin 0001, Qingyu Song 0004, Xuechen Liu 0004, Simon Graham, Raja Muhammad Saad Bashir, Canqian Yang, Shaofei Qin, Xinmei Tian 0001, Jie Zhao 0014, Dimitris N. Metaxas, Hongsheng Li 0001, Chaofu Wang, Shaoting Zhang 0001 |
Medical Image Anal. | 21 |
| 2022 | Boundary regression-based reep neural network for thyroid nodule segmentation in ultrasound images
Zhihao Jin, Xuechen Li 0001, Yudi Zhang 0005, LinLin Shen, Zhihui Lai 0001, Heng Kong |
Neural Comput. Appl. | 4 |
| 2022 | GuidedStyle: Attribute knowledge guided style manipulation for semantic face editing
Xianxu Hou, Hanbang Liang, LinLin Shen, Zhihui Lai 0001, Jun Wan 0005 |
Neural Networks | 4 |
| 2022 | Spectral Representation of Behaviour Primitives for Depression AnalysisabstractDepression is a serious mental disorder affecting millions of people all over the world. Traditional clinical diagnosis methods are subjective, complicated and require extensive participation of clinicians. Recent advances in automatic depression analysis systems promise a future where these shortcomings are addressed by objective, repeatable, and readily available diagnostic tools to aid health professionals in their work. Yet there remain a number of barriers to the development of such tools. One barrier is that existing automatic depression analysis algorithms base their predictions on very brief sequential segments, sometimes as little as one frame. Another barrier is that existing methods do not take into account what the context of the measured behaviour is. In this article, we extract multi-scale video-level features for video-based automatic depression analysis. We propose to use automatically detected human behaviour primitives as the low-dimensional descriptor for each frame. We also propose two novel spectral representations, i.e., spectral heatmaps and spectral vectors, to represent video-level multi-scale temporal dynamics of expressive behaviour. Constructed spectral representations are fed to Convolution Neural Networks (CNNs) and Artificial Neural Networks (ANNs) for depression analysis. We conducted experiments on the AVEC 2013 and AVEC 2014 benchmark datasets to investigate the influence of interview tasks on depression analysis. In addition to achieving state of the art accuracy in severity of depression estimation, we show that the task conducted by the user matters, that fusion of a combination of tasks reaches highest accuracy, and that longer tasks are more informative than shorter tasks, up to a point. Siyang Song, Shashank Jaiswal, LinLin Shen, Michel F. Valstar |
IEEE Trans. Affect. Comput. | 3 |
| 2022 | Triplet Loss With Multistage Outlier Suppression and Class-Pair Margins for Facial Expression RecognitionabstractDeep metric based triplet loss has been widely used to enhance inter-class separability and intra-class compactness of network features. However, the margin parameters in the triplet loss for current approaches are usually fixed and not adaptive to the variations among different expression pairs. Meanwhile, outlier samples like faces with confusing expressions, occlusion and large head poses may be introduced during the selection of the hard triplets, which may deteriorate the generalization performance of the learned features for normal testing samples. In this work, a new triplet loss based on class-pair margins and multistage outlier suppression is proposed for facial expression recognition (FER). In this approach, each expression pair is assigned with an order-insensitive or two order-aware adaptive margin parameters. While expression samples with large head poses or occlusion are firstly detected and excluded, abnormal hard triplets are discarded if their feature distances do not fit the model of normal feature distance distribution. Extensive experiments on seven public benchmark expression databases show that the network using the proposed loss achieves much better accuracy than that using the original triplet loss and the network without using the proposed strategies, and the most balanced performances among state-of-the-art algorithms in the literature. Weicheng Xie 0001, Haoqian Wu, Mengchao Bai, LinLin Shen |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2022 | Fingerprint Presentation Attack Detector Using Global-Local ModelabstractThe vulnerability of automated fingerprint recognition systems (AFRSs) to presentation attacks (PAs) promotes the vigorous development of PA detection (PAD) technology. However, PAD methods have been limited by information loss and poor generalization ability, resulting in new PA materials and fingerprint sensors. This article thus proposes a global-local model-based PAD (RTK-PAD) method to overcome those limitations to some extent. The proposed method consists of three modules, called: 1) the global module; 2) the local module; and 3) the rethinking module. By adopting the cut-out-based global module, a global spoofness score predicted from nonlocal features of the entire fingerprint images can be achieved. While by using the texture in-painting-based local module, a local spoofness score predicted from fingerprint patches is obtained. The two modules are not independent but connected through our proposed rethinking module by localizing two discriminative patches for the local module based on the global spoofness score. Finally, the fusion spoofness score by averaging the global and local spoofness scores is used for PAD. Our experimental results evaluated on LivDet 2017 show that the proposed RTK-PAD can achieve an average classification error (ACE) of 2.28% and a true detection rate (TDR) of 91.19% when the false detection rate (FDR) equals 1.0%, which significantly outperformed the state-of-the-art methods by ~10% in terms of TDR (91.19% versus 80.74%). Wentian Zhang, Feng Liu 0013, Haoqian Wu, LinLin Shen |
IEEE Trans. Cybern. | 5 |
| 2022 | Fingerprint Presentation Attack Detection by Channel-Wise Feature DenoisingabstractDue to the diversity of attack materials, fingerprint recognition systems (AFRSs) are vulnerable to malicious attacks. It is thus important to propose effective fingerprint presentation attack detection (PAD) methods for the safety and reliability of AFRSs. However, current PAD methods often exhibit poor robustness under new attack types settings. This paper thus proposes a novel channel-wise feature denoising fingerprint PAD (CFD-PAD) method by handling the redundant noise information ignored in previous studies. The proposed method learns important features of fingerprint images by weighing the importance of each channel and identifying discriminative channels and “noise” channels. Then, the propagation of “noise” channels is suppressed in the feature map to reduce interference. Specifically, a PA-Adaptation loss is designed to constrain the feature distribution to make the feature distribution of live fingerprints more aggregate and that of spoof fingerprints more disperse. Experimental results evaluated on the LivDet 2017 dataset showed that the proposed CFD-PAD can achieve 2.53% average classification error (ACE) and a 93.83% true detection rate when the false detection rate equals 1.0% (TDR@FDR=1%). Also, the proposed method markedly outperforms the best single-model-based methods in terms of ACE (2.53% vs. 4.56%) and TDR@FDR=1%(93.83% vs. 73.32%), which demonstrates its effectiveness. Although we have achieved a comparable result with the state-of-the-art multiple-model-based methods, there still is an increase in TDR@FDR=1% from 91.19% to 93.83%. In addition, the proposed model is simpler, lighter and more efficient and has achieved a 74.76% reduction in computation time compared with the state-of-the-art multiple-model-based method.The source code is available athttps://github.com/kongzhecn/cfd-pad. Feng Liu 0013, Zhe Kong, Wentian Zhang, LinLin Shen |
IEEE Trans. Inf. Forensics Secur. | 5 |
| 2022 | Canonical Correlation Analysis With Low-Rank Learning for Image RepresentationabstractAs a multivariate data analysis tool, canonical correlation analysis (CCA) has been widely used in computer vision and pattern recognition. However, CCA uses Euclidean distance as a metric, which is sensitive to noise or outliers in the data. Furthermore, CCA demands that the two training sets must have the same number of training samples, which limits the performance of CCA-based methods. To overcome these limitations of CCA, two novel canonical correlation learning methods based on low-rank learning are proposed in this paper for image representation, named robust canonical correlation analysis (robust-CCA) and low-rank representation canonical correlation analysis (LRR-CCA). By introducing two regular matrices, the training sample numbers of the two training datasets can be set as any values without any limitation in the two proposed methods. Specifically, robust-CCA uses low-rank learning to remove the noise in the data and extracts the maximization correlation features from the two learned clean data matrices. The nuclear norm andL1-norm are used as constraints for the learned clean matrices and noise matrices, respectively. LRR-CCA introduces low-rank representation into CCA to ensure that the correlative features can be obtained in low-rank representation. To verify the performance of the proposed methods, five publicly image databases are used to conduct extensive experiments. The experimental results demonstrate the proposed methods outperform state-of-the-art CCA-based and low-rank learning methods. Yuwu Lu, Zhihui Lai 0001, LinLin Shen, Xuelong Li 0001 |
IEEE Trans. Image Process. | 5 |
| 2022 | Learning a Model-Driven Variational Network for Deformable Image RegistrationabstractData-driven deep learning approaches to image registration can be less accurate than conventional iterative approaches, especially when training data is limited. To address this issue and meanwhile retain the fast inference speed of deep learning, we propose VR-Net, a novel cascaded variational network for unsupervised deformable image registration. Using a variable splitting optimization scheme, we first convert the image registration problem, established in a generic variational framework, into two sub-problems, one with a point-wise, closed-form solution and the other one being a denoising problem. We then propose two neural layers (i.e. warping layer and intensity consistency layer) to model the analytical solution and a residual U-Net (termed generalized denoising layer) to formulate the denoising problem. Finally, we cascade the three neural layers multiple times to form our VR-Net. Extensive experiments on three (two 2D and one 3D) cardiac magnetic resonance imaging datasets show that VR-Net outperforms state-of-the-art deep learning methods on registration accuracy, whilst maintaining the fast inference speed of deep learning and the data-efficiency of variational models. Xi Jia, Alexander Thorley, Wei Chen 0092, Huaqi Qiu, LinLin Shen, Iain B. Styles, Hyung Jin Chang, Ales Leonardis, Antonio M. Simoes Monteiro de Marvao, Declan P. O'Regan, Daniel Rueckert, Jinming Duan 0001 |
IEEE Trans. Medical Imaging | 5 |
| 2022 | AniGAN: Style-Guided Generative Adversarial Networks for Unsupervised Anime Face GenerationabstractIn this paper, we propose a novel framework to translate a portrait photo-face into an anime appearance. Different from existing translation methods which do not designate specific styles, we aim to synthesize anime-faces which are style-consistent with a given reference anime-face. However, unlike typical translation tasks, such anime-face translation is particularly challenging due to the large and complex variations of appearances among anime-faces. Existing methods often fail to transfer the styles of reference anime-faces to the generated anime-faces, or introduce noticeable artifacts/distortions in the local shapes of their generated anime-faces. We propose a novel GAN-based anime-face translator, called AniGAN, to synthesize high-quality anime-faces. Specifically, a new generator architecture is proposed to simultaneously transfer color/texture styles and transform local facial shapes into anime-like counterparts based on the style of a reference anime-face, while preserving the global structure of the source photo-face. New normalization functions are designed for the generator to further improve local shape transformation and color/texture style transfer. Besides, we propose a double-branch discriminator to learn domain-specific distributions through individual branches and learn cross-domain shared distributions via shared layers, helping generate visually pleasing anime-faces and effectively mitigate artifacts/distortions. Extensive experiments on benchmark datasets qualitatively and quantitatively demonstrate the superiority of our method over state-of-the-art methods. Bing Li 0024, Yuanlue Zhu, Chia-Wen Lin, Bernard Ghanem, LinLin Shen |
IEEE Trans. Multim. | 6 |
| 2022 | Gated SwitchGAN for Multi-Domain Facial Image TranslationabstractRecent studies on multi-domain facial image translation have achieved impressive results. The existing methods generally provide a discriminator with an auxiliary classifier to impose domain translation. However, these methods neglect important information regarding domain distribution matching. To solve this problem, we propose a switch generative adversarial network (SwitchGAN) with a more adaptive discriminator structure and a matched generator to perform delicate image translation among multiple domains. A feature-switching operation is proposed to achieve feature selection and fusion in our conditional modules. We demonstrate the effectiveness of our model. Furthermore, we also introduce a new capability of our generator that represents attribute intensity control and extracts content information without tailored training. Experiments on the Morph, RaFD and CelebA databases visually and quantitatively show that our extended SwitchGAN (i.e., Gated SwitchGAN) can achieve better translation results than StarGAN, AttGAN and STGAN. The attribute classification accuracy achieved using the trained ResNet-18 model and the FID score obtained using the ImageNet pretrained Inception-v3 model also quantitatively demonstrate the superior performance of our models. Yuanlue Zhu, Wenting Chen, Wenshuang Liu, LinLin Shen |
IEEE Trans. Multim. | 5 |
| 2022 | Generalized Embedding Regression: A Framework for Supervised Feature ExtractionabstractSparse discriminative projection learning has attracted much attention due to its good performance in recognition tasks. In this article, a framework called generalized embedding regression (GER) is proposed, which can simultaneously perform low-dimensional embedding and sparse projection learning in a joint objective function with a generalized orthogonal constraint. Moreover, the label information is integrated into the model to preserve the global structure of data, and a rank constraint is imposed on the regression matrix to explore the underlying correlation structure of classes. Theoretical analysis shows that GER can obtain the same or approximate solution as some related methods with special settings. By utilizing this framework as a general platform, we design a novel supervised feature extraction approach called jointly sparse embedding regression (JSER). In JSER, we construct an intrinsic graph to characterize the intraclass similarity and a penalty graph to indicate the interclass separability. Then, the penalty graph Laplacian is used as the constraint matrix in the generalized orthogonal constraint to deal with interclass marginal points. Moreover, the$L_{2,1}$-norm is imposed on the regression terms for robustness to outliers and data’s variations and the regularization term for jointly sparse projection learning, leading to interesting semantic interpretability. An effective iterative algorithm is elaborately designed to solve the optimization problem of JSER. Theoretically, we prove that the subproblem of JSER is essentially an unbalanced Procrustes problem and can be solved iteratively. The convergence of the designed algorithm is also proved. Experimental results on six well-known data sets indicate the competitive performance and latent properties of JSER. Jianglin Lu, Zhihui Lai 0001, Yudong Chen 0002, Jie Zhou 0009, LinLin Shen |
IEEE Trans. Neural Networks Learn. Syst. | 6 |
| 2021 | Adversarial Defence by Diversified Simultaneous Training of Deep EnsemblesabstractLearning-based classifiers are susceptible to adversarial examples. Existing defence methods are mostly devised on individual classifiers. Recent studies showed that it is viable to increase adversarial robustness by promoting diversity over an ensemble of models. In this paper, we propose adversarial defence by encouraging ensemble diversity on learning high-level feature representations and gradient dispersion in simultaneous training of deep ensemble networks. We perform extensive evaluations under white-box and black-box attacks including transferred examples and adaptive attacks. Our approach achieves a significant gain of up to 52% in adversarial robustness, compared with the baseline and the state-of-the-art method on image benchmarks with complex data scenes. The proposed approach complements the defence paradigm of adversarial training, and can further boost the performance. The source code is available at https://github.com/ALIS-Lab/AAAI2021-PDD. Bo Huang 0017, Zhiwei Ke, Yi Wang 0017, Wei Wang 0011, LinLin Shen, Feng Liu 0013 |
AAAI | 5 |
| 2021 | Translate the Facial Regions You Like Using Self-Adaptive Region TranslationabstractWith the progression of Generative Adversarial Networks (GANs), image translation methods has achieved increasingly remarkable performance. However, most available methods can only achieve image level translation, which is unable to precisely control the regions to be translated. In this paper, we propose a novel self-adaptive region translation network (SART) for region-level translation, which uses region-adaptive instance normalization (RIN) and a region matching loss (RML) for this task. We first encode the style and content image for each region with style and content encoder. To translate both shape and texture of the target region, we inject region-adaptive style features into the decoder by RIN. To ensure independent translation among different regions, RML is proposed to measure the similarity between the non-translated/translated regions of content and translated images. Extensive experiments on three publicly available datasets, i.e. Morph, RaFD and CelebAMask-HQ, suggest that our approach demonstrate obvious improvement over state-of-the-art methods like StarGAN, SEAN and FUNIT. Our approach has further advantages in precise control of the regions to be translated. As a result, region level expression changes and step-by-step make-up can be achieved. The video demo is available at (https://youtu.be/DvIdmcR2LEc). Wenshuang Liu, Wenting Chen, Zhanjia Yang, LinLin Shen |
AAAI | 4 |
| 2021 | A Contrastive Learning-based PPC-UNet for Colorectal Histopathology Whole Slide Image SegmentationabstractColorectal cancer (CRC) is the third most common cancer and is usually diagnosed using colonoscopy and biopsy. Diagnosis of pathological biopsy requires professional knowledge and technology. Computer-aided gland and lesion segmentation systems have been proposed to help pathologists in diagnosis of CRC. However, to the best of our knowledge, there has not been a literature work trying to segment different levels of intraepithelial neoplasia in CRC pathological image. To reduce such a research gap, in this paper, we firstly collect a colorectal cancer biopsy histopathology whole slide image (WSI) dataset, named Histo-CRC Biopsy dataset, for algorithm evaluation. We further propose a PPC-UNet network to segment high level, low level intraepithelial neoplasia and normal tissues. The proposed PPC-UNet consists of two modules i.e., a UNet-based network for segmentation, and a pixel-to-propagation consistency (PPC) contrastive learning-based network for UNet encoder pre-training. As the important feature can be learned from the unannotated data during pre-training, our approach can consistently improve the Dice of UNet by around 2% when different ratios of the training data are labeled. Xuechen Li 0001, Jingxin Liu 0005, LinLin Shen, Kunming Sun, Suying Wang |
BIBM | 4 |
| 2021 | Classification and Localization Consistency Regularized Student-Teacher Network for Semi-supervised Cervical Cell DetectionabstractCytopathology image analysis gives an important indication of the cervical carcinoma. Automation-assisted diagnosis has received more and more attention because of its high efficiency. Thanks to the development of artificial intelligence, supervised deep learning methods have shown promising results for cervical cell detection task. However, large amounts of labeled data are quite expensive and time-consuming for acquisition. In this paper, we propose a Classification and Localization Consistency Regularized Student-Teacher Network (CLCR-STNet) with online pseudo label mining to leverage both labeled and unlabeled data for semi-supervised cervical cell detection. Both classification and localization consistency regularization are introduced to ensure that the bounding boxes predicted by the student and teacher networks are consistent. Instead of sharing the network parameters with student model, our teacher model is updated using exponential moving average (EMA). Moreover, the teacher network is used to generate high-confidence pseudo labels for unlabeled data to provide student network with more supervised information. The experiment results show that the proposed method outperforms the supervised methods learned using labeled data only. Menglu Zhang, Xuechen Li 0001, LinLin Shen |
CBMS | 3 |
| 2021 | FS-Net: Fast Shape-Based Network for Category-Level 6D Object Pose Estimation With Decoupled Rotation MechanismabstractIn this paper, we focus on category-level 6D pose and size estimation from a monocular RGB-D image. Previous methods suffer from inefficient category-level pose feature extraction, which leads to low accuracy and inference speed. To tackle this problem, we propose a fast shape-based network (FS-Net) with efficient category-level feature extraction for 6D pose estimation. First, we design an orientation aware autoencoder with 3D graph convolution for latent feature extraction. Thanks to the shift and scale-invariance properties of 3D graph convolution, the learned latent feature is insensitive to point shift and object size. Then, to efficiently decode category-level rotation information from the latent feature, we propose a novel decoupled rotation mechanism that employs two decoders to complementarily access the rotation information. For translation and size, we estimate them by two residuals: the difference between the mean of object points and ground truth translation, and the difference between the mean size of the category and ground truth size, respectively. Finally, to increase the generalization ability of the FS-Net, we propose an on-line box-cage based 3D deformation mechanism to augment the training data. Extensive experiments on two benchmark datasets show that the proposed method achieves state-of-the-art performance in both category- and instance-level 6D object pose estimation. Especially in category-level pose estimation, without extra synthetic data, our method outperforms existing methods by 6.3% on the NOCS-REAL dataset1. Wei Chen 0092, Xi Jia, Hyung Jin Chang, Jinming Duan 0001, LinLin Shen, Ales Leonardis |
CVPR | 5 |
| 2021 | Self-Attention Based Text Knowledge Mining for Text DetectionabstractPre-trained models play an important role in deep learning based text detectors. However, most methods ignore the gap between natural images and scene text images and directly apply ImageNet for pre-training. To address such a problem, some of them firstly pre-train the model using a large amount of synthetic data and then fine-tune it on target datasets, which is task-specific and has limited generalization capability. In this paper, we focus on providing general pre-trained models for text detectors. Considering the importance of exploring text contents for text detection, we propose STKM (Self-attention based Text Knowledge Mining), which consists of a CNN Encoder and a Self-attention Decoder, to learn general prior knowledge for text detection from SynthText. Given only image level text labels, Self-attention Decoder directly decodes features extracted from CNN Encoder to texts without requirement of detection, which guides the CNN backbone to explicitly learn discriminative semantic representations ignored by previous approaches. After that, the text knowledge learned by the backbone can be transferred to various text detectors to significantly improve their detection performance (e.g., 5.89% higher F-measure for EAST on ICDAR15 dataset) without bells and whistles. Pre-trained model is available at: https://github.com/CVI-SZU/STKM Qi Wan, Haoqin Ji, LinLin Shen |
CVPR | 3 |
| 2021 | Selective Multi-scale Learning for Object Detection
Junliang Chen 0002, Weizeng Lu, LinLin Shen |
ICANN (2) | 3 |
| 2021 | PointFace: Point Set Based Feature Learning for 3D Face RecognitionabstractThough 2D face recognition (FR) has achieved great success due to powerful 2D CNNs and large-scale training data, it is still challenged by extreme poses and illumination conditions. On the other hand, 3D FR has the potential to deal with aforementioned challenges in the 2D domain. However, most of available 3D FR works transform 3D surfaces to 2D maps and utilize 2D CNNs to extract features. The works directly processing point clouds for 3D FR is very limited in literature. To bridge this gap, in this paper, we propose a light-weight framework, named PointFace, to directly process point set data for 3D FR. Inspired by contrastive learning, our PointFace use two weight-shared encoders to directly extract features from a pair of 3D faces. A feature similarity loss is designed to guide the encoders to obtain discriminative face representations. We also present a pair selection strategy to generate positive and negative pairs to boost training. Extensive experiments on Lock3DFace and Bosphorus show that the proposed PointFace outperforms state-of-the-art 2D CNN based methods. Changyuan Jiang, Shisong Lin, Wei Chen 0092, Feng Liu 0013, LinLin Shen |
IJCB | 5 |
| 2021 | High Quality Facial Data Synthesis and Fusion for 3D Low-quality Face Recognitionabstract3D face recognition (FR) is a popular topic in computer vision, since 3D face data is invariant to pose and illumination condition changes which easily affect the performance of 2D FR. Though many 3D solutions have achieved impressive performances on public high-quality 3D face databases, few works concentrate on low-quality 3D FR. As the quality of 3D face acquired by widely used low-cost RGB-D sensors is really low, more robust methods are required to achieve satisfying performance on these 3D face data. To address this issue, we propose a novel two-stage pipeline to improve the performance of 3D FR. In the first stage, we utilize pix2pix network to restore the quality of low-quality face. In the second stage, we launch a multi-quality fusion network (MQFNet) to fuse the features from different qualities and enhance FR performance. Our proposed network achieves the state-of-the-art performance on the Lock3DFace database. Furthermore, extensive controlled experiments are conducted to demonstrate the effectiveness of each model of our network. Shisong Lin, Changyuan Jiang, Feng Liu 0013, LinLin Shen |
IJCB | 4 |
| 2021 | RamFace: Race Adaptive Margin Based Face Recognition for Racial Bias MitigationabstractRecent studies show that there exist significant racial bias among state-of-the-art (SOTA) face recognition algorithms, i.e., the accuracy for Caucasian is consistently higher than that for other races like African and Asian. To mitigate racial bias, we propose the race adaptive margin based face recognition (RamFace) model, designed under the multi-task learning framework with the race classification as the auxiliary task. The experiments show that the race classification task can enforce the model to learn the racial features and thus improve the discriminability of the extracted feature representations. In addition, a racial bias robust loss function, i.e., race adaptive margin loss, is proposed such that different optimal margins can be automatically derived for different races in training the model, which further mitigates the racial bias. The experimental results show that on RFW dataset, our model not only achieves SOTA face recognition accuracy but also mitigates the racial bias problem. Besides, RamFace is also tested on several public face recognition evaluation benchmarks, i.e., LFW, CPLFW and CALFW, and achieves better performance than the commonly used face recognition methods, which justifies the generalization capability of RamFace. Zhanjia Yang, Xiangping Zhu, Changyuan Jiang, Wenshuang Liu, LinLin Shen |
IJCB | 5 |
| 2021 | Group-wise Inhibition based Feature Regularization for Robust ClassificationabstractThe convolutional neural network (CNN) is vulnerable to degraded images with even very small variations (e.g. corrupted and adversarial samples). One of the possible reasons is that CNN pays more attention to the most discriminative regions, but ignores the auxiliary features when learning, leading to the lack of feature diversity for final judgment. In our method, we propose to dynamically suppress significant activation values of CNN by group-wise inhibition, but not fixedly or randomly handle them when training. The feature maps with different activation distribution are then processed separately to take the feature independence into account. CNN is finally guided to learn richer discriminative features hierarchically for robust classification according to the proposed regularization. Our method is comprehensively evaluated under multiple settings, including classification against corruptions, adversarial attacks and low data regime. Extensive experimental results show that the proposed method can achieve significant improvements in terms of both robustness and generalization performances, when compared with the state-of-the-art methods. Code is available at https://github.com/LinusWu/TENET_Training. Haoqian Wu, Weicheng Xie 0001, Feng Liu 0013, LinLin Shen |
ICCV | 5 |
| 2021 | Online Refinement of Low-level Feature Based Activation Map for Weakly Supervised Object LocalizationabstractWe present a two-stage learning framework for weakly supervised object localization (WSOL). While most previous efforts rely on high-level feature based CAMs (Class Activation Maps), this paper proposes to localize objects using the low-level feature based activation maps. In the first stage, an activation map generator produces activation maps based on the low-level feature maps in the classifier, such that rich contextual object information is included in an online manner. In the second stage, we employ an evaluator to evaluate the activation maps predicted by the activation map generator. Based on this, we further propose a weighted entropy loss, an attentive erasing, and an area loss to drive the activation map generator to substantially reduce the uncertainty of activations between object and background, and explore less discriminative regions. Based on the low-level object information preserved in the first stage, the second stage model gradually generates a well-separated, complete, and compact activation map of object in the image, which can be easily thresholded for accurate localization. Extensive experiments on CUB-200-2011 and ImageNet-1K datasets show that our framework surpasses previous methods by a large margin, which sets a new state-of-the-art for WSOL. Code will be available soon. Jinheng Xie, Xiangping Zhu, Ziqi Jin, Weizeng Lu, LinLin Shen |
ICCV | 6 |
| 2021 | Multi-scale Attention-Based Feature Pyramid Networks for Object Detection
Junliang Chen 0002, Minmin Liu, Kai Ye 0004, LinLin Shen |
ICIG (1) | 5 |
| 2021 | Delving into the Scale Variance Problem in Object DetectionabstractObject detection has made substantial progress in the last decade, due to the capability of convolution in extracting local context of objects. However, the scales of objects are diverse and current convolution can only process single-scale input. The capability of traditional convolution with a fixed receptive field in dealing with such a scale variance problem, is thus limited. Multiscale feature representation has been proven to be an effective way to mitigate the scale variance problem. Recent researches mainly adopt partial connection with certain scales, or aggregate features from all scales and focus on the global information across the scales. However, the information across spatial and depth dimensions is ignored. Inspired by this, we propose the multi-scale convolution (MSConv) to handle this problem. Taking into consideration scale, spatial and depth information at the same time, MSConv is able to process multi-scale input more comprehensively. MSConv is effective and computationally efficient, with only a small increase of computational cost. For most of the single-stage object detectors, replacing the traditional convolutions with MSConvs in the detection head can bring more than 2.5% improvement in AP (on COCO 2017 dataset), with only 3% increase of FLOPs. MSConv is also flexible and effective for two-stage object detectors. When extended to the mainstream two-stage object detectors, MSConv can bring up to 3.0% improvement in AP. Our best model under single-scale testing achieves 48.9% AP on COCO 2017 test-dev split, which surpasses many state-of-the-art methods. Junliang Chen 0002, LinLin Shen |
ICTAI | 3 |
| 2021 | Selective Adversarial Adaptation Learning via Exclusive Regularization for Partial Domain AdaptationabstractIn consideration of the suitability for the application scenario, partial domain adaptation is more significant and more valuable than traditional domain adaptation. Most existing partial domain adaptation methods adopt weighting mechanism to avoid negative migration which is caused by outlier classes samples. However, these methods give the equal consideration of each category in the source domain and determine the classes weight by classifier or discriminator, and they do not consider the possible misprediction of the similar samples from classes which are difficult to distinguish in the source domain. This situation may cause the misalignment of the outlier source classes and target classes, and the wrong alignment of the discriminators. In this work, we propose a selective adversarial adaptation learning method via exclusive regularization for partial domain adaptation (ERPDA) to solve these problems. Specifically, we utilize the exclusive regularization to extend the distance between samples of different classes in source domain to learn an inter-class separable discriminant representation to avoid negative transfer. Meanwhile, the positive transfer is performed by Joint Maximum Mean Discrepancy (JMMD) based on selective adaptation adversarial learning via multi-discriminator. Extensive experiments show that ERPDA achieves state-of-the-art results on several partial domain adaptation benchmark datasets. Ping Li 0021, LinLin Shen, Lei Wu 0010, Qian Wang 0001, Chuang Zhao 0001 |
IJCNN | 2 |
| 2021 | GazeFlow: Gaze Redirection with Normalizing FlowsabstractGaze estimation often requires a large scale datasets with well annotated gaze information to train the estimator. However, such a dataset requires costive annotation and is usually very difficult to collect. Therefore, a number of gaze redirection approaches have been proposed to address such a problem. However, existing methods lack the ability to precisely synthesize images with target gaze and head pose in complex lighting scenes. As a powerful technique to model the distribution of given data, normalizing flows have the ability to generate photo-realistic images and provide flexible latent space manipulation. In this work, we present a novel flow-based generative model, GazeFlow11The code will be made available at https://github.com/CVI-SZU/GazeFlow, for gaze redirection. The visual results of gaze redirection show that the quality of eye images synthesized by GazeFlow is significantly higher than that of other approaches like Deep Warp and PRGAN. Our approach has also been applied to augment the training data to improve the accuracy of gaze estimators and significant improvement has been achieved for both within dataset and cross dataset experiments. Hanbang Liang, Xianxu Hou, LinLin Shen |
IJCNN | 4 |
| 2021 | Think About Boundary: Fusing Multi-level Boundary Information for Landmark Heatmap RegressionabstractAlthough current face alignment algorithms have obtained pretty good performances at predicting the location of facial landmarks, huge challenges remain for faces with severe occlusion and large pose variations, etc. On the contrary, semantic location of facial boundary is more likely to be reserved and estimated on these scenes. Therefore, we study a two-stage but end-to-end approach for exploring the relationship between the facial boundary and landmarks to get boundary-aware landmark predictions, which consists of two modules: the self-calibrated boundary estimation (SCBE) module and the boundary-aware landmark transform (BALT) module. In the SCBE module, we modify the stem layers and employ intermediate supervision to help generate high-quality facial boundary heatmaps. Boundary-aware features inherited from the SCBE module are integrated into the BALT module in a multi-scale fusion framework to better model the transformation from boundary to landmark heatmap. Experimental results conducted on the challenging benchmark datasets demonstrate that our approach outperforms state-of-the-art methods in the literature. Code will be available soon. Jinheng Xie, Jun Wan 0005, LinLin Shen, Zhihui Lai 0001 |
IJCNN | 3 |
| 2021 | SSFlow: Style-guided Neural Spline Flows for Face Image ManipulationabstractSignificant progress has been made in high-resolution and photo-realistic image generation by Generative Adversarial Networks (GANs). However, the generation process is still lack of control, which is crucial for semantic face editing. Furthermore, it remains challenging to edit target attributes and preserve the identity at the same time. In this paper, we propose SSFlow to achieve identity-preserved semantic face manipulation in StyleGAN latent space based on conditional Neural Spline Flows. To further improve the performance of Neural Spline Flows on such task, we also propose Constractive Squash component and Blockwise 1 x 1 Convolution layer. Moreover, unlike other conditional flow-based approaches that require facial attribute labels during inference, our method can achieve label-free manipulation in a more flexible way. As a result, our methods are able to perform well-disentangled edits along various attributes, and generalize well for both real and artistic face image manipulation. Qualitative and quantitative evaluations show the advantages of our method for semantic face manipulation over state-of-the-art approaches. Hanbang Liang, Xianxu Hou, LinLin Shen |
ACM Multimedia | 3 |
| 2021 | Personality Recognition by Modelling Person-specific Cognitive Processes using Graph RepresentationabstractRecent research shows that in dyadic and group interactions individuals' nonverbal behaviours are influenced by the behaviours of their conversational partner(s). Therefore, in this work we hypothesise that during a dyadic interaction, the target subject's facial reactions are driven by two main factors: (i) their internal (person-specific) cognition, and (ii) the externalised nonverbal behaviours of their conversational partner. Subsequently, our novel proposition is to simulate and represent the target subject's (i.e., the listener) cognitive process in the form of a person-specific CNN architecture whose input is the audio-visual non-verbal cues displayed by the conversational partner (i.e., the speaker), and the output is the target subject's (i.e., the listener) facial reactions. We then undertake a search for the optimal CNN architecture whose results are used to create a person-specific graph representation for recognising the target subject's personality. The graph representation, fortified with a novel end-to-end edge feature learning strategy, helps with retaining both the unique parameters of the person-specific CNN and the geometrical relationship between its layers. Consequently, the proposed approach is the first work that aims to recognize the true (self-reported) personality of a target subject (i.e., the listener) from the learned simulation of their cognitive process (i.e., parameters of the person-specific CNN). The experimental results show that the CNN architectures are well associated with target subjects' personality traits and the proposed approach clearly outperforms multiple existing approaches that predict personality directly from non-verbal behaviours. In light of these findings, this work opens up a new avenue of research for predicting and recognizing socio-emotional phenomena (personality, affect, engagement etc.) from simulations of person-specific cognitive processes. Zilong Shao, Siyang Song, Shashank Jaiswal, LinLin Shen, Michel F. Valstar, Hatice Gunes |
ACM Multimedia | 4 |
| 2021 | Imbalanced data learning by minority class augmentation using capsule adversarial networks
Pourya Shamsolmoali, Masoumeh Zareapoor, LinLin Shen, Abdul Hamid Sadka, Jie Yang 0002 |
Neurocomputing | 3 |
| 2021 | Robust facial landmark detection by cross-order cross-semantic deep network
Jun Wan 0005, Zhihui Lai 0001, LinLin Shen, Jie Zhou 0009, Can Gao, Xianxu Hou |
Neural Networks | 3 |
| 2021 | Surrogate network-based sparseness hyper-parameter optimization for deep expression recognition
Weicheng Xie 0001, Wenting Chen, LinLin Shen, Jinming Duan 0001, Meng Yang 0001 |
Pattern Recognit. | 3 |
| 2021 | Two-dimensional jointly sparse robust discriminant regression
Zhihui Lai 0001, Zhuozhen Yu, Heng Kong, LinLin Shen |
Signal Process. Image Commun. | 4 |
| 2021 | Locality Preserving Robust Regression for Jointly Sparse Subspace LearningabstractAs the extended version of conventional Ridge Regression, L2,1-norm based ridge regression learning methods have been widely used in subspace learning since they are more robust than Frobenius norm based regression and meanwhile guarantee joint sparsity. However, conventional L2,1-norm regression methods encounter the small-class problem and meanwhile ignore the local geometric structures, which degrade their performances. To address these problems, we propose a novel regression method called Locality Preserving Robust Regression (LPRR). In addition to using the L2,1-norm for jointly sparse regression, we also utilize capped L2-norm in loss function to further enhance the robustness of the proposed algorithm. Moreover, to make use of local structure information, we also integrate the property of locality preservation into our model since it is of great importance in dimensionality reduction. The convergence analysis and computational complexity of the proposed iterative algorithm are presented. Experimental results on four datasets indicate that the proposed LPRR performs better than some famous subspace learning methods in classification tasks. Zhihui Lai 0001, Xuechen Li 0001, Yudong Chen 0002, Dongmei Mo, Heng Kong, LinLin Shen |
IEEE Trans. Circuits Syst. Video Technol. | 7 |
| 2021 | Adaptive Weighting of Handcrafted Feature Losses for Facial Expression RecognitionabstractDue to the importance of facial expressions in human-machine interaction, a number of handcrafted features and deep neural networks have been developed for facial expression recognition. While a few studies have shown the similarity between the handcrafted features and the features learned by deep network, a new feature loss is proposed to use feature bias constraint of handcrafted and deep features to guide the deep feature learning during the early training of network. The feature maps learned with and without the proposed feature loss for a toy network suggest that our approach can fully explore the complementarity between handcrafted features and deep features. Based on the feature loss, a general framework for embedding the traditional feature information into deep network training was developed and tested using the FER2013, CK+, Oulu-CASIA, and MMI datasets. Moreover, adaptive loss weighting strategies are proposed to balance the influence of different losses for different expression databases. The experimental results show that the proposed feature loss with adaptive weighting achieves much better accuracy than the original handcrafted feature and the network trained without using our feature loss. Meanwhile, the feature loss with adaptive weighting can provide complementary information to compensate for the deficiency of a single feature. Weicheng Xie 0001, LinLin Shen, Jinming Duan 0001 |
IEEE Trans. Cybern. | 2 |
| 2021 | Memristive Quantized Neural Networks: A Novel Approach to Accelerate Deep Learning On-ChipabstractExisting deep neural networks (DNNs) are computationally expensive and memory intensive, which hinder their further deployment in novel nanoscale devices and applications with lower memory resources or strict latency requirements. In this paper, a novel approach to accelerate on-chip learning systems using memristive quantized neural networks (M-QNNs) is presented. A real problem of multilevel memristive synaptic weights due to device-to-device (D2D) and cycle-to-cycle (C2C) variations is considered. Different levels of Gaussian noise are added to the memristive model during each adjustment. Another method of using memristors with binary states to build M-QNNs is presented, which suffers from fewer D2D and C2C variations compared with using multilevel memristors. Furthermore, methods of solving the sneak path issues in the memristive crossbar arrays are proposed. The M-QNN approach is evaluated on two image classification datasets, that is, ten-digit number and handwritten images of mixed National Institute of Standards and Technology (MNIST). In addition, input images with different levels of zero-mean Gaussian noise are tested to verify the robustness of the proposed method. Another highlight of the proposed method is that it can significantly reduce computational time and memory during the process of image recognition. Yang Zhang 0012, Menglin Cui, LinLin Shen, Zhigang Zeng |
IEEE Trans. Cybern. | 3 |
| 2021 | Memristive Fuzzy Deep Learning SystemsabstractAs a novel nanoscale device, the memristor has elicited widespread interest in implementing compact and efficient neurocomputing systems in the hardware. In this article, fuzzy deep learning systems using fuzzy memristive modeling methods are presented. One key issue in memristive modeling is the device variation issue due to device-to-device and cycle-to-cycle variations. It is very difficult to distinguish the intermediate memristive states in a multilevel memristor. A fuzzy modeling method is therefore proposed to define the memristive states in a dynamic way. In addition, fuzzy deep learning systems are presented for fuzzy pattern recognition such as image recognition. To the authors' best knowledge, this is the first work utilizing memristive fuzzy deep learning systems to realize image recognition. The expected input and output in the memristive fuzzy deep learning systems are refined via a bidirectional fuzzy rule. The effectiveness of the proposed fuzzy methods has been verified with comprehensive methods, such as single-layer neural networks, multi-layer neural networks, convolutional neural networks, and k-nearest neighbor method. Another highlight of the proposed fuzzy deep learning system is that there is a great reduction in memory and a significant increase in the speed for image recognition tasks with the same level of testing accuracy. Yang Zhang 0012, Menglin Cui, LinLin Shen, Zhigang Zeng |
IEEE Trans. Fuzzy Syst. | 3 |
| 2021 | WaveCNet: Wavelet Integrated CNNs to Suppress Aliasing Effect for Noise-Robust Image ClassificationabstractThough widely used in image classification, convolutional neural networks (CNNs) are prone to noise interruptions, i.e. the CNN output can be drastically changed by small image noise. To improve the noise robustness, we try to integrate CNNs with wavelet by replacing the common down-sampling (max-pooling, strided-convolution, and average pooling) with discrete wavelet transform (DWT). We firstly propose general DWT and inverse DWT (IDWT) layers applicable to various orthogonal and biorthogonal discrete wavelets like Haar, Daubechies, and Cohen, etc., and then design wavelet integrated CNNs (WaveCNets) by integrating DWT into the commonly used CNNs (VGG, ResNets, and DenseNet). During the down-sampling, WaveCNets apply DWT to decompose the feature maps into the low-frequency and high-frequency components. Containing the main information including the basic object structures, the low-frequency component is transmitted into the following layers to generate robust high-level features. The high-frequency components are dropped to remove most of the data noises. The experimental results show that WaveCNets achieve higher accuracy on ImageNet than various vanilla CNNs. We have also tested the performance of WaveCNets on the noisy version of ImageNet, ImageNet-C and six adversarial attacks, the results suggest that the proposed DWT/IDWT layers could provide better noise-robustness and adversarial robustness. When applying WaveCNets as backbones, the performance of object detectors (i.e., faster R-CNN and RetinaNet) on COCO detection dataset are consistently improved. We believe that suppression of aliasing effect, i.e. separation of low frequency and high frequency information, is the main advantages of our approach. The code of our DWT/IDWT layer and different WaveCNets are available at https://github.com/CVI-SZU/WaveCNet. Qiufu Li, LinLin Shen, Sheng Guo 0005, Zhihui Lai 0001 |
IEEE Trans. Image Process. | 2 |
| 2021 | One-Class Fingerprint Presentation Attack Detection Using Auto-Encoder NetworkabstractAutomated Fingerprint Recognition Systems (AFRSs) have been threatened by Presentation Attack (PA) since its existence. It is thus desirable to develop effective presentation attack detection (PAD) methods. However, the unpredictable PAs make PAD be a challenging problem. This paper proposes a novel One-Class PAD (OCPAD) method for Optical Coherence Technology (OCT) images based fingerprint PA detection. The proposed OCPAD model is learned from a training set only consists of Bonafides (i.e. real fingerprints). The reconstruction error and latent code obtained from the trained auto-encoder network in the proposed model is taken as the basis for the following spoofness score calculation. To get more accurate reconstruction error, we propose an activation map based weighting model to further refine the accuracy of reconstruction error. We test different statistics and distance measures and finally use a decision level fusion to make the final prediction. Our experiments are performed using a dataset with 93200 bonafide scans and 48400 PA scans. The results show that the proposed OCPAD can achieve a True Positive Rate (TPR) of 99.43% when the False Positive Rate (FPR) equals to 10% and a TPR of 96.59% when FPR=5%, which significantly outperformed a feature based approach and a supervised learning based model requiring PAs for training. Feng Liu 0013, Wentian Zhang, Guojie Liu, LinLin Shen |
IEEE Trans. Image Process. | 5 |
| 2021 | Orthogonalization-Guided Feature Fusion Network for Multimodal 2D+3D Facial Expression RecognitionabstractAs 2D and 3D data present different views of the same face, the features extracted from them can be both complementary and redundant. In this paper, we present a novel and efficient orthogonalization-guided feature fusion network, namely OGF$^2$Net, to fuse the features extracted from 2D and 3D faces for facial expression recognition. While 2D texture maps are fed into a 2D feature extraction pipeline (FE2DNet), the attribute maps generated from 3D data are concatenated as input of the 3D feature extraction pipeline (FE3DNet). The two networks are separately trained at the first stage and frozen in the second stage for late feature fusion, which can well address the unavailability of a large number of 3D+2D face pairs. To reduce the redundancies among features extracted from 2D and 3D streams, we design an orthogonal loss-guided feature fusion network to orthogonalize the features before fusing them. Experimental results show that the proposed method significantly outperforms the state-of-the-art algorithms on both the BU-3DFE and Bosphorus databases. While accuracies as high as 89.05% (P1 protocol) and 89.07% (P2 protocol) are achieved on the BU-3DFE database, an accuracy of 89.28% is achieved on the Bosphorus database. The complexity analysis also suggests that our approach achieves a higher processing speed while simultaneously requiring lower memory costs. Shisong Lin, Mengchao Bai, Feng Liu 0013, LinLin Shen, Yicong Zhou |
IEEE Trans. Multim. | 4 |
| 2020 | Group-Wise Dynamic Dropout Based on Latent Semantic VariationsabstractDropout regularization has been widely used in various deep neural networks to combat overfitting. It works by training a network to be more robust on information-degraded data points for better generalization. Conventional dropout and variants are often applied to individual hidden units in a layer to break up co-adaptations of feature detectors. In this paper, we propose an adaptive dropout to reduce the co-adaptations in a group-wise manner by coarse semantic information to improve feature discriminability. In particular, we showed that adjusting the dropout probability based on local feature densities can not only improve the classification performance significantly but also enhance the network robustness against adversarial examples in some cases. The proposed approach was evaluated in comparison with the baseline and several state-of-the-art adaptive dropouts over four public datasets of Fashion-MNIST, CIFAR-10, CIFAR-100 and SVHN. Zhiwei Ke, Zhiwei Wen, Weicheng Xie 0001, Yi Wang 0017, LinLin Shen |
AAAI | 5 |
| 2020 | Backbone Based Feature Enhancement for Object Detection
Haoqin Ji, Weizeng Lu, LinLin Shen |
ACCV (3) | 3 |
| 2020 | Automatic Primary Gross Tumor Volume Segmentation for Nasopharyngeal Carcinoma using ResSE-UNetabstractNasopharyngeal carcinoma (NPC) is an endemic disease within specific regions in the world. Radiotherapy is the standard treatment for NPC and accurate segmentation of primary gross tumor volume (GTV) is a critical process of continue therapy. In this paper we proposed a ResSE-UNet network and a Ternary Cross-Entropy (TCE) loss function for delineation of GTV. ResSE-UNet employed ResSE blocks to replace convolutional blocks in the original UNet to extract better features, and reduced the number of down-sampling processing to keep relatively high resolution of the images. TCE combined dice loss and Binary cross-entropy loss for larger gradient and better stability in training. The experimental results showed that among all combinations of networks and loss functions, the ResSE-UNet with TCE loss achieved the best segmentation performance, i.e. about 0.84 DSC can be obtained. Zhihao Jin, Xuechen Li 0001, LinLin Shen, Jinyi Lang, Junxiang Wu, Jiang Duan |
CBMS | 3 |
| 2020 | Wavelet Integrated CNNs for Noise-Robust Image ClassificationabstractConvolutional Neural Networks (CNNs) are generally prone to noise interruptions, i.e., small image noise can cause drastic changes in the output. To suppress the noise effect to the final predication, we enhance CNNs by replacing max-pooling, strided-convolution, and average-pooling with Discrete Wavelet Transform (DWT). We present general DWT and Inverse DWT (IDWT) layers applicable to various wavelets like Haar, Daubechies, and Cohen, etc., and design wavelet integrated CNNs (WaveCNets) using these layers for image classification. In WaveCNets, feature maps are decomposed into the low-frequency and high-frequency components during the down-sampling. The low-frequency component stores main information including the basic object structures, which is transmitted into the subsequent layers to extract robust high-level features. The high-frequency components, containing most of the data noise, are dropped during inference to improve the noise-robustness of the WaveCNets. Our experimental results on ImageNet and ImageNet-C (the noisy version of ImageNet) show that WaveCNets, the wavelet integrated versions of VGG, ResNets, and DenseNet, achieve higher accuracy and better noise-robustness than their vanilla versions. Qiufu Li, LinLin Shen, Sheng Guo 0005, Zhihui Lai 0001 |
CVPR | 2 |
| 2020 | Geometry Constrained Weakly Supervised Object Localization
Weizeng Lu, Xi Jia, Weicheng Xie 0001, LinLin Shen, Yicong Zhou, Jinming Duan 0001 |
ECCV (26) | 4 |
| 2020 | Self-Supervised CycleGAN for Object-Preserving Image-to-Image Domain Adaptation
Xinpeng Xie, Jiawei Chen 0009, Yuexiang Li, LinLin Shen, Kai Ma 0002, Yefeng Zheng 0001 |
ECCV (20) | 4 |
| 2020 | Feature map masking based single-stage face detectionabstractAlthough great progress has been made in face detection, a trade-off between speed and accuracy is still a great challenge. We propose in this paper a feature map masking based approach for single-stage face detection. As feature maps extracted from feature pyramid network might contain face unrelated features, we propose a mask generation branch to predict those significant units for face detection. The masked feature maps, where only important features are left, are then passed through the following detection process. Ground truth masks, directly generated from the training images, based on the face bounding boxes, are used to train the feature mask generation module. A mask constrained dropout module has also been proposed to drop out significant units of the shared feature maps, such that the detection performance can be further improved. The proposed approach is extensively tested using the WIDER FACE dataset. The results suggest that our detector with ResNet-152 backbone, achieves the best precision-recall performance among competing methods. As high as 95.4%, 94.0% and 86.9% accuracies have been achieved on the easy, medium and hard subsets, respectively. Junliang Chen 0002, Weicheng Xie 0001, LinLin Shen |
IJCB | 4 |
| 2020 | Feature Embedding Based Text Instance Grouping for Largely Spaced and Occluded Text DetectionabstractA text instance can be easily detected as multiple ones due to the large space between texts/characters, curved shape and partial occlusion. In this paper, a feature embedding based text instance grouping algorithm is proposed to solve this problem. To learn the feature space, a TIEM (Text Instance Embedding Module) is trained to minimize the within instance scatter and maximize the between instance scatter. Similarity between different text instances are measured in the feature space and merged if they meet certain conditions. Experimental results show that our approach can effectively connect text regions that belong to the same text instance. Competitive performance of our approach has been achieved on CTW1500, Total-Text, IC15 and a subset consists of texts selected from the three datasets, with large spacing and occlusions. Pan Gao 0002, Qi Wan, Renwu Gao, LinLin Shen |
ICPR | 4 |
| 2020 | SATGAN: Augmenting Age Biased Dataset for Cross-Age Face RecognitionabstractIn this paper, we propose a Stable Age Translation GAN (SATGAN) to generate fake face images at different ages to augment age biased face datasets for Cross-Age Face Recognition (CAFR). The proposed SATGAN consists of both generator and discriminator. As a part of the generator, a novel Mask Attention Module (MAM) is introduced to make the generator focus on the face area. In addition, the generator employs a Uniform Distribution Discriminator (UDD) to supervise the learning of latent feature map and enforce the uniform distribution. Besides, the discriminator employs a Feature Separation Module (FSM) to disentangle identity information from the age information. The quantitative and qualitative evaluations on Morph dataset prove that SATGAN achieves much better performance than existing methods. The face recognition model trained using dataset (VGGFace2 and MS-Celeb-lM) augmented using our SATGAN achieves better accuracy on cross age dataset like Cross-Age LFW and AgeDB-30. Wenshuang Liu, Wenting Chen, Yuanlue Zhu, LinLin Shen |
ICPR | 4 |
| 2020 | One-stage Multi-task Detector for 3D Cardiac MR ImagingabstractFast and accurate landmark location and bounding box detection are important steps in 3D medical imaging. In this paper, we propose a novel multi-task learning framework, for real-time, simultaneous landmark location and bounding box detection in 3D space. Our method extends the famous single-shot multibox detector (SSD) from single-task learning to multitask learning and from 2D to 3D. Furthermore, we propose a post-processing approach to refine the network landmark output, by averaging the candidate landmarks. Owing to these settings, the proposed framework is fast and accurate. For 3D cardiac magnetic resonance (MR) images with size 224×224×64, our framework runs ~128 volumes per second (VPS) on GPU and achieves 6.75mm average point-to-point distance error for landmark location, which outperforms both state-of-the-art and baseline methods. We also show that segmenting the 3D image cropped with the bounding box results in both improved performance and efficiency. Weizeng Lu, Xi Jia, Wei Chen 0092, Nicolò Savioli, Antonio M. Simoes Monteiro de Marvao, LinLin Shen, Declan P. O'Regan, Jinming Duan 0001 |
ICPR | 6 |
| 2020 | Self-supervised learning of Dynamic Representations for Static ImagesabstractFacial actions are spatio-temporal signals by nature, and therefore their modeling is crucially dependent on the availability of temporal information. In this paper, we focus on inferring such temporal dynamics of facial actions when no explicit temporal information is available, i.e. from still images. We present a novel self-supervised learning approach to capture multiple scales of temporal dynamics, with an application to facial Action Unit (AU) intensity estimation and dimensional affect estimation. In particular: 1. We propose a framework that infers a dynamic representation (DR) from a still image, capturing the bi-directional flow of time within a short time-window centered at the input image; 2. We show that the proposed rank loss can apply facial temporal evolution to self-supervise the training process without using target representations, allowing the network to represent dynamics more broadly; 3. We propose a multiple temporal scale approach that infers DRs for different window lengths (MDR) from a still image. We empirically validate the value of our approach on the task of frame ranking, and show how our proposed MDR attains state of the art results on BP4D for AU intensity estimation and on SEMAINE for dimensional affect estimation, using only still images at test time. Siyang Song, Enrique Sanchez, LinLin Shen, Michel F. Valstar |
ICPR | 3 |
| 2020 | Group-wise Feature Orthogonalization and Suppression for GAN based Facial Attribute TranslationabstractGenerative Adversarial Network (GAN) has been widely used for object attribute editing. However, the semantic correlation, resulted from the feature map interaction in the generative network of GAN, may impair the generalization ability of the generative network. In this work, semantic disentanglement is introduced in GAN to reduce the attribute correlation. The feature maps of the generative network are first grouped with an efficient clustering algorithm based on hash encoding, which are used to excavate hidden semantic attributes and calculate the group-wise orthogonality loss for the reduction of attribute entanglement. Meanwhile, the feature maps falling in the intersection regions of different groups are further suppressed to reduce the attribute-wise interaction. Extensive experiments reveal that the proposed GAN generated more genuine objects than the state of the arts. Quantitative results of classification accuracy, inception score and FID score further justify the effectiveness of the proposed GAN. Zhiwei Wen, Haoqian Wu, Weicheng Xie 0001, LinLin Shen |
ICPR | 4 |
| 2020 | Pose-aware Multi-feature Fusion Network for Driver Distraction RecognitionabstractTraffic accidents caused by distracted driving have gradually increased in recent years. In this work, we propose a novel multi-feature fusion network based on pose estimation, for image based distracted driving detection. Since hand is the most important part of driver to infer the distracted actions, our proposed method firstly detects hands using the human body posture information. In addition to the features extracted from the whole image, our network also include the important information of hand and human body posture. The global feature, hand and pose features are finally fused by weighted combination of probability vectors and concatenation of feature maps. The experimental results show that our method achieves state-of-the-art performance on our own SZ Bus Driver dataset and the public AUC Distracted Driver dataset. Mingyan Wu, LinLin Shen |
ICPR | 3 |
| 2020 | TR-GAN: Topology Ranking GAN with Triplet Loss for Retinal Artery/Vein Classification
Wenting Chen, Kai Ma 0002, Cheng Bian, Chunyan Chu, LinLin Shen, Yefeng Zheng 0001 |
MICCAI (5) | 7 |
| 2020 | MI2GAN: Generative Adversarial Network for Medical Image Domain Adaptation Using Mutual Information Constraint
Xinpeng Xie, Jiawei Chen 0009, Yuexiang Li, LinLin Shen, Kai Ma 0002, Yefeng Zheng 0001 |
MICCAI (2) | 4 |
| 2020 | Instance-Aware Self-supervised Learning for Nuclei Segmentation
Xinpeng Xie, Jiawei Chen 0009, Yuexiang Li, LinLin Shen, Kai Ma 0002, Yefeng Zheng 0001 |
MICCAI (5) | 4 |
| 2020 | Multi-resolution convolutional networks for chest X-ray radiograph based lung nodule detection
Xuechen Li 0001, LinLin Shen, Xinpeng Xie, Shiyun Huang, Zhien Xie, Xian Hong |
Artif. Intell. Medicine | 2 |
| 2020 | Robust and high-security fingerprint recognition system using optical coherence tomography
Feng Liu 0013, Guojie Liu, Qijun Zhao, LinLin Shen |
Neurocomputing | 4 |
| 2020 | Fingerprint pore matching using deep features
Feng Liu 0013, Yuanhao Zhao, Guojie Liu, LinLin Shen |
Pattern Recognit. | 4 |
| 2020 | Adaptive Convolution Local and Global Learning for Class-Level Joint Representation of Facial Recognition With a Single Sample Per Data SubjectabstractDue to the absence of training samples and intraclass variation, the extraction of discriminative facial features and construction of powerful classifiers have bottlenecks in improving the performance of facial recognition (FR) with a single sample per data subject (SSPDS). In this paper, we propose to learn regional adaptive convolution features that are locally and globally discriminative to facial identity and robust to facial variation. Then, a novel class-level joint representation framework is presented to exploit the distinctiveness and class-level commonality of different facial features. In the proposed class-level joint representation with regional adaptive convolution features (CJR-RACF), both discriminative facial features that are robust to facial variations and powerful representations for classification with generic facial variations have been fully exploited. Furthermore, the gallery discrimination is extracted by our proposed weight-embedded supervision in the training phase (denoted by CJR-RACFw), which is conducive to more specific features for FR with SSPDS. CJR-RACF and CJR-RACFw have been evaluated on several popular databases, including the large-scale CMU Multi-PIE, LFW, Megaface, and VGGFace datasets. Experimental results demonstrate the much higher robustness and effectiveness of the proposed methods compared to the state-of-the-art methods. Meng Yang 0001, Xing Wang 0012, LinLin Shen, Guangwei Gao |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2020 | A Multi-Organ Nucleus Segmentation ChallengeabstractGeneralized nucleus segmentation techniques can contribute greatly to reducing the time to develop and validate visual biomarkers for new digital pathology datasets. We summarize the results of MoNuSeg 2018 Challenge whose objective was to develop generalizable nuclei segmentation techniques in digital pathology. The challenge was an official satellite event of the MICCAI 2018 conference in which 32 teams with more than 80 participants from geographically diverse institutes participated. Contestants were given a training set with 30 images from seven organs with annotations of 21,623 individual nuclei. A test dataset with 14 images taken from seven organs, including two organs that did not appear in the training set was released without annotations. Entries were evaluated based on average aggregated Jaccard index (AJI) on the test set to prioritize accurate instance segmentation as opposed to mere semantic segmentation. More than half the teams that completed the challenge outperformed a previous baseline. Among the trends observed that contributed to increased accuracy were the use of color normalization as well as heavy data augmentation. Additionally, fully convolutional networks inspired by variants of U-Net, FCN, and Mask-RCNN were popularly used, typically based on ResNet or VGG base architectures. Watershed segmentation on predicted semantic segmentation maps was a popular post-processing strategy. Several of the top techniques compared favorably to an individual human annotator and can be used with confidence for nuclear morphometrics. Neeraj Kumar 0002, Ruchika Verma, Deepak Anand, Yanning Zhou 0001, Omer Fahri Onder, Efstratios Tsougenis, Hao Chen 0011, Pheng-Ann Heng, Jiahui Li 0005, Navid Alemi Koohbanani, Mostafa Jahanifar, Neda Zamani Tajeddin, Ali Gooya, Nasir M. Rajpoot, Xuhua Ren, Sihang Zhou 0001, Qian Wang 0001, Dinggang Shen, Cheng-Kun Yang, Chi-Hung Weng, Wei-Hsiang Yu, Chao-Yuan Yeh, Shuoyu Xu, Pak-Hei Yeung, Amirreza Mahbod, Gerald Schaefer, Isabella Ellinger, Rupert Ecker, Örjan Smedby, Chunliang Wang, Benjamin Chidester, Vinh Ton-That, Minh-Triet Tran, Jian Ma 0004, Minh N. Do, Simon Graham, Quoc Dang Vu, Jin Tae Kwak, Akshaykumar Gunda, Raviteja Chunduri, Corey Hu, Dariush Lotfi, Reza Safdari, Antanas Kascenas, Alison O'Neil, Dennis Eschweiler, Johannes Stegmaier, Yanping Cui, Kailin Chen, Xinmei Tian 0001, Philipp Grüning, Erhardt Barth, Elad Arbel, Itay Remer, Amir Ben-Dor, Ekaterina Sirazitdinova, Matthias Kohl, Stefan Braunewell, Yuexiang Li, Xinpeng Xie, LinLin Shen, Jun Ma 0016, Krishanu Das Baksi, Mohammad Azam Khan, Jaegul Choo, Adrián Colomer, Valery Naranjo, Linmin Pei, Khan M. Iftekharuddin, Kaushiki Roy, Debotosh Bhattacharjee, Aníbal Pedraza, Gloria Bueno García, Sabarinathan Devanathan, Saravanan Radhakrishnan, Praveen Koduganty, Zihan Wu 0001, Guanyu Cai, Amit Sethi |
IEEE Trans. Medical Imaging | 67 |
| 2020 | 3D Neuron Reconstruction in Tangled Neuronal Image With Deep NetworksabstractDigital reconstruction or tracing of 3D neuron is essential for understanding the brain functions. While existing automatic tracing algorithms work well for the clean neuronal image with a single neuron, they are not robust to trace the neuron surrounded by nerve fibers. We propose a 3D U-Net-based network, namely 3D U-Net Plus, to segment the neuron from the surrounding fibers before the application of tracing algorithms. All the images in BigNeuron, the biggest available neuronal image dataset, contain clean neurons with no interference of nerve fibers, which are not practical to train the segmentation network. Based upon the BigNeuron images, we synthesize a SYNethic TAngled NEuronal Image dataset (SYNTANEI) to train the proposed network, by fusing the neurons with extracted nerve fibers. Due to the adoption of dropout, àtrous convolution and Àtrous Spatial Pyramid Pooling (ASPP), experimental results on the synthetic and real tangled neuronal images show that the proposed 3D U-Net Plus network achieved very promising segmentation results. The neurons reconstructed by the tracing algorithm using the segmentation result match significantly better with the ground truth than that using the original images. Qiufu Li, LinLin Shen |
IEEE Trans. Medical Imaging | 2 |
| 2020 | Attention by Selection: A Deep Selective Attention Approach to Breast Cancer ClassificationabstractDeep learning approaches are widely applied to histopathological image analysis due to the impressive levels of performance achieved. However, when dealing with high-resolution histopathological images, utilizing the original image as input to the deep learning model is computationally expensive, while resizing the original image to achieve low resolution incurs information loss. Some hard-attention based approaches have emerged to select possible lesion regions from images to avoid processing the original image. However, these hard-attention based approaches usually take a long time to converge with weak guidance, and valueless patches may be trained by the classifier. To overcome this problem, we propose a deep selective attention approach that aims to select valuable regions in the original images for classification. In our approach, a decision network is developed to decide where to crop and whether the cropped patch is necessary for classification. These selected patches are then trained by the classification network, which then provides feedback to the decision network to update its selection policy. With such a co-evolution training strategy, we show that our approach can achieve a fast convergence rate and high classification accuracy. Our approach is evaluated on a public breast cancer histopathological image database, where it demonstrates superior performance compared to state-of-the-art deep learning approaches, achieving approximately 98% classification accuracy while only taking 50% of the training time of the previous hard-attention approach. Bolei Xu, Jingxin Liu 0005, Xianxu Hou, Jonathan M. Garibaldi, Ian O. Ellis, Andrew R. Green, LinLin Shen, Guoping Qiu |
IEEE Trans. Medical Imaging | 8 |
| 2019 | Local Feature Tensor Based Deep Learning for 3D Face RecognitionabstractA local feature tensor similarity based deep learning approach is proposed in this paper for 3D face recognition. Once a set of salient points on the 3D mesh are detected, three scale and rotation invariant features are extracted to represent local surface around each salient point. The local features of all the salient points are concatenated to produce a 3rdorder feature tensor to represent a 3D face. Similarity of two 3D faces can thus be measured by a similarity tensor calculated using the two feature tensors. To address the unavailability of large 3D face samples, a feature tensor based data augmentation approach is proposed to augment the number of feature tensors. Experimental results show that the ResNet model trained using the augmented feature tensors achieves the best performance among state of the art competitors, i.e. 99.71% and 96.2% accuracy are achieved for Bosphorus and BU3DFE database, respectively. Shisong Lin, Feng Liu 0013, LinLin Shen |
FG | 4 |
| 2019 | Local Normalization Based BN Layer Pruning
Xi Jia, LinLin Shen, Zhong Ming 0001, Jinming Duan 0001 |
ICANN (2) | 3 |
| 2019 | Disentangled Feature Based Adversarial Learning for Facial Expression RecognitionabstractA facial expression image can be considered as an addition of expressive component to a neutral expression face. With this in mind, in this paper, we propose a novel end-to-end adversarial disentangled feature learning (ADFL) framework for facial expression recognition. The ADFL framework is mainly composed of three branches: expression disentangling branch ADFL-d, neutral expression branch ADFL-n and residual expression branch ADFL-r. The ADFL-d and ADFL-n aim to extract the expressive component and neutral component, respectively. The ADFL-r extracts the residual expression by calculating the difference between feature maps of ADFL-d and ADFL-n, and uses the residual expression feature for expression classification. Experimental results on several benchmark databases (CK+, MMI and Oulu-CASIA) show that the proposed method has remarkable performance compared to state-of-the-art methods. Mengchao Bai, Weicheng Xie 0001, LinLin Shen |
ICIP | 3 |
| 2019 | Introspective Gan for Meshface RecognitionabstractMajority of face recognition systems can only retrieve protected ID photos from the government agency in China. These facial images covered with mesh-like curves are termed meshface. Meshface could significantly affect the performance of face recognition systems. Although some GAN based methods are proposed to address this issue by translating meshface images to clean face images, their capacities are limited. In this paper, we introduce the introspective modules to encourage the generator to reconstruct face images and accelerate the learning process of the discriminator. Both a private meshface dataset and the public LFW dataset are used for experiments. The quantitative evaluation on both datasets proves that the introspective GAN can recover face image with better quality. Additionally, face recognition performance is also significantly improved. Wenting Chen, LinLin Shen, Zhihui Lai 0001 |
ICIP | 2 |
| 2019 | Outlier-Suppressed Triplet Loss with Adaptive Class-Aware Margins for Facial Expression RecognitionabstractTriplet loss has been proposed to increase the inter-class distance and decrease the intra-class distance for various tasks of image recognition. However, for facial expression recognition (FER) problem, the fixed margin parameter does not fit the diversity of scales between different expressions. Meanwhile, the strategy of selecting the hardest triplets can introduce noisy guidance information since various persons may present significantly different expressions. In this work, we propose a new triplet loss based on class-aware margins and outlier-suppressed triplet for FER, where each pair of expressions, e.g. 'happy' and 'fear', is assigned with an adaptive margin parameter and the abnormal hard triplets are discarded according to the feature distance distribution. Experimental results of the proposed triplet loss on the FER2013 and CK+ expression databases show that the proposed network achieves much better accuracy than the original triplet loss and the network without using the proposed strategies, and competitive performance compared with the state-of-the-art algorithms. Zhiwei Wen, Weicheng Xie 0001, LinLin Shen, Jinming Duan 0001 |
ICIP | 5 |
| 2019 | DEEPPOREID: An Effective Pore Representation Descriptor in Direct Pore MatchingabstractThis paper proposes an effective pore representation descriptor based on Convolutional Neural Networks (CNNs). We make full use of the diversity and large quantities of sweat pores in fingerprints to learn a deep feature, denoted as DeepPoreID. The DeepPoreID is then used to describe the local feature for each pore and finally integrated into the classical direct pore matching method. Experiments carried on the challenge public high-resolution fingerprint database with small image size of 320 × 240 shows the effectiveness of the proposed DeepPoreID. The results also have shown that the proposed method outperforms other existing state-of-the-art methods in the aspect of recognition accuracy. About ~35% rise in accuracy can be obtained when compared with the best result achieved by existing methods. Yuanhao Zhao, Guojie Liu, Feng Liu 0013, LinLin Shen, Qin Li 0001 |
ICIP | 4 |
| 2019 | SwitchGAN for Multi-domain Facial Image TranslationabstractRecent studies for multi-domain facial image translation have achieved an impressive performance. However, the existing methods still have limitations for some tasks, such as translating a facial image into different age groups, or translating a facial expression into other expressions. To address this problem, we propose a Switch Generative Adversarial Network (SwitchGAN) to perform delicate image translation among multiple domains. A feature switching operation is proposed to achieve features selection and fusion in our conditional modules. Experiments on Morph, RaFD and CelebA databases show that our SwitchGAN can achieve visually better translation effects than StarGAN. The attribute classification results using the trained ResNet-18 classifier also quantitatively suggest that the face images generated by SwitchGAN achieved much higher accuracy than that generated by StarGAN. Yuanlue Zhu, Mengchao Bai, LinLin Shen, Zhiwei Wen |
ICME | 3 |
| 2019 | Adversarial Feature Distillation for Facial Expression Recognition
Mengchao Bai, Xi Jia, Weicheng Xie 0001, LinLin Shen |
PRICAI (3) | 4 |
| 2019 | Texture Deformation Based Generative Adversarial Networks for Multi-domain Face Editing
Wenting Chen, Xinpeng Xie, Xi Jia, LinLin Shen |
PRICAI (1) | 4 |
| 2019 | Reverse active learning based atrous DenseNet for pathological image classificationabstractBACKGROUND: Due to the recent advances in deep learning, this model attracted researchers who have applied it to medical image analysis. However, pathological image analysis based on deep learning networks faces a number of challenges, such as the high resolution (gigapixel) of pathological images and the lack of annotation capabilities. To address these challenges, we propose a training strategy called deep-reverse active learning (DRAL) and atrous DenseNet (ADN) for pathological image classification. The proposed DRAL can improve the classification accuracy of widely used deep learning networks such as VGG-16 and ResNet by removing mislabeled patches in the training set. As the size of a cancer area varies widely in pathological images, the proposed ADN integrates the atrous convolutions with the dense block for multiscale feature extraction. RESULTS: The proposed DRAL and ADN are evaluated using the following three pathological datasets: BACH, CCG, and UCSB. The experiment results demonstrate the excellent performance of the proposed DRAL + ADN framework, achieving patch-level average classification accuracies (ACA) of 94.10%, 92.05% and 97.63% on the BACH, CCG, and UCSB validation sets, respectively. CONCLUSIONS: The DRAL + ADN framework is a potential candidate for boosting the performance of deep learning models for partially mislabeled training datasets. Yuexiang Li, Xinpeng Xie, LinLin Shen, Shaoxiong Liu |
BMC Bioinform. | 3 |
| 2019 | Improving variational autoencoder with deep feature consistent and generative adversarial training
Xianxu Hou, Ke Sun 0006, LinLin Shen, Guoping Qiu |
Neurocomputing | 3 |
| 2019 | Deep learning based early stage diabetic retinopathy detection using optical coherence tomography
Xuechen Li 0001, LinLin Shen, Meixiao Shen, Fan Tan, Connor S. Qiu |
Neurocomputing | 2 |
| 2019 | Hybrid CMOS-Memristive Convolutional computation for on-chip learning
Yang Zhang 0012, Menglin Cui, LinLin Shen |
Neurocomputing | 4 |
| 2019 | Binary sparse signal recovery algorithms based on logic observation
Xiao-Li Hu, Jiajun Wen 0001, Zhihui Lai 0001, Wai Keung Wong, LinLin Shen |
Pattern Recognit. | 5 |
| 2019 | Sparse deep feature learning for facial expression recognition
Weicheng Xie 0001, Xi Jia, LinLin Shen, Meng Yang 0001 |
Pattern Recognit. | 3 |
| 2019 | UP-CNN: Un-pooling augmented convolutional neural network
Chunyan Xu, Jian Yang 0003, Hanjiang Lai, Junbin Gao, LinLin Shen, Shuicheng Yan |
Pattern Recognit. Lett. | 5 |
| 2019 | Generalized Robust Regression for Jointly Sparse Subspace LearningabstractRidge regression is widely used in multiple variable data analysis. However, in very high-dimensional cases such as image feature extraction and recognition, conventional ridge regression or its extensions have the small-class problem, that is, the number of the projections obtained by ridge regression is limited by the number of the classes. In this paper, we proposed a novel method called generalized robust regression (GRR) for jointly sparse subspace learning which can address the problem. GRR not only imposes L2,1-norm penalty on both loss function and regularization term to guarantee the joint sparsity and the robustness to outliers for effective feature selection, but also utilizes L2,1-norm as the measurement to take the intrinsic local geometric structure of the data into consideration to improve the performance. Moreover, by incorporating the elastic factor on the loss function, GRR can enhance the robustness to obtain more projections for feature selection or classification. To obtain the optimal solution of GRR, an iterative algorithm was proposed and the convergence was also proved. Experiments on six wellknown data sets demonstrate the merits of the proposed method. The result indicates that GRR is a robust and efficient regression method for face recognition. Zhihui Lai 0001, Dongmei Mo, Jiajun Wen 0001, LinLin Shen, Wai Keung Wong |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2019 | Structurally Incoherent Low-Rank 2DLPP for Image ClassificationabstractPreserving projection-based methods are good for finding the manifold structure embedded in data. As they use the Euclidean distance as a metric, which is sensitive to noise and outliers in data, nuclear norm-based 2D locality preserving projection (NN-2DLPP) is thus proposed to improve the robustness of 2DLPP. However, NN-2DLPP does not consider the discriminant ability of data. In order to improve the discriminant ability of preserving projection methods, in this paper, we use preserving projection learning with structurally incoherence of data and propose structurally incoherent low-rank 2DLPP (SILR-2DLPP) for image classification. This approach provides a discriminative representation of preserving projection learning by recovering the distinct different classes of the data. SILR-2DLPP searches the optimal subspace and low-rank representation simultaneously. We further extend SILR-2DLPP to a kernel case and propose kernel SILR-2DLPP (KSILR-2DLPP) to obtain a nonlinear representation. The theoretical analysis including the convergence and computational complexity of SILR-2DLPP are presented. To verify the performance of SILR-2DLPP and KSILR-2DLPP, six well-known image databases were used in the experiments. The experimental results show that the proposed methods are superior to the previous preserving projection methods for image classification. Yuwu Lu, Chun Yuan 0003, Xuelong Li 0001, Zhihui Lai 0001, David Zhang 0001, LinLin Shen |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2018 | Human Behaviour-Based Automatic Depression Analysis Using Hand-Crafted Statistics and Deep Learned Spectral FeaturesabstractDepression is a serious mental disorder that affects millions of people all over the world. Traditional clinical diagnosis methods are subjective, complicated and need extensive participation of experts. Audio-visual automatic depression analysis systems predominantly base their predictions on very brief sequential segments, sometimes as little as one frame. Such data contains much redundant information, causes a high computational load, and negatively affects the detection accuracy. Final decision making at the sequence level is then based on the fusion of frame or segment level predictions. However, this approach loses longer term behavioural correlations, as the behaviours themselves are abstracted away by the frame-level predictions. We propose to on the one hand use automatically detected human behaviour primitives such as Gaze directions, Facial action units (AU), etc. as low-dimensional multi-channel time series data, which can then be used to create two sequence descriptors. The first calculates the sequence-level statistics of the behaviour primitives and the second casts the problem as a Convolutional Neural Network problem operating on a spectral representation of the multichannel behaviour signals. The results of depression detection (binary classification) and severity estimation (regression) experiments conducted on the AVEC 2016 DAIC-WOZ database show that both methods achieved significant improvement compared to the previous state of the art in terms of the depression severity estimation. Siyang Song, LinLin Shen, Michel F. Valstar |
FG | 2 |
| 2018 | Hand-Crafted Feature Guided Deep Learning for Facial Expression RecognitionabstractA number of facial expression recognition algorithms based on hand-crafted features and deep neutral networks have been developed. Motivated by the similarity between the hand-crafted features and features learned by deep network, a new feature loss is proposed to embed the information of hand-crafted features into the training process of network, which tries to reduce the difference between the two features. Based on the feature loss, a general framework for embedding the traditional feature information was developed and tested using CK+, JAFFE and FER2013 datasets. Experimental results show that the proposed network achieves much better accuracy than the original hand-crafted feature and the network without using our feature loss. When compared with other algorithms in literature, our network also achieved the best performance on CK+ dataset, i.e. 97.35% accuracy has been achieved. Guohang Zeng, Jiancan Zhou, Xi Jia, Weicheng Xie 0001, LinLin Shen |
FG | 5 |
| 2018 | A Self-Organizing Tensor Architecture for Multi-view ClusteringabstractIn many real-world applications, data are often unlabeled and comprised of different representations/views which often provide information complementary to each other. Although several multi-view clustering methods have been proposed, most of them routinely assume one weight for one view of features, and thus inter-view correlations are only considered at the view-level. These approaches, however, fail to explore the explicit correlations between features across multiple views. In this paper, we introduce a tensor-based approach to incorporate the higher-order interactions among multiple views as a tensor structure. Specifically, we propose a multi-linear multi-view clustering (MMC) method that can efficiently explore the full-order structural information among all views and reveal the underlying subspace structure embedded within the tensor. Extensive experiments on realworld datasets demonstrate that our proposed MMC algorithm clearly outperforms other related state-of-the-art methods. Lifang He 0001, Chun-Ta Lu, Yong Chen 0016, Jiawei Zhang 0001, LinLin Shen, Philip S. Yu, Fei Wang 0001 |
ICDM | 5 |
| 2018 | Adaptive convolution local and global learning for class-level joint representation of face recognition with single sample per personabstractDue to the absence of samples with intra-class variation, extracting discriminative facial features and building powerful classifiers are the bottlenecks of improving the performance of face recognition (FR) with single sample per person (SSPP). In this paper, we propose to learn regional adaptive convolution features which are locally and globally discriminative to face identity and robust to face variation. With collected generic facial variations, a novel class-level joint representation framework is presented to exploit the distinctiveness and class-level commonality of different facial features. In the proposed class-level joint representation with regional adaptive convolution feature (CJR-RACF), both discriminative facial features robust to various facial variations and powerful representation for classification with generic facial variations that can overcome the small-sample-size problem are fully exploited. CJR-RACF has been evaluated on several popular databases, including large-scale CMU Multi-PIE and LFW databases. Experimental results demonstrate the much higher robustness and effectiveness of CJR-RACF to complex facial variations compared to the state-of-the-art methods. Xing Wang 0012, LinLin Shen, Meng Yang 0001 |
ICPR | 3 |
| 2018 | GT-Net: A Deep Learning Network for Gastric Tumor DiagnosisabstractGastric cancer is one of the most common cancers, which causes the second largest number of deaths in the world. Traditional diagnosis approach requires pathologists to manually annotate the gastric tumor in gastric slice for cancer identification, which is laborious and time-consuming. In this paper, we proposed a deep learning based framework, namely GT-Net, for automatic segmentation of gastric tumor. The proposed GT-Net adopts different architectures for shallow and deep layers for better feature extraction. We evaluate the proposed framework on publicly available BOT gastric slice dataset. The experimental results show that our GT-Net performs better than state-of-the-art networks like FCN-8s, U-net, and achieved a new state-of-the-art F1 score of 90.88% for gastric tumor segmentation. Yuexiang Li, Xinpeng Xie, Shaoxiong Liu, Xuechen Li 0001, LinLin Shen |
ICTAI | 5 |
| 2018 | Noise Invariant Frame Selection: A Simple Method to Address the Background Noise Problem for Text-independent Speaker VerificationabstractThe performance of speaker-related systems usually degrades heavily in practical applications largely due to the presence of background noise. To improve the robustness of such systems in unknown noisy environments, this paper proposes a simple pre-processing method called Noise Invariant Frame Selection (NIFS). Based on several noisy constraints, it selects noise invariant frames from utterances to represent speakers. Experiments conducted on the TIMIT database showed that the NIFS can significantly improve the performance of Vector Quantization (VQ), Gaussian Mixture Model-Universal Background Model (GMM-UBM) and i-vector-based speaker verification systems in different unknown noisy environments with different SNRs, in comparison to their baselines. Meanwhile, the proposed NIFS-based speaker verification systems achieves similar performance when we change the constraints (hyper-parameters) or features, which indicates that it is robust and easy to reproduce. Since NIFS is designed as a general algorithm, it could be further applied to other similar tasks. Siyang Song, Shuimei Zhang, Björn W. Schuller, LinLin Shen, Michel F. Valstar |
IJCNN | 4 |
| 2018 | Open snake model based on global guidance field for embryo vessel locationabstractThe development of vessels can provide important information about the growth status of animal embryos. It is, therefore, important to automatically locate the deformed vessel branches from the embryo images. However, very few vessel detectors can accurately locate all vessel branches when the captured images are low quality and the implied vessel shapes are complex. In this study, a new framework consisting of vessel region extraction and snake shape optimisation is proposed. The main contribution in this detector is a novel open snake model based on the global guidance field and deformation template initialisation. Experimental results on a specific application of an embryo vessel database [Database and source codes: https://github.com/wcxie/Egg‐embryro‐vessel‐location/ .] demonstrate that the proposed algorithm not only locates the vessel shape properly but also obtains the orientations of embryo vessel branches accurately. Comparison to traditional guidance fields and the active appearance model illustrates the effectiveness and competitiveness of the proposed model. Weicheng Xie 0001, Jinming Duan 0001, LinLin Shen, Yuexiang Li, Meng Yang 0001, Guojun Lin |
IET Comput. Vis. | 3 |
| 2018 | Joint Bayesian guided metric learning for end-to-end face verification
Chunyan Xu, Jian Yang 0003, Jianjun Qian, Yuhui Zheng, LinLin Shen |
Neurocomputing | 6 |
| 2018 | Facial expression synthesis with direction field preservation based mesh deformation and lighting fitting based wrinkle mapping
Weicheng Xie 0001, LinLin Shen, Meng Yang 0001, Jianmin Jiang |
Multim. Tools Appl. | 2 |
| 2018 | Robust, discriminative and comprehensive dictionary learning for face recognition
Guojun Lin, Meng Yang 0001, Jian Yang 0003, LinLin Shen, Weicheng Xie 0001 |
Pattern Recognit. | 4 |
| 2018 | Deep cross residual network for HEp-2 cell staining pattern classification
LinLin Shen, Xi Jia, Yuexiang Li |
Pattern Recognit. | 1 |
| 2018 | A 3-D Gabor Phase-Based Coding and Matching Framework for Hyperspectral Imagery ClassificationabstractAs manual labeling is very difficult and time-consuming, the labeled samples used to train a supervised classifier are generally limited, which become one of the biggest challenge for hyperspectral imagery classification. In order to tackle this issue, a recent trend is to exploit the structure information of materials, as which reflects the region homogeneity in the spatial domain and offers an invaluable complement to the spectral information. In this respect, 3-D Gabor wavelets have been introduced to extract joint spectral-spatial features for hyperspectral images. One the one hand, the features extracted by 3-D Gabor wavelets lead to very good performance for classification. On the other hand, its drawbacks, i.e., big number of features and high computational cost, limit its applicability. In this paper, a 3-D Gabor-wavelet-based phase coding and Hamming distance-based matching (3DGPC-HDM) framework is developed for hyperspectral imagery classification. The proposed method, instead of taking into account the large volume of Gabor magnitude features, exploits the Gabor phase features with certain orientations (i.e., the direction parallel to the spectral axis), which are then encoded by a simple quadrant bit coding scheme. After that, a normalized Hamming distance matching (HDM) method is adopted to determine the similarity of two samples, and the nearest neighbor classifier is routinely utilized for pixelwise recognition. Finally, experiments on three real hyperspectral data sets show that the proposed 3DGPC-HDM leads to very good performance. Comparisons with the state-of-the-art methods in the literature, in terms of both classifier complexity and generalization ability from very small training sets, are also included. Sen Jia 0001, LinLin Shen, Jiasong Zhu, Qingquan Li 0001 |
IEEE Trans. Cybern. | 2 |
| 2018 | A Solitary Feature-Based Lung Nodule Detection Approach for Chest X-Ray RadiographsabstractLung cancer is one of the most deadly diseases. It has a high death rate and its incidence rate has been increasing all over the world. Lung cancer appears as a solitary nodule in chest x-ray radiograph (CXR). Therefore, lung nodule detection in CXR could have a significant impact on early detection of lung cancer. Radiologists define a lung nodule in CXR as "solitary white nodule-like blob." However, the solitary feature has not been employed for lung nodule detection before. In this paper, a solitary feature-based lung nodule detection method was proposed. We employed stationary wavelet transform and convergence index filter to extract the texture features and used AdaBoost to generate white nodule-likeness map. A solitary feature was defined to evaluate the isolation degree of candidates. Both the isolation degree and the white nodule likeness were used as final evaluation of lung nodule candidates. The proposed method shows better performance and robustness than those reported in previous research. More than 80% and 93% of lung nodules in the lung field in the Japanese Society of Radiological Technology (JSRT) database were detected when the false positives per image were two and five, respectively. The proposed approach has the potential of being used in clinical practice. Xuechen Li 0001, LinLin Shen, Suhuai Luo |
IEEE J. Biomed. Health Informatics | 2 |
| 2018 | Low-Rank Linear Embedding for Image RecognitionabstractLocality preserving projections (LPP) has been widely studied and extended in recent years, because of its promising performance in feature extraction. In this paper, we propose a modified version of the LPP by constructing a novel regression model. To improve the performance of the model, we impose a low-rank constraint on the regression matrix to discover the latent relations between different neighbors. By using the L2,1-norm as a metric for the loss function, we can further minimize the reconstruction error and derive a robust model. Furthermore, the L2,1-norm regularization term is added to obtain a jointly sparse regression matrix for feature selection. An iterative algorithm with guaranteed convergence is designed to solve the optimization problem. To validate the recognition efficiency, we apply the algorithm to a series of benchmark datasets containing face and character images for feature extraction. The experimental results show that the proposed method is better than some existing methods. The code of this paper can be downloaded from http://www.scholat.com/laizhihui. Yudong Chen 0002, Zhihui Lai 0001, Wai Keung Wong, LinLin Shen, Qinghua Hu |
IEEE Trans. Multim. | 4 |
| 2017 | Multi-way Multi-level Kernel Modeling for Neuroimaging ClassificationabstractOwing to prominence as a diagnostic tool for probing the neural correlates of cognition, neuroimaging tensor data has been the focus of intense investigation. Although many supervised tensor learning approaches have been proposed, they either cannot capture the nonlinear relationships of tensor data or cannot preserve the complex multi-way structural information. In this paper, we propose a Multi-way Multi-level Kernel (MMK) model that can extract discriminative, nonlinear and structural preserving representations of tensor data. Specifically, we introduce a kernelized CP tensor factorization technique, which is equivalent to performing the low-rank tensor factorization in a possibly much higher dimensional space that is implicitly defined by the kernel function. We further employ a multi-way nonlinear feature mapping to derive the dual structural preserving kernels, which are used in conjunction with kernel machines (e.g., SVM). Extensive experiments on real-world neuroimages demonstrate that the proposed MMK method can effectively boost the classification performance on diverse brain disorders (i.e., Alzheimers disease, ADHD, and HIV). Lifang He 0001, Chun-Ta Lu, Hao Ding 0003, Shen Wang 0005, LinLin Shen, Philip S. Yu, Ann B. Ragin |
CVPR | 5 |
| 2017 | Gabor phase feature-based hyperspectral imagery classificationabstractIn this paper, a three-dimensional (3D) Gabor phase coding and Hamming distance matching approach, called 3DGPC-HDM, is proposed for hyperspectral imagery classification. Specifically, the Gabor phase features with certain orientations are utilized, which are then encoded by a simple quadrant bit coding scheme. Next, a normalized Hamming distance matching method has been introduced to determine the similarity of two samples, and the nearest neighbor classifier is routinely used for recognition. The extensive experiments on two real hyper-spectral data sets have demonstrated superior performance of the proposed 3DGPC-HDM approach over the state-of-the-art methods in the literature. Sen Jia 0001, Huimin Xie, LinLin Shen |
ICME | 4 |
| 2017 | Kernelized Support Tensor MachinesabstractIn the context of supervised tensor learning, preserving the structural information and exploiting the discriminative nonlinear relationships of tensor data are crucial for improving the performance of learning tasks. Based on tensor factorization theory and kernel methods, we propose a novel Kernelized Support Tensor Machine (KSTM) which integrates kernelized tensor factorization with maximum-margin criterion. Specifically, the kernelized factorization technique is introduced to approximate the tensor data in kernel space such that the complex nonlinear relationships within tensor data can be explored. Further, dual structural preserving kernels are devised to learn the nonlinear boundary between tensor data. As a result of joint optimization, the kernels obtained in KSTM exhibit better generalization power to discriminative analysis. The experimental results on real-world neuroimaging datasets show the superiority of KSTM over the state-of-the-art techniques. Lifang He 0001, Chun-Ta Lu, Guixiang Ma, Shen Wang 0005, LinLin Shen, Philip S. Yu, Ann B. Ragin |
ICML | 5 |
| 2017 | Deep Feature Consistent Variational AutoencoderabstractWe present a novel method for constructing Variational Autoencoder (VAE). Instead of using pixel-by-pixel loss, we enforce deep feature consistency between the input and the output of a VAE, which ensures the VAE's output to preserve the spatial correlation characteristics of the input, thus leading the output to have a more natural visual appearance and better perceptual quality. Based on recent deep learning works such as style transfer, we employ a pre-trained deep convolutional neural network (CNN) and use its hidden features to define a feature perceptual loss for VAE training. Evaluated on the CelebA face dataset, we show that our model produces better results than other methods in the literature. We also show that our method can produce latent vectors that can capture the semantic information of face expressions and can be used to achieve state-of-the-art performance in facial attribute prediction. Xianxu Hou, LinLin Shen, Ke Sun 0006, Guoping Qiu |
WACV | 2 |
| 2017 | Invariant feature extraction for gait recognition using only one uniform model
Shiqi Yu 0001, LinLin Shen, Yongzhen Huang |
Neurocomputing | 4 |
| 2017 | Joint regularized nearest points for image set based face recognition
Meng Yang 0001, Xing Wang 0012, Weiyang Liu, LinLin Shen |
Image Vis. Comput. | 4 |
| 2017 | Joint and collaborative representation with local adaptive convolution feature for face recognition with single sample per person
Meng Yang 0001, Xing Wang 0012, Guohang Zeng, LinLin Shen |
Pattern Recognit. | 4 |
| 2017 | Rotational Invariant Dimensionality Reduction AlgorithmsabstractA common intrinsic limitation of the traditional subspace learning methods is the sensitivity to the outliers and the image variations of the object since they use the norm as the metric. In this paper, a series of methods based on the -norm are proposed for linear dimensionality reduction. Since the -norm based objective function is robust to the image variations, the proposed algorithms can perform robust image feature extraction for classification. We use different ideas to design different algorithms and obtain a unified rotational invariant (RI) dimensionality reduction framework, which extends the well-known graph embedding algorithm framework to a more generalized form. We provide the comprehensive analyses to show the essential properties of the proposed algorithm framework. This paper indicates that the optimization problems have global optimal solutions when all the orthogonal projections of the data space are computed and used. Experimental results on popular image datasets indicate that the proposed RI dimensionality reduction algorithms can obtain competitive performance compared with the previous norm based subspace learning algorithms. Zhihui Lai 0001, Yong Xu 0001, Jian Yang 0003, LinLin Shen, David Zhang 0001 |
IEEE Trans. Cybern. | 4 |
| 2017 | HEp-2 Specimen Image Segmentation and Classification Using Very Deep Fully Convolutional NetworkabstractReliable identification of Human Epithelial-2 (HEp-2) cell patterns can facilitate the diagnosis of systemic autoimmune diseases. However, traditional approach requires experienced experts to manually recognize the cell patterns, which suffers from the inter-observer variability. In this paper, an automatic pattern recognition system using fully convolutional network (FCN) was proposed to simultaneously address the segmentation and classification problem of HEp-2 specimen images. The proposed system transforms the residual network (ResNet) to fully convolutional ResNet (FCRN) enabling the network to perform semantic segmentation task. A sand-clock shape residual module is proposed to effectively and economically improve the performance of FCRN. The publicly available I3A-2014 data set was used to train the FCRN model to classify HEp-2 specimen images into seven catalogs: homogeneous, speckled, nucleolar, centromere, golgi, nuclear membrane, and mitotic spindle. The proposed system achieves a mean class accuracy of 94.94% for leave-one-out tests, which outperforms the winner of ICPR 2014, i.e., 89.93%. At the same time, our model also achieves a segmentation accuracy of 89.03%, which is 19.05% higher than that of the benchmark approach, i.e., 69.98%. Yuexiang Li, LinLin Shen, Shiqi Yu 0001 |
IEEE Trans. Medical Imaging | 2 |
| 2017 | A Novel Transient Wrinkle Detection Algorithm and Its Application for Expression SynthesisabstractBecause facial wrinkle is a representative feature of facial expression, automatic wrinkle detection has been an important and challenging topic for expression simulation, recognition, and animation. Recently, most works about wrinkle detection have focused on permanent wrinkles (e.g., age wrinkles), which are usually linear shapes, whereas the detection of transient wrinkles (e.g., expression wrinkles) has not been sufficiently studied because of their shape diversity and complexity. In this work, a novel algorithm for automatic detection of transient wrinkles with linear, fixed, and chaotic shapes is proposed, which largely consists of edge pair matching, active-appearance-model-based wrinkle structure location, and support-vector-machine-based wrinkle classification. The proposed wrinkle detector is applied for expression synthesis and an improved Poisson wrinkle mapping approach is proposed. Experimental results illustrate the competitiveness of the proposed wrinkle detector in detecting different transient wrinkles. Compared with state-of-the-art algorithms, the proposed approach yields complete and accurate wrinkle centers. The expression synthesized by the improved wrinkle mapping is also much more realistic. Weicheng Xie 0001, LinLin Shen, Jianmin Jiang |
IEEE Trans. Multim. | 2 |
| 2017 | Output Constraint Transfer for Kernelized Correlation Filter in TrackingabstractThe kernelized correlation filter (KCF) is one of the state-of-the-art object trackers. However, it does not reasonably model the distribution of correlation response during tracking process, which might cause the drifting problem, especially when targets undergo significant appearance changes due to occlusion, camera shaking, and/or deformation. In this paper, we propose an output constraint transfer (OCT) method that by modeling the distribution of correlation response in a Bayesian optimization framework is able to mitigate the drifting problem. OCT builds upon the reasonable assumption that the correlation response to the target image follows a Gaussian distribution, which we exploit to select training samples and reduce model uncertainty. OCT is rooted in a new theory which transfers data distribution to a constraint of the optimized variable, leading to an efficient framework to calculate correlation filters. Extensive experiments on a commonly used tracking benchmark show that the proposed method significantly improves KCF, and achieves better performance than other state-of-the-art trackers. To encourage further developments, the source code is made available. Baochang Zhang 0001, Xianbin Cao 0001, Qixiang Ye, Chen Chen 0001, LinLin Shen, Alessandro Perina, Rongrong Ji |
IEEE Trans. Syst. Man Cybern. Syst. | 6 |
| 2016 | Analysis-Synthesis Dictionary Learning for Universality-Particularity Representation Based ClassificationabstractDictionary learning has played an important role in the success of sparse representation. Although synthesis dictionary learning for sparse representation has been well studied for universality representation (i.e., the dictionary is universal to all classes) and particularity representation (i.e., the dictionary is class-particular), jointly learning an analysis dictionary and a synthesis dictionary is still in its infant stage. Universality-particularity representation can well match the intrinsic characteristics of data (i.e., different classes share commonality and distinctness), while analysis-synthesis dictionary can give a more complete view of data representation (i.e., analysis dictionary is a dual-viewpoint of synthesis dictionary). In this paper, we proposed a novel model of analysis-synthesis dictionary learning for universality-particularity (ASDL-UP) representation based classification. The discrimination of universality and particularity representation is jointly exploited by simultaneously learning a pair of analysis dictionary and synthesis dictionary. More specifically, we impose a label preserving term to analysis coding coefficients for universality representation. Fisher-like regularizations for analysis coding coefficients and the subsequent synthesis representation are introduced to particularity representation. Compared with other state-of-the-art dictionary learning methods, ASDL-UP has shown better or competitive performance in various classification tasks. Meng Yang 0001, Weiyang Liu, Weixin Luo, LinLin Shen |
AAAI | 4 |
| 2016 | Deep convolutional neural network based HEp-2 cell classificationabstractAs different staining patterns of HEp-2 cells indicate different diseases, the classification of Indirect Immune Fluorescence (IIF) images on Human Epithelial-2 (HEp-2) cell is important for clinical applications. Different from traditional pattern recognition techniques, we use CNN to extract more high-level features for cell images classification. Compared to the existing CNN based HEp-2 classification methods, we proposed a network with deeper architecture. A class-balanced approach is also proposed to augment the HEp-2 cell dataset for network training. The proposed framework achieves an average class accuracy of 79.29% on ICPR 2012 HEp-2 dataset and a mean class accuracy of 98.26% on ICPR 2016 HEp-2 training set. Xi Jia, LinLin Shen, Xiande Zhou, Shiqi Yu 0001 |
ICPR | 2 |
| 2016 | HEp-2 specimen classification with fully convolutional networkabstractReliable automatic system for Human Epithelial-2 (HEp-2) cell image classification can facilitate the diagnosis of systemic autoimmune diseases. In this paper, an automatic pattern recognition system using fully convolutional network (FCN) was proposed to address the HEp-2 specimen classification problem. The FCN in the proposed framework was adapted from VGG-16, which was trained with ICPR 2016 dataset to classify specimen images into seven catalogs: homogeneous, speckled, nucleolar, centromere, golgi, nuclear membrane, and mitotic spindle. The proposed system achieves a mean class accuracy of 90.89% for 5 fold-cross-validation tests using the I3A Contest Task 2 dataset, which is comparable to the winner of ICPR 2014, i.e. 89.93%. Furthermore, since the FCN was firstly developed for semantic segmentation, the proposed framework can simultaneously solve Task 4, Cell segmentation, newly suggested in I3A Contest 2016. The segmentation accuracy of the system is 87.38% on Task 4 dataset which is 17.4% higher than that of the traditional approach, Otsu, i.e. 69.98%. Yuexiang Li, LinLin Shen, Xiande Zhou, Shiqi Yu 0001 |
ICPR | 2 |
| 2016 | View invariant gait recognition using only one uniform modelabstractGait recognition has been proved useful in human identification at a distance. But view variance of gait feature is always a great challenge because of the difference in appearance. If the view of the probe is different from that of the gallery, one view transformation model can be employed to convert the gait feature from one view to another. But most existing models need to estimate the view angle first, and can work for only one view pair. They can not convert multi-view data to one specific view efficiently. We employ one deep model based on auto-encoder for view invariant gait extraction. The model can synthesize gait feature in a progressive way by stacked multi-layer auto-encoders. The unique advantage is that it can extract view invariant feature from any view using only one model, and view estimation is not needed. The proposed method is evaluated on a large dataset, CASIA Gait Dataset B. The experimental results show that it can achieve state-of-the-art performance, and the improvement is more obvious when the view variance is larger. Shiqi Yu 0001, LinLin Shen, Yongzhen Huang |
ICPR | 3 |
| 2016 | Joint Community and Structural Hole Spanner Detection via Harmonic ModularityabstractDetecting communities (or modular structures) and structural hole spanners, the nodes bridging different communities in a network, are two essential tasks in the realm of network analytics. Due to the topological nature of communities and structural hole spanners, these two tasks are naturally tangled with each other, while there has been little synergy between them. In this paper, we propose a novel harmonic modularity method to tackle both tasks simultaneously. Specifically, we apply a harmonic function to measure the smoothness of community structure and to obtain the community indicator. We then investigate the sparsity level of the interactions between communities, with particular emphasis on the nodes connecting to multiple communities, to discriminate the indicator of SH spanners and assist the community guidance. Extensive experiments on real-world networks demonstrate that our proposed method outperforms several state-of-the-art methods in the community detection task and also in the SH spanner identification task (even the methods that require the supervised community information). Furthermore, by removing the SH spanners spotted by our method, we show that the quality of other community detection methods can be further improved. Lifang He 0001, Chun-Ta Lu, Jiaqi W. Ma, Jianping Cao, LinLin Shen, Philip S. Yu |
KDD | 5 |
| 2016 | Spatio-Temporal Tensor Analysis for Whole-Brain fMRI ClassificationabstractOwing to prominence as a research and diagnostic tool in human brain mapping, whole-brain fMRI image analysis has been the focus of intense investigation. Conventionally, input fMRI brain images are converted into vectors or matrices and adapted in kernel based classifiers. fMRI data, however, are inherently coupled with sophisticated spatio-temporal tensor structure (i.e., 3D space × time). Valuable structural information will be lost if the tensors are converted into vectors. Furthermore, time series fMRI data are noisy, involving time shift and low temporal resolution. To address these analytic challenges, more compact and discriminative representations for kernel modeling are needed. In this paper, we propose a novel spatio-temporal tensor kernel (STTK) approach for whole-brain fMRI image analysis. Specifically, we design a volumetric time series extraction approach to model the temporal data, and propose a spatio-temporal tensor based factorization for feature extraction. We further leverage the tensor structure to encode prior knowledge in the kernel. Extensive experiments using real-world datasets demonstrate that our proposed approach effectively boosts the fMRI classification performance in diverse brain disorders (i.e., Alzheimer's disease, ADHD and HIV). Guixiang Ma, Lifang He 0001, Chun-Ta Lu, Philip S. Yu, LinLin Shen, Ann B. Ragin |
SDM | 5 |
| 2016 | LOAD: Local orientation adaptive descriptor for texture and material classification
Xianbiao Qi, Guoying Zhao 0001, LinLin Shen, Qingquan Li 0001, Matti Pietikäinen |
Neurocomputing | 3 |
| 2016 | Reconstruction of normal and albedo of convex Lambertian objects by solving ambiguity matrices using SVD and optimization method
Yujuan Sun, Muwei Jian, Xiaofeng Zhang 0003, Junyu Dong, LinLin Shen, Beijing Chen |
Neurocomputing | 5 |
| 2016 | Structured regularized robust coding for face recognition
Xing Wang 0012, Meng Yang 0001, LinLin Shen |
Neurocomputing | 3 |
| 2016 | Robust object representation by boosting-like deep learning architecture
Lei Wang 0018, Baochang Zhang 0001, Jungong Han, LinLin Shen, Chengshan Qian |
Signal Process. Image Commun. | 4 |
| 2016 | Gabor Cube Selection Based Multitask Joint Sparse Representation for Hyperspectral Image ClassificationabstractThe large amount of spectral and spatial information contained in hyperspectral imagery has provided great opportunity to effectively characterize and identify the surface materials of interest. As a novel feature extraction technique, a series of Gabor wavelet filters with different scales and frequencies was applied on hyperspectral data to extract spectral-spatial-combined features, which produced impressive performance on pixel-oriented classification. However, the incredibly large number of Gabor features could cause too much burden for onboard computation, limiting the efficiency of the method. To make matters worse, due to the nonhomogeneous spatial distribution of materials as well as the different characteristics of the constructed Gabor filters, some Gabor features could have a smaller or even negative impact on material representation, deteriorating the classification accuracy eventually. In this paper, a Gabor cube selection based multitask joint sparse representation approach, abbreviated as GS-MTJSRC, was proposed for hyperspectral image classification. First, based on the Fisher discrimination criterion, the most representative Gabor cubes for each class were picked out. Next, under multitask joint sparse representation framework, a coefficient vector could be obtained for each test sample with the selected Gabor cube features, which could be directly used for the following residual-based classification. Experimental results on three real hyperspectral data sets with different characteristics and spatial resolutions demonstrated the feasibility and efficiency of the proposed method. Sen Jia 0001, Jie Hu 0004, Yao Xie 0001, LinLin Shen, Xiuping Jia, Qingquan Li 0001 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2015 | Orthogonal self-guided similarity preserving projectionsabstractIn this paper, we propose a novel unsupervised dimensionality reduction (DR) method called orthogonal self-guided similarity preserving projections (OSSPP), which seamlessly integrates the procedures of an adjacency graph learning and DR into a one step. Specifically, OSSPP projects the data into a low-dimensional subspace and simultaneously performs similarity preserving learning by using the similarity preserving regularization term in which the reconstruction coefficients of the projected data are used to encode the similarity structure information. An interesting finding is that the problem to determine the reconstruction coefficients can be converted into a weighted non-negative sparse coding problem without any explicit sparsity constraint. Thus the projections obtained by OSSPP contain natural discriminating information. Experimental results demonstrate that OSSPP outperforms state-of-the-art methods in DR. Xiaozhao Fang, Yong Xu 0001, Zheng Zhang 0006, Zhihui Lai 0001, LinLin Shen |
ICIP | 5 |
| 2015 | Gabor feature based dictionary fusion for hyperspectral imagery classificationabstractMultiple kinds of features extracted from hyperspectral imagery (HSI) have shown great potential for pixel-oriented classification. However, two difficulties can be encountered during the classification process. Firstly, it is time consuming to directly utilize the large amount of features. Secondly, because each kind of feature is usually processed individually, the high-level relationship among different features is not completely configured, decreasing the performance eventually. In this paper, a new strategy to fuse the features and exploit dictionary learning for HSI classification is proposed. Based on the high-level relationship, the extracted Gabor features have been integrated into a more compact and more discriminative representation through a Fisher-based criterion. Experimental results have shown that the fused features can not only produce competitive performance for HSI classification, but also greatly reduce the computational complexity. Sen Jia 0001, Jie Hu 0004, Guihua Tang, LinLin Shen |
IGARSS | 4 |
| 2015 | Study on novel Curvature Features for 3D fingerprint recognition
Feng Liu 0013, David Zhang 0001, LinLin Shen |
Neurocomputing | 3 |
| 2015 | Joint representation and pattern learning for robust face recognition
Meng Yang 0001, Pengfei Zhu 0001, Feng Liu 0013, LinLin Shen |
Neurocomputing | 4 |
| 2015 | Three-dimensional Gabor feature extraction for hyperspectral imagery classification using a memetic framework
Zexuan Zhu 0001, Sen Jia 0001, Shan He 0001, Zhen Ji, LinLin Shen |
Inf. Sci. | 6 |
| 2015 | Globally rotation invariant multi-scale co-occurrence local binary pattern
Xianbiao Qi, LinLin Shen, Guoying Zhao 0001, Qingquan Li 0001, Matti Pietikäinen |
Image Vis. Comput. | 2 |
| 2015 | Visual-Patch-Attention-Aware Saliency DetectionabstractThe human visual system (HVS) can reliably perceive salient objects in an image, but, it remains a challenge to computationally model the process of detecting salient objects without prior knowledge of the image contents. This paper proposes a visual-attention-aware model to mimic the HVS for salient-object detection. The informative and directional patches can be seen as visual stimuli, and used as neuronal cues for humans to interpret and detect salient objects. In order to simulate this process, two typical patches are extracted individually and in parallel from the intensity channel and the discriminant color channel, respectively, as the primitives. In our algorithm, an improved wavelet-based salient-patch detector is used to extract the visually informative patches. In addition, as humans are sensitive to orientation features, and as directional patches are reliable cues, we also propose a method for extracting directional patches. These two different types of patches are then combined to form the most important patches, which are called preferential patches and are considered as the visual stimuli applied to the HVS for salient-object detection. Compared with the state-of-the-art methods for salient-object detection, experimental results using publicly available datasets show that our produced algorithm is reliable and effective. Muwei Jian, Kin-Man Lam 0001, Junyu Dong, LinLin Shen |
IEEE Trans. Cybern. | 4 |
| 2015 | Gabor Feature-Based Collaborative Representation for Hyperspectral Imagery ClassificationabstractSparse-representation-based classification (SRC) assigns a test sample to the class with minimum representation error via a sparse linear combination of all the training samples, which has successfully been applied to several pattern recognition problems. According to compressive sensing theory, the l1-norm minimization could yield the same sparse solution as the l0norm under certain conditions. However, the computational complexity of the l1-norm optimization process is often too high for large-scale high-dimensional data, such as hyperspectral imagery (HSI). To make matter worse, a large number of training data are required to cover the whole sample space, which is difficult to obtain for hyperspectral data in practice. Recent advances have revealed that it is the collaborative representation but not the l1-norm sparsity that makes the SRC scheme powerful. Therefore, in this paper, a 3-D Gabor feature-based collaborative representation (3GCR) approach is proposed for HSI classification. When 3-D Gabor transformation could significantly increase the discrimination power of material features, a nonparametric and effective l2-norm collaborative representation method is developed to calculate the coefficients. Due to the simplicity of the method, the computational cost has been substantially reduced; thus, all the extracted Gabor features can be directly utilized to code the test sample, which conversely makes the l2-norm collaborative representation robust to noise and greatly improves the classification accuracy. The extensive experiments on two real hyperspectral data sets have shown higher performance of the proposed 3GCR over the state-of-the-art methods in the literature, in terms of both the classifier complexity and generalization ability from very small training sets. Sen Jia 0001, LinLin Shen, Qingquan Li 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2014 | An effective collaborative representation algorithm for hyperspectral image classificationabstractIn this paper, an effective l2-norm collaborative representation algorithm based on 3D discrete wavelet transform (3D-DWT) features, called CR_DWT, is proposed for hyperspec-tral image classification. By using the discriminative 3D-DWT features extracted from the original spectral space, a non-parametric and efficient l2-norm CR method is developed to calculate the representation coefficients. Due to the simplicity of the method, the computational cost has been substantially reduced, thus all the extracted 3D-DWT texture features can be directly utilized to code the test sample, which greatly improves the classification accuracy of the l2-norm CR mechanism. The extensive experiments on two real hy-perspectral data sets have shown higher performance of the proposed CR_DWT approach over the state-of-the-art methods in the literature, in terms of both the accuracy and classifier complexity. Sen Jia 0001, LinLin Shen |
ICME | 3 |
| 2014 | The BeiHang Keystroke Dynamics Systems, Databases and baselines
Baochang Zhang 0001, Haoran Zeng, LinLin Shen, Jianzhuang Liu, Jason Zhao |
Neurocomputing | 4 |
| 2014 | HEp-2 image classification using intensity order pooling based features and bag of words
LinLin Shen, Jiaming Lin, Shengyin Wu, Shiqi Yu 0001 |
Pattern Recognit. | 1 |
| 2013 | Discriminative Gabor Feature Selection for Hyperspectral Image ClassificationabstractThree-dimensional Gabor wavelets have recently been successfully applied for hyperspectral image classification due to their ability to extract joint spatial and spectrum information. However, the dimension of the extracted Gabor feature is incredibly huge. In this letter, we propose a symmetrical-uncertainty-based and Markov-blanket-based approach to select informative and nonredundant Gabor features for hyperspectral image classification. The extracted Gabor features with large dimension are first ranked by their information contained for classification and then added one by one after investigating the redundancy with already selected features. The proposed approach was fully tested on the widely used Indian Pine site data. The results show that the selected features are much more efficient and can achieve similar performance with previous approach using only hundreds of features. LinLin Shen, Zexuan Zhu 0001, Sen Jia 0001, Jiasong Zhu |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2012 | Hyperspectral face recognition using 3D Gabor wavelets
LinLin Shen, Songhao Zheng |
ICPR | 1 |
| 2011 | Face Recognition from Visible and Near-Infrared Images Using Boosted Directional Binary Code
LinLin Shen, Jinwen He, Shipei Wu, Songhao Zheng |
ICIC (2) | 1 |
| 2011 | Differentiated security levels for personal identifiable information in identity management system
Jianyong Chen, Guihua Wu, LinLin Shen, Zhen Ji |
Expert Syst. Appl. | 3 |
| 2011 | Fpcode: an Efficient Approach for Multi-Modal BiometricsabstractAlthough face recognition technology has progressed substantially, its performance is still not satisfactory due to the challenges of great variations in illumination, expression and occlusion. This paper aims to improve the accuracy of personal identification, when only few samples are registered as templates, by integrating multiple modal biometrics, i.e. face and palmprint. We developed in this paper a feature code, namely FPCode, to represent the features of both face and palmprint. Though feature code has been used for palmprint recognition in literature, it is first applied in this paper for face recognition and multi-modal biometrics. As the same feature is used, fusion is much easier. Experimental results show that both feature level and decision level fusion strategies achieve much better performance than single modal biometrics. The proposed approach uses fixed length 1/0 bits coding scheme that is very efficient in matching, and at the same time achieves higher accuracy than other fusion methods available in literature. LinLin Shen, Li Bai 0001, Zhen Ji |
Int. J. Pattern Recognit. Artif. Intell. | 1 |
| 2011 | Three-Dimensional Gabor Wavelets for Pixel-Based Hyperspectral Imagery ClassificationabstractThe rich information available in hyperspectral imagery not only poses significant opportunities but also makes big challenges for material classification. Discriminative features seem to be crucial for the system to achieve accurate and robust performance. In this paper, we propose a 3-D Gabor-wavelet-based approach for pixel-based hyperspectral imagery classification. A set of complex Gabor wavelets with different frequencies and orientations is first designed to extract signal variances in space, spectrum, and joint spatial/spectral domains. The magnitude of the response at each sampled location (x, y) for spectral band b contains rich information about the signal variances in the local region. Each pixel can be well represented by the rich information extracted by Gabor wavelets. A feature selection and fusion process has also been developed to reduce the redundancy among Gabor features and make the fused feature more discriminative. The proposed approach was fully tested on two real-world hyperspectral data sets, i.e., the widely used Indian Pine site and Kennedy Space Center. The results show that our method achieves as high as 96.04% and 95.36% accuracies, respectively, even when only few samples, i.e., 5% of the total samples per class, are labeled. LinLin Shen, Sen Jia 0001 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2010 | Sparse nonnegative matrix factorization with the elastic netabstractNonnegative matrix factorization is used extensively for feature extraction and clustering analysis. Recently many sparsity/sparseness constraints, such as L1penalty, are introduced for sparse nonnegative matrix factorization. Inspired by sparsity measures from linear regression model, this paper proposes to integrate nonnegative matrix factorization with another sparsity constraint, the elastic net. The experimental results of clustering analysis on three gene expression datasets demonstrate the effectiveness of the proposed method. Weixiang Liu, Songfeng Zheng, Sen Jia 0001, LinLin Shen, Xianghua Fu |
BIBM | 4 |
| 2010 | Directional binary code with application to PolyU near-infrared face database
Baochang Zhang 0001, Lei Zhang 0006, David Zhang 0001, LinLin Shen |
Pattern Recognit. Lett. | 4 |
| 2009 | Erratum to "3D Gabor wavelets for evaluating SPM normalization algorithm"
LinLin Shen, Li Bai 0001, Dorothee Auer |
Medical Image Anal. | 1 |
| 2008 | 3D Gabor wavelets for evaluating SPM normalization algorithm
LinLin Shen, Li Bai 0001 |
Medical Image Anal. | 1 |
| 2007 | Tuning Kernel Parameters with Different Gabor Features for Face Recognition
LinLin Shen, Zhen Ji, Li Bai 0001 |
ICIC (2) | 1 |
| 2007 | Influence of Wavelet Frequency and Orientation in an SVM-Based Parallel Gabor PCA Face Verification System
Ángel Serrano Sánchez de León, Isaac Martín de Diego, Cristina Conde, Enrique Cabello, LinLin Shen, Li Bai 0001 |
IDEAL | 5 |
| 2007 | Gabor wavelets and General Discriminant Analysis for face identification and verification
LinLin Shen, Li Bai 0001, Michael C. Fairhurst |
Image Vis. Comput. | 1 |
| 2006 | A review on Gabor wavelets for face recognition
LinLin Shen, Li Bai 0001 |
Pattern Anal. Appl. | 1 |
| 2006 | MutualBoost learning for selecting Gabor features for face recognition
LinLin Shen, Li Bai 0001 |
Pattern Recognit. Lett. | 1 |
| 2005 | Kernel Enhanced Informative Gabor Features for Face Recognition
LinLin Shen, Li Bai 0001 |
BMVC | 1 |
| 2005 | InfoBoost for Selecting Discriminative Gabor Features
Li Bai 0001, LinLin Shen |
CAIP | 2 |
| 2004 | Facial recognition/verification using gabor wavelets and kernel methods
LinLin Shen, Li Bai 0001, Phil D. Picton |
ICIP | 1 |