Rushi Lan

dblp:44/8845 · DBLP profile ↗
← Back
114ranked-venue papers
19as first author
90since 2021 · last 2026
0000-0002-9488-8236ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 60 · 9 first-author · 47 since 2021Artificial intelligence and machine learning · 30 · 6 first-author · 24 since 2021Applied, interdisciplinary, general and emerging computing · 18 · 3 first-author · 15 since 2021Computer networks · 14 · 1 first-author · 11 since 2021Databases, data management, data science and information retrieval · 4 · 4 since 2021Human-computer interaction and ubiquitous computing · 3 · 1 first-author · 1 since 2021Security and privacy · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Integral-based Knockoffs Inference for Partially Linear Models
abstract
Partial linear models (PLM) have attracted much attention for regression estimation and variable selection due to their feasibility on utilizing linear and nonlinear approximations jointly. However, theoretical understanding of how they control the false discovery rate (FDR) during variable selection remains limited. To address this issue, we formulate a new integral-based knockoffs (IKO) inference scheme for controlled variable selection in PLM, where integral-based knockoff statistics are used to measure the variable importance and B-splines (or random Fourier features) are employed for approximating nonlinear components. In theory, FDR control is guaranteed for both linear and nonlinear parts, and the statistical analysis for its power is established. Empirical evaluations validate the effectiveness of our proposed approach.
Biqin Song, Rushi Lan, Hong Chen 0004
AAAI3
2026 Illuminating the Shadows: Enhanced Low-Light Image via a Retinex-based Model with Color Equalization
Zhenbing Liu, Weidong Zhang 0007, Rushi Lan, Haoxiang Lu
Expert Syst. Appl.6
2026 A generative 3D asset encryption scheme with tiered visual effect after decryption
Zhongshuai Wang, Yushu Zhang 0001, Rushi Lan
Frontiers Comput. Sci.6
2026 Efficient and Adjustable 3-D Ancient Architecture Encryption Scheme With Multilevel Decryption for IoT
Zhongshuai Wang, Yilun Li, Rushi Lan, Yannan Xiao, Yushu Zhang 0001
IEEE Internet Things J.5
2026 Maximum likelihood neural additive models
Peipei Yuan, Rushi Lan, Hong Chen 0004
Inf. Sci.4
2026 Retinal vessel segmentation via bifurcation intensity driven and heterogeneous graph optimization
Huadeng Wang, Ning Kang 0019, Zhenwei Shi 0002, Xipeng Pan, Rushi Lan
Knowl. Based Syst.5
2026 Federated cross-source learning for lung nodule segmentation with data characteristic-aware weight optimization
Xinjun Bian, Lingqiao Li, Zhenbing Liu, Huadeng Wang, Zhenwei Shi 0002, Zaiyi Liu, Rushi Lan, Xipeng Pan
Pattern Recognit.10
2026 Federated semi-supervised medical image segmentation with temporal fluctuation aggregation and pseudo-label relation mining
Junchang Kuang, Xinjun Bian, Siyang Feng, Shufang Pei, Zhenbing Liu, Rushi Lan, Xipeng Pan
Pattern Recognit.6
2026 GMM-based discriminative feature embedding framework for unsupervised fine-grained image retrieval
Kaiyuan Song, Rushi Lan, Kaichen Li
Pattern Recognit.3
2026 A Topology-Aware Authenticated Encryption Transformation for 3D Meshes
abstract
Secure transmission and integrity protection of 3D mesh signals are essential in multimedia, virtual reality, and cloud-based rendering applications. Existing 3D mesh encryption methods mainly focus on geometric confidentiality, while largely neglecting the critical role of mesh topology in authentication and tamper detection. In this letter, a topology-aware authenticated encryption transformation for 3D meshes is proposed. Vertex coordinates are encrypted using AES-CFB to ensure geometric confidentiality, while 3D mesh connectivity information is incorporated into a GMAC-based authentication process to enable topology-aware integrity verification. Any modification to geometry or topology can be reliably detected during the decryption process. Experiments demonstrate near-lossless reconstruction with errors on the order of$10^{-8}$and a 100% detection rate against ciphertext and topology tampering.
Zhongshuai Wang, Rushi Lan, Yushu Zhang 0001
IEEE Signal Process. Lett.3
2026 Chain Reaction: A Triple-Chain Architecture for Sensitive Thumbnail-Preserving Image Encryption
abstract
Recently, thumbnail-preserving encryption (TPE) has attracted significant attention because it maintains certain visual information in encrypted images, thereby achieving a trade-off between visual usability and privacy protection. However, existing TPE schemes utilize independent encryption between pixels, blocks, or channels, which results in excessive robustness and consequently introduces security vulnerabilities. To address this issue, we propose a triple-chain architecture to achieve sensitive TPE. This architecture integrates three interdependent encryption chains at the pixel, block, and channel levels, establishing cryptographic dependencies such that any minor modification to either the original or encrypted image leads to completely divergent encryption/decryption outcomes. Specifically, the architecture introduces three chained encryption mechanisms: (1) pixel-level chained encryption propagates encrypted pixels outputs as the part of subsequent inputs; (2) block-level hashing derives new keys from plaintext-key concatenations; and (3) channel-level dependencies are established via iterative key derivation. This cascading structure ensures cryptographic interdependency across all levels. The experimental results verify that the proposed architecture achieves sensitive TPE for the first time while preserving the TPE characteristics of visual usability and privacy protection.
Wenying Wen, William Puech, Rushi Lan, Yushu Zhang 0001
IEEE Trans. Inf. Forensics Secur.4
2026 Long-Tailed and Inter-Class Homogeneity Matters in Multi-Class Weakly Supervised Tissue Segmentation of Histopathology Images
abstract
Using image-level weakly supervised semantic segmentation (WSSS) techniques to segment tissue regions in giga-pixel histopathological whole slide images (WSI) has garnered widespread attention, as it can reduce many annotation workloads for pathologists. Most recent studies are based on class activation mapping (CAM) to generate pseudo masks, which are then used to train segmentation model in a fully supervised manner. However, it is still a challenge to accurately segment non-predominant tissue categories due to the existence of long-tailed and inter-class homogeneity matters. For these matters, we propose three designs to solve them: 1) Diffusion-based Data Generation to synthesis new images of tail class to expand data distribution; 2) Feature Recalibration to reassign the logits in CAM to narrow the feature-level prediction gap between predominant and non-predominant classes; 3) Grade-skip Learning to correct the under-fitting tendency of hard samples during the segmentation phase. Moreover, we also design a powerful pipeline LoHo for histopathology tissue segmentation. Extensive experiments demonstrate that our method not only achieves new state-of-the-art performances but also significantly improves segmentation of tail classes. In addition, our methods are plug-and-play, making it easily integrable into many mainstream WSSS frameworks.
Siyang Feng, Xipeng Pan, Huadeng Wang, Zhenbing Liu, Weidong Zhang 0007, Rushi Lan
IEEE Trans. Image Process.6
2026 Enabling Reliable and Anonymous Data Collection for Fog-Assisted Mobile Crowdsensing With Malicious User Detection
Zhongyun Hua, Yifeng Zheng 0001, Rushi Lan, Qing Liao 0001, Guoai Xu
IEEE Trans. Mob. Comput.4
2026 QuPaS: SAM-Based Semi-Supervised Histopathological Image Segmentation With Quantum Force Field Finetuning and Adversarial Estimation
abstract
Semi-supervised segmentation (S3) is one of the preferred choices for histopathological image segmentation tasks, while how to improve model’s learning capability for unlabeled data remains a key challenge in S3. The remarkable feature extraction abilities of Segment Anything Model (SAM) offers a potential opportunity. However, SAM’s performance on contextual complex histopathological images is not so desirable due to its limitations in finely capture structural relationships. To address this issue, we propose a novel SAM-based S3framework QuPaS, which consists of Quantum Force Field (QFF) Finetuning and Adversarial Estimation (AE). QFF covers the shortage of SAM’s limited understanding of spatial structure by simulating intermolecular forces to explore the structural topological relationships between pixel-level features. AE introduces an adversarial estimation network to align the consistency of confidence distributions between different outputs, thereby reducing the interference of incompatible semantic features on the model. Extensive experiments across three challenging histopathological segmentation scenarios have demonstrate that our QuPaS completely outperforms the state-of-the-art S3methods. Furthermore, QuPaS is able to maintain stable generalization performance on previously unseen domains. The code will be released at: https://github.com/director87/QuPaS.
Siyang Feng, Xipeng Pan, Weidong Zhang 0007, Minghua Pan, Chu Han, Rushi Lan
IEEE Trans. Medical Imaging6
2026 Wave-Aware Weakly Supervised Histopathological Tissue Segmentation With Cross-Scale Logits Distillation
abstract
Weakly supervised learning based on image-level labels can effectively reduce annotation costs, making it a popular choice for histopathological tissue segmentation. However, this pattern still face some challenges: 1) inaccurate class activation maps (CAM) make pseudo masks quality insufficient; 2) noisy pixels in pseudo masks will mislead the segmentation model's decision-making. To deal with these problems, we propose a novel weakly supervised semantic segmentation (WSSS) framework. First, we introduce Local Spatial Affine Perturbation to strengthen the model's utilization of weak supervision signals and improve its robustness to noisy regions within CAM. Second, we propose Wave-aware Dynamic Feature Aggregation to adaptively enhance the information-aware representation of target regions to obtain fine-grained pseudo masks enriched with positive semantic information. Third, we train a segmentation model with a noise-suppression scheme called Cross-scale Logits Distillation to reduce the inevitable false positive pixels in pseudo masks. We conduct extensive experiments to validate our method and set new state-of-the-art segmentation performances on five histopathological tissue segmentation datasets. Moreover, we will introduce a new dataset, GCSS-WSSS for gastric cancer, to promote the diversification for the research community of computational pathology. Code and data will be released at: https://github.com/director87/WaWeHis.
Siyang Feng, Hualong Zhang, Xianjing Zhao, Liting Shi, Zhenbing Liu, Rushi Lan, Xipeng Pan
IEEE Trans. Medical Imaging6
2025 CA-MLIF: Cross-Attention and Multimodal Low-Rank Interaction Fusion Framework for Tumor Prognostic Prediction
abstract
Cancer is a leading cause of death worldwide due to its aggressive nature and complex variability. Accurate prognosis is therefore challenging but essential for guiding personalized treatment and follow-up. Previous research often relied on single data sources, missing the opportunity to combine various types of patient information for more comprehensive survival predictions. To address these challenges, we propose a two-stage fusion method named Cross-Attention and Multimodal Low-Rank Interaction Fusion Framework (CA-MLIF). In the first stage, we propose a CA mechanism for real-time feature updates and cross-modal mutual learning to capture rich semantic information. In the second stage, we design a novel multimodal low-rank interaction fusion method for survival prediction. Specifically, we present modal attention mechanism (MAM) for feature filtration, low-rank multimodal fusion (LMF) for model complexity reduction, and optimal weight concatenation (OWC) for maximizing feature integration. Extensive experiments on two public datasets TCGA-GBMLGG and TCGA-KIRC, as well as a multi-center in-house lung adenocarcinoma (LUAD) dataset validate the effectiveness of CA-MLIF, which demonstrate that our method outperforms existing approaches in survival prediction under both pathology-gene fusion and CT-pathology fusion scenarios.
Yajun An, Zhenbing Liu, Siyang Feng, Hualong Zhang, Rushi Lan, Zaiyi Liu, Xipeng Pan
AAAI7
2025 Weakly Supervised Gland Segmentation with Class Semantic Consistency and Purified Labels Filtration
abstract
Image-level weakly supervised semantic segmentation (WSSS) reduces the dependence on high-quality data annotation, which plays a crucial role in computational pathology. Benefit from the ability to localize the objects with only binary labels, Class Activation Map (CAM) is a widely used method to initial pseudo masks. However, due to the low contrast among different tissues in histopathological images, most existing CAM-based methods perform poorly in gland segmentation. We retrospect this process and find that class consistency and semantic consistency can guide the network to effectively distinguish confusing pixels and generate fine-grained pseudo masks. Specifically, for class consistency, we propose Consistency Correlation Attention (CCA) to encourage the network to focus on the contribution of class features to semantic dependencies. For semantic consistency, we propose Multi-scale Pyramid Fusion Pooling (MPFP) to aggregate coarse-to-fine global semantic information from CAMs at multiple spatial resolutions, thus identifying class localization. Additionally, we introduce a Purified Labels Filtration (PLF) strategy during the segmentation phase to mitigate the noisy supervision signal and improve the segmentation quality of the model. Extensive experiments show that the our method achieves new state-of-the-art results on three publicly available gland datasets. Furthermore, our method demonstrates impressive domain adaptation capability, achieving satisfactory results with only a small portion of samples when faced with unseen domain data.
Siyang Feng, Huadeng Wang, Chu Han, Zhenbing Liu, Hualong Zhang, Rushi Lan, Xipeng Pan
AAAI6
2025 Phoneme-Level Feature Discrepancies: A Key to Detecting Sophisticated Speech Deepfakes
abstract
Recent advancements in text-to-speech and speech conversion technologies have enabled the creation of highly convincing synthetic speech. While these innovations offer numerous practical benefits, they also cause significant security challenges when maliciously misused. Therefore, there is an urgent need to detect these synthetic speech signals. Phoneme features provide a powerful speech representation for deepfake detection. However, previous phoneme-based detection approaches typically focused on specific phonemes, overlooking temporal inconsistencies across the entire phoneme sequence. In this paper, we develop a new mechanism for detecting speech deepfakes by identifying the inconsistencies of phoneme-level speech features. We design an adaptive phoneme pooling technique that extracts sample-specific phoneme-level features from frame-level speech data. By applying this technique to features extracted by pre-trained audio models on previously unseen deepfake datasets, we demonstrate that deepfake samples often exhibit phoneme-level inconsistencies when compared to genuine speech. To further enhance detection accuracy, we propose a deepfake detector that uses a graph attention network to model the temporal dependencies of phoneme-level features. Additionally, we introduce a random phoneme substitution augmentation technique to increase feature diversity during training. Extensive experiments on four benchmark datasets demonstrate the superior performance of our method over existing state-of-the-art detection methods.
Zhongyun Hua, Rushi Lan, Yushu Zhang 0001, Yifang Guo
AAAI3
2025 Multi-View Collaborative Learning Network for Speech Deepfake Detection
abstract
As deep learning techniques advance rapidly, deepfake speech synthesized through text-to-speech or voice conversion networks is becoming increasingly realistic, posing significant challenges for detection and raising potential threats to social security. This growing realism has prompted extensive research in speech deepfake detection. However, current detection methods primarily focus on extracting features from either the raw waveform or the spectrogram, often overlooking the valuable correspondences between these two modalities that could enhance the detection of previously unseen types of deepfakes. In this work, we propose a multi-view collaborative learning network for speech deepfake detection, which jointly learns robust speech representations from both raw waveforms and spectrograms. Specifically, we first design a Dual-Branch Contrastive Learning (DBCL) framework for learning different view features. DBCL consists of two branches that learn representations from the raw waveform or the spectrogram and utilizes contrastive learning to enhance inter- and inner-view correlations. Additionally, we introduce a Waveform-Spectrogram Fusion Module (WSFM) to exchange multi-view information for collaborative learning. In the feature learning process, WSFM converts features between views and merges them adaptively using waveform-spectrogram cross-attention. The final detection is conducted based on the concatenation of the waveform and spectrogram features. We conduct extensive experiments on four benchmark deepfake speech detection datasets, and the experimental results demonstrate that our method can achieve better detection performance than current state-of-the-art detection methods.
Zhongyun Hua, Rushi Lan, Yifang Guo, Yushu Zhang 0001, Guoai Xu
AAAI3
2025 DeepShield: Fortifying Deepfake Video Detection with Local and Global Forgery Analysis
Yinqi Cai, Jichang Li, Zhaolun Li, Weikai Chen 0001, Rushi Lan, Guanbin Li
ICCV5
2025 FakeRadar: Probing Forgery Outliers to Detect Unknown Deepfake Videos
abstract
In this paper, we propose FakeRadar, a novel deepfake video detection framework designed to address the challenges of cross-domain generalization in real-world scenarios. Existing detection methods typically rely on manipulation-specific cues, performing well on known forgery types but exhibiting severe limitations against emerging manipulation techniques. This poor generalization stems from their inability to adapt effectively to unseen forgery patterns. To overcome this, we leverage large-scale pretrained models (e.g. CLIP) to proactively probe the feature space, explicitly highlighting distributional gaps between real videos, known forgeries, and unseen manipulations. Specifically, FakeRadar introduces Forgery Outlier Probing, which employs dynamic subcluster modeling and cluster-conditional outlier generation to synthesize outlier samples near boundaries of estimated subclusters, simulating novel forgery artifacts beyond known manipulation types. Additionally, we design Outlier-Guided Tri-Training, which optimizes the detector to distinguish real, fake, and outlier samples using proposed outlier-driven contrastive learning and outlier-conditioned cross-entropy losses. Experiments show that FakeRadar outperforms existing methods across various benchmark datasets for deepfake video detection, particularly in cross-domain evaluations, by handling the variety of emerging manipulation techniques.
Zhaolun Li, Jichang Li, Yinqi Cai, Junye Chen, Guanbin Li, Rushi Lan
ICCV7
2025 Selective Kernel and Offset Prediction Network for Video Super-Resolution
Tengjie Hu, Jiheng Hong, Chaoyi Huang, Rushi Lan
ICIG (3)5
2025 Spatially-Aware Framework for Sequential Deepfake Detection
Chaoyi Huang, Rui Yang 0018, Rushi Lan, Zhanghui Wu, Tengjie Hu
ICIG (3)3
2025 A Non-contiguous 3D Object Encryption Method with Multi-level Visual Access Control
Zhongshuai Wang, Yushu Zhang 0001, Rushi Lan
ICIG (3)5
2025 Edge-Semantic Synergy Fusion and Adaptive Noise-Aware for Weakly Supervised Pathological Tissue Segmentation
Hualong Zhang, Siyang Feng, Zihan Huan, Huadeng Wang, Zhenbing Liu, Rushi Lan, Xipeng Pan
MICCAI (8)6
2025 Weakly supervised nuclei segmentation based on pseudo label correction and uncertainty denoising
Xipeng Pan, Shilong Song, Zhenbing Liu, Huadeng Wang, Lingqiao Li, Haoxiang Lu, Rushi Lan
Artif. Intell. Medicine7
2025 Weakly supervised histopathology tissue semantic segmentation with multi-scale voting and online noise suppression
Xipeng Pan, Hualong Zhang, Huahu Deng, Huadeng Wang, Lingqiao Li, Zhenbing Liu, Yajun An, Cheng Lu 0001, Zaiyi Liu, Chu Han, Rushi Lan
Eng. Appl. Artif. Intell.12
2025 Privacy-preserving face attribute classification via differential privacy
Tao Wang 0084, Junhao Ji, Yushu Zhang 0001, Rushi Lan
Neurocomputing5
2025 Differentially Private and Communication-Efficient Federated Learning for AIoT: The Perspective of Denoising and Sparsification
abstract
As public awareness of privacy protection increases and data become more valuable, the applications of federated learning (FL) in the emerging field of Artificial Internet of Things (AIoT) has received widespread attention. Meanwhile, differential privacy (DP), providing strict privacy guarantees, has been introduced to meet users’ stringent privacy protection needs and increasingly sound laws and regulations. However, the implementations of DP in multiple iterations and rounds of FL training, as well as adding noise to all parameters without differentiation, will cause noise accumulation, resulting in the FL system to decline in performance or even fail to converge. To address the issue, FL with denoising DP and sparsification (DDPS-FL) is proposed in this article. First, a local denoising mechanism (LDM) suitable for DP with arbitrary noise adding mechanism is proposed. By removing the previously added noise from global models, LDM achieves direct noise reduction for clients. Second, sparsification based on parameter variation (SPV) is proposed to reduce noise indirectly by deleting nonsignificant parameters without compromising the level of privacy protection. Besides, SPV is able to multiple beneficial effects, such as saving privacy budget, stimulating the dynamism of FL training, amplifying privacy protection effect, and improving communication efficiency. Third, theoretical analysis is performed to prove that DDPS-FL can guarantee user privacy and has ideal convergence, and to analyze the impact of parameters, such as the number of training rounds and iterations on the system performance. Evaluation experiments based on four real-world datasets are elaborated to show that DDPS-FL outperforms state-of-the-art schemes in terms of training stability, model accuracy, and communication efficiency, and its performance becomes relatively better when more noise is added.
Long Li 0005, Zhenshen Liu, Xiyan Sun, Liang Chang 0003, Rushi Lan, Jingjing Li 0003, Jun Wang 0002
IEEE Internet Things J.5
2025 Multi-layer Feature Fusion and Coarse-to-fine Label Learning for Semi-supervised Lesion Segmentation of Lung Cancer
Siyang Feng, Yanfen Cui, Chuansong Fan, Xinjun Bian, Lingqiao Li, Zhenbing Liu, Zaiyi Liu, Rushi Lan, Xipeng Pan
Knowl. Based Syst.10
2025 Perceptual stretch and multi-feature fusion for enhancing nighttime images
Haoxiang Lu, Tianle Fang, Zhenbing Liu, Weidong Zhang 0007, Rushi Lan
Knowl. Based Syst.5
2025 Feature reinforcement meets feature suppression: a hierarchical bilateral method for fine-grained visual classification
Dingzhou Xie, Rushi Lan
Multim. Tools Appl.4
2025 TransBranch: A transformer branch architecture for fine-grained recognition
Dingzhou Xie, Rushi Lan
Pattern Recognit. Lett.4
2025 Multimodal Fusion Framework Based on Low-Rank Interaction for Tumor Prognostic Prediction
abstract
To improve the overall survival rate of cancer patients, we propose an innovative approach named Multimodal Fusion Framework based on Low-rank Interaction (MF2LI), which aims to overcome the current limitations of relying solely on single-modal data prediction and the excessive complexity of fusion. By harnessing low-rank multimodal fusion (LMF) and optimal weight integration (OWI), MF2LI maximizes the integration of pathological images and genomic data. The model incorporates a parallel decomposition strategy, reducing complexity and facilitating fusion based on the contributions of each component. We validate our method using the GBMLGG and KIRC datasets from The Cancer Genome Atlas (TCGA). The C-index of the proposed model stands at $0.895 \pm 0.007$ and $0.728 \pm 0.030$ for the two datasets, respectively, outperforming existing methods. Furthermore, we generate visualizations of the risk ratios, which demonstrate a strong alignment with the actual grade classifications. Extensive experiments have shown that our model improves the prognosis prediction of tumor patients and has considerable clinical value.
Yajun An, Rushi Lan, Huahu Deng, Zhenbing Liu, Zaiyi Liu, Cheng Lu 0001, Xipeng Pan
IEEE Trans. Comput. Biol. Bioinform.2
2025 MASFNet: Multiscale Adaptive Sampling Fusion Network for Object Detection in Adverse Weather
abstract
Object detection methods using deep convolutional neural networks (CNNs) have derived major advances in normal images. However, such success is hardly achieved with adverse weather due to a lack of visibility . To tackle this problem, we propose a Multi-scale Adaptive Sampling Fusion Network, named MASFNet. In this paper, we design a Feature Adaptive Enhancement Network (FAENet) consisting of three modules to adaptively perform feature enhancement on feature maps in adverse scenarios. These modules in FAENet are integrated by the Laplace pyramid, which can perform receptive field fusion, attention perception, and affine transformation for image feature enhancement. To improve the detection performance, we propose a Multi-scale Sampling Fusion Pyramid Network (MSFNet), which is capable of fusing different scale features to improve the semantic information. Experimental results demonstrate that MASFNet achieves 73.68% and 30.95% mAP on the real scene fog dataset (RTTS) and foggy driving dataset (FDD) respectively. Additionally, on the real-world scenario low illumination dataset (ExDark), MASFNet attains a substantial mAP of 63.80%, surpassing current state-of-the-art object detectors while retaining lightweight and high-speed. The source code will be released at https://github.com/PolarisFTL/MASFNet.
Zhenbing Liu, Tianle Fang, Haoxiang Lu, Weidong Zhang 0007, Rushi Lan
IEEE Trans. Geosci. Remote. Sens.5
2025 All Roads Lead to Rome: Achieving 3D Object Encryption Through 2D Image Encryption Methods
abstract
In this paper, we explore a new road for format-compatible 3D object encryption by proposing a novel mechanism of leveraging 2D image encryption methods. It alleviates the difficulty of designing 3D object encryption schemes coming from the intrinsic intricacy of the data structure, and implements the flexible and diverse 3D object encryption designs. First, turning complexity into simplicity, the vertex values, real numbers with continuous values, are converted into integers ranging from 0 to 255. The simplification result for a 3D object is a 2D numerical matrix. Second, six prototypes for three encryption patterns (permutation, diffusion, and permutation-diffusion) are designed as exemplifications to encrypt the 2D matrix. Third, the integer-valued elements in the encrypted numeric matrix are converted into real numbers complying with the syntax of the 3D object. In addition, some experiments are conducted to verify the effectiveness of the proposed mechanism.
Yushu Zhang 0001, Rushi Lan, Zhongyun Hua, Jian Weng 0001
IEEE Trans. Image Process.3
2025 Semi-Supervised Gland Segmentation via Feature-Enhanced Contrastive Learning and Dual-Consistency Strategy
abstract
In the field of gland segmentation in histopathology, deep-learning methods have made significant progress. However, most existing methods not only require a large amount of high-quality annotated data but also tend to confuse the internal of the gland with the background. To address this challenge, we propose a new semi-supervised method named DCCL-Seg for gland segmentation, which follows the teacher-student framework. Our approach can be divided into follows steps. First, we design a contrastive learning module to improve the ability of the student model's feature extractor to distinguish between gland and background features. Then, we introduce a Signed Distance Field (SDF) prediction task and employ dual-consistency strategy (across tasks and models) to better reinforce the learning of gland internal. Next, we proposed a pseudo label filtering and reweighting mechanism, which filters and reweights the pseudo labels generated by the teacher model based on confidence. However, even after reweighting, the pseudo labels may still be influenced by unreliable pixels. Finally, we further designed an assistant predictor to learn the reweighted pseudo labels, which do not interfere with the student model's predictor and ensure the reliability of the student model's predictions. Experimental results on the publicly available GlaS and CRAG datasets demonstrate that our method outperforms other semi-supervised medical image segmentation methods.
Jiejiang Yu, Xipeng Pan, Zhenwei Shi 0002, Huadeng Wang, Rushi Lan
IEEE J. Biomed. Health Informatics6
2025 Learning Distinguishable Degradation Maps for Unknown Image Super-Resolution
abstract
Most existing super-resolution (SR) methods assume that the degradation is fixed (e.g., bicubic downsampling), whereas their performance would be degraded if the actual degradation differs from this assumption. To deal with unknown degradations, existing unknown SR methods are committed to learning degradation representation to generate high-resolution images. Nevertheless, they ignore that the impact of degradations on images is related to image content, or they learn degradation representations without any constraints. In this article, we propose a degradation maps extractor for unknown SR. Specifically, we learn degradation maps and condense them into a one-dimensional representation space to distinguish various degradations, which obtains distinguishable degradation maps and preserves the connection with the image contents. Furthermore, we propose a degradation map-guided SR (DMGSR) network, in which the degradation maps adaptively influence the SR process by applying channel attention and spatial attention to middle features. With the cooperation of the degradation maps extractor and the degradation maps-guided SR network, our network can flexibly handle various degradations. Experimental results show that our model achieves state-of-the-art performance in quantitative and qualitative metrics for the unknown SR task.
Zhenbing Liu, Haoxiang Lu, Rushi Lan
IEEE Trans. Multim.5
2025 AES-AUDIO: An Encryption Scheme for Audio Supporting Differentiated Decryption
abstract
In this paper, we propose an audio encryption scheme that supports differentiated decryption, called AES-AUDIO, in which an audio only needs to be encrypted once and can be decrypted into different resolutions as needed. First, we design four security levels, confidential, harsh, noisy, and clear, based on the audio resolution perceived by human auditory perception. Second, the audio data in decimal floating-point numbers (D-FPNs) are unfolded to 32 bits (B-FPNs). Third, we design a region of interest (RoI) encryption algorithm for the audio with the B-FPN format, where the result preserves some perceptual information as needed. Fourth, we construct the AES-AUDIO scheme based on the RoI encryption algorithm, which allows the audio to be encrypted once and then decrypted into different security levels. It supports changing parameters to alter the perception effect corresponding to the security level. Overall, it achieves a balance between the security and usability of the protected audio. User experiments verify that the audios produced by differential decryption can achieve the expected security levels. Some security tests also yielded excellent results, such as an NSCR value of 1.
Yushu Zhang 0001, Junhao Ji, Wenying Wen, Rushi Lan
IEEE Trans. Multim.6
2025 Deepfake Video Detection Using Facial Feature Points and Ch-Transformer
abstract
With the development of Metaverse technology, the avatar in Metaverse has faced serious security and privacy concerns. Analyzing facial features to distinguish between genuine and manipulated facial videos holds significant research importance for ensuring the authenticity of characters in the virtual world and for mitigating discrimination as well as preventing malicious use of facial data. To address this issue, the Facial Feature Points and Class-head-Transformer (FFP-ChT) deepfake video detection model is designed based on the clues of different FFPs distribution in real and fake videos and different displacement distances of real and fake FFPs between frames. The face video input is first detected by the BlazeFace model, and the face detection results are fed into the FaceMesh model to extract 468 FFPs. Then, the Lucas–Kanade (LK) optical flow method is used to track the points of the face, the face calibration algorithm is introduced to re-calibrate the FFPs, and the jitter displacement is calculated by tracking the FFPs between frames. Finally, the Ch is designed in the transformer, and the FFPs and FFP displacement are jointly classified through the ChT model. In this way, the designed ChT classifier is able to accurately and effectively identify deepfake videos. Experiments on open datasets clearly demonstrate the effectiveness and generalization capabilities of our approach.
Rui Yang 0018, Rushi Lan, Zhenrong Deng, Xiyan Sun
ACM Trans. Multim. Comput. Commun. Appl.2
2024 EOFD-Net: Edge Optimization and Feature Denoising for Weakly Supervised Deep Nuclei Segmentation with Point Annotations
abstract
Nuclei segmentation is a fundamental and critical step in digital pathological image analysis. Fully supervised nuclei segmentation requires a lot of pixel-by-pixel manual annotation by pathologists, which is very time-consuming and laborious. To minimize the labeling burden of pathologists, this paper uses only point annotations of nuclei data for weakly supervised learning. Specifically, a two-stage model named EOFD-Net with feature denoising and edge optimization is proposed. In the first stage, three weak labels (K-means cluster labels, Voronoi labels, and superpixel labels) with complementary information are used to train the encoder-decoder network to achieve coarse segmentation of nuclei. A feature denoising module(FDM) is designed in the encoder part, which can effectively reduce noise interference. In the second stage, we designed an edge optimization strategy using the prior knowledge of the trained model in the first stage. Confident learning is employed to denoise pseudo-label and rectify the mislabel. These optimized labels are input into the second stage to obtain the final segmentation results. The performance of our method outperforms current state-of-the-art methods on two publicly nuclei segmentation datasets, MoNuSeg and TNBC.
Xipeng Pan, Feihu Hou, Zhenbing Liu, Siyang Feng, Rushi Lan
ICASSP5
2024 Local Optimization Networks for Multi-View Multi-Person Human Posture Estimation
abstract
With the growing applicability of multi-view multi-person 3D human pose estimation across diverse scenarios, the impact of external environmental factors and occlusion on accuracy has garnered substantial attention. In this research, we introduce a novel approach to multi-view multi-person 3D human pose estimation, leveraging a localized optimization strategy. Specifically, our method enhances the interplay of feature information from different channels and fine-tunes the optimal feature weights to capture intricate dependencies among joints. This refinement leads to improved accuracy in handling external environmental factors. Experimental evaluations were conducted on two prominent benchmark datasets, namely Campus and Shelf. The proposed method achieved a remarkable performance, with a Percentage of Correct Parts (PCP) score of 97.4% and 98.2% for the Campus and Shelf datasets, respectively.
Jucheng Song, Chi-Man Pun, Haolun Li 0001, Rushi Lan, Jiucheng Xie, Hao Gao 0005
ICASSP4
2024 Gland Segmentation Via Dual Encoders and Boundary-Enhanced Attention
abstract
Accurate and automated gland segmentation on pathological images can assist pathologists in diagnosing the malignancy of colorectal adenocarcinoma. However, due to various gland shapes, severe deformation of malignant glands, and overlapping adhesions between glands. Gland segmentation has always been very challenging. To address these problems, we propose a DEA model. This model consists of two branches: the backbone encoding and decoding network and the local semantic extraction network. The backbone encoding and decoding network extracts advanced Semantic features, uses the proposed feature decoder to restore feature space information, and then enhances the boundary features of the gland through boundary enhancement attention. The local semantic extraction network uses the pre-trained DeepLabv3+ as a Local semantic-guided encoder to realize the extraction of edge features. Experimental results on two public datasets, GlaS and CRAG, confirm that the performance of our method is better than other gland segmentation methods.
Huadeng Wang, Jiejiang Yu, Xipeng Pan, Zhenbing Liu, Rushi Lan
ICASSP6
2024 Mining Gold from the Sand: Weakly Supervised Histological Tissue Segmentation with Activation Relocalization and Mutual Learning
Siyang Feng, Zhenbing Liu, Wentao Liu 0004, Zimin Wang, Rushi Lan, Xipeng Pan
MICCAI (8)6
2024 PG-MLIF: Multimodal Low-Rank Interaction Fusion Framework Integrating Pathological Images and Genomic Data for Cancer Prognosis Prediction
Xipeng Pan, Yajun An, Rushi Lan, Zhenbing Liu, Zaiyi Liu, Cheng Lu 0001
MICCAI (3)3
2024 Enhanced Spatial Adaptive Fusion Network For Video Super-Resolution
Boyue Li, Shiqian Yuan, Rushi Lan
PRCV (6)4
2024 Semi-supervised Gland Segmentation via Label Purification and Reliable Pixel Learning
Huadeng Wang, Lingqi Zeng, Jiejiang Yu, Xipeng Pan, Rushi Lan
PRCV (15)6
2024 Residual Hybrid Attention Enhanced Video Super-Resolution with Cross Convolution
Shiqian Yuan, Boyue Li, Rushi Lan
PRCV (6)4
2024 Real-Time Detection Transformer with Bi-Level Routing Attention
Shiqian Yuan, Boyue Li, Rushi Lan
PRCV (4)4
2024 3D mesh encryption with differentiated visual effect and high efficiency based on chaotic system
Yushu Zhang 0001, Wenying Wen, Rushi Lan
Expert Syst. Appl.6
2024 Brighten up Images via Dual-Branch Structure-Texture Awareness Feature Interaction
abstract
Images captured under low-light conditions suffer from inevitable degradation leading to the missing global structure and detailed local texture. However, existing methods consider these two components as a single entity or perform a similar convolutional operation, which can yield suboptimal results. In this letter, we propose a dual-branch structure-texture awareness feature interaction network named DFINet to tackle the above problems. First, we generate structure and texture components through the Gaussian operator. Subsequently, we conduct CNN-based and Transformer-based branches to cope with the texture and structure components separately. Among them, we design a Feature Interaction Block that leverages local-global information to enrich features in the encoding phase. Then, we generate queries with the potential structural-texture cues for the Transformer blocks in the decoding phase. Finally, we develop a Fusion Block to progressively integrate cross-layer features from two branches for the reconstruction. Our extensive experiment indicates the proposed method outperforms several representative methods in terms of both visual quality and objective assessment.
Yingxin Huang, Zhenbing Liu, Haoxiang Lu, Rushi Lan
IEEE Signal Process. Lett.5
2024 Blood Vessel Segmentation via Topology Interaction and Contrast
abstract
The topology integrity of blood vessel segmentation is crucial for clinical disease diagnosis. However, existing methods for enhancing vessel topology have overlooked the degree and type of topology interaction, and have not incorporated topology contrast relation into consideration. In this letter, we propose confindence-based topolopy interaction and topology contrast loss method to enhance the interaction and contrast relations for intra-class and inter-class of topology structure. Additionally, considering the influence of multiscale feature on topology learning, we propose lightweight global feature extraction and align fusion to better capture global features and mitigate feature misalignment. The quantitative and qualitative comparisons on the DRIVE and STARE datasets show the superiority of the proposed model. Furthermore, the improvements experiments of the advanced topology enhancement networks using our proposed methods, and ablation experiments on DCA1 datasets provide evidence of the effectiveness of the proposed method.
Huadeng Wang, Wenbin Zuo, Xipeng Pan, Rushi Lan
IEEE Signal Process. Lett.4
2024 Client-Side Watermarking for Images With Solid-Color Backgrounds
abstract
The ever-growing popularity of image sharing underscores the claim for copyright protection. Digital watermarking plays a pivotal role in combating illegal redistribution and safeguarding copyright. This methodology includes two embedding modes, namely owner-side and client-side embedding, with the latter having an advantage in terms of owner-side efficiency. However, all current client-side watermarking schemes overlook the applicability to images with solid-color backgrounds. Building upon existing researches, this letter uses the lookup table method to realize the client-side embedding of watermarking, and puts forward two schemes specifically tailored for images with solid-color backgrounds. Both two schemes achieve global encryption while successfully limiting the residual watermark effects after decryption to the subject matter of the image. Separately, one scheme adopts the principle of spread spectrum, while the other relies on the principle of spread transform dither modulation. These two schemes exhibit distinct performance trade-offs, and through experimental verification, we demonstrate their decent visual performance, robustness, and efficiency.
Xiangli Xiao, Rushi Lan, Wenying Wen, Yushu Zhang 0001
IEEE Signal Process. Lett.2
2024 4DPM: Deepfake Detection With a Denoising Diffusion Probabilistic Mask
abstract
In the face of increasingly realistic fake human faces, research on enhancing the differences between real and fake images is valuable for improving the generalization capabilities of fake face detection models. In this letter, we propose a method called DPMask (Diffusion Probabilistic Mask) to amplify the distinctions between authentic and counterfeit human facial images. Specifically, we use a dataset consisting of real human facial images and Simplex noise to train a denoising diffusion probabilistic model for the proposed DPMask. Subsequently, we separately apply the DPMask and U-Net to real and fake human facial images to create noticeably distinct genuine and counterfeit human facial images. A lightweight classification network blue is further designed based on RepVGG to classify the newly generated real and fake human faces. Experimental results demonstrate that our model achieves high accuracy on a manually created fake face dataset (RFFD), a GAN-generated fake face dataset (Seq-DeepFake), and a DDPM-generated face dataset (HiFi-IFDL). Furthermore, the addition of DPMask significantly improves the performance of some public fake face detection models.
Rui Yang 0018, Zhenrong Deng, Yushu Zhang 0001, Rushi Lan
IEEE Signal Process. Lett.5
2024 A Consumer-Oriented Image Transformation Scheme With a Secret Key for Privacy Protection
abstract
Images in electronic devices may pose privacy threats since they can capture sensitive information about consumers. Meanwhile, face recognition (FR) systems are widely used, exacerbating these concerns. Some learning-based schemes have been proposed to protect the privacy of facial images. However, consumers might encounter challenges in implementing them due to specific requirements related to computing power and professional background. The non-learning semantic adversarial perturbation schemes address the aforementioned issues, but they are either irreversible or necessitate additional storage space. To this end, we propose a consumer-oriented image transformation scheme that can prevent the recognition of images by FR systems. It offers a consumer-friendly method compared to schemes that need high-performance equipment and specialized knowledge. The proposed scheme is implemented by reversible embedding the key to coefficients during the JPEG compression. The proposed scheme can be operated in general computing devices by ordinary consumers. The experiments demonstrated that facial images can be protected by the proposed scheme.
Wenying Wen, Youwen Zhu, Rushi Lan
IEEE Signal Process. Lett.4
2024 Revamping Blood Vessel Edge-Buffer Labels: A Self-Correcting Region Supervision
abstract
Deep learning-based vessel segmentation tasks serve as important auxiliary tools for disease diagnosis. However, the region connecting the foreground and background, which is named the edge-buffer region in this letter, suffers from noisy labels and a lack of discriminative features due to low contrast and limitations of imaging devices. To address these limitations, we propose a self-correcting region supervision to revamp the noisy labels in the edge-buffer region. Furthermore, we introduce the concept of treating the edge-buffer region independently from the foreground and background, leveraging the designed contrastive learning method and edge-blur-guided module to enhance discriminative learning ability and collaborative learning ability across different regions, respectively. The experimental results comparison with other classical and state-of-the-art methods on DRIVE, CHASEDB1, and DCA1 datasets has proven the effectiveness of the proposed methods.
Wenbin Zuo, Huadeng Wang, Xipeng Pan, Rushi Lan
IEEE Signal Process. Lett.4
2024 Spatial-Frequency Discriminability for Revealing Adversarial Perturbations
abstract
The vulnerability of deep neural networks to adversarial perturbations has been widely perceived in the computer vision community. From a security perspective, it poses a critical risk for modern vision systems, e.g., the popular Deep Learning as a Service (DLaaS) frameworks. For protecting deep models while not modifying them, current algorithms typically detect adversarial patterns through discriminative decomposition for natural and adversarial data. However, these decompositions are either biased towards frequency resolution or spatial resolution, thus failing to capture adversarial patterns comprehensively. Also, when the detector relies on few fixed features, it is practical for an adversary to fool the model while evading the detector (i.e., defense-aware attack). Motivated by such facts, we propose a discriminative detector relying on a spatial-frequency Krawtchouk decomposition. It expands the above works from two aspects: 1) the introduced Krawtchouk basis provides better spatial-frequency discriminability, capturing the differences between natural and adversarial data comprehensively in both spatial and frequency distributions, w.r.t. the common trigonometric or wavelet basis; 2) the extensive features formed by the Krawtchouk decomposition allows for adaptive feature selection and secrecy mechanism, significantly increasing the difficulty of the defense-aware attack, w.r.t. the detector with few fixed features. Theoretical and numerical analyses demonstrate the uniqueness and usefulness of our detector, exhibiting competitive scores on several deep models and image sets against a variety of adversarial attacks.
Chao Wang 0028, Yushu Zhang 0001, Rushi Lan, Xiaochun Cao, Fenglei Fan
IEEE Trans. Circuits Syst. Video Technol.5
2024 Visually Semantics-Aware Color Image Encryption Based on Cross-Plane Substitution and Permutation
abstract
The concept of thumbnail-preserving encryption (TPE) is proposed to strike a balance between image privacy and visual usability. It allows users to have both without having to choose between privacy and usability, but it processes the information within a channel and not between channels. Recently, two cross-plane TPE schemes have been proposed in an attempt to process the information between channels. However, in a sense, they are not completely cross-plane. In addition, as with other TPE schemes, they do not respect the syntax of the color image. They simply treat each channel as a regular numerical matrix for the substitution encryption, and thus there are issues of out of bounds and some elements being left unencrypted. In this article, we propose a TPE scheme for cross-plane substitution and permutation. As with other TPE schemes, a certain amount of visually semantics-aware information is preserved in the encrypted image. Meanwhile, this TPE scheme is the first to implement a color image compliant syntax. In general, a color image composes of three channels, meaning a (pixel) point has three elements. In the proposed scheme, substitution encryption can directly be used for a point (three elements across planes). It eliminates possibly going out of bounds and ensures that all elements in the image are encrypted during substitution.
Yushu Zhang 0001, Rushi Lan
IEEE Trans. Ind. Informatics5
2024 Single Traffic Image Deraining via Similarity-Diversity Model
abstract
Single traffic image deraining technology based on deep learning is a vital branch of image preprocessing, which is of great help to intelligent monitoring systems and driving navigation system. It is well understood that established deraining methods are derived based on one specific imaging model, neglecting the underlying correlations between different weather models and thereby limiting the applicability of these standard methods in real scenarios. To ameliorate this issue, in this work, we first explore the inherent relationship between a rain model and the haze one established up to date. We discover that these two models experience similar degradations in the low-frequency components (i.e., similarity) but diverse degradations in the high-frequency areas (i.e., diversity). Based on these observations, we develop a Similarity-Diversity model to describe these characteristics. Afterwards, we introduce a novel deep neural network to restore the rain-free background embedding the similarity-diversity model, namely deep similarity-diversity network (DSDNet). Extensive experiments have been conducted to evaluate our proposed method that outperforms the other state of the art deraining techniques. On the other hand, we deploy the proposed algorithm with Google Vision API for object recognition, which also obtains satisfactory results both qualitatively and quantitatively.
Youxing Li, Rushi Lan, Huiwen Huang, Huiyu Zhou 0001, Zhenbing Liu
IEEE Trans. Intell. Transp. Syst.2
2024 User Behavior Threat Detection Based on Adaptive Sliding Window GAN
abstract
User behavior threat detection is important for the protection of network system security. Traditional supervised modeling methods and unbalanced sample data lead to a high false positive rate in user behavior detection. In addition, network user behaviors are complex, changeable, and difficult to predict, and existing detection methods are facing ever greater challenges. Effectively detecting user behavior remains a challenge. In this paper, we propose a user behavior threat detection method based on an Adaptive Sliding Window Generative Adversarial Network(ASW-GAN). This method designs an adaptive sliding window mechanism to process behavior data and uses the GAN model to detect threat behavior, finally uses the maximum interclass variance algorithm Otsu to optimize test detection result. Compared with other typical methods, the proposed method achieves a higher accuracy rate and a markedly lower false positive rate, and can effectively evaluate user threat behaviors.
Xiaoling Tao, Shen Lu, Feng Zhao 0002, Rushi Lan, Longsheng Chen, Lianyou Fu, Ruchun Jia
IEEE Trans. Netw. Serv. Manag.4
2024 E-TPE: Efficient Thumbnail-Preserving Encryption for Privacy Protection in Visual Sensor Networks
abstract
When visual sensor networks (VSNs) enter daily life in society, they not only bring great convenience but also cause people to worry about privacy. Traditional image encryption uses the snowflake effect to protect the privacy, but this can compromise the usability of directly browsing images captured by nodes in VSNs. Recently, thumbnail-preserving encryption (TPE) has been proposed to balance image usability and privacy. However, the Markov chain’s limited connectivity and high time cost, which result in low efficiency, may preclude its application in VSNs. Motivated by this, we propose a novel TPE scheme to protect the image collected in VSNs’ privacy efficiently. We first deeply study the bijective relationship between the three-pixel group and its rank, and then propose a new number method for both, based on which we design the efficient rank mapping. Subsequently, an efficient thumbnail-preserving image encryption via three-pixel rank mapping (E-TPE) is designed, achieving NR security and enhancing the Markov chain’s connectivity. The experiments demonstrate the scheme’s efficiency; that is, it is currently the quickest ideal-TPE scheme with NR security, and a good balance is achieved between usability and privacy.
Yushu Zhang 0001, Wenying Wen, Rushi Lan, Yong Xiang 0001
ACM Trans. Sens. Networks4
2024 Enhancing infrared images via multi-resolution contrast stretching and adaptive multi-scale detail boosting
Haoxiang Lu, Zhenbing Liu, Xipeng Pan, Rushi Lan
Vis. Comput.4
2023 Retinex-inspired contrast stretch and detail boosting for lowlight image enhancement
abstract
Abstract Lowlight images with low brightness and contrast, blurry details usually bring us an uncomfortable visual experience. To promote the quality of these deviation images, this paper presents a new and efficient approach, named MFMR, for enhancing lowlight images in the hue‐saturation‐value (HSV) colour space. Concretely, the multi‐angle filter is first applied to estimate the artifact‐free illumination and reflection component of the V‐channel. Afterward, the adaptive bi‐interval histogram with human visual characteristics and morphological operations is employed to process the former, adaptive gamma correction to process the latter for generating various feature maps. In the end, these feature maps are united via adaptive multi‐scale fusion strategy to reconstruct high‐quality images, which are characterized by high contrast and brightness, vivid colour, and clearer details. Extensive experiments show that this method is a well‐proven low‐light image enhancement approach, which outperforms the state‐of‐the‐art comparison methods. Furthermore, the proposed method also can yield satisfying images in the heavy foggy, yellow sand, underwater, and other severe conditions.
Haoxiang Lu, Zhenbing Liu, Rushi Lan, Xipeng Pan, Junming Gong
IET Image Process.3
2023 Heterogeneous and Customized Cost-Efficient Reversible Image Degradation for Green IoT
abstract
With the large-scale deployment of the Internet of Things (IoT) in daily life, more and more privacy data are collected by IoT devices. These data are not directly physically controlled by users, which may cause privacy concerns. In fact, privacy has become one of the significant problems faced by IoT. In this article, we mainly study the protection of image privacy under the green IoT. We have conducted an in-depth analysis of the green IoT scenario and put forward the scope and corresponding goals that the scheme should have. Motivated by this, a novel image privacy protection scheme is proposed, i.e., heterogeneous and customized cost-efficient reversible image degradation for green IoT. This scheme fully considers the characteristics of privacy and the various users’ diverse requirements to achieve a heterogeneous and customized privacy protection. Meanwhile, cost effectiveness cannot be confined to the efficiency of the direct image processing at the expense of greatly increasing costs in other aspects, such as transmission and reversion. It is mitigated by preserving some visual content in the privacy-protected image. It also improves the image compression efficiency and ensures that the user can select the desired image according to the visual content for reversion. Some experiments have been carried out to demonstrate that this work has achieved the proposed scope and corresponding goals.
Yushu Zhang 0001, Rushi Lan, Zhongyun Hua, Yong Xiang 0001
IEEE Internet Things J.3
2023 PCRTAM-Net: A Novel Pre-Activated Convolution Residual and Triple Attention Mechanism Network for Retinal Vessel Segmentation
Huadeng Wang, Zi-Zheng Li, Idowu Paul Okuwobi, Xipeng Pan, Zhenbing Liu, Rushi Lan
J. Comput. Sci. Technol.7
2023 SMILE: Cost-sensitive multi-task learning for nuclear segmentation and classification with imbalanced annotations
Xipeng Pan, Jijun Cheng, Feihu Hou, Rushi Lan, Cheng Lu 0001, Lingqiao Li, Zhengyun Feng, Huadeng Wang, Changhong Liang, Zhenbing Liu, Xin Chen 0058, Chu Han, Zaiyi Liu
Medical Image Anal.4
2023 Identifiable Face Privacy Protection via Virtual Identity Transformation
abstract
Massive face images collected in smart surveillance and social networks are vulnerable to malicious access, thus compromising individual privacy. Existing schemes have been able to protect face privacy while preserving a certain level of identifiability, but have different limitations, e.g., the lack of strong transferability or the inability to retain irrelevant attributes. This letter proposes a novel face privacy protection scheme via virtual identity transformation, which guarantees strong privacy protection and high identifiability. We first solve a specific identity mask for the user, which ensures that the identity features extracted only from the user's faces can be approximated to the given virtual identity. Based on it, the identity transformation networks transform the original face into the protected form, which belongs to the virtual identity while retaining irrelevant attributes. Lastly, the virtual identity of the protected face is extracted for face recognition. Adequate experiments show that our scheme has satisfactory privacy protection, high identifiability, and strong transferability.
Tao Wang 0084, Yushu Zhang 0001, Wenying Wen, Rushi Lan
IEEE Signal Process. Lett.5
2023 Usability Enhanced Thumbnail-Preserving Encryption Based on Data Hiding for JPEG Images
abstract
As an increasing number of people tend to upload personal images to cloud platforms such as iCloud, Google Drive, and Baidu Cloud, the issue of image privacy protection in clouds has attracted widespread attention. Thumbnail-preserving encryption (TPE) is a newly proposed technology especially for balancing usability and privacy of images stored in clouds. The TPE-encrypted image does not reveal anything other than the thumbnail and size of the original one. Although many well-designed TPE schemes have been put forward, they almost all operate in the spatial domain and are not compatible with JPEG images. The only two schemes that can be applied in the frequency domain were proposed by Wright et al. and Marohn et al., but both of them have disadvantages and do not consider non-visual usability. In this paper, we propose a usability enhanced TPE scheme especially for JPEG images by utilizing reversible data hiding and image encryption in the frequency domain. We first transform the original image into the frequency domain and select some specific locations for extra information recording in every block. The original coefficient values at the selected locations are then collected and embedded into the rest of the block before encryption. Finally, the extra information is expected to be recorded into the selected locations of every block for non-visual usability. Experiments confirm that the proposed scheme achieves the target of balancing enhanced usability and privacy.
Xi Ye 0004, Yushu Zhang 0001, Xiangli Xiao, Rushi Lan
IEEE Signal Process. Lett.5
2023 Adversarial Thumbnail-Preserving Transformation for Facial Images Based on GAN
abstract
People are used to capturing their daily lives with images and uploading them to the cloud for easy archiving. While cloud storage services bring convenience to users, malicious humans and artificial intelligence machines may pose privacy threats to these uploaded images. Thumbnail-preserving encryption (TPE) is a solution for handling privacy threats posed by the usage of cloud storage and it is able to balance the privacy against humans and usability for users. However, the present TPE schemes display certain vulnerabilities when subjected to machine-based attacks and when the size of the thumbnail block is small, the facial recognition (FR) models exhibit strong capability in recognizing the identity within the protected facial images. To this end, we propose an adversarial thumbnail-preserving transform scheme for facial images, which can not only balance privacy and usability but also resist recognition from FR models. The experiments demonstrate that the proposed scheme has a better ability to resist machine recognition compared with existing schemes.
Yushu Zhang 0001, Rushi Lan
IEEE Signal Process. Lett.5
2022 Superpixel-guided locality quaternion representation for color face hallucination
Licheng Liu, Xiaoqin Tang, C. L. Philip Chen, Luyang Cai, Rushi Lan
Inf. Sci.5
2022 Primitively visually meaningful image encryption: A new paradigm
Yushu Zhang 0001, Yu Nan, Wenying Wen, Xiu-Li Chai, Rushi Lan
Inf. Sci.6
2022 Noise-free thumbnail-preserving image encryption based on MSB prediction
Ye Zhu 0002, Yushu Zhang 0001, Xiangli Xiao, Rushi Lan, Yong Xiang 0001
Inf. Sci.5
2022 PRA-TPE: Perfectly Recoverable Approximate Thumbnail-Preserving Image Encryption
Xi Ye 0004, Yushu Zhang 0001, Rushi Lan, Yong Xiang 0001
J. Vis. Commun. Image Represent.4
2022 Diagnosis of Alzheimer's disease via an attention-based multi-scale convolutional neural network
Zhenbing Liu, Haoxiang Lu, Xipeng Pan, Mingchang Xu, Rushi Lan
Knowl. Based Syst.5
2022 Meta multi-task nuclei segmentation with fewer training samples
Chu Han, Huasheng Yao, Bingchao Zhao, Zhenhui Li, Zhenwei Shi 0002, Xin Chen 0058, Jinrong Qu, Rushi Lan, Changhong Liang, Xipeng Pan, Zaiyi Liu
Medical Image Anal.10
2022 HF-TPE: High-Fidelity Thumbnail- Preserving Encryption
abstract
With the popularity of cloud storage services, people are increasingly accustomed to storing images in the cloud. However, cloud storage services raise privacy concerns, e.g., leakage of images to unauthorized third parties and service providers may exploit image detection technologies to portrait users without permission. Although privacy concerns can be solved by encrypting images before they are uploaded to the cloud, traditional encryption methods significantly affect the usability and user experience, for example, users cannot preview images in the cloud. Recently, Marohnet al.proposed two approximate thumbnail-preserving encryption schemes, called DRPE and TPE-LSB, to balance the privacy and usability of images in the cloud. However, both schemes have defects that either the decryption may fail or the ciphertext images have poor performance in perceived quality and too many noise points after decryption. To this end, we pertinently propose a high-fidelity thumbnail-preserving encryption scheme (HF-TPE). Compared with the previous works, on the one hand, the HF-TPE scheme not only ensures the correct decryption of ciphertext images, but also makes the ciphertext thumbnails more close to the plaintext images perceptually. On the other hand, the decrypted thumbnails have lower noise intensity and upper limit of the number of noises. In addition, simulation experiments further show that the HF-TPE scheme can guarantee users’ usability.
Yushu Zhang 0001, Xiangli Xiao, Rushi Lan, Zhe Liu 0001, Xinpeng Zhang 0001
IEEE Trans. Circuits Syst. Video Technol.4
2022 Label Guided Discrete Hashing for Cross-Modal Retrieval
abstract
Due to their low storage capacity and fast retrieval speed, hashing techniques have received much attention in cross-modal retrieval. However, there are some issues that need to be further explored. First, some existing hashing methods use the labels to construct the semantic similarity matrix between pairwise data, ignoring the potential manifold structure between heterogeneous data. Second, some existing methods underestimate the importance of multi-label and the gaps between different class labels, making the learned hash codes less discriminative. Third, few of them embed both manifold and balanced structures within the same model, and the relaxation of discrete constraints will lead to an increasing quantization error. To mitigate these problems, this paper proposes a novel supervised hashing method, termed Label Guided Discrete Hashing (LGDH), which simultaneously preserves the comprehensive manifold structure and discriminative balanced codes that are both constructed by label information into Hamming space. We develop a local category distribution of the nearest neighbors, to excavate the underlying manifold structure of heterogeneous data. To maximize the gaps of different categories, a balanced matrix is constructed by labels to generate hash codes with balanced bits. For multi-label data, we also design a novel multi-label manifold and balanced structure matrix to adapt the real-world scenarios. An effective discrete optimization method is used to optimize our proposed objective function instead of the relaxation one. Extensive experiments on three benchmark datasets verify the effectiveness of LGDH. The comparison results demonstrat that LGDH achieves about 2% and 3% improved to different cross-modal tasks on average.
Rushi Lan, Yu Tan, Zhenbing Liu
IEEE Trans. Intell. Transp. Syst.1
2022 Virtual Reality Aided High-Quality 3D Reconstruction by Remote Drones
abstract
Artificial intelligence including deep learning and 3D reconstruction methods is changing the daily life of people. Now, an unmanned aerial vehicle that can move freely in the air and avoid harsh ground conditions has been commonly adopted as a suitable tool for 3D reconstruction. The traditional 3D reconstruction mission based on drones usually consists of two steps: image collection and offline post-processing. But there are two problems: one is the uncertainty of whether all parts of the target object are covered, and another is the tedious post-processing time. Inspired by modern deep learning methods, we build a telexistence drone system with an onboard deep learning computation module and a wireless data transmission module that perform incremental real-time dense reconstruction of urban cities by itself. Two technical contributions are proposed to solve the preceding issues. First, based on the popular depth fusion surface reconstruction framework, we combine it with a visual-inertial odometry estimator that integrates the inertial measurement unit and allows for robust camera tracking as well as high-accuracy online 3D scan. Second, the capability of real-time 3D reconstruction enables a new rendering technique that can visualize the reconstructed geometry of the target as navigation guidance in the HMD. Therefore, it turns the traditional path-planning-based modeling process into an interactive one, leading to a higher level of scan completeness. The experiments in the simulation system and our real prototype demonstrate an improved quality of the 3D model using our artificial intelligence leveraged drone system.
Feng Xu 0005, Chi-Man Pun, Yang Yang 0002, Rushi Lan, Yujie Li 0001, Hao Gao 0005
ACM Trans. Internet Techn.5
2022 Binary Representation via Jointly Personalized Sparse Hashing
abstract
Unsupervised hashing has attracted much attention for binary representation learning due to the requirement of economical storage and efficiency of binary codes. It aims to encode high-dimensional features in the Hamming space with similarity preservation between instances. However, most existing methods learn hash functions in manifold-based approaches. Those methods capture the local geometric structures (i.e., pairwise relationships) of data, and lack satisfactory performance in dealing with real-world scenarios that produce similar features (e.g., color and shape) with different semantic information. To address this challenge, in this work, we propose an effective unsupervised method, namely, Jointly Personalized Sparse Hashing (JPSH), for binary representation learning. To be specific, first, we propose a novel personalized hashing module, i.e., Personalized Sparse Hashing (PSH). Different personalized subspaces are constructed to reflect category-specific attributes for different clusters, adaptively mapping instances within the same cluster to the same Hamming space. In addition, we deploy sparse constraints for different personalized subspaces to select important features. We also collect the strengths of the other clusters to build the PSH module with avoiding over-fitting. Then, to simultaneously preserve semantic and pairwise similarities in our proposed JPSH, we incorporate the proposed PSH and manifold-based hash learning into the seamless formulation. As such, JPSH not only distinguishes the instances from different clusters but also preserves local neighborhood structures within the cluster. Finally, an alternating optimization algorithm is adopted to iteratively capture analytical solutions of the JPSH model. We apply the proposed representation learning algorithm JPSH to the similarity search task. Extensive experiments on four benchmark datasets verify that the proposed JPSH outperforms several state-of-the-art unsupervised hashing algorithms.
Chen Chen 0151, Rushi Lan, Licheng Liu, Zhenbing Liu, Huiyu Zhou 0001
ACM Trans. Multim. Comput. Commun. Appl.3
2021 A Two-stage Chinese text summarization algorithm using keyword information and adversarial learning
Zhenrong Deng, Ma Fuxin, Rushi Lan, Wenming Huang
Neurocomputing3
2021 Sensorineural hearing loss classification via deep-HLNet and few-shot learning
Rushi Lan, Shuihua Wang, Yudong Zhang 0001
Multim. Tools Appl.3
2021 Bilinear pyramid network for flower species categorization
Rushi Lan, Zhuo Shi
Multim. Tools Appl.3
2021 TPE2: Three-Pixel Exact Thumbnail-Preserving Image Encryption
Yushu Zhang 0001, Xiangli Xiao, Xi Ye 0004, Rushi Lan
Signal Process.5
2021 Infrared Image Super-Resolution via Transfer Learning and PSRGAN
abstract
Recent advances in single image super-resolution (SISR) demonstrate the power of deep learning for achieving better performance. Because it is costly to recollect the training data and retrain the model for infrared (IR) image super-resolution, the availability of only a few samples for restoring IR images presents an important challenge in the field of SISR. To solve this problem, we first propose the progressive super-resolution generative adversarial network (PSRGAN) that includes the main path and branch path. The depthwise residual block (DWRB) is used to represent the features of the IR image in the main path. Then, the novel shallow lightweight distillation residual block (SLDRB) is used to extract the features of the readily available visible image in the other path. Furthermore, inspired by transfer learning, we propose the multistage transfer learning strategy for bridging the gap between different high-dimensional feature spaces that can improve the PSGAN performance. Finally, quantitative and qualitative evaluations of two public datasets show that PSRGAN can achieve better results compared to the SR methods.
Yongsong Huang, Zetao Jiang, Rushi Lan, Shaoqin Zhang, Kui Pi
IEEE Signal Process. Lett.3
2021 MADNet: A Fast and Lightweight Network for Single-Image Super Resolution
abstract
Recently, deep convolutional neural networks (CNNs) have been successfully applied to the single-image super-resolution (SISR) task with great improvement in terms of both peak signal-to-noise ratio (PSNR) and structural similarity (SSIM). However, most of the existing CNN-based SR models require high computing power, which considerably limits their real-world applications. In addition, most CNN-based methods rarely explore the intermediate features that are helpful for final image recovery. To address these issues, in this article, we propose a dense lightweight network, called MADNet, for stronger multiscale feature expression and feature correlation learning. Specifically, a residual multiscale module with an attention mechanism (RMAM) is developed to enhance the informative multiscale feature representation ability. Furthermore, we present a dual residual-path block (DRPB) that utilizes the hierarchical features from original low-resolution images. To take advantage of the multilevel features, dense connections are employed among blocks. The comparative results demonstrate the superior performance of our MADNet model while employing considerably fewer multiadds and parameters.
Rushi Lan, Zhenbing Liu, Huimin Lu 0001
IEEE Trans. Cybern.1
2021 Cascading and Enhanced Residual Networks for Accurate Single-Image Super-Resolution
abstract
Deep convolutional neural networks (CNNs) have contributed to the significant progress of the single-image super-resolution (SISR) field. However, the majority of existing CNN-based models maintain high performance with massive parameters and exceedingly deeper structures. Moreover, several algorithms essentially have underused the low-level features, thus causing relatively low performance. In this article, we address these problems by exploring two strategies based on novel local wider residual blocks (LWRBs) to effectively extract the image features for SISR. We propose a cascading residual network (CRN) that contains several locally sharing groups (LSGs), in which the cascading mechanism not only promotes the propagation of features and the gradient but also eases the model training. Besides, we present another enhanced residual network (ERN) for image resolution enhancement. ERN employs a dual global pathway structure that incorporates nonlocal operations to catch long-distance spatial features from the the original low-resolution (LR) input. To obtain the feature representation of the input at different scales, we further introduce a multiscale block (MSB) to directly detect low-level features from the LR image. The experimental results on four benchmark datasets have demonstrated that our models outperform most of the advanced methods while still retaining a reasonable number of parameters.
Rushi Lan, Zhenbing Liu, Huimin Lu 0001, Zhixun Su
IEEE Trans. Cybern.1
2021 A Two-Phase Learning-Based Swarm Optimizer for Large-Scale Optimization
abstract
In this article, a simple yet effective method, called a two-phase learning-based swarm optimizer (TPLSO), is proposed for large-scale optimization. Inspired by the cooperative learning behavior in human society, mass learning and elite learning are involved in TPLSO. In the mass learning phase, TPLSO randomly selects three particles to form a study group and then adopts a competitive mechanism to update the members of the study group. Then, we sort all of the particles in the swarm and pick out the elite particles that have better fitness values. In the elite learning phase, the elite particles learn from each other to further search for more promising areas. The theoretical analysis of TPLSO exploration and exploitation abilities is performed and compared with several popular particle swarm optimizers. Comparative experiments on two widely used large-scale benchmark datasets demonstrate that the proposed TPLSO achieves better performance on diverse large-scale problems than several state-of-the-art algorithms.
Rushi Lan, Yu Zhu 0004, Huimin Lu 0001, Zhenbing Liu
IEEE Trans. Cybern.1
2021 Chinese Emotional Dialogue Response Generation via Reinforcement Learning
abstract
In an open-domain dialogue system, recognition and expression of emotions are the key factors for success. Most of the existing research related to Chinese dialogue systems aims at improving the quality of content but ignores the expression of human emotions. In this article, we propose a Chinese emotional dialogue response generation algorithm based on reinforcement learning that can generate responses not only according to content but also according to emotion. In the proposed method, a multi-emotion classification model is first used to add emotion labels to the corpus of post-response pairs. Then, with the help of reinforcement learning, the reward function is constructed based on two aspects, namely, emotion and content. Among the generated candidates, the system selects the one with long-term success as the best reply. At the same time, to avoid safe responses and diversify dialogue, a diversity beam search algorithm is applied in the decoding process. The comparative experiments demonstrate that the proposed model achieves satisfactory results according to both automatic and human evaluations.
Rushi Lan, Wenming Huang, Zhenrong Deng, Xiyan Sun
ACM Trans. Internet Techn.1
2021 Chinese Image Captioning via Fuzzy Attention-based DenseNet-BiLSTM
abstract
Chinese image description generation tasks usually have some challenges, such as single-feature extraction, lack of global information, and lack of detailed description of the image content. To address these limitations, we propose a fuzzy attention-based DenseNet-BiLSTM Chinese image captioning method in this article. In the proposed method, we first improve the densely connected network to extract features of the image at different scales and to enhance the model’s ability to capture the weak features. At the same time, a bidirectional LSTM is used as the decoder to enhance the use of context information. The introduction of an improved fuzzy attention mechanism effectively improves the problem of correspondence between image features and contextual information. We conduct experiments on the AI Challenger dataset to evaluate the performance of the model. The results show that compared with other models, our proposed model achieves higher scores in objective quantitative evaluation indicators, including BLEU , BLEU , METEOR, ROUGEl, and CIDEr. The generated description sentence can accurately express the image content.
Huimin Lu 0001, Rui Yang 0018, Zhenrong Deng, Yonglin Zhang, Guangwei Gao, Rushi Lan
ACM Trans. Multim. Comput. Commun. Appl.6
2021 Discriminative Face Hallucination via Locality-Constrained and Category Embedding Representation
abstract
Recent years have witnessed the rapid development of face image hallucination techniques. However, the previous face hallucination methods are unsupervised and ignore the label information of training samples, leading to undesirable results. This article proposes a locality-constrained and category embedding representation (LCER) method to super-resolve face image in a supervised manner by embedding the label information in data representation. The proposed LCER incorporates the locality prior and category information into one unified framework, which aims to learn both the advantages of locality in preserving the true typologic structure of data manifold and the discriminability in exposing the class subspace information. Such strategy allows the LCER not only to preserve more sharpen image details but also to guarantee the face structure pattern be transferred mainly from the same subject in super-resolution reconstruction. Extensive experiments were conducted to evaluate the proposed LCER, and the comparative results demonstrate that it achieved superior face hallucination performance in both the quantitative measurements and visual impressions compared to several state-of-the-art.
Licheng Liu, Rushi Lan, Yaonan Wang 0001
IEEE Trans. Syst. Man Cybern. Syst.2
2020 An LBP encoding scheme jointly using quaternionic representation and angular information
Rushi Lan, Huimin Lu 0001, Yicong Zhou, Zhenbing Liu
Neural Comput. Appl.1
2020 Image captioning using DenseNet network and adaptive attention
Zhenrong Deng, Zhouqin Jiang, Rushi Lan, Wenming Huang
Signal Process. Image Commun.3
2020 Prior Knowledge-Based Probabilistic Collaborative Representation for Visual Recognition
abstract
Collaborative representation is an effective way to design classifiers for many practical applications. In this paper, we propose a novel classifier, called the prior knowledge-based probabilistic collaborative representation-based classifier (PKPCRC), for visual recognition. Compared with existing classifiers which use the collaborative representation strategy, the proposed PKPCRC further includes characteristics of training samples of each class as prior knowledge. Four types of prior knowledge are developed from the perspectives of image distance and representation capacity. They adaptively accommodate the contribution of each class and result in an accurate representation to classify a query sample. Experiments and comparisons on four challenging databases demonstrate that PKPCRC outperforms several state-of-the-art classifiers.
Rushi Lan, Yicong Zhou, Zhenbing Liu
IEEE Trans. Cybern.1
2020 Emotional Dialogue Generation Based on Conditional Variational Autoencoder and Dual Emotion Framework
abstract
An excellent dialogue system needs to not only generate rich and diverse logical responses but also meet the needs of users for emotional communication. However, despite much work, these two problems have not been solved. In this paper, we propose a model based on conditional variational autoencoder and dual emotion framework (CVAE-DE) to generate emotional responses. In our model, latent variables of the conditional variational autoencoder are adopted to promote the diversity of conversation. A dual emotion framework is adopted to control the explicit emotion of the response and prevent the conversation from generating emotion drift indicating that the emotion of the response is not related to the input sentence. A multiclass emotion classifier based on the Bidirectional Encoder Representations from Transformers (BERT) model is employed to obtain emotion labels, which promotes the accuracy of emotion recognition and emotion expression. A large number of experiments show that our model not only generates rich and diverse responses but also is emotionally coherent and controllable.
Zhenrong Deng, Hongquan Lin, Wenming Huang, Rushi Lan
Wirel. Commun. Mob. Comput.4
2020 A Novel Ray-Casting Algorithm Using Dynamic Adaptive Sampling
abstract
Ray-casting algorithm is an important volume rendering algorithm, which is widely used in medical image processing. Aiming to address the shortcomings of the current ray-casting algorithms in 3D reconstruction of medical images, such as slow rendering speed and low sampling efficiency, an improved algorithm based on dynamic adaptive sampling is proposed. By using the central difference gradient method, the corresponding sampling interval is obtained dynamically according to the different sampling points. Meanwhile, a new rendering operator is proposed based on the color value and opacity changes before and after the ray enters the volume element, and the resistance luminosity. Compared with the state of other algorithms, experimental results show that the method proposed in this paper has a faster rendering speed while ensuring the quality of the generated image.
Huadeng Wang, Xipeng Pan, Zhenbing Liu, Rushi Lan
Wirel. Commun. Mob. Comput.5
2020 SK-FMYOLOV3: A Novel Detection Method for Urine Test Strips
abstract
To accurately detect small defects in urine test strips, the SK-FMYOLOV3 defect detection algorithm is proposed. First, the prediction box clustering algorithm of YOLOV3 is improved. The fuzzy C-means clustering algorithm is used to generate the initial clustering centers, and then, the clustering center is passed to the K-means algorithm to cluster the prediction boxes. To better detect smaller defects, the YOLOV3 feature map fusion is increased from the original three-scale prediction to a four-scale prediction. At the same time, 23 convolutional layers of size 3 × 3 in the YOLOV3 network are replaced with SkNet structures, so that different feature maps can independently select different convolution kernels for training, improving the accuracy of defect classification. We collected and enhanced urine test strip images in industrial production and labeled the small defects in the images. A total of 11634 image sets were used for training and testing. The experimental results show that the algorithm can obtain an anchor frame with an average cross ratio of 86.57, while the accuracy rate and recall rate of nonconforming products are 96.8 and 94.5, respectively. The algorithm can also accurately identify the category of defects in nonconforming products.
Rui Yang 0018, Yonglin Zhang, Zhenrong Deng, Wenming Huang, Rushi Lan
Wirel. Commun. Mob. Comput.5
2019 EDCNN: A Novel Network for Image Denoising
abstract
In recent years, deep convolutional neural network (DCNN) has achieved impressive performance in image denoising. However, the existing CNN-based methods cannot work very well on those images with high-level noise. In order to solve this problem, we propose a novel method, named enhanced deep convolution neural network (EDCNN), for image de-noising in this work. Compared with existing models, ED-CNN adopts the residual learning in both global and local manners. In particular, we further apply a residual excitation strategy that enables a short path to be built directly from the input image to output layer. The final model, composed of 52 weight layers, is much deeper than existing ones. Experimental results on standard test images have demonstrated that the proposed method outperforms several state-of-the-art de-noising algorithms in terms of both quantitative measure and visual perception quality.
Haizhang Zou, Rushi Lan, Yanru Zhong, Zhenbing Liu
ICIP2
2018 A New Two-Phase Classifier for Face Recognition
abstract
Sparse representation plays an important role in many applications, such as face recognition and image denoising. The sparse representation based on regularization is proposed by many researchers. And it has not only very high computational efficiency but also has a simple solution. In this paper, we propose the sparse representation method that bases on the regularization and tikhonov regularization. Face recognition is divided into two steps: First, we use the Sparse representation based representation to Select the top M of Suspected classes for each test samples. Then we use the sparse representation method that bases on the regularization and tikhonov regularization to train the M Suspected classes. At last, testing our method with test samples. The experiment's results demonstrate that the improved algorithm has a better performance in image classification.
Rushi Lan
SMC2
2018 A simple texture feature for retrieval of medical images
Rushi Lan, Si Zhong, Zhenbing Liu, Zhuo Shi
Multim. Tools Appl.1
2018 Cost-sensitive collaborative representation based classification via probability estimation with addressing the class imbalance
Zhenbing Liu, Chao Ma 0006, Chunyang Gao, Rushi Lan
Multim. Tools Appl.5
2018 Integrated chaotic systems for image encryption
Rushi Lan, Jinwen He, Shouhua Wang, Tianlong Gu
Signal Process.1
2017 An extended probabilistic collaborative representation based classifier for image classification
abstract
Collaborative representation based classifier (CRC) and its probabilistic improvement ProCRC have achieved satisfactory performance in many image classification applications. They, however, do not comprehensively take account of the structure characteristics of the training samples. In this paper, we present an extended probabilistic collaborative representation based classifier (EProCRC) for image classification. Compared with CRC and ProCRC, the proposed EProCRC further considers a prior information that describes the distribution of each class in the training data. This prior information enlarges the margin between different classes to enhance the discriminative capacity of EProCRC. Experiments on two challenging databases, namely CUB200-2011 and Caltech-256, are conducted to evaluate EProCRC, and comparison results demonstrate that it outperforms several state-of-the-art classifiers.
Rushi Lan, Yicong Zhou
ICME1
2017 An Integrated Chaotic System with Application to Image Encryption
Jinwen He, Rushi Lan, Shouhua Wang
ICONIP (5)2
2017 Quaternionic Weber Local Descriptor of Color Images
abstract
This paper proposes a simple but effective framework named quaternionic Weber local descriptor (QWLD) for color image feature extraction. Integrating quaternionic representation (QR) of the color image and Weber's law (WL), QWLD possesses both their superiorities. It uses QR to handle all color channels of the image in a holistic way while preserving their relations, and applies WL to ensure that the derived descriptors are robust and discriminative. Using the QWLD framework, we further develop the quaternionic-increment-based Weber descriptor and quaternionic-distance-based Weber descriptor in terms of different perspectives. Extensive experiments on different color image recognition problems demonstrate that the proposed framework and descriptors outperform state-of-the-art local descriptors.
Rushi Lan, Yicong Zhou, Yuan Yan Tang
IEEE Trans. Circuits Syst. Video Technol.1
2017 Medical Image Retrieval via Histogram of Compressed Scattering Coefficients
abstract
The features used in many current medical image retrieval systems are usually low-level hand-crafted features. This limitation may adversely affect the retrieval performance. To address this problem, this paper proposes a simple yet discriminative feature, called histogram of compressed scattering coefficients (HCSC), for medical image retrieval. In the proposed work, the scattering transform, a particular variation of deep convolutional networks, is first performed to yield more abstract representations of a medical image. A projection operation is then conducted to compress the obtained scattering coefficients for efficient processing. Finally, a bag-of-words (BoW) histogram is derived from the compressed scattering coefficients as the features of the medical image. The proposed HCSC takes the advantages of both scattering transform and BoW model. Experiments on three benchmark medical computer tomography image databases demonstrate that HCSC outperforms several state-of-the-art features.
Rushi Lan, Yicong Zhou
IEEE J. Biomed. Health Informatics1
2016 Pairwise Linear Regression Classification for Image Set Retrieval
abstract
This paper proposes the pairwise linear regression classification (PLRC) for image set retrieval. In PLRC, we first define a new concept of the unrelated subspace and introduce two strategies to constitute the unrelated subspace. In order to increase the information of maximizing the query set and the unrelated image set, we introduce a combination metric for two new classifiers based on two constitution strategies of the unrelated subspace. Extensive experiments on six well-known databases prove that the performance of PLRC is better than that of DLRC and several state-of-theart classifiers for different vision recognition tasks: clusterbased face recognition, video-based face recognition, object recognition and action recognition.
Qingxiang Feng, Yicong Zhou, Rushi Lan
CVPR3
2016 Quaternion decomposition based discriminant analysis for color face recognition
abstract
In this paper, we propose a novel quaternion decomposition based discriminant analysis (QDDA) method for color face recognition. Unlike traditional approaches that handle color face images by vector representation or by each color channel individually, QDDA makes use of the quaternion to encode all color channels such that we can process all these channels in a holistic way and consider their relations simultaneously. In order to extract more discriminant color information from the image, a decomposition operation is performed to the quaternion matrix. A linear discriminant analysis is finally implemented to the obtained subcomponents for feature extraction. Experimental results have demonstrated the effectiveness of QDDA by comparing with other quaternion based methods.
Rushi Lan, Yicong Zhou
SMC1
2016 Quaternion-Michelson Descriptor for Color Image Classification
abstract
In this paper, we develop a simple yet powerful framework called quaternion-Michelson descriptor (QMD) to extract local features for color image classification. Unlike traditional local descriptors extracted directly from the original (raw) image space, QMD is derived from the Michelson contrast law and the quaternionic representation (QR) of color images. The Michelson contrast is a stable measurement of image contents from the viewpoint of human perception, while QR is able to handle all the color information of the image holisticly and to preserve the interactions among different color channels. In this way, QMD integrates both the merits of Michelson contrast and QR. Based on the QMD framework, we further propose two novel quaternionic Michelson contrast binary pattern descriptors from different perspectives. Experiments and comparisons on different color image classification databases demonstrate that the proposed framework and descriptors outperform several state-of-the-art methods.
Rushi Lan, Yicong Zhou
IEEE Trans. Image Process.1
2016 Quaternionic Local Ranking Binary Pattern: A Local Descriptor of Color Images
abstract
This paper proposes a local descriptor called quaternionic local ranking binary pattern (QLRBP) for color images. Different from traditional descriptors that are extracted from each color channel separately or from vector representations, QLRBP works on the quaternionic representation (QR) of the color image that encodes a color pixel using a quaternion. QLRBP is able to handle all color channels directly in the quaternionic domain and include their relations simultaneously. Applying a Clifford translation to QR of the color image, QLRBP uses a reference quaternion to rank QRs of two color pixels, and performs a local binary coding on the phase of the transformed result to generate local descriptors of the color image. Experiments demonstrate that the QLRBP outperforms several state-of-the-art methods.
Rushi Lan, Yicong Zhou, Yuan Yan Tang
IEEE Trans. Image Process.1
2016 Generalization Performance of Regularized Ranking With Multiscale Kernels
abstract
The regularized kernel method for the ranking problem has attracted increasing attentions in machine learning. The previous regularized ranking algorithms are usually based on reproducing kernel Hilbert spaces with a single kernel. In this paper, we go beyond this framework by investigating the generalization performance of the regularized ranking with multiscale kernels. A novel ranking algorithm with multiscale kernels is proposed and its representer theorem is proved. We establish the upper bound of the generalization error in terms of the complexity of hypothesis spaces. It shows that the multiscale ranking algorithm can achieve satisfactory learning rates under mild conditions. Experiments demonstrate the effectiveness of the proposed method for drug discovery and recommendation tasks.
Yicong Zhou, Hong Chen 0004, Rushi Lan, Zhibin Pan
IEEE Trans. Neural Networks Learn. Syst.3
2014 Person reidentification using quaternionic local binary pattern
abstract
Person reidentification is to identify the persons observed in nonoverlapping camera networks. Most existing methods usually extract features from the red, green, and blue color channels of images individually. They, however, neglect the connections between each color component in the image. To overcome this problem, a novel quaternionic local binary pattern (QLBP) is proposed for person reidentification in this paper. In the proposed QLBP, each pixel in a color image is represented by a quaternion so that we can handle all color components in a holistic way. A novel pseudo-rotation of quaternion (PRQ) is proposed to rank two quaternions. Some properties of PRQ are also discussed. After a QLBP coding, the local histograms are extracted and used as features. Experiments on two public benchmarking datasets, ETHZ and i-LIDS MCTS, are carried out to evaluate the QLBP performance. Comparison results show that the QLBP outperforms several stat-of-art methods for person reidentification.
Rushi Lan, Yicong Zhou, Yuan Yan Tang, C. L. Philip Chen
ICME1
2013 Whitening central projection descriptor for affine-invariant shape description
abstract
A novel descriptor, referred to as the whitening central projection predictor (WCPD), is developed for affine‐invariant shape description. The proposed descriptor is based on central projection transform (CPT) and whitening transform (WT). Dislike contour‐based or region‐based approaches, an object is first converted to a closed curve by CPT, which is called the general curve (GC). The derived GC not only keeps the affine transform information, but also is very robust to noise. Then WT is performed to the GC with the purpose that the affine transformation is simplified to a rotation only. Finally, Fourier descriptors are employed to remove the rotation, and WCPD is obtained. One advantage of using WCPD for affine‐invariant description lies in that it is applicable to objects consisting of several components. Furthermore, the approach used on the GC is contour‐based, and is of small computational complexity. Several experiments have been conducted to evaluate the performance of the proposed method. Experimental results show that the proposed method has a powerful discrimination ability, and is more robust to noise.
Rushi Lan, Colin Fyfe, Zhan Song
IET Image Process.1
2012 An affine invariant discriminate analysis with canonical correlation analysis
Rushi Lan, Zhan Song, Yuan Yan Tang
Neurocomputing1
2010 Orthogonal projection transform with application to shape description
abstract
In this paper, an approach called orthogonal projection transform (OPT) is proposed for invariant shape description. It is inspired by the construction of orthogonal Fourier-Mellin moments (OFMMs), but integral is only performed along lines with different polar angles in the proposed approach. By performing OPT, any object can be converted to a set of closed curves. In comparison with moment based methods, such as OFMMs, complex moments (CMs) etc., these projected curves not only preserve the information along lines with different polar angles, but also derive the information of the global intensity distributions. Descriptors, which are invariant to translation, rotation and scaling, are constructed from these closed curves. Applying the descriptors to images taken from standard image datasets, the numerical experiments show that the proposed method is more robust than some moment based methods such as OFMMs and CMs to noise, occlusions and real view angle disturbances.
Rushi Lan
ICIP1