Zhe Sun 0009

dblp:43/8664-9 · DBLP profile ↗
← Back
39ranked-venue papers
3as first author
37since 2021 · last 2027
0000-0002-6531-0769ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 17 · 15 since 2021Graphics, computer vision, multimedia, augmented reality and games · 17 · 17 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 3 first-author · 7 since 2021Systems, architecture and hardware · 3 · 3 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Computer networks · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2027 Corrigendum to "Short-time variational mode decomposition" [Signal Processing 238 (2026) 110203]
Tong Liang, Cesar F. Caiafa, Zhe Sun 0009, Yasuhiro Kushihashi, Antoni Grau-Saldes, Yolanda Bolea, Feng Duan 0006, Jordi Solé i Casals
Signal Process.5
2026 Adaptive knowledge selection in dialogue systems: Accommodating diverse knowledge types, requirements, and generation models
Zhongtian Bao, Hongru Liang, Jun Wang 0023, Zhenglu Yang, Zhe Sun 0009, Andrzej Cichocki
Neural Networks7
2026 A2VAD: Attribute-augmented prompt learning for weakly supervised video anomaly detection
Zheng Wang 0044, Xing Xu 0001, Jingkuan Song, Zhe Sun 0009, Andrzej Cichocki
Pattern Recognit.5
2026 Short-time variational mode decomposition
Tong Liang, Cesar F. Caiafa, Zhe Sun 0009, Yasuhiro Kushihashi, Antoni Grau-Saldes, Yolanda Bolea, Feng Duan 0006, Jordi Solé i Casals
Signal Process.5
2026 A Non-Torque Sensing Predefined-Time Sliding Mode Adaptive Admittance Control Approach for Lower Limb Exoskeleton Robots
Zhe Sun 0009, Guodong Wu, Xiaohua Ge, Jinchuan Zheng, Zhihong Man
IEEE Trans. Ind. Informatics1
2026 Egocentric Online Action Segmentation via Parametric Context Memory Learning
abstract
To facilitate smart wearable devices or human-like robotics with real-time first-person perspective perception ability, recent researchers proposed the Egocentric Online Action Segmentation (EOAS) task. It requires models to recognize what is happening in egocentric streaming videos and discriminate the starting and ending times of an activity in a real-time manner. However, compared with offline-recorded exocentric videos, egocentric streaming videos cannot provide equivalent sufficient temporal-spatial cues due to the limited perspective and unknown coming frames. Hence, it raises a high demand for the long-term episodic memory ability of models. To this end, most previous approaches work on compressing long-term memory into feature representations. In this paper, we propose a novel EOAS paradigm, termed Parametric Context Memory Learning (PCML), which integrates episodic memory into learnable parameters and keeps dynamic updates according to real-time frames. Concretely, we design the Parametric Context Perception layer and construct a novel Episodic Semantic Memorization Network (ESMN) based on it, which integrates episodic memory into learnable parameters and keeps dynamic updates with real-time frames. We evaluate our proposed method on three public egocentric streaming video benchmarks including EgoPER, EgoProceL, and GTEA. Extensive experiments demonstrate the ESMN model significantly outperforms recent state-of-the-art methods. Our code is available at https://github.com/XunCHN/PCML.
Xun Jiang 0001, Xing Xu 0001, Zheng Wang 0044, Jingkuan Song, Zhe Sun 0009, Andrzej Cichocki, Heng Tao Shen
IEEE Trans. Image Process.6
2026 Collaborated With Hallucination: Enhancing Egocentric Grounded Question Answering via Error Demonstrations
abstract
The grounded question answering in egocentric videos (Ego-GQA) aims to identify the relevant temporal window and generate corresponding responses in natural language given a textual question. Compared with third-person videos, egocentric video understanding requires more advanced human-centric thinking capability. However, existing Ego-GQA approaches often fail to distinguish the inherent limitations of dynamic egocentric context understanding, treating both first-person and third-person perspectives equally. This oversight leads to hallucinations and a lack of proper egocentric reasoning in first-person video understanding. To address this issue, we propose a novel Collaborated with Hallucination (CoHa) framework for the Ego-GQA, which quantifies the hallucinations generated by an Ego-GQA model and further leverages them as error demonstrations to constrain the model's reasoning process, encouraging it to ground predictions in egocentric visual cues instead of relying on biased pretraining priors. Specifically, we first employ Subjective Logic to quantify the degree of uncertainty in unreliable answers. We then generate diffusion-based noisy visual inputs to amplify the hallucinations as error demonstrations, which are used to append appropriate constraints to the model according to the uncertainty. These constraints effectively steer predictions away from the unreliable semantics induced by inherent drawbacks in egocentric thinking. Additionally, we incorporate an interactive refinement module to facilitate the model to explore more fine-grained cues observed from the first-person view. Extensive experiments on two widely used benchmarks demonstrate that our CoHa method outperforms recent state-of-the-art methods. Our code is available at https://github.com/Mrshenshen/CoHa.
Shenshen Li, Xing Xu 0001, Fumin Shen, Zhe Sun 0009, Andrzej Cichocki, Heng Tao Shen
IEEE Trans. Image Process.4
2025 So Far Yet So Near: Time Series Data Augmentation with Exploring non-Semantic Boundaries based on Reinforcement Learning
abstract
Data augmentation effectively expands feature distribution in time series classification, enhancing downstream task performance. However, existing techniques often fail to maintain semantic consistency between augmented and original time series data, causing label noise and thereby degrading downstream task performance. We argue that data augmentation should preserve time series semantic consistency and expand the non-semantic information space. In this paper, we reformulate data augmentation as a semantic path planning problem between original data and augmented data, modeled as a Markov Decision Process (MDP). We propose a reinforcement learning-based algorithm (RL) named FreqSYN, where the action space is defined by a set of learnable Gaussian kernels that perturbs the frequency domain of the original data to generate augmented samples. The confidence coefficients of augmented data in semantically relevant classification tasks are used as a reward to iteratively refine the FreqSYN. Our method is validated across four datasets, achieving state-of-the-art performance, with a 2% improvement in F1 score over the SimPSI method. The code and models are available at https://github.com/NKU-EmbeddedSystem/FreqSYN.
Haoran Li 0014, Jiarong Kang, Xun Jiang 0001, Xiaoli Gong, Jin Zhang 0003, Zhe Sun 0009, Andrzej Cichocki
ICASSP7
2025 Essentia: Boosting Artifact Removal from EEG through Semantic Guidance Utilizing Diffusion Model
abstract
Electroencephalography (EEG) is a time-series signal containing semantic information that can be used to determine human brain activities. Artifacts within EEG data can interfere with the intrinsic distribution of this semantic information, so removing artifacts is crucial for improving EEG analysis performance on downstream tasks. In this paper, we redefine the efficacy of the artifact removal model by evaluating the performance of the noisy EEG data in downstream tasks before and after artifact removal. Currently, most artifact removal models fail to ensure semantic consistency, rendering them ineffective. To solve it, we propose an artifact removal model based on the 1-dimensional diffusion model utilizing the U-Net, referred to as Essentia. Moreover, we find that the skip-connection layer in U-Net contains mid-to-high-frequency information that interferes with the semantic representation. We introduce a semantic guidance module (SGM) that leverages contrastive learning to generate semantic distribution weights, boosting semantic representation. We evaluate Essentia on three datasets with six solutions. The accuracy of downstream tasks from the denoised EEG data increased by 4% compared with the DeepSeparetor. The code and models are available at https://github.com/NKU-EmbeddedSystem/Essentia.
Haoran Li 0014, Xiaoli Gong, Jin Zhang 0003, Tingjuan Lu, Zhe Sun 0009, Andrzej Cichocki
ICASSP8
2025 DSDIR: A Two-Stage Method for Addressing Noisy Long-Tailed Problems in Malicious Traffic Detection
abstract
In recent years, deep learning based malicious traffic detection (MTD) systems have demonstrated remarkable success. However, their effectiveness tend to decrease because most malicious traffic datasets are suffered from noisy-labeled and long-tailed problems. While numerous approaches have been developed to address these two problems individually, they become inefficient when confronting the combined challenge of noisy long-tailed data, as they typically tackle only a single adverse factor at a time. This paper proposes a two-stage method called Distribution-aware sample Selection and Dynamic Instance-based Relabeling (DSDIR), which simultaneously addresses the impacts of noisy-labeled and long-tailed problems. In the first stage, a noise-independent clean sample selection method is designed to obtain a clean dataset, which converts negative effects of the long-tailed problem into positive ones. In the second stage, dynamic instance-based relabeling is designed to train a model and improve the dataset’s quality simultaneously. Eventually, DSDIR not only produces a balanced and noise-tolerant model but also obtains a clean dataset. Experimental results demonstrate that in the high noise condition of 60% and 80%, the accuracy rate of DSDIR is 5% higher than the state-of-the-art methods. Our code is available at https://github.com/nku-ligl/DSDIR.
Zhe Sun 0009, Lingkai Xing, Yu Zhang 0009
ICASSP3
2025 Social Optimum Assisted Gradient Modulation for Imbalanced Multimodal Learning
abstract
The imbalanced modality problem in multimodal learning is a vicious phenomenon that leads to the sub-optimization of the modalities due to the gradient conflicts among the modalities and the model’s preference for the easier learning modalities. Recent studies have been dedicated to modulating the gradient from the uni-modal perspective, boosting the learning of the single modality. Nevertheless, they overlook the similarity between multimodal gradient optimization and multi-objective learning, while overemphasizing competition between modalities’ gradients to improve optimization. Therefore, we perceive the imbalanced multimodal optimization as Multi-Objective Optimization, and propose a novel training method: Social Optimum Assisted Gradient Modulation (SOA-GM). In detail, we implemented social Optimum to guide the multimodal model to achieve the trade-off state in which the collaborations of modalities can be maximized. We also proposed envy to reduce the model’s preference to a particular modality. Finally, experiments across multiple datasets indicate our superior and extendable method performance on multimodal learning.
Disen Hu, Xun Jiang 0001, Zhe Sun 0009, Hao Yang 0015, Xing Xu 0001
ICME3
2025 Real-Time Anomaly Detection and Completion in Data Streams Based on RRCF and RTHaLRTC Tensor Completion
Tong Liang, Binghua Li 0001, Ziqing Chang, Chao Li 0013, Jiahe Guo, Jordi Solé i Casals, Yasuhiro Kushihashi, Ryutaro Himeno, Zhe Sun 0009
ICONIP (3)11
2025 Parameter-Efficient Fine-Tuning of 3D DDPM for MRI Image Generation Using Tensor Networks
Binghua Li 0001, Ziqing Chang, Tong Liang, Chao Li 0013, Toshihisa Tanaka 0001, Shigeki Aoki, Qibin Zhao, Zhe Sun 0009
MICCAI (4)8
2025 Heterogeneous Graph Embedding for Multimodal Multi-Label Emotion Recognition
abstract
Multimodal Multi-label Emotion Recognition (MMER) aims to identify human emotions through various modalities. Previous studies mainly focus on aligning cross-modal data to extract discriminative emotion-dependent features using attention or reconstruction-based strategies, while omitting the fact that the MMER task is also subjected to the multi-label noises that exist in the multi-label classifications, which disturb the modality-to-label correlations. Besides, most of the research also failed to balance the strategy of finding internal label correlations and label dependency of modalities in noisy conditions. In this paper, we proposed a novel Heterogeneous Graph Embedding (HGE) method for the MMER task, which exploits heterogeneous graphs to extract emotional commonality in the modal and temporal levels, and explicitly models cross-modal correlations among heterogeneous modalities. Additionally, it also captures uncertainty brought by multi-label noise and leverages the unevenness of multi-label to overcome potential data issues. Experimental results demonstrate that our HGE method achieves state-of-the-art performance on two widely used multimodal multi-label emotion recognition datasets under both noise-free and noisy circumstances.
Disen Hu, Xun Jiang 0001, Zhe Sun 0009, Fumin Shen, Xing Xu 0001
ICMR3
2025 Geometric Gradient Divergence Modulation for Imbalanced Multimodal Learning
abstract
Multimodal learning, which has been given great significance recently, may face the challenge of the imbalanced multimodal phenomenon, which leads to the insufficient optimization of both multimodal and unimodal objectives. The core problem lies in the optimization conflicts between the above optimization objectives, resulting in the diverse updating directions and strengths that cause antagonism between them. In this paper, we mathematically analyze the optimization processes of imbalanced multimodal learning in the hyperspaces from a novel geometric perspective. Additionally, based on our theoretical analysis, we defined the volumes of the gradients constructed parallel polyhedron in the hyperspace to quantify the misalignment between the optimization objectives. Subsequently, we proposed the Geometric Gradient Divergence Modulation (GGDM), which leverages the volumes of gradient polyhedron to perform gradient modulation, encouraging alignment among gradients and promoting a synergistic optimization effect. Lastly, we evaluate our GGDM on five widely used multimodal benchmarks, where RGB image, optical flow, text, image, video and audio are involved. Our method achieved state-of-the-art performance compared to other imbalanced multimodal learning methods. Our code is available at: https://github.com/ConstantineWayne/GGDM.
Disen Hu, Xun Jiang 0001, Zhe Sun 0009, Hao Yang 0015, Heng Tao Shen, Xing Xu 0001
ACM Multimedia3
2025 AnoOnly: Semi-supervised anomaly detection with the only loss on anomalies
Yixuan Zhou 0001, Peiyu Yang, Xing Xu 0001, Zhe Sun 0009, Andrzej Cichocki
Expert Syst. Appl.5
2025 Non-Force-Sensing Variable Admittance Control of Lower Limb Rehabilitation Exoskeleton Robots Using Class κ∞ Function-Based Adaptive Sliding Mode
abstract
In this paper, a non-force-sensing variable admittance control approach is proposed for lower limb rehabilitation exoskeleton robots. This approach aims to provide satisfactory assistance and rehabilitation-training performance for users of such exoskeletons. Our method initiates with a novel fixed-time sliding mode observer that estimates human-machine interaction torques according to generalized momentum. This observer-based torque estimation strategy can estimate the interaction torque precisely, thereby enabling a non-force-sensing effect. Next, a variable admittance control framework is formulated for lower limb exoskeleton robots to ensure superior compliance. This framework comprises two essential components. First, a newly designed adaptive fixed-time sliding mode controller based on a class$\kappa _{\infty } $function for the inner loop, which operates without requiring prior knowledge of the upper bound of lumped perturbations and guarantees precise gait trajectory-tracking performance. Second, a variable-parameter admittance model for the outer loop, which utilizes an exponential function to dynamically adjust the admittance parameters, thereby achieving a balance between the exoskeleton’s compliance and gait-correction efficacy. Finally, both simulation and experimental results are presented and analyzed to validate the effectiveness and superiority of the proposed non-force-sensing variable admittance control approach. Specifically, simulation results demonstrate that the root-mean-square (RMS) values of the inner-loop tracking errors for the proposed method are reduced by 26.1% and 20.4% at the hip and knee joints, respectively, compared with the top-performing benchmark algorithm. Meanwhile, the precision of the outer-loop observation is improved by 21% and 18% at these joints. Experimental validation further shows reductions of 14.2% and 20.1% in the RMS inner-loop tracking errors at the hip and knee joints, respectively, versus this benchmark algorithm.Note to Practitioners—Motivated by the problem of how to realize effective control of lower limb exoskeleton robots to provide the wearers with appropriate comfort and gait-correction effect, this paper formulates an adaptive variable admittance control framework. In the inner loop of this framework, a novel gait trajectory-tracking controller using adaptive sliding mode is designed to ensure high tracking accuracy. In the outer loop, a new adaptive observer is designed to estimate the human-robot interaction torque, and an adaptive admittance model is formulated to provide appropriate compliance for the robot. The proposed strategies are experimentally validated and can be utilized in real-word applications. The work of this paper has reference significance for the development of lower limb exoskeleton technologies.
Zhe Sun 0009, Tianyu Chai, Bo Chen 0003, Hai Wang 0004, Jinchuan Zheng, Zhihong Man
IEEE Trans Autom. Sci. Eng.1
2025 VQ-Flow: Taming Normalizing Flows for Multi-Class Anomaly Detection via Hierarchical Vector Quantization
abstract
Normalizing flows, a category of probabilistic models famed for their capabilities in modeling complex data distributions, have exhibited remarkable efficacy in unsupervised anomaly detection. This paper explores the potential of normalizing flows in multi-class anomaly detection, wherein the normal data is compounded with multiple classes without providing class labels. Through the integration of vector quantization (VQ), we empower the flow models to distinguish different concepts of multi-class normal data in an unsupervised manner, resulting in a novel flow-based unified method, named VQ-Flow. Specifically, our VQ-Flow leverages hierarchical vector quantization to estimate two relative codebooks: a Conceptual Prototype Codebook (CPC) for concept distinction and its concomitant Concept-Specific Pattern Codebook (CSPC) to capture concept-specific normal patterns. The flow models in VQ-Flow are conditioned on the concept-specific patterns captured in CSPC, capable of modeling specific normal patterns associated with different concepts. Moreover, CPC further enables our VQ-Flow for concept-aware distribution modeling, faithfully mimicking the intricate multi-class normal distribution through a mixed Gaussian distribution reparametrized on the conceptual prototypes. Through the introduction of vector quantization, the proposed VQ-Flow advances the state-of-the-art in multi-class anomaly detection within a unified training scheme, yielding the Det./Loc. AUROC of 99.5%/98.3% on MVTec AD.
Yixuan Zhou 0001, Xing Xu 0001, Zhe Sun 0009, Jingkuan Song, Andrzej Cichocki, Heng Tao Shen
IEEE Trans. Multim.3
2025 Resisting Noise in Pseudo Labels: Audible Video Event Parsing With Evidential Learning
abstract
Perceiving temporal events and discriminating their modality types in audible videos, which is also called audio-visual video parsing (AVVP), is becoming a research hotspot in multimodal video understanding. The AVVP task generally follows weakly supervised learning settings, since only video-level labels are provided. Most existing works usually generate modalitywise pseudo labels (PLs) first and then learn to parse audio or visual events from the audible videos. However, this paradigm inevitably results in two defects: 1) the generated PLs for each modality are not fully reliable, which may confuse models if they are adopted as supervision signals for discriminating modalities; and 2) the absence of temporal annotations increases the ambiguities in localizing foregrounds in videos, furtherly causing models prone to being disturbed by noisy labels. To tackle these problems, we propose a novel AVVP framework termed noise-resistant event parsing (NREP), which introduces evidential deep learning (EDL) to overcome the limitations of noisy pseudo supervision. Specifically, our NREP framework consists of three key components: 1) modalitywise evidential learning (MEL) that discriminates the modality-class dependency; 2) temporalwise evidential learning (TEL) that explores meaningful foregrounds; and 3) foreground-background consistency learning (FBCL) for collaborating two evidential learning branches above. Through perceiving meaningful video content and learning evidence for modality dependencies, our method suppresses the disturbance of noise in generated PLs thus achieving remarkable performance with different PL generation strategies. We evaluate our NREP method on two AVVP benchmark datasets and demonstrate it consistently to establish new state-of-the-art. Our implementation codes are available at https://github.com/CFM-MSG/NREP.
Xun Jiang 0001, Xing Xu 0001, Liqing Zhu, Zhe Sun 0009, Andrzej Cichocki, Heng Tao Shen
IEEE Trans. Neural Networks Learn. Syst.4
2024 Cross View Capture for Distributed Image Compression with Decoder Side Information
abstract
Image compression is increasingly important in applications like intelligent driving and smart surveillance systems. This study presents a novel cross view capture distributed image compression network (CVCDIC) to improve the compression quality by using decoder side information. The CVCDIC’s decoder utilizes feature extraction networks to extract features from both the primary image and the side information. Furthermore, a multi-level cross view attention module is designed to capture interrelated details between images at multiple hierarchical levels. Finally, a spatial refinement module, constructed on the foundation of information distillation networks, is designed to further refine the quality of reconstructed images. The results show that CVCDIC can achieve an MS-SSIM of 0.978 at 0.15 bpp, surpassing DSIN (0.925), NDIC (0.956), and ATN (0.955) on the KITTI Stereo dataset.
Yankai Yin, Zhe Sun 0009, Peiying Ruan, Feng Duan 0006, Ruidong Li 0001, Chi Zhu 0001
ICRA2
2024 Online Hand Movement Recognition System with EEG-EMG Fusion Using One-Dimensional Convolutional Neural Network
abstract
Upper limb amputees face significant challenges in their daily lives due to the loss of hand or arm functionality. Researchers have developed upper limb prostheses to restore normal hand movements for them. Most hand movement recognition systems of prostheses use electromyography (EMG) as the input signal source, but ignore the interrelationship with electroencephalography (EEG), which may contain valuable movement-related information as well. In order to enhance the accuracy of hand movement classification, we proposed a hand movement recognition system based on a one-dimensional convolutional neural network (1D-CNN) that combines EEG and EMG as the input signal sources to increase the quantity of accessible information. In this work, we collected the EEG and EMG of five subjects during the hand movements and used a 1D-CNN based model to classify the preprocessed signals. The average accuracy of using EEG-EMG fusion is 96.59±2.63%, significantly higher than 74.99±8.24% of using single EEG and 90.31±7.16% of using single EMG. Then, we applied the model trained by offline experiment for online recognition, and controlled the Pepper robot to complete the corresponding hand movements. The average accuracy of online recognition can reach 93.00±4.85% by using majority voting method. The results indicate that the method of EEG-EMG fusion can effectively enhance the performance of hand movement recognition system, which promote the development of upper limb prostheses and contribute to the rehabilitation of upper limb amputees.
Haozheng Wang, Zhe Sun 0009, Feng Duan 0006
IROS3
2024 Enabling temporal-spectral decoding in multi-class single-side upper limb classification
abstract
This manuscript presents a novel approach for decoding pre-movement patterns from brain signals using a two-stage-training temporal–spectral neural network (TTSNet). The TTSNet employs a combination of filter bank task-related component analysis (FBTRCA) and convolutional neural network (CNN) techniques to enhance the classification of single-upper limb movements in non-invasive brain–computer interfaces (BCIs). In our previous work, we introduced the FBTRCA method which utilized filter banks and spatial filters to handle spectral and spatial information, respectively. However, we observed limitations in the temporal decoding phase, where correlation features failed to effectively utilize temporal information because of misaligned onset and noisy spikes. To address this issue, our proposed method focuses on analyzing multi-channel signals in the temporal–spectral domain. The TTSNet first divides the signals into various filter banks, employing task-related component analysis to reduce dimensionality and eliminate noise, respectively. Subsequently, a CNN is employed to optimize the temporal characteristics of the signals and extract class-related features. Finally, the class-related features from all filter banks are concatenated and classified using the fully connected layer. To evaluate the effectiveness of our proposed method, we conducted experiments on two publicly available datasets. In binary classification tasks, the TTSNet achieved an improved accuracy of 0.7707 ± 0.1168, surpassing the performance of EEGNet (accuracy: 0.7340 ± 0.1246) and FBTRCA (accuracy: 0.7487 ± 0.1250). In multi-class tasks, TTSNet achieved an accuracy of 0.4588 ± 0.0724, exhibiting a 4.27% and 3.95% accuracy increase over EEGNet and FBTRCA, respectively. The findings of this study suggest that the proposed TTSNet method holds promise for detecting limb movements and assisting in the rehabilitation of stroke patients. The classification of single-side limb movements is expected to facilitate the interaction between patients and external environment by increasing the number of control commands in BCIs.
Shuning Han, Cesar F. Caiafa, Feng Duan 0006, Yu Zhang 0009, Zhe Sun 0009, Jordi Solé i Casals
Eng. Appl. Artif. Intell.6
2024 Adaptive fuzzy sliding mode control of uncertain nonholonomic wheeled mobile robot with external disturbance and actuator saturation
Yunjun Zheng, Jinchuan Zheng, Han Zhao 0007, Zhihong Man, Zhe Sun 0009
Inf. Sci.6
2024 Cross-Modal Attention Preservation with Self-Contrastive Learning for Composed Query-Based Image Retrieval
abstract
In this article, we study the challenging cross-modal image retrieval task,Composed Query-Based Image Retrieval (CQBIR), in which the query is not a single text query but a composed query, i.e., a reference image, and a modification text. Compared with the conventional cross-modal image-text retrieval task, the CQBIR is more challenging as it requires properly preserving and modifying the specific image region according to the multi-level semantic information learned from the multi-modal query. Most recent works focus on extracting preserved and modified information and compositing it into a unified representation. However, we observe that the preserved regions learned by the existing methods contain redundant modified information, inevitably degrading the overall retrieval performance. To this end, we propose a novel method termedCross-ModalAttentionPreservation (CMAP). Specifically, we first leverage the cross-level interaction to fully account for multi-granular semantic information, which aims to supplement the high-level semantics for effective image retrieval. Furthermore, different from conventional contrastive learning, our method introduces self-contrastive learning into learning preserved information, to prevent the model from confusing the attention for the preserved part with the modified part. Extensive experiments on three widely used CQBIR datasets, i.e., FashionIQ, Shoes, and Fashion200k, demonstrate that our proposed CMAP method significantly outperforms the current state-of-the-art methods on all the datasets. The anonymous implementation code of our CMAP method is available at https://github.com/CFM-MSG/Code_CMAP.
Shenshen Li, Xing Xu 0001, Xun Jiang 0001, Fumin Shen, Zhe Sun 0009, Andrzej Cichocki
ACM Trans. Multim. Comput. Commun. Appl.5
2023 TMOVF: A Task-Agnostic Model Ownership Verification Framework
abstract
The protection of model intellectual property is becoming an increasingly important issue. However, the existing methods for protecting model ownership, although effective, have limitations. Firstly, they primarily focus on classification models, and secondly, most of the proposed methods reduce the model's utility. To overcome these shortcomings, this paper proposes a task-agnostic model ownership verification framework based on feature fingerprint, called TMOVF, which separates ownership verification from model task. Our key idea is that model knowledge can be uniquely characterized by the extracted features, which may be high-dimensional, complicated, and difficult to compare for each input sample. Nevertheless, these features contain inherent information that cannot be ignored in cases of piracy. To measure the inheritance of our fingerprint, we introduce outlier detection into model ownership verification, which is a first in the field. By reconstructing the outlier detection algorithm, we extract the feature fingerprints of the victim model and the suspicious model, and compute the outliers of their feature fingerprints. By comparing the results, we can verify the ownership of the models. We conduct extensive experiments to evaluate our framework and demonstrate the inheritability of feature fingerprints in stolen models. Our experiments show that the framework is effective in verifying ownership, regardless of the model task. Additionally, our results demonstrate that our framework is more effective than existing methods.
Zhe Sun 0009, Zhongyu Huang, Wangqi Zhao, Yu Zhang 0009
SMC1
2023 Underwater sEMG-based recognition of hand gestures using tensor decomposition
Jianing Xue, Zhe Sun 0009, Feng Duan 0006, Cesar F. Caiafa, Jordi Solé i Casals
Pattern Recognit. Lett.2
2023 Privacy-Preserving Multi-Source Domain Adaptation for Medical Data
abstract
Great progress has been made in diagnosing medical diseases based on deep learning. Large-scale medical data are expected to improve deep learning performance further. It is almost impossible for a single institution to collect so much data due to the time-consuming and costly collection and labeling of medical data. Many studies have turned attention to data sharing among multiple medical institutions. However, due to different data acquiring and processing procedures, multiple institutions' medical data is characterized by distribution heterogeneity. Besides, the protection of patient privacy in medical data sharing has also been a common concern. To simultaneously address the problems of heterogeneous data distribution and privacy protection, we propose a novel multi-source source free domain adaptation. When aligning distributed heterogeneous data, our method only require to transfer the pre-trained source models rather than the direct source domain data, thus protecting patients' privacy. In addition, it has the advantages of being efficient and less costly in network resources. The proposed method is evaluated on the multi-site fMRI database Autism Brain Imaging Data Exchange (ABIDE) and yields an average accuracy of 69.37%. We also analyzed its effectiveness on network resource-saving and conducted additional experiments on Camelyon17 to validate the generalization.
Xiaoli Gong, Jin Zhang 0003, Zhe Sun 0009, Yu Zhang 0009
IEEE J. Biomed. Health Informatics5
2023 Multi-Class Classification of Upper Limb Movements With Filter Bank Task-Related Component Analysis
abstract
The classification of limb movements can provide with control commands in non-invasive brain-computer interface. Previous studies on the classification of limb movements have focused on the classification of left/right limbs; however, the classification of different types of upper limb movements has often been ignored despite that it provides more active-evoked control commands in the brain-computer interface. Nevertheless, few machine learning method can be used as the state-of-the-art method in the multi-class classification of limb movements. This work focuses on the multi-class classification of upper limb movements and proposes the multi-class filter bank task-related component analysis (mFBTRCA) method, which consists of three steps: spatial filtering, similarity measuring and filter bank selection. The spatial filter, namely the task-related component analysis, is first used to remove noise from EEG signals. The canonical correlation measures the similarity of the spatial-filtered signals and is used for feature extraction. The correlation features are extracted from multiple low-frequency filter banks. The minimum-redundancy maximum-relevance selects the essential features from all the correlation features, and finally, the support vector machine is used to classify the selected features. The proposed method compared against previously used models is evaluated using two datasets. mFBTRCA achieved a classification accuracy of 0.4193 ± 0.0780 (7 classes) and 0.4032 ± 0.0714 (5 classes), respectively, which improves on the best accuracies achieved using the compared methods (0.3590 ± 0.0645 and 0.3159 ± 0.0736, respectively). The proposed method is expected to provide more control commands in the applications of non-invasive brain-computer interfaces.
Cesar F. Caiafa, Feng Duan 0006, Yu Zhang 0009, Zhe Sun 0009, Jordi Solé i Casals
IEEE J. Biomed. Health Informatics6
2022 Preliminary Results on the Generation of Artificial Handwriting Data Using a Decomposition-Recombination Strategy
abstract
Deep learning techniques are able to extract the characteristics of temporal signals to study their patterns and diagnose diseases such as essential tremor. However, these techniques require a large amount of data to train the neural network and achieve good results, and the more data the network has, the more accurate the final model implemented will be. This work proposes the use of data augmentation techniques to improve the accuracy of a Long short-term memory system in the diagnosis of essential tremor. For this purpose, the Empirical Modal Decomposition method will be used to decompose the original temporal signals collected from control subjects and patients with essential tremor. The time series obtained from the decomposition, covering different frequency ranges, will be randomly shuffled and combined to generate new artificial samples for each group. Then, both the generated artificial samples and part of the real samples will be used to train the LSTM network, and the remaining original samples will be used to test the model. Experimental results demonstrate the capability of the proposed method, increasing the classifier accuracy from 83.2% to almost 93% when artificial samples are used.
José Fernando Adrán Otero, Oscar Soláns Caballer, Pere Martí-Puig, Zhe Sun 0009, Toshihisa Tanaka 0001, Jordi Solé i Casals
ICASSP4
2022 Multi-scale Learning for Multimodal Neurophysiological Signals: Gait Pattern Classification as an Example
Feng Duan 0006, Yizhi Lv, Zhe Sun 0009
Neural Process. Lett.3
2022 Domain classifier-based transfer learning for visual attention prediction
Zhiwen Zhang 0004, Feng Duan 0006, Cesar F. Caiafa, Jordi Solé i Casals, Zhenglu Yang, Zhe Sun 0009
World Wide Web6
2021 News Content Completion with Location-Aware Image Selection
Zhengkun Zhang, Jun Wang 0023, Adam Jatowt, Zhe Sun 0009, Shao-Ping Lu, Zhenglu Yang
AAAI4
2021 LAMS: A Location-aware Approach for Multimodal Summarization (Student Abstract)
abstract
Multimodal summarization aims to refine salient information from multiple modalities, among which texts and images are two mostly discussed ones. In recent years, many fantastic works have emerged in this field by modeling image-text interactions; however, they neglect the fact that most of multimodal documents have been elaborately organized by their writers. This means that a critical organized factor has long been short of enough attention, that is, image locations, which may carry illuminating information and imply the key contents of a document. To address this issue, we propose a location-aware approach for multimodal summarization (LAMS) based on Transformer. We investigate image locations for multimodal summarization via a stack of multimodal fusion block, which can formulate the high-order interactions among images and texts. An extensive experimental study on an extended multimodal dataset validates the superior summarization performance of the proposed model.
Zhengkun Zhang, Jun Wang 0023, Zhe Sun 0009, Zhenglu Yang
AAAI3
2021 Generalized Relation Learning with Semantic Correlation Awareness for Link Prediction
abstract
Developing link prediction models to automatically complete knowledge graphs has recently been the focus of significant research interest. The current methods for the link prediction task have two natural problems: 1) the relation distributions in KGs are usually unbalanced, and 2) there are many unseen relations that occur in practical situations. These two problems limit the training effectiveness and practical applications of the existing link prediction models. We advocate a holistic understanding of KGs and we propose in this work a unified Generalized Relation Learning framework GRL to address the above two problems, which can be plugged into existing link prediction models. GRL conducts a generalized relation learning, which is aware of semantic correlations between relations that serve as a bridge to connect semantically similar relations. After training with GRL, the closeness of semantically similar relations in vector space and the discrimination of dissimilar relations are improved. We perform comprehensive experiments on six benchmarks to demonstrate the superior capability of GRL in the link prediction task. In particular, GRL is found to enhance the existing link prediction models making them insensitive to unbalanced relation distributions and capable of learning unseen relations.
Jun Wang 0023, Hongru Liang, Wenqiang Lei, Zhe Sun 0009, Adam Jatowt, Zhenglu Yang
AAAI6
2021 Mind Control of a Service Robot with Visual Servoing
abstract
In the growing elderly population globally, patients with severe movement disorders account for a large proportion. Moreover, the development of intelligent service equipment can better assist them in their daily. This paper proposes a new service robot control system. The brain-computer interface (BCI) based on Steady-State Visual Evoked Potentials (SSVEP) is used to acquire and process electroencephalogram(EEG) signals and output various control commands accordingly. Then, considering the visual fatigue of SSVEP-BCI, we added an object detection method based on Yolov3-tiny and saliency prediction to identify the patient’s selection intention intelligently. The results show that the subject can successfully complete the object delivery task with an average accuracy of 90.3%. The proposed control system can help the patients control a service robot in a more intelligent and friendly way to realize some daily tasks.
Zhe Sun 0009, Feng Duan 0006, Chi Zhu 0001, Hiroshi Yokoi
IROS2
2021 Component-mixing strategy: A decomposition-based data augmentation algorithm for motor imagery signals
Binghua Li 0001, Zhiwen Zhang 0004, Feng Duan 0006, Zhenglu Yang, Qibin Zhao, Zhe Sun 0009, Jordi Solé i Casals
Neurocomputing6
2021 Serial-EMD: Fast empirical mode decomposition method for multi-dimensional signals based on serialization
abstract
Empirical mode decomposition (EMD) has developed into a prominent tool for adaptive, scale-based signal analysis in various fields like robotics, security and biomedical engineering. Since the dramatic increase in amount of data puts forward higher requirements for the capability of real-time signal analysis, it is difficult for existing EMD and its variants to trade off the growth of data dimension and the speed of signal analysis. In order to decompose multi-dimensional signals at a faster speed, we present a novel signal-serialization method (serial-EMD), which concatenates multi-variate or multi-dimensional signals into a one-dimensional signal and uses various one-dimensional EMD algorithms to decompose it. To verify the effects of the proposed method, synthetic multi-variate time series, artificial 2D images with various textures and real-world facial images are tested. Compared with existing multi-EMD algorithms, the decomposition time becomes significantly reduced. In addition, the results of facial recognition with Intrinsic Mode Functions (IMFs) extracted using our method can achieve a higher accuracy than those obtained by existing multi-EMD algorithms, which demonstrates the superior performance of our method in terms of the quality of IMFs. Furthermore, this method can provide a new perspective to optimize the existing EMD algorithms, that is, transforming the structure of the input signal rather than being constrained by developing envelope computation techniques or signal decomposition methods. In summary, the study suggests that the serial-EMD technique is a highly competitive and fast alternative for multi-dimensional signal analysis.
Jin Zhang 0003, Pere Martí-Puig, Cesar F. Caiafa, Zhe Sun 0009, Feng Duan 0006, Jordi Solé i Casals
Inf. Sci.5
2018 JTAV: Jointly Learning Social Media Content Representation by Fusing Textual, Acoustic, and Visual Features
abstract
Learning social media content is the basis of many real-world applications, including information retrieval and recommendation systems, among others. In contrast with previous works that focus mainly on single modal or bi-modal learning, we propose to learn social media content by fusing jointly textual, acoustic, and visual information (JTAV). Effective strategies are proposed to extract fine-grained features of each modality, that is, attBiGRU and DCRNN. We also introduce cross-modal fusion and attentive pooling techniques to integrate multi-modal information comprehensively. Extensive experimental evaluation conducted on real-world datasets demonstrate our proposed model outperforms the state-of-the-art approaches by a large margin.
Hongru Liang, Haozheng Wang, Jun Wang 0023, Shaodi You, Zhe Sun 0009, Jinmao Wei 0001, Zhenglu Yang
COLING5
2018 HAVAE: Learning Prosodic-Enhanced Representations of Rap Lyrics
Hongru Liang, Qian Li 0016, Haozheng Wang, Jun Wang 0023, Zhe Sun 0009, Jinmao Wei 0001, Zhenglu Yang
PRICAI (1)6