EDBT 2026 Demo / reviewers in the wild / expert
Jiasong Wu
dblp:63/8859
· DBLP profile ↗
52ranked-venue papers
7as first author
34since 2021 · last 2026
0000-0001-7171-1318ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 25 · 5 first-author · 14 since 2021Artificial intelligence and machine learning · 23 · 2 first-author · 16 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 4 since 2021Databases, data management, data science and information retrieval · 4 · 4 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | HMTE: Memory-transformer representation learning for knowledge hypergraph completion
Wanqiang Cai, Yingyao Ma, Lotfi Senhadji, Huazhong Shu, Jiasong Wu |
Neurocomputing | 7 |
| 2026 | Seeing cracks in frequency: FD-Mamba accurately segments cracks via frequency-difference priors
Wanqiang Cai, Junwen Zheng, Jiasong Wu, ZongYuan Ge, Laurent D. Cohen |
Pattern Recognit. | 6 |
| 2026 | FocusKG: A Novel Multimodal Knowledge Graph Dataset With Temporal Information for Link PredictionabstractKnowledge graphs (KGs) play a central role in enabling structured reasoning for AI applications. Recent research has advanced two prominent extensions of KGs: multimodal KGs, which integrate diverse data sources such as text, images, audio, and video; and temporal KGs, which capture the dynamic evolution of knowledge over time. However, these two paradigms remain largely disjoint—multimodal KGs typically ignore temporal dynamics, while temporal KGs overlook rich multimodal context. To bridge this gap, we introduceFocusKG, the first discrete-time multimodal knowledge graph that unifies four modalities (text, image, audio, and video) with temporal annotations. Built from the Focus news program, FocusKG captures real-world events as temporally grounded multimodal facts, offering new expressiveness for time-sensitive tasks. We further propose aDiscrete-TimeMultimodal Knowledge GraphEmbedding (DTME) method, a novel framework tailored for learning representations over temporal multimodal KGs. DTME is composed of three key modules: (1) Multimodal Preprocessing, which encodes raw inputs from each modality into vector representations using modality-specific encoders; (2) Cross-Modal Graph Enhancer, which captures semantic interplay between entity and relation modalities through role-aware subgraph construction and a modality interaction network; and (3) Multi-Branch Time Aggregation, which injects temporal signals by modeling interactions between time and modality-aware features. Extensive experiments demonstrate FocusKG's value and DTME's efficacy. On link prediction, DTME outperforms state-of-the-art multimodal KG models and temporal KG models. We release FocusKG as a benchmark to foster research in multimodal temporal reasoning. Yingyao Ma, Wanqiang Cai, Rubing Duan, Jiasong Wu, Lotfi Senhadji, Huazhong Shu |
IEEE Trans. Knowl. Data Eng. | 5 |
| 2026 | SCL: Semantic Coherence Learning for Video Question AnsweringabstractVideo Question Answering (VideoQA) requires clarifying the true question intent and understanding dynamic video content to predict the answer. However, most existing methods overlook the temporal semantic exploration of fine-grained clues across frames, and neglect the interrogative properties of question text, leading to suboptimal grounding of video frames. To address these limitations, in this paper, we propose a novel Semantic Coherence Learning (SCL) framework, which explicitly models the fine-grained spatio-temporal coherence of critical visual regions and progressively mines multi-modal semantic clues aligned with the reasoning intent. Specifically, we design a spatio-temporal clue reasoning module to adaptively model the fine-grained coherent transition semantics of critical regions across temporal frames while suppressing irrelevant noisy regions within frames. Additionally, we propose a multi-modal clue reasoning module to progressively mine multi-modal coherent semantics associated with reasoning intent, which iteratively grounds critical frames in videos based on refined question semantics and mines critical phrases in questions based on refined video semantics. Finally, the answer is derived through an answer decoder with intra and inter-sample contrastive strategies. Extensive experimental results on six widely used benchmarks verify our effectiveness and superiority. The experimental source codes are available at https://github.com/XizeWu/SCL. Xize Wu, Zheng Wang 0044, Jiasong Wu, Lei Zhu 0002, Heng Tao Shen |
IEEE Trans. Multim. | 3 |
| 2026 | SSWMNet: Solving the Speech Separation Problem While the Target is Wearing a MaskabstractSingle-channel speech separation remains one of the most challenging tasks in the field of speech signal processing. In many situations, such as during epidemics that involve respiratory diseases (e.g., COVID-19 or influenza A), individuals are required to wear masks while communicating. Is it possible to address the challenge of speech separation when the target speaker is wearing a mask? Can audio–visual approaches achieve better speech separation performance than that of audio-only approaches in scenarios where speakers are wearing masks? To address the aforementioned questions, we first construct a large-scale multimodal dataset, termed Speech Separation while Wearing a Mask (SSWM), which includes both the audio modality and the visual modality with masked faces. We explore two strategies for addressing the problem of facial occlusion. One strategy involves utilizing occluded faces—which lack critical visual cues such as mouth movements—directly as supervisory information for self-supervised speech separation; the other strategy involves the use of Wav2Lip to first generate visual information, which is then used as supervisory guidance for self-supervised speech separation. Building upon these two strategies, we propose the SSWM network (SSWMNet), which can flexibly choose to either utilize occluded facial images directly or employ Wav2Lip to generate visual information. The experimental results demonstrate that the proposed speech separation method in which Wav2Lip is used for visual information generation outperforms the approach of utilizing occluded faces directly for self-supervised speech separation. Both proposed audio–visual methods outperform the audio-only speech separation approach, which operates without the aid of visual information. Availability—SSWMNet is available at https://github.com/fanmanqian/SSWMNetwork . Fanman Meng, Kang Qin, Huazhong Shu, Lotfi Senhadji, Jiasong Wu |
ACM Trans. Internet Techn. | 6 |
| 2025 | BiGuidedPrompt: Dynamic Bidirectional Guided Multimodal Prompt Learning
Jiacheng Zhong, Xinguo Zhang, Bingxue Zhang, Jiasong Wu |
ICIC (21) | 4 |
| 2025 | CrackMamba with Normalized Soft-Frangi-Filter Enhancement towards Accurate Crack SegmentationabstractCrack segmentation is crucial in monitoring infrastructure degradation. However, the strong contextual dependencies of long-spanned morphology and numerous fine-grained branches hinder high-precision segmentation performance. To tackle the above obstacles, we provide a pure Mamba-based model termed CrackMamba for accurate crack segmentation. CrackMamba utilizes the Mamba-based Feature Extractor (MFE) to effectively model global dependencies of long-spanned cracks with linear computational complexity for powerful representation. To capture fine-grained crack branches, we design a novel texture refinement module, which employs a multi-scale aggregation strategy and utilizes the MFE and a Normalized Soft-Frangi-Filter (NSFF) module to integrate hierarchical features from the decoder. The CrackMamba and NSFF exhibit strong complementarity. The NSFF module enhances the capability of CrackMamba to segment fine-grained crack textures, while CrackMamba effectively eliminating crack-unrelated curves extracted by NSFF module. Extensive experiments are conducted on two benchmark crack datasets, and the results demonstrate that the proposed CrackMamba achieves the state-of-the-art (SOTA) performance with fewer parameters and higher computational efficiency. Wanqiang Cai, Yingyao Ma, Jiasong Wu, ZongYuan Ge, Bin Wang 0041 |
ICMR | 5 |
| 2025 | BFC-Net: Boundary-Frame cross graph attention network for partially spoofed audio localization
Zhaodong Xue, Lotfi Senhadji, Huazhong Shu, Jiasong Wu |
Neurocomputing | 5 |
| 2025 | HSAE: Hierarchical structure augment embedding for various knowledge graph completion
Wanqiang Cai, Yingyao Ma, Lotfi Senhadji, Huazhong Shu, Jiasong Wu |
Knowl. Based Syst. | 6 |
| 2025 | Make your choice for multimodal knowledge graph completion
Shuoyan Ren, Wanqiang Cai, Yingyao Ma, Lotfi Senhadji, Huazhong Shu, Jiasong Wu |
Knowl. Based Syst. | 7 |
| 2025 | Collaborative Aware Bidirectional Semantic Reasoning for Video Question AnsweringabstractVideo question answering (VideoQA) is the challenging task of accurately responding to natural language questions based on a given video. Most previous methods focus on designing complex cross-modal interactions to perform question-oriented video scene mining and semantic reasoning, and utilize straightforward classification and matching strategies with different decoders to forcibly associate the predicted representation with ground-truth answer. However, the limitations of question-oriented reasoning and the overlapping semantic co-occurrences between questions and candidates may cause them to fall into spurious correlation reasoning. In this paper, we propose a Collaborative aware Bidirectional Semantic Reasoning (CBSR) model to alleviate this challenging problem. Specifically, we first propose a collaborative aware adaptive correlation reasoning module to collaboratively mine multi-granularity text-aware critical video scenes and reason about the complex intrinsic correlations between them via bottom-up cross-granularity adaptive aggregation. By progressively performing video reasoning from object-level to frame-level, we can obtain a set of semantically rich critical video representations. Then, we collaboratively decode it together with question and knowledge semantics into an implicit representation through the proposed unified answer semantic collaborated decoding module. Finally, a novel bidirectional semantic reasoning learning strategy is proposed to bridge and strengthen the unique positive semantic correlation between the learned implicit representation and the ground-truth answer, and explicitly alleviate the challenge of overlapping semantic co-occurrence. Benefiting from the same model structure and learning strategy, our method can achieve seamless transfer between Open-Ended and Multi-Choice tasks. Extensive experimental results on seven commonly tested datasets (i.e. MSVD-QA, MSRVTT-QA, NExT-QA, Causal-VidQA, NExT-OOD, ActivityNet-QA and EgoSchema) verify the superior performance of our method and the effectiveness of each reasoning module. We provide our source codes and experimental datasets athttps://github.com/XizeWu/CBSR. Xize Wu, Jiasong Wu, Lei Zhu 0002, Lotfi Senhadji, Huazhong Shu |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2025 | TTFNet: Temporal-Frequency Features Fusion Network for Speech Based Automatic Depression Recognition and AssessmentabstractRelated studies have revealed that the phonological features of depressed patients are different from those of healthy individuals. With the increasing prevalence of depression, an objective and convenient approach for early screening is necessary. To this end, we propose an automatic depression detection method based on hybrid speech features extracted by deep learning, dubbed as TTFNet. Firstly, to effectively excavate the intrinsic relationship among multidimensional dynamic features in the frequency domain, the log-Mel spectrogram of raw speech and its related derivatives are encoded into quaternion representation. Then, the innovatively designed quaternion VisionLSTM is utilized to capture their synergistic effects. Simultaneously, we integrate sLSTM with the pre-trained wav2vec 2.0 model to fully acquire the temporal features. In addition, to further exploit the complementarity between temporal and frequency features, we design an XConformer block for cross-sequence interactions, which ingeniously combines self-attention mechanisms and convolutional modules. Based on this block, the dual-path fusion module closely utilizes the mutual promotion of features from different domains, thereby enhancing generalization capability of the proposed model. Extensive experiments conducted on the AVEC 2013, AVEC 2014, DAIC-WOZ and E-DAIC datasets demonstrate that our method outperforms current state-of-the-art methods in both depression recognition and severity prediction tasks. Xiyuan Chen 0004, Zhuhong Shao, Yinan Jiang, Runsen Chen, Bicao Li, Mingyue Niu, Hongguang Chen, Jiasong Wu |
IEEE J. Biomed. Health Informatics | 10 |
| 2025 | Multimodal Entity Linking With Dynamic Modality Selection and Interactive Prompt LearningabstractRecent advances in Multimodal Entity Linking leverage multimodal information to link target mentions to corresponding entities. However, existing methods uniformly adopt a “one-size-fits-all” approach, which overlooks the unique requirements of individual samples and fails to adequately balance modality-assisted disambiguation and modality-induced noise. Also, the commonly used separate large-scale visual and text pretrained models for feature extraction do not address inter-modal heterogeneity and the high computational cost of fine-tuning. To resolve these two issues, we introduce a novel approach named Multimodal Entity Linking with Dynamic Modality Selection and Interactive Prompt Learning (DSMIP). First, we design three expert networks that utilize different subsets of modalities tailored to the task and train them individually. Specifically, for the multimodal expert network, we enhance entity and mention feature extraction by updating multimodal prompts and setting up a coupling function to realize the interaction of prompts between modalities. Subsequently, to select the best-suited expert network for each specific sample, we devise a Modality Selection Gating Network to gain the optimal one-hot selection vector by applying a specialized reparameterization technique and a two-stage training process. Experimental results on three public benchmark datasets demonstrate that the proposed DSMIP outperforms all state-of-the-art baselines. The code is released on https://github.com/mayy-seu/DSMIP-code. Yingyao Ma, Jiasong Wu, Lotfi Senhadji, Huazhong Shu, Jian Yang 0009 |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2025 | Wavelet-Based Dual-Task NetworkabstractIn image processing, wavelet transform (WT) offers multiscale image decomposition, generating a blend of low-resolution approximation images and high-resolution detail components. Drawing parallels to this concept, we view feature maps in convolutional neural networks (CNNs) as a similar mix, but uniquely within the channel domain. Inspired by multitask learning (MTL) principles, we propose a wavelet-based dual-task (WDT) framework. This novel framework employs WT in the channel domain to split a single task into two parallel tasks, thereby reforming traditional single-task CNNs into dynamic dual-task networks. Our WDT framework integrates seamlessly with various popular network architectures, enhancing their versatility and efficiency. It offers a more rational approach to resource allocation in CNNs, balancing between low-frequency and high-frequency information. Rigorous experiments on Cifar10, ImageNet, HMDB51, and UCF101 validate our approach's effectiveness. Results reveal significant improvements in the performance of traditional CNNs on classification tasks, and notably, these enhancements are achieved with fewer parameters and computations. In summary, our work presents a pioneering step toward redefining the performance and efficiency of CNN-based tasks through WT. Fuzhi Wu, Jiasong Wu, Chen Zhang 0024, Youyong Kong, Guanyu Yang 0001, Huazhong Shu, Guy Carrault, Lotfi Senhadji |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2025 | LowDDAWP-Net: Low-Resolution Double Deep Audio Waveform Prior Network for Audio Systems Reliability DefenceabstractImproving the information reliability of the audio system is critical to safeguarding the security of the audio system. Adversarial samples crafted by in-the-wild attackers by introducing perturbations to the audio become a severe threat to the trustworthiness of deep learning-based classifiers. To achieve dynamic defence against audio adversarial sample attacks, a low-resolution double deep audio waveform prior network (LowDDAWP-Net) for audio systems reliability defence is proposed. Specifically, LowDDAWP-Net consists of a noise audio prior extraction module ($\mathbf{DAWP}_{\mathbf{noise}}$), an speech prior extraction module ($\mathbf{DAWP}_{\mathbf{speech}}$), a low-resolution extraction module (LREM), and a voice activity detection module (VADM). The role of the VADM is to automatically detect voice activity signals and silent signals from the audio signal.$\mathbf{DAWP}_{\mathbf{speech}}$and$\mathbf{DAWP}_{\mathbf{noise}}$are encoder–decoders with the same architecture. The encoder extracts the superficial features of the input audio, and the decoder performs temporal fusion to form high-dimensional features and reconstructs them into waveform signals. A LREM is employed to extract low-resolution audio to facilitate the encoder–decoder to perform detail on low-resolution audio and to speed up the recovery of DAWP networks to high resolution. The adversarial samples generated by several diverse attack state-of-the-art on three different datasets and their corresponding benign samples form a novel private dataset. The qualitative and quantitative results of the novel private dataset demonstrate the effectiveness and superiority of LowDDAWP-Net. Kai Chen 0039, Yikun Zhang 0001, Jiasong Wu, Jean-Louis Coatrieux, Yang Chen 0008 |
IEEE Trans. Reliab. | 3 |
| 2024 | Multiscale Low-Frequency Memory Network for Improved Feature Extraction in Convolutional Neural NetworksabstractDeep learning and Convolutional Neural Networks (CNNs) have driven major transformations in diverse research areas. However, their limitations in handling low-frequency in-formation present obstacles in certain tasks like interpreting global structures or managing smooth transition images. Despite the promising performance of transformer struc-tures in numerous tasks, their intricate optimization com-plexities highlight the persistent need for refined CNN en-hancements using limited resources. Responding to these complexities, we introduce a novel framework, the Mul-tiscale Low-Frequency Memory (MLFM) Network, with the goal to harness the full potential of CNNs while keep-ing their complexity unchanged. The MLFM efficiently preserves low-frequency information, enhancing perfor-mance in targeted computer vision tasks. Central to our MLFM is the Low-Frequency Memory Unit (LFMU), which stores various low-frequency data and forms a parallel channel to the core network. A key advantage of MLFM is its seamless compatibility with various prevalent networks, requiring no alterations to their original core structure. Testing on ImageNet demonstrated substantial accuracy improvements in multiple 2D CNNs, including ResNet, MobileNet, EfficientNet, and ConvNeXt. Furthermore, we showcase MLFM's versatility beyond traditional image classification by successfully integrating it into image-to-image translation tasks, specifically in semantic segmenta-tion networks like FCN and U-Net. In conclusion, our work signifies a pivotal stride in the journey of optimizing the ef-ficacy and efficiency of CNNs with limited resources. This research builds upon the existing CNN foundations and paves the way for future advancements in computer vision. Our codes are available at https://github.com/AlphaWuSeu/MLFM. Fuzhi Wu, Jiasong Wu, Youyong Kong, Guanyu Yang 0001, Huazhong Shu, Guy Carrault, Lotfi Senhadji |
AAAI | 2 |
| 2024 | ST-LDM: A Universal Framework for Text-Grounded Object Generation in Real Images
Xiangtian Xue, Jiasong Wu, Youyong Kong, Lotfi Senhadji, Huazhong Shu |
ECCV (46) | 2 |
| 2024 | AHMN: A multi-modal network for long MOOC videos chapter segmentation
Jiasong Wu, Youyong Kong, Huazhong Shu, Lotfi Senhadji |
Multim. Tools Appl. | 1 |
| 2024 | CSLNSpeech: Solving the extended speech separation problem with the help of Chinese sign language
Jiasong Wu, Taotao Li, Fanman Meng, Youyong Kong, Guanyu Yang 0001, Lotfi Senhadji, Huazhong Shu |
Speech Commun. | 1 |
| 2024 | Spatial-Enhanced Multi-Level Wavelet Patching in Vision TransformersabstractBy seamlessly integrating wavelet transforms into the image patching stage of ViT, we leverage the power of multi-level wavelet transforms to decompose images into a diverse array of frequency-domain features. These features, integrated with spatial characteristics at equivalent scales, enrich image details, enhancing ViT's proficiency in delineating intricate textures and distinct edges. Consequently, we registered a notable 2.7% accuracy enhancement on the ImageNet100 dataset in ViT. Our wavelet patching module, designed for versatility, seamlessly fits into various ViT derivatives without necessitating architecture modifications. This advancement has uplifted the performance of several leading vision transformers by 0.46–4.3%, preserving parameter efficiency without notable FLOPs increment. Fuzhi Wu, Jiasong Wu, Huazhong Shu, Guy Carrault, Lotfi Senhadji |
IEEE Signal Process. Lett. | 2 |
| 2024 | Improving End-to-End Sign Language Translation With Adaptive Video Representation Enhanced TransformerabstractThe aim of end-to-end sign language translation (SLT) is to interpret continuous sign language (SL) video sequences into coherent natural language sentences without any intermediary annotations, i.e., glosses. However, end-to-end SLT suffers several intractable issues: (i) the temporal correspondence constraint loss problem between SL videos and glosses, and (ii) the weakly supervised sequence labeling problem between long SL videos and sentences. To address these issues, we propose an adaptive video representation enhanced Transformer (AVRET), with three extra modules: adaptive masking (AM), local clip self-attention (LCSA) and adaptive fusion (AF). Specifically, we utilize the first AM module to generate a special mask that adaptively drops out temporally important SL video frame representations to enhance the SL video features. Then, we pass the masked video feature to the Transformer encoder consisting of LCSA and masked self-attention to learn clip-level and continuous video-level feature information. Finally, the output feature of encoder is fused with the temporal feature of AM module via the AF module and use the second AM module to generate more robust feature representations. Besides, we add weakly supervised loss terms to constrain these two AM modules. To promote the Chinese SLT research, we further construct CSL-FocusOn, a Chinese continuous SLT dataset, and share its collection method. It involves many common scenarios, and provides SL sentence annotations and multi-cue images of signers. Our experiments on the CSL-FocusOn, PHOENIX14T, and CSL-Daily datasets show that the proposed method achieves the competitive performance on the end-to-end SLT task without using glosses in training. The code is available at https://github.com/LzDddd/AVRET. Jiasong Wu, Xin Chen 0086, Qianyu Wu, Zhiguo Gui, Lotfi Senhadji, Huazhong Shu |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2024 | RED-Net: Residual and Enhanced Discriminative Network for Image Steganalysis in the Internet of Medical Things and TelemedicineabstractInternet of Medical Things (IoMT) and telemedicine technologies utilize computers, communications, and medical devices to facilitate off-site exchanges between specialists and patients, specialists, and medical staff. If the information communicated in IoMT is illegally steganography, tampered or leaked during transmission and storage, it will directly impact patient privacy or the consultation results with possible serious medical incidents. Steganalysis is of great significance for the identification of medical images transmitted illegally in IoMT and telemedicine. In this article, we propose a Residual and Enhanced Discriminative Network (RED-Net) for image steganalysis in the internet of medical things and telemedicine. RED-Net consists of a steganographic information enhancement module, a deep residual network, and steganographic information discriminative mechanism. Specifically, a steganographic information enhancement module is adopted by the RED-Net to boost the illegal steganographic signal in texturally complex high-dimensional medical image features. A deep residual network is utilized for steganographic feature extraction and compression. A steganographic information discriminative mechanism is employed by the deep residual network to enable it to recalibrate the steganographic features and drop high-frequency features that are mistaken for steganographic information. Experiments conducted on public and private datasets with data hiding payloads ranging from 0.1bpp/bpnzac-0.5bpp/bpnzac in the spatial and JPEG domain led to RED-Net's steganalysis error$P_{\mathrm{E}}$in the range of 0.0732-0.0010 and 0.231-0.026, respectively. In general, qualitative and quantitative results on public and private datasets demonstrate that the RED-Net outperforms 8 state-of-art steganography detectors. Kai Chen 0039, Zhengyuan Zhou, Jiasong Wu, Jean-Louis Coatrieux, Yang Chen 0008, Gouenou Coatrieux |
IEEE J. Biomed. Health Informatics | 5 |
| 2024 | Learning Better Registration to Learn Better Few-Shot Medical Image Segmentation: Authenticity, Diversity, and RobustnessabstractIn this work, we address the task of few-shot medical image segmentation (MIS) with a novel proposed framework based on the learning registration to learn segmentation (LRLS) paradigm. To cope with the limitations of lack of authenticity, diversity, and robustness in the existing LRLS frameworks, we propose the better registration better segmentation (BRBS) framework with three main contributions that are experimentally shown to have substantial practical merit. First, we improve the authenticity in the registration-based generation program and propose the knowledge consistency constraint strategy that constrains the registration network to learn according to the domain knowledge. It brings the semantic-aligned and topology-preserved registration, thus allowing the generation program to output new data with great space and style authenticity. Second, we deeply studied the diversity of the generation process and propose the space-style sampling program, which introduces the modeling of the transformation path of style and space change between few atlases and numerous unlabeled images into the generation program. Therefore, the sampling on the transformation paths provides much more diverse space and style features to the generated data effectively improving the diversity. Third, we first highlight the robustness in the learning of segmentation in the LRLS paradigm and propose the mix misalignment regularization, which simulates the misalignment distortion and constrains the network to reduce the fitting degree of misaligned regions. Therefore, it builds regularization for these regions improving the robustness of segmentation learning. Without any bells and whistles, our approach achieves a new state-of-the-art performance in few-shot MIS on two challenging tasks that outperform the existing LRLS-based few-shot methods. We believe that this novel and effective framework will provide a powerful few-shot benchmark for the field of medical image and efficiently reduce the costs of medical image research. All of our code will be made publicly available online. Yuting He 0001, Rongjun Ge, Xiaoming Qi, Yang Chen 0008, Jiasong Wu, Jean-Louis Coatrieux, Guanyu Yang 0001, Shuo Li 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2023 | Residual shuffle attention network for image super-resolution
Xuanyi Li, Zhuhong Shao, Bicao Li, Jiasong Wu, Yuping Duan |
Mach. Vis. Appl. | 5 |
| 2023 | Randomized nonlinear two-dimensional principal component analysis network for object recognition
Zhijian Sun, Zhuhong Shao, Bicao Li, Jiasong Wu |
Mach. Vis. Appl. | 5 |
| 2023 | Multi-scale self-attention mixup for graph classification
Youyong Kong, Jiasong Wu |
Pattern Recognit. Lett. | 4 |
| 2023 | Self-supervised speech denoising using only noisy audio signals
Jiasong Wu, Qingchun Li, Guanyu Yang 0001, Lei Li 0020, Lotfi Senhadji, Huazhong Shu |
Speech Commun. | 1 |
| 2022 | Temporal Cross-Graph Network for Brain Functional Activity PredictionabstractPrediction of brain functional activity is of great significance for neuroscience research. The brain functional activities at different regions are highly related, and their relationships can be captured with functional connectivity and structural connectivity. The existing works are challenging to integrate two connectivity information for functional activity prediction. In this paper, we propose a Temporal Cross-Graph Network (TCGN) for predicting brain functional activity, which can comprehensively exploit multi-modal spatial dependence and temporal patterns. In particular, a novel cross-graph convolution module is developed to capture the spatial features of brain structural and functional connectivity. A temporal fusion module is designed to learn the pattern of dynamic functional connectivity to guide the prediction. Specially, a multi-task loss function is proposed to incorporate functional activity and dynamic functional connectivity. Extensive experiments on the Human Connectome Project dataset demonstrate the effectiveness of the proposed framework. Xinyu Yuan, Wenhan Wang, Youyong Kong, Jiasong Wu, Guanyu Yang 0001, Huazhong Shu |
ICASSP | 4 |
| 2022 | Hierarchical Diffusion Scattering Graph Neural NetworkabstractGraph neural network (GNN) is popular now to solve the tasks in non-Euclidean space and most of them learn deep embeddings by aggregating the neighboring nodes. However, these methods are prone to some problems such as over-smoothing because of the single-scale perspective field and the nature of low-pass filter. To address these limitations, we introduce diffusion scattering network (DSN) to exploit high-order patterns. With observing the complementary relationship between multi-layer GNN and DSN, we propose Hierarchical Diffusion Scattering Graph Neural Network (HDS-GNN) to efficiently bridge DSN and GNN layer by layer to supplement GNN with multi-scale information and band-pass signals. Our model extracts node-level scattering representations by intercepting the low-pass filtering, and adaptively tunes the different scales to regularize multi-scale information. Then we apply hierarchical representation enhancement to improve GNN with the scattering features. We benchmark our model on nine real-world networks on the transductive semi-supervised node classification task. The experimental results demonstrate the effectiveness of our method. Xinyan Pu, Jiasong Wu, Huazhong Shu, Youyong Kong |
IJCAI | 4 |
| 2022 | Convolutional modulation theory: A bridge between convolutional neural networks and signal modulation theory
Fuzhi Wu, Jiasong Wu, Youyong Kong, Guanyu Yang 0001, Huazhong Shu, Guy Carrault, Lotfi Senhadji |
Neurocomputing | 2 |
| 2021 | Thin Semantics Enhancement via High-Frequency Priori Rule for Thin Structures SegmentationabstractReceptive field-based segmentation models represent features in receptive fields having weak perception for thin semantics in thin structures segmentation, due to the challenges in small local size and large global variation. High-frequency (HiFe) components have strong thin perception ability and is stable for global variation, but its weak adaptability limits its direct application. We propose a HiFe priori rule which enables the network to adaptively extract and fuse HiFe components, enhancing the thin semantics and making the network naturally prefer thin structures for their segmentation. We further propose High-Frequency Semantics Enhancement Network (HiFeNet) based on our HiFe priori rule, boosting the SOTA methods in thin structures segmentation: 1) Our Deep High Frequency (DHiFe) block learns to extract task-dependent HiFe components and adds them to feature maps, achieving great perception of thin structures. 2) Our Latent Residual Denoising (LRD) block progressively weakens task-independent features via hierarchical residuals and learns to fuse HiFe components back to feature maps, further enhancing the thin semantics and weakening the interference of global variation. Extensive experiments on the retinal vessel [1], [2], [3] and Massachusetts road [4] segmentation datasets show great superiority of our HiFeNet. Yuting He 0001, Rongjun Ge, Jiasong Wu, Jean-Louis Coatrieux, Huazhong Shu, Yang Chen 0008, Guanyu Yang 0001, Shuo Li 0001 |
ICDM | 3 |
| 2021 | Semi-Supervised Medical Image Semantic Segmentation with Multi-scale Graph Cut LossabstractMost semantic segmentation methods are based on supervised convolutional neural networks which require large amounts of labeled data. However, the acquisition of a large number of high-quality labels is time-consuming and of high annotation cost for medical images. In this paper, we propose a semi-supervised learning framework based on a novel multi-scale graph cut loss function. Firstly, the multi-scale features obtained from the segmentation network are utilized to construct the graph in non-Euclidean space. Then the long-distance information between voxels at different scales can be captured through the graph embedding module. After that, the graph cut loss is calculated according to the final latent features. Only a few labeled data is needed in our proposed method, which is of significance in the practical clinic. The experiments on the BrainWeb20 dataset and the IBSR18 dataset demonstrate the effectiveness of the proposed method compared to the well-known state-of-the-art methods. Junxiao Sun, Yan Zhang 0094, Jiasong Wu, Youyong Kong |
ICIP | 4 |
| 2021 | GSCFN: A graph self-construction and fusion network for semi-supervised brain tissue segmentation in MRI
Yan Zhang 0094, Youyong Kong, Jiasong Wu, Jian Yang 0009, Huazhong Shu, Gouenou Coatrieux |
Neurocomputing | 4 |
| 2021 | Quaternion discrete fractional Krawtchouk transform and its application in color image encryption and watermarking
Xilin Liu 0003, Yongfei Wu, Hao Zhang 0061, Jiasong Wu, Liming Zhang 0002 |
Signal Process. | 4 |
| 2020 | Deep octonion networks
Jiasong Wu, Fuzhi Wu, Youyong Kong, Lotfi Senhadji, Huazhong Shu |
Neurocomputing | 1 |
| 2020 | Dense biased networks with deep priori anatomy and hard region adaptation: Semi-supervised learning for fine renal artery segmentation
Yuting He 0001, Guanyu Yang 0001, Jian Yang 0009, Yang Chen 0008, Youyong Kong, Jiasong Wu, Lijun Tang, Xiaomei Zhu, Jean-Louis Dillenseger, Shaobo Zhang 0008, Huazhong Shu, Jean-Louis Coatrieux, Shuo Li 0001 |
Medical Image Anal. | 6 |
| 2020 | The modified generic polar harmonic transforms for image representation
Xilin Liu 0003, Yongfei Wu, Zhuhong Shao, Jiasong Wu |
Pattern Anal. Appl. | 4 |
| 2020 | Anisotropic tubular minimal path model with fast marching front freezing scheme
Li Liu 0065, Da Chen 0002, Laurent D. Cohen, Jiasong Wu, Michel Pâques, Huazhong Shu |
Pattern Recognit. | 4 |
| 2019 | Brain Tissue Segmentation based on Graph Convolutional NetworksabstractIn neuroscience research, brain tissue segmentation from magnetic resonance imaging is of significant importance. A challenging issue is to provide an accurate segmentation due to the tissue heterogeneity, which is caused by noise, bias filed and partial volume effects. To overcome these problems, we propose a novel brain MRI segmentation algorithm, the originality of which stands on the combination of supervoxels with graph convolutional networks. Supervoxels are generated from the 3D MRI image with the help of an improved simple linear iterative clustering algorithm. A graph is then built from these supervoxels through the K nearest neighbor algorithm, before being sent to GCNs for tissues classification. The proposed method is evaluated on the two common datasets- the BrainWeb18 dataset and the Internet Brain Segmentation Repository 18 dataset. Experiments demonstrate the performance of our method and that it is better than well-known state-of-the-art methods such as FMRIB software library, statistical parametric mapping, adaptive graph filter. Yan Zhang 0094, Youyong Kong, Jiasong Wu, Gouenou Coatrieux, Huazhong Shu |
ICIP | 3 |
| 2019 | DPA-DenseBiasNet: Semi-supervised 3D Fine Renal Artery Segmentation with Dense Biased Network and Deep Priori Anatomy
Yuting He 0001, Guanyu Yang 0001, Yang Chen 0008, Youyong Kong, Jiasong Wu, Lijun Tang, Xiaomei Zhu, Jean-Louis Dillenseger, Shaobo Zhang 0008, Huazhong Shu, Jean-Louis Coatrieux, Shuo Li 0001 |
MICCAI (6) | 5 |
| 2018 | Automatic Segmentation of Kidney and Renal Tumor in CT Images Based on 3D Fully Convolutional Neural Network with Pyramid Pooling ModuleabstractRenal cancer is one of ten most common cancers in human beings. The laparoscopic partial nephrectomy (LPN) becomes the main therapeutic approach in treating renal cancer. Accurate kidney and tumor segmentation in CT images is a prerequisite step in the surgery planning. However, automatic and accurate kidney and renal tumor segmentation in CT images remains a challenge. In this paper, we propose a new method to perform a precise segmentation of kidney and renal tumor in CT angiography images. This method relies on a three-dimensional (3D) fully convolutional network (FCN) which combines a pyramid pooling module (PPM). The proposed network is implemented as an end-to-end learning system directly on 3D volumetric images. It can make use of the 3D spatial contextual information to improve the segmentation of the kidney as well as the tumor lesion. The experiments conducted on 140 patients show that these target structures can be segmented with a high accuracy. The resulting average dice coefficients obtained for kidney and renal tumor are equal to 0.931 and 0.802 respectively. These values are higher than those obtained from the other two neural networks. Guanyu Yang 0001, Tan Pan, Youyong Kong, Jiasong Wu, Huazhong Shu, Limin Luo 0001, Jean-Louis Dillenseger, Jean-Louis Coatrieux, Lijun Tang, Xiaomei Zhu |
ICPR | 5 |
| 2018 | PCANet: An energy perspective
Jiasong Wu, Shijie Qiu, Youyong Kong, Longyu Jiang, Yang Chen 0008, Wankou Yang, Lotfi Senhadji, Huazhong Shu |
Neurocomputing | 1 |
| 2017 | MomentsNet: A simple learning-free method for binary image recognitionabstractIn this paper, we propose a new simple and learning-free deep learning network named MomentsNet, whose convolution layer, nonlinear processing layer and pooling layer are constructed by Moments kernels, binary hashing and block-wise histogram, respectively. Twelve typical moments (including geometrical moment, Zernike moment, Tchebichef moment, etc.) are used to construct the MomentsNet whose recognition performance for binary image is studied. The results reveal that MomentsNet has better recognition performance than its corresponding moments in almost all cases and ZernikeNet achieves the best recognition performance among MomentsNet constructed by twelve moments. ZernikeNet also shows better recognition performance on a binary image database than that of PCANet, which is a learning-based deep learning network. Jiasong Wu, Shijie Qiu, Youyong Kong, Yang Chen 0008, Lotfi Senhadji, Huazhong Shu |
ICIP | 1 |
| 2016 | Color image classification via quaternion principal component analysis network
Jiasong Wu, Zhuhong Shao, Yang Chen 0008, Beijing Chen, Lotfi Senhadji, Huazhong Shu |
Neurocomputing | 2 |
| 2016 | Robust watermarking scheme for color image based on quaternion-type moment invariants and visual cryptography
Zhuhong Shao, Huazhong Shu, Gouenou Coatrieux, Jiasong Wu |
Signal Process. Image Commun. | 6 |
| 2015 | Tensor object classification via multilinear discriminant analysis networkabstractThis paper proposes an multilinear discriminant analysis network (MLDANet) for the recognition of multidimensional objects, knows as tensor objects. The MLDANet is a variation of linear discriminant analysis network (LDANet) and principal component analysis network (PCANet), both of which are the recently proposed deep learning algorithms. The MLDANet consists of three parts: 1) The encoder learned by MLDA from tensor data. 2) Features maps obtained from decoder. 3) The use of binary hashing and histogram for feature pooling. A learning algorithm for MLDANet is described. Evaluations on UCF11 database indicate that the proposed MLDANet outperforms the PCANet, LDANet, MPCA+LDA, and MLDA in terms of classification for tensor objects. Jiasong Wu, Lotfi Senhadji, Huazhong Shu |
ICASSP | 2 |
| 2014 | Quaternion Bessel-Fourier moments and their invariant descriptors for object reconstruction and recognition
Zhuhong Shao, Huazhong Shu, Jiasong Wu, Beijing Chen, Jean-Louis Coatrieux |
Pattern Recognit. | 3 |
| 2013 | Quaternion gyrator transform and its application to color image encryptionabstractThe gyrator transform has been proposed in optics a few years ago. By using the theory of quaternion numbers, this paper presents the quaternion gyrator transform (QGT). It is shown that the QGT can be computed via the left-side type of quaternion Fourier transforms. The new transform is applied to color image encryption for validation, where the rotation angles are used as encryption keys making it more secure compared to a recent method using discrete quaternion Fourier transforms (DQFTs). Experimental results show that the proposed encryption algorithm for color image performs as well as the DQFTs method in terms of noise robustness, so that it could be a useful tool for color image encryption. Zhuhong Shao, Jiasong Wu, Jean-Louis Coatrieux, Gouenou Coatrieux, Huazhong Shu |
ICIP | 2 |
| 2012 | Fast Radix-3 Algorithm for the Generalized Discrete Hartley Transform of Type IIabstractWe present a new fast radix-3 algorithm for the computation of the length-Ngeneralized discrete Hartley transform of type-II (GDHT-II), whereN= 3m,m≥ 2. Then we apply this algorithm to the direct computation of length-NGDHT-II coefficients when given three adjacent length-N/3 GDHT-II coefficients. The computational complexity of the proposed method is lower than that of the traditional approach for lengthN≥ 9. The arithmetic operations can be saved from 19% to 29% forN= 3mvarying from 9 to 243 and from 17% to 29% forN= 3×2mvarying from 12 to 384. Furthermore, the new approach can be easily implemented. Huazhong Shu, Jiasong Wu, Lotfi Senhadji |
IEEE Signal Process. Lett. | 2 |
| 2009 | New Fast Algorithm for Modulated Complex Lapped Transform With Sine Windowing FunctionabstractA novel algorithm for fast computation of the modulated complex lapped transform (MCLT) with sine windowing function is presented. For the MCLT of length-2Minput data sequence, the proposed algorithm is based on computing a length-2Mtype-II generalized discrete Hartley transform. Comparison with existing algorithms shows that the proposed method achieves the minimal number of arithmetic operations. Huazhong Shu, Jiasong Wu, Lotfi Senhadji, Limin Luo 0001 |
IEEE Signal Process. Lett. | 2 |
| 2008 | Radix-2 algorithm for the fast computation of type-III 3-D discrete W transform
Huazhong Shu, Jiasong Wu, Lotfi Senhadji, Limin Luo 0001 |
Signal Process. | 2 |
| 2008 | A fast algorithm for the computation of 2-D forward and inverse MDCT
Jiasong Wu, Huazhong Shu, Lotfi Senhadji, Limin Luo 0001 |
Signal Process. | 1 |