Jianguo Wei

dblp:89/4446 · DBLP profile ↗
← Back
122ranked-venue papers
8as first author
83since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 72 · 7 first-author · 37 since 2021Artificial intelligence and machine learning · 53 · 3 first-author · 33 since 2021Databases, data management, data science and information retrieval · 11 · 11 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 7 since 2021Computer networks · 4 · 4 since 2021Security and privacy · 4 · 3 since 2021Systems, architecture and hardware · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 since 2021
YearPublicationVenuePosition
2026 EA-VAE: Learning to Reconstruct Dysarthric Speech via Variational Autoencoder with Encoding Alignment
abstract
Dysarthric speech reconstruction (DSR) aims to enhance the intelligibility of dysarthric speech. Compared with normal speech, the dysarthric speech is characterized by its pathological features, including discontinuous pronunciation, slow speech, hoarseness, and improper pauses. Significant disparities in the feature space between normal and dysarthric speech may result in suboptimal speech reconstruction, thereby degrading speech intelligibility. To enhance the reconstruction ability of speech feature spaces, this paper proposes a DSR model named the Encoding-Aligned Variational Autoencoder (EA-VAE). By incorporating alignment modules of frame-level embedding features, prior distributions, and duration into the encoder of the VAE, the model explicitly aligns the dysarthric speech encoding with a representation of the parallel normal speech. A shared decoder is then used to generate speech with improved intelligibility. Experimental results on the UASpeech benchmark confirm that EA-VAE achieves state-of-the-art performance, with a 31.7% relative word error rate reduction and the highest subjective MOS score (4.48), thoroughly validating the effectiveness and advancements of the proposed method in dysarthric speech reconstruction.
Daipeng Zhang 0001, Wenhuan Lu, Xianghu Yue, Hongcheng Zhang, Jianguo Wei
AAAI5
2026 HalluAudio: A Comprehensive Benchmark for Hallucination Detection in Large Audio-Language Models
abstract
Large Audio-Language Models (LALMs) have recently achieved strong performance across various audio-centric tasks.However, hallucination, where models generate responses that are semantically incorrect or acoustically unsupported, remains largely underexplored in the audio domain.Existing hallucination benchmarks mainly focus on text or vision, while the few audio-oriented studies are limited in scale, modality coverage, and diagnostic depth.We therefore introduce HalluAudio, the first large-scale benchmark for evaluating hallucinations across speech, environmental sound, and music.HalluAudio comprises over 5K humanverified QA pairs and spans diverse task types, including binary judgments, multi-choice reasoning, attribute verification, and open-ended QA.To systematically induce hallucinations, we design adversarial prompts and mixed-audio conditions.Beyond accuracy, our evaluation protocol measures hallucination rate, yes/no bias, error-type analysis, and refusal rate, enabling a fine-grained analysis of LALM failure modes.We benchmark a broad range of open-source and proprietary models, providing the first large-scale comparison across speech, sound, and music.Our results reveal significant deficiencies in acoustic grounding, temporal reasoning, and music attribute understanding, underscoring the need for reliable and robust LALMs.
Feiyu Zhao, Wenhuan Lu, Daipeng Zhang 0001, Xianghu Yue, Jianguo Wei
ACL (1)6
2026 Uncertainty-Gated Generative Compression for Structure-Preserving Multimedia Retrieval
abstract
Large-scale multimedia retrieval systems increasingly rely on compressed visual representations for storage and transmission efficiency. However, aggressive lossy compression risks degrading the structural cues (edges, boundaries, and semantic layouts) on which retrieval models depend. Generative image compression can synthesize perceptually plausible details at low bitrates, yet existing methods separate structure-critical from synthesizable information either statically or implicitly, leaving retrieval-relevant structure unprotected. We introduce Uncertainty-Gated Variational Inference (UG-VI), a framework that embeds deterministic gating into hierarchical VAEs. Its ELBO derivation yields a decoupling KL term that explicitly penalizes withholding predictable, structure-critical information from transmission. The resulting codec, Uncertainty-gated Structural Compression (USC), ranks latent elements by hyperprior-predicted uncertainty (available at the decoder without side information), transmitting low-uncertainty structural latents and synthesizing high-uncertainty stochastic details from the conditional prior. On CLIC2020, USC achieves 41.38% FID BD-Rate reduction over MS-ILLM while maintaining distortion (PSNR BD-Rate: \(-1.13\%\)). Retrieval-after-compression evaluation on Oxford5k and RParis6k shows that USC preserves downstream retrieval accuracy (mAP) more effectively than both traditional and generative baselines, validating uncertainty-guided structure preservation as a retrieval-aware compression strategy.
Jianguo Wei
ICMR3
2026 Integrated Mixture of Neighborhood and Community Experts for Graph-Based Fraud Detection
abstract
Graph-based fraud detection (GFD) aims to identify fraud nodes within graph-structured data that significantly deviate from the majority of benign nodes. However, existing graph neural networks (GNNs) often struggle in GFD scenarios due to their reliance on homophily assumption, which is frequently violated by the inherent homophily-heterophily mixture of fraud graphs. Moreover, most methods focus primarily on local topology, overlooking mesoscopic community structures, making them less efficient in detecting suspicious patterns like densely connected subgraphs. To address the aforementioned issues, we present NeCo, a novel approach that integrates mixture of neighborhood and community experts for graph-based fraud detection. Specifically, we first introduce a fraud-discriminative representation preservation mechanism from a neighborhood perspective, leveraging the empirical finding that fraud nodes tend to exhibit larger feature propagation discrepancies compared to benign nodes. We then design a community-oriented node representation module that models structural compactness among nodes, enabling the detection of suspicious topological patterns associated with fraud behaviors. By integrating these two complementary perspectives, NeCo can effectively captures both local inconsistency and global structural irregularity. Extensive experiments across five real-world datasets demonstrate the effectiveness of our proposed NeCo over state-of-the-art baselines.
Zhizhi Yu, Di Jin 0001, Dongxiao He, Wenhuan Lu, Jianguo Wei
WWW5
2026 ResiGuide: Residual-Guided Hierarchical Entropy Modeling for Edge-IoT Devices
abstract
Learned image compression (LIC) for edge Internet of Things (Edge-IoT) systems is constrained not just by rate-distortion (RD) performance, but critically by on-device encoding latency and memory. We identify a primary cause in the feature competition within mainstream entropy models, where modeling statistically diverse information in a unified space incurs computational and representational conflicts. To address this, we propose ResiGuide, a residual-guided hierarchical entropy model. Its core mechanism, asymmetric residual guidance, transforms this competition into efficient, guided collaboration. ResiGuide first employs an efficient Frequency-aware Global Context Module (FGCM) to generate a global prediction. Its prediction residual—an explicit error map isolating the global model’s blind spots—then directs a Residual-guided Local Context Module (RLCM) for targeted refinement. This decoupled architecture is evaluated in a CPU-only ARM Cortex-A-class edge setting: FGCM employsO(NlogN) frequency-domain operations instead ofO(N2) attention, avoiding scattered memory access, while RLCM operates solely on residuals with narrower activation ranges that are more quantization-friendly. Experiments on public benchmarks validate ResiGuide’s dual advantages: achieving a 0.73 dB BD-PSNR gain over VTM-17.0 on Kodak while reducing encoding latency and peak memory by 44% and 63% respectively against a Transformer-based baseline. This efficiency translates to smaller bitstreams that indicate lower payload-dominated latency in constrained-uplink experiments at 200 to 1000 kbps.
Jianguo Wei
IEEE Internet Things J.3
2026 Spatial-Temporal Fuzzy Logic-Driven Topological Decoupling Mechanism With Balanced Supply and Demand in Data and Energy for BSSN
abstract
To meet the sustainable connectivity requirements of 6G terrestrial networks, energy self-sufficiency and large-scale scalability have emerged as critical challenges. The Simultaneous Wireless Information and Power Transfer (SWIPT) technology provides a promising technical route for Battery-free SWIPT-enabled Sensor Networks (BSSN). However, due to the dynamic coupling and uncertainty inherent in spatial-temporal data and energy flows, existing solutions struggle to balance scalability and energy sustainability. To tackle this issue, this paper proposes a Spatial-temporal Fuzzy logic-driven Topological Decoupling mechanism with Balanced Supply and Demand (SFTD-BSD). A Multidimensional Spatial-Temporal Fuzzy prediction (MSTF-prediction) model combining Analytic Hierarchy Process (AHP) and Sparse Bayesian Learning (SBL) is developed to predict data and energy flows and capture spatial-temporal correlations. An Optimized Case-based Reasoning Inference Rule Generation (OCR-IRG) method is designed to enable online learning and adaptive inference, thereby smoothing and suppressing prediction errors. A lightweight fuzzy output scheme is further proposed to achieve robust Master Node (MN) election, thereby realizing dynamic topology decoupling and energy self-sustainability with low computational overhead. Extensive experiments demonstrate that SFTD-BSD significantly outperforms LEACH, LEACH-R, HHCA, and GWOA-CH in energy sustainability and scalability, validating its effectiveness for energy-harvesting BSSN and future 6G terrestrial networks.
Deyu Lin, Chan Su, Jianguo Wei, Yong Liang Guan 0001
IEEE J. Sel. Areas Commun.5
2026 Text-to-graph query using semantic subgraph retrieval
Yongzhe Jia, Xin Wang 0030, Jianguo Wei, Yurong Qian, Wushour Slamu
Knowl. Based Syst.7
2026 Listening for "You": Enhancing Speech Image Retrieval via Target Speaker Extraction
abstract
Image retrieval using spoken language cues has emerged as a promising direction in multimodal perception, yet leveraging speech in multi-speaker scenarios remains challenging. We propose a novel Target Speaker Speech-Image Retrieval task and a framework that learns the relationship between images and multi-speaker speech signals in the presence of a target speaker. Our method integrates pre-trained self-supervised audio encoders with vision models via target speaker-aware contrastive learning, conditioned on a Target Speaker Retrieval Extractor (TSRE) module. This method enables the extraction of semantic content from the target speaker's speech and aligns it with images representing the corresponding semantic meaning. Experiments on SpokenCOCO2Mix and SpokenCOCO3Mix show that TSRE significantly outperforms existing methods, achieving 36.3% and 29.9% Recall@1 in 2- and 3-speaker scenarios, respectively-substantial improvements over single-speaker baselines and state-of-the-art models. Our approach demonstrates potential for real-world deployment in assistive robotics and multimodal interaction systems.
Jianguo Wei, Wenhuan Lu, Xinyue Song, Xianghu Yue
IEEE Signal Process. Lett.2
2026 Domain Adaptation for Speaker Verification Using Optimal Transport With Pseudo Label
abstract
Domain gap often degrades the performance of speaker verification (SV) systems when the statistical distributions of training data and real-world test speech are mismatched. Channel variation is a primary factor causing this gap, including bandwidth changes, background noise and encoding, etc. Although various domain adaptation algorithms could be applied to handle this domain gap problem, most algorithms could not take the complex distribution structure in domain alignment with discriminative learning. In this paper, we propose a novel unsupervised domain adaptation method for speaker verification, i.e., Joint Partial Optimal Transport with Pseudo Label (JPOT-PL), to alleviate the domain mismatch problem. Leveraging the geometric-aware distance metric of optimal transport in distribution alignment and speaker consistency in speech distribution, we further design a pseudo label-based discriminative learning where the pseudo label can be regarded as a new type of speaker label derived from the optimal coupling. With the JPOT-PL, we carry out experiments on the SV channel and lingual domain adaptation with VoxCeleb, LibriSpeech, CNCeleb, and AISHELL-2. Experiments show our method reduces EER by up to 30% compared with several state-of-the-art domain adaptation algorithms.
Jianguo Wei, Wenhuan Lu, Lei Li 0050, Xugang Lu
IEEE Trans. Inf. Forensics Secur.2
2026 CoNR-Miner: Self-Adaptive Co-Occurrence Nonoverlapping Sequential Rule Mining
Yan Li 0087, Mengyao He, Jianguo Wei, Youxi Wu
IEEE Trans. Knowl. Data Eng.3
2026 Domain Adaptive Multiple Instance Self-Training for Intraoperative Anomaly Detection
abstract
Intraoperative anomalies cause deviations from the ideal surgical workflow, heightening the risk of consequential errors and complications. Their reliable recognition has traditionally relied on continuous surgeon monitoring, yet automated anomaly detection systems are now indispensable for the safe advancement of assistive and autonomous surgery. However, existing approaches struggle with domain shifts across surgical platforms and unpredictable scenarios in deformable surgical environments. To address this, we propose DA-MIST, a Domain Adaptive Multiple Instance Self-Training framework for weakly supervised anomaly detection. DA-MIST adopts a two-stage training strategy that combines multiple instance learning with self-training, enhanced by a scene-decoupled memory mechanism that disentangles state-irrelevant scene variations from memory banks, preserving only state-discriminative features for robust anomaly identification. Additionally, a state-aware dual-branch attention module integrates Gaussian dynamic and global self-attention for effective temporal reasoning. Evaluated on our newly compiled large-scale endoscopic video dataset encompassing seven representative anomalies, DA-MIST demonstrates strong adaptability across heterogeneous surgical domains, consistently reducing false alarms and enhancing anomaly localization accuracy. Our code and dataset will be available at: https://github.com/iamziang/DA-MISThttps://github.com/iamziang/DA-MIST.
Jianchang Zhao, Jianguo Wei
IEEE Trans. Medical Imaging5
2025 Dynamic Neighborhood Modeling via Node-Subgraph Contrastive Learning for Graph-Based Fraud Detection
abstract
Fraud detection that aims to discern frauds from the majority of benigns has become an increasingly prominent research field. Recently, Graph Neural Networks (GNNs) have been widely applied in graph-based fraud detection due to their outstanding data analysis and mining capabilities. However, owing to the inherent homophily-heterophily mixture and class imbalance of fraud graphs, most GNNs with homophily assumption inevitably suffer from local abnormal signal loss during information propagation, posing significant challenges in situations where frauds are rare and valuable. To address the aforementioned issues, we present a novel dynamic neighborhood modeling via node-subgraph contrastive learning for graph-based fraud detection, dubbed DCL-GFD. Specifically, we first design a node abnormality estimation module from the perspective of feature, which analyses the likelihood of a node belonging to fraud or benign by comparing the feature similarity between the target node and its corresponding subgraph. We then present a dynamic neighborhood modeling mechanism guided by the abnormal probability of a node to adaptively group and aggregate neighborhood information. By this means, the target node can effectively aggregate the neighbor information from the perspective of fraud or benign, thereby preserving as much fraud characteristics that occupy minority population as possible. Extensive experiments across four real-world fraud detection datasets demonstrate the superiority and effectiveness of our proposed DCL-GFD over state-of-the-art baselines.
Zhizhi Yu, Chundong Liang, Xinglong Chang, Dongxiao He, Di Jin 0001, Jianguo Wei
AAAI6
2025 Continual Unsupervised Domain Adaptation for Audio Deepfake Detection
abstract
Audio deepfake detection (ADD) aims to verify the authenticity of audio. However, its performance declines sharply when facing significant domain discrepancies caused by unknown datasets. Unsupervised domain adaptation (UDA) has been applied to mitigate domain mismatch. However, as generative models evolve, existing UDA methods struggle with catastrophic forgetting when facing continuously emerging spoofing methods. To address this challenge, we introduce continual UDA for ADD, which involves sequentially training across multiple target domains with continual learning. We propose a causality-distillation-based continual domain adversarial training framework for continual UDA, called CD-DAT. Specifically, we employ the domain adversarial training (DAT) framework to learn both spoofing-discriminative and domain-invariant deep features. In addition, we design a continual learning algorithm utilizing causality distillation to capture the mapping between utterances and classes, effectively mitigating forgetting and maintaining generalization. Experiments demonstrated that CD-DAT improved detection performance across all domains, confirming its memory stability and learning plasticity.
Xiaohuan Chen, Wenhuan Lu, Ruiteng Zhang, Junhai Xu, Xugang Lu, Lin Zhang 0054, Jianguo Wei
ICASSP7
2025 Neighborhood Attention Transformer with Progressive Channel Fusion for Speaker Verification
abstract
Transformer-based architectures for speaker verification typically require more training data than ECAPA-TDNN. Therefore, recent work has generally been trained on VoxCeleb1&2. We propose a backbone network based on self-attention, which can achieve competitive results when trained on VoxCeleb2 alone. The network alternates between neighborhood attention and global attention to capture local and global features, then aggregates features of different hierarchical levels, and finally performs attentive statistics pooling. Additionally, we employ a progressive channel fusion strategy to expand the receptive field in the channel dimension as the network deepens. We trained the proposed PCF-NAT model on VoxCeleb2 and evaluated it on VoxCeleb1 and the validation sets of VoxSRC. The EER and minDCF of the shallow PCF-NAT are on average more than 20% lower than those of similarly sized ECAPA-TDNN. Deep PCF-NAT achieves an EER lower than 0.5% on VoxCeleb1-O. We have also released the code1.
Jianguo Wei
ICASSP2
2025 You Only Speak Once to See
abstract
Grounding objects in images using visual cues is a well-established approach in computer vision, yet the potential of audio as a modality for object recognition and grounding remains underexplored. We introduce YOSS, "You Only Speak Once to See," to leverage audio for grounding objects in visual scenes, termed Audio Grounding. By integrating pre-trained audio models with visual models using contrastive learning and multi-modal alignment, our approach captures speech commands or descriptions and maps them directly to corresponding objects within images. Experimental results indicate that audio guidance can be effectively applied to object grounding, suggesting that incorporating audio guidance may enhance the precision and robustness of current object grounding methods and improve the performance of robotic systems and computer vision applications. This finding opens new possibilities for advanced object recognition, scene understanding, and the development of more intuitive and capable robotic systems.
Jianguo Wei, Wenhuan Lu
ICASSP2
2025 Weak Semantic-Guided Entropy Model for Image Compression
abstract
Neural Image Codecs (NICs) have made significant strides, but current entropy models struggle to efficiently leverage semantic redundancies and capture cross-dimensional correlations. We propose a novel weak semantic-guided entropy model (WSEM) to address these limitations by introducing the concept of "weak semantic", a flexible, self-learned representation that captures multi-level semantic correlations without external supervision. Acting as a dynamic bridge across pixels, channels and features, weak semantic enables the coordination of multi-dimensional redundancies during compression. WSEM introduces two key innovations: (1) a Parallel Weak Semantic Prior (PWSP) module that extracts multi-level priors from fine-grained and global semantic correlations, and (2) an Adaptive Conditional Entropy Modeling (ACEM) method that adjusts probability distributions based on these priors. Experimental results show that WSEM outperforms the state-of-the-art (SOTA) traditional codec VTM-17.0 by 0.3 dB in rate-distortion performance, while being 14% faster than the SOTA NIC, demonstrating a superior balance between compression efficiency and speed.
Jianguo Wei
ICME2
2025 Attribute Association Driven Multi-Task Learning for Session-based Recommendation
abstract
Session-based Recommendation (SBR) aims to predict users’ next interaction based on their current session without relying on long-term profiles. Despite its effectiveness in privacy-preserving and real-time scenarios, SBR remains challenging due to limited behavioral signals. Prior methods often overfit co-occurrence patterns, neglecting semantic priors like item attributes. Recent studies have attempted to incorporate item attributes (e.g., category) by assigning fixed embeddings shared across all sessions. However, such approaches suffer from two key limitations: 1) Static attribute encoding fails to reflect semantic shifts under different session contexts. 2) Semantic misalignment between attribute and item ID embeddings. To address these issues, we propose attribute association driven multi-task learning for SBR, dubbed A²D-MTL. It explicitly models item categories using cross-session context to capture user potential interests and designs an adaptive sparse attention mechanism to suppress noise. Experimental results on three public datasets demonstrate the superiority of our method in recommendation accuracy (P@20) and ranking quality (MRR@20), validating the model’s effectiveness.
Zhizhi Yu, Dongxiao He, Liang Yang 0002, Jianguo Wei, Di Jin 0001
IJCAI5
2025 DNVC-FC: A Low-Latency Distributed Neural Video Codec for Resource-Constrained Multimedia Applications
abstract
High-quality, low-latency video compression is essential for real-time multimedia applications, particularly in resource-constrained edge computing scenarios and large-scale systems. However, improving rate-distortion (RD) performance in neural video codecs (NVCs) often increases encoding complexity, hindering their adoption in latency-sensitive applications such as real-time video retrieval, robot navigation, and large-scale video indexing. To address this, we propose a novel distributed neural video codec (DNVC) that significantly improves RD performance while reducing encoder-side complexity. Our approach introduces a novel feature-channel conditional coding paradigm, integrating two key components: (1) a feature-channel conditioned entropy model that leverages implicit feature extraction to capture complex patterns and exploits cross-channel dependencies for efficient compression; (2) a high-precision side information generator that enables low-latency encoding and enhances decoder-side information quality by leveraging multi-frame reference. Experimental results show that our DNVC outperforms current state-of-the-art (SOTA) distributed video codecs and several NVCs in RD performance. Specifically, our codec achieves an average PSNR improvement of 1.2 dB compared to the current SOTA DNVC and 2 dB more than the widely-used H.264 codec. In low-latency scenarios, our method achieves a 5.5× to 14.3× speedup in encoding compared to previous SOTA NVCs.
Jianguo Wei
ICMR2
2025 LLGformer: Learnable Long-range Graph Transformer for Traffic Flow Prediction
abstract
Traffic prediction plays a pivotal role in intelligent transportation systems. Most existing studies only predict traffic flow for a specific time period based on traffic data from a short period, such as an hour, overlooking the influence of periodicity present in traffic data. Moreover, most of the existing advanced methods rely on manually constructed spatio-temporal graphs for joint modeling, or use pure spatial and pure temporal modules to separately model spatial and temporal features, which limits the learning of complex spatio-temporal patterns in traffic data due to structural inadequacies in the model. To address these issues, we propose a novel approach by constructing a learnable long-range spatio-temporal graph, which can better capture complex patterns in traffic data. We introduce a new model, LLGformer, which improves upon traditional Transformer-style models, facilitating more efficient learning of traffic flow data by integrating long-range historical information. Leveraging attention mechanisms on a spatiotemporal graph enables direct interaction of information across different time slices and locations. Additionally, we propose two optimization strategies to further boost the speed of training and inference. Extensive experiments on four real-world datasets show that the new model significantly outperforms state-of-the-art methods.
Di Jin 0001, Cuiying Huo, Dongxiao He, Jianguo Wei, Philip S. Yu
WWW5
2025 Integrated registration and utility of mobile AR Human-Machine collaborative assembly in rail transit
Jiu Yong, Jianguo Wei, Xiaomei Lei, Yangping Wang, Wenhuan Lu
Adv. Eng. Informatics2
2025 Synergizing multimodal data and fingerprint space exploration for mechanism of action prediction
abstract
MOTIVATION: Effective computational methods for predicting the mechanism of action (MoA) of compounds are essential in drug discovery. Current MoA prediction models mainly utilize the structural information of compounds. However, high-throughput screening technologies have generated more targeted cell perturbation data for MoA prediction, a factor frequently disregarded by the majority of current approaches. Moreover, exploring the commonalities and specificities among different fingerprint representations remains challenging. RESULTS: In this paper, we propose IFMoAP, a model integrating cell perturbation image and fingerprint data for MoA prediction. Firstly, we modify the Res-Net to accommodate the feature extraction of five-channel cell perturbation images and establish a granularity-level attention mechanism to combine coarse- and fine-grained features. To learn both common and specific fingerprint features, we introduce an FP-CS module, projecting four fingerprint embeddings into distinct spaces and incorporating two loss functions for effective learning. Finally, we construct two independent classifiers based on image and fingerprint features for prediction and for weighting the two prediction scores. Experimental results demonstrate that our model achieves highest accuracy of 0.941 when using multimodal data. The comparison with other methods and explorations further highlights the superiority of our proposed model and the complementary characteristics of multimodal data. AVAILABILITY AND IMPLEMENTATION: The source code is available at https://github.com/ s1mplehu/IFMoAP. The raw image data of Cell Painting can be accessed from Figshare (https://doi.org/10.17044/scilifelab.21378906).
Kaimiao Hu, Jianguo Wei, Changming Sun, Jie Geng 0001, Leyi Wei, Ran Su
Bioinform.2
2025 Efficient dehazing network based on mix structure for single image with uneven haze distribution
Kangle Yuan, Jianguo Wei, Wenhuan Lu
Eng. Appl. Artif. Intell.2
2025 Qibo: A Large Language Model for traditional Chinese medicine
Yongzhe Jia, Xin Wang 0030, Heyi Zhang, Zhaopeng Meng, Pengwei Zhuang, Jianguo Wei
Expert Syst. Appl.12
2025 Research on speech synthesis technology based on Tibetan rhythmic features
abstract
Text-to-speech(TTS) synthesis technology is one of the core technologies in the field of human–computer interaction, playing an important role in this area. This article starts from the theory of Tibetan grammar and the phonetic characteristics of Tibetan language, and designs a rhythm boundary automatic annotation method for Tibetan text and acoustic features based on the phonetic characteristics of Tibetan language. By predicting the rhythm structure hierarchy through a rhythm prediction model, and after modeling Tibetan rhythm on the acoustic model, the Tibetan synthesis model based on the improved Tacotron2 is constructed to obtain the final synthetic Tibetan speech. Experimental results show that the deep model for Tibetan speech synthesis , which utilizes Tibetan rhythm features, can further improve the naturalness and comprehensibility of the synthesized speech .
Kuntharrgyal Khysru, Yangzom, Jianguo Wei
Expert Syst. Appl.4
2025 Abnormal Behavior Detection Based on D-S Evidence Theory for Air-Ground-Integrated Vehicular Networks
abstract
The advancement of intelligent connected vehicles and aerial computing has garnered extensive attention from scholars worldwide. In particular, high-altitude platforms (HAPs) and autonomous aerial vehicles (AAVs) have emerged as effective tools to extend vehicular network coverage and enhance real-time monitoring capabilities. This study, grounded in the data from the Yizhuang Intelligent Connected Autonomous Driving Demonstration Zone, introduces an air-ground integrated vehicular network model for abnormal driving behavior detection using D-S evidence theory. The model integrates aerial computing and vehicular networks in a cohesive manner, scrutinizes the data associated with routine driving behaviors, and integrates the outcomes of various analyses through evidence theory at the decision-making level, culminating in the estimation of the probability of abnormal driving behaviors. Simulation experiments results demonstrate that this algorithm not only enhances the precision of abnormal behavior identification to a remarkable 97.1% but also significantly accelerates the detection process and improves the robustness.
Jian Jun Zeng, Han-Chieh Chao, Jianguo Wei
IEEE Internet Things J.3
2025 Self-distillation-based domain exploration for source speaker verification under spoofed speech from unknown voice conversion
Xinlei Ma, Ruiteng Zhang, Jianguo Wei, Xugang Lu, Junhai Xu, Lin Zhang 0054, Wenhuan Lu
Speech Commun.3
2025 SHDA: Sinkhorn Domain Attention for Cross-Domain Audio Anti-Spoofing
abstract
Audio anti-spoofing algorithms struggle with fake samples from unseen spoofing techniques, even when trained with diverse data sets or data augmentation strategies. Unsupervised domain adaptation (UDA) algorithms have the potential to mitigate this challenge. Typically, UDA assumes that the source and target domains are distinct distributions with clear boundaries and seeks to align model representations between them. However, in anti-spoofing, various spoofing algorithms could cause the distributions of the generated samples to overlap, resulting in unclear domain boundaries. This hinders UDA algorithms from effectively measuring and aligning domain discrepancies. Moreover, forcibly aligning samples with significant discrepancies could diminish the model’s discriminative capability. To solve this problem, we propose a domain attention algorithm with optimal transport (OT), termed Sinkhorn Domain Attention (SHDA). Unlike traditional attention mechanisms, SHDA identifies the optimal transfer plan by analyzing the global probability differences among cross-domain samples. Specifically, we first extract audio representations from various domains to compute the overall cost matrix between the source and target domains. Next, we employ Sinkhorn’s iteration to calculate the OT coupling matrix, where cross-domain samples with minor differences receive higher transfer weights, while those with substantial differences receive lower weights. Finally, we use the coupling and cost matrices to compute the adaptation loss, effectively transferring the anti-spoofing model from multiple sources to the target domain. We conducted eight cross-domain experiments using eleven well-known anti-spoofing corpora. The results indicate that our label-free SHDA surpassed the state-of-the-art model by 40%.
Ruiteng Zhang, Jianguo Wei, Xugang Lu, Lin Zhang 0054, Di Jin 0001, Junhai Xu, Wenhuan Lu
IEEE Trans. Inf. Forensics Secur.2
2025 BFGTP: A BERT-Guided Two-Stage Molecular Representation Learning Framework for Toxicity Prediction
abstract
Accurate prediction of molecular toxicity is vital for drug development. Most mainstream methods rely on fingerprints or graph-based feature extraction, the emergence of large language models (LLMs) offers new prospects for molecular representation learning in toxicity prediction. Although several studies attempt to leverage LLMs to integrate molecular sequence data for pretraining molecular representations, certain limitations remain. Current LLM-based approaches usually utilize solely on class embedding features, overlooking the rich information in sequence embedding. Moreover, integrating pre-trained molecular representations with multi-modal molecular data may further enhance performance in toxicity prediction. To address these challenges, we propose BFGTP, a BERT-guided two-stage molecular representation learning framework for toxicity prediction. Firstly, we design independent encoders for molecular descriptions of three modalities, where the fingerprint encoder with dual level attention mechanisms effectively integrates multi-category fingerprints. Then, the two-stage guide strategy is introduced to fully utilize the prior knowledge of LLMs, employing contrastive learning to align and fuse the tri-modal representations and knowledge distillation to align predicted value distributions. BFGTP ultimately combines fingerprint and graph representations to predict molecular toxicity. Experiments on seven toxicity datasets show that BFGTP outperforms baselines, achieving the highest AUC on five datasets and the best average performance across five evaluation metrics. Ablation studies, t-SNE visualization and case study confirm the effectiveness of BFGTP's components and its ability to capture meaningful molecular representations.
Kaimiao Hu, Yuan He 0016, Jianguo Wei, Changming Sun, Jie Geng 0001, Leyi Wei, Ran Su
IEEE J. Biomed. Health Informatics3
2024 Graphologue: Bridging RDBMS and Graph Databases with Natural Language Interfaces
Yongzhe Jia, Jianguo Wei, Xin Wang 0030, Xintian Zuo, Yuxuan Yang 0006
DASFAA (7)2
2024 Breaking the Corpus Bottleneck for Multi-dialect Speech Recognition with Flexible Adapters
Tengyue Deng, Jianguo Wei, Wenjun Ke 0001, Xiaokang Yang 0003, Wenhuan Lu
ICANN (7)2
2024 SSR-GPCsT: Deep Learning Models Based on Functional Connectivity Maps in Autism Research
abstract
Autism is a neurodevelopmental disorder characterized by difficulties in social interaction, communication, and sensory sensitivity. Functional magnetic resonance imaging (fMRI) is a commonly used brain imaging technique to obtain functional connectivity information in individuals with autism. However, the heterogeneity of functional MRI data from different imaging centers poses a modeling challenge, and the lack of interpretability in the models hinders the identification of biomarkers. To address these issues, this study proposes a deep learning model based on a dynamic functional connectivity matrix that extracts shared features across multiple centers, mitigates the multicentre problem, enhances the representation of the data, and combines features from multiple pathways to capture a comprehensive view of brain connectivity patterns. In addition, two complementary approaches are designed to enhance the interpretability of the model and further explore the presence of biomarkers by analyzing changes in important features of the model. The code will be released soon.
Jiacheng Hao, Junhai Xu, Jianguo Wei
ICASSP4
2024 Evaluation of an Improved Ultrasonic Imaging Helmet for Observing Articulatory Data
abstract
Ultrasonic imaging is one of the most popular methods for tracking tongue motion. Imaging plane shift and contact variation are crucial factors affecting the consistency of the obtained ultrasonic images. To solve this issue, researchers proposed many different helmets. In this study, we propose an evaluation framework to quantitatively assess the helmet’s imaging plane shift and contact variation. The framework is applied to one helmet we designed. Compared to the baseline, the imaging plane shift using our helmet is less than 0.1°, and the contact variation is reduced by 68.5%, decreasing the mean difference of the extracted contours by 50.3% and increasing the contrast and sharpness of the obtained images by 9.0% and 28.4%. The proposed framework provides an approach for quantitatively proofing helmet structure with data accuracy and consistency.
Jianguo Wei, Qiang Fang 0003, Xugang Lu
ICASSP2
2024 EEG-Based Fast Auditory Attention Detection in Real-Life Scenarios Using Time-Frequency Attention Mechanism
abstract
Auditory attention detection (AAD) based on electroencephalogram (EEG) helps recognize the target speaker in a cocktail party scenario, advancing auditory brain-computer interface development. Previous EEG studies on AAD were largely based on data collected in laboratory settings. In this study, we investigated the AAD with EEG data collected when subjects were walking and sitting in real-life scenarios. To improve the detection accuracy, we proposed the time-frequency attention mechanism to the convolution neural network on EEG data. Experimental results show that the proposed model outperforms the state-of-the-art models, with an accuracy of 98.1% on a decision window of 2s. When we used a 0.1s time window for fast decoding, the accuracy remained at 91.8%, suggesting the potential for real application. Further study on ablation experiments demonstrates the effectiveness of the proposed time-frequency attention mechanism. Analysis of the key EEG features indicates that the β band plays a vital role in AAD.
Zhuang Xie, Jianguo Wei, Wenhuan Lu, Chunli Wang, Gaoyan Zhang
ICASSP2
2024 Self-Supervised Domain Exploration with an Optimal Transport Regularization for Open Set Cross-Domain Speech Emotion Recognition
abstract
In the tasks of domain adaptation (DA) for speech emotion recognition (SER), self-supervised learning (SSL) algorithms could effectively explore domain and structural information from target domain samples, thereby mitigating domain discrepancies. However, in a general setting, when the target domain contains emotions that are never observed in the source domain, namely in open-set DA, existing SSL-based DA methods cannot maintain the robustness because of the interference of the extra unknown classes. To address this challenge, we propose the self-supervised domain exploration with an optimal transport (OT) regularization (SDEOTR) algorithm. First, we integrate the SSL algorithm into the SER model to mitigate the domain differences. Further, we categorize target domain samples into known and unknown groups based on the network’s prediction confidence. Finally, we employ OT to maximize the global probability distance between the two groups, aiming to decrease the impact of unknown emotions on the SER model. Cross-domain SER experimental results showed that our label-free SDEOTR significantly improved the performance of existing adaptive SER algorithms in open-set scenarios.
Ruiteng Zhang, Jianguo Wei, Xugang Lu, Wenhuan Lu, Di Jin 0001, Junhai Xu
ICASSP2
2024 Generalized Taxonomy-Guided Graph Neural Networks
Yu Zhou 0050, Di Jin 0001, Jianguo Wei, Dongxiao He, Zhizhi Yu, Weixiong Zhang
IJCAI3
2024 Synergistic Dual Spatial-aware Generation of Image-to-text and Text-to-image
abstract
In the visual spatial understanding (VSU) field, spatial image-to-text (SI2T) and spatial text-to-image (ST2I) are two fundamental tasks that appear in dual form. Existing methods for standalone SI2T or ST2I perform imperfectly in spatial understanding, due to the difficulty of 3D-wise spatial feature modeling. In this work, we consider modeling the SI2T and ST2I together under a dual learning framework. During the dual framework, we then propose to represent the 3D spatial scene features with a novel 3D scene graph (3DSG) representation that can be shared and beneficial to both tasks. Further, inspired by the intuition that the easier 3D$\to$image and 3D$\to$text processes also exist symmetrically in the ST2I and SI2T, respectively, we propose the Spatial Dual Discrete Diffusion (SD$^3$) framework, which utilizes the intermediate features of the 3D$\to$X processes to guide the hard X$\to$3D processes, such that the overall ST2I and SI2T will benefit each other. On the visual spatial understanding dataset VSD, our system outperforms the mainstream T2I and I2T methods significantly. Further in-depth analysis reveals how our dual learning strategy advances.
Yu Zhao 0043, Hao Fei 0001, Xiangtai Li, Libo Qin 0004, Jiayi Ji, Hongyuan Zhu 0002, Meishan Zhang, Min Zhang 0005, Jianguo Wei
NeurIPS9
2024 Distillation-Based Feature Extraction Algorithm For Source Speaker Verification
abstract
Automatic speaker verification (ASV) systems face significant challenges when exposed to spoofing attacks, necessitating robust countermeasures. In this work, we focus on the source speaker verification (SSV) task, which aims to identify the source speaker hidden in spoofed speech generated by voice conversion (VC) systems. We propose a distillation-based feature extraction algorithm to enhance the model’s ability to verify source speakers. Our method employs a pretraining ASV model as a teacher network and the SSV model as a student network, using bona fide speech to guide the learning process. However, the improvements were marginal, particularly on the development set, indicating the complexity and resource demands of fine-tuning the distillation parameters. Our findings underscore the inherent difficulties in SSV and highlight the need for further research to develop more effective solutions. Besides, our submission won fourth place in the 2024 Source Speaker Tracking Challenge.
Xinlei Ma, Wenhuan Lu, Ruiteng Zhang, Junhai Xu, Xugang Lu, Jianguo Wei
SLT6
2024 End-To-End Speaker Anonymization Based on Location-Variable Convolution and Multi-Head Self-Attention
abstract
Speaker anonymization, a user-centric solution for voice privacy, aims to conceal the speaker’s identity while maintaining clarity and naturalness. The prevalent approach involves cascading modules of automatic speech recognition (ASR) and text-to-speech (TTS) models for speaker anonymization through speech synthesis. However, the inherent multimodal cascade nature of this approach leads to high error rates and unclear speech due to inaccuracies propagated from the ASR system to the TTS system. To address these issues, this paper proposes an end-to-end method for achieving zero-shot speaker anonymization. This method improves the shortcomings of traditional speech models that use a large number of fixed convolution kernels to capture the internal dependencies of speech sequences. It captures the internal dependencies of speech from both long-time and long-distance perspectives by combining a network with variable kernels, namely location-variable convolutions (LVCs), with a multi-head self-attention mechanism and dynamically adjusting weights. Besides, it learns anonymized speaker features flexibly through an enhanced cycle-consistency loss, iteratively aligning speaker information for restructured speech with that of anonymized speakers indefinitely. The efficacy of our proposed speaker anonymization model was demonstrated on the English dataset VCTK.
Feiyu Zhao, Jianguo Wei, Wenhuan Lu
TrustCom2
2024 SARN: Script-Aware Recognition Network for scene multilingual text recognition
Wenjun Ke 0001, Darcy Qingzhi Hou, Yutian Liu 0003, Xinyue Song, Jianguo Wei
Expert Syst. Appl.5
2024 MVIB-DVA: Learning minimum sufficient multi-feature speech emotion embeddings under dual-view aware
Guoyan Li, Junjie Hou, Jianguo Wei
Expert Syst. Appl.4
2024 Semisupervised Medical Image Segmentation through Prototype-Based Mutual Consistency Learning
abstract
Medical image segmentation is a critical task in the healthcare field. While deep learning techniques have shown promise in this area, they often require a large number of accurately labeled images. To address this issue, semisupervised learning has emerged as a potential solution by reducing the reliance on precise annotations. Among these approaches, the student-teacher framework has garnered attention, but it is limited in its reliance solely on the teacher model for information. To overcome this limitation, we propose a prototype-based mutual consistency learning (PMCL) framework. This framework utilizes two branches that learn from each other, incorporating supervision loss and consistency loss to adapt to minor data perturbations and structural differences. By employing prototype consistency learning, we are able to achieve reliable consistency loss. Our experiments on three public medical image datasets demonstrate that PMCL outperforms other state-of-the-art methods, indicating its potential in semisupervised medical image segmentation. Our framework has the potential to assist medical professionals in enhancing their diagnoses and delivering improved patient care.
Xinqiang Wang, Wenhuan Lu, Junhai Xu, Jianguo Wei
Int. J. Intell. Syst.6
2024 Task Offloading and Resource Allocation for Fog Computing in NG Wireless Networks: A Federated Deep Reinforcement Learning Approach
abstract
Task offloading (TO) is beneficial to reducing the delay and energy consumption for the prosperity of the applications in next generation (NG) wireless networks. However, existing TO approaches are inability to exhibit low complexity and stable performance. To this end, a novel federated hierarchical deep deterministic policy gradient (FHDDPG) algorithm for TO and resource allocation (RA) is proposed in this article. To be specific, three deep deterministic policy gradient (DDPG) modules are deployed in parallel to make offloading decision on the execution mode of tasks and the proportion allocation of the transmission rate. Subsequently, a federated learning method is proposed to collaboratively train the HDDPG model by means of sharing models’ weights. Meanwhile, the delay and the energy consumption are comprehensively considered as the average system consumption, which is defined as a reward metric of FHDDPG. Finally, extensive simulations are conducted to demonstrate the effectiveness of our proposal. The experimental results indicate that the average system consumption of FHDDPG is cut down by 11.4% and 18% compare with HDDPG and DDPG, respectively, which means FHDDPG can achieve a better performance effectively.
Chan Su, Jianguo Wei, Deyu Lin, Linghe Kong, Yong Liang Guan 0001
IEEE Internet Things J.2
2024 A novel model for fall detection and action recognition combined lightweight 3D-CNN and convolutional LSTM networks
Chan Su, Jianguo Wei, Deyu Lin, Linghe Kong, Yong Liang Guan 0001
Pattern Anal. Appl.2
2024 Zero-shot voice conversion based on feature disentanglement
Jianguo Wei, Wenhuan Lu, Jianhua Tao 0001
Speech Commun.2
2024 Multi-modal co-learning for silent speech recognition based on ultrasound tongue images
Jianguo Wei, Ruiteng Zhang, Qiang Fang 0003
Speech Commun.2
2024 Unsupervised Adaptive Speaker Recognition by Coupling-Regularized Optimal Transport
abstract
Cross-domain speaker recognition (SR) can be improved by unsupervised domain adaptation (UDA) algorithms. UDA algorithms often reduce domain mismatch at the cost of decreasing the discrimination of speaker features. In contrast, optimal transport (OT) has the potential to achieve domain alignment while preserving the speaker discrimination capability in UDA applications; however, naively applying OT to measure global probability distribution discrepancies between the source and target domains may induce negative transports where samples belonging to different speakers are coupled in transportation. These negative transports reduce the SR model's discriminative power, degrading the SR performance. This paper proposes a coupling-regularized optimal transport (CROT) algorithm for cross-domain SR to reduce the negative transport during UDA. In the proposed CROT, two consecutive processing modules regularize the coupling paths for the OT solution: a progressive inter-speaker constraint (PISC) module and a coupling-smoothed regularization (CSR) module. The PISC, designed as a pseudo-label memory bank with curriculum learning, is first applied to select valid samples to guarantee that coupling samples are from the same speaker. The CSR, designed to control the information entropy of the coupling paths further, reduces the effect of negative transport in UDA. To evaluate the effectiveness of the proposed algorithm, cross-domain SR experiments were conducted under different target domains, speaker encoders, corpora, and acoustic features. Experimental results showed that CROT achieved a 50% relative reduction in equal error rates compared to conventional OT-based UDAs, outperforming the state-of-the-art UDAs.
Ruiteng Zhang, Jianguo Wei, Xugang Lu, Wenhuan Lu, Di Jin 0001, Lin Zhang 0054, Junhai Xu
IEEE ACM Trans. Audio Speech Lang. Process.2
2024 TeKo: Text-Rich Graph Neural Networks With External Knowledge
abstract
Graph neural networks (GNNs) have gained great prevalence in tackling various analytical tasks on graph-structured data (i.e., networks). Typical GNNs and their variants adopt a message-passing principle that obtains network representations by the attribute propagates along network topology, which however ignores the rich textual semantics (e.g., local word-sequence) that exist in numerous real-world networks. Existing methods for text-rich networks integrate textual semantics by mainly using internal information such as topics or phrases/words, which often suffer from an inability to comprehensively mine the textual semantics, limiting the reciprocal guidance between network structure and textual semantics. To address these problems, we present a novel text-rich GNN with external knowledge (TeKo), in order to make full use of both structural and textual information within text-rich networks. Specifically, we first present a flexible heterogeneous semantic network that integrates high-quality entities as well as interactions among documents and entities. We then introduce two types of external knowledge, that is, structured triplets and unstructured entity descriptions, to gain a deeper insight into textual semantics. Furthermore, we devise a reciprocal convolutional mechanism for the constructed heterogeneous semantic network, enabling network structure and textual semantics to collaboratively enhance each other and learn high-level network representations. Extensive experiments illustrate that TeKo achieves state-of-the-art performance on a variety of text-rich networks as well as a large-scale e-commerce searching dataset.
Zhizhi Yu, Di Jin 0001, Jianguo Wei, Yawen Li 0001, Ziyang Liu 0004, Jiawei Han 0001, Lingfei Wu 0001
IEEE Trans. Neural Networks Learn. Syst.3
2023 Local-Global Defense against Unsupervised Adversarial Attacks on Graphs
abstract
Unsupervised pre-training algorithms for graph representation learning are vulnerable to adversarial attacks, such as first-order perturbations on graphs, which will have an impact on particular downstream applications. Designing an effective representation learning strategy against white-box attacks remains a crucial open topic. Prior research attempts to improve representation robustness by maximizing mutual information between the representation and the perturbed graph, which is sub-optimal because it does not adapt its defense techniques to the severity of the attack. To address this issue, we propose an unsupervised defense method that combines local and global defense to improve the robustness of representation. Note that we put forward the Perturbed Edges Harmfulness (PEH) metric to determine the riskiness of the attack. Thus, when the edges are attacked, the model can automatically identify the risk of attack. We present a method of attention-based protection against high-risk attacks that penalizes attention coefficients of perturbed edges to encoders. Extensive experiments demonstrate that our strategies can enhance the robustness of representation against various adversarial attacks on three benchmark graphs.
Di Jin 0001, Bingdao Feng, Siqi Guo 0002, Xiaobao Wang, Jianguo Wei, Zhen Wang 0004
AAAI5
2023 Generating Visual Spatial Description via Holistic 3D Scene Understanding
abstract
Visual spatial description (VSD) aims to generate texts that describe the spatial relations of the given objects within images.Existing VSD work merely models the 2D geometrical vision features, thus inevitably falling prey to the problem of skewed spatial understanding of target objects.In this work, we investigate the incorporation of 3D scene features for VSD.With an external 3D scene extractor, we obtain the 3D objects and scene features for input images, based on which we construct a target object-centered 3D spatial scene graph (GO3D-S 2 G), such that we model the spatial semantics of target objects within the holistic 3D scenes.Besides, we propose a scene subgraph selecting mechanism, sampling topologically-diverse subgraphs from GO3D-S 2 G, where the diverse local structure features are navigated to yield spatially-diversified text generation.Experimental results on two VSD datasets demonstrate that our framework outperforms the baselines significantly, especially improving on the cases with complex visual spatial relations.Meanwhile, our method can produce more spatially-diversified generation.Code is available at https://github.com/zhaoyucs/VSD.
Yu Zhao 0043, Hao Fei 0001, Wei Ji 0008, Jianguo Wei, Meishan Zhang, Min Zhang 0005, Tat-Seng Chua
ACL (1)4
2023 HyperMatch: Knowledge Hypergraph Question Answering Based on Sequence Matching
Yongzhe Jia, Jianguo Wei, Lifan Han
DASFAA (4)2
2023 An Improved GPU Acceleration Framework for Smoothed Particle Hydrodynamics
Yuejin Cai, Jianguo Wei, Jiyou Duan, Darcy Qingzhi Hou
ICA3PP (6)2
2023 Optimal Transport with a Diversified Memory Bank for Cross-Domain Speaker Verification
abstract
Optimal transport (OT) can be applied to cross-domain adaptation in speaker verification (SV) by converting speakers' probability distributions from source to target domains. However, in scenarios involving over-massive categories (speakers) or difficult samples in discrimination, OT often has difficulty computing effective transports. To address this challenge, we propose an OT-based unsupervised domain adaptation (UDA) framework for SV, OT with a diversified memory bank, called DMB-OT, which ensures the accuracy of transfers by two strategies: (1) It regularizes the solution space of OT, which attempts to plan transformations between audio samples from the same speaker with high confidence; (2) it integrates a dynamic curriculum learning algorithm, preventing OT from calculating transport couplings based on hard-discriminative samples in the early stage of UDA. Experiments under different target domains showed that our unsupervised DMB-OT could significantly improve the performance of OT-based UDA and could even match the performance of the supervised PLDA-based adaptation.
Ruiteng Zhang, Jianguo Wei, Xugang Lu, Wenhuan Lu, Di Jin 0001, Lin Zhang 0054, Junhai Xu
ICASSP2
2023 Frequency Patterns of Individual Speaker Characteristics at Higher and Lower Spectral Ranges
Ju Zhang 0001, Yujie Chi, Kiyoshi Honda, Jianguo Wei
INTERSPEECH6
2023 SOT: Self-supervised Learning-Assisted Optimal Transport for Unsupervised Adaptive Speech Emotion Recognition
Ruiteng Zhang, Jianguo Wei, Xugang Lu, Junhai Xu, Di Jin 0001, Jianhua Tao 0001
INTERSPEECH2
2023 Transvelar Nasal Coupling Contributing to Speaker Characteristics in Non-nasal Vowels
Yujie Chi, Kiyoshi Honda, Jianguo Wei
INTERSPEECH5
2023 Constructing Holistic Spatio-Temporal Scene Graph for Video Semantic Role Labeling
abstract
As one of the core video semantic understanding tasks, Video Semantic Role Labeling (VidSRL) aims to detect the salient events from given videos, by recognizing the predict-argument event structures and the interrelationships between events. While recent endeavors have put forth methods for VidSRL, they can be mostly subject to two key drawbacks, including the lack of fine-grained spatial scene perception and the insufficiently modeling of video temporality. Towards this end, this work explores a novel holistic spatio-temporal scene graph (namely HostSG) representation based on the existing dynamic scene graph structures, which well model both the fine-grained spatial semantics and temporal dynamics of videos for VidSRL. Built upon the HostSG, we present a nichetargeting VidSRL framework. A scene-event mapping mechanism is first designed to bridge the gap between the underlying scene structure and the high-level event semantic structure, resulting in an overall hierarchical scene-event (termed ICE) graph structure. We further perform iterative structure refinement to optimize the ICE graph, e.g., filtering noisy branches and newly building informative connections, such that the overall structure representation can best coincide with end task demand. Finally, three subtask predictions of VidSRL are jointly decoded, where the end-to-end paradigm effectively avoids error propagation. On the benchmark dataset, our framework boosts significantly over the current best-performing model. Further analyses are shown for a better understanding of the advances of our methods. Our HostSG representation shows greater potential to facilitate a broader range of other video understanding tasks.
Yu Zhao 0043, Hao Fei 0001, Yixin Cao 0002, Bobo Li 0001, Meishan Zhang, Jianguo Wei, Min Zhang 0005, Tat-Seng Chua
ACM Multimedia6
2023 TMS: Temporal multi-scale in time-delay neural network for speaker verification
Ruiteng Zhang, Jianguo Wei, Xugang Lu, Wenhuan Lu, Di Jin 0001, Lin Zhang 0054, Junhai Xu, Jianwu Dang 0001
Appl. Intell.2
2023 A multi-scale multi-model deep neural network via ensemble strategy on high-throughput microscopy image for protein subcellular localization
Jiaqi Ding, Junhai Xu, Jianguo Wei, Jijun Tang, Fei Guo 0001
Expert Syst. Appl.3
2023 Rethinking text rectification for scene text recognition
Wenjun Ke 0001, Jianguo Wei, Darcy Qingzhi Hou
Expert Syst. Appl.2
2023 PINN-CDR: A Neural Network-Based Simulation Tool for Convection-Diffusion-Reaction Systems
abstract
In this paper, a discretization‐free approach based on the physics‐informed neural network (PINN) is proposed for solving the forward and inverse problems governed by the nonlinear convection‐diffusion‐reaction (CDR) systems. By embedding physical information described by the CDR system in the feedforward neural networks, PINN is trained to approximate the solution of the system without the need of labeled data. The good performance of PINN in solving the forward problem of the nonlinear CDR systems is verified by studying the problems of gas‐solid adsorption and autocatalytic reacting flow. For CDR systems with different Péclet number, PINN can largely eliminate the numerical diffusion and unphysical oscillations in traditional numerical methods caused by high Péclet number. Meanwhile, the PINN framework is implemented to solve the inverse problem of nonlinear CDR systems and the results show that the unknown parameters can be effectively recognized even with high noisy data. It is concluded that the established PINN algorithm has good accuracy, convergence, and robustness for both the forward and inverse problems of CDR systems.
Darcy Qingzhi Hou, Honghan Du, Zewei Sun, Jianping Wang 0009, Jianguo Wei
Int. J. Intell. Syst.6
2023 GSS: A group similarity system based on unsupervised outlier detection for big data computing
Wenjun Ke 0001, Jianguo Wei, Naixue Xiong, Darcy Qingzhi Hou
Inf. Sci.2
2023 A watermark detection scheme based on non-parametric model applied to mute machine voice
Yangxia Hu, Wenhuan Lu, Jianguo Wei, Junhai Xu, Maode Ma
Multim. Tools Appl.3
2023 Using attention LSGB network for facial expression recognition
Chan Su, Jianguo Wei, Deyu Lin, Linghe Kong
Pattern Anal. Appl.2
2023 Self-supervised learning based domain regularization for mask-wearing speaker verification
Ruiteng Zhang, Jianguo Wei, Xugang Lu, Wenhuan Lu, Di Jin 0001, Lin Zhang 0054, Yantao Ji, Junhai Xu
Speech Commun.2
2023 An Intelligent Classification Diagnosis Based on Blood Oxygen Saturation Signals for Medical Data Security Including COVID-19 in Industry 5.0
abstract
Obstructive sleep apnea-hypopnea syndrome (OSAHS) is gradually valued due to its high prevalence, high risk, and high mortality. Alternative to the polysomnography (PSG) diagnosis, the proposed method assesses the subject's degree of illness considering the supply chain and Industry 5.0 requirement efficiently and accurately. This article uses the blood oxygen saturation (SpO2) signal count of the number of apnea or hypoventilation events during the sleep of the subject, calculating the apnea-hypopnea index (AHI) and the subject's disease level. SpO2signals are used to extract 35-D features based on the time domain, including approximate entropy, central tendency measure, and Lempel–Ziv complexity to accelerate the diagnosis process in supply chains. The feature selection process is reduced from 35 to 7 dimensions that benefits to the implementation in the practical supply chains in Industry 5.0 by extracting the extracted features. This article applies Pearson correlation coefficient selection, based on minimum redundancy-maximum correlation algorithm selection, and a wrapper based on the backward search algorithm. The accuracy rate is 86.92%, and the specificity is 90.7% under the selected random forest classifier. A random forest classifier was used to calculate the AHI index, and a linear regression analysis was performed with the AHI index obtained from the PSG. The result reaches a 92% accuracy rate in assessing the prevalence of OSAHS, satisfying the industrial deployment.
Mingdong Zhang, Chaoyu Dong, Ming-Lang Tseng, Jianguo Wei
IEEE Trans. Ind. Informatics5
2023 Improving Multispike Learning With Plastic Synaptic Delays
abstract
Emulating the spike-based processing in the brain, spiking neural networks (SNNs) are developed and act as a promising candidate for the new generation of artificial neural networks that aim to produce efficient cognitions as the brain. Due to the complex dynamics and nonlinearity of SNNs, designing efficient learning algorithms has remained a major difficulty, which attracts great research attention. Most existing ones focus on the adjustment of synaptic weights. However, other components, such as synaptic delays, are found to be adaptive and important in modulating neural behavior. How could plasticity on different components cooperate to improve the learning of SNNs remains as an interesting question. Advancing our previous multispike learning, we propose a new joint weight-delay plasticity rule, named TDP-DL, in this article. Plastic delays are integrated into the learning framework, and as a result, the performance of multispike learning is significantly improved. Simulation results highlight the effectiveness and efficiency of our TDP-DL rule compared to baseline ones. Moreover, we reveal the underlying principle of how synaptic weights and delays cooperate with each other through a synthetic task of interval selectivity and show that plastic delays can enhance the selectivity and flexibility of neurons by shifting information across time. Due to this capability, useful information distributed away in the time domain can be effectively integrated for a better accuracy performance, as highlighted in our generalization tasks of the image, speech, and event-based object recognitions. Our work is thus valuable and significant to improve the performance of spike-based neuromorphic computing.
Qiang Yu 0005, Jialu Gao, Jianguo Wei, Kay Chen Tan, Tiejun Huang 0001
IEEE Trans. Neural Networks Learn. Syst.3
2022 Visual Spatial Description: Controlled Spatial-Oriented Image-to-Text Generation
abstract
Image-to-text tasks, such as open-ended image captioning and controllable image description, have received extensive attention for decades.Here, we further advance this line of work by presenting Visual Spatial Description (VSD), a new perspective for image-to-text toward spatial semantics.Given an image and two objects inside it, VSD aims to produce one description focusing on the spatial perspective between the two objects.Accordingly, we manually annotate a dataset to facilitate the investigation of the newly-introduced task and build several benchmark encoder-decoder models by using VL-BART and VL-T5 as backbones.In addition, we investigate pipeline and joint end-to-end architectures for incorporating visual spatial relationship classification (VSRC) information into our model.Finally, we conduct experiments on our benchmark dataset to evaluate all our models.Results show that our models are impressive, providing accurate and human-like spatial-oriented text descriptions.Meanwhile, VSRC has great potential for VSD, and the joint end-to-end architecture is the better choice for their integration.We make the dataset and codes public for research purposes.
Yu Zhao 0043, Jianguo Wei, Zhichao Lin, Yueheng Sun, Meishan Zhang, Min Zhang 0005
EMNLP2
2022 DMANET: Deep Learning-Based Differential Microphone Arrays for Multi-Channel Speech Separation
abstract
In this paper, we develop a novel differential microphone arrays network (DMANet) for solving the multi-channel speech separation problem. In DMANet we explore a neural network combined to differential microphone arrays (DMAs) beamforming technique. Specifically, a sequence of differential operation is introduced alternately into network. Based on the filter-and-sum network (FaSNet), we show how DMANet significantly improves the separation performance. Numerical experiments demonstrate that the proposed network has a clearly advantageous improvement on SI-SNR with a smaller model.
Xiaokang Yang 0003, Jianguo Wei
ICASSP2
2022 Joint and Adversarial Training with ASR for Expressive Speech Synthesis
abstract
Style modeling is an important issue and has been proposed in expressive speech synthesis. In existing unsupervised methods, the style encoder extracts the latent representation from the reference audio as style information. However, the style information extracted from the style encoder will entangle some content information, which will cause conflicts with the real input content, and the synthesized speech will be influenced. In this study, we propose to alleviate the entanglement problem by integrating Text-To-Speech (TTS) model and Automatic Speech Recognition (ASR) model with a share layer network for joint training, and using ASR adversarial training to eliminate the content information in the style information. At the same time, we propose an adaptive adversarial weight learning strategy to prevent the model from collapsing. The objective evaluation using word error rate(WER) demonstrates that our method can effectively alleviate the entanglement between style and content information. Subjective evaluation indicates that the method improves the quality of synthesized speech and enhances the ability of style transfer compared with the baseline models.
Wenhuan Lu, Longbiao Wang, Jianguo Wei
ICASSP5
2022 CS-REP: Making Speaker Verification Networks Embracing Re-Parameterization
abstract
Automatic speaker verification (ASV) systems, which determine whether two speeches are from the same speaker, mainly focus on verification accuracy while ignoring inference speed. However, in real applications, both inference speed and verification accuracy are essential. This study proposes cross-sequential re-parameterization (CS-Rep), a novel topology re-parameterization strategy for multi-type networks, to increase the inference speed and verification accuracy of models. CS-Rep solves the problem that existing re-parameterization methods are not suitable for typical ASV backbones. When a model applies CS-Rep, the training-period network utilizes a multi-branch topology to capture speaker information, whereas the inference-period model converts to a time-delay neural network (TDNN)-like plain backbone with stacked TDNN layers to achieve the fast inference speed. Based on CS-Rep, an improved TDNN with friendly test and deployment called Rep-TDNN is proposed. Compared with the state-of-the-art model ECAPA-TDNN, Rep-TDNN increases the actual inference speed by about 50% and reduces the EER by 10%. The code and trained models are available at https://github.com/zrtlemontree/CS-Rep.
Ruiteng Zhang, Jianguo Wei, Wenhuan Lu, Lin Zhang 0054, Yantao Ji, Junhai Xu, Xugang Lu
ICASSP2
2022 Double Noise Mean Teacher Self-Ensembling Model for Semi-Supervised Tumor Segmentation
abstract
Accurate tumor segmentation of tumor images can assist doctors to diagnose diseases. However, achieving very high precision in tumor segmentation requires a large amount of annotated data, which is not easy for medical image data. In this paper, we present a novel double noise mean teacher self-ensembling model for semi-supervised 2D tumor segmentation. Concretely, the network is serialized by two groups of student-teacher networks. We design an auxiliary student-teacher module to learn the consistency regularity between the unlabeled image feature maps. In order to improve the robustness of the network, we add the random Gaussian noise to the student model every time the teacher model is updated. We test our model on the small cell lung tumor dataset and CVC-ClinicDB, and our model achieves the performance of nearly fully supervised segmentation. Moreover, the performance of our method outperforms the existing semi-supervised methods in four indicators.
Junhai Xu, Jianguo Wei
ICASSP3
2022 Vocal-Tract Area Functions with Articulatory Reality for Tract Opening
Ju Zhang 0001, Jianguo Wei, Kiyoshi Honda, Tatsuya Kitamura
INTERSPEECH3
2022 A semi fragile watermarking algorithm based on compressed sensing applied for audio tampering detection and recovery
Yangxia Hu, Wenhuan Lu, Maode Ma, Qilong Sun, Jianguo Wei
Multim. Tools Appl.5
2022 wUnet: A new network used for ultrasonic tongue contour extraction
Guoyan Li, Jianguo Wei
Speech Commun.4
2022 One-shot emotional voice conversion based on feature separation
Wenhuan Lu, Xinyue Zhao, Jianguo Wei, Jianhua Tao 0001, Jianwu Dang 0001
Speech Commun.5
2022 Temporal Encoding and Multispike Learning Framework for Efficient Recognition of Visual Patterns
abstract
Biological systems under a parallel and spike-based computation endow individuals with abilities to have prompt and reliable responses to different stimuli. Spiking neural networks (SNNs) have thus been developed to emulate their efficiency and to explore principles of spike-based processing. However, the design of a biologically plausible and efficient SNN for image classification still remains as a challenging task. Previous efforts can be generally clustered into two major categories in terms of coding schemes being employed: rate and temporal. The rate-based schemes suffer inefficiency, whereas the temporal-based ones typically end with a relatively poor performance in accuracy. It is intriguing and important to develop an SNN with both efficiency and efficacy being considered. In this article, we focus on the temporal-based approaches in a way to advance their accuracy performance by a great margin while keeping the efficiency on the other hand. A new temporal-based framework integrated with the multispike learning is developed for efficient recognition of visual patterns. Different approaches of encoding and learning under our framework are evaluated with the MNIST and Fashion-MNIST data sets. Experimental results demonstrate the efficient and effective performance of our temporal-based approaches across a variety of conditions, improving accuracies to higher levels that are even comparable to rate-based ones but importantly with a lighter network structure and far less number of spikes. This article attempts to extend the advanced multispike learning to the challenging task of image recognition and bring state of the arts in temporal-based approaches to a novel level. The experimental results could be potentially favorable to low-power and high-speed requirements in the field of artificial intelligence and contribute to attract more efforts toward brain-like computing.
Qiang Yu 0005, Shiming Song 0001, Chenxiang Ma, Jianguo Wei, Shengyong Chen, Kay Chen Tan
IEEE Trans. Neural Networks Learn. Syst.4
2022 ATU: An Aggregate-Then-Update Diffusion Intelligent Estimation Scheme for Adaptive Networked Systems
abstract
For distributed estimation arising in the nonlinear least squares (NLLSs) problems over adaptive networks, where every node has the abilities of data processing and learning, only the incomplete local data are exploited by the traditional noncooperative method, thereby resulting in the degradation on estimation performance. In this article, a cooperative diffusion strategy is proposed by using a Gauss–Newton (GN) method in order to fully utilize the diversity of temporal–spatial data on local updates. The proposed algorithm includes two steps, i.e., aggregate then update (ATU), where the aggregating step collects in real time the global information instead of local information due to the diffusion strategy, and the updating step implements the local GN iteration. The resulting ATU diffusion algorithm is a distributed and cooperative system without any increase on communication cost, as compared with the noncooperative version. Based on the detailed convergence analysis for ATU, which is fundamental to the promotion of this algorithm, the sufficient conditions for convergence are derived and the evidences of faster convergence than the noncooperative version are provided. The simulation results confirm the obtained theoretical derivations by applying the ATU algorithm to an NLLS-based target localization problem and show the cooperation gains in many aspects, such as the convergence rate, steady-state accuracy, and robustness to noisy range, step size, node, and link failures.
Mou Wu, Naixue Xiong, Liansheng Tan, Jianguo Wei, Jie Wu 0001
IEEE Trans. Syst. Man Cybern. Syst.4
2021 An Optimized GPU Implementation of Weakly-Compressible SPH Using CUDA-Based Strategies
Yuejin Cai, Jianguo Wei, Darcy Qingzhi Hou, Ruixue Gao
ICA3PP (1)2
2021 Portable Photoglottography for Monitoring Vocal Fold Vibrations in Speech Production
abstract
Photoglottography (PGG) is an effective method to monitor vocal fold vibrations via measuring light transmission across the glottis. The difficulty in operation however limits its wide use in speech studies. This paper is to realize a portable PGG (P-PGG) module with an audio interface to record glottal and speech waveforms simultaneously with ease. Near Infrared (NIR) LEDs are driven as a light source and an extremely high-gain photodetector circuit is employed. The output PGG signal is subsequently band-pass filtered and amplified for recording. The whole system is minimized, battery powered with well-controlled heat radiation. In experiments, P-PGG, EGG and microphone are worn together by speakers. The results verify that the NIR lighting P-PGG is successful at recording complete information of glottal cycles in comparison to EGG. Thus, the P-PGG is an effective method for investigating phonation types and consonant-vowel interactions in speech.
Yujie Chi, Kiyoshi Honda, Jianguo Wei
ICASSP3
2021 Zero-Shot Voice Conversion with Adjusted Speaker Embeddings and Simple Acoustic Features
abstract
Zero-shot voice conversion (VC) where both source and target speakers are unseen in the training dataset has become a new research direction. Using speaker embeddings instead of one-hot vectors to represent speaker identity is a key point, which makes VC models work on unseen speakers. In our work, a newly designed neural network was used to adjust the speaker embeddings of unseen speakers. This enables speaker embeddings to perform better on zero-shot VC. In addition, disentangled representation of features is the mainstream method to achieve zero-shot VC. In terms of input features of VC model, we use Mel-cepstral and F0 as simple acoustic features (SAF) rather than Mel-spectrograms. This avoids F0 conflicts in decoder that existed in the previous methods. The evaluations demonstrate that our proposed methods improve the quality of converted speech in terms of naturalness and similarity.
Zhiyuan Tan 0003, Jianguo Wei, Junhai Xu, Wenhuan Lu
ICASSP2
2021 Multi-Modal Emotion Recognition Based On deep Learning Of EEG And Audio Signals
abstract
Automatic recognition of human emotional states has attracted many researchers' attention in Human-Computer Interactions and emotional brain-computer interface recently. However, the accuracy of emotion recognition is not satisfying. Considering the advantage of information supplement based on deep learning of multi-modal signals related to emotion, this study proposed a novel emotion recognition architecture to fuse emotional features from brain electroencephalography (EEG) signal and the corresponding audio signal in emotion recognition on DEAP dataset. We used convolutional neural network (CNN) to extract EEG features and bidirectional long short term memory (BiLSTM) neural networks to extract audio features. After that, we combine the multi-modal features into a deep learning architecture to recognize arousal and valence levels. Results showed an improved accuracy compared with previous studies that merely used the EEG signals in both arousal level and valence level, which suggests the effectiveness of our proposed multi-modal fused emotion recognition model. In future work, multi-modal data from nature interaction scenes will be collected and inputted into this architecture to further validate the effectiveness of the method.
Gaoyan Zhang, Jianwu Dang 0001, Longbiao Wang, Jianguo Wei
IJCNN5
2021 Reconstruction of natural images from evoked brain activity with a dictionary-based invertible encoding procedure
Chao Li 0052, Baolin Liu 0001, Jianguo Wei
Neurocomputing3
2021 Residual Learning Diagnosis Detection: An Advanced Residual Learning Diagnosis Detection System for COVID-19 in Industrial Internet of Things
abstract
Due to the fast transmission speed and severe health damage, COVID-19 has attracted global attention. Early diagnosis and isolation are effective and imperative strategies for epidemic prevention and control. Most diagnostic methods for the COVID-19 is based on nucleic acid testing (NAT), which is expensive and time-consuming. To build an efficient and valid alternative of NAT, this article investigates the feasibility of employing computed tomography images of lungs as the diagnostic signals. Unlike normal lungs, parts of the lungs infected with the COVID-19 developed lesions, ground-glass opacity, and bronchiectasis became apparent. Through a public dataset, in this article, we propose an advanced residual learning diagnosis detection (RLDD) scheme for the COVID-19 technique, which is designed to distinguish positive COVID-19 cases from heterogeneous lung images. Besides the advantage of high diagnosis effectiveness, the designed residual-based COVID-19 detection network can efficiently extract the lung features through small COVID-19 samples, which removes the pretraining requirement on other medical datasets. In the test set, we achieve an accuracy of 91.33%, a precision of 91.30%, and a recall of 90%. For the batch of 150 samples, the assessment time is only 4.7 s. Therefore, RLDD can be integrated into the application programming interface and embedded into the medical instrument to improve the detection efficiency of COVID-19.
Mingdong Zhang, Ronghe Chu, Chaoyu Dong, Jianguo Wei, Wenhuan Lu, Naixue Xiong
IEEE Trans. Ind. Informatics4
2020 Retrieving Vocal-Tract Resonance and anti-Resonance From High-Pitched Vowels Using a Rahmonic Subtraction Technique
abstract
Vocal tract resonances give rise to core spectral information of speech signals. Linear prediction and cepstral methods are widely used for this purpose. However, both approaches are prone to fail as the fundamental frequency (F0) rises. In this study, a new cepstral method is developed combined with a refined rahmonic subtraction technique (RS-CEPS) to extract spectral envelopes excited by glottal noise sources. A vowel synthesis system based on 3D-printed solid vocal tract models is used to obtain reference transfer functions for accuracy verification. A series of stable vowels /a/ was synthesized for a wide F0 range. By analyzing the synthetic vowels, the results showed that the RS-CEPS yields accurate estimates of resonance-peak and anti-resonance frequencies in comparison to those from the conventional methods. The RS-CEPS is simple and stable, offering a potential for expanding speech analysis applications.
Kiyoshi Honda, Jianguo Wei
ICASSP3
2020 Visual Encoding and Decoding of the Human Brain Based on Shared Features
abstract
Using a convolutional neural network to build visual encoding and decoding models of the human brain is a good starting point for the study on relationship between deep learning and human visual cognitive mechanism. However, related studies have not fully considered their differences. In this paper, we assume that only a portion of neural network features is directly related to human brain signals, which we call shared features. In the encoding process, we extract shared features from the lower and higher layers of the neural network, and then build a non-negative sparse map to predict brain activities. In the decoding process, we use back-propagation to reconstruct visual stimuli, and use dictionary learning and a deep image prior to improve the robustness and accuracy of the algorithm. Experiments on a public fMRI dataset confirm the rationality of the encoding models, and comparing with a recently proposed method, our reconstruction results obtain significantly higher accuracy.
Chao Li 0052, Baolin Liu 0001, Jianguo Wei
IJCAI3
2020 Regional Resonance of the Lower Vocal Tract and its Contribution to Speaker Characteristics
abstract
S.1391-1395
Lin Zhang 0054, Kiyoshi Honda, Jianguo Wei, Seiji Adachi
INTERSPEECH3
2020 ARET: Aggregated Residual Extended Time-Delay Neural Networks for Speaker Verification
Ruiteng Zhang, Jianguo Wei, Wenhuan Lu, Longbiao Wang, Meng Liu 0017, Lin Zhang 0054, Jiayu Jin, Junhai Xu
INTERSPEECH2
2020 Adversarial Separation Network for Speaker Recognition
Hanyi Zhang, Longbiao Wang, Yunchun Zhang, Meng Liu 0017, Kong-Aik Lee, Jianguo Wei
INTERSPEECH6
2020 Dynamic Margin Softmax Loss for Speaker Verification
Dao Zhou, Longbiao Wang, Kong-Aik Lee, Meng Liu 0017, Jianwu Dang 0001, Jianguo Wei
INTERSPEECH7
2020 The online estimation of the joint angle based on the gravity acceleration using the accelerometer and gyroscope in the wireless networks
Zhen Ding, Chifu Yang, Jiantao Ma, Jianguo Wei, Feng Jiang 0001
Multim. Tools Appl.4
2020 A Novel Image Encryption Algorithm Based on Hybrid Chaotic Mapping and Intelligent Learning in Financial Security System
Shuang Pan, Jianguo Wei, Shaobo Hu
Multim. Tools Appl.2
2019 Glottographic and Aerodynamic Analysis on Consonant Aspiration and Onset F0 in Mandarin Chinese
abstract
Stop consonants in Mandarin Chinese are all voiceless at word-initial positions only showing aspirated and unaspirated distinctions. Between the two phonation types, voice onset time (VOT) shows a clear contrast in duration, whereas voice onset fundamental frequency (onset F0) does not, as seen in previous studies. This study reports our first use of improved instrumentation techniques to record speech sound, oral airflow and glottal activity to examine consonant aspiration and onset F0. Glottal abduction, adduction and vibration cycles are monitored by a new external photo-glottographic system (ePGG) refined for better signal quality to detect accurate voice onset. Experimental data on Mandarin stops obtained from two male subjects suggests that consonant aspiration results in a large variation of VOT and oral airflow at voice onset. The onset F0 shows individual variation of falling and rising contours, and it is higher in aspirated stops than in unaspirated ones.
Yujie Chi, Kiyoshi Honda, Jianguo Wei
ICASSP3
2019 Breast Cancer Detection Based on Merging Four Modes MRI Using Convolutional Neural Networks
abstract
The objective of the study is to develop a framework for automatic breast cancer detection with merging four imaging modes. Attempts were made for tumor classification and segmentation; using a multi-parametric Magnetic Resonance Imaging (MRI) method on breast tumors. MRI data of the breast were obtained from 67 subjects with a 1.5T-MRI scanner. Four imaging modes: were T1 weighted, T2 weighted, Diffusion Weighted and eTHRIVE sequences, and dynamic-contrast-enhanced(DCE)-MRI parameters are acquired. The proposed four-mode linkage backbone in tumor classification, which overcomes the limitations of single-modality image detection and simulates actual diagnosis processes by clinicians, achieves the accuracy of 0.942. The proposed automatic segmentation approach is performed by a refined U-Net architecture, and the result improved segmentation performance significantly. The combination of four-mode linkage classification backbone and improved segmentation network for breast cancer detection forms a computer-aided detection (CAD) system that corresponds to the actual clinical diagnosis work.
Wenhuan Lu, Hong Yu 0017, Naixue Xiong, Jianguo Wei
ICASSP6
2019 Acoustic and Articulatory Study of Ewe Vowels: A Comparative Study of Male and Female
Kowovi Comivi Alowonou, Jianguo Wei, Wenhuan Lu, Kiyoshi Honda, Jianwu Dang 0001
INTERSPEECH2
2019 Individual Difference of Relative Tongue Size and its Acoustic Effects
Chongke Bi, Kiyoshi Honda, Wenhuan Lu, Jianguo Wei
INTERSPEECH5
2018 A Nonlinear 3D Geometric Tongue Model
abstract
This study describes a nonlinear geometric tongue model based on MRI and Cone-beam CT (CBCT) data. Comparing with the conventional geometric tongue model, the proposed tongue model is controlled by several prototype vertices, and the relationship between tongue mesh vertices and prototype vertices are modeled with quadratic functions. The results indicate that: i) quadratic models do improve the reconstruction performance of tongue mesh, especially in the tongue root region; ii) the quadratic model which use the cross-prototype-vertex information achieves the best performance of tongue mesh reconstruction; iii) the reconstruction performance can be further improved if an extra prototype vertex TP in the tongue root region is taken into account, even if TP is estimated from the measured prototype vertices.
Qiang Fang 0003, Hequn Li, Jianguo Wei, Jianrong Wang, Xiyu Wu
ICASSP3
2018 RE-CNN: A Robust Convolutional Neural Networks for Image Recognition
Wenhuan Lu, Naixue Xiong, Jianguo Wei
ICONIP (1)5
2018 Three-Dimensional Joint Geometric-Physiologic Feature for Lip-Reading
abstract
Lip-reading has been successfully demonstrated that it can improve the performance of automatic speech recognition system especially in the presence of acoustic noise. However, the information about lip movement is still insufficient as the lip features are obtained from discrete three-dimensional points and planar images. The internal mechanisms of lip movement are not described and reflected. In this paper, we employed a novel deepening technique, namely densely connected convolutional networks (DenseNets), to obtain visual representation from color images. In addition, a new 3D lip physiologic feature based on the position and structure of facial muscles was extracted to represent the similarity of the way people speak. The color image feature and 3D lip geometric-physiologic feature were coupled together in the last fully-connected layer of DenseNets. The experimental results show that DenseNets can handle spatial-temporal information of a whole image sequence and the lip feature integrating our proposed 3D geometric-physiological feature is sufficient to improve the recognition rate by as much as 3.91% (from 94.84%, with the color images only, to 98.75%).
Jianguo Wei, Ju Zhang 0001, Mei Yu 0004, Jianrong Wang
ICTAI1
2018 Tongue Segmentation with Geometrically Constrained Snake Model
Zhihua Su, Jianguo Wei, Qiang Fang 0003, Jianrong Wang, Kiyoshi Honda
INTERSPEECH2
2018 Study of articulators' contribution and compensation during speech by articulatory speech recognition
Jianguo Wei, Jingshu Zhang, Qiang Fang 0003, Wenhuan Lu, Kiyoshi Honda, Xugang Lu
Multim. Tools Appl.1
2018 Tooth visualization in vowel production MR images for three-dimensional vocal tract modeling
Ju Zhang 0001, Kiyoshi Honda, Jianguo Wei
Speech Commun.3
2017 Acoustic VR in the mouth: A real-time speech-driven visual tongue system
abstract
We propose an acoustic-VR system that converts acoustic signals of human language (Chinese) to realistic 3D tongue animation sequences in real time. It is known that directly capturing the 3D geometry of the tongue at a frame rate that matches the tongue's swift movement during the language production is challenging. This difficulty is handled by utilizing the electromagnetic articulography (EMA) sensor as the intermediate medium linking the acoustic data to the simulated virtual reality. We leverage Deep Neural Networks to train a model that maps the input acoustic signals to the positional information of pre-defined EMA sensors based on 1,108 utterances. Afterwards, we develop a novel reduced physics-based dynamics model for simulating the tongue's motion. Unlike the existing methods, our deformable model is nonlinear, volume-preserving, and accommodates collision between the tongue and the oral cavity (mostly with the jaw). The tongue's deformation could be highly localized which imposes extra difficulties for existing spectral model reduction methods. Alternatively, we adopt a spatial reduction method that allows an expressive subspace representation of the tongue's deformation. We systematically evaluate the simulated tongue shapes with real-world shapes acquired by MRI/CT. Our experiment demonstrates that the proposed system is able to deliver a realistic visual tongue animation corresponding to a user's speech signal.
Ran Luo 0001, Qiang Fang 0003, Jianguo Wei, Wenhuan Lu, Weiwei Xu 0003, Yin Yang 0002
VR3
2017 Parameterization of LSB in Self-Recovery Speech Watermarking Framework in Big Data Mining
abstract
The privacy is a major concern in big data mining approach. In this paper, we propose a novel self-recovery speech watermarking framework with consideration of trustable communication in big data mining. In the framework, the watermark is the compressed version of the original speech. The watermark is embedded into the least significant bit (LSB) layers. At the receiver end, the watermark is used to detect the tampered area and recover the tampered speech. To fit the complexity of the scenes in big data infrastructures, the LSB is treated as a parameter. This work discusses the relationship between LSB and other parameters in terms of explicit mathematical formulations. Once the LSB layer has been chosen, the best choices of other parameters are then deduced using the exclusive method. Additionally, we observed that six LSB layers are the limit for watermark embedding when the total bit layers equaled sixteen. Experimental results indicated that when the LSB layers changed from six to three, the imperceptibility of watermark increased, while the quality of the recovered signal decreased accordingly. This result was a trade-off and different LSB layers should be chosen according to different application conditions in big data infrastructures.
Zhanjie Song, Wenhuan Lu, Daniel Sun 0004, Jianguo Wei
Secur. Commun. Networks5
2016 Continuous ultrasound based tongue movement video synthesis from speech
abstract
The movement of tongue plays an important role in pronunciation. Visualizing the movement of tongue can improve speech intelligibility and also helps learning a second language. However, hardly any research has been investigated for this topic. In this paper, a framework to synthesize continuous ultrasound tongue movement video from speech is presented. Two different mapping methods are introduced as the most important parts of the framework. The objective evaluation and subjective opinions show that the Gaussian Mixture Model (GMM) based method has a better result for synthesizing static image and Vector Quantization (VQ) based method produces more stable continuous video. Meanwhile, the participants of evaluation state that the results of both methods are visual-understandable.
Jianrong Wang, Yalong Yang 0001, Jianguo Wei, Ju Zhang 0001
ICASSP3
2016 An Improved 3D Geometric Tongue Model
Qiang Fang 0003, Jianguo Wei, Jianrong Wang, Xiyu Wu
INTERSPEECH4
2016 A New Model for Acoustic Wave Propagation and Scattering in the Vocal Tract
Jianguo Wei, Wendan Guan, Darcy Qingzhi Hou, Dingyi Pan, Wenhuan Lu, Jianwu Dang 0001
INTERSPEECH1
2016 Semi-fragile watermarking for image authentication based on compressive sensing
Xiaochun Cao, Wei Zhang 0031, Xinpeng Zhang 0001, Jianguo Wei
Sci. China Inf. Sci.6
2016 Morphological normalization of vowel images for articulatory speech recognition
Jianguo Wei, Jingshu Zhang, Qiang Fang 0003, Wenhuan Lu
J. Vis. Commun. Image Represent.1
2016 Audio-visual speech recognition integrating 3D lip information obtained from the Kinect
Jianrong Wang, Ju Zhang 0001, Kiyoshi Honda, Jianguo Wei, Jianwu Dang 0001
Multim. Syst.4
2016 Mapping ultrasound-based articulatory images and vowel sounds with a deep neural network framework
Jianguo Wei, Qiang Fang 0003, Xinyuan Zheng, Wenhuan Lu, Jianwu Dang 0001
Multim. Tools Appl.1
2016 Multi-modal recording and modeling of vocal tract movements
Jianguo Wei, Song Wang 0005, Wenhuan Lu, Darcy Qingzhi Hou, Qiang Fang 0003, Jianwu Dang 0001
Multim. Tools Appl.1
2015 Vocal responses to frequency modulated composite sinewaves via auditory and vibrotactile pathways
abstract
Feedback control mechanisms for speaking have been examined using the transformed auditory feedback (TAF) technique. Previous studies have shown that speakers demonstrate fundamental frequency (F0) changes when they monitor their voice with artificial alterations of F0. However, those studies underestimate the role of vibrotactile information involved in feedback F0 control. This pilot study aims at exploring whether and how vibrotactile information from the larynx influences vowel F0. Participants in our experiment were asked to sustain vowel with their F0 adjusted to composite sinewave stimuli, which were given via auditory and vibrotactile channels using a headset on the ears or a bone-conduction transducer on the larynx. Results revealed the greater compensatory responses to combined vibrotactile-auditory stimuli than to the responses to auditory-only stimuli. The effect of vibrotactile stimuli on feedback F0 adjustment was also observed with the shorter latency of the responses.
Kiyoshi Honda, Jianwu Dang 0001, Jianguo Wei
ICASSP4
2015 Combined cine- and tagged-MRI for tracking landmarks on the tongue surface
Honghao Bao, Wenhuan Lu, Kiyoshi Honda, Jianguo Wei, Qiang Fang 0003, Jianwu Dang 0001
INTERSPEECH4
2015 Measuring oral and nasal airflow in production of Chinese plosive
Yujie Chi, Kiyoshi Honda, Jianguo Wei, Jianwu Dang 0001
INTERSPEECH3
2014 Reconstruction of mistracked articulatory trajectories
Qiang Fang 0003, Jianguo Wei
INTERSPEECH2
2013 An anisotropic diffusion filter based on multidirectional separability
Jianguo Wei, Xin Wang 0037, Wenhuan Lu, Qiang Fang 0003, Jianwu Dang 0001
INTERSPEECH2
2013 An MRI-based acoustic study of Mandarin vowels
Yuguang Wang 0003, Jianwu Dang 0001, Jianguo Wei, Hongcui Wang, Kiyoshi Honda
INTERSPEECH4
2013 A PMC-driven methodology for energy estimation in RVC-CAL video codec specifications
Rong Ren, Jianguo Wei, Eduardo Juárez Martínez, Matías J. Garrido, César Sanz, Fernando Pescador
Signal Process. Image Commun.2
2012 A method of speaker identification based on phoneme mean F-ratio contribution
Songgun Hyon, Hongcui Wang, Jianguo Wei, Jianwu Dang 0001
INTERSPEECH4
2010 Morphological normalization of vocal tract shape
abstract
The articulatory databases are not utilized so widely as acoustic databases. One of the reasons is the difficulty of reducing morphological variations among subjects. To reduce morphological differences in speech organs among speakers and remain their speech dynamics, this study proposed a framework of normalizing vocal tract by using a Thin-plate spline method. Electromagnetic Midsagittal Articulographic data for three subjects have been used in this research. The template for normalization was obtained by averaging all three subjects' palates and tongue shapes. The landmarks of the template and subjects have been defined according to a gridline system of the vocal tract. The results show that the variances among subjects were reduced 0.8 mm in horizontal and 2.4 mm in vertical direction. The similar vowel structure of pre/post-normalization data indicates that speaker specific characteristics can be maintained by this framework. The effects of the normalization in acoustic space are also investigated by using a physiological articulatory model. Results show that the variations have also been reduced in acoustic space.
Jianguo Wei, Jianwu Dang 0001
ICASSP1
2006 A simulation based parameter optimization for a coarticulation model
abstract
A coarticulation model, namely ‘carrier model’, has been proposed previously by Dang et al. to improve the performance of a physiological articulatory model based speech synthesizer. The carrier model offers a good framework to account for coarticulation in the planning stage, while its parameters need to be refined for improving the performance of the model. This study is to refine the parameters of the carrier model and estimate typical phonetic targets by minimizing the differences between model simulations and observations. A simulation based optimization framework is proposed for this purpose. The framework consists of two layers: obtaining planned targets in a low layer; estimating phonetic targets and optimizing the parameters in a high layer. A direct search method was applied to the low layer due to the non-analytic nature of the articulation model, while the high layer adopts bilevel optimization strategy to decompose the complicated problem into a set of subproblems. A general evaluation was conducted by combining the refined carrier model and the learned phonetic targets together using the physiological articulatory model and the average error between observations and simulations was 0.15 cm over 103 VCV combinations on the jaw, tongue tip and tongue dorsum. Index Terms: speech production, coarticulation, optimization
Jianguo Wei, Xugang Lu, Jianwu Dang 0001
INTERSPEECH1
2005 Investigation and modeling of coarticulation during speech
Jianwu Dang 0001, Jianguo Wei, Takeharu Suzuki, Pascal Perrier
INTERSPEECH2