VLDB 2026 Research / reviewers in the wild / expert
Wenhuan Lu
dblp:01/3219
· DBLP profile ↗
66ranked-venue papers
5as first author
45since 2021 · last 2026
0000-0002-7951-8907ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 37 · 2 first-author · 25 since 2021Artificial intelligence and machine learning · 19 · 10 since 2021Applied, interdisciplinary, general and emerging computing · 9 · 2 first-author · 7 since 2021Databases, data management, data science and information retrieval · 7 · 7 since 2021Security and privacy · 5 · 4 since 2021Systems, architecture and hardware · 2 · 1 first-authorHuman-computer interaction and ubiquitous computing · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | DuoKD: Dual Knowledge Distillation from Large Language Models for Robust Graph Neural NetworksabstractGraph neural networks (GNNs) have become a dominant modeling paradigm for graph-structured data, and the emergence of large language models (LLMs) has spurred growing interest in integrating external semantic knowledge into GNNs. Current LLM-based GNNs are devoted to extracting semantically similar information from LLMs to enhance representation learning. However, they generally overlook key signals that are semantically dissimilar but exhibit stronger inter-class discriminative ability. Especially when the original graph data contains noise or semantic ambiguity, a single similarity-based semantic augmentation strategy not only fails to provide effective enhancement, but may also amplify misleading signals generated by the LLM in response to low-quality inputs or its own hallucinations, further degrading the discriminative power and robustness of GNNs. To this end, we propose a dual positive-negative knowledge extraction strategy based on LLMs, and integrate it with a knowledge distillation mechanism to dynamically transfer multi-dimensional enhanced signals to GNNs, thereby achieving fine-grained and robust graph representation learning. Specifically, we design personalized prompts to guide LLMs in generating semantically similar positive signals and semantically dissimilar negative signals, which help the model capture intra-class consistency and inter-class distinction. Then, we further generate structural and semantic reasoning as supplementary knowledge to support the rationality and guidance of supervision signals. To identify high-confidence transferred knowledge, we introduce a language-based evaluation mechanism to filter low-confidence or hallucinated outputs. Finally, under a unified distillation framework, our method uses both positive and negative knowledge to guide GNN training, achieving adaptive and robust representation learning. Extensive experiments on benchmark datasets verify the superior performance of our approach across various tasks. Cuiying Huo, Xiaotong Huang, Dongxiao He, Wenhuan Lu, Di Jin 0001 |
AAAI | 5 |
| 2026 | EA-VAE: Learning to Reconstruct Dysarthric Speech via Variational Autoencoder with Encoding AlignmentabstractDysarthric speech reconstruction (DSR) aims to enhance the intelligibility of dysarthric speech. Compared with normal speech, the dysarthric speech is characterized by its pathological features, including discontinuous pronunciation, slow speech, hoarseness, and improper pauses. Significant disparities in the feature space between normal and dysarthric speech may result in suboptimal speech reconstruction, thereby degrading speech intelligibility. To enhance the reconstruction ability of speech feature spaces, this paper proposes a DSR model named the Encoding-Aligned Variational Autoencoder (EA-VAE). By incorporating alignment modules of frame-level embedding features, prior distributions, and duration into the encoder of the VAE, the model explicitly aligns the dysarthric speech encoding with a representation of the parallel normal speech. A shared decoder is then used to generate speech with improved intelligibility. Experimental results on the UASpeech benchmark confirm that EA-VAE achieves state-of-the-art performance, with a 31.7% relative word error rate reduction and the highest subjective MOS score (4.48), thoroughly validating the effectiveness and advancements of the proposed method in dysarthric speech reconstruction. Daipeng Zhang 0001, Wenhuan Lu, Xianghu Yue, Hongcheng Zhang, Jianguo Wei |
AAAI | 2 |
| 2026 | HalluAudio: A Comprehensive Benchmark for Hallucination Detection in Large Audio-Language ModelsabstractLarge Audio-Language Models (LALMs) have recently achieved strong performance across various audio-centric tasks.However, hallucination, where models generate responses that are semantically incorrect or acoustically unsupported, remains largely underexplored in the audio domain.Existing hallucination benchmarks mainly focus on text or vision, while the few audio-oriented studies are limited in scale, modality coverage, and diagnostic depth.We therefore introduce HalluAudio, the first large-scale benchmark for evaluating hallucinations across speech, environmental sound, and music.HalluAudio comprises over 5K humanverified QA pairs and spans diverse task types, including binary judgments, multi-choice reasoning, attribute verification, and open-ended QA.To systematically induce hallucinations, we design adversarial prompts and mixed-audio conditions.Beyond accuracy, our evaluation protocol measures hallucination rate, yes/no bias, error-type analysis, and refusal rate, enabling a fine-grained analysis of LALM failure modes.We benchmark a broad range of open-source and proprietary models, providing the first large-scale comparison across speech, sound, and music.Our results reveal significant deficiencies in acoustic grounding, temporal reasoning, and music attribute understanding, underscoring the need for reliable and robust LALMs. Feiyu Zhao, Wenhuan Lu, Daipeng Zhang 0001, Xianghu Yue, Jianguo Wei |
ACL (1) | 3 |
| 2026 | Integrated Mixture of Neighborhood and Community Experts for Graph-Based Fraud DetectionabstractGraph-based fraud detection (GFD) aims to identify fraud nodes within graph-structured data that significantly deviate from the majority of benign nodes. However, existing graph neural networks (GNNs) often struggle in GFD scenarios due to their reliance on homophily assumption, which is frequently violated by the inherent homophily-heterophily mixture of fraud graphs. Moreover, most methods focus primarily on local topology, overlooking mesoscopic community structures, making them less efficient in detecting suspicious patterns like densely connected subgraphs. To address the aforementioned issues, we present NeCo, a novel approach that integrates mixture of neighborhood and community experts for graph-based fraud detection. Specifically, we first introduce a fraud-discriminative representation preservation mechanism from a neighborhood perspective, leveraging the empirical finding that fraud nodes tend to exhibit larger feature propagation discrepancies compared to benign nodes. We then design a community-oriented node representation module that models structural compactness among nodes, enabling the detection of suspicious topological patterns associated with fraud behaviors. By integrating these two complementary perspectives, NeCo can effectively captures both local inconsistency and global structural irregularity. Extensive experiments across five real-world datasets demonstrate the effectiveness of our proposed NeCo over state-of-the-art baselines. Zhizhi Yu, Di Jin 0001, Dongxiao He, Wenhuan Lu, Jianguo Wei |
WWW | 4 |
| 2026 | Listening for "You": Enhancing Speech Image Retrieval via Target Speaker ExtractionabstractImage retrieval using spoken language cues has emerged as a promising direction in multimodal perception, yet leveraging speech in multi-speaker scenarios remains challenging. We propose a novel Target Speaker Speech-Image Retrieval task and a framework that learns the relationship between images and multi-speaker speech signals in the presence of a target speaker. Our method integrates pre-trained self-supervised audio encoders with vision models via target speaker-aware contrastive learning, conditioned on a Target Speaker Retrieval Extractor (TSRE) module. This method enables the extraction of semantic content from the target speaker's speech and aligns it with images representing the corresponding semantic meaning. Experiments on SpokenCOCO2Mix and SpokenCOCO3Mix show that TSRE significantly outperforms existing methods, achieving 36.3% and 29.9% Recall@1 in 2- and 3-speaker scenarios, respectively-substantial improvements over single-speaker baselines and state-of-the-art models. Our approach demonstrates potential for real-world deployment in assistive robotics and multimodal interaction systems. Jianguo Wei, Wenhuan Lu, Xinyue Song, Xianghu Yue |
IEEE Signal Process. Lett. | 3 |
| 2026 | Domain Adaptation for Speaker Verification Using Optimal Transport With Pseudo LabelabstractDomain gap often degrades the performance of speaker verification (SV) systems when the statistical distributions of training data and real-world test speech are mismatched. Channel variation is a primary factor causing this gap, including bandwidth changes, background noise and encoding, etc. Although various domain adaptation algorithms could be applied to handle this domain gap problem, most algorithms could not take the complex distribution structure in domain alignment with discriminative learning. In this paper, we propose a novel unsupervised domain adaptation method for speaker verification, i.e., Joint Partial Optimal Transport with Pseudo Label (JPOT-PL), to alleviate the domain mismatch problem. Leveraging the geometric-aware distance metric of optimal transport in distribution alignment and speaker consistency in speech distribution, we further design a pseudo label-based discriminative learning where the pseudo label can be regarded as a new type of speaker label derived from the optimal coupling. With the JPOT-PL, we carry out experiments on the SV channel and lingual domain adaptation with VoxCeleb, LibriSpeech, CNCeleb, and AISHELL-2. Experiments show our method reduces EER by up to 30% compared with several state-of-the-art domain adaptation algorithms. Jianguo Wei, Wenhuan Lu, Lei Li 0050, Xugang Lu |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2025 | Continual Unsupervised Domain Adaptation for Audio Deepfake DetectionabstractAudio deepfake detection (ADD) aims to verify the authenticity of audio. However, its performance declines sharply when facing significant domain discrepancies caused by unknown datasets. Unsupervised domain adaptation (UDA) has been applied to mitigate domain mismatch. However, as generative models evolve, existing UDA methods struggle with catastrophic forgetting when facing continuously emerging spoofing methods. To address this challenge, we introduce continual UDA for ADD, which involves sequentially training across multiple target domains with continual learning. We propose a causality-distillation-based continual domain adversarial training framework for continual UDA, called CD-DAT. Specifically, we employ the domain adversarial training (DAT) framework to learn both spoofing-discriminative and domain-invariant deep features. In addition, we design a continual learning algorithm utilizing causality distillation to capture the mapping between utterances and classes, effectively mitigating forgetting and maintaining generalization. Experiments demonstrated that CD-DAT improved detection performance across all domains, confirming its memory stability and learning plasticity. Xiaohuan Chen, Wenhuan Lu, Ruiteng Zhang, Junhai Xu, Xugang Lu, Lin Zhang 0054, Jianguo Wei |
ICASSP | 2 |
| 2025 | Adaptive Multi-Scale Local Correction for Semi-Supervised 3D Medical Image SegmentationabstractIn recent years, semi-supervised 3D medical image segmentation has gained significant attention. However, current methods often struggle with multi-scale voxel differences and overlook the importance of loss weight balancing. To address these issues, we propose an adaptive multi-scale local correction method (AMLC). Our key contributions are: (1) a multi-scale local correction module to accurately capture differences between 3D voxel blocks at various scales; (2) an adaptive weighting adjustment module that dynamically balances losses, improving robustness and training. Experiments show that AMLC significantly outperforms existing methods on the public dataset, demonstrating its effectiveness. Xinqiang Wang, Wenhuan Lu, Junhai Xu |
ICASSP | 2 |
| 2025 | You Only Speak Once to SeeabstractGrounding objects in images using visual cues is a well-established approach in computer vision, yet the potential of audio as a modality for object recognition and grounding remains underexplored. We introduce YOSS, "You Only Speak Once to See," to leverage audio for grounding objects in visual scenes, termed Audio Grounding. By integrating pre-trained audio models with visual models using contrastive learning and multi-modal alignment, our approach captures speech commands or descriptions and maps them directly to corresponding objects within images. Experimental results indicate that audio guidance can be effectively applied to object grounding, suggesting that incorporating audio guidance may enhance the precision and robustness of current object grounding methods and improve the performance of robotic systems and computer vision applications. This finding opens new possibilities for advanced object recognition, scene understanding, and the development of more intuitive and capable robotic systems. Jianguo Wei, Wenhuan Lu |
ICASSP | 3 |
| 2025 | DCHT-Net: Medical Object Detection Based on Dynamic Deep Circular Hough Transform
Wanling Liu, Wenhuan Lu |
ICIC (9) | 2 |
| 2025 | Adaptive Capsule Graph Neural Network with Attention Mechanism for Parathyroid Glands Detection
Wanling Liu, Wenhuan Lu, Fei Chen 0012, Wenxin Zhao |
KSEM (4) | 2 |
| 2025 | Integrated registration and utility of mobile AR Human-Machine collaborative assembly in rail transit
Jiu Yong, Jianguo Wei, Xiaomei Lei, Yangping Wang, Wenhuan Lu |
Adv. Eng. Informatics | 6 |
| 2025 | Efficient dehazing network based on mix structure for single image with uneven haze distribution
Kangle Yuan, Jianguo Wei, Wenhuan Lu |
Eng. Appl. Artif. Intell. | 3 |
| 2025 | Block sparse Bayes-based fuzzy system for RNA N6-methyladenosine sites predictionabstractN6-methyladenosine (m6A) can significantly affect RNA expression, gene regulation, and determination of cell fate. As a common and abundant post-transcriptional modification (PTM) of RNA, m6A is also closely associated with the occurrence of numerous diseases. Thus, identifying the m6A modification site in the RNA sequence is a prerequisite for related research. High-throughput sequencing technology has high requirements and low cost performance. Computational methods have made encouraging progress in site prediction. However, most models only consider the effects of different species, ignoring the simultaneous exploration of RNA modifications in different tissues within the same species. We develop and validate a fuzzy system based on Block Sparse Bayesian Learning (BSBL), named BSBL-TSK-FS, which is a powerful sequence-level m6A prediction model. We introduce a Bayesian method that provides a posterior probability output to produce more sparse solutions so that the model has higher accuracy. The model classifies the m6A sites in several tissues of mouse, human, and rat. Under the five-fold cross-validation method (5-CV), the precision of the BSBL-TSK-FS model is 0.84∼0.95. The accuracy of our model improves by 9.4% over the existing SOTA predictors. BSBL-TSK-FS achieves superior performance over current SOTA methods. Finally, in order to verify the generalizability of the model, we carry out cross-species tests, and the results prove the robustness and adaptability of the model. An accurate and reliable sequence modification prediction model is developed to better understand the complex landscape of methylation modification. Yuqing Qian, Wenhuan Lu, Yijie Ding, Fei Guo 0001 |
PLoS Comput. Biol. | 5 |
| 2025 | Self-distillation-based domain exploration for source speaker verification under spoofed speech from unknown voice conversion
Xinlei Ma, Ruiteng Zhang, Jianguo Wei, Xugang Lu, Junhai Xu, Lin Zhang 0054, Wenhuan Lu |
Speech Commun. | 7 |
| 2025 | SHDA: Sinkhorn Domain Attention for Cross-Domain Audio Anti-SpoofingabstractAudio anti-spoofing algorithms struggle with fake samples from unseen spoofing techniques, even when trained with diverse data sets or data augmentation strategies. Unsupervised domain adaptation (UDA) algorithms have the potential to mitigate this challenge. Typically, UDA assumes that the source and target domains are distinct distributions with clear boundaries and seeks to align model representations between them. However, in anti-spoofing, various spoofing algorithms could cause the distributions of the generated samples to overlap, resulting in unclear domain boundaries. This hinders UDA algorithms from effectively measuring and aligning domain discrepancies. Moreover, forcibly aligning samples with significant discrepancies could diminish the model’s discriminative capability. To solve this problem, we propose a domain attention algorithm with optimal transport (OT), termed Sinkhorn Domain Attention (SHDA). Unlike traditional attention mechanisms, SHDA identifies the optimal transfer plan by analyzing the global probability differences among cross-domain samples. Specifically, we first extract audio representations from various domains to compute the overall cost matrix between the source and target domains. Next, we employ Sinkhorn’s iteration to calculate the OT coupling matrix, where cross-domain samples with minor differences receive higher transfer weights, while those with substantial differences receive lower weights. Finally, we use the coupling and cost matrices to compute the adaptation loss, effectively transferring the anti-spoofing model from multiple sources to the target domain. We conducted eight cross-domain experiments using eleven well-known anti-spoofing corpora. The results indicate that our label-free SHDA surpassed the state-of-the-art model by 40%. Ruiteng Zhang, Jianguo Wei, Xugang Lu, Lin Zhang 0054, Di Jin 0001, Junhai Xu, Wenhuan Lu |
IEEE Trans. Inf. Forensics Secur. | 7 |
| 2024 | Real-Time Double-Layer Graph Attention Networks for Parathyroid DetectionabstractSince parathyroid glands (PG) regulate the body’s calcium levels and significantly impact human health, developing methods for their automatic detection during endoscopic thyroid surgery is of utmost clinical significance. However, existing parathyroid detection works suffer from color variations, target deformation, blur, and lighting effects in intricate surgical environments. To address the above shortcomings, in this paper, we propose a novel double-layer graph attention network for PG detection, which explicitly facilitates local augmentation via key visual features (e.g., texture and shape) identification and global interactions. It can robustly combat image blur and better differentiate the PG targets and background parts, thus improving the detection precision. Furthermore, we observe most prior works fail to deeply understand the spatial relation among targets and unavoidably suffer from false or missed detection, which is heavily due to total ignorance or insufficient utilization of depth information, especially under lighting variations and occlusions. To fill the gap, we propose a depth relation augmentation component to adaptively capture the prominent relative positional relations between targets based on depth information and incorporate it into the proposed GNN framework, significantly deepening spatial understandings and naturally enhancing generalizability. Due to lacking a thyroid endoscopy surgery benchmark for evaluating this task, we meticulously established a novel dataset from 838 actual surgeries conducted (via the fully laparoscopic thoracic-breast approach) at the Fujian Medical University Union Hospital. Extensive experiments show that our framework achieves superior PG detection accuracy compared to current state-of-the-art counterparts while keeping real-time efficiency. Wanling Liu, Wenhuan Lu, Fei Chen 0012, Wenxin Zhao |
BIBM | 2 |
| 2024 | Breaking the Corpus Bottleneck for Multi-dialect Speech Recognition with Flexible Adapters
Tengyue Deng, Jianguo Wei, Wenjun Ke 0001, Xiaokang Yang 0003, Wenhuan Lu |
ICANN (7) | 7 |
| 2024 | EEG-Based Fast Auditory Attention Detection in Real-Life Scenarios Using Time-Frequency Attention MechanismabstractAuditory attention detection (AAD) based on electroencephalogram (EEG) helps recognize the target speaker in a cocktail party scenario, advancing auditory brain-computer interface development. Previous EEG studies on AAD were largely based on data collected in laboratory settings. In this study, we investigated the AAD with EEG data collected when subjects were walking and sitting in real-life scenarios. To improve the detection accuracy, we proposed the time-frequency attention mechanism to the convolution neural network on EEG data. Experimental results show that the proposed model outperforms the state-of-the-art models, with an accuracy of 98.1% on a decision window of 2s. When we used a 0.1s time window for fast decoding, the accuracy remained at 91.8%, suggesting the potential for real application. Further study on ablation experiments demonstrates the effectiveness of the proposed time-frequency attention mechanism. Analysis of the key EEG features indicates that the β band plays a vital role in AAD. Zhuang Xie, Jianguo Wei, Wenhuan Lu, Chunli Wang, Gaoyan Zhang |
ICASSP | 3 |
| 2024 | Self-Supervised Domain Exploration with an Optimal Transport Regularization for Open Set Cross-Domain Speech Emotion RecognitionabstractIn the tasks of domain adaptation (DA) for speech emotion recognition (SER), self-supervised learning (SSL) algorithms could effectively explore domain and structural information from target domain samples, thereby mitigating domain discrepancies. However, in a general setting, when the target domain contains emotions that are never observed in the source domain, namely in open-set DA, existing SSL-based DA methods cannot maintain the robustness because of the interference of the extra unknown classes. To address this challenge, we propose the self-supervised domain exploration with an optimal transport (OT) regularization (SDEOTR) algorithm. First, we integrate the SSL algorithm into the SER model to mitigate the domain differences. Further, we categorize target domain samples into known and unknown groups based on the network’s prediction confidence. Finally, we employ OT to maximize the global probability distance between the two groups, aiming to decrease the impact of unknown emotions on the SER model. Cross-domain SER experimental results showed that our label-free SDEOTR significantly improved the performance of existing adaptive SER algorithms in open-set scenarios. Ruiteng Zhang, Jianguo Wei, Xugang Lu, Wenhuan Lu, Di Jin 0001, Junhai Xu |
ICASSP | 5 |
| 2024 | SHAN: Shape Guided Network for Thyroid Nodule Ultrasound Cross-Domain Segmentation
Wenhuan Lu, Cuntai Guan, Jie Gao 0008, Xi Wei 0002, Xuewei Li 0001 |
MICCAI (4) | 2 |
| 2024 | Distillation-Based Feature Extraction Algorithm For Source Speaker VerificationabstractAutomatic speaker verification (ASV) systems face significant challenges when exposed to spoofing attacks, necessitating robust countermeasures. In this work, we focus on the source speaker verification (SSV) task, which aims to identify the source speaker hidden in spoofed speech generated by voice conversion (VC) systems. We propose a distillation-based feature extraction algorithm to enhance the model’s ability to verify source speakers. Our method employs a pretraining ASV model as a teacher network and the SSV model as a student network, using bona fide speech to guide the learning process. However, the improvements were marginal, particularly on the development set, indicating the complexity and resource demands of fine-tuning the distillation parameters. Our findings underscore the inherent difficulties in SSV and highlight the need for further research to develop more effective solutions. Besides, our submission won fourth place in the 2024 Source Speaker Tracking Challenge. Xinlei Ma, Wenhuan Lu, Ruiteng Zhang, Junhai Xu, Xugang Lu, Jianguo Wei |
SLT | 2 |
| 2024 | End-To-End Speaker Anonymization Based on Location-Variable Convolution and Multi-Head Self-AttentionabstractSpeaker anonymization, a user-centric solution for voice privacy, aims to conceal the speaker’s identity while maintaining clarity and naturalness. The prevalent approach involves cascading modules of automatic speech recognition (ASR) and text-to-speech (TTS) models for speaker anonymization through speech synthesis. However, the inherent multimodal cascade nature of this approach leads to high error rates and unclear speech due to inaccuracies propagated from the ASR system to the TTS system. To address these issues, this paper proposes an end-to-end method for achieving zero-shot speaker anonymization. This method improves the shortcomings of traditional speech models that use a large number of fixed convolution kernels to capture the internal dependencies of speech sequences. It captures the internal dependencies of speech from both long-time and long-distance perspectives by combining a network with variable kernels, namely location-variable convolutions (LVCs), with a multi-head self-attention mechanism and dynamically adjusting weights. Besides, it learns anonymized speaker features flexibly through an enhanced cycle-consistency loss, iteratively aligning speaker information for restructured speech with that of anonymized speakers indefinitely. The efficacy of our proposed speaker anonymization model was demonstrated on the English dataset VCTK. Feiyu Zhao, Jianguo Wei, Wenhuan Lu |
TrustCom | 3 |
| 2024 | Semisupervised Medical Image Segmentation through Prototype-Based Mutual Consistency LearningabstractMedical image segmentation is a critical task in the healthcare field. While deep learning techniques have shown promise in this area, they often require a large number of accurately labeled images. To address this issue, semisupervised learning has emerged as a potential solution by reducing the reliance on precise annotations. Among these approaches, the student-teacher framework has garnered attention, but it is limited in its reliance solely on the teacher model for information. To overcome this limitation, we propose a prototype-based mutual consistency learning (PMCL) framework. This framework utilizes two branches that learn from each other, incorporating supervision loss and consistency loss to adapt to minor data perturbations and structural differences. By employing prototype consistency learning, we are able to achieve reliable consistency loss. Our experiments on three public medical image datasets demonstrate that PMCL outperforms other state-of-the-art methods, indicating its potential in semisupervised medical image segmentation. Our framework has the potential to assist medical professionals in enhancing their diagnoses and delivering improved patient care. Xinqiang Wang, Wenhuan Lu, Junhai Xu, Jianguo Wei |
Int. J. Intell. Syst. | 2 |
| 2024 | Zero-shot voice conversion based on feature disentanglement
Jianguo Wei, Wenhuan Lu, Jianhua Tao 0001 |
Speech Commun. | 4 |
| 2024 | Unsupervised Adaptive Speaker Recognition by Coupling-Regularized Optimal TransportabstractCross-domain speaker recognition (SR) can be improved by unsupervised domain adaptation (UDA) algorithms. UDA algorithms often reduce domain mismatch at the cost of decreasing the discrimination of speaker features. In contrast, optimal transport (OT) has the potential to achieve domain alignment while preserving the speaker discrimination capability in UDA applications; however, naively applying OT to measure global probability distribution discrepancies between the source and target domains may induce negative transports where samples belonging to different speakers are coupled in transportation. These negative transports reduce the SR model's discriminative power, degrading the SR performance. This paper proposes a coupling-regularized optimal transport (CROT) algorithm for cross-domain SR to reduce the negative transport during UDA. In the proposed CROT, two consecutive processing modules regularize the coupling paths for the OT solution: a progressive inter-speaker constraint (PISC) module and a coupling-smoothed regularization (CSR) module. The PISC, designed as a pseudo-label memory bank with curriculum learning, is first applied to select valid samples to guarantee that coupling samples are from the same speaker. The CSR, designed to control the information entropy of the coupling paths further, reduces the effect of negative transport in UDA. To evaluate the effectiveness of the proposed algorithm, cross-domain SR experiments were conducted under different target domains, speaker encoders, corpora, and acoustic features. Experimental results showed that CROT achieved a 50% relative reduction in equal error rates compared to conventional OT-based UDAs, outperforming the state-of-the-art UDAs. Ruiteng Zhang, Jianguo Wei, Xugang Lu, Wenhuan Lu, Di Jin 0001, Lin Zhang 0054, Junhai Xu |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2024 | Weakly-Supervised Video Anomaly Detection With Snippet Anomalous AttentionabstractWith a focus on abnormal events contained within untrimmed videos, there is increasing interest among researchers in video anomaly detection. Among different video anomaly detection scenarios, weakly-supervised video anomaly detection poses a significant challenge as it lacks frame-wise labels during the training stage, only relying on video-level labels as coarse supervision. Previous methods have made attempts to either learn discriminative features in an end-to-end manner or employ a two-stage self-training strategy to generate snippet-level pseudo labels. However, both approaches have certain limitations. The former tends to overlook informative features at the snippet level, while the latter can be susceptible to noises. In this paper, we propose an Anomalous Attention mechanism for weakly-supervised anomaly detection to tackle the aforementioned problems. Our approach takes into account snippet-level encoded features without the supervision of pseudo labels. Specifically, our approach first generates snippet-level anomalous attention and then feeds it together with original anomaly scores into a Multi-branch Supervision Module. The module learns different areas of the video, including areas that are challenging to detect, and also assists the attention optimization. Experiments on benchmark datasets XD-Violence and UCF-Crime verify the effectiveness of our method. Besides, thanks to the proposed snippet-level attention, we obtain a more precise anomaly localization. Yidan Fan, Yongxin Yu, Wenhuan Lu, Yahong Han |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2023 | Optimal Transport with a Diversified Memory Bank for Cross-Domain Speaker VerificationabstractOptimal transport (OT) can be applied to cross-domain adaptation in speaker verification (SV) by converting speakers' probability distributions from source to target domains. However, in scenarios involving over-massive categories (speakers) or difficult samples in discrimination, OT often has difficulty computing effective transports. To address this challenge, we propose an OT-based unsupervised domain adaptation (UDA) framework for SV, OT with a diversified memory bank, called DMB-OT, which ensures the accuracy of transfers by two strategies: (1) It regularizes the solution space of OT, which attempts to plan transformations between audio samples from the same speaker with high confidence; (2) it integrates a dynamic curriculum learning algorithm, preventing OT from calculating transport couplings based on hard-discriminative samples in the early stage of UDA. Experiments under different target domains showed that our unsupervised DMB-OT could significantly improve the performance of OT-based UDA and could even match the performance of the supervised PLDA-based adaptation. Ruiteng Zhang, Jianguo Wei, Xugang Lu, Wenhuan Lu, Di Jin 0001, Lin Zhang 0054, Junhai Xu |
ICASSP | 4 |
| 2023 | Discriminative and Contrastive Consistency for Semi-supervised Domain Adaptive Image ClassificationabstractWith sufficient source and limited target supervised information, semi-supervised domain adaptation (SSDA) aims to perform well on unlabeled target domain. Although various strategies have been proposed in SSDA field, they fail to fully exploit limited target labels and adequately explore domain-invariant knowledge. In this study, we propose a framework that first introduces consistent processing of augmented training data based on contrastive learning. Specifically, supervised contrastive learning is introduced to assist the classical cross-entropy iteration to make full use of the limited target labels. Additionally, traditional unsupervised contrastive learning and pseudo-labeling are utilized to further minimize the intra-domain discrepancy. Besides, an adversarial loss is then combined with a sharpening function to acquire a more certain category center that is domain-invariant. Experimental results on DomainNet, Office-Home, and Office show the effectiveness of our method. Particularly, for 1-shot case of Office-Home with AlexNet as backbone, our method outperforms the previous state-of-the-art by 5.6% in terms of mean accuracy. Yidan Fan, Wenhuan Lu, Yahong Han |
ICME | 2 |
| 2023 | A Cross-modal and Redundancy-reduced Network for Weakly-Supervised Audio-Visual Violence DetectionabstractMultimodal learning using audio and visual information has improved Violence Detection tasks. However, previous studies overlook the gap between pre-trained networks and the final violence detection task, as well as the semantic inconsistency between audio and visual features. We consider task-irrelevant information caused by the former situation and semantic noise due to the latter as redundancy, negatively affecting overall detection performance. Besides, the prevailing visual modality-centric approach with audio features as guidance may be biased. We contend that both modalities are crucial in violence detection. To address these issues, we propose a Cross-modal and Redundancy-reduced Network for Weakly-Supervised Audio-Visual Violence Detection. Our framework integrates a relation-ware module with a bi-directional cross-modal attention mechanism to explore interactions between modalities. Then, we introduce a feature filter gate to reduce redundancy. Finally, a multi-branch classification module is proposed for better utilization of both modalities. Extensive experiments demonstrate the effectiveness of our approach, surpassing previous methods with state-of-the-art performance in violence detection. Yidan Fan, Yongxin Yu, Wenhuan Lu, Yahong Han |
MMAsia | 3 |
| 2023 | TMS: Temporal multi-scale in time-delay neural network for speaker verification
Ruiteng Zhang, Jianguo Wei, Xugang Lu, Wenhuan Lu, Di Jin 0001, Lin Zhang 0054, Junhai Xu, Jianwu Dang 0001 |
Appl. Intell. | 4 |
| 2023 | Super-Resolution Reconstruction of Remote Sensing Images Based on Symmetric Local Fusion BlocksabstractIn view of the rich information and strong autocorrelation of remote sensing images, a super-resolution reconstruction algorithm based on symmetric local fusion blocks is proposed using a convolutional neural network based on local fusion blocks, which improves the effect of high-frequency information reconstruction. By setting local fusion in the residual block, the problem of insufficient high-frequency feature extraction is alleviated, and the reconstruction accuracy of remote sensing images of deep networks is improved. To improve the utilization of global features and reduce the computational complexity of the network, a residual method is used to set the symmetric jump connection between the local fusion blocks to form the symmetry between them. Experimental results show that the reconstruction results of 2-, 3-, and 4-fold sampling factors on the UC Merced and nwpu-resisc45 remote sensing datasets are better than those of comparison algorithms in image clarity and edge sharpness, and the reconstruction results are better in objective evaluation and subjective vision. Xinqiang Wang, Wenhuan Lu |
Int. J. Inf. Secur. Priv. | 2 |
| 2023 | A deep multiple kernel learning-based higher-order fuzzy inference system for identifying DNA N4-methylcytosine sites
Yijie Ding, Prayag Tiwari, Junhai Xu, Wenhuan Lu, Khan Muhammad 0001, Victor Hugo C. de Albuquerque, Fei Guo 0001 |
Inf. Sci. | 5 |
| 2023 | RFI-GAN: A reference-guided fuzzy integral network for ultrasound image augmentation
Wenhuan Lu, Jie Gao 0008, Xi Wei 0002, Chenhan Wang, Xuewei Li 0001, Mei Yu 0004 |
Inf. Sci. | 2 |
| 2023 | A watermark detection scheme based on non-parametric model applied to mute machine voice
Yangxia Hu, Wenhuan Lu, Jianguo Wei, Junhai Xu, Maode Ma |
Multim. Tools Appl. | 2 |
| 2023 | Self-supervised learning based domain regularization for mask-wearing speaker verification
Ruiteng Zhang, Jianguo Wei, Xugang Lu, Wenhuan Lu, Di Jin 0001, Lin Zhang 0054, Yantao Ji, Junhai Xu |
Speech Commun. | 4 |
| 2022 | Soft Label Mining and Average Expression Anchoring for Facial Expression Recognition
Haipeng Ming, Wenhuan Lu |
ACCV (4) | 2 |
| 2022 | SemiPainter: Learning to Draw Semi-realistic Paintings from the Manga Line Drawings and Flat Shadow
Keyue Fan, Shiguang Liu, Wenhuan Lu |
CGI | 3 |
| 2022 | Joint and Adversarial Training with ASR for Expressive Speech SynthesisabstractStyle modeling is an important issue and has been proposed in expressive speech synthesis. In existing unsupervised methods, the style encoder extracts the latent representation from the reference audio as style information. However, the style information extracted from the style encoder will entangle some content information, which will cause conflicts with the real input content, and the synthesized speech will be influenced. In this study, we propose to alleviate the entanglement problem by integrating Text-To-Speech (TTS) model and Automatic Speech Recognition (ASR) model with a share layer network for joint training, and using ASR adversarial training to eliminate the content information in the style information. At the same time, we propose an adaptive adversarial weight learning strategy to prevent the model from collapsing. The objective evaluation using word error rate(WER) demonstrates that our method can effectively alleviate the entanglement between style and content information. Subjective evaluation indicates that the method improves the quality of synthesized speech and enhances the ability of style transfer compared with the baseline models. Wenhuan Lu, Longbiao Wang, Jianguo Wei |
ICASSP | 3 |
| 2022 | CS-REP: Making Speaker Verification Networks Embracing Re-ParameterizationabstractAutomatic speaker verification (ASV) systems, which determine whether two speeches are from the same speaker, mainly focus on verification accuracy while ignoring inference speed. However, in real applications, both inference speed and verification accuracy are essential. This study proposes cross-sequential re-parameterization (CS-Rep), a novel topology re-parameterization strategy for multi-type networks, to increase the inference speed and verification accuracy of models. CS-Rep solves the problem that existing re-parameterization methods are not suitable for typical ASV backbones. When a model applies CS-Rep, the training-period network utilizes a multi-branch topology to capture speaker information, whereas the inference-period model converts to a time-delay neural network (TDNN)-like plain backbone with stacked TDNN layers to achieve the fast inference speed. Based on CS-Rep, an improved TDNN with friendly test and deployment called Rep-TDNN is proposed. Compared with the state-of-the-art model ECAPA-TDNN, Rep-TDNN increases the actual inference speed by about 50% and reduces the EER by 10%. The code and trained models are available at https://github.com/zrtlemontree/CS-Rep. Ruiteng Zhang, Jianguo Wei, Wenhuan Lu, Lin Zhang 0054, Yantao Ji, Junhai Xu, Xugang Lu |
ICASSP | 3 |
| 2022 | A semi fragile watermarking algorithm based on compressed sensing applied for audio tampering detection and recovery
Yangxia Hu, Wenhuan Lu, Maode Ma, Qilong Sun, Jianguo Wei |
Multim. Tools Appl. | 2 |
| 2022 | One-shot emotional voice conversion based on feature separation
Wenhuan Lu, Xinyue Zhao, Jianguo Wei, Jianhua Tao 0001, Jianwu Dang 0001 |
Speech Commun. | 1 |
| 2022 | A Progressive Generative Adversarial Method for Structurally Inadequate Medical Image Data AugmentationabstractThe generation-based data augmentation method can overcome the challenge caused by the imbalance of medical image data to a certain extent. However, most of the current research focus on images with unified structure which are easy to learn. What is different is that ultrasound images are structurally inadequate, making it difficult for the structure to be captured by the generative network, resulting in the generated image lacks structural legitimacy. Therefore, a Progressive Generative Adversarial Method for Structurally Inadequate Medical Image Data Augmentation is proposed in this paper, including a network and a strategy. Our Progressive Texture Generative Adversarial Network alleviates the adverse effect of completely truncating the reconstruction of structure and texture during the generation process and enhances the implicit association between structure and texture. The Image Data Augmentation Strategy based on Mask-Reconstruction overcomes data imbalance from a novel perspective, maintains the legitimacy of the structure in the generated data, as well as increases the diversity of disease data interpretably. The experiments prove the effectiveness of our method on data augmentation and image reconstruction on Structurally Inadequate Medical Image both qualitatively and quantitatively. Finally, the weakly supervised segmentation of the lesion is the additional contribution of our method. Wenhuan Lu, Xi Wei 0002, Han Jiang 0004, Zhiqiang Liu 0002, Jie Gao 0008, Xuewei Li 0001, Jian Yu 0003, Mei Yu 0004 |
IEEE J. Biomed. Health Informatics | 2 |
| 2021 | Zero-Shot Voice Conversion with Adjusted Speaker Embeddings and Simple Acoustic FeaturesabstractZero-shot voice conversion (VC) where both source and target speakers are unseen in the training dataset has become a new research direction. Using speaker embeddings instead of one-hot vectors to represent speaker identity is a key point, which makes VC models work on unseen speakers. In our work, a newly designed neural network was used to adjust the speaker embeddings of unseen speakers. This enables speaker embeddings to perform better on zero-shot VC. In addition, disentangled representation of features is the mainstream method to achieve zero-shot VC. In terms of input features of VC model, we use Mel-cepstral and F0 as simple acoustic features (SAF) rather than Mel-spectrograms. This avoids F0 conflicts in decoder that existed in the previous methods. The evaluations demonstrate that our proposed methods improve the quality of converted speech in terms of naturalness and similarity. Zhiyuan Tan 0003, Jianguo Wei, Junhai Xu, Wenhuan Lu |
ICASSP | 5 |
| 2021 | Residual Learning Diagnosis Detection: An Advanced Residual Learning Diagnosis Detection System for COVID-19 in Industrial Internet of ThingsabstractDue to the fast transmission speed and severe health damage, COVID-19 has attracted global attention. Early diagnosis and isolation are effective and imperative strategies for epidemic prevention and control. Most diagnostic methods for the COVID-19 is based on nucleic acid testing (NAT), which is expensive and time-consuming. To build an efficient and valid alternative of NAT, this article investigates the feasibility of employing computed tomography images of lungs as the diagnostic signals. Unlike normal lungs, parts of the lungs infected with the COVID-19 developed lesions, ground-glass opacity, and bronchiectasis became apparent. Through a public dataset, in this article, we propose an advanced residual learning diagnosis detection (RLDD) scheme for the COVID-19 technique, which is designed to distinguish positive COVID-19 cases from heterogeneous lung images. Besides the advantage of high diagnosis effectiveness, the designed residual-based COVID-19 detection network can efficiently extract the lung features through small COVID-19 samples, which removes the pretraining requirement on other medical datasets. In the test set, we achieve an accuracy of 91.33%, a precision of 91.30%, and a recall of 90%. For the batch of 150 samples, the assessment time is only 4.7 s. Therefore, RLDD can be integrated into the application programming interface and embedded into the medical instrument to improve the detection efficiency of COVID-19. Mingdong Zhang, Ronghe Chu, Chaoyu Dong, Jianguo Wei, Wenhuan Lu, Naixue Xiong |
IEEE Trans. Ind. Informatics | 5 |
| 2020 | ARET: Aggregated Residual Extended Time-Delay Neural Networks for Speaker Verification
Ruiteng Zhang, Jianguo Wei, Wenhuan Lu, Longbiao Wang, Meng Liu 0017, Lin Zhang 0054, Jiayu Jin, Junhai Xu |
INTERSPEECH | 3 |
| 2019 | Breast Cancer Detection Based on Merging Four Modes MRI Using Convolutional Neural NetworksabstractThe objective of the study is to develop a framework for automatic breast cancer detection with merging four imaging modes. Attempts were made for tumor classification and segmentation; using a multi-parametric Magnetic Resonance Imaging (MRI) method on breast tumors. MRI data of the breast were obtained from 67 subjects with a 1.5T-MRI scanner. Four imaging modes: were T1 weighted, T2 weighted, Diffusion Weighted and eTHRIVE sequences, and dynamic-contrast-enhanced(DCE)-MRI parameters are acquired. The proposed four-mode linkage backbone in tumor classification, which overcomes the limitations of single-modality image detection and simulates actual diagnosis processes by clinicians, achieves the accuracy of 0.942. The proposed automatic segmentation approach is performed by a refined U-Net architecture, and the result improved segmentation performance significantly. The combination of four-mode linkage classification backbone and improved segmentation network for breast cancer detection forms a computer-aided detection (CAD) system that corresponds to the actual clinical diagnosis work. Wenhuan Lu, Hong Yu 0017, Naixue Xiong, Jianguo Wei |
ICASSP | 1 |
| 2019 | Acoustic and Articulatory Study of Ewe Vowels: A Comparative Study of Male and Female
Kowovi Comivi Alowonou, Jianguo Wei, Wenhuan Lu, Kiyoshi Honda, Jianwu Dang 0001 |
INTERSPEECH | 3 |
| 2019 | Individual Difference of Relative Tongue Size and its Acoustic Effects
Chongke Bi, Kiyoshi Honda, Wenhuan Lu, Jianguo Wei |
INTERSPEECH | 4 |
| 2019 | How to Improve Semantics Understanding of Word CloudsabstractWord cloud is a text visualization technique which is widely applied in helping improve semantic understanding about target materials. One of the most important features is the font size, which represents words frequencies of a document. As the result, in this paper, we explore how to set font sizes of words, and its influence on semantic understanding through people's performance with qualitative and controlled experiments. Adopting an machine learning algorithm LDA (Latent Dirichlet Allocation) topic model, we quantify semantics of the document and judge participants' accuracy performance. The experimental results show the influence of different font size on semantic understanding performance and provide insights for ways in promoting semantic understanding of word cloud. Jie Li 0006, Wenhuan Lu, Yi Chen 0007, Kang Zhang 0001, Yan Li 0080 |
VINCI | 3 |
| 2019 | LSTM-EFG for wind power forecasting based on sequential correlation features
Jie Gao 0008, Mei Yu 0004, Wenhuan Lu, Mankun Zhao, Jie Zhang 0003, Zhuo Zhang 0003 |
Future Gener. Comput. Syst. | 4 |
| 2018 | RE-CNN: A Robust Convolutional Neural Networks for Image Recognition
Wenhuan Lu, Naixue Xiong, Jianguo Wei |
ICONIP (1) | 2 |
| 2018 | Study of articulators' contribution and compensation during speech by articulatory speech recognition
Jianguo Wei, Jingshu Zhang, Qiang Fang 0003, Wenhuan Lu, Kiyoshi Honda, Xugang Lu |
Multim. Tools Appl. | 5 |
| 2017 | Using Deep Learning for Community Discovery in Social NetworksabstractCommunity detection is an important task in social network analysis. Existing methods typically use the topological information alone, and ignore the rich information available in the content data. Recently, some researchers have noticed that user profiles can also benefit to community detection, and hence the combination of topology and node contents has become a new hot topic. Some methods using both topology and content have been proposed. However, they often suffer from two drawbacks: 1) they cannot extract a potential deep representation of the network; 2) they cannot automatically weight different information sources with adequate balance parameters. To overcome these issues, we propose a deep integration representation (DIR) algorithm via deep joint reconstruction, which is motivated by the similarity between deep feedforward auto-encoders and spectral clustering in terms of matrix reconstruction. Thanks to spectral clustering which is one of the best community detection methods, the proposed new method is also good at community discovery task. In addition, DIR has further benefit because it not only provides a nonlinear and deep representation of the network, but also learns the most suitable balance between different components automatically. We compare the proposed new approach with nine state-of-the-art community detection methods on eight real relatively large networks. The experimental results show the definite superiority of this new approach. Di Jin 0001, Meng Ge, Wenhuan Lu, Dongxiao He, Françoise Fogelman-Soulié |
ICTAI | 4 |
| 2017 | Identification of Generalized Communities with Semantics in Networks with ContentabstractDiscovery of communities in networks is a fundamental data analysis task. Recently, researchers have tried to improve its performance by exploiting node contents, and further interpret the communities using the derived semantics. However, the existing methods typically assume that the communities are assortative (i.e. members of each group are mostly connected to other members of the same group), and are unable to find the generalized community structure, e.g. structures with either assortative or disassortative communities (i.e. vertices of the same group have most of their connections outside their group), or a combination. In addition, these methods often assume that the network topology and node contents share the same group memberships, and thus cannot perform well when the contents mismatch with network structure. Also, they are limited to using only one topic to interpret each community. To address these two issues, we propose a new generative probabilistic model which is learned by using a nested expectation-maximization algorithm. It describes the generalized communities (based on network) and the content clusters (based on contents) separately, and further explores and models their correlation to improve as much as possible each of the communities and clusters based on the other. By depicting and utilizing this correlation, our model is not only robust with respect to the above problems, but is also able to interpret each community using more than one topic, which provides richer explanations. We validate the robustness of this proposed new approach on an artificial benchmark, and test its interpretability using a case study analysis. We finally show its definite superiority for community detection by comparing with seven state-of-the-art algorithms on eight real networks. Di Jin 0001, Xiaobao Wang, Dongxiao He, Wenhuan Lu, Françoise Fogelman-Soulié, Jianwu Dang 0001 |
ICTAI | 4 |
| 2017 | Acoustic VR in the mouth: A real-time speech-driven visual tongue systemabstractWe propose an acoustic-VR system that converts acoustic signals of human language (Chinese) to realistic 3D tongue animation sequences in real time. It is known that directly capturing the 3D geometry of the tongue at a frame rate that matches the tongue's swift movement during the language production is challenging. This difficulty is handled by utilizing the electromagnetic articulography (EMA) sensor as the intermediate medium linking the acoustic data to the simulated virtual reality. We leverage Deep Neural Networks to train a model that maps the input acoustic signals to the positional information of pre-defined EMA sensors based on 1,108 utterances. Afterwards, we develop a novel reduced physics-based dynamics model for simulating the tongue's motion. Unlike the existing methods, our deformable model is nonlinear, volume-preserving, and accommodates collision between the tongue and the oral cavity (mostly with the jaw). The tongue's deformation could be highly localized which imposes extra difficulties for existing spectral model reduction methods. Alternatively, we adopt a spatial reduction method that allows an expressive subspace representation of the tongue's deformation. We systematically evaluate the simulated tongue shapes with real-world shapes acquired by MRI/CT. Our experiment demonstrates that the proposed system is able to deliver a realistic visual tongue animation corresponding to a user's speech signal. Ran Luo 0001, Qiang Fang 0003, Jianguo Wei, Wenhuan Lu, Weiwei Xu 0003, Yin Yang 0002 |
VR | 4 |
| 2017 | Parameterization of LSB in Self-Recovery Speech Watermarking Framework in Big Data MiningabstractThe privacy is a major concern in big data mining approach. In this paper, we propose a novel self-recovery speech watermarking framework with consideration of trustable communication in big data mining. In the framework, the watermark is the compressed version of the original speech. The watermark is embedded into the least significant bit (LSB) layers. At the receiver end, the watermark is used to detect the tampered area and recover the tampered speech. To fit the complexity of the scenes in big data infrastructures, the LSB is treated as a parameter. This work discusses the relationship between LSB and other parameters in terms of explicit mathematical formulations. Once the LSB layer has been chosen, the best choices of other parameters are then deduced using the exclusive method. Additionally, we observed that six LSB layers are the limit for watermark embedding when the total bit layers equaled sixteen. Experimental results indicated that when the LSB layers changed from six to three, the imperceptibility of watermark increased, while the quality of the recovered signal decreased accordingly. This result was a trade-off and different LSB layers should be chosen according to different application conditions in big data infrastructures. Zhanjie Song, Wenhuan Lu, Daniel Sun 0004, Jianguo Wei |
Secur. Commun. Networks | 3 |
| 2016 | A New Model for Acoustic Wave Propagation and Scattering in the Vocal Tract
Jianguo Wei, Wendan Guan, Darcy Qingzhi Hou, Dingyi Pan, Wenhuan Lu, Jianwu Dang 0001 |
INTERSPEECH | 5 |
| 2016 | Morphological normalization of vowel images for articulatory speech recognition
Jianguo Wei, Jingshu Zhang, Qiang Fang 0003, Wenhuan Lu |
J. Vis. Commun. Image Represent. | 5 |
| 2016 | Mapping ultrasound-based articulatory images and vowel sounds with a deep neural network framework
Jianguo Wei, Qiang Fang 0003, Xinyuan Zheng, Wenhuan Lu, Jianwu Dang 0001 |
Multim. Tools Appl. | 4 |
| 2016 | Multi-modal recording and modeling of vocal tract movements
Jianguo Wei, Song Wang 0005, Wenhuan Lu, Darcy Qingzhi Hou, Qiang Fang 0003, Jianwu Dang 0001 |
Multim. Tools Appl. | 3 |
| 2015 | Combined cine- and tagged-MRI for tracking landmarks on the tongue surface
Honghao Bao, Wenhuan Lu, Kiyoshi Honda, Jianguo Wei, Qiang Fang 0003, Jianwu Dang 0001 |
INTERSPEECH | 2 |
| 2013 | An anisotropic diffusion filter based on multidirectional separability
Jianguo Wei, Xin Wang 0037, Wenhuan Lu, Qiang Fang 0003, Jianwu Dang 0001 |
INTERSPEECH | 4 |
| 2012 | An ontological approach to support legal information modeling
Wenhuan Lu, Naixue Xiong, Doo-Soon Park |
J. Supercomput. | 1 |
| 2007 | A Conceptual Model Makes Curriculum Evolution in Higher Education Smooth
Wenhuan Lu, Mitsuru Ikeda, Koichiro Ochimizu, Shintaro Kitayama |
ICCE | 1 |
| 2005 | An Intention-oriented Model of Copyright Law for e-Learning: - International Semantic Mapping of Copyright Laws Based on A Copyright Ontology
Wenhuan Lu, Mitsuru Ikeda |
ICCE | 1 |