EDBT 2026 Demo / reviewers in the wild / expert
Taihao Li
dblp:128/6519
· DBLP profile ↗
44ranked-venue papers
3as first author
40since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 23 · 1 first-author · 22 since 2021Graphics, computer vision, multimedia, augmented reality and games · 22 · 1 first-author · 20 since 2021Databases, data management, data science and information retrieval · 5 · 5 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Microblog sentiment classification via a multilayer graph with social and semantic representations using hyperbolic learning
Xiaomei Zou, Taihao Li, Shoukang Han |
Inf. Sci. | 2 |
| 2026 | GAN semantics for personalized facial beauty synthesis and enhancement
Irina Lebedeva 0001, Fangli Ying, Yi Guo 0009, Taihao Li |
J. Vis. Commun. Image Represent. | 4 |
| 2026 | Think-Before-Draw: Decomposing emotion semantics for fine-grained controllable generation of expressive talking heads
Hanlei Shi, Leyuan Qu, Yu Liu 0132, Linlin Gong, Yuhua Zheng, Taihao Li |
Pattern Recognit. | 7 |
| 2025 | Label Semantic-Driven Contrastive Learning for Speech Emotion Recognition
Jiaxi Hu, Leyuan Qu, Haoxun Li, Taihao Li |
INTERSPEECH | 4 |
| 2025 | EME-TTS: Unlocking the Emphasis and Emotion Link in Speech Synthesis
Haoxun Li, Leyuan Qu, Jiaxi Hu, Taihao Li |
INTERSPEECH | 4 |
| 2025 | HOPE: Hierarchical Fusion for Optimized and Personality-Aware Estimation of DepressionabstractDepression detection remains challenged by generalized modeling approaches that fail to account for individual heterogeneity. To address this, the Multimodal Personality-aware Depression Detection (MPDD) Challenge introduced personalized features into the modeling process, aiming to better capture individual variability. However, the baseline models still exhibit two critical limitations: the neglect of textual semantics embedded in audio, and inconsistent predictions for the same subject across tasks and samples. Motivated by these limitations, we introduce HOPE (Hierarchical fusion for Optimized and Personality-aware Estimation of Depression), a unified framework for consistent, subject-level depression estimation. HOPE first employs a Latent Semantic Projection (LSP) module to reconstruct textual semantics from audio features when transcripts are unavailable. It then introduces a consistency-aware integration mechanism that hierarchically fuses multi-branch predictions to resolve inter-task and inter-sample contradictions. HOPE achieved first place in the MPDD Challenge Young Track, demonstrating strong cross-modal learning capabilities and consistent, subject-level depression prediction. Hanlei Shi, Yu Liu 0132, Haoxun Li, Jiaxi Hu, Leyuan Qu, Taihao Li |
ACM Multimedia | 7 |
| 2025 | tanh As a robust feature scaling method in training deep learning models with imbalanced data
Aijia Yang, Huai Chen, Taihao Li, Shupeng Liu, Xiaoyin Xu |
Pattern Recognit. | 4 |
| 2025 | A Residual Multi-Scale Convolutional Neural Network With Transformers for Speech Emotion RecognitionabstractThe great variety of human emotional expression as well as the differences in the ways they perceive and annotate them make Speech Emotion Recognition (SER) an ambiguous and challenging task. With the development of deep learning, long-term progress has been made in SER systems. However, the existing convolutional neural networks present certain limitations, such as their inability to well capture global features, which contain important emotional information. Moreover, the position encoding in the Transformer structure is relatively fixed and only encodes the time domain dimension, which cannot effectively obtain the position information of discriminative features in the frequency domain dimension. In order to overtake these limitations, we propose an end-to-end Residual Multi-Scale Convolutional Neural Networks (RMSCNN) with Transformer model network. Simultaneously, to further validate the effectivenessof RMSCNN in extracting multi-scale features and delivering pertinent emotion localization data, we developed the RMSC_down network in conjunction with the Wav2Vec 2.0 model. The results of the prediction of Arousal, Valenceand Dominanceon the popular corpora demonstrate the superiority and robustness of our approach for SER, showing an improvement of the recognition accuracy in the public dataset MSP-Podcast 1.9 version. Tianhao Yan, Emilia Parada-Cabaleiro, Jianhua Tao 0001, Taihao Li, Björn W. Schuller |
IEEE Trans. Affect. Comput. | 5 |
| 2025 | Disentanglement of Prosody Representations via Diffusion Models and Scheduled Gradient ReversalabstractProsody plays a fundamental role in human speech and communication, facilitating intelligibility and conveying emotional and cognitive states. Extracting accurate prosodic information from speech is vital for building assistive technology, such as controllable speech synthesis, speaking style transfer, and speech emotion recognition (SER). However, it is challenging to disentangle speaker-independent prosody representations since prosodic attributes, such as intonation, excessively entangle with speaker-specific attributes, e.g., pitch. In this article, we propose a novel model, called Diffsody, to disentangle and refine prosody representations: 1) to disentangle prosody representations, we leverage the expressive generative ability of a diffusion model by conditioning it on quantified semantic information and pretrained speaker embeddings. Additionally, a prosody encoder automatically learns prosody representations used for spectrogram reconstruction in an unsupervised fashion; and 2) to refine and learn speaker-invariant prosody representations, a scheduled gradient reversal layer (sGRL) is proposed and integrated into the prosody encoder of Diffsody. We extensively evaluate Diffsody through qualitative and quantitative means. t-SNE visualization and speaker verification experiments demonstrate the efficacy of the sGRL method in preventing speaker-specific information leakage. Experimental results on speaker-independent SER and automatic depression detection (ADD) tasks demonstrate that Diffsody can efficiently factorize speaker-independent prosody representations, resulting in a significant boost in SER and ADD. In addition, Diffsody synergistically integrates with the semantic representation model WavLM, which leads to a discernibly elevated performance, outperforming contemporary methods in both SER and ADD tasks. Furthermore, the Diffsody model exhibits promising potential for various practical applications, such as voice or style conversion. Some audio samples can be found on our https://leyuanqu.github.io/Diffsody/demo website. Leyuan Qu, Cornelius Weber, Wei Wang 0310, Jia Jin, Yingming Gao, Taihao Li, Stefan Wermter |
IEEE Trans. Neural Networks Learn. Syst. | 6 |
| 2024 | CORECODE: A Common Sense Annotated Dialogue Dataset with Benchmark Tasks for Chinese Large Language ModelsabstractAs an indispensable ingredient of intelligence, commonsense reasoning is crucial for large language models (LLMs) in real-world scenarios. In this paper, we propose CORECODE, a dataset that contains abundant commonsense knowledge manually annotated on dyadic dialogues, to evaluate the commonsense reasoning and commonsense conflict detection capabilities of Chinese LLMs. We categorize commonsense knowledge in everyday conversations into three dimensions: entity, event, and social interaction. For easy and consistent annotation, we standardize the form of commonsense knowledge annotation in open-domain dialogues as "domain: slot = value". A total of 9 domains and 37 slots are defined to capture diverse commonsense knowledge. With these pre-defined domains and slots, we collect 76,787 commonsense knowledge annotations from 19,700 dialogues through crowdsourcing. To evaluate and enhance the commonsense reasoning capability for LLMs on the curated dataset, we establish a series of dialogue-level reasoning and detection tasks, including commonsense knowledge filling, commonsense knowledge generation, commonsense conflict phrase detection, domain identification, slot identification, and event causal inference. A wide variety of existing open-source Chinese LLMs are evaluated with these tasks on our dataset. Experimental results demonstrate that these models are not competent to predict CORECODE's plentiful reasoning content, and even ChatGPT could only achieve 0.275 and 0.084 accuracy on the domain identification and slot identification tasks under the zero-shot setting. We release the data and codes of CORECODE at https://github.com/danshi777/CORECODE to promote commonsense reasoning evaluation and study of LLMs in the context of daily conversations. Dan Shi 0001, Chaobin You, Jiantao Huang, Taihao Li, Deyi Xiong |
AAAI | 4 |
| 2024 | RedCore: Relative Advantage Aware Cross-Modal Representation Learning for Missing Modalities with Imbalanced Missing RatesabstractMultimodal learning is susceptible to modality missing, which poses a major obstacle for its practical applications and, thus, invigorates increasing research interest. In this paper, we investigate two challenging problems: 1) when modality missing exists in the training data, how to exploit the incomplete samples while guaranteeing that they are properly supervised? 2) when the missing rates of different modalities vary, causing or exacerbating the imbalance among modalities, how to address the imbalance and ensure all modalities are well-trained. To tackle these two challenges, we first introduce the variational information bottleneck (VIB) method for the cross-modal representation learning of missing modalities, which capitalizes on the available modalities and the labels as supervision. Then, accounting for the imbalanced missing rates, we define relative advantage to quantify the advantage of each modality over others. Accordingly, a bi-level optimization problem is formulated to adaptively regulate the supervision of all modalities during training. As a whole, the proposed approach features Relative advantage aware Cross-modal representation learning (abbreviated as RedCore) for missing modalities with imbalanced missing rates. Extensive empirical results demonstrate that RedCore outperforms competing models in that it exhibits superior robustness against either large or imbalanced missing rates. The code is available at: https://github.com/sunjunaimer/RedCore. Shoukang Han, Yu-Ping Ruan, Taihao Li |
AAAI | 5 |
| 2024 | Least-Effort Adversarial Attack Against Gait-Based Identity Recognition SystemabstractIn this paper, we propose a least-effort adversarial attack against a gait-based identity recognition system (GIRS). Specifically, we leverage a bilevel optimization framework to characterize this leader-follower Stackelberg game between the attacker and gait recognizer to pursue the Nash equilibrium state with hybrid attack intentions of maximum effectiveness and minimum cost. Then, to tractably solve this NP-hard bilevel problem, we present the duality theory and linearization representation technique to reformulate a computable mixed-integer program and derive a globally optima. Finally, we perform comparison experiments on two public datasets to verify the validity of our attack strategy in stealthily inducing mistaken identities. Empirical results can also shed light on an attack mitigation measure to secure a GIRS for the privacy protection. Datian Peng, Taihao Li |
ICASSP | 3 |
| 2024 | Improving Speech Emotion Recognition with Unsupervised Speaking Style TransferabstractHumans can effortlessly modify various prosodic attributes, such as the placement of stress and the intensity of sentiment, to convey a specific emotion while maintaining consistent linguistic content. Motivated by this capability, we propose EmoAug, a novel style transfer model designed to enhance emotional expression and tackle the data scarcity issue in speech emotion recognition tasks. EmoAug consists of a semantic encoder and a paralinguistic encoder that represent verbal and non-verbal information respectively. Additionally, a decoder reconstructs speech signals by conditioning on the aforementioned two information flows in an unsupervised fashion. Once training is completed, EmoAug enriches expressions of emotional speech with different prosodic attributes, such as stress, rhythm and intensity, by feeding different styles into the paralinguistic encoder. EmoAug enables us to generate similar numbers of samples for each class to tackle the data imbalance issue as well. Experimental results on the IEMOCAP dataset demonstrate that EmoAug can successfully transfer different speaking styles while retaining the speaker identity and semantic content. Furthermore, we train a SER model with data augmented by EmoAug and show that the augmented model not only surpasses the state-of-the-art supervised and self-supervised methods but also overcomes overfitting problems caused by data imbalance. Some audio samples can be found on our demo website1. Leyuan Qu, Wei Wang 0310, Cornelius Weber, Pengcheng Yue, Taihao Li, Stefan Wermter |
ICASSP | 5 |
| 2024 | Fusing Modality-Specific Representations and Decisions for Multimodal Emotion RecognitionabstractMultimodal emotion recognition (MER) is important for building humanoid chatbots and has gained increasing attention in recent years. Existing studies have proven that extracting better modality-specific representations, which keep both commonality and individuality information of different modalities, is important for the MER task. However, all these works are restricted in making final predictions based on fusing modality-specific representations, and the effectiveness of the modality-specific decisions has not been studied. In this paper, we propose for the first time to fuse both the modality-specific representations and decisions for the MER task and design a bi-channel fusing network (BCFN). Specifically, a BCFN model first extracts and mixes the modality-specific representations and decisions in two convolutional blocks respectively, and then fuses the two joint multimodal features for the final decision. Extensive experiments are conducted on two MER benchmark datasets with state-of-the-art models employed as baselines for comparison. The empirical results demonstrate the superiority of our proposed BCFN model and confirm the effectiveness of incorporating modality-specific decisions for the MER task. Yu-Ping Ruan, Shoukang Han, Taihao Li |
ICASSP | 3 |
| 2024 | Graph-Enhanced Hybrid Sampling for Multi-Armed Bandit RecommendationabstractGraph-based multi-armed bandit algorithms utilize the relationship between users to select the best item to recommend for maximal reward, which is decided by items’ features and un-known users’ preferences. Therefore, the precise estimation of users’ preferences is fairly important and indispensable for bandit sampling, though it is not the ultimate target. However, existing algorithms generally neglect this crucial point and utilize reward maximization as objective in the first beginning, using inaccurate estimation as input, which deteriorates the performance from a long-term perspective. In this paper, we will propose one hybrid sampling framework for bandit selection, which at first purely focuses on the performance of estimation and then on the performance of reward maximization. Specifically, we propose an ‘unsupervised’ bandit selection objective to minimize expected estimation error, which doesn’t take users’ preferences as input and suppresses an approximate upper-bound of cumulative regret. Then, we design a low-complexity selection algorithm to optimize this formulated problem with simple multiplications between items’ features and users’ graphical relations. Subsequently, for reward maximization, we cascade one graph-based algorithm to find the following bandits on the basis of our proposed warm-starts. Extensive experiments on different graphs indicate that our proposed hybrid framework is substantially better than existing popular methods in terms of recommendation performance. Taihao Li, Wuyue Zhang, Xue Zhang 0008, Cheng Yang 0003 |
ICASSP | 2 |
| 2024 | Multi-Modal Emotion Recognition Using Multiple Acoustic Features and Dual Cross-Modal TransformerabstractMulti-modal emotion recognition (MER) using speech and text has attracted extensive attention because of the easy availability of data for these two modalities. Recently, the self-surprised learning (SSL) pre-trained model has become the state-of-the-art (SOTA) method for the extraction of acoustic and textual features. However, the SSL speech representation may lose some important paralinguistic information, resulting in limited speech knowledge for MER. In this paper, we propose to adopt two kinds of acoustic features (i.e., the SSL representation and the spectral feature) as inputs to comprehensively extract speech characteristics. In addition, a dual cross-modal Transformer module is presented to model the interaction on the unaligned sequences between the textual feature and two acoustic features. Moreover, we introduce a blended loss including two uni-modal losses to better extract the uni-modal information. Experiments conducted on the widely used IEMOCAP dataset indicate that our proposed method achieves the SOTA performance compared with previous methods. Pengcheng Yue, Leyuan Qu, Taihao Li, Yu-Ping Ruan |
ICASSP | 4 |
| 2024 | Domain Knowledge Enhanced Vision-Language Pretrained Model for Dynamic Facial Expression RecognitionabstractDynamic facial expression recognition (DFER) is a rapidly developing field that focuses on recognizing facial expressions in video sequences. However, the complex temporal modeling caused by noisy frames, along with the limited training data significantly hinder the further development of DFER. Previous efforts in this domain have been limited as they tackled these issues separately. Inspired by recent advances of pretrained vision-language models (e.g., CLIP), we propose to leverage it to jointly address the two limitations in DFER. Since the raw CLIP model lacks the ability to model temporal relationships and determine the optimal task-related textual prompts, we utilize DFER-specific domain knowledge, including characteristics of temporal correlations and relationships between facial behavior descriptions at different levels, to guide the adaptation of CLIP to DFER. Specifically, we propose enhancements to CLIP's visual encoder through the design of a hierarchical video encoder that captures both short- and long-term temporal correlations in DFER. Meanwhile, we align facial expressions with action units through prior knowledge to construct semantically rich textual prompts, which are further enhanced with visual contents. Furthermore, we introduce a class-aware consistency regularization mechanism that adaptively filters out noisy frames, bolstering the model's robustness against interference. Extensive experiments on three in-the-wild dynamic facial expression datasets demonstrate that our method outperforms the state-of-the-art DFER approaches. The code is available at https://github.com/liliupeng28/DK-CLIP. Liupeng Li, Yuhua Zheng, Shupeng Liu, Xiaoyin Xu, Taihao Li |
ACM Multimedia | 5 |
| 2024 | Light-weight residual convolution-based capsule network for EEG emotion recognition
Cunhang Fan, Jinqin Wang, Xiaoke Yang, Guanxiong Pei, Taihao Li, Zhao Lv |
Adv. Eng. Informatics | 6 |
| 2024 | Revisiting 3D visual grounding with Context-aware Feature Aggregation
Peng Guo 0011, Hongyuan Zhu 0002, Hancheng Ye, Taihao Li, Tao Chen 0003 |
Neurocomputing | 4 |
| 2024 | Multistage guidance on the diffusion model inspired by human artists' creative thinkingabstract目前文本生成图像的研究已显示出与普通画家类似的水平,但与艺术家绘画水平相比仍有很大改进空间;艺术家水平的绘画通常将多个意象的特征融合到一个意象中,以表示多层次语义信息。在预实验中,我们证实了这一点,并咨询了3个具有不同艺术欣赏能力的群体的意见,以确定画家和艺术家之间绘画水平的区别。之后,利用这些观点帮助人工智能绘画系统从普通画家水平的图像生成改进为艺术家水平的图像生成。具体来说,提出一种无需任何进一步预训练的、基于文本的多阶段引导方法,帮助扩散模型在生成的图像中向多层次语义表示迈进。实验中的机器和人工评估都验证了所提方法的有效性。此外,与之前单阶段引导方法不同,该方法能够通过控制不同阶段之间的指导步数来控制各个意象特征在绘画中的表现程度。 Huanghuang Deng, Taihao Li |
Frontiers Inf. Technol. Electron. Eng. | 3 |
| 2024 | VLP2MSA: Expanding vision-language pre-training to multimodal sentiment analysis
Guofeng Yi, Cunhang Fan, Kang Zhu, Zhao Lv, Shan Liang 0007, Zhengqi Wen, Guanxiong Pei, Taihao Li, Jianhua Tao 0001 |
Knowl. Based Syst. | 8 |
| 2024 | Ecarnet: enhanced clue-ambiguity reasoning network for multimodal fake news detection
Shannan Zhong, Shu-Juan Peng, Xin Liu 0011, Lei Zhu 0002, Xing Xu 0001, Taihao Li |
Multim. Syst. | 6 |
| 2024 | Vote2Cap-DETR++: Decoupling Localization and Describing for End-to-End 3D Dense Captioningabstract3D dense captioning requires a model to translate its understanding of an input 3D scene into several captions associated with different object regions. Existing methods adopt a sophisticated "detect-then-describe" pipeline, which builds explicit relation modules upon a 3D detector with numerous hand-crafted components. While these methods have achieved initial success, the cascade pipeline tends to accumulate errors because of duplicated and inaccurate box estimations and messy 3D scenes. In this paper, we first propose Vote2Cap-DETR, a simple-yet-effective transformer framework that decouples the decoding process of caption generation and object localization through parallel decoding. Moreover, we argue that object localization and description generation require different levels of scene understanding, which could be challenging for a shared set of queries to capture. To this end, we propose an advanced version, Vote2Cap-DETR++, which decouples the queries into localization and caption queries to capture task-specific features. Additionally, we introduce the iterative spatial refinement strategy to vote queries for faster convergence and better localization performance. We also insert additional spatial information to the caption head for more accurate descriptions. Without bells and whistles, extensive experiments on two commonly used datasets, ScanRefer and Nr3D, demonstrate Vote2Cap-DETR and Vote2Cap-DETR++ surpass conventional "detect-then-describe" methods by a large margin. Sijin Chen, Hongyuan Zhu 0002, Mingsheng Li, Xin Chen 0040, Peng Guo 0011, Yinjie Lei, Gang Yu 0002, Taihao Li, Tao Chen 0003 |
IEEE Trans. Pattern Anal. Mach. Intell. | 8 |
| 2024 | OLCH: Online Label Consistent Hashing for streaming cross-modal retrieval
Shu-Juan Peng, Jinhan Yi, Xin Liu 0011, Yiu-Ming Cheung, Zhen Cui 0001, Taihao Li |
Pattern Recognit. | 6 |
| 2024 | Use estimated signal and noise to adjust step size for image restoration
Shupeng Liu, Taihao Li, Huai Chen, Xiaoyin Xu |
Pattern Recognit. Lett. | 3 |
| 2024 | Design of a differentiable L-1 norm for pattern recognition and machine learning
Min Zhang 0069, Taihao Li, Shupeng Liu, Xianfeng Gu, Xiaoyin Xu |
Pattern Recognit. Lett. | 4 |
| 2024 | Disentangling Prosody Representations With Unsupervised Speech ReconstructionabstractHuman speech can be characterized by different components, including semantic content, speaker identity and prosodic information. Significant progress has been made in disentangling representations for semantic content and speaker identity in Automatic Speech Recognition (ASR) and speaker verification tasks respectively. However, it is still an open challenging research question to extract prosodic information because of the intrinsic association of different attributes, such as timbre and rhythm, and because of the need for supervised training schemes to achieve robust large-scale and speaker-independent ASR. The aim of this paper is to address the disentanglement of emotional prosody from speech based on unsupervised reconstruction. Specifically, we identify, design, implement and integrate three crucial components in our proposed speech reconstruction model Prosody2Vec: (1) a unit encoder that transforms speech signals into discrete units for semantic content, (2) a pretrained speaker verification model to generate speaker identity embeddings, and (3) a trainable prosody encoder to learn prosody representations. We first pretrain the Prosody2Vec representations on unlabelled emotional speech corpora, then fine-tune the model on specific datasets to perform Speech Emotion Recognition (SER) and Emotional Voice Conversion (EVC) tasks. Both objective (weighted and unweighted accuracies) and subjective (mean opinion score) evaluations on the EVC task suggest that Prosody2Vec effectively captures general prosodic features that can be smoothly transferred to other emotional speech. In addition, our SER experiments on the IEMOCAP dataset reveal that the prosody features learned by Prosody2Vec are complementary and beneficial for the performance of widely used speech pretraining models and surpass the state-of-the-art methods when combining Prosody2Vec with HuBERT representations. Some audio samples can be found on our demo website Leyuan Qu, Taihao Li, Cornelius Weber, Theresa Pekarek-Rosin, Fuji Ren, Stefan Wermter |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2023 | Layer-wise Fusion with Modality Independence Modeling for Multi-modal Emotion RecognitionabstractJun Sun, Shoukang Han, Yu-Ping Ruan, Xiaoning Zhang, Shu-Kai Zheng, Yulong Liu, Yuxin Huang, Taihao Li. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Shoukang Han, Yu-Ping Ruan, Shukai Zheng, Taihao Li |
ACL (1) | 8 |
| 2023 | Social Links Enhanced Microblog Sentiment Analysis: Integrating Link Prediction and Sentiment Connection Weights
Xiaomei Zou, Taihao Li |
DEXA (1) | 2 |
| 2023 | Capsule Network with Label Dependency Modeling for Multi-Label Emotion ClassificationabstractThis paper proposes a simple-yet-efficient model, called capsule network with label dependency modeling (CapsLDM), for the task of multi-label emotion classification (MLEC) in text, in which multiple emotion categories can be assigned to the input data instance (e.g., a sentence). Unlike the traditional single-label emotion classification, the modeling of label (i.e., emotion) dependency plays an important role in MLEC, since the co-existing emotions in an utterance are not independent of each other. The capsule network has been successfully applied to many multi-label classification scenarios, however, the modeling of label dependency has not been considered in existing work. In our proposed CapsLDM model, we add similarity regularization terms on both the dynamic routing weights and the instance vectors of emotion capsules by exploiting the co-occurrence information of emotion labels, which resembles the dependency between different emotion categories for a certain input instance. Extensive experiments are conducted on four MLEC benchmark datasets with state-of-the-art models employed as baselines for comparison. The empirical results demonstrate the superiority of our proposed CapsLDM model and confirm the effectiveness of label dependency modeling in CapsLDM for the MLEC task. Yu-Ping Ruan, Taihao Li |
ECAI | 2 |
| 2023 | Revisit Sampling Theory of Bandlimited Graph Signals: One Bridge Between GSP and DSPabstractSampling of bandlimited (BL) graph signals is one fundamental problem in graph signal processing (GSP), whose underlying kernel is an irregular graph rather than regular 1-D time-series kernel in classical discrete signal processing (DSP). Though there were amounts of sampling objectives and algorithms proposed for BL graph signals, the essential relationship between those sampling objectives in GSP and Nyquist sampling theorem in DSP is still undiscovered. In this paper, we bridge this gap by revisiting sampling theory in GSP thoroughly. In specific, we first figure out that the minimal sample size used in GSP for unique recovery can derive exact Nyquist sampling frequency in DSP when eigenvector matrix is discrete Fourier transform (DFT) matrix. Then, we propose a graph sampling objective for directed cyclic graph as one bridge, which leads to uniform sampling pattern in DSP and has the same optimal solution as other popular graph sampling objectives in GSP, thus connecting GSP and DSP closely via sampling. Finally, we present simulations to demonstrate potential applications inspired by our study, such as fast reconstruction of BL graph signals. Taihao Li, Xue Zhang 0008 |
ICASSP | 2 |
| 2023 | Label-Semantic-Enhanced Online Hashing for Efficient Cross-modal RetrievalabstractExisting online cross-modal hashing methods often treat the semantic label categories independently to correlate the semantically similar data instances, which intrinsically ignore the potential dependency between the label categories and thus fail to capture the discriminative information in the hash code learning process. To alleviate this concern, we explore the inter-dependency between the label categories through their co-occurrence correlation from the label set, and present an efficient Label-Semantic-Enhanced Online Hashing (LSE-OH) method for various cross-modal retrieval task. To be specific, the proposed framework integrates the instance-wise similarity and label-category affinity to incrementally learn the discriminative hash codes for the current arriving data, while updating the hash functions at a streaming manner. Further, an iterative discrete optimization algorithm is derived to mine the inter-dependency between the label categories and discriminatively learn the hash codes without relaxation. Accordingly, the hash codes are adaptively learned online with the high discriminative capability and inter-dependency, while avoiding high computation complexity to process the streaming data. Experimental results show its outstanding performance in comparison with the-state-of-arts. Xueting Jiang, Xin Liu 0011, Yiu-Ming Cheung, Xing Xu 0001, Shukai Zheng, Taihao Li |
ICME | 6 |
| 2023 | Taking a Part for the Whole: An Archetype-agnostic Framework for Voice-Face AssociationabstractVoice-face association is generally specialized as a cross-modal cognitive matching problem, and recent attention has been paid on the feasibility of devising the computational mechanisms for recognizing such associations. Existing works are commonly resorting to the combination of contrastive learning and classification-based loss to correlate the heterogeneous datas. Nevertheless, the reliance on typical features of each category, known as archetypes, derived from the combination suffer from the weak invariance of modality-specific features within the same identity, which might induce a cross-modal joint feature space with calibration deviations. To tackle these problems, this paper presents an efficient Archetype-agnostic framework for reliable voice-face association. First, an Archetype-agnostic Subspace Merging (AaSM) method is carefully designed to perform feature calibration which can well get rid of the archetype dependence to facilitate the mutual perception of datas. Further, an efficient Bilateral Connection Re-gauging scheme is proposed to quantitatively screen and calibrate the biased datas, namely loose pairs that deviate from joint feature space. Besides, an Instance Equilibrium strategy is dynamically derived to optimize the training process on loose data pairs and significantly improve the data utilization. Through the joint exploitation of the above, the proposed framework can well associate the voice-face data to benefit various kinds of cross-modal cognitive tasks. Extensive experiments verify the superiorities of the proposed voice-face association framework and show its competitive performances with the state-of-the-arts. Guancheng Chen, Xin Liu 0011, Xing Xu 0001, Yiu-Ming Cheung, Taihao Li |
ACM Multimedia | 5 |
| 2023 | Optimal Targeted Attacks Against Gait-Based Identity RecognitionabstractTo adequately understand endogenetic vulnerabilities of gait-based identity recognition, we propose an optimal generation strategy for launching the targeted attacks. Specifically, we formulate a bilevel optimization problem to model a Stackelberg game, during which an adversarial interaction process is activated between the leader, i.e., attacker, and the follower, i.e., recognizer. The former intends to undermine the gait signal features for fooling the recognizer, resulting that an unauthenticated user can be misrecognized as a specific authenticated victim. The latter usually adapts a common sparse presentation optimization to exactly match the similar gait signals for recognizing the genuine identities. Further, to tractably solve this NP-hard bilevel problem, we absorb the duality theory to reformulate a calculable second-order cone program. Finally, we perform experiments on public datasets to verify the validity of our strategy in adversarial effectiveness and attack costs. Empirical results shed light on huge potential of our model in the vulnerability assessments of gait-based identity recognition. Taihao Li |
SMC | 3 |
| 2023 | A Fused Speech Enhancement Framework for Robust Speaker VerificationabstractRobust speaker verification (RSV) under noisy con- ditions is still a challenging task. Recently, some task-specific speech enhancement (SE) approaches are proposed and achieve excellent performance on RSV. However, all these works adopt only one kind of SE network and thus can not remove noise from different aspects, limiting the performance of the RSV task. In this letter, we propose a fused SE framework (FSEF) for RSV, which integrates both T-F masking-based and feature mapping- based SE networks to collect complementary information and improve the robustness against noise. Two FESF-RSV systems are constructed based on two kinds of fusion methods: score fusion and feature fusion. In addition, we present a Multi- Scale Attentive Context Aggregation Network (MSACAN) as the backbone structure in the FSEF. The MSACAN can not only extract and fuse multi-scale features adaptively but also enhance speaker characteristics against noise and interfering speakers. Experiments conducted on the noise-simulated VoxCeleb1 dataset demonstrate both the FSEF and the MSACAN can improve the performance of RSV compared to previous approaches. Taihao Li, Junan Zhao, Jing Xu 0008 |
IEEE Signal Process. Lett. | 2 |
| 2022 | Learning with Twin Noisy Labels for Visible-Infrared Person Re-IdentificationabstractIn this paper, we study an untouched problem in visible-infrared person re-identification (VI-ReID), namely, Twin Noise Labels (TNL) which refers to as noisy annotation and correspondence. In brief, on the one hand, it is inevitable to annotate some persons with the wrong identity due to the complexity in data collection and annotation, e.g., the poor recognizability in the infrared modality. On the other hand, the wrongly annotated data in a single modality will eventually contaminate the cross-modal correspondence, thus leading to noisy correspondence. To solve the TNL problem, we propose a novel method for robust VI-ReID, termed DuAlly Robust Training (DART). In brief, DART first computes the clean confidence of annotations by resorting to the memorization effect of deep neural networks. Then, the proposed method rectifies the noisy correspondence with the estimated confidence and further divides the data into four groups for further utilizations. Finally, DART employs a novel dually robust loss consisting of a soft identification loss and an adaptive quadruplet loss to achieve robustness on the noisy annotation and noisy correspondence. Extensive experiments on SYSU-MM01 and RegDB datasets verify the effectiveness of our method against the twin noisy labels compared with five state-of-the-art methods. The code could be accessed from https://github.com/XLearning-SCU/2022-CVPR-DART. Mouxing Yang, Zhenyu Huang 0005, Peng Hu 0002, Taihao Li, Jiancheng Lv 0001, Xi Peng 0001 |
CVPR | 4 |
| 2022 | Hierarchical and Multi-View Dependency Modelling Network for Conversational Emotion RecognitionabstractThis paper proposes a new model, called hierarchical and multi-view dependency modelling network (HMVDM), for the task of emotion recognition in conversations (ERC). The modelling of conversational context plays an important role in ERC, especially for the multi-turn and multi-speaker conversations which hold complex dependency between different speakers. In our proposed HMVDM1, we model the dependency between different speakers at both tokenlevel and utterance-level. Specifically, the HMVDM model has a hierarchical structure with two main modules: 1) token-level dependency modelling module (TDM), which aims to learn the long-range token-level dependency between different utterances in a speaker-aware manner and output the utterance representation; 2) utterance-level dependency modelling module (UDM), which accepts the utterance representation from TDM as inputs and aims to learn the utterance-level dependency from intra-, inter-, and global-speaker(s) view simultaneously. Extensive experiments are conducted on four ERC benchmark datasets with state-of-the-art models employed as baselines for comparison. The empirical results demonstrate the superiority of our proposed HMVDM model and confirm the importance of hierarchical and multi-view context dependency modelling for ERC. Yu-Ping Ruan, Shukai Zheng, Taihao Li, Guanxiong Pei |
ICASSP | 3 |
| 2022 | Inconsistency Distillation For Consistency: Enhancing Multi-View Clustering via Mutual Contrastive Teacher-Student LeaningabstractMulti-view clustering has attracted more attention recently since many real-world data are comprised of different representations or views. Recent multi-view clustering works mainly exploit the instance consistency to obtain the shared representations across different views, and apply a single-view clustering method to perform data partitions. However, these existing methods often ignore the inconsistency of instance associations within the views, which may enlarge the intra-class diversity among the views and therefore degrade the clustering performance. To address this issue, this paper proposes an efficient mutual contrastive teacher-student leaning (MC-TSL) model to enhance the multi-view clustering, which is the first attempt to study the inconsistency distillation for consistency learning. First, the proposed MC-TSL approach exploits a view-specific encoder with two heads, an instance encoding head and a semantic distillation head, respectively, for capturing the consistent and discriminative feature representations. To be specific, the former head exploits a cross-view contrastive learning method to obtain a redundancy-free consistent representation at the instance level, while the latter head designs a mutual teacher-student learning module to capture the intra-view information at semantic level. By training these two heads in an end-to-end manner, the discriminative multi-view embeddings are efficiently obtained and refined by minimizing the weighted sum of the reconstruction loss, contrastive loss and contrast distillation loss. Extensive experiments verify the superiorities of the proposed MC-TSL framework and show its competitive clustering performances. Dunqiang Liu, Shu-Juan Peng, Xin Liu 0011, Lei Zhu 0002, Zhen Cui 0001, Taihao Li |
ICDM | 6 |
| 2022 | Detach and Enhance: Learning Disentangled Cross-modal Latent Representation for Efficient Face-Voice Association and MatchingabstractMany researches in cognitive science have shown that humans often perform face-voice association for various perception tasks, and some recent data mining works have been designed in emulating such ability intelligently. Nevertheless, most methods often suffer from the degraded performance when there exist semantically irrelevant interference factors across different modalities. To alleviate this concern, this paper presents an efficient Disentangled Cross-modal Latent Representation (DCLR) method to adaptively detach the discriminative feature attributes and enhance the face-voice association. To be specific, the proposed DCLR framework consists of two-stage cross-modal disentangling process. First, the former stage employs the supervised contrastive learning to push the representations of face-voice data from the same person closer while pulling those representations of different person away. Then, the latter stage freezes all the parameters of the former stage, and further innovates a multi-layer orthogonal decoupling scheme to learn the disentangled latent representations, while filtering out the modality-dependent irrelevant factors. Besides, the cross-modal reconstruction loss is further utilized to narrow down the semantic gap between heterogeneous feature expressions. Through the joint exploitation of the above, the proposed framework can well associate the face-voice data to benefit various kinds of cross-modal perception tasks. Extensive experiments verify the superiorities of the proposed face-voice association framework and show its competitive performances. Zhenning Yu, Xin Liu 0011, Yiu-Ming Cheung, Minghang Zhu, Xing Xu 0001, Nannan Wang 0001, Taihao Li |
ICDM | 7 |
| 2022 | Twin Contrastive Learning for Online Clustering
Yunfan Li 0003, Mouxing Yang, Dezhong Peng, Taihao Li, Jiantao Huang, Xi Peng 0001 |
Int. J. Comput. Vis. | 4 |
| 2019 | A new design in iterative image deblurring for improved robustness and performanceabstractIn many applications, image deblurring is a pre-requisite to improve the sharpness of an image before it can be further processed. Iterative methods are widely used for deblurring images but care must be taken to ensure that the iterative process is robust, meaning that the process does not diverge and reaches the solution reasonably fast, two goals that sometimes compete against each other. In practice, it remains challenging to choose parameters for the iterative process to be robust. We propose a new approach consisting of relaxed initialization and pixel-wise updates of the step size for iterative methods to achieve robustness. The first novel design of the approach is to modify the initialization of existing iterative methods to stop a noise term from being propagated throughout the iterative process. The second novel design is the introduction of a vectorized step size that is adaptively determined through the iteration to achieve higher stability and accuracy in the whole iterative process. The vectorized step size aims to update each pixel of an image individually, instead of updating all the pixels by the same factor. In this work, we implemented the above designs based on the Landweber method to test and demonstrate the new approach. Test results showed that the new approach can deblur images from noisy observations and achieve a low mean squared error with a more robust performance. Taihao Li, Huai Chen, Shupeng Liu, Shunren Xia, Xinhua Cao, Geoffrey S. Young, Xiaoyin Xu |
Pattern Recognit. | 1 |
| 2018 | Using feature points and angles between them to recognise facial expression by a neural network approachabstractIn this study, the authors propose a neural network (NN) method that uses feature points and the angles formed between the points to recognise facial expressions. Accurate facial expression recognition is an important part of affective computing with many practical applications. Yet, achieving acceptable levels of facial recognition accuracy has proven difficult. Feature points and the distances between the points are used to model basic expressions in NN‐based approaches, but, in some cases, they cannot generate satisfactory performance. They expand on the characterisation of facial expression by considering the angles formed between feature points to augment the amount of information that is sent to the NNs. Furthermore, to circumvent a common challenge in facial expressions recognition, which is the difficulty of differentiating among several expressions, they designed a post‐processing step to assess the output of the NN against a threshold. The whole method makes a decision only when the output of the NN exceeds the threshold. Otherwise, the frame under consideration is assigned to a ‘no decision’ class. They tested our method on the widely used facial expression CK + database and found that it can achieve good accuracy. Taihao Li, Cuifen Du, Tuya Naren, Shupeng Liu, Jianshe Zhou, Xiaoyin Xu |
IET Image Process. | 1 |
| 2017 | Modeling wireless sensor networks radio frequency signal loss in corn environment
He Pan, Xin Wang 0050, Taihao Li |
Multim. Tools Appl. | 4 |
| 2003 | Machine translation method using super-function for mobile terminalabstractThe mobile telephone is the representation of mobile communication. In this paper we proposed a new machine translation method called Super-Function Based Machine Translation (SFBMT) to try to use for mobile terminal. SFBMT uses Super Function(SF) to translate the sentence without thorough syntactic and semantic analysis. And it can translate the sentence fast and high correct. Our experiment shows that this method can get 70% translation precision using the sentences that learned for extracting Super-Function, And 40% translation precision using the sentences that not learned for extracting Super-Function. Taihao Li, Fuji Ren |
SMC | 1 |