Mi-Xiao Hou

dblp:182/8438 · also Mixiao Hou · DBLP profile ↗
← Back
17ranked-venue papers
6as first author
9since 2021 · last 2024
0000-0002-6470-0351ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 8 · 3 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 2 first-authorArtificial intelligence and machine learning · 3 · 1 first-author · 3 since 2021
YearPublicationVenuePosition
2024 Efficient U-Shape Invertible Neural Network for Image Steganography
abstract
Currently, it is challenging to recover high-quality secret images from highly secure stego images while maintaining a limited computational cost for image steganography. This paper proposes an Efficient U-shape Invertible Neural Network (EUIN-Net) for image steganography. Due to the gradual fusion and separation properties of the U-shape invertible mechanism, our EUIN-Net comprehensively couples and decouples the secret-cover information on different scales and depths. Besides, using the skip connections between each pair of U-shape invertible blocks, the long-range dependency can be retrieved. Such above two factors can drive our EUIN-Net to promote the quality of both stego and revealed secret images. Furthermore, the shared and multi-scale characteristics of the U-shaped invertible blocks during the hiding and revealing stages contribute to significant reductions of our EUIN-Net in the model size and Flops. Extensive experiments demonstrate that the proposed EUIN-Net is efficient and can achieve state-of-the-art performances for image steganography.
Le Zhang 0016, Yao Lu 0008, Mi-Xiao Hou, Guangming Lu 0002
ICME4
2023 Multimodal Emotion Interaction and Visualization Platform
abstract
In this paper, we present a multimodal emotion analysis platform, which can flexibly capture, detect and analyze the emotions of video object with multiple modalities under different situations, including offline and online application scenarios. This system can visualize the dynamic effects of different types of emotions from both multimodal and unimodal circumstances. The presented emotion analysis results show instant and time series states in both specific modality and multiple modalities. Our system fills the current research and application gaps in multimodal emotion analysis with an interactive interface. Notably, the constructed system can adaptively process pre-recorded video clips as well as collected real-world data with excellent practicality and interactivity.
Zheng Zhang 0006, Songling Chen, Mi-Xiao Hou, Guangming Lu 0002
ACM Multimedia3
2023 Semantic Alignment Network for Multi-Modal Emotion Recognition
abstract
Modality alignment can maintain the consistency of semantics in multi-modal emotion recognition tasks, ensuring that features from different modalities accurately represent the emotion-related information in an encoding space. However, current alignment models either focus only on the local fusion of different modal representations or lack a mining process for unimodal specificity information. We design a Semantic Alignment network based on Multi-Spatial learning (SAMS) for multi-modal emotion recognition, which achieves local and global alignment between modalities using high-level emotion representations of different modalities as supervisory signals. SAMS builds a multi-spatial learning framework for each modality, and constructs a self-modal interaction module under this framework based on cross-modal semantic learning. SAMS provides two learning spaces for each modality, one to detect the affective information for a specific modality, and the other to learn semantic knowledge from other modalities. Subsequently, the features of these two spaces are aligned in temporal and utterance levels by homologous encoding and different target constraints. Based on the alignment characteristics of these two spaces, a self-modal interaction is built to investigate the fusion representation by exploring the global correlation between the alignment features in unimodal multi-spatial learning. In experiments, our proposed model yields consistent improvements on two standard multi-modal benchmarks, and outperforms state-of-the-art approaches. The code of our SAMS is available at:https://github.com/xiaomi1024/code_SAMS.
Mi-Xiao Hou, Zheng Zhang 0006, Guangming Lu 0002
IEEE Trans. Circuits Syst. Video Technol.1
2022 Multi-Modal Emotion Recognition with Self-Guided Modality Calibration
abstract
Multi-modal emotion recognition aims to extract sentiment-related information from multiple sources and integrate different modal representations for sentiment analysis. Alignment is an effective strategy to achieve semantically consistent representations for multi-modal emotion recognition, while the current alignment models are jointly unable to maintain the dependence of word-to-sentence and independence of unimodal learning. In this paper, we propose a Self-guided Modality Calibration Network (SMCN) to realize multi-modal alignment which can capture the global connections without interfering with unimodal learning. While preserving unimodal learning without interference, our model leverages semantic sentiment-related features to guide modality-specific representation learning. On one hand, SMCN simulates human thinking by deriving a branch for acquiring knowledge of other modalities in unimodal learning. This branch aims to lean high-level semantic information of other modalities for realizing semantic alignment between modalities. On the other hand, we also provide an indirect interaction manner to integrate unimodal feature and calibrate features in different levels for avoiding unimodal features mixed with other clues. Experiments demonstrate that our approach outperforms the state-of-the-art methods on both IEMOCAP and MELD datasets.
Mi-Xiao Hou, Zheng Zhang 0006, Guangming Lu 0002
ICASSP1
2022 Improved Deep Unsupervised Hashing with Fine-grained Semantic Similarity Mining for Multi-Label Image Retrieval
abstract
In this paper, we study deep unsupervised hashing, a critical problem for approximate nearest neighbor research. Most recent methods solve this problem by semantic similarity reconstruction for guiding hashing network learning or contrastive learning of hash codes. However, in multi-label scenarios, these methods usually either generate an inaccurate similarity matrix without reflection of similarity ranking or suffer from the violation of the underlying assumption in contrastive learning, resulting in limited retrieval performance. To tackle this issue, we propose a novel method termed HAMAN, which explores semantics from a fine-grained view to enhance the ability of multi-label image retrieval. In particular, we reconstruct the pairwise similarity structure by matching fine-grained patch features generated by the pre-trained neural network, serving as reliable guidance for similarity preserving of hash codes. Moreover, a novel conditional contrastive learning on hash codes is proposed to adopt self-supervised learning in multi-label scenarios. According to extensive experiments on three multi-label datasets, the proposed method outperforms a broad range of state-of-the-art methods.
Zeyu Ma 0001, Xiao Luo 0001, Yingjie Chen 0002, Mi-Xiao Hou, Jinxing Li 0003, Minghua Deng, Guangming Lu 0002
IJCAI4
2022 Recursive Feature Diversity Network for audio super-resolution
Bo Jiang 0017, Mi-Xiao Hou, Yao Lu 0008, David Zhang 0001, Guangming Lu 0002
Speech Commun.2
2022 Multi-View Speech Emotion Recognition Via Collective Relation Construction
abstract
Automatic emotion recognition from speech plays a fundamental role towards advanced emotional intelligence in human-machine interaction systems. The discriminative knowledge from speech for effective emotion recognition may come from multiple physical properties such as energy spectrum, frequency, prosody, which could be collected as multi-view representations. However, the current works fail to fully explore the underlying interactive relations among multiple speech representations for emotion recognition. In this paper, we propose a novel Collective Multi-view Relation Network (CMRN) to exploit the intrinsic characteristics of multi-view speech representations for discriminative speech emotion recognition. Generally, the proposed CMRN consists of three sub-networks,i.e.,view-specific attention network, multi-view shared attention network and collective relation network. Specifically, the view-specific attention network is designed to excavate the distinguishable view-specific features deduced from the original speech. By contrast, the multi-view shared attention network is conceived to capture the collaborative knowledge from multiple views. Moreover, a well-designed collective relation network is explicitly constructed to characterize the shared-specific correlations, which could reflect the underlying physical interaction capabilities. As such, the decision phase can comprehensively leverage the shared and view-specific information of multiple representations, such that the final privileged deciding principle can aggregate the heterogeneous information of multi-view features to make accurate emotion recognition. Extensive experiments on two benchmark datasets demonstrate the superb performance of the proposed method in comparison with some state-of-the-art methods.
Mi-Xiao Hou, Zheng Zhang 0006, David Zhang 0001, Guangming Lu 0002
IEEE ACM Trans. Audio Speech Lang. Process.1
2021 Hierarchical Network Based on the Fusion of Static and Dynamic Features for Speech Emotion Recognition
abstract
Many studies on automatic speech emotion recognition (SER) have been devoted to extracting meaningful emotional features for generating emotion-relevant representations. However, they generally ignore the complementary learning of static and dynamic features, leading to limited performances. In this paper, we propose a novel hierarchical network called HNSD that can efficiently integrate the static and dynamic features for SER. Specifically, the proposed HNSD framework consists of three different modules. To capture the discriminative features, an effective encoding module is firstly designed to simultaneously encode both static and dynamic features. By taking the obtained features as inputs, the Gated Multi-features Unit (GMU) is conducted to explicitly determine the emotional intermediate representations for frame-level features fusion, instead of directly fusing these acoustic features. In this way, the learned static and dynamic features can jointly and comprehensively generate the unified feature representations. Benefiting from a well-designed attention mechanism, the last classification module is applied to predict the emotional states at the utterance level. Extensive experiments on the IEMOCAP benchmark dataset demonstrate the superiority of our method in comparison with state-of-the-art baselines.
Mi-Xiao Hou, Bingzhi Chen, Zheng Zhang 0006, Guangming Lu 0002
ICASSP2
2021 Multimodal Emotion Recognition With Temporal and Semantic Consistency
abstract
Automated multimodal emotion recognition has become an emerging but challenging research topic in the fields of affective learning and sentiment analysis. The existing works mainly focus on developing multimodal fusion strategies to incorporate different emotion-related features. However, they fail to explore the inherent contextual consistency to reconcile the emotional information across modalities. In this paper, we propose a novel Time and Semantic Interaction Network (TSIN), which concurrently incorporates the advantages of temporal and semantic consistency into the multimodal emotion recognition task. Specifically, a well-designed Speech and Text Embedding (STE) module is devoted to formulating the initial embedding spaces by respectively building the modality-specific representations of speech and text. Instead of separately learning or directly fusing the acoustic and textual features, we propose a well-defined Time and Semantic Interaction (TSI) module to conduct the emotional parsing and sentiment refining by performing the fine-grained temporal alignment and cross-modal semantic interaction. Benefitting from temporal and semantic consistency constraints, both speech-text embeddings can be interactively optimized and fine-tuned in the learning process. In this way, the learnt acoustics and textual features can jointly and efficiently predict the final emotional state. Extensive experiments on the IEMOCAP dataset demonstrate the superiorities of our TSIN framework in comparison with state-of-the-art baselines.
Bingzhi Chen, Mi-Xiao Hou, Zheng Zhang 0006, Guangming Lu 0002, David Zhang 0001
IEEE ACM Trans. Audio Speech Lang. Process.3
2020 A supervised non-negative matrix factorization model for speech emotion recognition
Mi-Xiao Hou, Jinxing Li 0003, Guangming Lu 0002
Speech Commun.1
2019 PCA via joint graph Laplacian and sparse constraint: Identification of differentially expressed genes and sample clustering on gene expression data
abstract
BACKGROUND: In recent years, identification of differentially expressed genes and sample clustering have become hot topics in bioinformatics. Principal Component Analysis (PCA) is a widely used method in gene expression data. However, it has two limitations: first, the geometric structure hidden in data, e.g., pair-wise distance between data points, have not been explored. This information can facilitate sample clustering; second, the Principal Components (PCs) determined by PCA are dense, leading to hard interpretation. However, only a few of genes are related to the cancer. It is of great significance for the early diagnosis and treatment of cancer to identify a handful of the differentially expressed genes and find new cancer biomarkers. RESULTS: In this study, a new method gLSPCA is proposed to integrate both graph Laplacian and sparse constraint into PCA. gLSPCA on the one hand improves the clustering accuracy by exploring the internal geometric structure of the data, on the other hand identifies differentially expressed genes by imposing a sparsity constraint on the PCs. CONCLUSIONS: Experiments of gLSPCA and its comparison with existing methods, including Z-SPCA, GPower, PathSPCA, SPCArt, gLPCA, are performed on real datasets of both pancreatic cancer (PAAD) and head & neck squamous carcinoma (HNSC). The results demonstrate that gLSPCA is effective in identifying differentially expressed genes and sample clustering. In addition, the applications of gLSPCA on these datasets provide several new clues for the exploration of causative factors of PAAD and HNSC.
Chun-Mei Feng 0001, Yong Xu 0001, Mi-Xiao Hou, Ling-Yun Dai, Junliang Shang
BMC Bioinform.3
2018 Sparse Orthogonal Nonnegative Matrix Factorization for Identifying Differentially Expressed Genes and Clustering Tumor Samples
Ling-Yun Dai, Jin-Xing Liu 0001, Mi-Xiao Hou, Shasha Yuan
BIBM5
2018 Performance Analysis of Non-negative Matrix Factorization Methods on TCGA Data
Mi-Xiao Hou, Jin-Xing Liu 0001, Junliang Shang, Ying-Lian Gao, Ling-Yun Dai
ICIC (2)1
2018 Identifying Characteristic Genes and Clustering via an Lp-Norm Robust Feature Selection Method for Integrated Data
Shasha Wu, Mi-Xiao Hou, Jin-Xing Liu 0001, Juan Wang 0003, Shasha Yuan
ICIC (2)2
2016 Robust graph regularized discriminative nonnegative matrix factorization for characteristic gene selection
abstract
Recent research shows that characteristic gene selection based on gene expression data remains faced with considerable challenges. This is primarily because vast amount of gene expression data have been generated with the development of gene detection technology. Nonetheless, the recognition rate and reliability of gene selection still need to be improved. In this paper, we propose a novel constrained method: robust graph regularized discriminative nonnegative matrix factorization (RGDNMF) for characteristic gene selection. The method mainly includes two aspects: firstly, we incorporate both intrinsic geometrical structure and discriminative label information into the NMF model. Secondly, we adopt L2,1 -norm minimization to both the error function and the regularization term which is robust to noises and outliers in gene data. Furthermore we present the multiplicative update rules and the convergence proof. Our experiments demonstrate that RGDNMF is far more effective than other existing methods.
Ling-Yun Dai, Chun-Mei Feng 0001, Jin-Xing Liu 0001, Chun-Hou Zheng 0001, Mi-Xiao Hou, Jiguo Yu
BIBM5
2016 A p-norm singular value decomposition method for robust tumor clustering
abstract
Tumor clustering based on biomolecular data plays a very important role for cancer classifications discovery. To further improve the robustness, stability and accuracy of tumor clustering, we develop a novel dimension reduction method named p-norm singular value decomposition (PSVD) to seek a low-rank approximation matrix to the bimolecular data. To enhance the robustness to outliers, the Lp-norm is taken as the error function and the Schatten p-norm is used as the regularization function in our optimization model. To evaluate the performance of PSVD, Kmeans clustering method is then employed for tumor clustering based on the low-rank approximation matrix. The extensive experiments are performed on gene expression dataset and cancer genome dataset respectively. All experimental results demonstrate that the PSVD-based method outperforms many existing methods. Especially it is experimentally proved that the proposed method is efficient for processing higher dimensional data with good robustness and superior time performance.
Jin-Xing Liu 0001, Chun-Hou Zheng 0001, Mi-Xiao Hou, Yao Lu 0008
BIBM4
2016 Comparison of Non-negative Matrix Factorization Methods for Clustering Genomic Data
Mi-Xiao Hou, Ying-Lian Gao, Jin-Xing Liu 0001, Junliang Shang, Chun-Hou Zheng 0001
ICIC (2)1