Minhao Liu

dblp:79/10137 · DBLP profile ↗
← Back
14ranked-venue papers
2as first author
11since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 12 · 2 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 5 since 2021
YearPublicationVenuePosition
2026 TMDC: A Two-Stage Modality Denoising and Complementation Framework for Multimodal Sentiment Analysis with Missing and Noisy Modalities
abstract
Multimodal Sentiment Analysis (MSA) aims to infer human sentiment by integrating information from multiple modalities such as text, audio, and video. In real-world scenarios, however, the presence of missing modalities and noisy signals significantly hinders the robustness and accuracy of existing models. While prior works have made progress on these issues, they are typically addressed in isolation, limiting overall effectiveness in practical settings. To jointly mitigate the challenges posed by missing and noisy modalities, we propose a framework called Two-stage Modality Denoising and Complementation (TMDC). TMDC comprises two sequential training stages. In the Intra-Modality Denoising Stage, denoised modality-specific and modality-shared representations are extracted from complete data using dedicated denoising modules, reducing the impact of noise and enhancing representational robustness. In the Inter-Modality Complementation Stage, these representations are leveraged to compensate for missing modalities, thereby enriching the available information and further improving robustness. Extensive evaluations on MOSI, MOSEI, and IEMOCAP demonstrate that TMDC consistently achieves superior performance compared to existing methods, establishing new state-of-the-art results.
Yan Zhuang 0002, Minhao Liu, Yanru Zhang, Jiawen Deng 0006, Fuji Ren
AAAI2
2026 Anti-saturation prescribed performance control of steer-by-wire systems with zeroing neural dynamics
Siyuan Jia, Minhao Liu, Zhineng Long, Hui Pang, Huiyuan Xiong
Neurocomputing2
2025 CMAD: Correlation-Aware and Modalities-Aware Distillation for Multimodal Sentiment Analysis with Missing Modalities
Yan Zhuang 0002, Minhao Liu, Yanru Zhang, Jiawen Deng 0006, Fuji Ren
ICCV2
2025 EndoDUM: Unsupervised Endoscopic Depth Estimation with Uncertainty Mask
abstract
Unsupervised depth estimation plays a key role in endoscopic minimally invasive surgery. However, endoscopic images often suffer from abnormal lighting, with some areas being overexposed and others underexposed. This leads to poor performance of current depth estimation methods in these regions. In this paper, we propose Unsupervised Endoscopic Depth Estimation with Uncertainty Mask (EndoDUM), the first unsupervised method for depth estimation that addresses abnormal lighting regions in endoscopic images. EndoDUM addresses the abnormal lighting problem in endoscopic images by learning illumination and applying a soft mask to regions with abnormal lighting. Specifically, it estimates depth information from a single RGB image accurately through three novel components: 1) The illumination calibration network, which calibrates the raw endoscopic image by estimating lighting, 2) The uncertainty mask, which suppresses overexposed or underexposed regions based on the illumination calibration ratio, and 3) A Three-Dimensional dynamic convolution module that captures local fine-grained features by leveraging complementary attention across three dimensions. Experimental results on the SCARED and Hamlyn datasets demonstrate that EndoDUM significantly outperforms existing methods in depth estimation tasks, achieving state-of-the-art (SOTA) performance.
Xuanxuan Liu, Minhao Liu, Lixin Duan
IJCNN5
2025 FAME: Fusion-Aware Multi-modal Ensemble for Social Media Popularity Prediction
abstract
As social media becomes a dominant platform for sharing content, predicting the popularity of user posts has become increasingly important for applications such as content recommendation, trend forecasting, and user engagement. However, this task is challenging due to the diverse and multimodal nature of social media posts, which often include unstructured text, images, and structured metadata. To address this challenge, we propose Fusion-Aware Multi-modal Ensemble (FAME), a framework effectively captures and integrates diverse information sources within social media content. Unlike prior approaches that rely on a single model to process all modalities, FAME leverages four specialized predictors. Three of them-CatBoost, LightGBM, and AutoGluon-are tree-based models that excel at handling structured metadata and its interactions with unstructured features. The fourth is a denoising autoencoder (DAE), which learns robust joint representations from unstructured text and image data. These models are combined through a weighted ensemble strategy, allowing FAME to leverage the complementary strengths of different architectures. Experiments on the Social Media Prediction Dataset demonstrate that FAME significantly outperforms existing baselines, achieving state-of-the-art results and validating its effectiveness in modeling the complex, multimodal nature of social media content.
Yan Zhuang 0002, Yanru Zhang, Minhao Liu, Jiawen Deng 0006, Fuji Ren
ACM Multimedia4
2025 Hyper-Modality Enhancement for Multimodal Sentiment Analysis with Missing Modalities
abstract
Multimodal Sentiment Analysis (MSA) aims to infer human emotions by integrating complementary signals from diverse modalities. However, in real-world scenarios, missing modalities are common due to data corruption, sensor failure, or privacy concerns, which can significantly degrade model performance. To tackle this challenge, we propose Hyper-Modality Enhancement (HME), a novel framework that avoids explicit modality reconstruction by enriching each observed modality with semantically relevant cues retrieved from other samples. This cross-sample enhancement reduces reliance on fully observed data during training, making the method better suited to scenarios with inherently incomplete inputs. In addition, we introduce an uncertainty-aware fusion mechanism that adaptively balances original and enriched representations to improve robustness. Extensive experiments on three public benchmarks show that HME consistently outperforms state-of-the-art methods under various missing modality conditions, demonstrating its practicality in real-world MSA applications.
Yan Zhuang 0002, Minhao Liu, Yanru Zhang, Wei Li 0308, Jiawen Deng 0006, Fuji Ren
NeurIPS2
2022 T-WaveNet: A Tree-Structured Wavelet Neural Network for Time Series Signal Analysis
Minhao Liu, Ailing Zeng, Qiuxia Lai, Ruiyuan Gao 0001, Min Li 0019, Harry Qin, Qiang Xu 0001
ICLR1
2022 SCINet: Time Series Modeling and Forecasting with Sample Convolution and Interaction
abstract
One unique property of time series is that the temporal relations are largely preserved after downsampling into two sub-sequences. By taking advantage of this property, we propose a novel neural network architecture that conducts sample convolution and interaction for temporal modeling and forecasting, named SCINet. Specifically, SCINet is a recursive downsample-convolve-interact architecture. In each layer, we use multiple convolutional filters to extract distinct yet valuable temporal features from the downsampled sub-sequences or features. By combining these rich features aggregated from multiple resolutions, SCINet effectively models time series with complex temporal dynamics. Experimental results show that SCINet achieves significant forecasting accuracy improvements over both existing convolutional models and Transformer-based solutions across various real-world time series forecasting datasets. Our codes and data are available at https://github.com/cure-lab/SCINet.
Minhao Liu, Ailing Zeng, Muxi Chen, Qiuxia Lai, Lingna Ma, Qiang Xu 0001
NeurIPS1
2022 Adaptive sliding mode attitude control of two-wheel mobile robot with an integrated learning-based RBFNN approach
Hui Pang, Minhao Liu, Chuan Hu 0003, Fengqi Zhang
Neural Comput. Appl.2
2021 Learning Skeletal Graph Neural Networks for Hard 3D Pose Estimation
abstract
Various deep learning techniques have been proposed to solve the single-view 2D-to-3D pose estimation problem. While the average prediction accuracy has been improved significantly over the years, the performance on hard poses with depth ambiguity, self-occlusion, and complex or rare poses is still far from satisfactory. In this work, we target these hard poses and present a novel skeletal GNN learning solution. To be specific, we propose a hop-aware hierarchical channel-squeezing fusion layer to effectively extract relevant information from neighboring nodes while suppressing undesired noises in GNN learning. In addition, we propose a temporal-aware dynamic graph construction procedure that is robust and effective for 3D pose estimation. Experimental results on the Human3.6M dataset show that our solution achieves 10.3% average prediction accuracy improvement and greatly improves on hard poses over state-of-the-art techniques. We further apply the proposed technique on the skeleton-based action recognition task and also achieve state-of-the-art performance. Our code is available at https://github.com/ailingzengzzz/Skeletal-GNN.
Ailing Zeng, Xiao Sun 0001, Nanxuan Zhao, Minhao Liu, Qiang Xu 0001
ICCV5
2021 Information Bottleneck Approach to Spatial Attention Learning
abstract
The selective visual attention mechanism in the human visual system (HVS) restricts the amount of information to reach visual awareness for perceiving natural scenes, allowing near real-time information processing with limited computational capacity. This kind of selectivity acts as an ‘Information Bottleneck (IB)’, which seeks a trade-off between information compression and predictive accuracy. However, such information constraints are rarely explored in the attention mechanism for deep neural networks (DNNs). In this paper, we propose an IB-inspired spatial attention module for DNN structures built for visual recognition. The module takes as input an intermediate representation of the input image, and outputs a variational 2D attention map that minimizes the mutual information (MI) between the attention-modulated representation and the input, while maximizing the MI between the attention-modulated representation and the task label. To further restrict the information bypassed by the attention map, we quantize the continuous attention scores to a set of learnable anchor values during training. Extensive experiments show that the proposed IB-inspired spatial attention mechanism can yield attention maps that neatly highlight the regions of interest while suppressing backgrounds, and bootstrap standard DNN structures for visual recognition tasks (e.g., image classification, fine-grained recognition, cross-domain classification). The attention maps are interpretable for the decision making of the DNNs as verified in the experiments. Our code is available at this https URL.
Qiuxia Lai, Yu Li 0007, Ailing Zeng, Minhao Liu, Hanqiu Sun, Qiang Xu 0001
IJCAI4
2020 SRNet: Improving Generalization in 3D Human Pose Estimation with a Split-and-Recombine Approach
Ailing Zeng, Xiao Sun 0001, Fuyang Huang, Minhao Liu, Qiang Xu 0001, Stephen Lin 0001
ECCV (14)4
2020 DeepFuse: An IMU-Aware Network for Real-Time 3D Human Pose Estimation from Multi-View Image
abstract
In this paper, we propose a two-stage fully 3D network, namely DeepFuse, to estimate human pose in 3D space by fusing body-worn Inertial Measurement Unit (IMU) data and multi-view images deeply. The first stage is designed for pure vision estimation. To preserve data primitiveness of multi-view inputs, the vision stage uses multi-channel volume as data representation and 3D soft-argmax as activation layer. The second one is the IMU refinement stage which introduces an IMU-bone layer to fuse the IMU and vision data earlier at data level. without requiring a given skeleton model a priori, we can achieve a mean joint error of 28.9mm on TotalCapture dataset and 13.4mm on Human3.6M dataset under protocol 1, improving the SOTA result by a large margin. Finally, we discuss the effectiveness of a fully 3D network for 3D pose estimation experimentally which may benefit future research.
Fuyang Huang, Ailing Zeng, Minhao Liu, Qiuxia Lai, Qiang Xu 0001
WACV3
2018 Structure-Aware 3D Hourglass Network for Hand Pose Estimation from Single Depth Image
Fuyang Huang, Ailing Zeng, Minhao Liu, Harry Qin, Qiang Xu 0001
BMVC3