Hao Sun 0013

dblp:82/2248-13 · DBLP profile ↗
← Back
14ranked-venue papers
6as first author
14since 2021 · last 2026
0000-0001-8094-1991ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 8 · 2 first-author · 8 since 2021Artificial intelligence and machine learning · 4 · 4 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Computer networks · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 MIRTH: Mutual-Information Reasoning with Temporal Hubs for Vision-Language-Action Agents
abstract
VLA models have emerged as a powerful paradigm for transferring semantic knowledge from web-scale data to physical robotic control.However, current single-frame architectures suffer from intrinsic limitations: temporal myopia that discards historical dynamics, reasoning gaps between high-level instructions and low-level motor commands, and inference inefficiency due to autoregressive scalar decoding.In this work, we propose MIRTH, a unified framework designed to address these challenges.MIRTH augments a pretrained VLA backbone with three key innovations: (1) dualscale temporal memory hubs that compress long-term scene evolution and short-term motion trends into compact embeddings; (2) latent reasoning tokens optimized via a mutualinformation objective carving out a semantic plan space to align multimodal context with action trajectories; and (3) a parallel action decoding scheme that replaces autoregressive generation with vector-wise prediction to maximize control throughput.Extensive evaluations on the LIBERO simulation benchmark and a real-world LeRobot platform demonstrate that MIRTH achieves state-of-the-art performance and exhibiting emergent error recovery capabilities.The codes and collected datasets are released at http://github.com/kiva12138/mirth.
Hao Sun 0013, Yu Song 0008, Shiyu Teng, Ziwei Niu, Yen-Wei Chen 0001
ACL (1)1
2026 Enhancing Depression Detection Using Pretrained Multi-modal Sentiment Analysis Models with Deep Prefix Tuning
abstract
Depression, a pervasive mental health condition, affects millions globally, challenging early and accurate diagnosis due to its subtle and varied manifestations. Recognizing the critical link between emotional dysregulation and depressive symptoms, our research introduces a pioneering training paradigm that integrates sentiment analysis with depression detection. This approach is motivated by the potential of sentiment data to enrich models with a deeper understanding of emotional states, crucial for identifying depressive patterns. To leverage the nuanced sentiment information without compromising the pretrained model’s integrity, we employ deep prefix tuning. This novel technique allows for targeted model refinement, ensuring that the valuable pretrained structures are not overshadowed by the sparse and specific nature of depression-related data. The empirical results demonstrate superior performance across standard benchmarks, setting a new precedent for multimodal depression detection.
Shiyu Teng, Jiaqing Liu, Shurong Chai, Hao Sun 0013, Tomoko Tateyama, Lanfen Lin, Yen-Wei Chen 0001
ACM Trans. Comput. Heal.4
2026 One framework to rule them all: Unifying multimodal tasks with LLM neural-tuning
Hao Sun 0013, Yu Song 0008, Jiaqing Liu, Jihong Hu, Yen-Wei Chen 0001, Lanfen Lin
Pattern Recognit.1
2026 S2Match: Revisiting Weak-to-Strong Consistency From a Semantic Similarity Perspective for Semi-Supervised Medical Image Segmentation
abstract
Semi-supervised learning (SSL) for medical image segmentation is a challenging yet highly practical task, which reduces reliance on large-scale labeled datasets by leveraging unlabeled samples. Among SSL techniques, the weak-to-strong consistency framework, popularized by FixMatch, has emerged as a state-of-the-art method in classification tasks. Notably, such a simple pipeline has also shown competitive performance in medical image segmentation. However, two key limitations still persist, impeding its efficient adaptation: (1) the neglect of contextual dependencies results in inconsistent predictions for similar semantic features, leading to incomplete object segmentation; (2) the lack of exploitation on semantic similarity between labeled and unlabeled data induces considerable class-distribution discrepancy. To address these limitations, we propose a novel SSL framework for medical image segmentation, named S2Match, powered by two appealing designs from a semantic similarity perspective: (1) rectifying pixel-wise prediction by reasoning about the intra-image pair-wise affinity map, thus integrating contextual dependencies explicitly into the final prediction; (2) bridging labeled and unlabeled data via a feature querying mechanism for compact class representation learning, which fully considers cross-image anatomical similarities. As the reliable semantic similarity extraction depends on robust features, we further introduce an effective Spatial-aware Fusion Module (SFM) to explore distinctive information from multiple scales. Experiments show that S2Match yields consistent improvements over the state-of-the-art methods across five public medical image segmentation benchmarks, exhibiting competitive performance on both 2D and 3D tasks.
Shiao Xie, Hongyi Wang 0002, Ziwei Niu, Hao Sun 0013, Shuyi Ouyang, Yen-Wei Chen 0001, Lanfen Lin
IEEE J. Biomed. Health Informatics4
2026 Disentangled Multimodal Tuning and Interaction for Human Perception Understanding
abstract
Understanding human perceptions poses a significant multimodal challenge for computers, involving textual, acoustic, and visual signals. Recently, large language models (LLMs) have garnered great attention, leading to numerous methods aimed at efficiently fine-tuning pretrained models for multimodal downstream tasks. However, there remains a scarcity of techniques that prioritize modality-invariant and -specific information during parameter-efficient tuning, despite evidence from previous studies showcasing the effectiveness of modality disentangling. To address this gap, we propose a novel multimodal tuning approach for LLMs, termed Disentangled Multimodal Tuning and Interaction. Specifically, we evaluate the independence among different modalities and disentangle corresponding modality-invariant and specific components, which are subsequently leveraged for prompt tuning. Following tuning, a newly designed independence-guided cross-attention module is introduced for modality interaction, where the attention mechanism is decoupled and bolstered with independence from the modality-disentangling process. This approach not only enables LLMs to efficiently assimilate information from various modalities but also cultivates an awareness of both modality-invariant and specific information. Compared to previous methods, our approach facilitates modality interaction at a more granular level, resulting in enhanced performance. We validate our method through experiments on four public datasets, demonstrating significant performance improvements.
Hao Sun 0013, Ziwei Niu, Jiaqing Liu, Yen-Wei Chen 0001, Lanfen Lin
ACM Trans. Multim. Comput. Commun. Appl.1
2025 Enhanced Multimodal Depression Detection With Emotion Prompts
abstract
Depression is a pervasive mental health disorder that remains frequently undiagnosed and untreated due to societal barriers and the subjective nature of its symptoms. Leveraging recent advances in large language models (LLMs), we propose a novel depression detection pipeline that generates emotion prompts tailored to individual data, enhancing detection accuracy. Our approach integrates cross-modality fusion via cross attention mechanisms to combine depressive and emotional features, creating a comprehensive representation of depression indicators. Evaluated on the E-DAIC and EATD datasets, our method outperforms state-of-the-art techniques, demonstrating its potential for more precise emotion-based depression detection.
Shiyu Teng, Jiaqing Liu, Hao Sun 0013, Shurong Chai, Tomoko Tateyama, Lanfen Lin, Yen-Wei Chen 0001
ICASSP3
2025 EPIC: Efficient Prompt Interaction for Text-Image Classification
abstract
In recent years, large-scale pre-trained multimodal models (LMMs) generally emerge to integrate the vision and language modalities, achieving considerable success in multimodal tasks, such as text-image classification. The growing size of LMMs, however, results in a significant computational cost for fine-tuning these models for downstream tasks. Hence, prompt-based interaction strategy is studied to align modalities more efficiently. In this context, we propose a novel efficient prompt-based multimodal interaction strategy, namely Efficient Prompt Interaction for text-image Classification (EPIC). Specifically, we utilize temporal prompts on intermediate layers, and integrate different modalities with similarity-based prompt interaction, to leverage sufficient information exchange between modalities. Utilizing this approach, our method achieves reduced computational resource consumption and fewer trainable parameters (about 1% of the foundation model) compared to other fine-tuning strategies. Furthermore, it demonstrates superior performance on the UPMC-Food101 and SNLI-VE datasets, while achieving comparable performance on the MM-IMDB dataset.
Xinyao Yu 0003, Hao Sun 0013, Zeyu Ling, Ziwei Niu, Zhenjia Bai, Yen-Wei Chen 0001, Lanfen Lin
ICME2
2025 Multimodal Sentiment Analysis With Mutual Information-Based Disentangled Representation Learning
abstract
Multimodal sentiment analysis seeks to utilize various types of signals to identify underlying emotions and sentiments. A key challenge in this field lies in multimodal representation learning, which aims to develop effective methods for integrating multimodal features into cohesive representations. Recent advancements include two notable approaches: one focuses on decomposing multimodal features into modality-invariant and -specific components, while the other emphasizes the use of mutual information to enhance the fusion of modalities. Both strategies have demonstrated effectiveness and yielded remarkable results. In this paper, we propose a novel learning framework that combines the strengths of these two approaches, termed mutual information-based disentangled multimodal representation learning. Our approach involves estimating different types of information during feature extraction and fusion stages. Specifically, we quantitatively assess and adjust the proportions of modality-invariant, -specific, and -complementary information during feature extraction. Subsequently, during fusion, we evaluate the amount of information retained by each modality in the fused representation. We employ mutual information or conditional mutual information to estimate each type of information content. By reconciling the proportions of these different types of information, our approach achieves state-of-the-art performance on popular sentiment analysis benchmarks, including CMU-MOSI and CMU-MOSEI.
Hao Sun 0013, Ziwei Niu, Hongyi Wang 0002, Xinyao Yu 0003, Jiaqing Liu, Yen-Wei Chen 0001, Lanfen Lin
IEEE Trans. Affect. Comput.1
2024 IRLSG: Invariant Representation Learning for Single-Domain Generalization in Medical Image Segmentation
abstract
Single-domain generalization (SDG) can efficiently enhance model generalization while avoiding high annotation costs and privacy concerns. However, existing SDG methods are mainly based on data manipulation and meta-learning, which are not efficient enough due to the limited generalization performance and complex inference. In response to these challenges, we present a novel single domaininvariant representation learning approach for medical image segmentation, called IRLSG, with two appealing designs: (1) A Classscale Photo-metric Augmentation is first proposed to simulate unseen target domain that is sufficient in diversity and informativeness. After that, a Dual-Consistency Framework is further designed to constrain the consistency of intermediate features and segmentation results between the original and the augmented images, which helps to explore the domain-invariant representation. (2) A simple and effective Style Feature Whitening is designed to decouple and remove the domain-specific style from higher-order covariance statistics, which can further improve the modeling and generalization capability of the network. Experimental results on different benchmarks demonstrate that our IRLSG outperforms the current state-of-the-art methods in tackling single-domain generalization.
Ziwei Niu, Hao Sun 0013, Shuyi Ouyang, Shiao Xie, Yen-Wei Chen 0001, Ruofeng Tong 0001, Lanfen Lin
ICASSP2
2024 LGA: A Language Guide Adapter for Advancing the SAM Model's Capabilities in Medical Image Segmentation
Jihong Hu, Yinhao Li 0002, Hao Sun 0013, Yu Song 0008, Chujie Zhang, Lanfen Lin, Yen-Wei Chen 0001
MICCAI (12)3
2023 MCKD: Mutually Collaborative Knowledge Distillation For Federated Domain Adaptation And Generalization
abstract
Conventional unsupervised domain adaptation (UDA) and domain generalization (DG) methods rely on the assumption that all source domains can be directly accessed and combined for model training. However, this centralized training strategy may violate privacy policies in many real-world applications. A paradigm for tackling this problem is to train multiple local models and aggregate a generalized central model without data sharing. Recent methods have made remarkable advancements in this paradigm by exploiting parameter alignment and aggregation. But when sources domain variety increases, directly aligning and aggregating local parameters becomes more challenging. Adapting a different approach in this work, we devised a data-free semantic collaborative distillation strategy to learn domain-invariant representation for both federated UDA and DG. Each local model transmits its predictions to the central server and derives its target distribution from the average of other local models' distributions to facilitate the mutual transfer of domain-specific knowledge. When unlabeled target data is available, we introduce a novel UDA strategy termed knowledge filter to adapt the central model to the target data. Extensive experiments on four UDA and DG datasets demonstrate that our method has a competitive performance compared with the state-of-the-art methods.
Ziwei Niu, Hongyi Wang 0002, Hao Sun 0013, Shuyi Ouyang, Yen-Wei Chen 0001, Lanfen Lin
ICASSP3
2023 IS2Net: Intra-domain Semantic and Inter-domain Style Enhancement for Semi-supervised Medical Domain Generalization
abstract
Domain generalization (DG) demonstrates superior generalization ability in cross-center medical image segmentation. Despite its great success, existing fully supervised DG methods require collecting a large quantity of pixel-level annotations which is quite expensive and time-consuming. To address this challenge, several semi-supervised domain generalized (SSDG) methods have been proposed by simply coupling semi-supervised learning (SSL) with DG tasks, which give rise to two main concerns: (1) Intra-domain dubious semantic information: the quality of pseudo labels in each source domain suffers from the limited amount of labeled data and cross-domain discrepancy. (2) Inter-domain intangible style relationship: current models fail in integrating domain-level information and overlook the relationships among different domains, which degrades the generalization ability of model. In light of these two issues, we propose a novel SSDG framework, namely IS2Net, by arranging an inter-domain generalization branch and several intra-domain SSL branches in a parallel manner, powered by two appealing designs that build a positive interaction between them: (1) A style and semantic memory mechanism is designed to provide both high-quality class-wise representations for intra-domain semantic enhancement and stable domain-specific knowledge for inter-domain style relationship construction. (2) Confident pseudo labeling strategy aims at generating more reliable supervision for intra and inter domain branches, and thus facilitating the learning process of the whole framework. Extensive experiments show that IS2Net yields consistent improvements over the state-of- the-art methods in three public benchmarks.
Shiao Xie, Ziwei Niu, Huimin Huang 0002, Hao Sun 0013, Yen-Wei Chen 0001, Lanfen Lin
ACM Multimedia4
2023 TensorFormer: A Tensor-Based Multimodal Transformer for Multimodal Sentiment Analysis and Depression Detection
abstract
Sentiment analysis is an important research field aiming to extract and fuse sentimental information from human utterances. Due to the diversity of human sentiment, analyzing from multiple modalities is usually more accurate than from a single modality. To complement the information between related modalities, one effective approach is performing cross-modality interactions. Recently, Transformer-based frameworks have shown a strong ability to capture long-range dependencies, leading to the introduction of several Transformer-based approaches for multimodal processing. However, due to the built-in attention mechanism of the Transformers, only two modalities can be engaged at once. As a result, the complementary information flow in these Transformer-based techniques is partial and constrained. To mitigate this, we propose, TensorFormer, a tensor-based multimodal Transformer framework that takes into account all relevant modalities for interactions. More precisely, we first construct a tensor utilizing the features extracted from each modality, assuming one modality is the target while the remaining tensors serve as the sources. We can generate the corresponding interacted features by calculating source-target attention. This strategy interacts with all involved modalities and generates complementing global information. Experiments on multimodal sentiment analysis benchmark datasets demonstrated the effectiveness of TensorFormer. In addition, we also evaluate TensorFormer in another related area: depression detection and the results reveal significant improvements when compared to other state-of-the-art methods.
Hao Sun 0013, Yen-Wei Chen 0001, Lanfen Lin
IEEE Trans. Affect. Comput.1
2022 CubeMLP: An MLP-based Model for Multimodal Sentiment Analysis and Depression Estimation
abstract
Multimodal sentiment analysis and depression estimation are two important research topics that aim to predict human mental states using multimodal data. Previous research has focused on developing effective fusion strategies for exchanging and integrating mind-related information from different modalities. Some MLP-based techniques have recently achieved considerable success in a variety of computer vision tasks. Inspired by this, we explore multimodal approaches with a feature-mixing perspective in this study. To this end, we introduce CubeMLP, a multimodal feature processing framework based entirely on MLP. CubeMLP consists of three independent MLP units, each of which has two affine transformations. CubeMLP accepts all relevant modality features as input and mixes them across three axes. After extracting the characteristics using CubeMLP, the mixed multimodal features are flattened for task predictions. Our experiments are conducted on sentiment analysis datasets: CMU-MOSI and CMU-MOSEI, and depression estimation dataset: AVEC2019. The results show that CubeMLP can achieve state-of-the-art performance with a much lower computing cost.
Hao Sun 0013, Hongyi Wang 0002, Jiaqing Liu, Yen-Wei Chen 0001, Lanfen Lin
ACM Multimedia1