Haoyu Shi

dblp:245/5871 · DBLP profile ↗
← Back
7ranked-venue papers
2as first author
7since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 4 · 2 first-author · 4 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Systems, architecture and hardware · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Aspect-Aware Affective Focus Network For Joint Multimodal Aspect Sentiment Analysis
Xiangbo Ji, Haoyu Shi
DASFAA (2)2
2025 Dual-Branch Sentiment Enhancement Modeling For Joint Multimodal Aspect-Based Sentiment Analysis
abstract
Joint Multimodal Aspect-based Sentiment Analysis (JMASA) is a challenging task that seeks to identify all aspect-sentiment pairs from multimodal data.Existing methods typically utilize attention mechanisms to search for the associated dependencies between words. However, sentiment judgments can be easily influenced by sentiment-irrelevant words in a sentence.These words also potentially impact the model's sensitivity to specific aspects.Additionally, the redundant sentiment-irrelevant patches in an image hinder effective cross-modal fusion and increase computational complexity.To this end, we propose a novel Dual-Branch Sentiment Enhancement Modeling (DSEM) framework for JMASA.Specifically, at the textual level, we leverage syntactic and part-of-speech structures to enhance the model's sensitivity to aspect-relevant semantic information.At the visual level, we identify sentiment-related patches and adaptively aggregate them to obtain optimal visual semantic representations.Finally, we introduce a language-guided semantic interaction to further improve the relevance between image-text pairs and achieve efficient cross-modal fusion.Extensive experiments on two datasets validate the effectiveness of our proposed method.
Xiangbo Ji, Haoyu Shi
ICMR2
2025 Sequence-Event Semantic Consistent Learning for Text-to-Motion Retrieval
abstract
Text-to-Motion Retrieval (TMR) is a challenging task to retrieve relevant motion sequences with the natural language description. Existing TMR methods primarily utilize single embeddings to represent and align text and motion sequences. However, real-world motion sequences typically contain multiple sequential actions with intricate semantics, which are hard to precisely capture by single embedding. Additionally, relying solely on naive contrastive training to capture high-level semantics may struggle to perceive and capture fine-grained action details necessary for precise text-motion alignment. In this work, we propose a novel Sequence-Event Semantic Consistent Learning (SECL) framework for 3D human motion retrieval. Specifically, we introduce a self-supervised learning strategy to incorporate fine-grained action details into the motion representations via the generative feedback from the diffusion model. We design a parameter-free sequence-level interaction to explore coarse-grained alignment and an event-level interaction that utilizes several learnable queries to capture event semantics in a shared learning manner for fine-grained alignment. Furthermore, an inter-consistency loss is introduced to align the event semantics between the motion and corresponding text, and an intra-diversity loss is designed to encourage event features to attend to different contents, effectively capturing the rich action information. Finally, we modify the traditional contrastive alignment objective and propose an importance-sampling strategy to emphasize harder negatives for discriminative representation learning. Extensive experiments show that our method significantly outperforms existing methods in text-to-motion retrieval and other challenging tasks, e.g., human interaction recognition and motion temporal localization.
Haoyu Shi, Huaiwen Zhang
ACM Multimedia1
2025 Robust graph neural networks based on feature fusion
Haoyu Shi, Huaiye Meng
J. Supercomput.2
2024 Multi-Instance Multi-Label Learning for Text-motion Retrieval
abstract
Text-motion retrieval (TMR) is a significant cross-modal task that retrieves motion sequences semantically similar to a given query text. Existing TMR methods primarily utilize single embeddings to represent and align text and motion sequences. However, real-world motion sequences typically contain multiple atomic motions with complex semantics, which is hard to precisely capture by single embeddings. Additionally, the common co-occurring and coupling of atomic motions further post significant challenges in effective modeling and aligning text and motion sequences. In this paper, we regard TMR as a Multi-Instance Multi-Label (MIML) learning problem, where the motion sequence is viewed as a bag of atomic motions and the text is the bag of corresponding phrases. To address the MIML problem, we propose a novel Multi-Granularity Semantics Interaction (MGSI) approach, which effectively captures and aligns the semantics of text and motion sequences across various levels. Specifically, the MGSI approach initially decomposes both the query and motion sequences into three hierarchical levels: token, instance, and bag. Then, we utilize graph neural networks to explicitly model their semantics correlation and perform semantics interaction at these respective levels, precisely capturing the semantics at multiple granularities. To identify and model co-occurring atomic motions, we measure the frame-wise semantic consistency between motions and then fuse and interact the accordant ones to refine their representations. Finally, we exploit token, instance, and bag-wise semantics interaction to comprehensively align text and motion sequence. We evaluated our methods on two widely-used benchmark datasets, HumanML3D and KIT-ML. The proposed method achieves significant improvements, outperforming the state-of-the-art with a 23.09% increase in Rsum on HumanML3D and a 21.84% increase on KIT-ML.
Yang Yang 0121, Haoyu Shi, Huaiwen Zhang
ACM Multimedia3
2024 Modal-Enhanced Semantic Modeling for Fine-Grained 3D Human Motion Retrieval
abstract
Text to Motion Retrieval (TMR) is an emerging task to retrieve relevant motion sequences with the nature language description. The dominant approach learns a joint embedding space to measure global-level similarities. However, simple global embeddings are insufficient to represent complicated motion and textual details, such as the movement of specific body parts and the coordination among these body parts. In addition, most of the motion variations occur subtly and locally, resulting in semantic vagueness among these motions, which further presents considerable challenges in precisely aligning motion sequences with texts. To address these challenges, we propose a novel Modal-Enhanced Semantic Modeling (MESM) method, focusing on fine-grained alignment through enhanced modal semantics. Specifically, we develop a prompt-enhanced textual module (PTM) to generate detailed descriptions of specific body part movements, which comprehensively captures the fine-grained textual semantics for precise matching. We employ a skeleton-enhanced motion module (SMM) to effectively enhance the model's capability to represent intricate motions. This module leverages a graph convolutional network to meticulously model the intricate spatial dependencies among relevant body parts. To improve the sensitivity to the subtle motions, we further propose a text-driven semantics interaction module (TSIM). The TSIM assigns motion features into a set of aggregated descriptors and employs cross-attention to aggregate discriminative motion embeddings guided by text, enabling precise semantic alignment between subtle motions and corresponding texts. Extensive experiments conducted on two widely used benchmark datasets, HumanML3D and KIT-ML, demonstrate the effectiveness of our proposed method. Our approach outperforms existing state-of-the-art retrieval methods, achieving significant Rsum improvements of 24.28% on HumanML3D and 25.80% on KIT-ML.
Haoyu Shi, Huaiwen Zhang
ACM Multimedia1
2024 Hierarchical Semantics Alignment for 3D Human Motion Retrieval
abstract
Text to 3D human Motion Retrieval (TMR) is a challenging task in information retrieval, aiming to query relevant motion sequences with the natural language description. The conventional approach for TMR is to represent the data instances as point embeddings for alignment. However, in real-world scenarios, multiple motions often co-occur and superimpose on a single avatar. Simply aggregating text and motion sequences into a single global embedding may be inadequate for capturing the intricate semantics of superimposing motions. In addition, most of the motion variations occur locally and subtly, which further presents considerable challenges in precisely aligning motion sequences with their corresponding text. To address the aforementioned challenges, we propose a novel Hierarchical Semantics Alignment (HSA) framework for text-to-3D human motion retrieval. Beyond global alignment, we propose the Probabilistic-based Distribution Alignment (PDA) and a Descriptors-based Fine-grained Alignment (DFA) to achieve precise semantic matching. Specifically, the PDA encodes the text and motion sequences into multidimensional probabilistic distributions, effectively capturing the semantics of superimposing motions. By optimizing the problem of probabilistic distribution alignment, PDA achieves a precise match between superimposing motions and their corresponding text. The DFA first adopts a fine-grained feature gating by selectively filtering to the significant and representative local representations and meanwhile excluding the interferences of meaningless features. Then we adaptively assign local representations from text and motion into a set of cross-modal local aggregated descriptors, enabling local comparison and interaction between fine-grained text and motion features. Extensive experiments on two widely used benchmark datasets, HumanML3D and KIT-ML, demonstrate the effectiveness of the proposed method. It significantly outperforms existing state-of-the-art retrieval methods, achieving Rsum improvements of 24.74% on HumanML3D and 23.08% on KIT-ML.
Yang Yang 0121, Haoyu Shi, Huaiwen Zhang
SIGIR2