Shihao Gao

dblp:260/7027 · DBLP profile ↗
← Back
15ranked-venue papers
3as first author
15since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 9 · 1 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 2 first-author · 6 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Local augmentation and debiasing for graph contrastive learning
Shihao Gao, Zhongwei Xiong, Taisong Jin
Neurocomputing1
2025 K-hop Hypergraph Neural Network: A Comprehensive Aggregation Approach
abstract
The powerful capability of HyperGraph Neural Networks (HGNNs) in modeling intricate, high-order relationships among multiple data samples stems primarily from their ability to aggregate both the direct neighborhood features of individual nodes and those associated with hyperedges. However, the limited scope of feature propagation in existing HGNNs significantly reduces the utilization of hypergraph information, exacerbating over-squashing and over-smoothing issues. To this end, we propose a novel K-hop HyperGraph Neural Network (KHGNN) to facilitate the interactions of distant nodes and hyperedges. Specifically, the bisection nested convolution based on HyperGINE is employed to extract features from nodes, hyperedges, and structures along all shortest paths between nodes or hyperedges, providing representations of long-distance relationships. With these comprehensive path features, nodes and hyperedges are guided to aggregate distant information while learning their complex relationships. The extensive experiments, particularly on long-range graph datasets, demonstrate that the proposed method achieves SOTA performance compared to existing HGNNs and graph neural networks.
Linhuang Xie, Shihao Gao, Ming Yin 0002, Taisong Jin
AAAI2
2025 DASA-Trans-STM: Adaptive Efficient Transformer for Short Text Matching using Data Augmentation and Semantic Awareness
abstract
Rencent advancements in large language models (LLM) have shown impressive versatility across various tasks.Short text matching is one of the fundamental technologies in natural language processing.In previous studies, the common approach to applying them to Chinese is segmenting each sentence into words, and then taking these words as input.However, existing approaches have three limitations: 1) Some Chinese words are polysemous, and semantic information is not fully utilized.2) Some models suffer potential issues caused by word segmentation and incorrect recognition of negative words affects the semantic understanding of the whole sentence.3) Fuzzy negation words in ancient Chinese are difficult to recognize and match.In this work, we propose a novel adaptive Transformer for Chinese short text matching using Data Augmentation and Semantic Awareness (DASA), which can fully mine the information expressed in Chinese text to deal with word ambiguity.DASA is based on a Graph Attention Transformer Encoder that takes two word lattice graphs as input and integrates sense information from N-HowNet to moderate word ambiguity.Specially, we use an LLM to generate similar sentences for the optimal text representation.Experimental results show that the augmentation done using DASA can considerably boost the performance of our system and achieve significantly better results than previous state-of-theart methods on four available datasets, namely MNS, LCQMC, AFQMC, and BQ.
Jiguo Liu, Chao Liu 0020, Meimei Li, Shihao Gao, Dali Zhu
EMNLP5
2025 Graph Contrastive Learning with Decoupled Augmentation
abstract
Graph contrastive learning based on augmentation strategies has recently demonstrated remarkable performance. Existing methods typically jointly leverage attribute and structural augmentations to generate graph views, learning data invariance information through contrasting sample pairs. However, this joint approach may deviate from the expectation of semantically similar before and after augmentation. The propagation of attribute information in graphs usually occurs through their structure, meaning that structural and attribute augmentations can interfere with each other and potentially distort the graph’s semantics. To address this, we propose a decoupled augmentation framework for graph contrastive learning, which eliminates the mutual interference between the two levels of augmentation while fully exploring graph information. Specifically, our framework employs separate encoders to learn data invariance under different augmentation levels, and it considers the positive gains generated between these levels. Experimental results on five public datasets show that the proposed method is more competitive than state-of-the-art approaches.
Shihao Gao, Caoshuo Li, Cunli Mao, Xulong Zhang 0001, Xiaoyang Qu, Taisong Jin, Jianzong Wang
ICASSP1
2025 Homogeneous Graph Extraction: An Approach to Learning Heterogeneous Graph Embedding
abstract
Heterogeneous Graph Neural Networks (HGNNs) aim to embed rich structural and semantic information of heterogeneous graphs into low-dimensional node representations. While HGNNs extend the foundational work of homogeneous Graph Neural Networks, the methodology for effectively transforming heterogeneous graphs into homogeneous graphs and then learning node representations remains under-explored. In this paper, we propose a novel heterogeneous graph embedding method via the Homogeneous Graph Extraction strategy, termed HGE. Specifically, the proposed method ingeniously harnesses information clusters and metapaths to extract tailored homogeneous graphs from the complex heterogeneous graph. Subsequently, these distilled homogeneous graphs are fed into a weight-shared homogeneous graph encoder to obtain embeddings with diverse semantic information. Finally, we employ an attention mechanism, which adeptly fuses embeddings derived from distinct homogeneous graphs, resulting in the more expressive capability of the nodes. The effectiveness of the proposed architecture was demonstrated through experiments on three real heterogeneous graph datasets.
Shihao Gao, Xulong Zhang 0001, Jianzong Wang, Taisong Jin
ICASSP1
2025 Breaking the Fusion Barrier: An Online Algorithm for Fused Normalization and Linear Layers
abstract
The performance of Large Language Model (LLM) inference is critically hindered by memory-bound operations, among which normalization layers are a primary bottleneck, especially during the latency-sensitive decoding phase. While deep learning compilers fuse normalization operations into a single kernel, a fundamental fusion barrier prevents them from merging the ubiquitous Normalization and subsequent Linear layer ($N \& L$) pattern. This barrier, rooted in a core data dependency, forces the execution of two separate kernels, incurring prohibitive kernel launch overhead and costly data round-trips to global memory. In this paper, we break this barrier by introducing FlashFusion, a novel online algorithm that reformulates the$N \& L$pattern to be mathematically equivalent to a single-pass computation. Our key insight is to decompose the computation into a set of independent parallel sums, allowing the normalization statistics to be calculated concurrently with the matrix multiplication, thus eliminating the core data dependency. We co-design a high-performance, hardware-aware GPU kernel that efficiently maps this algorithm to modern architectures, leveraging a tiling strategy to maximize the utilization of Tensor Cores and the memory hierarchy. FlashFusion significantly outperforms state-of-the-art compilers like PyTorch Inductor and TensorRT, achieving speedups of up to$3.03 \times$for the$N \& L$pattern in LLM decoding and effectively eliminating the normalization bottleneck.
Hanghang Cao, Shihao Gao, Quanyi Li, Mingjie Xing
ICPADS2
2025 Exploring the Feasibility of End-to-End Large Language Model as a Compiler
abstract
In recent years, end-to-end Large Language Model (LLM) technology has shown substantial advantages across various domains. As critical system software and infrastructure, compilers are responsible for transforming source code into target code. While LLMs have been leveraged to assist in compiler development and maintenance, their potential as an end-to-end compiler remains largely unexplored. This paper explores the feasibility of LLM as a Compiler (LaaC) and its future directions. We designed the CompilerEval†dataset and framework specifically to evaluate the capabilities of mainstream LLMs in source code comprehension and assembly code generation. In the evaluation, we analyzed various errors, explored multiple methods to improve LLM-generated code, and evaluated cross-platform compilation capabilities. Experimental results demonstrate that LLMs exhibit basic capabilities as compilers but currently achieve low compilation success rates. By optimizing prompts, scaling up the model, and incorporating reasoning methods, the quality of assembly code generated by LLMs can be significantly enhanced. Based on these findings, we maintain an optimistic outlook for LaaC and propose practical architectural designs and future research directions. We believe that with targeted training, knowledge-rich prompts, and specialized infrastructure, LaaC has the potential to generate high-quality assembly code and drive a paradigm shift in the field of compilation.
Shihao Gao, Mingjie Xing
IJCNN2
2025 SWV: A Large-Scale Sensitive Word Variants Dataset for Semantic Text Matching
Jiguo Liu, Chao Liu 0020, Meimei Li, Shihao Gao, Dali Zhu
KSEM (1)5
2025 Structured Prompting and LLM Ensembling for Multimodal Conversational Aspect-based Sentiment Analysis
abstract
Understanding sentiment in multimodal conversations is a complex yet crucial challenge toward building emotionally intelligent AI systems. The Multimodal Conversational Aspect-based Sentiment Analysis (MCABSA) Challenge invited participants to tackle two demanding subtasks: (1) extracting a comprehensive sentiment sextuple-including holder, target, aspect, opinion, sentiment, and rationale-from multi-speaker dialogues, and (2) detecting sentiment flipping, which detects dynamic sentiment shifts and their underlying triggers. For Subtask-I, in the present paper, we designed a structured prompting pipeline that guided large language models (LLMs) to sequentially extract sentiment components with refined contextual understanding. For Subtask-II, we further leveraged the complementary strengths of three LLMs through ensembling to robustly identify sentiment transitions and their triggers. Our system achieved a 47.38% average score on Subtask-I and a 74.12% exact match F1 on Subtask-II, showing the effectiveness of step-wise refinement and ensemble strategies in rich, multimodal sentiment analysis tasks.
Shihao Gao, Zixing Zhang 0001, Jing Han 0010
ACM Multimedia2
2025 MLProf: A Multi-Level Runtime Performance Profiling Framework for MLIR Operations
abstract
Optimizing performance within complex MLIR compilation pipelines necessitates a precise understanding of runtime behavior at the operation level.However, most existing performance analysis tools and methodologies are not specifically designed to accurately capture execution time across MLIR's multi-level hierarchical structure, and they typically lack automated mechanisms for identifying operation-level performance bottlenecks in a developer-friendly manner.Consequently, developers often need to perform manual instrumentation, which is a labor-intensive and error-prone process.In this paper, we present MLProf, a framework specifically designed to support developers in runtime performance analysis of MLIR operations.MLProf employs an automated instrumentation mechanism based on the MLIR infrastructure and an independent runtime data capture method to achieve accurate measurement and finegrained analysis of operation execution time.MLProf provides automated analysis of operation execution time and generates intuitive visualization reports, effectively highlighting runtime bottlenecks for developers.We evaluate MLProf on representative machine learning workloads, demonstrating its capability for multi-level operation analysis, and its efficiency in identifying performance bottlenecks.
Shihao Gao, Hanghang Cao, Mingjie Xing
SEKE2
2025 COMPASS: An Agent for MLIR Compilation Pass Pipeline Generation
Shihao Gao, Mingjie Xing
TASE2
2025 Skeleton action recognition via group sparsity constrained variant graph auto-encoder
Hongjuan Pei, Shihao Gao, Taisong Jin
Image Vis. Comput.3
2024 MeshStyle: Text-driven Efficient and High-Quality 3D Mesh Stylization via Hypergraph Convolution
abstract
Text-driven 3D mesh stylization aims to transform unstylized meshes into vivid stylized 3D mesh according to the provided text prompts. Existing works have achieved impressive results in this task, but they lack a specific design for processing mesh features, such as internal interactions within mesh data, resulting in unsatisfactory stylization, and slower convergence rates. To overcome these limitations, we propose a text-driven 3D stylization framework called MeshStyle, including a novel mesh feature processing module named Mesh HyperGraph Neural Network (MHGNN). Hypergraph neural network is exploited to aggregate mesh spatial features, addressing the issue of the lack of inner connectivity in mesh data. Furthermore, we incorporate depth prior and an extra diffusion prior to enhance geometry and appearance optimization, respectively. We also constructed a new dataset collected from various public 3D datasets, along with the evaluation protocol. Through both qualitative and quantitative experiments, we validate the capability of our MeshStyle.
Shihao Gao, Songzhi Su, Xizhi Chen
ICME2
2023 LADA-Trans-NER: Adaptive Efficient Transformer for Chinese Named Entity Recognition Using Lexicon-Attention and Data-Augmentation
abstract
Recently, word enhancement has become very popular for Chinese Named Entity Recognition (NER), reducing segmentation errors and increasing the semantic and boundary information of Chinese words. However, these methods tend to ignore the semantic relationship before and after the sentence after integrating lexical information. Therefore, the regularity of word length information has not been fully explored in various word-character fusion methods. In this work, we propose a Lexicon-Attention and Data-Augmentation (LADA) method for Chinese NER. We discuss the challenges of using existing methods in incorporating word information for NER and show how our proposed methods could be leveraged to overcome those challenges. LADA is based on a Transformer Encoder that utilizes lexicon to construct a directed graph and fuses word information through updating the optimal edge of the graph. Specially, we introduce the advanced data augmentation method to obtain the optimal representation for the NER task. Experimental results show that the augmentation done using LADA can considerably boost the performance of our NER system and achieve significantly better results than previous state-of-the-art methods and variant models in the literature on four publicly available NER datasets, namely Resume, MSRA, Weibo, and OntoNotes v4. We also observe better generalization and application to a real-world setting from LADA on multi-source complex entities.
Jiguo Liu, Chao Liu 0020, Shihao Gao, Mingqi Liu, Dali Zhu
AAAI4
2023 GHGA-Net: Global Heterogeneous Graph Attention Network for Chinese Short Text Classification
Meimei Li, Yuzhi Bao, Jiguo Liu, Chao Liu 0020, Shihao Gao
PRICAI (2)6