EDBT 2026 Demo / reviewers in the wild / expert
Yingxue Gao
dblp:292/0980
· DBLP profile ↗
18ranked-venue papers
9as first author
18since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 9 · 7 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 1 first-author · 6 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Computer networks · 1 · 1 first-author · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Making Visual Dialogue More Engaging: A New Task, Method, and MetricabstractLarge language model (LLM)-based visual dialogue (VD) systems have made response generation for image-grounded conversations more correct and coherent. However, user engagement - the extent to which a user is interested, emotionally involved, and willing to continue the conversation - remains a challenge. To fully explore engaging VD, we propose: (i) a new task named Audio-enhanced VD (AVD), which introduces additional audio dialogue contexts that can more vividly convey the speaker's emotions as input, with the aim of generating correct but more engaging dialogue responses. Specifically, we employ a text-to-speech model as the modality translator to generate the paired acoustic utterances from the inputting textual utterances; (ii) an accompanying approach named Visually-grounded and Interleaved Text-Audio Dialogue Modeling (VITA-DM), which utilizes both image-grounded information and interleaved text-audio utterances for visual dialogue modeling, differentiating from previous multi-modal LLM (MLLM)-based methods that normally model text and audio modalities separately. We also present three pre-training tasks to better learn multi-modal interactions across language, vision, and audio; (iii) a novel metric named Multi-Modal Engagement (MME), which fills the gap of engagement estimation in VD and can provide a fine-grained assessment along emotional, attentional, and reply engagement dimensions (EE, AE, RE). We experiment on two popular datasets and provide extensive evaluations (automatic, engagement-specific, and human), supporting the validity of our approach. Furthermore, based on empirical results that reveal that emotions contribute the most to engagement, we justify our emphasis on the emotional aspect throughout the definition, solution, and evaluation of our task. Guanghui Ye, Huan Zhao 0003, Yingxue Gao, Zhixue Zhao, Xupeng Zha, Zhihua Jiang |
AAAI | 3 |
| 2026 | DyGIN: Modality-Unified Dynamic Graph Inception Network for Context-Aware Emotion RecognitionabstractGraph representation learning has attracted considerable attention for its ability to generate node representations by aggregating information from neighboring nodes. However, existing graph-based methods for sequential data merely focus on a static graph constructed on an entire unimodal sequence, largely ignoring their dynamic evolution. To overcome this limitation, we introduce a modality-unifiedDynamicGraphInceptionNetwork (DyGIN) that models dynamic evolutionary graphs for different modalities. DyGIN constructs dynamic graphs from temporal subsequences using a sliding window, and incorporates a Temporal Graph Evolution Gated Recurrent Unit (TGE-GRU) to update graph weights at each segment. Additionally, we propose a context-aware similarity matrix, which updates node representations based on the degree of neighboring nodes, replacing the traditional averaging method. Our model is optimized with a combination of graph structure loss, classification loss, and a learnable pooling function. We validate DyGIN on emotion recognition tasks across three public datasets—RML, IEMOCAP, and DEAP—where it outperforms state-of-the-art models by 1.6% on RML, 1.81% and 2.33% on IEMOCAP, and 2.78% and 6.81% on DEAP, based on accuracy and weighted F1 score. Our code is available at https://github.com/G22-web/DyGIN. Yingxue Gao, Huan Zhao 0003, Zixing Zhang 0001 |
IEEE Internet Things J. | 1 |
| 2026 | Geometric Contrastive Ensemble Distillation with calibration margin induction
Yang Yang 0080, Junyao Hou, Xiang Li 0067, Zhenghua Chen, Yingxue Gao, Chao Wang 0003, Min Wu 0008, Quanjun Yin |
Pattern Recognit. | 7 |
| 2025 | Enhanced Multimodal Emotion Recognition in Conversations via Contextual Filtering and Multi-Frequency Graph PropagationabstractMultimodal Emotion Recognition in Conversations (ERC) plays a crucial role in understanding human language and behavior in real-world scenarios. However, existing research tends to simply concatenate multimodal representations, failing to capture the complex relationships between modalities. Recent advances have shown that Graph Neural Networks (GNNs) are effective in capturing complex data relationships, offering a promising solution for multimodal ERC. Despite this, current GNN-based methods still face challenges, including weak interactions between modalities, neglecting the information entropy of utterances, and erasure of high-frequency signals that capture key variations and discrepancies between closely related nodes. To address these limitations, we propose a GNNs-based multi-frequency propagation method enhanced by contextual filtering for multimodal ERC. Our approach introduces a context filtering module that combines a similarity matrix and an information entropy matrix, enabling GNNs to effectively capture the inherent relationships among utterances and provide sufficient multimodal and contextual modeling. Additionally, our method explores multivariate relationships by recognizing the varying importance of emotional discrepancies and commonalities through multi-frequency signals. Experimental results on two benchmark datasets, IEMOCAP and MELD, demonstrate that our method outperforms the latest (non-)graph-based works. Our method is available at https://github.com/G22-web/ConFilMER. Huan Zhao 0003, Yingxue Gao, Haijiao Chen, Guanghui Ye, Zixing Zhang 0001 |
ICASSP | 2 |
| 2025 | Parameter-Efficient Federal-Tuning Enhances Privacy Preserving for Speech Emotion RecognitionabstractThe Pre-trained Speech Models (PSMs) generate universal speech representations using self-supervised or weakly-supervised learning from large-scale datasets. It achieves promising performance when fine-tuned for specific tasks such as Speech Emotion Recognition (SER). However, fine-tuning on various datasets requires storing the entire model’s weight parameters, complicating real-world deployment. Additionally, centralized fine-tuning relies on user data, posing significant privacy risks. To address these challenges, we propose employing Federated Learning (FL) for fine-tuning PSMs with a Parameter-Efficient Fine-Tuning (PEFT) method. By embedding trainable layers in the feed-forward layers of the pre-trained model, we keep the backbone model frozen and only update the trainable layer parameters during federated training, significantly reducing parameter transmission. Specifically, we evaluated the performance of downstream model fine-tuning, adapter tuning, embedding prompt tuning, and LoRA within a federated fine-tuning framework for PSMs, demonstrating the framework’s feasibility and effectiveness. Furthermore, attribute inference attack tests showed that gender inference results on three datasets were at chance levels. Haijiao Chen, Huan Zhao 0003, Yingxue Gao, Zixing Zhang 0001 |
ICASSP | 3 |
| 2025 | Hardware Accelerated Vision Transformer via Heterogeneous Architecture Design and Adaptive Dataflow MappingabstractVision transformer (ViT) models have demonstrated remarkable advantages in visual tasks. However, the ViT model contains various types of operators, and its sophisticated model structure imposes substantial computational complexity and storage burden. Existing hardware solutions still fail to fully unleash the ViT acceleration potential due to the mismatch between operators and hardware architectures, suffering from inefficient dataflow mapping. This work proposes HDViT, a full-fledged heterogeneous hardware accelerator on FPGA, to enhance the ViT acceleration by comprehensively analyzing and addressing the challenges of heterogeneous architecture design. Specifically, HDViT first develops a heterogeneous architecture design that is composed of multiple processing engines (PEs) to accelerate various operators in the ViT model. Then, HDViT devises a hybrid-oriented dataflow mapping strategy to reduce data transmission granularity and alleviate storage resource pressure. Lastly, to achieve the latency balancing among multiple PEs, we formulate the HDViT architecture and implement an automated exploration process to identify optimized parallelism parameters that satisfy computation and storage demands while enhancing the heterogeneous architectural performance. Experimental results indicate that HDViT achieves significant performance speedups of 2.16$\times$and 3.51$\times$compared to previous heterogeneous and unified accelerators, respectively. HDViT also achieves a maximum of 98.46% hardware utilization. Yingxue Gao, Lei Gong 0003, Chao Wang 0003, Dong Dai 0001, Yang Yang 0080, Xianglan Chen, Xi Li 0003, Xuehai Zhou |
IEEE Trans. Computers | 1 |
| 2025 | Advancing Neuromorphic Architecture Toward Emerging Spiking Neural Network on FPGAabstractSpiking neural networks (SNNs) replace the multiply-and-accumulate operations in traditional artificial neural networks (ANNs) with lightweight mask-and-accumulate operations, achieving greater performance. Existing SNN architectures are primarily designed based on fully-connected or convolutional SNN topologies and still struggle with low task accuracy, limiting their practical applications. Recently, transformer SNN (TSNN) models have shown promise in matching the accuracy of nonspiking ANNs and demonstrated potential application prospects. However, their diverse computation pattern and sophisticated network structure with high computation and memory footprints impede their efficient deployment. Thus, in this work, we move our attention to heterogeneous architecture design and propose SpikeTA, the first neuromorphic hardware accelerator explicitly designed for the TSNN model on FPGA. First, SpikeTA enables parameterizable hardware engines (HEs) designed for the network layers in TSNN, enhancing compatibility between HEs and network layers. Second, SpikeTA optimizes arithmetic operations between binary spikes and synaptic weights by presenting a DSP-efficient addition tree. By analyzing the inherent data characteristics, SpikeTA further introduces a depth-aware buffer management strategy to provide sufficient access ports. Third, SpikeTA employs a streaming dataflow mapping to optimize data transmission granularity and leverages a split-engine dataflow mapping to facilitate pipelined latency balancing. Experimental results demonstrate that SpikeTA achieves significant performance speedups of$140.73\times $–$1023.53\times $and$2.97\times $–$7.29\times $over architectures running on the AMD EPYC 7542 CPU and NVIDIA A100 GPU, respectively. SpikeTA also outperforms state-of-the-art SNN and Transformer accelerators by$2.79\times $and$2.66\times $in architecture performance while achieving a peak performance of 28.99 TOPs. Yingxue Gao, Yang Yang 0080, Lei Gong 0003, Xianglan Chen, Chao Wang 0003, Xi Li 0003, Xuehai Zhou |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2025 | Uncertainty-Aware Self-Knowledge DistillationabstractSelf-knowledge distillation has emerged as a powerful method, notably boosting the prediction accuracy of deep neural networks while being resource-efficient, setting it apart from traditional teacher-student knowledge distillation approaches. However, in safety-critical applications, high accuracy alone is not adequate; conveying uncertainty effectively holds equal importance. Regrettably, existing self-knowledge distillation methods have not met the need to improve both prediction accuracy and uncertainty quantification simultaneously. In response to this gap, we present an uncertainty-aware self-knowledge distillation method named UASKD. UASKD introduces an uncertainty-aware contrastive loss and a prediction synthesis technique within the self-knowledge distillation process, aiming to fully harness the potential of self-knowledge distillation for improving both prediction accuracy and uncertainty quantification. Extensive assessments illustrate that UASKD consistently surpasses other self-knowledge distillation techniques and numerous uncertainty calibration methods in both prediction accuracy and uncertainty quantification metrics across various classification and object detection tasks, highlighting its efficacy and adaptability. Yang Yang 0080, Chao Wang 0003, Lei Gong 0003, Min Wu 0008, Zhenghua Chen, Yingxue Gao, Xuehai Zhou |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2024 | Adaptive Speech Emotion Representation Learning Based On Dynamic GraphabstractGraph representation learning has become a hot research topic due to its powerful nonlinear fitting capability in extracting representative node embeddings. However, for sequential data such as speech signals, most traditional methods merely focus on the static graph created within a sequence, and largely overlook the intrinsic evolving patterns of these data. This may reduce the efficiency of graph representation learning for sequential data. For this reason, we propose an adaptive graph representation learning method based on dynamically evolved graphs, which are consecutively constructed on a series of subsequences segmented by a sliding window. In doing this, it is better to capture local and global context information within a long sequence. Moreover, we introduce a weighted approach to update the node representation rather than the conventional average one, where the weights are calculated by a novel matrix computation based on the degree of neighboring nodes. Finally, we construct a learnable graph convolutional layer that combines the graph structure loss and classification loss to optimize the graph structure. To verify the effectiveness of the proposed method, we conducted experiments for speech emotion recognition on the IEMOCAP and RAVDESS datasets. Experimental results show that the proposed method outperforms the latest (non-)graph-based models. Yingxue Gao, Huan Zhao 0003, Zixing Zhang 0001 |
ICASSP | 1 |
| 2024 | Bilevel Relational Graph Representation Learning-based Multimodal Emotion Recognition in ConversationabstractEmotion recognition in conversation (ERC) is a key research topic in natural language processing, helping computers understand human emotions. Although substantial strides have been made in deep learning methods, the exploration of graphs in multimodal ERC is still in its infancy. The bottleneck of graph-based methods lies in the neighborhood aggregation strategy, a mechanism through which node attributes are transmitted and gathered. Existing strategies suffer from the issue of redundant irrelevant information, causing interference with the discriminative information of nodes. Moreover, traditional single-layer graph convolutional networks face challenges in efficiently extracting long-range contextual information. To address these issues, we propose a multimodal conversational emotion recognition approach based on a bilevel relational graph (BiGraph). Specifically, we construct two graphs: a global affinity graph, clustered by assessing node similarity between the target node and its neighborhood nodes to preserve discriminative information. Another is a local context dependency graph based on information from different speakers. The edges of these graphs are mainly determined by speaker context and temporal relations. Our method yields compelling results in extensive experiments conducted on the IEMOCAP and MOSEI datasets, which demonstrate the effectiveness and superiority of the proposed model. Our code is available at https://github.com/LiMei0329/BiGraph. Huan Zhao 0003, Yi Ju, Yingxue Gao |
ICME | 3 |
| 2024 | Enhancing Graph Random Walk Acceleration via Efficient Dataflow and Hybrid Memory ArchitectureabstractGraph random walk sampling is becoming increasingly important with the widespread popularity of graph applications. It aims to capture the desirable graph properties by launching multiple walkers to collect feature paths. However, previous research suffers long sampling latency and severe memory access bottlenecks due to intrinsic data dependency and skewed vertex distribution. Thus, in this paper, we propose FastRW, a dedicated accelerator to boost graph random walk operation on FPGAs. Specifically, FastRW first integrates multiple parallel processing engines to achieve data-level parallelism, where each processing engine also leverages dataflow scheduling to resolve data dependency and hide long sampling latency. Secondly, FastRW leverages a combination of multiple storage resources to implement a hybrid memory architecture adapted to skewed vertex distribution. By integrating the above optimizations, FastRW develops a performance model to take advantage of the balance between computation parallelism and bandwidth demand. We evaluate FastRW with two classic sampling algorithms on a wide range of real-world graph datasets. The experimental results show that FastRW achieves a speedup of 37.52$\boldsymbol{\times}$on average over the system running on two 8-core Intel CPUs. FastRW also achieves an average of 28.04$\boldsymbol{\times}$speedup over the architecture implemented on V100 GPU. Yingxue Gao, Lei Gong 0003, Chao Wang 0003, Yiqing Hu, Zhongming Liu, Xi Li 0003, Xuehai Zhou |
IEEE Trans. Computers | 1 |
| 2023 | FastRW: A Dataflow-Efficient and Memory-Aware Accelerator for Graph Random Walk on FPGAsabstractGraph random walk (GRW) sampling is becoming increasingly important with the widespread popularity of graph applications. It involves some walkers that wander through the graph to capture the desirable properties and reduce the size of the original graph. However, previous research suffers long sampling latency and severe memory access bottlenecks due to intrinsic data dependency and irregular vertex distribution. This paper proposes FastRW, a dedicated accelerator to release GRW acceleration on FPGAs. FastRW first schedules walkers' execution to address data dependency and mask long sampling latency. Then, FastRW leverages pipeline specialization and bit-level optimization to customize a processing engine with five modules and achieve a pipelining dataflow. Finally, to alleviate the differential accesses caused by irregular vertex distribution, FastRW implements a hybrid memory architecture to provide parallel access ports according to the vertex's degree. We evaluate FastRW with two classic GRW algorithms on a wide range of real-world graph datasets. The experimental results show that FastRW achieves a speedup of 14.13× on average over the system running on two 8-core Intel CPUs. FastRW also achieves 3.28×∼198.24× energy efficiency over the architecture implemented on V100 GPU. Yingxue Gao, Lei Gong 0003, Chao Wang 0003, Xi Li 0003, Xuehai Zhou |
DATE | 1 |
| 2023 | Algorithm/Hardware Co-Optimization for Sparsity-Aware SpMM Acceleration of GNNsabstractIn recent years, graph neural networks (GNNs) have achieved impressive performance in various application fields by extracting information from graph-structured data. It contains extensive feature aggregation operations and has become a performance bottleneck, which can be abstracted as a specialized sparse-dense matrix multiplication (SpMM) operation. Previous works have leveraged the inner product or outer product to accelerate the feature aggregation process. However, inefficient execution leads to extremely unbalanced workloads and extensive intermediate data, hampering the performance of previous processors. So in this article, we demonstrate an algorithm/hardware co-optimization chance to enhance SpMM acceleration for GNNs. First, the algorithm part develops a dataflow-efficient SpMM algorithm that integrates three optimization methods to mitigate computation and memory access inefficiencies. Specifically, 1) the proposed equal-value partition method achieves fine-grained data partition and enables load balancing during data movement; 2) after observing the vertex aggregation phenomenon, a vertex-clustering optimization method is presented to enable significant data locality; and 3) the adaptive dataflow based on Gustavson’s algorithm is further implemented to enable the efficient distribution of sparse elements and improves computing resource utilization. Then, the hardware part features the proposed SpMM algorithm and customizes SDMA, a flexible and efficient accelerator to boost SpMM acceleration, which follows the adaptive dataflow to eliminate sparsity and explore the regular parallelism dimension. Finally, we prototype SDMA on the Xilinx Alveo U280 FPGA accelerator card. The results demonstrate that SDMA achieves$5.68\times $–$14.68\times $energy efficiency over the previous GPU implementations deployed on the Nvidia GTX 1080Ti and$1.32\times $higher throughput over the state-of-the-art FPGA prototype. Yingxue Gao, Lei Gong 0003, Chao Wang 0003, Xi Li 0003, Xuehai Zhou |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2022 | Work-in-Progress: HeteroRW: A Generalized and Efficient Framework for Random Walks in Graph AnalysisabstractRandom walk (RW) is a common graph analysis algorithm that consists of two phases: construction and sampling. The construction phase is responsible for generating the sampling table. The sampling phase contains many walkers which wander through the whole graph to sample. However, RW is notorious for its dynamic and sparse memory access pattern, which makes existing research suffer low throughput and memory bottleneck. In addition, the variety of RW algorithms in different scenarios also brings new design challenges.This paper proposes HeteroRW, a generalized framework to accelerate RWs on FPGAs. HeteroRW first identifies the two phases’ computation characteristics and presents corresponding hardware acceleration designs, respectively. Then, HeteroRW achieves the template-based design to support a variety of RW algorithms. Finally, HeteroRW integrates a novel scheduling layer to partition the input data and perform design space exploration (DSE). Experimental results show that HeteroRW achieves 4.3x speedup over the recent FPGA implementation while effectively simplifying the accelerator customization process. Yingxue Gao, Lei Gong 0003, Chao Wang 0003, Xuehai Zhou |
CODES+ISSS | 1 |
| 2022 | SDMA: An Efficient and Flexible Sparse-Dense Matrix-Multiplication Architecture for GNNsabstractIn recent years, graph neural networks (GNNs) as a deep learning model have emerged. Sparse-Dense Matrix Multiplication (SpMM) is the critical component of GNNs. However, SpMM involves many irregular calculations and random memory accesses, resulting in the inefficiency of general-purpose processors and dedicated accelerators. The highly sparse and uneven distribution of the graph further exacerbates the above problems. In this work, we propose SDMA, an efficient architecture to accelerate SpMM for GNNs. SDMA can collaboratively address the challenges of load imbalance and irregular memory accesses. We first present three hardware-oriented optimization methods: 1) The Equal-value partition method effectively divides the sparse matrix to achieve load balancing between tiles. 2) The vertex-clustering optimization method can explore more data locality. 3) An adaptive on-chip dataflow scheduling method is proposed to make full use of computing resources. Then, we combine and integrate the above optimization into SDMA to achieve a high-performance architecture. Finally, we prototype SDMA on the Xilinx Alveo U50 FPGA. The results demonstrate that SDMA achieves 2.19x-3.35x energy efficiency over the GPU implementation and 2.03x DSP efficiency over the FPGA implementation. Yingxue Gao, Lei Gong 0003, Chao Wang 0003, Xuehai Zhou |
FPL | 1 |
| 2022 | Multi-clusters: An Efficient Design Paradigm of NN Accelerator Architecture Based on FPGA
Lei Gong 0003, Chao Wang 0003, Yang Yang 0080, Yingxue Gao |
NPC | 5 |
| 2022 | ViA: A Novel Vision-Transformer Accelerator Based on FPGAabstractSince Google proposed Transformer in 2017, it has made significant natural language processing (NLP) development. However, the increasing cost is a large amount of calculation and parameters. Previous researchers designed and proposed some accelerator structures for transformer models in field-programmable gate array (FPGA) to deal with NLP tasks efficiently. Now, the development of Transformer has also affected computer vision (CV) and has rapidly surpassed convolution neural networks (CNNs) in various image tasks. And there are apparent differences between the image data used in CV and the sequence data in NLP. The details in the models contained with transformer units in these two fields are also different. The difference in terms of data brings about the problem of the locality. The difference in the model structure brings about the problem of path dependence, which is not noticed in the existing related accelerator design. Therefore, in this work, we propose the ViA, a novel vision transformer (ViT) accelerator architecture based on FPGA, to execute the transformer application efficiently and avoid the cost of these challenges. By analyzing the data structure in the ViT, we design an appropriate partition strategy to reduce the impact of data locality in the image and improve the efficiency of computation and memory access. Meanwhile, by observing the computing flow of the ViT, we use the half-layer mapping and throughput analysis to reduce the impact of path dependence caused by the shortcut mechanism and fully utilize hardware resources to execute the Transformer efficiently. Based on optimization strategies, we design two reuse processing engines with the internal stream, different from the previous overlap or stream design patterns. In the stage of the experiment, we implement the ViA architecture in Xilinx Alveo U50 FPGA and finally achieved ~5.2 times improvement of energy efficiency compared with NVIDIA Tesla V100, and 4–10 times improvement of performance compared with related accelerators based on FPGA, that obtained nearly 309.6 GOP/s computing performance in the peek. Lei Gong 0003, Chao Wang 0003, Yang Yang 0080, Yingxue Gao, Xuehai Zhou, Huaping Chen 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2021 | Upgraded Attention-Based Local Feature Learning Block for Speech Emotion Recognition
Huan Zhao 0003, Yingxue Gao, Yufeng Xiao |
PAKDD (2) | 2 |