EDBT 2026 Demo / reviewers in the wild / expert
Yumeng Shi
dblp:131/2053
· DBLP profile ↗
15ranked-venue papers
7as first author
13since 2021 · last 2026
0000-0002-9623-3778ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 7 · 5 first-author · 7 since 2021Systems, architecture and hardware · 5 · 1 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 2 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-authorDatabases, data management, data science and information retrieval · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | LLMC+: Benchmarking Vision-Language Model Compression with a plug-and-play ToolkitabstractLarge Vision-Language Models (VLMs) exhibit impressive multi-modal capabilities but suffer from prohibitive computational and memory demands, due to their long visual token sequences and massive parameter sizes. To address these issues, recent works have proposed training-free compression methods. However, existing efforts often suffer from three major limitations: (1) Current approaches do not decompose techniques into comparable modules, hindering fair evaluation across spatial and temporal redundancy. (2) Evaluation confined to simple single-turn tasks, failing to reflect performance in realistic scenarios. (3) Isolated use of individual compression techniques, without exploring their joint potential. To overcome these gaps, we introduce LLMC+, a comprehensive VLM compression benchmark with a versatile, plug-and-play toolkit. LLMC+ supports over 20 algorithms across five representative VLM families and enables systematic study of token-level and model-level compression. Our benchmark reveals that: (1) Spatial and temporal redundancies demand distinct technical strategies. (2) Token reduction methods degrade significantly in multi-turn dialogue and detail-sensitive tasks. (3) Combining token and model compression achieves extreme compression with minimal performance loss. We believe LLMC+ will facilitate fair evaluation and inspire future research in efficient VLM. Chengtao Lv, Bilang Zhang, Yang Yong, Ruihao Gong, Yushi Huang, Shiqiao Gu, Jiajun Wu 0024, Yumeng Shi, Wenya Wang 0001 |
AAAI | 8 |
| 2026 | Causality Matters: How Temporal Information Emerges in Video Language ModelsabstractVideo language models (VideoLMs) have made significant progress in multimodal understanding. However, temporal understanding, which involves identifying event order, duration, and relationships across time, still remains a core challenge. Prior works emphasize positional encodings (PEs) as a key mechanism for encoding temporal structure. Surprisingly, we find that removing or modifying PEs in video inputs yields minimal degradation in the performance of temporal understanding. In contrast, reversing the frame sequence while preserving the original PEs causes a substantial drop. To explain this behavior, we conduct substantial analysis experiments to trace how temporal information is integrated within the model. We uncover a causal information pathway: temporal cues are progressively synthesized through inter-frame attention, aggregated in the final frame, and subsequently integrated into the query tokens. This emergent mechanism shows that temporal reasoning emerges from inter-visual token interactions under the constraints of causal attention, which implicitly encodes temporal structure. Based on these insights, we propose two efficiency-oriented strategies: staged cross-modal attention and a temporal exit mechanism for early token truncation. Experiments on two benchmarks validate the effectiveness of both approaches. Yumeng Shi, Quanyu Long, Yin Wu 0001, Wenya Wang 0001 |
AAAI | 1 |
| 2026 | REDM: Regression-Guided Diffusion Modeling for Universal Soft Sensor Enhancement in Semiconductor Process ControlabstractIn semiconductor manufacturing, soft sensors play a key role in Advanced Process Control (APC) by enabling realtime wafer-to-wafer monitoring. However, their performance is often limited by sparse labeled data, process variability, and model-specific tuning. To address these challenges, we propose REDM: a Regression-Guided Diffusion Modeling framework designed to boost the accuracy and robustness of soft sensor prediction across diverse fabrication stages. REDM generates highfidelity virtual data guided by predictive regression objectives and incorporates a quality-aware filtering mechanism based on Sliced Wasserstein Distance and intra-subset Cosine Similarity. Through multi-objective selection techniques, REDM identifies informative virtual samples that balance distributional similarity and internal diversity, thereby enhancing downstream model training. We evaluate REDM on real-world datasets from three major semiconductor process stages: Chemical Vapor Deposition (CVD), Etching, and Chemical Mechanical Polishing (CMP). Across various regression models, REDM consistently enhances soft sensor performance, with an average $\mathbf{R}^{2}$ improvement of $3.27 \%$. Its independence from process-specific customization makes REDM a scalable and process-aware solution for soft sensor enhancement in smart manufacturing. Weiping Xie, Yumeng Shi, Pang Guo, Yining Chen 0001 |
ASP-DAC | 2 |
| 2025 | Static or Dynamic: Towards Query-Adaptive Token Selection for Video Question AnsweringabstractVideo question answering benefits from the rich information in videos, enabling various applications.However, the large volume of tokens generated from long videos presents challenges to memory efficiency and model performance.To alleviate this, existing works propose to compress video inputs, but often overlook the varying importance of static and dynamic information across different queries, leading to inefficient token usage within limited budgets.We propose a novel token selection strategy, EXPLORE-THEN-SELECT, that adaptively adjusts static and dynamic information based on question requirements.Our framework first explores different token allocations between key frames, which preserve spatial details, and delta frames, which capture temporal changes.Then it employs a query-aware attention-based metric to select the optimal token combination without model updates.Our framework is plug-and-play and can be seamlessly integrated within diverse video language models.Extensive experiments show that our method achieves significant performance improvements (up to 5.8%) on multiple video question answering benchmarks.Our code is available at https://github.com/ANDgate99/Explore- Then-Select. Yumeng Shi, Quanyu Long, Wenya Wang 0001 |
EMNLP | 1 |
| 2025 | CreditARF: A Framework for Corporate Credit Rating with Annual Report and Financial Feature IntegrationabstractCorporate credit rating serves as a crucial intermediary service in the market economy, playing a key role in maintaining economic order. Existing credit rating models rely on financial metrics and deep learning. However, they often overlook insights from non-financial data, such as corporate annual reports. To address this, this paper introduces a corporate credit rating framework that integrates financial data with features extracted from annual reports using FinBERT, aiming to fully leverage the potential value of unstructured text data. In addition, we have developed a large-scale dataset, the Comprehensive Corporate Rating Dataset (CCRD), which combines both traditional financial data and textual data from annual reports. The experimental results show that the proposed method improves the accuracy of the rating predictions by 8–12%, significantly improving the effectiveness and reliability of corporate credit ratings. Yumeng Shi, Zhongliang Yang, DiYang Lu, Yisi Wang, Yiting Zhou, Linna Zhou |
IJCNN | 1 |
| 2025 | GLCM: A Multimodal Framework for Credit Rating With Chain-of-Thought Reasoning
Sihan Hu, Yisi Wang, Zhongliang Yang, Yumeng Shi, Linna Zhou |
IEEE Signal Process. Lett. | 4 |
| 2024 | POSTER: ParGNN: Efficient Training for Large-Scale Graph Neural Network on GPU ClustersabstractFull-batch graph neural network (GNN) training is essential for interdisciplinary applications. Large-scale graph data is usually divided into subgraphs and distributed across multiple compute units to train GNN. The state-of-the-art load balancing method based on direct graph partition is too rough to effectively achieve true load balancing on GPU clusters. We propose ParGNN, which employs a profiler-guided load balance workflow in conjunction with graph repartition to alleviate load imbalance and minimize communication traffic. Experiments have verified that ParGNN has the capability to scale to larger clusters. Shunde Li, Junyu Gu, Jue Wang 0013, Tiechui Yao, Yumeng Shi, Shigang Li 0002, Weiting Xi, Shushen Li, Chunbao Zhou, Yangang Wang 0002, Xuebin Chi |
PPoPP | 6 |
| 2024 | Explainable prediction of deposited film thickness in IC fabrication with CatBoost and SHapley Additive exPlanations (SHAP) models
Yumeng Shi, Shunyuan Lou, Yining Chen 0001 |
Appl. Intell. | 1 |
| 2024 | HGNAS: Hardware-Aware Graph Neural Architecture Search for Edge DevicesabstractGraph Neural Networks (GNNs) are becoming increasingly popular for graph-based learning tasks such as point cloud processing due to their state-of-the-art (SOTA) performance. Nevertheless, the research community has primarily focused on improving model expressiveness, lacking consideration of how to design efficient GNN models for edge scenarios with real-time requirements and limited resources. Examining existing GNN models reveals varied execution across platforms and frequent Out-Of-Memory (OOM) problems, highlighting the need for hardware-aware GNN design. To address this challenge, this work proposes a novel hardware-aware graph neural architecture search framework tailored for resource constraint edge devices, namely HGNAS. To achieve hardware awareness, HGNAS integrates an efficient GNN hardware performance predictor that evaluates the latency and peak memory usage of GNNs in milliseconds. Meanwhile, we study GNN memory usage during inference and offer a peak memory estimation method, enhancing the robustness of architecture evaluations when combined with predictor outcomes. Furthermore, HGNAS constructs a fine-grained design space to enable the exploration of extreme performance architectures by decoupling the GNN paradigm. In addition, the multi-stage hierarchical search strategy is leveraged to facilitate the navigation of huge candidates, which can reduce the single search time to a few GPU hours. To the best of our knowledge, HGNAS is the first automated GNN design framework for edge devices, and also the first work to achieve hardware awareness of GNNs across different platforms. Extensive experiments across various applications and edge devices have proven the superiority of HGNAS. It can achieve up to a$10.6\boldsymbol{\times}$speedup and an$82.5\%$peak memory reduction with negligible accuracy loss compared to DGCNN on ModelNet40. Jianlei Yang 0001, Yingjie Qi, Yumeng Shi, Cenlin Duan, Weisheng Zhao 0001, Chunming Hu |
IEEE Trans. Computers | 5 |
| 2023 | Hardware-Aware Graph Neural Network Automated Design for Edge Computing PlatformsabstractGraph neural networks (GNNs) have emerged as a popular strategy for handling non-Euclidean data due to their state-of-the-art performance. However, most of the current GNN model designs mainly focus on task accuracy, lacking in considering hardware resources limitation and real-time requirements of edge application scenarios. Comprehensive profiling of typical GNN models indicates that their execution characteristics are significantly affected across different computing platforms, which demands hardware awareness for efficient GNN designs. In this work, HGNAS is proposed as the first Hardware-aware Graph Neural Architecture Search framework targeting resource constraint edge devices. By decoupling the GNN paradigm, HGNAS constructs a fine-grained design space and leverages an efficient multi-stage search strategy to explore optimal architectures within a few GPU hours. Moreover, HGNAS achieves hardware awareness during the GNN architecture design by leveraging a hardware performance predictor, which could balance the GNN model accuracy and efficiency corresponding to the characteristics of targeted devices. Experimental results show that HGNAS can achieve about 10.6× speedup and 88.2% peak memory reduction with a negligible accuracy loss compared to DGCNN on various edge devices, including Nvidia RTX3080, Jetson TX2, Intel i7-8700K and Raspberry Pi 3B+. Jianlei Yang 0001, Yingjie Qi, Yumeng Shi, Weisheng Zhao 0001, Chunming Hu |
DAC | 4 |
| 2023 | Lossy and Lossless (L2) Post-training Model Size CompressionabstractDeep neural networks have delivered remarkable performance and have been widely used in various visual tasks. However, their huge sizes cause significant inconvenience for transmission and storage. Many previous studies have explored model size compression. However, these studies often approach various lossy and lossless compression methods in isolation, leading to challenges in achieving high compression ratios efficiently. This work proposes a post-training model size compression method that combines lossy and lossless compression in a unified way. We first propose a unified parametric weight transformation, which ensures different lossy compression methods can be performed jointly in a post-training manner. Then, a dedicated differentiable counter is introduced to guide the optimization of lossy compression to arrive at a more suitable point for later lossless compression. Additionally, our method can easily control a desired global compression ratio and allocate adaptive ratios for different layers. Finally, our method can achieve a stable 10× compression ratio without sacrificing accuracy and a 20× compression ratio with minor accuracy loss in a short time. Our code is available at https://github.com/ModelTC/L2_Compression. Yumeng Shi, Shihao Bai, Xiuying Wei, Ruihao Gong, Jianlei Yang 0001 |
ICCV | 1 |
| 2023 | A Sparse Matrix Optimization Method for Graph Neural Networks Training
Tiechui Yao, Jue Wang 0013, Junyu Gu, Yumeng Shi, Yangang Wang 0002, Xuebin Chi |
KSEM (1) | 4 |
| 2023 | Large-Scale Simulation of Structural Dynamics Computing on GPU ClustersabstractStructural dynamics simulation plays an important role in research on reactor design and complex engineering. The Hybrid Total Finite Element Tearing and Interconnecting (HTFETI) method combined with Newmark method is an efficient way to solve large-scale structural dynamics problems. However, the sparse direct solver and the load imbalance caused by inconsistent density models are two critical issues limiting the performance and the scalability of structural dynamics computing. For the former, we propose an efficient variable-size batched method to accelerate SpMV on GPUs. For the latter, we establish an online performance prediction model, based on which we then design a novel inter-cluster subdomain fine-tuning algorithm to balance the workload of HTFETI parallel computing. We are the first to achieve the high-fidelity structural dynamics simulation of China Experimental Fast Reactor core assembly with up to 53.4 billion grids. The weak and strong scalability efficiencies reach 91.77% and 86.13% on 12,800 GPUs, respectively. Yumeng Shi, Ningming Nie, Jue Wang 0013, Kehao Lin, Chunbao Zhou, Shigang Li 0002, Kehan Yao, Shunde Li, Yangde Feng, Yangang Wang 0002 |
SC | 1 |
| 2020 | The Evaluation Model of Traditional Media Transformation Competency Based on the Grey System TheoryabstractConsidering the complexity of the traditional media transformation, a multi-criteria evaluation index system based on transformation competency is proposed using grey relational analysis(GRA) is presented. This paper demonstrates the procedure of determining the grey incidence coefficients of the evaluation indexes, the weight of each index is identified by Coefficient of variation(CV). Then, by adopting the GRA method, each factor is evaluated using the grey relational matrix. The main purpose of this study is to identify key factors of transformation competency. The result of GRA suggests a preliminary list of significant factors on traditional media transformation. Based on this paper, a comprehensive evaluation of the traditional media transformation competence is obtained. This study provide the plausible variable candidates or a good reference for analyzing traditional media transformation in future researches. Yumeng Shi, Yuming Zhu |
SMC | 1 |
| 2013 | Large-Area 2-D Electronics: Materials, Technology, and DevicesabstractRecent experiments since the discovery of monolayer graphite or graphene have led to an exciting revival in the interest in the electronic applications for graphene, as well as other 2-D materials such as hexagonal boron nitride (hBN) and molybdenum disulfide (MoS$_{2}$). These layered materials serve as an exciting new platform for flexible and transparent electronics where surfaces can be enriched with new functionality. This paper aims to provide an overview behind these new class of materials ranging upon important issues for electronic integration including synthesis all the way to current state-of-the-art circuits and devices made from these materials. Allen Hsu, Han Wang 0008, Yong Cheol Shin, Benjamin Mailly, Xu Zhang 0012, Lili Yu, Yumeng Shi, Yi Hsien Lee, Madan Dubey, Ki Kang Kim, Tomás Palacios |
Proc. IEEE | 7 |