Yiming Shi

dblp:246/9181 · DBLP profile ↗
← Back
14ranked-venue papers
3as first author
13since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 3 · 1 first-author · 3 since 2021Systems, architecture and hardware · 3 · 1 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 since 2021Computer networks · 2 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021Security and privacy · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 An FD-SOI-Based Compact In-Pixel Computing Architecture Enabling Real-Time Feature Extraction
abstract
To empower resource-limited edge devices in artificial intelligence (AI) and Internet of Things (IoT) applications, it is essential to overcome challenges posed by restricted area resources and the high latency demands of transmitting and processing substantial sensory data. In-pixel computing addresses these challenges effectively, and the Fully Depleted Silicon-On-Insulator (FD-SOI)-based pixel, which relies on an FD-SOI transistor whose current is made photosensitive to light by applying a negative back-gate voltage, shows significant potential with its compact structure and in-situ computation capability. In this paper, for the first time, we present an FD-SOI-based chip-level architecture for in-pixel computing. Our design implements programmable, massively parallel convolution with low latency using a pulse-width modulation (PWM) input encoding scheme. Furthermore, the proposed compact 1P1T (1 Phototransistor 1 Transistor) pixel design, integrated with an improved single-slope analog-to-digital converter (SS ADC), greatly enhances area efficiency. Validated through simulation in a 22nm FD-SOI process, the design achieves 990 frames/s under typical outdoor illumination conditions, with the figure of merit (FoM) of 16.86pJ/pixel/frame. In addition, the proposed architecture has been evaluated on hand gesture recognition (5697 training and 633 validation images across six categories), achieving an accuracy of 97.48%. The results demonstrate that, compared to state-of-the-art designs, our approach achieves a$6\times $improvement in in-pixel convolution speed and a$7.3\times $reduction in area overhead.
Yijiao Wang, Jiayao Wu, Zhongzhen Tong, Xinrui Duan, Yiming Shi, Weisheng Zhao 0001
IEEE Trans. Circuits Syst. I Regul. Pap.6
2026 LoLDU: Low-Rank Adaptation via Lower-Diag-Upper Decomposition for Parameter-Efficient Fine-Tuning
abstract
The rapid growth of model scale has necessitated substantial computational resources for fine-tuning. Existing approach such as low-rank adaptation (LoRA) has sought to address the problem of handling the large updated parameters in full fine-tuning (FT). However, LoRA utilize random initialization and optimization of low-rank matrices to approximate updated weights, which can result in suboptimal convergence and an accuracy gap compared to full fine-tuning (FT). To address these issues, we propose low-rank LDU (LoLDU), a parameter-efficient fine-tuning (PEFT) approach that significantly reduces trainable parameters by 2600 times compared to regular PEFT methods while maintaining comparable performance. LoLDU leverages lower-diag-upper (LDU) decomposition to initialize low-rank matrices for faster convergence and nonsingularity. We focus on optimizing the diagonal matrix for scaling transformations. To the best of our knowledge, LoLDU has the fewest parameters among all PEFT approaches. We conducted extensive experiments across 4 instruction-following datasets, six natural language understanding (NLU) datasets, eight image classification datasets, and image generation datasets with multiple model types [LLaMA2, RoBERTa, ViT, and stable diffusion (SD)], providing a comprehensive and detailed analysis. Our open-source code can be accessed at https://anonymous.4open.science/r/LoLDU-B5A6.
Yiming Shi, Yujia Wu, Jiwei Wei, Ran Ran 0001, Cheng-Wei Sun, Shiyuan He, Yang Yang 0002
IEEE Trans. Neural Networks Learn. Syst.1
2025 Med-2E3: A 2D-Enhanced 3D Medical Multimodal Large Language Model
abstract
3D medical image analysis is essential for modern healthcare, yet traditional task-specific models are inadequate due to limited generalizability across diverse clinical scenarios. Multimodal large language models (MLLMs) offer a promising solution to these challenges. However, existing MLLMs have limitations in fully leveraging the rich, hierarchical information embedded in 3D medical images. Inspired by clinical practice, where radiologists focus on both 3D spatial structure and 2D planar content, we propose Med-2E3, a 3D medical MLLM that integrates a dual 3D-2D encoder architecture. To aggregate 2D features effectively, we design a Text-Guided Inter-Slice (TG-IS) scoring module, which scores the attention of each 2D slice based on slice contents and task instructions. To the best of our knowledge, Med-2E3 is the first MLLM to integrate both 3D and 2D features for 3D medical image analysis. Experiments on large-scale, open-source 3D medical multimodal datasets demonstrate that TG- IS exhibits task-specific attention distribution and sig-nificantly outperforms current state-of-the-art models. The code is available at: https://github.com/MSIIPlMed-2E3
Yiming Shi, Chenyi Guo, Miao Li 0003, Ji Wu 0002
BIBM1
2025 3D-HSPA: Integrating 3D Spatial Information with Hierarchical Slice-Patch Attention for Knee MRI Analysis
abstract
Magnetic Resonance Imaging (MRI) is a crucial modality for diagnosing knee joint diseases. However, accurately extracting disease-relevant features from complex multi-slice, multi-sequence MRI scans remains a considerable challenge. To address this, we propose 3D-HSPA, a novel diagnostic framework for multi-slice, multi-sequence knee MRI, which integrates disease-specific information at both slice and patch levels and establishes intrinsic spatial connections among different sequences. Specifically, we introduce a patch-level and slice-level label attention mechanism, guiding the model to automatically learn a precise alignment between image regions and disease labels. Furthermore, by mapping 2D images from various sequences into a unified 3D spatial coordinate system, we enhance the spatial consistency and robustness of the attention distributions. We validated 3D-HSPA on a large-scale MRI dataset comprising 50 fine-grained types of knee joint diseases. The experimental results demonstrate that 3D-HSPA not only achieves superior diagnostic performance but also exhibits strong model interpretability.
Jingzhi Yang, Yiming Shi, Ji Wu 0002, Huishu Yuan, Miao Li 0003, Xiangling Fu
BIBM2
2025 USD: Unsupervised Soft Contrastive Learning for Fault Detection in Multivariate Time Series
abstract
Unsupervised fault detection in multivariate time series is critical for maintaining the integrity and efficiency of complex systems, with current methodologies largely focusing on statistical and machine learning techniques. However, these approaches often rest on the assumption that data distributions conform to Gaussian models, overlooking the diversity of patterns that can manifest in both normal and abnormal states, thereby diminishing discriminative performance. Our innovation addresses this limitation by introducing a combination of data augmentation and soft contrastive learning, specifically designed to capture the multifaceted nature of state behaviors more accurately. The data augmentation process enriches the dataset with varied representations of normal states, while soft contrastive learning fine-tunes the model’s sensitivity to the subtle differences between normal and abnormal patterns, enabling it to recognize a broader spectrum of anomalies. This dual strategy significantly boosts the model’s ability to distinguish between normal and abnormal states, leading to a marked improvement in fault detection performance across multiple datasets and settings, thereby setting a new benchmark for unsupervised fault detection in complex systems.
Xiuxiu Qiu, Yiming Shi, Zelin Zang
ICASSP3
2025 Connector-S: A Survey of Connectors in Multi-modal Large Language Models
abstract
With the rapid advancements in multi-modal large language models (MLLMs), connectors play a pivotal role in bridging diverse modalities and enhancing model performance. However, the design and evolution of connectors have not been comprehensively analyzed, leaving gaps in understanding how these components function and hindering the development of more powerful connectors. In this survey, we systematically review the current progress of connectors in MLLMs and present a structured taxonomy that categorizes connectors into atomic operations (mapping, compression, mixture of experts) and holistic designs (multi-layer, multi-encoder, multi-modal scenarios), highlighting their technical contributions and advancements. Furthermore, we discuss several promising research frontiers and challenges, including high-resolution input, dynamic compression, guide information selection, combination strategy, and interpretability. This survey is intended to serve as a foundational reference and a clear roadmap for researchers, providing valuable insights into the design and optimization of next-generation connectors to enhance the performance and adaptability of MLLMs.
Xi Chen 0009, Yiming Shi, Miao Li 0003, Ji Wu 0002
IJCAI4
2025 Machine Vision Quality Assessment for Image Restoration
abstract
In recent years, substantial progress has been made in the realm of No-Reference Image Quality Assessment (NR-IQA) for image restoration, where performance has been predominantly evaluated using metrics such as BRISQUE [1] and Hyper-IQA [2]. However, these NR-IQA metrics assess the perceptual quality of images without considering their utility in specific machine vision tasks, such as object detection and semantic segmentation. In this paper, we propose a Machine Vision Quality Assessment (MVQA) framework for image restoration. Specifically, we introduce the weighted Alternative Free-response Operating Characteristic (wAFROC) [3] as a metric to assess the machine vision quality of three image restoration sub-tasks: dehazing, denoising, and superresolution. By accounting for both detection sensitivity and spatial localization, wAFROC provides a more comprehensive evaluation of image quality in machine vision contexts. Its effectiveness is validated through downstream tasks, specifically object detection and semantic segmentation. We construct an IQA dataset for image restoration to explore the impact of various image restoration algorithms on the accuracy of object detection and semantic segmentation algorithms. Extensive experimental results demonstrate that the MVQA framework, leveraging wAFROC, effectively predicts the influence of image quality on machine vision tasks, bridging the gap between perceptual IQA and task-specific quality requirements in machine vision applications.
Yiming Shi, Xiongkuo Min, Guangtao Zhai
ISCAS1
2025 Enhancing Multi-task Learning Capability of Medical Generalist Foundation Model via Image-centric Multi-annotation Data
abstract
The emergence of medical generalist foundation models has revolutionized conventional task-specific model development paradigms, aiming to better handle multiple tasks through joint training on large-scale medical datasets. However, recent advances prioritize simple data scaling or architectural component enhancement, while neglecting to re-examine multi-task learning from a data-centric perspective. Critically, simply aggregating existing data resources leads to decentralized image-task alignment, which fails to cultivate comprehensive image understanding or align with clinical needs for multi-dimensional image interpretation. In this paper, we introduce the image-centric multi-annotation X-ray dataset (IMAX), the first attempt to enhance the multi-task learning capabilities of medical multi-modal large language models (MLLMs) from the data construction level. To be specific, IMAX is featured from the following attributes: 1) High-quality data curation. A comprehensive collection of more than 354K entries applicable to seven different medical tasks. 2) Image-centric dense annotation. Each X-ray image is associated with an average of 4.10 tasks and 7.46 training entries, ensuring multi-task representation richness per image. Compared to the general decentralized multi-annotation X-ray dataset (DMAX), IMAX consistently demonstrates significant multi-task average performance gains ranging from 3.20% to 21.05% across seven open-source state-of-the-art medical MLLMs. Moreover, we investigate differences in statistical patterns exhibited by IMAX and DMAX training processes, exploring potential correlations between optimization dynamics and multi-task performance. Finally, leveraging the core concept of IMAX data construction, we propose an optimized DMAX-based training strategy to alleviate the dilemma of obtaining high-quality IMAX data in practical scenarios. Related resources are available at https://github.com/MSIIP/IMAX.
Fanbin Mo, Yiming Shi, Ming Wu 0001, Miao Li 0003, Ji Wu 0002
ACM Multimedia5
2025 SPEVS-CC: Separated parameter estimation with variable selection based on canonical correlation analysis for multivariate functional regression
abstract
The rapid advancement of real-time data monitoring technologies has established functional data analysis as a crucial predictive tool across diverse applications. Despite progress in multivariate functional regression and principal component analysis, significant challenges persist in function-to-function regression and feature selection. These include the inaccurate selection of predictor variables from extensive predictors and flawed parameter estimation, which compromise the precision of function-based data predictions. This paper introduces the SPEVS-CC method, a novel canonical correlation-based feature selection technique specifically designed for function-on-function regression. By effectively decoupling variable selection from model fitting, the SPEVS-CC method enhances both adaptability and interpretability. Validated through rigorous experiments and real-world applications, including in the Shenzhen subway system, semiconductor manufacturing, and ocean climate, SPEVS-CC significantly reduces mean squared error, confirming its robustness and practical utility. This methodological breakthrough harmonizes variable selection with functional regression, providing unmatched interpretability and usability in industrial applications.
Xing Yang 0003, Haijie Xu, Yiming Shi, Chen Zhang 0007
Adv. Eng. Informatics3
2025 Deep Multimanifold Transformation-Based Multivariate Time Series Fault Detection
abstract
Unsupervised fault detection in multivariate time series (MTS) plays a vital role in ensuring the stable operation of complex systems. Traditional methods often assume that normal data follow a single Gaussian distribution and identify anomalies as deviations from this distribution. However, this simplified assumption fails to capture the diversity and structural complexity of real-world time series, which can lead to misjudgments and reduced detection performance in practical applications. To address this issue, we propose a new method that combines a neighborhood-driven data augmentation strategy with a multimanifold representation learning framework. By incorporating information from local neighborhoods, the augmentation module can simulate contextual variations of normal data, enhancing the model's adaptability to distributional changes. In addition, we design a structure-aware feature learning approach that encourages natural clustering of similar patterns in the feature space while maintaining sufficient distinction between different operational states. Extensive experiments on several public benchmark datasets demonstrate that our method achieves superior performance in terms of both accuracy and robustness, showing strong potential for generalization and real-world deployment.
Xiuxiu Qiu, Yiming Shi, Zelin Zang, Zhen Lei 0001
IEEE Trans. Neural Networks Learn. Syst.3
2023 Analyzing Multi-player Equalizer Strategy in Iterated Threshold Public Goods Game
abstract
In this paper, we investigate the multi-player equalizer strategy in the public goods game with threshold, a kind of key zero-determinant strategies in iterated games that can enforce opponents' long-term payoff to a fixed value. This paper studies the role of the equalizer strategy on the total amount of endowment in the particular scenario of the iterated multi-player game, where everyone will suffer a loss if the endowment does not exceed a threshold. Our results show that the threshold that the equalizer strategy can enforce to reach varies depending on the reward and risk factors. This also provides a potential clue on how humanity should reach a consensus to avoid the collective-risk social dilemma.
Yuelin Lyu, Yiming Shi, Zhihai Rong
ISCAS2
2023 Physical Layer Identity Information Protection against Malicious Millimeter Wave Sensing
abstract
Gait recognition based on millimeter waves (mmWave) can recognize people's identity information by sensing their walking posture, which has found versatile usages in many fields, such as smart home, intelligent security, and health monitoring. While this technology has gained extensive attention in recent years, its possibility of being misused is also increasing. The snooper who misuses the technology could monitor the victim's identity information, which is imperceptible due to the characteristics of mmWave-based gait recognition. In this paper, we propose an identity protector called WW-IDguard, which disrupts the snooper at the physical level. The key idea is that the protector sends a unique signal to interfere with not only the signal but also the gait feature of the person “seen” by the snooper. Experiments demonstrate that WW-IDguard can significantly reduce the accuracy of the mmWave-based gait recognition used by snoopers. We also perform a measurement analysis on the basic method.
Yiming Shi, Yumeng Liang, Xinzhe Wen, Anfu Zhou, Huadong Ma, Hairong Qian
ISCC2
2021 Fine-Grained Intra-domain Bandwidth Allocation Against DDoS Attack
Lijia Xie, Xiao Zhang 0004, Yiming Shi, Zhiming Zheng 0001
SecureComm (1)4
2020 Understanding Operational 5G: A First Measurement Study on Its Coverage, Performance and Energy Consumption
abstract
5G, as a monumental shift in cellular communication technology, holds tremendous potential for spurring innovations across many vertical industries, with its promised multi-Gbps speed, sub-10 ms low latency, and massive connectivity. On the other hand, as 5G has been deployed for only a few months, it is unclear how well and whether 5G can eventually meet its prospects. In this paper, we demystify operational 5G networks through a first-of-its-kind cross-layer measurement study. Our measurement focuses on four major perspectives: (i) Physical layer signal quality, coverage and hand-off performance; (ii) End-to-end throughput and latency; (iii) Quality of experience of 5G's niche applications (e.g., 4K/5.7K panoramic video telephony); (iv) Energy consumption on smartphones. The results reveal that the 5G link itself can approach Gbps throughput, but legacy TCP leads to surprisingly low capacity utilization (< 32%), latency remains too high to support tactile applications and power consumption escalates to 2 - 3x over 4G. Our analysis suggests that the wireline paths, upper-layer protocols, computing and radio hardware architecture need to co-evolve with 5G to form an ecosystem, in order to fully unleash its potential.
Dongzhu Xu, Anfu Zhou, Xinyu Zhang 0003, Guixian Wang, Congkai An, Yiming Shi, Liang Liu 0001, Huadong Ma
SIGCOMM7