Youpeng Zhao 0002

dblp:259/5839-2 · DBLP profile ↗
← Back
8ranked-venue papers
4as first author
8since 2021 · last 2026
0009-0006-9437-9788ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 4 · 2 first-author · 4 since 2021Systems, architecture and hardware · 3 · 2 first-author · 3 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Classifier Enhancement Using Extended Context and Domain Experts for Semantic Segmentation
abstract
Prevalent semantic segmentation methods generally adopt a vanilla classifier to categorize each pixel into specific classes. Although such a classifier learns global information from the training data, this information is represented by a set of fixed parameters (weights and biases). However, each image has a different class distribution, which prevents the classifier from addressing the unique characteristics of individual images. At the dataset level, class imbalance leads to segmentation results being biased towards majority classes, limiting the model's effectiveness in identifying and segmenting minority class regions. In this paper, we propose an Extended Context-Aware Classifier (ECAC) that dynamically adjusts the classifier using global (dataset-level) and local (image-level) contextual information. Specifically, we leverage a memory bank to learn dataset-level contextual information of each class, incorporating the class-specific contextual information from the current image to improve the classifier for precise pixel labeling. Additionally, a teacher-student network paradigm is adopted, where the domain expert (teacher network) dynamically adjusts contextual information with ground truth and transfers knowledge to the student network. Comprehensive experiments illustrate that the proposed ECAC can achieve state-of-the-art performance across several datasets, including ADE20K, COCO-Stuff10K, and Pascal-Context.
Huadong Tang, Youpeng Zhao 0002, Min Xu 0001, Jun Wang 0001, Qiang Wu 0001
IEEE Trans. Multim.2
2025 MeRino: Entropy-Driven Design for Generative Language Models on IoT Devices
abstract
Generative Large Language Models (LLMs) stand as a revolutionary advancement in the modern era of artificial intelligence (AI). However, scaling down LLMs for resource-constrained hardware, such as Internet-of-Things (IoT) devices requires non-trivial efforts and domain knowledge. In this paper, we propose a novel information-entropy framework for designing mobile-friendly generative language models. The whole design procedure involves solving a mathematical programming (MP) problem, which can be done on the CPU within minutes, making it nearly zero-cost. We evaluate our designed models, termed MeRino, across fourteen NLP downstream tasks, showing their competitive performance against the state-of-the-art autoregressive transformer models under the mobile setting. Notably, MeRino achieves similar or better performance on both language modeling and zero-shot learning tasks, compared to the 350M parameter OPT while being 4.9x faster on NVIDIA Jetson Nano with 5.5x reduction in model size.
Youpeng Zhao 0002, Huadong Tang, Qiang Wu 0001, Jun Wang 0001
AAAI1
2025 AIRES: Accelerating Out-of-Core GCNs via Algorithm-System Co-Design
abstract
Graph convolutional networks (GCNs) are fundamental in various scientific applications, ranging from biomedical protein-protein interactions (PPI) to large-scale recommendation systems. An essential component for modeling graph structures in GCNs is sparse general matrix-matrix multiplication (SpGEMM). As the size of graph data continues to scale up, SpGEMMs are often conducted in an out-of-core fashion due to limited GPU memory space in resource-constrained systems. Albeit recent efforts that aim to alleviate the memory constraints of out-of-core SpGEMM through either GPU feature caching, hybrid CPU-GPU memory layout, or performing the computation in sparse format, current systems suffer from both high I/O latency and GPU under-utilization issues. In this paper, we first identify the problems of existing systems, where sparse format data alignment and memory allocation are the main performance bottlenecks, and propose AIRES, a novel algorithm-system co-design solution to accelerate out-of-core SpGEMM computation for GCNs. Specifically, from the algorithm angle, AIRES proposes to alleviate the data alignment issues on the block level for matrices in sparse formats and develops a tiling algorithm to facilitate row block-wise alignment. On the system level, AIRES employs a three-phase dynamic scheduling that features a dual-way data transfer strategy utilizing a tiered memory system: integrating GPU memory, GPU Direct Storage (GDS), and host memory to reduce I/O latency and improve throughput. Evaluations show that AIRES significantly outperforms the state-of-the-art methods, achieving up to${1. 8} \times$lower latency in real-world graph processing benchmarks.
Shakya Jayakody, Youpeng Zhao 0002, Jun Wang 0001
ASAP2
2024 ALISE: Accelerating Large Language Model Serving with Speculative Scheduling
abstract
Large Language Models (LLMs) represent a revolutionary advancement in the contemporary landscape of artificial general intelligence (AGI). As exemplified by ChatGPT, LLM-based applications necessitate minimal response latency and maximal throughput for inference serving. However, due to the unpredictability of LLM execution, the first-come-first-serve (FCFS) scheduling policy employed by current LLM serving systems suffers from head-of-line (HoL) blocking issues and long job response times.
Youpeng Zhao 0002, Jun Wang 0001
ICCAD1
2024 ALISA: Accelerating Large Language Model Inference via Sparsity-Aware KV Caching
abstract
The Transformer architecture has significantly advanced natural language processing (NLP) and has been foundational in developing large language models (LLMs) such as LLaMA and OPT, which have come to dominate a broad range of NLP tasks. Despite their superior accuracy, LLMs present unique challenges in practical inference, concerning the compute and memory-intensive nature. Thanks to the autoregressive characteristic of LLM inference, KV caching for the attention layers in Transformers can effectively accelerate LLM inference by substituting quadratic-complexity computation with linear-complexity memory accesses. Yet, this approach requires increasing memory as demand grows for processing longer sequences. The overhead leads to reduced throughput due to I/O bottlenecks and even out-of-memory errors, particularly on resource-constrained systems like a single commodity GPU. In this paper, we propose ALISA, a novel algorithm-system co-design solution to address the challenges imposed by KV caching. On the algorithm level, ALISA prioritizes tokens that are most important in generating a new token via a Sparse Window Attention (SWA) algorithm. SWA introduces high sparsity in attention layers and reduces the memory footprint of KV caching at negligible accuracy loss. On the system level, ALISA employs three-phase token-level dynamical scheduling and optimizes the trade-off between caching and recomputation, thus maximizing the overall performance in resource-constrained systems. In a single GPU-CPU system, we demonstrate that under varying workloads, ALISA improves the throughput of baseline systems such as FlexGen and vLLM by up to $3 \times$ and $1.9 \times$, respectively.
Youpeng Zhao 0002, Di Wu 0016, Jun Wang 0001
ISCA1
2024 CAA: Class-Aware Affinity calculation add-on for semantic segmentation
abstract
Leveraging contextual dependencies is a commonly used technique to enhance the performance of image segmentation. However, existing solutions do not effectively catch the class-level association between the pixels along the boundary across the objects of the different classes but focus more on the local pixel-to-pixel relation. This work proposes a Class-Aware Affinity module (CAA) that considers both pixel-to-pixel relation and pixel-to-class association. We try to argue that the pixel-to-pixel relations still catch the relation (e.g. similarity, attention, or affiliation) on the local texture level. At the same time, it should also consider the association between the pixel and the class context produced by the given image. Pixel-to-class association can best reveal the co-occurrent dependency on the semantic level between the given pixels and their nearby context. Such pixel-to-class association combined with the pixel-to-pixel relations aggregating the local texture information will best mitigate the confusion caused in the boundary regions across the objects of the different classes. Moreover, the proposed framework can serve as a generic add-on to be integrated with the existing image segmentation solution to boost the current performance. Equipped with CAA, we achieve promising performance against the existing work with 54.59% mIoU on ADE20K, 49.96% mIoU on COCO-Stuff10k, and 64.38% mIoU on Pascal-Context.
Huadong Tang, Youpeng Zhao 0002, Chaofan Du, Min Xu 0001, Qiang Wu 0001
Knowl. Based Syst.2
2023 Class-Aware Contextual Information for Semantic Segmentation
abstract
Exploring spatial contextual information is a well-adopted approach to achieving better semantic segmentation performance. However, most existing methods neglect the class association between the neighboring pixels. In this paper, we propose a CACINet, which consists of a Semantic Affinity Module (SAM) and a Class Association Module (CAM), to generate class-aware contextual information among pixels on a fine-grained level. SAM analyzes the affiliation of any two given pixels belonging to the same or different class. It produces intra-class and inter-class pixel contextual information. CAM classifies the image into different class regions globally and then it encodes the pixel based on the degree of affiliation of the pixels with each class in the image. In this way, it augments the class affiliation of the pixels into the corresponding context calculation. Comprehensive experiments demonstrate that the proposed method achieves competitive performance on two semantic segmentation benchmarks: ADE20K and PASCAL-Context.
Huadong Tang, Youpeng Zhao 0002, Yingying Jiang 0001, Zhuoxin Gan, Qiang Wu 0001
ICASSP2
2023 Parameter-Efficient Vision Transformer with Linear Attention
abstract
Recent advances in vision transformers (ViTs) have achieved outstanding performance in visual recognition tasks, including image classification and detection. ViTs can learn global representations with their self-attention mechanism, but they are usually heavy-weight and unsuitable for resource-constrained devices. In this paper, we propose a novel linear feature attention (LFA) module to reduce computation costs for vision transformers and combine efficient mobile CNN modules to form a parameter-efficient and high-performance CNN-ViT hybrid model, called LightFormer, which can serve as a general-purpose backbone to learn both global and local representation. Comprehensive experiments demonstrate that LightFormer achieves competitive performance across different visual recognition tasks. On the ImageNet-1K dataset, LightFormer achieves top-1 accuracy of 78.5% with 5.5 million parameters. Our model also performs well when transferred to object detection and semantic segmentation tasks. On the MS COCO dataset, LightFormer attains mAP of 33.2 within the YOLOv3 framework, and on the Cityscapes dataset, with only a simple all-MLP decoder, LightFormer achieves mIoU of 78.5 and FPS of 15.3, surpassing state-of-the-art lightweight segmentation networks.
Youpeng Zhao 0002, Huadong Tang, Yingying Jiang 0001, Yong A, Qiang Wu 0001, Jun Wang 0001
ICIP1