Yan Wang 0022

dblp:59/2227-22 · DBLP profile ↗
← Back
14ranked-venue papers
8as first author
10since 2021 · last 2026
0000-0002-9462-3983ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 8 · 5 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021Computer networks · 1 · 1 first-author · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
YearPublicationVenuePosition
2026 WaveSculpt: Text-to-3D generation with wavelet-guided score distillation
Weilong Peng, Jianhui Huo, Keke Tang, Yangtao Wang, Yan Wang 0022, Meie Fang
Comput. Aided Geom. Des.5
2026 Cost-Optimized Periodic DAG-Structured Task Offloading in Multi-User MEC Systems Using Reinforcement Learning
abstract
Reinforcement Learning (RL) has emerged as a promising solution for task offloading due to its adaptability to dynamic environments and ability to reduce online computational overhead. Thereby, this article explores RL for optimizing periodic Directed Acyclic Graph (DAG) task offloading in multi-user Mobile Edge Computing (MEC) systems, aiming to minimize overall costs, including user device energy consumption and server computational charges. A key contribution of this work is the explicit modeling of user competition for limited edge resources, where concurrent access leads to dynamic contention, significantly affecting offloading latency and energy usage. However, this optimization task faces two main challenges: the high dimensionality of task states and the large action space, both of which increase learning complexity. To address this, we propose a dynamic and distributed Proximal Policy Optimization (PPO)-based offloading framework. An encoder is employed to map DAG node features and structural information into a lower-dimensional representation, reducing computational overhead and improving learning efficiency. Additionally, we incorporate behavioral cloning to imitate greedy policies as the PPO agent’s initial behavior, effectively narrowing the action space and accelerating convergence. By combining representation learning and imitation-based initialization, our method enables the PPO agent to quickly adapt to environmental dynamics, leveraging both prior knowledge and real-time feedback to make informed offloading decisions. Simulation results confirm that our approach achieves rapid convergence and outperforms existing baselines in cost reduction, demonstrating its effectiveness for periodic task offloading in MEC scenarios. The source code and implementation details are available at: https://github.com/xiaolutihua/GAT/tree/master .
Yan Wang 0022, Gang Liu 0038, Keqin Li 0001
ACM Trans. Internet Techn.1
2025 From Pixels to Shapes: Generative AI for 2D Images and 3D Models
Jianhui Huo, Shijian Xu, Weilong Peng, Yangtao Wang, Yan Wang 0022, Meie Fang
ICIC (19)6
2025 HGACF: A Homogeneous Neighbor Graph Contrastive Learning Framework for Enhanced Collaborative Filtering
Yan Wang 0022, Jinting Nie, Weilong Peng
ICIC (19)1
2025 Cherry: Breaking the GPU Memory Wall for Large-Scale GNN Training via Micro-Batching
abstract
Graph Neural Networks (GNNs) have shown remarkable performance across a variety of graph-related tasks.Recent efforts indicate that GNN performance can be enhanced through more sophisticated strategies, such as employing advanced aggregators, increasing aggregation depth, and utilizing larger sampling rates, etc.While these strategies yield promising results, it also incurs a significantly larger memory footprint that can easily surpass the GPU memory capacity.Micro-batching has emerged as a promising method to mitigate GPU memory bottleneck while preserving model accuracy.Nevertheless, integrating micro-batches into GNN
Yan Wang 0022, Haoran Kong, Hao Chen 0002, Weile Jia, Dingwen Tao, Xin He 0054
ICS1
2025 Adaptive Multi-Lens Phase Modulation for Scale-Aware Privacy-Preserving Human Pose Recognition
abstract
Recently, optical privacy protection has emerged as a promising approach for safeguarding visual privacy at the physical acquisition stage. However, existing methods often face a trade‐off between privacy strength and human pose recognition accuracy, particularly in long‐range and multi‐scale scenarios. To address this challenge, we propose a novel adaptive optical privacy‐preserving framework that integrates a learnable optical modulation system with a human pose recognition network. The core of our method lies in a sparse‐weighted multi‐lens model, where a lightweight multilayer perceptron (MLP) predicts a sparse set of coefficients to linearly combine predefined lens phase profiles based on facial region geometry. This enables dynamic control over the point spread function (PSF), adapting the degree of image degradation to subject scale in real time. Additionally, we introduce a privacy‐aware loss function that selectively reduces facial localization accuracy while preserving body pose information. Extensive experiments on MSCOCO and FLIC datasets demonstrate that the proposed method achieves a favorable balance between privacy protection and pose estimation, outperforming previous optical‐ and software‐based baselines.
Weilong Peng, Quanwei Deng, Mingjie Li 0004, Yangtao Wang, Yan Wang 0022, Lisheng Fan, Meie Fang
IET Softw.5
2025 DC-ORAM: An ORAM Scheme Based on Dynamic Compression of Data Blocks and Position Map
abstract
Oblivious RAM (ORAM) is an efficient cryptographic primitive that prevents leakage of memory access patterns. It has been referenced by modern secure processors and plays an important role in memory security protection. Although the most advanced ORAM has made great progress in performance optimization, the access overhead (i.e., data blocks) and on-chip (i.e., PosMap) storage overhead is still too high, which will lead to problems such as low system performance. To overcome the above challenges, in this paper, we propose a DC-ORAM system, which reduces the data access overhead and on-chip PosMap storage overhead by using dynamic compression technology. Specifically, we use byte stream redundancy compression technology to compress data blocks on the ORAM tree. And in PosMap, a high-bit multiplexing strategy is used to achieve data compression for binary high-bit repeated data of leaf labels (or path labels). By introducing the above compression technology, in this work, compared with conventional Path ORAM, the compression rate of the ORAM tree is$52.9\%$, and the compression rate of PosMap is$40.0\%$. In terms of performance, compared to conventional Path ORAM, our proposed DC-ORAM system reduces the average latency by$33.6\%$. In addition, we apply the compression technology proposed in this work to the Ring ORAM system. By comparison, it is found that with the same compression ratio as Path ORAM, our design can still reduce latency by an average of$21.5\%$.
Chuang Li 0004, Changyao Tan, Gang Liu 0038, Yanhua Wen, Yan Wang 0022, Kenli Li 0001
IEEE Trans. Computers5
2022 Optimizing data placement and size configuration for morphable NVM based SPM in embedded multicore systems
Linbo Long, Jinpei Du, Xuxu Deng, Renping Liu 0002, Yan Wang 0022
Future Gener. Comput. Syst.6
2022 Performance-oriented cache management scheme based on a retention state for energy-harvesting nonvolatile processors
Yan Wang 0022, Henian Fang, Linbo Long
Future Gener. Comput. Syst.1
2022 Performance-aware cache management for energy-harvesting nonvolatile processors
Yan Wang 0022, Kenli Li 0001, Xia Deng, Keqin Li 0001
J. Supercomput.1
2020 Communication-Aware Task Scheduling for Energy-Harvesting Nonvolatile Processors
abstract
With the advent of the Internet of Things (IoT), energy-harvesting nonvolatile processors (NVPs) have become promising platforms due to their durability when running on an intermittent power supply and fast read/write operations. However, the penalties caused by an increasing amount of data to be processed and growing communication demands pose critical challenges in the scheduling of tasks on energy-harvesting NVP platforms with tight energy and latency budgets. Together with the high power switching overhead that is induced under unstable power conditions, an increase in the amount of data to be processed significantly degrades the system performance. To overcome the problems of high communication and switching overheads for energy-harvesting NVP platforms, this article proposes a novel communication-aware task-scheduling technique. The algorithm first selects one or more executable tasks to be performed based on the task benefits and then calls a task partitioning algorithm to dynamically divide the scheduled tasks. We evaluate the performance of our proposed algorithm in comparison with the performance-aware task scheduling (PATH) and greedy iterative (GI) algorithms. Experimental results show that the proposed algorithm can reduce the execution time by 17.08% and 13.72% on average compared with the PATH and GI algorithms, respectively.
Yan Wang 0022, Jingtong Hu
IEEE Trans. Very Large Scale Integr. Syst.1
2016 Data-aware task scheduling on heterogeneous hybrid memory multiprocessor systems
abstract
Summary In this paper, we propose a method about task scheduling and data assignment on heterogeneous hybrid memory multiprocessor systems for real‐time applications. In a heterogeneous hybrid memory multiprocessor system, an important problem is how to schedule real‐time application tasks to processors and assign data to hybrid memories. The hybrid memory consists of dynamic random access memory and solid state drives when considering the performance of solid state drives into the scheduling policy. To solve this problem, we propose two heuristic algorithms called improvement greedy algorithm and the data assignment according to the task scheduling algorithm, which generate a near‐optimal solution for real‐time applications in polynomial time. We evaluate the performance of our algorithms by comparing them with a greedy algorithm, which is commonly used to solve heterogeneous task scheduling problem. Based on our extensive simulation study, we observe that our algorithms exhibit excellent performance and demonstrate that considering data allocation in task scheduling is significant for saving energy. We conduct experiments on two heterogeneous multiprocessor systems. Copyright © 2016 John Wiley & Sons, Ltd.
Kenli Li 0001, Zhuo Tang, Chubo Liu, Yan Wang 0022, Keqin Li 0001
Concurr. Comput. Pract. Exp.5
2015 Minimizing write operation for multi-dimensional DSP applications via a two-level partition technique with complete memory latency hiding
Yan Wang 0022, Kenli Li 0001, Keqin Li 0001
J. Syst. Archit.1
2012 Loop scheduling optimization for chip-multiprocessors with non-volatile main memory
abstract
Non-Volatile Memories (NVMs) have many advantages over traditional DRAM. It is desirable to apply NVM as main memory in embedded Chip Multi-Processor (CMP) systems. However, NVMs have drawbacks that need to be overcome. That is, a write to the NVMs is expensive. Loops are the most critical and time-consuming part in digital signal processing (DSP) applications. However, loops are difficult to parallelize on multi-processor systems due to the inter-iteration dependencies. This paper targets on embedded CMP systems and proposes techniques to improve loop parallelism while considering reducing the write activities to the NVMs when they are used as main memory. The experimental results show that the proposed algorithm can reduce the number of write activities on NVM by 21.1% on average. In other words, the average lifetime of NVM can be extended to at least 2 times longer than before and the total schedule length is reduced by 19.6% on average.
Yan Wang 0022, Jiayi Du, Jingtong Hu, Qingfeng Zhuge, Edwin H.-M. Sha
ICASSP1