Zhiyan Liu

dblp:90/2829 · DBLP profile ↗
← Back
17ranked-venue papers
9as first author
11since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 6 · 6 first-author · 6 since 2021Systems, architecture and hardware · 4 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 2 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-authorArtificial intelligence and machine learning · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Semantic-Relevance-Based Sensor Selection for Edge-AI Empowered Sensing Systems
abstract
Thesixth-generation(6G) mobile network is envisioned to incorporate sensing and edgeartificial intelligence(AI) as two key functions. Their natural convergence leads to the emergence ofIntegrated Sensing and Edge AI(ISEA), a novel paradigm enabling real-time acquisition and understanding of sensory information at the network edge. However, ISEA faces a communication bottleneck due to the large number of sensors and the high dimensionality of sensory features. Traditional approaches to communication-efficient ISEA lack awareness ofsemantic relevance, i.e., the level of relevance between sensor observations and the downstream task. To fill this gap, this paper presents a novel framework for semantic-relevance-aware sensor selection to achieve optimalend-to-end(E2E) task performance under heterogeneous sensor relevance and channel states. E2E sensing accuracy analysis is provided to characterize the sensing task performance in terms of selected sensors’ relevance scores and channel states. The analysis reveals that the contribution of each selected sensor to the accuracy is quantized by its expected classification margin, which is an increasing function of its relevance score. Building on the results, the sensor-selection problem for accuracy maximization is formulated as a 0-1 fractional programming problem, which is in general NP-hard. To solve this challenging problem, we exploit a problem-specific property of the objective function to develop its tight approximation. This allows to transform the original problem into a series of solvable sub-problems. The optimal solution of each sub-problem is proved to have a priority-based structure, wherein sensors are ranked according to a priority indicator that combines relevance scores and channel states, and the top-ranked sensors are selected. Based on the derived sensor priority, low-complexity algorithms are developed to determine the optimal numbers of selected sensors and features. Experimental results on both synthetic and real datasets show substantial accuracy gain achieved by the proposed selection scheme compared to existing benchmarks.
Zhiyan Liu, Kaibin Huang
IEEE Trans. Wirel. Commun.1
2025 Optimize Winograd Convolution for a Novel MIMD Many-core Architecture PEZY-SC3s
abstract
Optimizing convolution operations is critical for enhancing the performance of convolutional neural networks (CNNs). The Winograd convolution algorithm, renowned for its significant reduction in computational complexity, has been widely adopted in convolution acceleration. While the Winograd convolution algorithm has been highly optimized for traditional computing platforms such as SIMD-based CPUs and SIMT-based GPUs, research on alternative architectures better suited for convolution operations remains ongoing. MIMD architectures, characterized by their high parallelism and thread divergence mitigation capabilities, present promising potential; nevertheless, their applicability and performance in convolution operations remain underexplored. Additionally, GPU-based convolution acceleration, despite delivering exceptional performance, faces escalating energy consumption challenges, necessitating a balanced optimization of hardware performance and energy efficiency. To address these, we implement and optimize the Winograd convolution algorithm on a low-power MIMD many-core processor PEZY-SC3s, aiming to investigate its viability for convolution workloads. Experimental evaluations using convolutional layer parameters from widely adopted CNNs demonstrate: 1) 78.37% average bandwidth utilization in data-intensive stages; 2) 98% single-core and 92.52% system-wide computational efficiency in compute-intensive phases. The optimized Winograd algorithm achieves significantly superior computational efficiency compared to cuDNN-based implementation on Nvidia A100 GPU and oneDNN-based implementation on Intel Xeon Silver 4314 CPU, with energy efficiency ratios (GFLOPS/W) of $2.58 \times$ and $27.52 \times$, respectively.
Zhiyan Liu, Bingwei Wang, Feiming Liu, Xiangdong Pei
PACT4
2025 WideGate: Beyond Directed Acyclic Graph Learning in Subcircuit Boundary Prediction
abstract
Subcircuit boundary prediction is an important application of machine learning in logical analysis, effectively supporting tasks such as functional verification and logic optimization. Existing methods often convert circuits into and-inverter graphs and then use directed acyclic graph neural networks to perform this task. However, two key characteristics of subcircuit boundary prediction do not align with the fundamental assumptions of directed acyclic graph (DAG) learning, which limits the model's expressiveness and generalization capabilities. To break these assumptions, we propose WideGate, which includes a receptive field generation module that extends beyond the fanin cone and fanout cone, as well as an adaptive aggregation module that focuses on boundaries. Extensive experiments show that WideGate significantly outperforms existing methods in terms of prediction accuracy and training efficiency for sub circuit boundary prediction. The code is available at https://github.com/BUPT-GAMMA/WideGate.
Jiawei Liu 0006, Zhiyan Liu, Jianwang Zhai, Zhengyuan Shi, Qiang Xu 0001, Bei Yu 0001, Chuan Shi 0001
DATE2
2025 Optimizing Incomplete Cholesky Factorization on MIMD Many-core Architecture
abstract
Incomplete Cholesky (IC) factorization is widely used to precondition large sparse positive definite symmetric linear equations, since it can effectively reduce the number of iterations and enhance the solution efficiency compared to the unconditioned Conjugate Gradient (CG) iterative algorithm. In this paper, we have carried out optimization research on the IC factorization parallel algorithms based on the MIMD many-core architecture PEZY-SC3s processor. We propose a parallel algorithm for IC factorization under the MIMD architecture. This algorithm is based on the level-scheduling and graph coloring reordering algorithm. We optimize algorithm performance using various measures, including memory access optimization based on vector units and on-chip local memory, load balancing optimization based on dynamic thread scheduling, and task allocation optimization based on graph partitioning. Experimental results demonstrate that our IC parallel factorization achieves average speedups of 17.5x and 8.4x over the two cuSPARSE implementations on an NVIDIA A30 GPU, and 230.6x and 83.0x over Ginkgo’s serial and OpenMP implementations on a dual-socket Intel Xeon 4314 platform. Our study demonstrates that the PEZY architecture with effective algorithmic design still achieves good performance gains while maintaining low power consumption.
Yongzhen Shi, Jie Liu 0002, Zhiyan Liu, Bingwei Wang, Feiming Liu, Xiangdong Pei
ICPP5
2025 Can Subdaily LST Be Constructed in High-Latitude Regions Using Polar Orbiting Satellites?
abstract
This study introduces a new method for constructing subdaily land surface temperatures (LSTs) in northern high-latitude areas. The method is based on 1) an integration of observations from multiple satellites-sensors, including Moderate Resolution Imaging Spectroradiometer (MODIS) and Visible Infrared Imaging Radiometer Suite (VIIRS); and 2) the use of swath data to more comprehensively utilize the information content available from overlapping events from each platform within a day. The initial validation phase demonstrated a comparably high performance for LSTs among different products with a root mean square error (RMSE) at approximately$3~^{\circ }$C. Subsequently, the results confirmed that more than 23 observations per day at 70°N could be achieved, almost quadrupling the observation frequency achieved with twice-daily LST products combining MODIS and VIIRS. At one of the sample sites, the subdaily LST was clearly identified as a near-hourly temporal variation, and the gaps in the early morning and late afternoon were effectively filled.
Zhiyan Liu, Kazuhito Ichii, Yuhei Yamamoto, Masahito Ueyama, Hideki Kobayashi, Tetsuya Hiyama, Ayumi Kotani, Trofim Maximov, Ryan C. Sullivan, Sébastien C. Biraud
IEEE Geosci. Remote. Sens. Lett.1
2025 Over-the-Air Fusion of Sparse Spatial Features for Integrated Sensing and Edge AI Over Broadband Channels
abstract
The sixth-generation (6G) mobile networks feature two new usage scenarios – distributed sensing and edge artificial intelligence (AI). Their natural integration, termed integrated sensing and edge AI (ISEA), promises to create a platform that enables intelligent environment perception for wide-ranging applications. A basic operation in ISEA is for a fusion center to acquire and fuse features of spatial sensing data distributed at many edge devices (known as agents), which is confronted by a communication bottleneck due to multiple access over hostile wireless channels. To address this issue, we propose a novel framework, called Spatial Over-the-Air Fusion (Spatial AirFusion), which exploits radio waveform superposition to aggregate spatially sparse features over the air and thereby enables simultaneous access. The framework supports simultaneous aggregation over multiple voxels, which partition the 3D sensing region, and across multiple subcarriers. It exploits both spatial feature sparsity with channel diversity to pair voxel-level aggregation tasks and subcarriers to maximize the minimum receive signal-to-noise ratio among voxels. Optimally solving the resultant mixed-integer problem of Voxel-Carrier Pairing and Power Allocation (VoCa-PPA) is a focus of this work. The proposed approach hinges on derivations of optimal power allocation as a closed-form function of voxel-carrier pairing and a useful property of VoCa-PPA that allows dramatic solution space reduction. Both a low-complexity greedy algorithm and an optimal tree-search algorithm are then designed for VoCa-PPA. The latter is accelerated with a customised compact search tree, node pruning and agent ordering. Extensive simulations using real datasets demonstrate that Spatial AirFusion significantly reduces computation errors and improves sensing accuracy compared with conventional over-the-air computation without awareness of spatial sparsity.
Zhiyan Liu, Qiao Lan, Kaibin Huang
IEEE Trans. Wirel. Commun.1
2024 Advancing Free-Breathing Cardiac Cine MRI: Retrospective Respiratory Motion Correction Via Kspace-and-Image Guided Diffusion Model
Hongming Guo, Ziqing Huang, Hanbo Song, Zhiyan Liu, Xianzhao Feng, Ruixi Zhou
ICANN (8)5
2024 Over-the-Air Multi-View Pooling for Distributed Sensing
abstract
Sensing is envisioned as a key network function of thesixth-generation(6G) mobile networks.Artificial intelligence(AI)-empowered sensing fuses features of multiple sensing views from devices distributed in edge networks for the edge server to perform accurate inference. This process, known asmulti-view pooling, creates a communication bottleneck due to multi-access by many devices. To alleviate this issue, we propose a task-oriented simultaneous access scheme for distributed sensing calledOver-the-Air Pooling(AirPooling). The existingOver-the-Air Computing(AirComp) technique can be directly applied to enable Average-AirPooling, which exploits the waveform superposition property of a multi-access channel to implement fast over-the-air averaging of pooled features. However, despite being most popular in practice, the over-the-air maximization, called Max-AirPooling, is not AirComp realizable given the fact that AirComp addresses only a limited subset of functions. We tackle the challenge by proposing the novel generalized AirPooling framework that can be configured to support both Max- and Average-AirPooling by controlling a configuration parameter and extended to even other pooling functions. The former is realized by adding to AirComp the designed pre-processing at devices and post-processing at the server. To characterize theEnd-to-End(E2E) sensing performance in object recognition, the theory of classification margin is applied to relate the classification accuracy and the AirPooling error, which allows the latter to be a tractable surrogate of the former. Furthermore, the analysis reveals an inherent tradeoff of Max-AirPooling between the accuracy of the pooling-function approximation and the effectiveness of noise suppression. Using the tradeoff, we make an attempt to optimize the configuration parameter of Max-AirPooling, yielding a sub-optimal closed-form method of adaptive parametric control. Experimental results obtained on real-world datasets show that AirPooling provides sensing accuracies close to those achievable by the traditional digital air interface but dramatically reduces the communication latency, by up to an order of magnitude.
Zhiyan Liu, Qiao Lan, Anders E. Kalør, Petar Popovski, Kaibin Huang
IEEE Trans. Wirel. Commun.1
2023 Resource Allocation for Batched Multiuser Edge Inference with Early Exiting
abstract
This work considers multiuser edge inference for providing inference services to multiple users at the wireless edge. Multiple tasks are uploaded and grouped into a single batch for parallel processing at the edge server, while a task may exit early from the neural network without traversing the whole model. To efficiently grant users with heterogeneous requirements on accuracy and latency, we study in this paper the joint allocation of communication-and-computation (C2) resources. Two efficient algorithms are designed under the criterion of maximum throughput. First, consider the case with batching but without early exiting. The target problem is optimally solved using a proposed algorithm that nests a threshold-based scheme, which selects users with the best channels and meeting the computation-time constraints, in a sequential search for the maximum batch size. Next, consider the general case with batching and early exiting. A low-complexity sub-optimal algorithm for C2resource allocation is developed by modifying the preceding algorithm to exploit early exiting for latency reduction. Experimental results demonstrate that the proposed C2resource allocation algorithms can leverage batching and early exiting to achieve 1.95x throughput over conventional schemes.
Zhiyan Liu, Qiao Lan, Kaibin Huang
ICC1
2023 Resource Allocation for Multiuser Edge Inference With Batching and Early Exiting
abstract
The deployment of inference services at the network edge, called edge inference, offloads computation-intensive inference tasks from mobile devices to edge servers, thereby enhancing the former’s capabilities and battery lives. In a multiuser system, the joint allocation of communication-and-computation (C2) resources (i.e., scheduling and bandwidth allocation) is made challenging by adopting efficient inference techniques, batching and early exiting, and further complicated by the heterogeneity in users’ requirements on accuracy and latency. Batching groups multiple tasks into a single batch for parallel processing to reduce time-consuming memory access and thereby boosts the throughput (i.e., completed task per second). On the other hand, early exiting allows a task to exit from a deep-neural network without traversing the whole network, thereby supporting a tradeoff between accuracy and latency. In this work, we study optimal C2 resource allocation with batching and early exiting, which is an NP-complete integer programming problem. A set of efficient algorithms are designed under the criterion of maximum throughput by tackling the challenge. First, consider the case with batching but without early exiting. The target problem is solved optimally using a proposed best-shelf-packing algorithm that nests a threshold-based scheme, which selects users with the best channels and meeting the computation-time constraints, in a sequential search for the maximum batch size. Next, consider the general case with batching and early exiting. A low-complexity sub-optimal algorithm for C2 resource allocation is developed by modifying the preceding algorithm to exploit early exiting for latency reduction. On the other hand, the optimal approach is developed based on nesting a depth-first tree-search with intelligent online pruning into a sequential search for the maximum batch size. The key idea is to derive pruning criteria based on the simple greedy solution for the target problem without a bandwidth constraint and apply the result to designing an intelligent online pruning scheme. Experimental results demonstrate that both optimal and sub-optimal C2 resource allocation algorithms can leverage integrated batching and early exiting to double the inference throughput compared with conventional schemes.
Zhiyan Liu, Qiao Lan, Kaibin Huang
IEEE J. Sel. Areas Commun.1
2022 Deep Unsupervised Learning for Joint Antenna Selection and Hybrid Beamforming
abstract
In this paper, we propose a novel deep unsupervised learning-based approach that jointly optimizes antenna selection and hybrid beamforming to improve the hardware and spectral efficiencies of massive multiple-input-multiple-output (MIMO) downlink systems. By employing ResNet to extract features from the channel matrices, two neural networks, i.e., the antenna selection network (ASNet) and the hybrid beamforming network (BFNet), are respectively proposed for dynamic antenna selection and hybrid beamformer design. Furthermore, a deep probabilistic subsampling trick and a specially designed quantization function are respectively developed for ASNet and BFNet to preserve the differentiability while embedding discrete constraints into the network structures. With the aid of a flexibly designed loss function, ASNet and BFNet are jointly trained in a phased unsupervised way, which avoids the prohibitive computational cost of acquiring training labels in supervised learning. Simulation results demonstrate the advantage of the proposed approach over conventional optimization-based algorithms in terms of both the achieved rate and the computational complexity.
Zhiyan Liu, Yuwen Yang, Feifei Gao 0001, Hongbing Ma
IEEE Trans. Commun.1
2019 Updated Data-Driven GPP and NEE Estimation with Remote Sensing and Machine Learning Across Asia
abstract
Data-driven approach is effective for upscaling observation network data of terrestrial carbon fluxes. In this study, we estimated terrestrial gross primary productivity (GPP) and net ecosystem exchange (NEE) across Asia by re-assessment of input parameters and updated MODIS datasets (collection 6, C6) with a machine learning approach. Site-level experiments showed that the newly introduced lagged effect successfully improved estimation of NEE. Interannual variations in GPP and NEE across Asia is more consistent with independent model-based estimations compared with the previous data-driven estimation based on MODIS collection 5 data (C5). Our new estimation provides a good benchmark for understanding spatio-temporal variability in terrestrial GPP and NEE.
Zhiyan Liu, Kazuhito Ichii, Yusuke Hayashi, Riku Kawase, Kodai Hayashi, Masahito Ueyama, Yuji Kominami, Kireet Kumar, Sandipan Mukherjee
IGARSS1
2016 DisSetSim: An online system for calculating similarity between disease sets
abstract
Functional similarity between molecules results in similar phenotypes, such as diseases. Therefore, it is an effective way to reveal the function of molecules based on their induced diseases. However, the lack of a tool for obtaining the similarity score of pair-wise disease sets (SSDS) limits this type of application. Here, we introduce DisSetSim, an online system to solve this problem in this article. Five state-of-the-art methods involving Resnik's, Lin's, Wang's, PSB, and SemFunSim methods were implemented to measure the similarity score of pair-wise diseases (SSD) first. And then “pair-wise-best pairs-average” (PWBPA) method was implemented to calculated the SSDS by the SSD. The system was applied for calculating the functional similarity of miRNAs based on their induced disease sets. The results were further used to predict potential disease-miRNA relationships. The high area under the receiver operating characteristic curve AUC (0.9296) based on leave-one-out cross validation shows that the PWBPA method achieves a high true positive rate and a low false positive rate. The system can be accessed from http://bio-annotation.cn/DisSetSim.
Yang Hu 0008, Lingling Zhao, Zhiyan Liu, Hong Ju, Peigang Xu, Yadong Wang 0001, Liang Cheng 0006
BIBM3
2016 A novel method to identify pre-microRNA in various species knowledge base
abstract
More than 1/3 of human genes are regulated by microRNAs. The identification of microRNA (miRNA) is the precondition of discovering the regulatory mechanism of miRNA and developing the cure for genetic diseases. The traditional identification method is biological experiment, but it has the defects of long period, high cost, and missing the miRNAs that only exist in a specific period or low expression level. Therefore, to overcome these defects, machine learning method is applied to identify miRNAs. In this study, for identifying real and pseudo miRNAs and classifying different species, we extracted 98 dimensional features based on the primary and secondary structure, then we proposed the BP-Adaboost method to figure out the overfitting phenomenon of BP neural network by constructing multiple BP neural network classifiers and distributed weights to these classifiers. The novel method we proposed raised the accuracy and the stability. In this study, we verified the effectiveness and superiority over other methods by experiments.
Tianyi Zhao 0001, Ningyi Zhang, Peigang Xu, Zhiyan Liu, Liang Cheng 0006, Yang Hu 0008
BIBM5
2002 Improving progressive view-dependent isosurface propagation
Zhiyan Liu, Adam Finkelstein, Kai Li 0001
Comput. Graph.1
2001 Software Environments For Cluster-Based Display Systems
abstract
An inexpensive way to construct a scalable display wall system is to use a cluster of PCs with commodity graphics accelerators to drive an array of projectors. A challenge is to bring off-the-shelf sequential applications to run on such a display wall efficiently without using expensive, high-performance interconnects. We study two execution models for a scalable display wall system: master-slave and synchronized execution models. We have designed and implemented four software tools, two for each execution model, including VDD (Virtual Display Driver), GLP (GL-DLL Replacement), SSE (System-level Synchronized Execution), and ASE (Application-level Synchronized Execution). In order to support the synchronized execution model, we have also designed a broadcast, speculative file cache to provide scalable I/O performance. We report our experimental results with several 3D applications on the display wall to understand the performance implications and tradeoffs of these methods.
Douglas W. Clark, Zhiyan Liu, Grant Wallace, Kai Li 0001, Yuqun Chen
CCGRID3
2001 Data distribution strategies for high-resolution displays
Yuqun Chen, Adam Finkelstein, Thomas A. Funkhouser, Kai Li 0001, Zhiyan Liu, Rudrajit Samanta, Grant Wallace
Comput. Graph.6