EDBT 2026 Demo / reviewers in the wild / expert
Liansheng Liu
dblp:46/595
· DBLP profile ↗
19ranked-venue papers
0as first author
14since 2021 · last 2026
0000-0001-5834-9772ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 8 · 8 since 2021Artificial intelligence and machine learning · 4 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Computer networks · 1 · 1 since 2021Security and privacy · 1Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | State estimation for low earth orbit satellite constellations using dynamic graph neural networks
Pengming Wang 0002, Liansheng Liu, Yinghao Guan, Datong Liu |
Expert Syst. Appl. | 2 |
| 2026 | Resource-Optimized Time-Multiplexed Constant Multiplication via Adjacency Matrix ModelingabstractThis article presents TmCM-AM, a new time-multiplexed constant multiplication framework based on adjacency matrix modeling. The TmCM-AM framework provides a universal and efficient approach for digital signal processing applications using much fewer resources and with greater adaptability based on conventional methods. By transforming adder graphs into adjacency matrices and using an optimization algorithm, the proposed framework minimizes the number of required adders and multiplexers to a large degree. In particular, three mathematical properties of adjacency matrices based on properties of adder graphs are presented. Meanwhile, the adjacency matrix is employed to model time-multiplexed adder graphs in detail, making hardware architecture analysis possible through matrix computation. Finally, heuristic algorithms are used to generate the best possible solution from matrices calculated. Experimental verification through FPGA and ASIC implementations further confirms the feasibility of TmCM-AM, presenting enormous reductions in area and power dissipation, as well as delay metrics across random data and various real-life coefficient sets. Martin Kumm, Liansheng Liu, Zhixian Zhang, Yu Peng 0002 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2026 | Scaling-Free CORDIC Optimization Based on the Time-Multiplexed Constant MultiplicationabstractModern digital systems in communication, control, and sensing rely on fast and accurate hardware evaluation of trigonometric functions. The classic CORDIC algorithm provides a multiplier-less solution using iterative shifts and adds, but conventional CORDIC requires many iterations and a final scaling correction that increases latency and hardware cost. Scaling-Free CORDIC (SF-CORDIC) eliminates the scaling factor by adjusting micro-rotation angles to achieve a net gain of unity. However, this approach introduces fixed constant multiplications at each iteration, which can dominate hardware resources and critical path delay. Existing methods often restrict the set of rotation angles so that constants can be implemented with simple add-shift operations. Other approaches use high-radix or hybrid schemes, but these methods limit the angle accuracy, introduce extra complexity, and still rely on numerous add-shift networks. To address these issues, a new SF-CORDIC architecture is proposed to eliminate per-iteration multipliers by using the time-multiplexed constant multiplier units implemented across all iterations. Bit-width growth is controlled via shift propagation and a minor compensation step, maintaining precision without a final scaling stage. This compact, low-latency design significantly reduces hardware area and delay compared to prior SF-CORDIC implementations while preserving high output accuracy. In a six-stage FPGA pipeline, the design reaches a maximum clock frequency of 216.03 MHz with an end-to-end latency of 27.78 ns. It uses about 1200 LUTs and 300 registers. The root-mean-square error is 6.8 × 10−5for sine and 6.4 × 10−5for cosine. Yu Peng 0002, Xuejing Wei, Liansheng Liu |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2026 | MultiSky: Dynamic Resource Allocation Framework for High-Throughput CGRA Multitask ExecutionabstractCoarse-grained reconfigurable arrays (CGRAs) offer a promising balance between high performance and flexibility, yet dynamic resource allocation in multi-task scenarios remains challenging due to unpredictable task creation/destruction. Existing static approaches lack flexibility, while dynamic methods suffer from high latency or limited applicability. This paper presents MultiSky, a framework for CGRA multi-task dynamic resource allocation, combining a hardware controller and a software pre-mapper. The hardware controller dynamically allocates resources within hundreds of cycles by calculating tile allocation for each task via weighted averaging, and generating tile shapes using a lightweight heuristic algorithm. The software pre-mapper employs incremental compilation to pre-generate configurations, avoiding online transformation overhead. Evaluations on a real-world multi-task scenario demonstrate that MultiSky achieves 1.72× higher throughput than baselines by maintaining 82.7% average resource utilization. The framework scales efficiently with larger CGRAs and task counts, with hardware overhead decreasing to 1% for 16×16 CGRAs. These results highlight MultiSky’s ability to balance flexibility, efficiency, and practicality in dynamic computing environments. Chenhao Xie 0001, Rui Wang 0014, Liansheng Liu, Xiyuan Peng, Yu Peng 0002 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2025 | FexMo: Enabling Fuse Execution Mode for Multi-task CGRAs
Chenhao Xie 0001, Chuliang Guo, Liansheng Liu, Xiyuan Peng, Datong Liu, Yu Peng 0002 |
MICRO | 4 |
| 2025 | Enhancing explainability in medical image classification and analyzing osteonecrosis X-ray images using shadow learner system
Yaoyang Wu, Simon Fong 0001, Liansheng Liu |
Appl. Intell. | 3 |
| 2025 | Shadow learner system: implementation of CNN with explainable AI model for bone radiology image classification
Yaoyang Wu, Simon Fong 0001, Liansheng Liu |
Soft Comput. | 3 |
| 2025 | DynMap: A Heuristic Dynamic Mapper for CGRA Multitask Dynamic Resource AllocationabstractCoarse-grained reconfigurable architecture (CGRA) has received increasing attention in both industry and academia due to its comprehensive advantages of performance, energy efficiency, and flexibility. To improve the resource utilization and handle the mixing workloads in the real-world, multiple tasks sharing the whole CGRA has became an important technical trend, and the varying resource requirements throughout their life cycles also makes run-time dynamic resource allocation (DRA) necessary for higher-multitask throughput. As the key stage of DRA, dynamic mapping (DM) is responsible for mapping kernels within each task to the dynamically allocated CGRA resources. However, existing DM methods have difficulty to balance the mapping time and the mapping quality, resulting in a significant gap between the actual and the optimal task throughput. To address the challenge, we propose DynMap, a heuristic dynamic mapper for CGRA multitask DRA. With the support of specialized scheduling and routing schemes, DynMap heuristically references the placement tendency in the static mapping result to dramatically save the mapping time, while maintaining the high-mapping quality by minimizing the possibility of resource conflicts. Experimental evaluation demonstrates DynMap not only achieves the average 1.17 ms mapping time and average 98.33% of the optimal mapping quality on different CGRA architectures, but also reaches average 98.85% of the optimal task throughput expected by different CGRA multitask DRA scenarios, reducing the gap between actual and optimal task throughput average$31.75\times $smaller than that of the current methods. Chenhao Xie 0001, Liansheng Liu, Xiyuan Peng |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2025 | Eyelet: A Cross-Mesh NoC-Based Fine-Grained Sparse CNN Accelerator for Spatio-Temporal Parallel Computing OptimizationabstractFine-grained sparse convolutional neural networks (CNNs) achieve a better trade-off between model accuracy and size than coarse-grained sparse CNNs. Due to irregular data structures and unbalanced computation loads, fine-grained sparse CNNs struggle to fully leverage the performance advantages of computation and storage on general-purpose edge hardware. However, existing custom sparse accelerators are designed from the perspective of emulating a balanced load by software or computational strategies, neglecting the exploration of the computing architecture’s adaptability and parallelism for fine-grained sparse models. To address these challenges, a cross-mesh NoC-based accelerator architecture is proposed. This architecture aligns with the irregular characteristics of fine-grained sparse CNN weights and enhances the spatio-temporal parallelism of fine-grained sparse CNNs. First, a sparse multiplier unit (SMU) array and an adder array are designed to enable parallel execution of convolution multiplication and accumulation operations. Then, element-wise unroll-based nonzero weight multiplication is mapped to the SMU array to provide more flexible spatial parallelism. A horizontal and vertical cross-mesh NoC is proposed for flexible dataflow scheduling between the SMU and adder arrays to further improve temporal parallelism. This architecture allows the multiplication and accumulation operations in convolution to be decoupled and pipelined with negligible latency. Finally, the proposed accelerator architecture is implemented on the ZU9EG platform. The experimental results show that the proposed accelerator achieves frame rates of 509.9, 249.3, 100.7, 48.4, and 168.9 frames per second (FPS) for AlexNet, VGG-16, ResNet-18, MobileNet-v2, and EfficientNet, respectively. Compared with related works, this accelerator achieves inference speed and energy efficiency improvements of$1.1\times \sim 36.1\times $and$2.4\times \sim 13.4\times $, respectively. Liansheng Liu, Yu Peng 0002, Xiyuan Peng, Heming Liu |
IEEE Trans. Circuits Syst. I Regul. Pap. | 2 |
| 2025 | PreTrans: Enabling Efficient CGRA Multi-Task Context Switch Through Config Pre-Mapping and Data TransceivingabstractDynamic resource allocation guarantees the performance of CGRA multi-task, but incurs a wide range of incompatible contexts (config & data) to the CGRA architecture. However, traditional context switch approaches including online config transformation and data reloading may significantly block the task to process inputs under new resource allocation decisions, resulting in the limited task throughput. To address this issue, online config transformation can be avoided if compatible configs have been prepared through offline pre-mapping, but traditional CGRA mappers require days to achieve comprehensive pre-mapping with considerable quality. Besides, online data reloading can also be eliminated through memory sharing, but the traditional arbiter-based approach has the difficulty of trading off physical complexity and memory access parallelism. PreTrans is the first system design to achieve the efficient CGRA multi-task context switch. PreTrans first avoids the online config transformation through a software incremental pre-mapper, which re-utilizes the previously finished pre-mapping results to dramatically accelerate the pre-mapping of subsequent resource allocation decisions with negligible mapping quality loss. Secondly, PreTrans replaces the traditional arbiter with a hardware data transceiver to better support the memory sharing that eliminates data reloading, which allows each tile to possess an individual memory that maximizes the access parallelism without introducing significant physical overhead. The overall evaluation demonstrates that PreTrans achieves 1.13$\sim 2.46\times$throughput improvement on pipeline and parallel multi-task scenarios, and can reach the target throughput immediately after the new resource allocation decision takes effect. Ablation study further shows that the pre-mapper is more than 3 magnitudes faster than the traditional CGRA mapper while maintaining more than 99% of the optimal mapping quality, and the data transceiver only introduces 9.02% hardware area overhead under 16×16 CGRA. Chenhao Xie 0001, Liansheng Liu, Xiyuan Peng, Yu Peng 0002, Hailong Yang 0002, Depei Qian 0001 |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2025 | Resource Optimization in Polyphase-Filter STFT Based on Time-Multiplexed Constant Multiplication
Yu Peng 0002, Zhixian Zhang, Tongrui Zhang, Liansheng Liu |
IEEE Trans. Very Large Scale Integr. Syst. | 5 |
| 2024 | Enhancing Bone Abnormality Classification Through Background and Irrelevant Part Processing: A Deep Learning and Explainable AI ApproachabstractMedical image classification has beenan important application for Deep Learning techniques for over a decade, and since the emergence of Explainable AI (XAI), researchers have started using XAI to validate the results produced by these black box models. In the research field, it has become clear that accuracy and efficiency are not the only crucial factors for developing medical deep learning models; the authenticity of results and the accountability of the model and its creator also matter greatly. The novelty of this paper lies in its specific application of extended preprocessing techniques-namely, the removal of background and irrelevant parts-to medical images for improving the performance of deep learning models in classification tasks. While the concept of preprocessing images has been explored by many researchers, applying such targeted preprocessing steps to medical images, combined with the use of XAI to validate and illustrate the benefits, is a novel approach. This paper highlights the unique requirements of medical image data and proposes an innovative method to enhance model accuracy and reliability in medical diagnostics by removing background and redundant features from the images. Yaoyang Wu, Simon Fong 0001, Qun Song 0007, Liansheng Liu |
HealthCom | 5 |
| 2024 | A model-driven dual-derivation framework for quantitative fault detection in satellite power system
Pengming Wang 0002, Liansheng Liu, Zhidong Li, Datong Liu |
Adv. Eng. Informatics | 2 |
| 2024 | Efficient Radius Search for Adaptive Foveal Sizing Mechanism in Collaborative Foveated Rendering FrameworkabstractCollaborative Foveated Rendering (CFR) is the latest collaborative rendering framework proposed to enable high frame rate VR applications on mobile devices. Compared with the strategies adopted in conventional collaborative rendering, the pixel-based Adaptive Foveal Sizing (AFS) mechanism in CFR offers a more flexible and intelligent workload trade-off by predicting the radius. However, the performance of the AFS mechanism in actual deployment depends on its adaptability to two factors, including the Sudden Environmental Variations (SEV) and the Random Discrete Latency (RDL). Guaranteeing the performance of the AFS mechanism by adapting to these two factors is of great significance to guaranteeing users' immersive experience.This paper identifies the existence of the SEV and RDL phenomenon in the AFS mechanism for the first time, and contributes the first method that offers the effective and real-time AFS mechanism implementation for the practical deployment, namely the Efficient Radius Search (ERS).The ERS method efficiently searches the largest radius online that controls the rendering workload within the foveated layer just below the offline baked threshold, thereby achieving the immediate response to SEV and reducing the oscillating frame rendering latency led by RDL.Through the experiments on 3 VR applications and 4 mobile devices, the resulting 2.44× to 9.07× higher frame rate precision compared with the state-of-the-art method demonstrate the superiority of the ERS method. Chenhao Xie 0001, Liansheng Liu, Philip H. W. Leong, Shuaiwen Song |
IEEE Trans. Mob. Comput. | 3 |
| 2020 | Multi-view convolutional neural network with leader and long-tail particle swarm optimizer for enhancing heart disease and breast cancer detection
Kun Lan, Liansheng Liu, Tengyue Li, Simon Fong 0001, João Alexandre Lôbo Marques, Raymond K. Wong 0001, Rui Tang 0010 |
Neural Comput. Appl. | 2 |
| 2019 | Dual feature selection and rebalancing strategy using metaheuristic optimization algorithms in X-ray image datasets
Jinyan Li 0002, Simon Fong 0001, Liansheng Liu, Nilanjan Dey, Amira S. Ashour, Luminita Moraru |
Multim. Tools Appl. | 3 |
| 2019 | Cross-Domain Noise Impact Evaluation for Black Box Two-Level Control CPSabstractControl Cyber-Physical Systems (CPSs) constitute a major category of CPS. In control CPSs, in addition to the well-studied noises within the physical subsystem, we are interested in evaluating the impact of cross-domain noise : the noise that comes from the physical subsystem, propagates through the cyber subsystem, and goes back to the physical subsystem. Impact of cross-domain noise is hard to evaluate when the cyber subsystem is a black box, which cannot be explicitly modeled. To address this challenge, this article focuses on the two-level control CPS, a widely adopted control CPS architecture, and proposes an emulation based evaluation methodology framework. The framework uses hybrid model reachability to quantify the cross-domain noise impact, and exploits Lyapunov stability theories to reduce the evaluation benchmark size. We validated the effectiveness and efficiency of our proposed framework on a representative control CPS testbed. Particularly, 24.1% of evaluation effort is saved using the proposed benchmark shrinking technology. Liansheng Liu, Stefan Winter 0001, Qixin Wang 0001, Neeraj Suri, Lei Bu, Yu Peng 0002, Xue (Steve) Liu, Xiyuan Peng |
ACM Trans. Cyber Phys. Syst. | 2 |
| 2005 | Identity Verification System Using Data Hiding and Fingerprint RecognitionabstractThis paper proposes an identity verification system using data hiding and fingerprint recognition. At user's home, the client's account information is encrypted and embedded into the fingerprint image via data hiding method secretly. Then the fingerprint image with embedded data is transferred to the bank over Internet. At bank side, the client's account information is extracted. It is used to retrieve the client's registered fingerprint from central database, which is then matched with extracted fingerprint via fingerprint recognition method to verify user's identity. This system is more reliable and secure than transferring password alone. The data are embedded with quantization watermark in the JPEG 2000 coding pipeline. Compare to our previous proposed system, the interaction time can be reduced because less data will be transmitted. When the fingerprint image is compressed to 1/4~1/20 of its original size, the embedded watermark can still be recovered. This system has been used in a bank pension distribution system. It can also be used in other E-business applications Guorong Xuan, Hongfei Ji, Yun Q. Shi 0001, Dekun Zou, Liansheng Liu, Heisheng Liu, Weichao Bai |
MMSP | 6 |
| 2004 | A Secure Internet-Based Personal Identity Verification System Using Lossless Watermarking and Fingerprint Recognition
Guorong Xuan, Junxiang Zheng 0001, Chengyun Yang, Yun Q. Shi 0001, Dekun Zou, Liansheng Liu, Weichao Bai |
IWDW | 6 |