EDBT 2026 Demo / reviewers in the wild / expert
Zhong Liu 0003
dblp:30/2371-3
· DBLP profile ↗
11ranked-venue papers
2as first author
9since 2021 · last 2025
0000-0002-6425-7791ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 10 · 2 first-author · 8 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | MetaCAN: Improving Generalizability of Few-shot Anomaly Detection with Meta-learningabstractFew-shot Anomaly Detection (AD) for images aims to detect anomalies with few-shot normal samples from the target dataset. It is a crucial task when only few samples can be obtained, and it is challenging since it needs to be generalized to different domains. Existing methods try to enhance the generalizability of AD by incorporating large vision-language models (LVLMs).However, how to transform category semantic information in LVLMs into anomaly information to improve the generalizability of AD remains a challenge facing existing methods.To address the challenge, we propose a few-shot AD method called MetaCAN, a novel category-to-anomaly network trained with AD meta-learning scheme based on an LVLM. Specifically, MetaCAN constructs the auxiliary training data and multiple tasks based on different categories to perform AD meta-learning, which ensures that the optimization toward the achievement of optimal anomaly detection across all categories. Moreover, MetaCAN introduces an image-image anomaly discriminator and an image-text anomaly detector to fully exploit the powerful multimodal semantic representations during auxiliary training. Once trained on auxiliary datasets, MetaCAN can be applied directly to other target datasets without retraining. Extensive experiments on six real-world datasets demonstrate that MetaCAN achieves state-of-the-art performance on cross-domain and cross-category anomaly detection tasks compared with existing methods. Zhisheng Lv, Songlei Jian, Chenlin Huang, Guansong Pang, Zhong Liu 0003 |
CIKM | 7 |
| 2024 | VCNN: A compiler of CNNs based on MLIR for multi-core vector acceleratorsabstractConvolutional Neural Network (CNN) is one of the representative algorithms of machine learning and deep learning. Multi-core vector accelerators are becoming increasingly popular due to their high performance and low power consumption. In the previous methods of deploying CNNs to vector accelerators, three issues are of concern. (1) Multi-core tasks are divided according to the image data dimension, which has limitations in low-latency real-time detection application scenarios. (2) Many memory optimization strategies only consider the impact of tensor size or tensor life, ignoring the inherent computational characteristics of operator types. (3) Manual mapping methods require a lot of engineering effort and are prone to introducing errors. Therefore, the VCNN proposed in this paper is a compiler based on MLIR for vector accelerators. The VCNN can automatically map CNN models to vector accelerators. It includes (1) a Multi-core Parallel Convolution algorithm (MPC) that is more suitable for low-latency real-time detection application scenarios. (2) An Adaptive Memory Reuse method (AMR). It not only considers the size or lifetime of tensors but also the inherent computational characteristics of operator types. (3) A quantization pass that quantizes the model and reshuffles the data. Experimental results show that the VCNN can effectively utilize the parallel processing capabilities of hardware when performing convolutions of different sizes, and achieve a parallel computing efficiency of up to 95.75%. In addition, we evaluated the performance of AlexNet, VGG16, and Yolov5s models. The results show that VCNN has a computational efficiency of more than 2X+ that of TVM, TensorRT, and OnnxRuntime, and a power efficiency of more than 7X+ that of them. Additionally, the VCNN inference model using Float16 is twice as energy efficient as that using Float32. Furthermore, the VCNN achieves a higher memory reuse rate compared to traditional memory reuse methods. Xiaorong Chen, Zhong Liu 0003 |
HPCC | 3 |
| 2023 | An Adaptive Instruction Set Encoding Automatic Generation Method for VLIW
Xin Xiao 0008, Zhong Liu 0003 |
ICA3PP (1) | 2 |
| 2023 | ISADL: An Instruction Set Architecture Description Language for VLIWabstractThis paper presents a novel architecture description language for VLIW, which called Instruction Set Architecture Description Language (ISADL). The design of VLIW instruction sets is an iterative process, the addition and deletion of instructions leads to issues such as encoding inefficiency, format adjustment, and decoding complexity. Previous architecture description languages have overlooked the stage of instruction format design, resulting in the need for a large amount of hard coding to support automatic tool chain generation. We propose two algorithms for automatic generation of instruction formats and variable length encoding for instruction format design. Our method solves a large number of hard coding problems by automatically generating tool chain models, while also optimizing the convenience and assembly speed of the automatically generated assembler. This paper delineates this technology and evaluates it on the MT-3000, MT-7004 and XT-4000 ISA. The experimental results demonstrate that it can automatically and quickly generate encoding schemes and optimize the issues in instruction format design. This can effectively accelerate the progress of the instruction set design. Xin Xiao 0008, Zhong Liu 0003 |
ICPADS | 2 |
| 2022 | Long-life Sensitive Modulo Scheduling with Adaptive Loop ExpansionabstractThis paper presents a novel modulo scheduling method, which is called Expanded Iterative Modulo Scheduling (EIMS). EIMS integrates an adaptive loop expansion mechanism, and the overhead of analyzing data dependencies depends only on the initial loop, which means that it expands the search space with lower overhead compared to other methods that separate loop unrolling and scheduling. EIMS focuses on the criticality of operations and makes those interdependent operations as close as possible, thereby reducing register requirements. Notably, Long-lived Transfer Mechanism (LTM) is proposed to address scheduling failures caused by long-lived variables, which has not been mentioned in previous papers. The paper describes the technique and evaluates it on the MT-3000, achieving over 18x performance improvement on 7 classical assemblies and better resource utilization against other methods. Hongli Zhong, Zhong Liu 0003 |
ICPADS | 2 |
| 2022 | SADD: A Novel Systolic Array Accelerator with Dynamic Dataflow for Sparse GEMM in Deep Learning
Sheng Ma, Zhong Liu 0003, Libo Huang 0002, Yuan Yuan 0034 |
NPC | 3 |
| 2022 | Adaptive Low-Cost Loop Expansion for Modulo Scheduling
Hongli Zhong, Zhong Liu 0003, Sheng Liu 0001, Sheng Ma, Chen Li 0015 |
NPC | 2 |
| 2022 | MT-3000: a heterogeneous multi-zone processor for HPC
Kai Lu 0001, Yang Guo 0003, Chun Huang 0006, Sheng Liu 0001, Ruibo Wang, Jianbin Fang, Tao Tang 0001, Zhaoyun Chen, Biwei Liu, Zhong Liu 0003, Yuanwu Lei, Haiyan Sun |
CCF Trans. High Perform. Comput. | 11 |
| 2022 | Optimizing convolutional neural networks on multi-core vector accelerator
Zhong Liu 0003, Xin Xiao 0008, Chen Li 0015, Sheng Ma, Rangyu Deng |
Parallel Comput. | 1 |
| 2020 | Accelerating Large-Scale Deep Convolutional Neural Networks on Multi-core Vector Accelerators
Zhong Liu 0003, Sheng Ma |
NPC | 1 |
| 2019 | Coordinated DMA: Improving the DRAM Access Efficiency for Matrix MultiplicationabstractHigh performance implementation of matrix multiplication is essential for scientific computing. The memory access procedure is quite possible to be the bottleneck of matrix multiplication. The widely used GotoBLAS GEMM implementation divides the integral matrix into several partitions to be assigned to different cores for parallelization. Traditionally, each core deploys a DMA transfer to access its own partition in the DRAM memory. However, deploying an independent DMA transfer for each core cannot efficiently exploit the inter-core locality. Also, multiple concurrent DMA transfers interfere with each other, further reducing the DRAM access efficiency. We observe that the same row of neighboring partitions is in the same DRAM page, which means that there is significant locality inherent in the address layout. We propose the coordinated DMA to efficiently exploit the locality. It invokes one transfer to serve all cores and moves data in a row-major manner to improve the DRAM access efficiency. Compared with a baseline design, the coordinated DMA improves the bandwidth by 84.8 percent and reduces DRAM energy consumption by 43.1 percent for micro-benchmarks. It achieves higher performance for the GEMM and Linpack benchmark. With much less hardware costs, the coordinated DMA significantly outperforms an out-of-order memory controller. Sheng Ma, Zhong Liu 0003, Shenggang Chen, Libo Huang 0002, Yang Guo 0003, Zhiying Wang 0003, Meidi Zhang |
IEEE Trans. Parallel Distributed Syst. | 2 |