EDBT 2026 Demo / reviewers in the wild / expert
Hao Liang 0003
dblp:62/5181-3
· DBLP profile ↗
18ranked-venue papers
5as first author
4since 2021 · last 2026
0000-0002-8097-4707ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 17 · 5 first-author · 3 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MoSs: Mixture of Scales for Efficient High-Resolution Autoregressive Image GenerationabstractSince next-scale prediction was introduced as a new paradigm for autoregressive image generation, it has attracted extensive research interest. By progressively increasing resolution in a draft-to-refinement process, next-scale prediction demonstrates great potential in both generation quality and efficiency. However, at high resolutions, this paradigm faces a fundamental challenge: token sequences grow quadratically and accumulate across multiple scales, resulting in a key performance bottleneck. Our systematic study uncovers two critical observations: (1) most image regions have stabilized during early drafting stages, making later refinement across the full-scale image token-inefficient; (2) different scales inherently trade off efficiency and fidelity, suggesting that adaptive token dispatch on different scales can focus resources where they yield the greatest quality gains. Motivated by these insights, we propose a training-free Mixture of Scales (MoSs) method for efficient high-resolution autoregressive image generation. MoSs breaks the strict causal dependency across scales in the final refinement steps by parallelizing scales of different resolutions, each responsible for a subset of spatial regions. A lightweight frequency-based token dispatcher analyzes the drafted image and assigns regions to the appropriate scale. The outputs are then composited over the draft to produce the final high-resolution image. The scale-mixture method exhibits remarkable efficiency with little impact on generation quality on various models. For instance, our implementation achieves 2.05-4.96x speedup on the transformer backbone, up to 85.62% KV cache reduction, incurring only 0.1-2.4% loss on GenEval quality, based on the state-of-the-art Infinity model. Yaoxiu Lian, Hao Liang 0003, Zhihong Gou, Guohao Dai 0001, Ningyi Xu |
AAAI | 2 |
| 2026 | ROMA: A Read-Only-Memory-based Accelerator for QLoRA-based On-Device LLM
Guanting Huo, Hao Liang 0003, Shijie Cao, Ningyi Xu |
ASP-DAC | 5 |
| 2023 | RECom: A Compiler Approach to Accelerating Recommendation Model Inference with Massive Embedding ColumnsabstractEmbedding columns are important for deep recommendation models to achieve high accuracy, but they can be very time-consuming during inference. Machine learning (ML) compilers are used broadly in real businesses to optimize ML models automatically. Unfortunately, no existing work uses compilers to automatically accelerate the heavy embedding column computations during recommendation model inferences. To fill this gap, we propose RECom, the first ML compiler that aims at optimizing the massive embedding columns in recommendation models on the GPU. RECom addresses three major challenges. First, generating an efficient schedule on the GPU for the massive operators within embedding columns is difficult. Existing solutions usually lead to numerous small kernels and also lack inter-subgraph parallelism. We adopt a novel codegen strategy that fuses massive embedding columns into a single kernel and maps each column into a separate thread block on the GPU. Second, the complex shape computations under dynamic shape scenarios impede further graph optimizations. We develop a symbolic expression-based module to reconstruct all shape computations. Third, ML frameworks inevitably introduce redundant computations due to robustness considerations. We develop a subgraph optimization module that performs graph-level simplifications based on the entire embedding column context. Experiments on both in-house and open-source models show that RECom can achieve 6.61X and 1.91X over state-of-the-art baselines in terms of end-to-end inference latency and throughput, respectively. RECom's source code is publicly available at https://github.com/AlibabaResearch/recom. Zaifeng Pan, Zhen Zheng, Feng Zhang 0007, Hao Liang 0003, Dalin Wang, Xiafei Qiu, Wei Lin 0016, Xiaoyong Du 0001 |
ASPLOS (4) | 5 |
| 2021 | Graph Sampling with Fast Random Walker on HBM-enabled FPGA AcceleratorsabstractGraph neural networks (GNNs) have gained increasing popularity among researchers recently and have been employed in many applications. Training GNNs introduce a crucial stage called graph sampling. One of the most important sampling algorithms is Random Walk. However, Random Walk and many of its variants share and suffer from the same performance problem caused by random and fragmented memory access patterns, leading to significant system performance degradation. In this work, we present an efficient graph sampling engine on modern FPGAs integrated with in-package high bandwidth memory (HBM), which brings data closer and faster to the core logic. The hardware walker design is modular and easily scalable for massive parallelism, to fully utilize the available HBM channels. Our design also provides the flexibility to support Random Walk and two of its variants on both homogeneous and heterogeneous graphs. On real-world graph datasets, we achieve a 1.39 × -3.74 × speedup with a 2.42 × -6.69 × higher energy efficiency over highly optimized parallel baselines on a Xeon CPU. We also implement these algorithms on a NVIDIA Tesla VIOO GPU and achieve comparable dynamic energy consumption. Chunyou Su, Hao Liang 0003, Wei Zhang 0012, Baole Ai, Wenting Shen, Zeke Wang |
FPL | 2 |
| 2020 | Optimizing OpenCL-Based CNN Design on FPGA with Comprehensive Design Space Exploration and Collaborative Performance ModelingabstractRecent success in applying convolutional neural networks (CNNs) to object detection and classification has sparked great interest in accelerating CNNs using hardware-like field-programmable gate arrays (FPGAs). However, finding an efficient FPGA design for a given CNN model and FPGA board is not trivial since a strong background in hardware design and detailed knowledge of the target board are required. In this work, we try to solve this problem by design space exploration with a collaborative framework. Our framework consists of three main parts: FPGA design generation, coarse-grained modeling, and fine-grained modeling. In the FPGA design generation, we propose a novel data structure, LoopTree, to capture the details of the FPGA design for CNN applications without writing down the source code. Different LoopTrees, which indicate different FPGA designs, are automatically generated in this process. A coarse-grained model will evaluate LoopTrees at the operation level, e.g., add, mult, and so on, so that the most efficient LoopTrees can be selected. A fine-grained model, which is based on the source code, will then refine the selected design in a cycle-accurate manner. A set of comprehensive OpenCL-based designs have been implemented on board to verify our framework. An average estimation error of 8.87% and 4.8% has been observed for our coarse-grained model and fine-grained model, respectively. This is much lower than the prevalent operation-statistics-based estimation, which is obtained according to a predefined formula for specific loop schedules. Jiandong Mu, Wei Zhang 0012, Hao Liang 0003, Sharad Sinha |
ACM Trans. Reconfigurable Technol. Syst. | 3 |
| 2019 | PAI-FCNN: FPGA Based Inference System for Complex CNN ModelsabstractConvolutional Neural Network (CNN) models are becoming complex with advanced OPs and structures, which introduces design challenges for FPGA-based system. In this paper, we present the design of an FPGA-based CNN inference system, PAI-FCNN, to support modern complex CNN models. PAI-FCNN consists of scalable hardware design and a model reconstruction flow in software compiler. In this way, advanced OPs like Deconv, Conv with upsampling, Dilated Conv, Concatenation can be processed by PAI-FCNN with high performance and hardware efficiency. PAI-FCNN also incorporates reduced precision to boost computing capacity, and the emerging CNN-RNN (Recurrent Neural Network) hybrid models are supported. Our experiments on both PC and embedded FPGA platforms show that the system consistently performs in an efficient manner. PAI-FCNN achieves better throughput and power efficiency than GPU solutions. Lixue Xia, Lansong Diao, Zhao Jiang, Hao Liang 0003, Kai Chen 0008, Shunli Dou, Zibin Su, Jiansong Zhang 0001, Wei Lin 0016 |
ASAP | 4 |
| 2019 | PAI-FCNN: FPGA Based CNN Inference SystemabstractWe describe the FPGA subsystem of the Platform of Artificial Intelligence (PAI) in Alibaba Group, called PAI-FCNN. PAI-FCNN plays the role of a heterogeneous back-end for CNN inference, together with other CPU, GPU and ASIC subsystems in PAI. Driven by various business needs, we built PAI-FCNN from scratch since two years ago. We present our experience from FPGA/compiler design and implementation, to system evaluation and deployment. In particular, in order to address three practical challenges: (1) Efficient processing for diverse operators and model structure such as Deconv, Dilated Conv, Up-sampling, PReLu and Concatenation. (2) Serving multiple highly-different models on single FPGA hardware. (3) Competitive performance with alternative GPU or ASIC solutions, we extensively perform joint software & hardware design to optimize system efficiency across multiple CNN models, which includes model reconstruction in compiler software and flexible data access in data-flow CNN processor. We also incorporate reduced precision and model retraining to boost system capacity. Using U-net as an example, on Xilinx KU115 chip, with the help of 74.9% efficiency on Int16-precision hardware (with 3.226TOPS capacity) and 72.9% efficiency on mixed-int8/int3-precision hardware (with 14.746TOPS capacity), we achieve slightly better throughput and 2X higher power efficiency than P4. Lansong Diao, Zhao Jiang, Hao Liang 0003, Chang'an Ye, Kai Chen 0008, Shunli Dou, Lixue Xia, Jiansong Zhang 0001, Wei Lin 0016 |
FPGA | 3 |
| 2019 | Ouroboros: An Inference Engine for Deep Learning Based TTS on Embedded DevicesabstractThis article consists of a collection of slides from the author's conference presentation. Jiansong Zhang 0001, Lixue Xia, Zhao Jiang, Hao Liang 0003, Shouda Liu, Wei Lin 0016, Yuan Xie 0001 |
Hot Chips Symposium | 4 |
| 2018 | A Collaborative Framework for FPGA-based CNN Design Modeling and OptimizationabstractConvolutional neural network (CNN) has presented a great success in numerous areas and has sparked an increasing interest in accelerating CNN using hardware like FPGAs. However, efficient FPGA design for CNN applications requires a long development time and a strong background in hardware details. Consequently, an easy-to-use yet powerful auto CNN design optimization framework is required. In this work, we propose a collaborative framework to model and optimize the OpenCL based FPGA design for CNN applications according to the device resource limitation and the CNN specification. Our framework mainly consists of LoopTree, a novel data structure we propose to capture the structure of OpenCL based CNN design; a LoopTree based coarse-grained model, which will estimate the performance of the CNN design at the module level; and a source code based fine-grained model, which will estimate the CNN design performance in a cycle-accurate manner. Efficient designs can be achieved by collaborating the two models in a search and refined manner. A variety of OpenCL based designs have been implemented on board to verify our framework. The results show that our coarse-grained model and fine-grained model have an average estimation error of 10.2% and 4.7% which are much lower than prevalent operation statistics based estimation calculated by the predefined formula for specific loop schedules. Jiandong Mu, Wei Zhang 0012, Hao Liang 0003, Sharad Sinha |
FPL | 3 |
| 2018 | Parallelizing Hardware Tasks on Multicontext FPGA With Efficient Placement and Scheduling AlgorithmsabstractField programmable gate arrays (FPGAs) are often used to accelerate multiple tasks simultaneously, working in a tightly coupled processor-coprocessor architecture. Recently, with the fast development of emerging memory technologies, multicontext FPGAs with high-density memories that support fast dynamic reconfiguration have become feasible. Compared with single-context FPGAs, multicontext FPGAs have a much higher on-chip configuration memory capacity but have not been thoroughly investigated to exploit their capabilities. In this paper, we investigate how to best utilize the capacity advantage of the multicontext FPGAs. We first propose a static placement strategy to place the requested hardware tasks with minimal area on the FPGA. We then optimize the running time of the static placement without sacrificing its solution quality. Along with the static placement, we propose collaborated online placement and scheduling strategies to manage the actual execution and reconfiguration of hardware tasks on a multicontext FPGA. Our experiments show that the static placement algorithm generates high quality placement solutions within a short time. Starting from the static placement solution, our collaborated online placer and scheduler schedules and places simultaneous acceleration tasks and reduces the acceleration task rejection rate significantly compared to a baseline design. Hao Liang 0003, Sharad Sinha, Wei Zhang 0012 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2017 | FP-DNN: An Automated Framework for Mapping Deep Neural Networks onto FPGAs with RTL-HLS Hybrid TemplatesabstractDNNs (Deep Neural Networks) have demonstrated great success in numerous applications such as image classification, speech recognition, video analysis, etc. However, DNNs are much more computation-intensive and memory-intensive than previous shallow models. Thus, it is challenging to deploy DNNs in both large-scale data centers and real-time embedded systems. Considering performance, flexibility, and energy efficiency, FPGA-based accelerator for DNNs is a promising solution. Unfortunately, conventional accelerator design flows make it difficult for FPGA developers to keep up with the fast pace of innovations in DNNs. To overcome this problem, we propose FP-DNN (Field Programmable DNN), an end-to-end framework that takes TensorFlow-described DNNs as input, and automatically generates the hardware implementations on FPGA boards with RTL-HLS hybrid templates. FP-DNN performs model inference of DNNs with our high-performance computation engine and carefully-designed communication optimization strategies. We implement CNNs, LSTM-RNNs, and Residual Nets with FPDNN, and experimental results show the great performance and flexibility provided by our proposed FP-DNN framework. Yijin Guan, Hao Liang 0003, Ningyi Xu, Shaoshuai Shi, Xi Chen 0107, Guangyu Sun 0003, Wei Zhang 0012, Jason Cong |
FCCM | 2 |
| 2016 | HeteroSim: A heterogeneous CPU-FPGA simulatorabstractHeterogeneous computing is rapidly gaining increased attention due to the promise it holds in overcoming power and performance walls in traditional computing systems. With its focus on customized processing nodes dedicated to the different tasks in an application, it is hoped that these walls will be overcome. Therefore, CPU-FPGA co-architectures are also gaining ground in application areas like recognition, mining, search, datacenter etc. However, research in CPU-FPGA co-architecture is constrained by the available synthesis and simulation tools which do not provide an integrated system level simulation and architectural exploration environment. This becomes critical when we incorporate novel memory hierarchies, multi-processor chip architectures, hardware level cache coherence etc. In this paper, we describe our open source and integrated system level simulator and architecture exploration tool called HeteroSim. It supports x86 based multi-core processor combined with a FPGA via bus-based architecture. It allows integrated system level simulation and returns performance metrics to understand application performance with respect to the simulated architectural configuration. Liang Feng 0001, Hao Liang 0003, Sharad Sinha, Wei Zhang 0012 |
FPL | 2 |
| 2015 | Static hardware task placement on multi-context FPGA using hybrid genetic algorithmabstractField Programmable Gate Arrays (FPGAs) are becoming pervasive in various kinds of computationally demanding applications. Working in a tightly coupled processor-coprocessor architecture, FPGAs are often anticipated to accelerate multiple fine-grained or coarse-grained tasks simultaneously. Single-context FPGAs are commonly used in such systems. With the recent development of emerging memory technologies, multi-context FPGAs that support dynamic reconfiguration with high-density non-volatile memories become feasible. Compared to single-context FPGAs, multi-context FPGAs are able to accelerate significantly more tasks with only moderate area and power overhead. However, the best way to utilize the computation capacity advantage of multi-context FPGAs for hardware task mapping remains an interesting and unexploited problem. In this paper, we first propose the framework of a processor-coprocessor architecture with multi-context FPGA as the coprocessor for multiple-task acceleration. Under the framework, a hybrid placement strategy based on genetic and greedy algorithms is proposed to efficiently place a set of tasks onto the multi-context FPGA to achieve the best logic capacity utilization. Experiments on real and synthetic benchmarks demonstrate the efficiency of the proposed algorithm compared with other general approaches. Hao Liang 0003, Sharad Sinha, Rakesh Warrier, Wei Zhang 0012 |
FPL | 1 |
| 2015 | Hierarchical library based power estimator for versatile FPGAsabstractFPGA is a promising hardware accelerator in modern high-performance computing systems. In such a system, power is a key factor in the design requiring thermal and energy-saving considerations. Modern power estimators for FPGA either support specific hardware provided by FPGA vendors or contain power models for certain types of conventional FPGA architectures. However, with technology advancement, novel versatile FPGA architectures are introduced to further augment current FPGA architecture, such as emerging FPGA with non-volatile memory, nano-wire interconnection of reconfigurable array, etc. To evaluate the power consumption of various FPGA designs, the power estimator has to be made more flexible and extensible to support new devices and architectures. We introduce a novel power estimator with hierarchical library supporting power models at different levels, e.g. novel circuits of components, emerging memory devices, time-multiplexing architecture etc. The power estimator also supports coarse-grain or fine-grain power estimation defined by users for achieving complexity-accuracy trade-off. Simulation results of benchmarks on the proposed power estimator against commercial estimators demonstrate accuracy of our tool. The proposed tool demonstrates flexibility to estimate power for both existing FPGA architectures and new architectures. Hao Liang 0003, Wei Zhang 0012, Sharad Sinha, Yi-Chung Chen, Hai Li 0001 |
FPL | 1 |
| 2015 | Leveraging Hotspots and Improving Chip Reliability via Carbon Nanotube Grid Thermal StructureabstractThe increasing power consumption of integrated circuits (ICs) enabled by technology scaling requires more efficient heat dissipation solutions to improve overall chip reliability and reduce hotspots. Rapidly growing 3-D IC technology strengthens the requirement with more devices stacked per unit area. Thermal interface material (TIM) and MicroChannel are widely adopted strategies to resolve the heat dissipation problem. In recent years, carbon nanotubes (CNTs) have been proposed as a promising TIM due to their superior thermal conductivity. Several CNT-based thermal structures for improving chip heat dissipation have been proposed and demonstrated significant temperature reduction. In this project, we developed an improved CNT TIM structure which includes a CNT grid and thermal vias. It collaborates with MicroChannel to dissipate heat more efficiently in 3-D chips and at the same time, obtain more uniform chip thermal profiles. We present simulation-based experimental results that indicate up to 19.88% peak temperature reduction, 7.81% average temperature reduction, over 66% maximum temperature difference reduction on chip and 17.26% improvement in chip reliability for IBM-PLACE 2.0 circuit benchmarks, showing the effectiveness of our proposed thermal structure for resolving thermal challenge and improving chip reliability in 3-D IC. Hao Liang 0003, Wei Zhang 0012, Shengqi Yang, Pallav Gupta |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2014 | Hierarchical library-based power estimator for versatile FPGAs (abstract only)abstractFPGAs are becoming promising hardware accelerators for high performance computing systems, such as cloud computing, big-data processing, etc., where power is a key factor due to thermal and energy saving considerations. Current CAD tools for FPGA power estimation either support specific hardware provided by vendors or contain power models for mainly conventional FPGA architectures. However, with technology advancement, versatile novel FPGA architectures are being proposed to further augment current FPGA architecture at various aspects, such as emerging FPGA based on non-volatile memory, improved logic and DSP design, etc. In order to evaluate the power consumption of versatile FPGA designs, the power estimator has to be made more flexible and extendable to support new devices and architectures. In this work, we proposed such a tool that the power estimation can be performed based on a hierarchical library which contains power models at different levels, such as circuit components or devices. The tool can collect resource utilization of FPGA for the implemented circuit, and then perform power estimation at coarse-grain or fine-grain levels based on the hierarchical library to achieve the desired complexity-accuracy trade-off. The flexibility is provided that users can customize the hierarchical library for new circuit components or devices with power number of their own study. Hao Liang 0003, Yi-Chung Chen, Wei Zhang 0012, Hai Li 0001 |
FPGA | 1 |
| 2014 | Reconfigurable DSP block design for dynamically reconfigurable architectureabstractReconfigurable architectures, such as Field-Programmable Gate Arrays (FPGAs), have become one of the key digital circuit implementation platform over the last decade due to its short time-to-market and low design cost. However, the major bottlenecks of FPGAs are their low logic utilization rate and long reconfiguration latency. In order to overcome these limitations, novel dynamically reconfigurable architectures, such as NATURE architecture, have been proposed. It enables runtime reconfiguration and reuse of hardware resources. Significant improvements on logic density, power reduction and reconfiguration flexibility are achieved. However, the previous architectures mainly focus on fine-grain logic. Since modern FPGAs are widely used in computation intensive applications, coarse-grain DSP blocks are needed to further enhance the performance. In this paper, we propose the design of a dynamically reconfigurable DSP block, which can be run-time reconfigured to implement different arithmetic functions in different clock cycles. We first demonstrate its efficiency through implementing typical DSP functions. Then based on NATURE design, simulations on seven benchmarks are performed to show that with DSP blocks, the performance is improved by 58.6% compared to fine-grain NATURE architecture. Then we demonstrate the efficiency reduction of DSP block number by enabling run-time reconfiguration. Rakesh Warrier, Hao Liang 0003, Wei Zhang 0012 |
ISCAS | 2 |
| 2013 | Thermal simulator of 3D-IC with modeling of anisotropic TSV conductance and microchannel entrance effectsabstractThis paper presents a fast and accurate steady state thermal simulator for heatsink and microfluid-cooled 3D-ICs. This model considers the thermal effect of TSVs at fine-granularity by calculating the anisotropic equivalent thermal conductances of a solid grid cell if TSVs are inserted. Entrance effect of microchannels is also investigated for accurate modeling of microfluidic cooling. The proposed thermal simulator is verified against commercial multiphysics solver COMSOL and compared with Hotspot and 3D-ICE. Simulation results shows that for heatsink cooling, the proposed simulator is as accurate as Hotspot but runs much faster at moderate granularity. For microfluidic cooling, our proposed simulator is much more accurate than 3D-ICE in its estimation of steady state temperature and thermal distribution. Hanhua Qian, Hao Liang 0003, Chip-Hong Chang, Wei Zhang 0012, Hao Yu 0001 |
ASP-DAC | 2 |