Sihao Liu

dblp:175/5266 · DBLP profile ↗
← Back
19ranked-venue papers
5as first author
12since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 8 · 1 first-author · 4 since 2021Artificial intelligence and machine learning · 5 · 2 first-author · 3 since 2021Software engineering, systems software and programming languages · 3 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 first-author · 2 since 2021Computer networks · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Double Graph Attention Network for predicting non-alcoholic fatty liver disease in patients with type 2 diabetes
Tianbin Chen, Yongbin Zeng, Sihao Liu, Ya Fu, Qishui Ou, Zhiheng Zhou 0003
Artif. Intell. Medicine5
2026 CyBond Net: Rethinking message passing mechanism via graph edge space
Sihao Liu, Zhiheng Zhou 0003, Weihua He, Guiying Yan
Knowl. Based Syst.1
2026 Traffic Engineering in Large-Scale Networks With Generalizable Graph Neural Networks
abstract
Traffic Engineering (TE) in large-scale networks like cloud Wide Area Networks (WANs) and Low Earth Orbit (LEO) satellite constellations is a critical challenge. Although learning-based approaches have been proposed to address the scalability of traditional TE algorithms, their practical application is often hindered by a lack of generalization, high training overhead, and a failure to respect link capacities. This paper proposes TELGEN, a novel TE algorithm that learns to solve TE problems efficiently in large-scale network scenarios, while achieving superior generalizability across diverse network conditions. TELGEN is based on the novel idea of transforming the problem of “predicting the optimal TE solution” into “predicting the optimal TE algorithm”, which enables TELGEN to learn and efficiently approximate the end-to-end solving process of classical optimal TE algorithms. The learned algorithm is agnostic to the exact underlying network topology or traffic patterns, and is able to very efficiently solve TE problems given arbitrary inputs and generalize well to unseen topologies and demands. We train and evaluate TELGEN with random and real-world topologies, with networks of up to 5000 nodes and 3.6×106links in testing. TELGEN shows less than 3% optimality gap while ensuring feasibility in all testing scenarios, even when the test network has 2-20× more nodes than the largest training network. It also saves up to 84% TE solving time than traditional interior-point method, and reduces up to 79.6% training time per epoch than the state-of-the-art learning-based algorithm.
Fangtong Zhou, Sihao Liu, Ruozhou Yu, Guoliang Xue
IEEE Trans. Netw.2
2025 Enhancing Online Continual Learning with Plug-and-Play State Space Model and Class-Conditional Mixture of Discretization
abstract
Online continual learning (OCL) seeks to learn new tasks from data streams that appear only once, while retaining knowledge of previously learned tasks. Most existing methods rely on replay, focusing on enhancing memory retention through regularization or distillation. However, they often overlook the adaptability of the model, limiting the ability to learn generalizable and discriminative features incrementally from online training data. To address this, we introduce a plug-and-play module, S6MOD, which can be integrated into most existing methods and directly improve adaptability. Specifically, S6MOD introduces an extra branch after the backbone, where a mixture of discretization selectively adjusts parameters in a selective state space model, enriching selective scan patterns such that the model can adaptively select the most sensitive discretization method for current dynamics. We further design a class-conditional routing algorithm for dynamic, uncertainty-based adjustment and implement a contrastive discretization loss to optimize it. Extensive experiments combining our module with various models demonstrate that S6MOD significantly enhances model adaptability, leading to substantial performance gains and achieving the state-of-the-art results. The code is available at https://github.com/MyToumaKazusa/S6MOD.
Sihao Liu, David A. Clifton, Bernard Ghanem
CVPR1
2025 NoH: NoC Compilation in High-Level Synthesis
abstract
In FPGAs, high communication latency in multi-die chips has driven the integration of hardened networks-on-chip (NoCs) in commercial devices. However, for programming FPGAs with high-level synthesis (HLS), existing tools only provide low-level cumbersome abstractions, and only work for offloading memory accesses. Furthermore, these abstractions remain inaccessible to programmers due to their reliance on placement knowledge. While automatically leveraging the NoC without manual intervention is ideal, it poses several challenges: 1. Managing the trade-off in resource utilization between the hard NoC and the Programmable Logic (PL). 2. Allocating limited hard NoC resources between different communication in the designs. 3. Aligning hard NoC and PL placement even though the actual PL placement cannot be determined beforehand. We address these challenges by developing NoH, the first HLS flow that automates hard NoC offloading. First, we develop a formal NoC-aware placement algorithm that leverages integer linear programming (ILP) and considers the first two challenges for offloading external memory accesses and latency-insensitive communication between modules. Then, we arrange the ports synergistically with PL modules via a port-affinity model that approximates the PL placement. Finally, NoH is integrated into an end-to-end HLS flow and evaluated on 4 workloads with diverse communication patterns. NoH gains 20% FPGA frequency over AMD tools by leveraging the hard NoC. Compared to AutoBridge [1], a recent high-level physical synthesis technique that optimizes frequency but does not consider the hard NoC, NoH never fails place-and-route by offloading inter-die crossings (AutoBridge fails in 31% of workload configurations tested) and is faster (6%) for the rest.
Huifeng Ke, Sihao Liu, Licheng Guo, Zifan He, Linghao Song, Suhail Basalama, Yuze Chi, Tony Nowatzki, Jason Cong
FCCM2
2025 KGCE: Knowledge-Augmented Dual-Graph Evaluator for Cross-Platform Educational Agent Benchmarking with Multimodal Language Models
abstract
With the rapid adoption of multimodal large language models (MLMs) in autonomous agents, cross-platform task execution capabilities in educational settings have garnered significant attention. However, existing benchmark frameworks still exhibit notable deficiencies in supporting cross-platform tasks in educational contexts, especially when dealing with school-specific software (such as XiaoYa Intelligent Assistant, HuaShi XiaZi, etc.), where the efficiency of agents often significantly decreases due to a lack of understanding of the structural specifics of these private-domain software. Additionally, current evaluation methods heavily rely on coarse-grained metrics like goal orientation or trajectory matching, making it challenging to capture the detailed execution and efficiency of agents in complex tasks. To address these issues, we propose KGCE (Knowledge-Augmented Dual-Graph Evaluator for Cross-Platform Educational Agent Benchmarking with Multimodal Language Models), a novel benchmarking platform that integrates knowledge base enhancement and a dual-graph evaluation framework. We first constructed a dataset comprising 104 education-related tasks, covering Windows, Android, and cross-platform collaborative tasks. KGCE introduces a dual-graph evaluation framework that decomposes tasks into multiple sub-goals and verifies their completion status, providing fine-grained evaluation metrics. To overcome the execution bottlenecks of existing agents in private-domain tasks, we developed an enhanced agent system incorporating a knowledge base specific to school-specific software. The code can be found at https://github.com/Kinginlife/KGCE.
Zixian Liu, Sihao Liu, Yuqi Zhao 0001
SMC2
2025 Develop a Deep-Learning Model to Predict Cancer Immunotherapy Response Using In-Born Genomes
abstract
The emergence of immune checkpoint inhibitors (ICIs) has significantly advanced cancer treatment. However, only 15-30% of the cancer patients respond to ICI treatment, which stimulates and enhances host immunity to eliminate tumor cells. ICI treatment is very expensive and has potential adverse reactions; therefore, it is crucial to develop a method which enables to accurately and rapidly assess a patient's suitability before ICI treatment. We complied germline whole-genome sequencing (WES) data of 37 melanoma patients who have been treated with ICIs and sequenced in our lab previously, and the WES data of other 700 ICI-treated cancer patients in public domain. Using these data, we proposed a novel double-channel attention neural network (DANN) model to predict cancer ICI-response and validate the predictions. DANN achieved a mean accuracy and AUC of 0.95 and 0.98, respectively, which outperformed traditional machine learning methods. Enrichment analysis of the DANN-identified genes indicated that cancer patients whose in-born genomic variants might mainly affect host immune system in a wide-ranging manner, and then affect ICI response. Finally, we found a set of 12 genes bearing genomic variants were significantly associated with cancer patient survivals after ICI treatment.
Zhiheng Zhou 0003, Sihao Liu, Guanghui Wang 0002, Guiying Yan, Edwin Wang
IEEE J. Biomed. Health Informatics3
2024 GSL-Mash: Enhancing Mashup Creation Service Recommendations Through Graph Structure Learning
Sihao Liu, Tianyu Jiang 0003, Hanchuan Xu, Zhongjie Wang 0003
ICSOC (2)1
2023 Efficient Few-Shot Image Generation via Lightweight Octave Generative Adversarial Networks
Sihao Liu
ICIG (2)1
2022 Near-Stream Computing: General and Transparent Near-Cache Acceleration
abstract
Data movement and communication have become the primary bottlenecks in large multicore systems. The near-data computing paradigm provides a solution: move computation to where the data resides on-chip. Two challenges keep near-data computing from the mainstream: lack of programmer transparency and applicability. Programmer transparency requires providing sequential memory semantics with distributed computation, which requires burdensome coordination. Broad applicability requires support for combinations of address patterns (e.g. affine, indirect, multi-operand) and computation types (loads, stores, reductions, atomics).We find that streams – coarse grain memory access patterns – are a powerful ISA abstraction for near data offloading. Tracking data access at stream-granularity heavily reduces the burden of coordination for providing sequential semantics. Decomposing the problem using streams means that arbitrary combinations of address and computation patterns can be combined for broad generality.With this insight, we develop a paradigm called near-stream computing, comprising a compiler, CPU ISA extension, and a microarchitecture that facilitate programmer transparent computation offloading to shared caches. We evaluate our system on OpenMP kernels that stress broad addressing and compute behavior, and find that 46% of dynamic instructions can be offloaded to remote banks, reducing the network traffic by 76%. Overall it achieves 2.13× speedup over a state-of-the-art near-data computing technique, with a 1.90× energy efficiency gain.
Zhengrong Wang, Jian Weng 0002, Sihao Liu, Tony Nowatzki
HPCA3
2022 OverGen: Improving FPGA Usability through Domain-specific Overlay Generation
abstract
FPGAs have been proven to be powerful computational accelerators across many types of workloads. The mainstream programming approach is high level synthesis (HLS), which maps high-level languages (e.g. C+ #pragmas) to hardware. Unfortunately, HLS leaves a significant programmability gap in terms of reconfigurability, customization and versatility: Although HLS compilation is fast, the downstream physical design takes hours to days; FPGA reconfiguration time limits the time-multiplexing ability of hardware, and tools do not reason about cross-workload flexibility. Overlay architectures mitigate the above by mapping a programmable design (e.g. CPU, GPU, etc.) on top of FPGAs. However, the abstraction gap between overlay and FPGA leads to low efficiency/utilization. Our essential idea is to develop a hardware generation framework targeting a highly-customizable overlay, so that the abstraction gap can be lowered by tuning the design instance to applications of interest. We leverage and extend prior work on customizable spatial architectures, SoC generation, accelerator compilers, and design space explorers to create an end-to-end FPGA acceleration system. Our novel techniques address inefficient networks between on-chip memories and processing elements, as well as improving DSE by reducing the amount of recompilation required. Our framework, OverGen, is highly competitive with fixed-function HLS-based designs, even though the generated designs are programmable with fast reconfiguration. We compared to a state-of-the-art DSE-based HLS framework, AutoDSE. Without kernel-tuning for AutoDSE, OverGen gets 1.2$\times$ geomean performance, and even with manual kernel-tuning for the baseline, OverGen still gets 0.55$\times$ geomean performance--all while providing runtime flexibility across workloads.
Sihao Liu, Jian Weng 0002, Dylan Kupsh, Atefeh Sohrabizadeh, Zhengrong Wang, Licheng Guo, Jiuyang Liu, Maxim Zhulin, Rishabh Mani, Lucheng Zhang, Jason Cong, Tony Nowatzki
MICRO1
2021 PolyGraph: Exposing the Value of Flexibility for Graph Processing Accelerators
abstract
Because of the importance of graph workloads and the limitations of CPUs/GPUs, many graph processing accelerators have been proposed. The basic approach of prior accelerators is to focus on a single graph algorithm variant (eg. bulk-synchronous + slicing). While helpful for specialization, this leaves performance potential from flexibility on the table and also complicates understanding the relationship between graph types, workloads, algorithms, and specialization.In this work, we explore the value of flexibility in graph processing accelerators. First, we identify a taxonomy of key algorithm variants. Then we develop a template architecture (PolyGraph) that is flexible across these variants while being able to modularly integrate specialization features for each.Overall we find that flexibility in graph acceleration is critical. If only one variant can be supported, asynchronous-updates/priority-vertex-scheduling/graph-slicing is the best design, achieving 1.93× speedup over the best-performing accelerator, GraphPulse. However, static flexibility per-workload can further improve performance by 2.71×. With dynamic flexibility per-phase, performance further improves by up to 50%.
Vidushi Dadu, Sihao Liu, Tony Nowatzki
ISCA2
2020 A Hybrid Systolic-Dataflow Architecture for Inductive Matrix Algorithms
abstract
Dense linear algebra kernels are critical for wireless, and the oncoming proliferation of 5G only amplifies their importance. Due to the inductive nature of many such algorithms, parallelism is difficult to exploit: parallel regions have fine-grain producer/consumer interaction with iteratively changing depen-dence distance, reuse rate, and memory access patterns. This makes multi-threading impractical due to fine-grain synchronization, and vectorization ineffective due to the non-rectangular iteration domain. CPUs, DSPs, and GPUs perform order-of-magnitude below peak. Our insight is that if the nature of inductive dependences and memory accesses were explicit in the hardware/software interface, then a spatial architecture could efficiently execute parallel code regions. To this end, we first develop a novel execution model, inductive dataflow, where inductive dependence patterns and memory access patterns (streams) are first-order primitives. Second, we develop a hybrid spatial architecture combining systolic and tagged dataflow execution to attain high utilization at low energy and area cost. Finally, we create a scalable design through a novel vector-stream control model which amortizes control overhead both in time and spatially across architecture lanes. We evaluate our design, REVEL, with a full stack (compiler, ISA, simulator, RTL). Across a suite of linear algebra kernels, REVEL outperforms equally-provisioned DSPs by 4.6×-37×. Compared to state-of-the-art spatial architectures, REVEL is mean 3× faster. Compared to a set of ASICs, REVEL is only 2× the power and half the area.
Jian Weng 0002, Sihao Liu, Zhengrong Wang, Vidushi Dadu, Tony Nowatzki
HPCA2
2020 DSAGEN: Synthesizing Programmable Spatial Accelerators
abstract
Domain-specific hardware accelerators can provide orders of magnitude speedup and energy efficiency over general purpose processors. However, they require extensive manual effort in hardware design and software stack development. Automated ASIC generation (eg. HLS) can be insufficient, because the hardware becomes inflexible. An ideal accelerator generation framework would be automatable, enable deep specialization to the domain, and maintain a uniform programming interface. Our insight is that many prior accelerator architectures can be approximated by composing a small number of hardware primitives, specifically those from spatial architectures. With careful design, a compiler can understand how to use available primitives, with modular and composable transformations, to take advantage of the features of a given program. This suggests a paradigm where accelerators can be generated by searching within such a rich accelerator design space, guided by the affinity of input programs for hardware primitives and their interactions. We use this approach to develop the DSAGEN framework, which automates the hardware/software co-design process for reconfigurable accelerators. For several existing accelerators, our evaluation demonstrates that the compiler can achieve 89% of the performance of manually tuned versions. For automated design space exploration, we target multiple sets of workloads which prior accelerators are design for; the generated hardware has mean 1.3× perf2/mm2over prior programmable accelerators.
Jian Weng 0002, Sihao Liu, Vidushi Dadu, Zhengrong Wang, Preyas Shah, Tony Nowatzki
ISCA2
2019 Towards General Purpose Acceleration by Exploiting Common Data-Dependence Forms
abstract
With slowing technology scaling, specialized accelerators are increasingly attractive solutions to continue expected generational scaling of performance. However, in order to accelerate more advanced algorithms or those from challenging domains, supporting data-dependence becomes necessary. This manifests as either data-dependent control (eg. join two sparse lists), or data-dependent memory accesses (eg. hash-table access). These forms of data-dependence inherently couple compute with memory, and also preclude efficient vectorization -- defeating the traditional mechanisms of programmable accelerators (eg. GPUs).
Vidushi Dadu, Jian Weng 0002, Sihao Liu, Tony Nowatzki
MICRO3
2019 μIR -An intermediate representation for transforming and optimizing the microarchitecture of application accelerators
abstract
Creating high quality application-specific accelerators requires us to make iterative changes to both algorithm behavior and microarchitecture, and this is a tedious and error-prone process. High-Level Synthesis (HLS) tools [5, 10] generate RTL for application accelerators from annotated software. Unfortunately, the generated RTL is challenging to change and optimize. The primary limitation of HLS is that the functionality and microarchitecture are conflated together in a single language (such as C++). Making changes to the accelerator design may require code restructuring, and microarchitecture optimizations are tied with program correctness.
Amirali Sharifian, Reza Hojabr, Navid Rahimi, Sihao Liu, Apala Guha, Tony Nowatzki, Arrvindh Shriraman
MICRO4
2019 A hybrid instance-intensive workflow scheduling method in private cloud environment
Xin Ye 0004, Sihao Liu, Jiwei Liang, Yaochu Jin
Nat. Comput.3
2018 Ex Vivo and In Vivo Monitoring and Characterization of Thermal Lesions by High-Intensity Focused Ultrasound and Microwave Ablation Using Ultrasonic Nakagami Imaging
abstract
The feasibility of ultrasonic Nakagami imaging to evaluate thermal lesions by high-intensity focused ultrasound and microwave ablation was explored in ex vivo and in vivo liver models. Dynamic changes of the ultrasonic Nakagami parameter in thermal lesions were calculated, and ultrasonic B-mode and Nakagami images were reconstructed simultaneously. The contrast-to-noise ratio (CNR) between thermal lesions and normal tissue was used to estimate the contrast resolution of the monitoring images. After thermal ablation, a bright hyper-echoic region appeared in the ultrasonic B-mode and Nakagami images, identifying the thermal lesion. During thermal ablation, mean values of Nakagami parameter showed an increasing trend from 0.72 to 1.01 for the ex vivo model and 0.54 to 0.72 for the in vivo model. After thermal ablation, mean CNR values of the ultrasonic Nakagami images were 1.29 dB (ex vivo) and 0.80 dB (in vivo), significantly higher ( ) than those for B-mode images. Thermal lesion size, assessed using ultrasonic Nakagami images, shows a good correlation to those obtained from the gross-pathology images (for the ex vivo model: length, = 0.96; width, = 0.90; for the in vivo model: length, = 0.95; width, = 0.85). This preliminary study suggests that ultrasonic Nakagami parameter may have a potential use in evaluating the formation of thermal lesions with better image contrast. Moreover, ultrasonic Nakagami imaging combined with B-mode imaging may be utilized as an alternative modality in developing monitoring systems for image-guided thermal ablation treatments.
Shaoqiang Shang, Yuqiang Han, Chunming Gu, Sihao Liu, Gang Niu 0003, Ayache Bouakaz, Mingxi Wan
IEEE Trans. Medical Imaging6
2017 User-oriented many-objective cloud workflow scheduling based on an improved knee point driven evolutionary algorithm
Xin Ye 0004, Sihao Liu, Yanli Yin, Yaochu Jin
Knowl. Based Syst.2