VLDB 2026 Research / reviewers in the wild / expert
Hui Kou
dblp:52/1735
· DBLP profile ↗
13ranked-venue papers
2as first author
7since 2021 · last 2026
0009-0007-4874-0867ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Theory of computation · 6 · 1 first-author · 2 since 2021Artificial intelligence and machine learning · 3 · 1 since 2021Systems, architecture and hardware · 3 · 1 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | 3D-TANoC: Thermal-Aware 3D LLM Accelerator with Hierarchical NoC and Operator-Aware Dataflow Mapping Strategy
Xingyu Xu 0008, Dengke Liu, Zihan Zou, Xilong Kang, Hui Kou, Hao Cai 0001, Bo Liu 0019 |
ISCAS | 6 |
| 2025 | TWDP: A Vision Transformer Accelerator with Token-Weight Dual-Pruning Strategy for Edge Device DeploymentabstractVision Transformers (ViTs) have attracted significant attention due to their superior accuracy compared to convolutional neural networks (CNNs) in various computer vision tasks. However, their substantial computational load and significant memory footprint lead to excessive delay and considerable data storage overhead, posing challenges for resource-limited edge device deployment. To address these issues, we present TWDP, a vision transformer accelerator employing a Token-Weight Dual-Pruning strategy to enhance the efficiency of the inference process. Firstly, we propose a parameter-free self-adaptive token pruning method to skip redundant computations in an image-dependent manner. Secondly, we apply a Hessian-aware layer-wise N:M weight pruning approach to minimize storage overhead, memory access, and computational power consumption. Additionally, to manage the complex computing patterns in ViTs, an overlapping dataflow is utilized to further reduce temporal storage and inference latency. Implemented and evaluated under an industrial 28nm technology, the proposed TWDP framework reduces 66.1% weight storage requirements and achieves an energy efficiency of 2070.9 FPS/W. Compared to state-of-the-art architectures, TWDP obtains a 1.6× energy efficiency improvement with negligible accuracy loss, demonstrating the superiority of TWDP in edge device deployment scenarios. Guang Yang 0036, Xinming Yan, Hui Kou, Zihan Zou, Qingwen Wei, Hao Cai 0001, Bo Liu 0019 |
ASP-DAC | 3 |
| 2025 | H3D-LLM: Heterogeneous 3D Chiplet Design for LLM Inference with Dynamic Task Scheduling and Memory-Aware OrchestrationabstractThe exponential growth of Large Language Model (LLM) intensifies hardware demands for energy-efficient, low-latency architectures with scalable memory bandwidth. While 3D chiplet integration addresses conventional systems’ memory wall limitations, three critical challenges persist: asymmetric compression constraints from divergent sparsity-precision requirements across attention/projection layers, tier-level load imbalance from static resource allocation in dynamic computation patterns, and coupling-induced signal degradation in high-density TSV networks, especially under LLM-phase-specific traffic with spatiotemporal burstiness. To address these, we present H3D-LLM, a vertically heterogeneous architecture combining analog/digital Computing-in-Memory (CIM) and Neural Processing Unit (NPU) chiplets through three innovations. First, a Sparse-Aware Dynamic Execution Framework (SADEF) with Precision-Adaptive Quantization Mechanism (PAQM) enables hardware-aware compression via layer-wise unstructured sparsity detection and INT4/8-FP/BF16 mixed precision. Second, a 3D Spatio-Temporal Interleaved Parallelism (3D-STIP) with Semantic-Aware Tiered Storage (SATS) eliminates resource imbalance and improves memory efficiency through dynamic sub-batch partitioning and Key-Value (KV) cache aware management. Third, a Phase-Adaptive TSV Management (PATM) scheme with dynamic encoding and cluster-based allocation enhances inter-connect efficiency and signal integrity through runtime-aware partitioning and phase-specific dataflow scheduling. Evaluations on Llama-7B demonstrate that H3D-LLM achieves 12.3× higher energy efficiency and 8.4× faster inference than the A800 GPU, while its TSV strategy increases eye height by 12% and reduces bit error rate by up to 60× compared to naïve 3D accelerators. Hui Kou, Chenjie Xia, Liyi Li 0004, Hao Cai 0001, Xin Si, Bo Liu 0019 |
ICCAD | 1 |
| 2024 | Small-Footprint Automatic Speech Recognition System using Two-Stage Transfer Learning based Symmetrized Ternary Weight NetworkabstractTraditional automatic speech recognition (ASR) models face challenges when deployed on edge devices due to their high computational requirements and storage demands. To address this issue, we present a novel ASR system specifically designed for edge applications, encompassing both keyword spotting (KWS) and speaker verification (SV) functionalities with on chip learning for speaker registration. Our proposed system employs a compact model trained using a two-stage transfer learning method for on-the-fly small-sample speaker registration. In the proposed model, sparsity-controllable weights are symmetrically ternary-quantized to further exploit data reuse. Additionally, we introduce a Huffman-coding based weight lossy compression method to achieve efficient storage compaction. Moreover, we propose a specialized classifier taking the signal-to-noise ratio into account to enhance the accuracy of SV. The proposed ASR system has been successfully deployed on a 1.65mm2custom chip fabricated under 28-nm technology, with only 10.84KB of on-chip memory. This compact system effectively handles KWS and SV tasks, as well as on-chip speaker registration. Xuanhao Zhang, Hui Kou, Chenjie Xia, Hao Cai 0001, Bo Liu 0019 |
ICASSP | 2 |
| 2024 | A construction of free dcpo-conesabstractAbstract We give a construction of the free dcpo-cone over any dcpo. There are two steps for getting this result. Firstly, we extend the notion of power domain to directed spaces which are equivalent to $T_0$ monotone-determined spaces introduced by Erné, and we construct the probabilistic powerspace of the monotone determined space, which is defined as a free monotone determined cone. Secondly, we take D-completion of the free monotone determined cone over the dcpo with its Scott topology. In addition, we show that generally the valuation power domain of any dcpo is not the free dcpo-cone. Yuxu Chen, Hui Kou, Zhenchao Lyu, Xiaolin Xie |
Math. Struct. Comput. Sci. | 2 |
| 2024 | Semi-supervised clustering guided by pairwise constraints and local density structures
Zhiguo Long, Hua Meng 0001, Yuxu Chen, Hui Kou |
Pattern Recognit. | 5 |
| 2024 | Upper powerdomains of quasicontinuous dcpos
Yuxu Chen, Hui Kou, Zhenchao Lyu |
Theor. Comput. Sci. | 2 |
| 2018 | A partial solution to an open problem of Amadio and Curien
Xiaoyong Xi, Jinbo Yang, Hui Kou |
Inf. Comput. | 3 |
| 2017 | A stable universal domain related to $\mathbb{T}$ωabstractIn 1978, G. Plotkin noticed that $\mathbb{T}$ ω , the cartesian product of ω copies of the three element flat domain of Booleans, is a universal domain, where ‘universal’ means that the retracts of $\mathbb{T}$ ω for Scott's continuous semantics are exactly all the ωCC -domains, which with Scott continuous functions form a cartesian closed category. As usual, ‘ ω ’ is for ‘countably based,’ and here ‘CC’ is for ‘conditionally complete,’ which essentially means that any subset which is pairwise bounded has a least upper bound. Since $\mathbb{T}$ ω is also an ωDI -domain (an important structure in stable domain theory), the following problem arises naturally: is there a cartesian closed category C of domains with stable functions such that $\mathbb{T}$ ω , or a related structure, is universal in C for Berry’s stable semantics? The aim of this paper is to answer this question. We first investigate the properties of stable retracts. We introduce a new class of domains called conditionally complete DI -domains ( CCDI -domain for short) and show that, (1) $\mathbb{T}$ ω is an ωCCDI -domain and the category of CCDI -domains (resp. ωCCDI -domains) with stable functions is cartesian closed; (2) [ $\mathbb{T}$ ω → st $\mathbb{T}$ ω ] is a stable universal domain in the sense that every ωCCDI -domain is a stable retract of [ $\mathbb{T}$ ω → st $\mathbb{T}$ ω ], where [ $\mathbb{T}$ ω → st $\mathbb{T}$ ω ] is the stable function space of $\mathbb{T}$ ω ; (3) in particular, [ $\mathbb{T}$ ω → st $\mathbb{T}$ Hui Kou |
Math. Struct. Comput. Sci. | 2 |
| 2015 | Belief Revision with General Epistemic StatesabstractIn order to properly regulate iterated belief revision, Darwiche and Pearl (1997) model belief revision as revising epistemic states by propositions. An epistemic state in their sense consists of a belief set and a set of conditional beliefs. Although the denotation of an epistemic state can be indirectly captured by a total preorder on the set of worlds, it is unclear how to directly capture the structure in terms of the beliefs and conditional beliefs it contains. In this paper, we first provide an axiomatic characterisation for epistemic states by using nine rules about beliefs and conditional beliefs, and then argue that the last two rules are too strong and should be eliminated for characterising the belief state of an agent. We call a structure which satisfies the first seven rules a general epistemic state (GEP). To provide a semantical characterisation of GEPs, we introduce a mathematical structure called belief algebra, which is in essence a certain binary relation defined on the power set of worlds.We then establish a 1-1 correspondence between GEPs and belief algebras, and show that total preorders on worlds are special cases of belief algebras. Furthermore, using the notion of belief algebras, we extend the classical iterated belief revision rules of Darwiche and Pearl to our setting of general epistemic states. Hua Meng 0001, Hui Kou, Sanjiang Li |
AAAI | 2 |
| 2015 | All cartesian closed categories of quasicontinuous domains consist of domains
Xiaodong Jia 0002, Achim Jung, Hui Kou, Qingguo Li |
Theor. Comput. Sci. | 3 |
| 2005 | Discrete dynamical systems in L-topological spaces
Hui Kou, Maokang Luo, Weinian Zhang |
Fuzzy Sets Syst. | 2 |
| 2003 | On (weakly) local approximation spaces of information systems
Hui Kou, Maokang Luo |
Fundam. Informaticae | 1 |