EDBT 2026 Demo / reviewers in the wild / expert
Kaixin Huang
dblp:219/2228
· DBLP profile ↗
27ranked-venue papers
10as first author
16since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 16 · 4 first-author · 8 since 2021Artificial intelligence and machine learning · 3 · 2 first-author · 3 since 2021Software engineering, systems software and programming languages · 3 · 2 first-authorDatabases, data management, data science and information retrieval · 3 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021Computer networks · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | vToP: VLM-Guided Transparent Object Perception for Robotic Manipulation
Kaixin Huang, Jiayan Zhuang, Sichao Ye, Kangkang Song, Jiangjian Xiao, Cong Zhai, Chendong Yao |
IEEE Trans Autom. Sci. Eng. | 1 |
| 2025 | Beyond Boundaries: Learning a Universal Entity Taxonomy across Datasets and Languages for Open Named Entity RecognitionabstractOpen Named Entity Recognition (NER), which involves identifying arbitrary types of entities from arbitrary domains, remains challenging for Large Language Models (LLMs). Recent studies suggest that fine-tuning LLMs on extensive NER data can boost their performance. However, training directly on existing datasets neglects their inconsistent entity definitions and redundant data, limiting LLMs to dataset-specific learning and hindering out-of-domain adaptation. To address this, we present B2NERD, a compact dataset designed to guide LLMs’ generalization in Open NER under a universal entity taxonomy. B2NERD is refined from 54 existing English and Chinese datasets using a two-step process. First, we detect inconsistent entity definitions across datasets and clarify them by distinguishable label names to construct a universal taxonomy of 400+ entity types. Second, we address redundancy using a data pruning strategy that selects fewer samples with greater category and semantic diversity. Comprehensive evaluation shows that B2NERD significantly enhances LLMs’ Open NER capabilities. Our B2NER models, trained on B2NERD, outperform GPT-4 by 6.8-12.0 F1 points and surpass previous methods in 3 out-of-domain benchmarks across 15 datasets and 6 languages. The data, models, and code are publicly available at https://github.com/UmeanNever/B2NER. Yuming Yang 0001, Wantong Zhao, Caishuang Huang, Junjie Ye 0005, Xiao Wang 0042, Huiyuan Zheng, Xueying Xu, Kaixin Huang, Yunke Zhang, Tao Gui, Qi Zhang 0001, Xuanjing Huang 0001 |
COLING | 10 |
| 2025 | Hyte: A Hotness-Aware Hybrid DRAM-PM Native Table Storage EngineabstractEmerging persistent memory (PM) technologies offer data persistence at close-to-DRAM latency, causing the legacy multi-layer storage stack to become a bottleneck. On the other hand, PM improves I/O efficiency while also introducing unique hardware features that require software modifications to exploit them. In this paper, we propose HYTE, a native table storage engine for hybrid memory, which abstracts the data service for the database’s table. HYTE performs a cross-layer design across the file system and storage engine to avoid the pitfalls of layered abstraction and provides SQL-compatible APIs. Meanwhile, it co-designs with PM’s properties to unlock the hardware’s performance potential. Furthermore, HYTE is hotness-aware and equipped with a suite of lightweight, customized fetching and eviction strategies. Our evaluation shows that HYTE can perform up to 1.4-7.2× better than existing state-of-the-art PM-based storage systems in industry and academia. Xiaopeng Fan 0003, Xiaoshuang Peng, Kaixin Huang, Chuliang Weng |
IEEE Trans. Computers | 3 |
| 2024 | SEDIT: Space-Efficient Discriminative Bit Tree for Hybrid Memory Indexing
Yuanjin Lin, Kaixin Huang, Kuankuan Guo, Linpeng Huang |
DASFAA (1) | 3 |
| 2024 | Automatic bug assignments without texts: a study
Kaixin Huang |
Frontiers Comput. Sci. | 2 |
| 2024 | A read-efficient and write-optimized hash table for Intel Optane DC Persistent Memory
Kaixin Huang |
Future Gener. Comput. Syst. | 2 |
| 2024 | Efficiently Adapt to New Dynamic via Meta-ModelabstractWe delve into the realm of offline meta-reinforcement learning (OMRL), a practical paradigm in the field of reinforcement learning that leverages offline data to adapt to new tasks. While prior approaches have not explored the utilization of context-based dynamical models to tackle OMRL problems, our research endeavors to fill this gap. Our investigation uncovers shortcomings in existing context-based methods, primarily related to distribution shifts during offline learning and challenges in establishing stable task representations. To address these issues, we formulate the problem as Hidden-Parameter MDPs and propose a framework for effective model adaptation using meta-models plus latent variables, which is inferred by the transformer-based system recognition module trained in an unsupervised fashion. Through extensive experimentation encompassing diverse simulated robotics and control tasks, we validate the efficacy of our approach and demonstrate its superior generalization ability compared to existing schemes, and explore multiple strategies for obtaining policies with personalized models. Our method achieves a model with reduced prediction error, outperforming previous methods in policy performance, and facilitating efficient adaptation when compared to prior dynamic model generalization methods and OMRL algorithms. Kaixin Huang, Chen Zhao 0019, Chun Yuan 0003 |
J. Artif. Intell. Res. | 1 |
| 2024 | RADAR: A Skew-Resistant and Hotness-Aware Ordered Index Design for Processing-in-Memory SystemsabstractPointer chasing becomes the performance bottleneck for today's in-memory indexes due to the memory wall. Emerging processing-in-memory (PIM) technologies are promising to mitigate this bottleneck, by enabling low-latency memory access and aggregated memory bandwidth scaling with the number of PIM modules. Prior PIM-based indexes adopt a fixed granularity to partition the key space and maintain static heights of skiplist nodes among PIM modules to accelerate index operations on skiplist, neglecting the changes in skewness and hotness of data access patterns during runtime. In this article, we present RADAR, an innovative PIM-friendly skiplist that dynamically partitions the key space among PIM modules to adapt to varying skewness. An offline learning-based model is employed to catch hotness changes to adjust the heights of skiplist nodes. In multiple datasets, RADAR achieves up to 198.2x performance improvement and consumes 47.4% less memory than state-of-the-art designs on real PIM hardware. Yifan Hua, Shengan Zheng, Weihan Kong, Kaixin Huang, Ruoyan Ma, Linpeng Huang |
IEEE Trans. Parallel Distributed Syst. | 5 |
| 2023 | TOD: Trend-Oriented Delay-Based Congestion Control in Lossless Datacenter NetworkabstractIn current high-speed data center networks, congestion control is crucial for ensuring consistent high performance. Over the past decade, researchers and developers have explored several congestion signals such as ECN, RTT, and INT. However, most of the existing congestion control algorithms suffer from either imprecise congestion detection due to ambiguous signals or excessive bandwidth loss due to aggressive rate decrease. This paper proposes a novel congestion control mechanism called TOD, which is a trend-oriented delay-based approach designed for lossless data center networks. TOD leverages the change in RTT to learn the congestion trend and adjusts the sending rate accordingly. By analyzing the congestion trend, the sender reacts by adjusting the sending rate to a reasonable level, while still maintaining high bandwidth utilization to dismiss congestion. The sender uses a reference rate, which is calculated by the receiver and communicated back to the sender, to achieve this target. Therefore, TOD is a sender-receiver cooperative congestion control mechanism. We evaluate TOD extensively in NS-3 simulations using both microbenchmark and macrobenchmark. Our experiments demonstrate that TOD outperforms DCQCN and Timely in terms of FCT and convergence speed. Kaixin Huang, Weihang Li, Lang Cheng |
APNet | 1 |
| 2023 | A server bypass architecture for hopscotch hashing key-value store on DRAM-NVM memories
Rulin Huang, Kaixin Huang |
J. Syst. Archit. | 3 |
| 2021 | Redesigning the Sorting Engine for Persistent Memory
Yifan Hua, Kaixin Huang, Shengan Zheng, Linpeng Huang |
DASFAA (3) | 2 |
| 2021 | Reno: An RDMA-Enabled, Non-Volatile Memory-Optimized Key-Value StoreabstractRemote direct memory access (RDMA) has been employed to boost remote data access for key-value stores, since it provides kernel-bypass, zero-copy and low-latency features. Meanwhile, existing RDMA-enabled key-value stores still have performance bottlenecks, since extra costs need to be spent on processing requests and keeping data consistency. This paper introduces Reno, an RDMA-enabled, NVM-optimized key-value store that supports fast remote access of persistent data. On the server-side, Reno is built atop a bucket-based hopscotch hash table, where the actual key-value items are stored in NVM (non-volatile memories), and metadata are managed by an in-DRAM index. On the client-side, Reno adopts a fully server-bypass paradigm for both remote read and write requests to achieve low latency and high throughput. We evaluate Reno on an Intel's Optane DC Persistent Memory platform with Infiniband network support. The results show the strengths of Reno. In particular, Reno outperforms its counterparts by 1.1~3.3x for remote reads and 1.9~4.8 x for remote writes in terms of latency; the speedups of concurrent throughput are up to 2.33 x, 2.36 x, 3.09 x and 5.08 x for read-only, read-heavy, write-heavy and write-only YCSB workloads, respectively. Rulin Huang, Kaixin Huang, Yuting Chen 0001 |
ICPADS | 2 |
| 2021 | HDNH: a read-efficient and write-optimized hashing scheme for hybrid DRAM-NVM memoryabstractWith high memory density, non-volatility and DRAM-scale latency, non-volatile memory (NVM) brings evolution to storage systems and durable data structures. And Intel Optane DC persistent memory module (AEP), the first commercial product of NVM, shows some features that are different from previous assumptions: higher read latency, lower bandwidth and block access granularity compared with DRAM. It is reasonable to build up hybrid memory to give full play to the complementary advantages of DRAM and NVM. In this paper, we present a read-efficient and write-optimized hashing scheme for hybrid DRAM-NVM memory, named HDNH (Hybrid DRAM-NVM Hashing). Our design can be summarized into three key points. First, we decouple the storage for data and metadata by placing key-value items in non-volatile table for persistence while placing metadata in Optimistic Compression Filter (OCF) to reduce excessive NVM accesses. Second, we design hot table in DRAM to speed up search requests and propose an efficient replacement strategy called RAFL. Third, we develop a fine-grained optimistic concurrency mechanism to enable high-performance concurrent accesses on multi-core systems. Experimental results on the AEP platform show that HDNH outperforms its counterparts by up to 2.9x under various YCSB workloads. Kaixin Huang, Xiaomin Zou, Nuo Xu 0001, Liang Fang 0008 |
ICPP | 2 |
| 2021 | A novel chromosome instance segmentation method based on geometry and deep learningabstractIn medicine, any abnormalities in the number of chromosomes or the structure of chromosomes may cause the newborn baby to suffer from genetic diseases, such as Edward syndrome and so on. Chromosome karyotype analysis is the most important and common method for prenatal diagnosis to determine whether a newborn baby has chromosome defects refers to segment chromosome instances from stained cell images and arrange chromosome instances according to their categories. However, due to the non-rigid nature of chromosomes, chromosome instances may overlap and adhere to each other, which makes the task of segmenting chromosome instances time-consuming and error-prone. This paper proposes a novel chromosome instance segmentation method that includes three stages. First, we segment a given stained cell image into several segments using geometric connectivity. Second, a machine learning method is proposed to distinguish chromosome of individual instances and clusters. Finally, a deep learning-based method is applied to separate chromosome instances from clusters. It shows that the proposed method achieves 97.61% instance segmentation accuracy in a hold-out clinical dataset with 162 cell images consisting of 7,452 chromosome instances, which is a promising result in clinical application. The innovation of this work is to combine geometry and deep learning to handle tasks for different stages of chromosome instance segmentation issue. The benefit of this innovation is that it can obtain a much better performance than existing geometric-based methods with a small number of training samples. Meanwhile, the segmentation performance of the proposed method is superior to existing methods fully based on deep learning. Kaixin Huang, Chengchuang Lin, Runhua Huang, Gansen Zhao, Aihua Yin, Hanbiao Chen, Li Guo 0019, Chun Shan, Ruihua Nie, Shuangyin Li |
IJCNN | 1 |
| 2021 | PMSort: An adaptive sorting engine for persistent memory
Yifan Hua, Kaixin Huang, Shengan Zheng, Linpeng Huang |
J. Syst. Archit. | 2 |
| 2021 | Lewat: A Lightweight, Efficient, and Wear-Aware Transactional Persistent Memory SystemabstractEmerging non-volatile memory (also termed as persistent memory, PM) technologies promise persistence, byte-addressability, and DRAM-like read/write latency. A proliferation of persistent memory systems have been proposed to leverage PM for fast data persistence and expose malloc-like persistent APIs. By eliminating disk I/Os, these systems gain low-latency and high-throughput access performance for persistent data. However, there still exist non-negligible limitations in these systems, such as frequent context switches, inefficient allocation, heavy logging overhead, and lack of wear-leveling techniques. To solve these problems, we develop Lewat, a lightweight, efficient, and wear-aware transactional persistent memory system. Lewat is built in user-layer to avoid kernel/user layer context switches and enables lightweight persistent data access. We decouple the data space into slot zone and page zone. Based on this, we design different allocators in these two zones to achieve efficient allocation performance for both small-sized data and large-sized data. To minimize logging overhead, we propose an efficient adaptive logging framework. The main idea is to utilize different logging techniques for different workloads. We also propose a suite of system-coupled wear-leveling techniques that contain wear-aware allocation, wear-aware update, and write reduction. We evaluate Lewat on a real non-volatile memory platform and the experimental results show that compared with state-of-the-art persistent memory systems, Lewat has much lower latency and higher throughput. Kaixin Huang, Sumin Li, Linpeng Huang, Kian-Lee Tan, Hong Mei 0001 |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2020 | Revisiting Persistent Hash Table Design for Commercial Non-Volatile MemoryabstractEmerging non-volatile memory technologies bring evolution to storage systems and durable data structures. Among them, a proliferation of researches on persistent hash table employ NVM as the storage layer for both fast access and efficient persistence. Most of them are based on the assumptions that NVM has cacheline access granularity, poor write endurance, DRAM-comparable read latency and much higher write latency. However, a commercial non-volatile memory product, named Intel Optane DC Persistent Memory (AEP), has a few interesting features that are different from previous assumptions, such as 1) block access granularity 2) hardware-layer wear-leveling and 3) much higher read latency than DRAM and DRAM-comparable write latency. Confronted with the new challenges brought by AEP, we propose Rewo-Hash, a novel read-efficient and write-optimized hash table for commercial non-volatile memory. Our design can be summarized into three key points. First, we keep a hash table copy in DRAM as a cached table to speed up search requests. Second, we design a log-free atomic mechanism to support fast writes. Third, we devise an efficient synchronization scheme between persistent table and cached table to mask the data synchronization overhead. We conduct extensive experiments using real NVM platform and the results show that compared with state-of-the-art NVM-Optimized hash tables, Rewo-Hash gains improvement of 1.73x-2.70x and 1.46x-3.11x in read latency and write latency, respectively. Rewo-Hash also outperforms its counterparts by 1.65x-4.24x in throughput for various YCSB workloads. Kaixin Huang, Linpeng Huang |
DATE | 1 |
| 2020 | CANRT: A Client-Active NVM-Based Radix Tree for Fast Remote Access
Yaoyao Ying, Kaixin Huang, Shengan Zheng, Yaofeng Tu, Linpeng Huang |
ICA3PP (1) | 2 |
| 2020 | ReoFS: A Read-Efficient and Write-Optimized File System for Persistent MemoryabstractIn this paper, we present ReoFS, a read-efficient and write-optimized PM-aware (Persistent Memory-Aware) file system that provides strong consistency guarantee. Different from traditional journaling techniques, ReoFS optimizes write requests on the write I/O path and allows concurrent reads to be served opportunistically. To reduce the write overhead for metadata-related operations, we devise a backup buffering mechanism, which keeps a replica of metadata in the backup buffer and supports fast in-place update. To minimize the write overhead for file data modifications, we design a fine-grained short logging mechanism to migrate writes to persistent log blocks and make new writes visible immediately after the log is persisted. With these two designs, ReoFS removes the double write overhead off the critical path of request execution for both metadata and data operations. We conduct extensive experiments on Intel Optane DC PM platform and the results show that ReoFS outperforms existing PM -aware file systems by 1.3x-4x. Kaixin Huang, Shengan Zheng, Dongliang Xue, Linpeng Huang |
ICECCS | 2 |
| 2020 | Quail: Using NVM write monitor to enable transparent wear-leveling
Kaixin Huang, Yijie Mei, Linpeng Huang |
J. Syst. Archit. | 1 |
| 2019 | EFLightPM: An Efficient and Lightweight Persistent Memory SystemabstractEmerging non-volatile memory (also termed as persistent memory, PM) technologies promise persistence, byte addressability and DRAM-like read/write latency. A proliferation of persistent memory systems such as Mnemosyne, NVHeaps, PMDK and HEAPO have been proposed to leverage PM for fast data persistence. However, their performance may suffer from inefficiency issues, mainly caused by kernel/user layer context switches and heavy transaction logging overhead. Concretely, getting a persistent region in Mnemosyne, NV-Heaps and PMDK needs two kernel/user layer context switches since the mmap-like system calls are used, which leads to high latency. To guarantee data consistency, existing systems employ redo or undo logging techniques but they bring non-negligible overhead due to double writes and persistence ordering. In this paper, we develop EFlightPM, an efficient and lightweight persistent memory system to manage data in a fine-grained style. We decouple the data organization for persistent regions by placing large regions in the kernel layer while exposing small regions in the user layer. We also design a lightweight transaction mechanism using hybrid logging with high efficiency by minimizing the writes in the critical path. The experimental results show that compared with state-of-the-art persistent memory systems, EFlightPM manipulates fine-grained persistent data with less persistent region operation overhead and more transaction throughput. Kaixin Huang, Linpeng Huang |
ICECCS | 1 |
| 2019 | Spindle: A Write-Optimized NVM Cache for Journaling File System
Kaixin Huang, Linpeng Huang |
NPC | 2 |
| 2019 | LiwePMS: A Lightweight Persistent Memory with Wear-aware Memory ManagementabstractNext-generation Storage Class Memory (SCM) offers low-latency, high-density, byte-addressable access and persistency. The potent combination of these attractive characteristics makes it possible for SCM to unify the main memory and storage to reduce the storage hierarchy. Aiming for this, several persistent memory systems were designed. However, the heavy metadata and transaction cost degrade the system performance. Moreover, neither of them pays attention to wear-leveling strategy. In this article, we present a lightweight persistent memory system, LiwePMS, which allows a fast access to persistent data stored in SCM with wear-aware memory management. LiwePMS makes performance improvement by simplifying the metadata management and the consistency method. LiwePMS abstracts SCM as heap space with container-based dynamic address mapping. Also, LiwePMS implements efficient wear-aware dynamic memory allocator and lightweight transaction mechanism for data consistency in user-space library. The experiments showed that LiwePMS persists key-value records 1.5× faster than Redis RDB mechanism. LiwePMS improves the performance of persistent region operation by more than 45%, 63%, and 1.1× comparing with HEAPO, Mnemosyne, and NVML, respectively. Also, the wear-leveling policy of memory allocator outperforms that of NVMalloc from 35% to 30%, and the transaction method promotes the transaction performance to 1.8× compared to NVML. Sumin Li, Kaixin Huang, Linpeng Huang, Jiashun Zhu |
ACM J. Emerg. Technol. Comput. Syst. | 2 |
| 2018 | An Adaptive Eviction Framework for Anti-caching Based In-Memory Databases
Kaixin Huang, Shengan Zheng, Yanyan Shen, Yanmin Zhu 0006, Linpeng Huang |
DASFAA (2) | 1 |
| 2018 | Forca: Fast and Atomic Remote Direct Access to Persistent MemoryabstractFor promising performance boost, recent trends of modern data centers tend to use Persistent Memory (PM) as storage and utilize the one-sided feature of Remote Direct Memory Access (RDMA) to directly access PM for I/O requests. However, accessing PM through one-sided RDMA faces two challenges: one-sided data races and remote data crash consistency. Existing systems that employ server-bypass mechanism either abandon server-bypass write to avoid the challenges at the cost of extra server loads, or support both server-bypass read/write through inefficient concurrency control without ensuring remote crash consistency. In this paper, we propose a novel server-bypass RDMA-to-PM framework, named Forca, to provide high concurrency and guarantee remote crash consistency at the same time. In Forca, we design an optimized log-structured mechanism to eliminate in-place updates, removing race conditions of one-sided RDMA and providing atomicity for each update simultaneously. We implement Forca as a generic module to support server-bypass RDMA to PM, and conduct experiments on a Forca-based key-value store called ForcaKV. The experiments show that Forca achieves higher throughput than state-of-the-art techniques by up to 1.4x under high concurrency scenarios, and its crash consistency mechanism incurs only 14.7%-19.1% time overhead. Haixin Huang, Kaixin Huang, Litong You, Linpeng Huang |
ICCD | 2 |
| 2018 | Statistical Monitoring for NVM WriteabstractEmerging non-volatile memory (NVM) technologies promise persistence, byte-addressability, low power consumption and high density. The potential of NVM implies that it can improve persistence performance and availability for high-performance computing tasks in data centers. However, NVM suffers from limited write endurance; skewed writes can extremely curtail the lifetime of NVM. Although several wear-leveling algorithms have been proposed to address the wear-out problem, based on estimation of write counts, they come with limitations in accuracy and compatibility with existing hardware and softwares. In this paper, we propose several software-based ways of monitoring NVM writes and obtaining the write counts, which are more accurate and non-intrusive for existing programs. We compare our proposed monitoring approaches in the aspects of granularity and portability; we conduct extensive experiments to evaluate them and the results show that our proposed methods can obtain good accuracy with an affordable time overhead. Yijie Mei, Kaixin Huang, Yanmin Zhu 0006, Linpeng Huang |
ICPADS | 2 |
| 2018 | NVHT: An efficient key-value storage library for non-volatile memory
Kaixin Huang, Linpeng Huang, Yanyan Shen |
J. Parallel Distributed Comput. | 1 |