EDBT 2026 Demo / reviewers in the wild / expert
Kaibo Wang
dblp:53/2141
· DBLP profile ↗
29ranked-venue papers
8as first author
11since 2021 · last 2026
0000-0001-9888-4323ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 8 · 4 first-author · 2 since 2021Artificial intelligence and machine learning · 7 · 2 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 1 first-author · 1 since 2021Systems, architecture and hardware · 5 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 3 · 1 since 2021Software engineering, systems software and programming languages · 2 · 1 first-author · 1 since 2021Security and privacy · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
4 papers |
Generative modeling · 92% Optimization for machine learning · 5% Speech recognition and synthesis · 3% | |
| Network and information security
2 papers |
Security and privacy of machine learning · 100% | |
| Computer architecture, parallel and distributed computing, and storage systems
8 papers |
Storage systems · 38% GPUs and heterogeneous computing · 23% Memory systems · 19% | |
| Databases, data mining, and information retrieval
4 papers |
Data integration and cleaning · 30% Indexing and storage engines · 30% Transaction processing and concurrency control · 22% | |
| Computer graphics and multimedia
3 papers |
Visual content generation and editing · 84% Image and video processing · 12% Image and video coding · 4% | |
| Computer networks
3 papers |
Internet of things and sensor networks · 95% Cellular and mobile networks · 4% Transport protocols and congestion control · 2% |
Topics — the 30 heaviest of 45, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Generative modeling
diffusion model |
2.1 | 3 | 2026 | TweezeEdit: Consistent and Efficient Image Editing with Path Regularization · AAAI 2026 Towards a Golden Classifier-Free Guidance Path via Foresight Fixed Point Iterations · NeurIPS 2025 DiffHammer: Rethinking the Robustness of Diffusion-Based Adversarial Purification · NeurIPS 2024 |
Security and privacy of machine learning
adversarial attack |
1.3 | 2 | 2024 | DiffHammer: Rethinking the Robustness of Diffusion-Based Adversarial Purification · NeurIPS 2024 FakeWake: Understanding and Mitigating Fake Wake-up Words of Voice Assistants · CCS 2021 |
Machine learning › Generative modeling › diffusion model › image editing
training-free image editing |
1.0 | 1 | 2026 | TweezeEdit: Consistent and Efficient Image Editing with Path Regularization · AAAI 2026 |
Visual content generation and editing
image editing |
1.0 | 1 | 2026 | TweezeEdit: Consistent and Efficient Image Editing with Path Regularization · AAAI 2026 |
Visual content generation and editing › image editing
text-guided image editing |
1.0 | 1 | 2026 | TweezeEdit: Consistent and Efficient Image Editing with Path Regularization · AAAI 2026 |
Machine learning › Generative modeling › diffusion model › guided diffusion
classifier-free guidance |
0.9 | 1 | 2025 | Towards a Golden Classifier-Free Guidance Path via Foresight Fixed Point Iterations · NeurIPS 2025 |
Machine learning › Generative modeling › diffusion model
text-to-image generation |
0.9 | 1 | 2025 | Towards a Golden Classifier-Free Guidance Path via Foresight Fixed Point Iterations · NeurIPS 2025 |
Security and privacy of machine learning
adversarial defense |
0.8 | 1 | 2024 | DiffHammer: Rethinking the Robustness of Diffusion-Based Adversarial Purification · NeurIPS 2024 |
Security and privacy of machine learning › adversarial defense
diffusion-based purification |
0.8 | 1 | 2024 | DiffHammer: Rethinking the Robustness of Diffusion-Based Adversarial Purification · NeurIPS 2024 |
Storage systems
key-value storage |
0.5 | 2 | 2017 | A distributed in-memory key-value store system on heterogeneous CPU-GPU cluster · VLDB J. 2017 Mega-KV: A Case for GPUs to Maximize the Throughput of In-Memory Key-Value Stores · Proc. VLDB Endow. 2015 |
Image and video processing
image restoration |
0.3 | 1 | 2026 | TweezeEdit: Consistent and Efficient Image Editing with Path Regularization · AAAI 2026 |
Storage systems › key-value storage
distributed in-memory key-value store |
0.3 | 1 | 2017 | A distributed in-memory key-value store system on heterogeneous CPU-GPU cluster · VLDB J. 2017 |
Machine learning › Optimization for machine learning
fixed-point iteration |
0.3 | 1 | 2025 | Towards a Golden Classifier-Free Guidance Path via Foresight Fixed Point Iterations · NeurIPS 2025 |
Transaction processing and concurrency control
concurrency control |
0.2 | 1 | 2016 | BCC: Reducing False Aborts in Optimistic Concurrency Control with Low Cost for In-Memory Databases · Proc. VLDB Endow. 2016 |
Transaction processing and concurrency control › concurrency control
optimistic concurrency control |
0.2 | 1 | 2016 | BCC: Reducing False Aborts in Optimistic Concurrency Control with Low Cost for In-Memory Databases · Proc. VLDB Endow. 2016 |
Memory systems
cache management |
0.2 | 2 | 2011 | ULCC: a user-level facility for optimizing shared cache performance on multicores · PPoPP 2011 SRM-buffer: an OS buffer management technique to prevent last level cache from thrashing in multicores · EuroSys 2011 |
Storage systems › key-value storage
in-memory key-value store |
0.2 | 1 | 2015 | Mega-KV: A Case for GPUs to Maximize the Throughput of In-Memory Key-Value Stores · Proc. VLDB Endow. 2015 |
Query processing and optimization
analytical query processing |
0.2 | 1 | 2014 | Concurrent Analytical Query Processing with GPUs · Proc. VLDB Endow. 2014 |
GPUs and heterogeneous computing
GPU computing |
0.2 | 1 | 2014 | GDM: device memory management for gpgpu computing · SIGMETRICS 2014 |
GPUs and heterogeneous computing › GPU query processing
GPU query engine |
0.2 | 1 | 2014 | Concurrent Analytical Query Processing with GPUs · Proc. VLDB Endow. 2014 |
Natural language and speech › Speech recognition and synthesis › automatic speech recognition
keyword spotting |
0.1 | 1 | 2021 | FakeWake: Understanding and Mitigating Fake Wake-up Words of Voice Assistants · CCS 2021 |
GPUs and heterogeneous computing
CPU-GPU heterogeneous computing |
0.1 | 1 | 2012 | Accelerating Pathology Image Data Cross-Comparison on CPU-GPU Hybrid Systems · Proc. VLDB Endow. 2012 |
Parallel and multicore computing
task scheduling |
0.1 | 1 | 2012 | BWS: balanced work stealing for time-sharing multicores · EuroSys 2012 |
Parallel and multicore computing › load balancing › dynamic load balancing
work stealing |
0.1 | 1 | 2012 | BWS: balanced work stealing for time-sharing multicores · EuroSys 2012 |
Operating systems › resource management › memory management
buffer cache |
0.1 | 1 | 2011 | SRM-buffer: an OS buffer management technique to prevent last level cache from thrashing in multicores · EuroSys 2011 |
Memory systems › cache management › cache partitioning
last-level cache partitioning |
0.1 | 1 | 2011 | ULCC: a user-level facility for optimizing shared cache performance on multicores · PPoPP 2011 |
High-performance computing
performance optimization |
0.1 | 1 | 2011 | ULCC: a user-level facility for optimizing shared cache performance on multicores · PPoPP 2011 |
GPUs and heterogeneous computing › heterogeneous cluster computing
heterogeneous CPU-GPU cluster |
0.1 | 1 | 2017 | A distributed in-memory key-value store system on heterogeneous CPU-GPU cluster · VLDB J. 2017 |
Database system architecture and tuning
main-memory database |
0.1 | 1 | 2016 | BCC: Reducing False Aborts in Optimistic Concurrency Control with Low Cost for In-Memory Databases · Proc. VLDB Endow. 2016 |
Transaction processing and concurrency control › transaction performance
transaction throughput |
0.1 | 1 | 2016 | BCC: Reducing False Aborts in Optimistic Concurrency Control with Low Cost for In-Memory Databases · Proc. VLDB Endow. 2016 |
Methods — techniques the papers use, named apart from their topics
path regularization · 2.0diffusion model · 2.0consistency model · 2.0gradient grafting · 1.5EOT-based attack · 1.5interpretable tree-based decision model · 1.5foresight guidance · 0.9fixed-point iteration · 0.9indexing · 0.8compression · 0.8device memory swapping · 0.4GPU query scheduling · 0.4task migration · 0.3pipelining · 0.3two-phase locking · 0.2dependency pattern detection · 0.2cache partitioning · 0.2buffer replacement policy · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | TweezeEdit: Consistent and Efficient Image Editing with Path RegularizationabstractRecent progress in training-free image editing has enabled existing text-to-image diffusion models to be directly adapted into text-guided image editors without additional training. However, existing methods often over-align with target prompts while inadequately preserving source image semantics. These approaches generate target images explicitly or implicitly from the inversion noise of the source images, termed the inversion anchors. We identify this strategy as suboptimal for semantic preservation and inefficient due to elongated editing paths. We propose TweezeEdit, a tuning- and inversion-free framework for consistent and efficient image editing. Our method addresses these limitations by regularizing the entire denoising path rather than relying solely on the inversion anchors, ensuring source semantic retention and shortening editing paths. Guided by gradient-driven regularization, we efficiently inject target prompt semantics along a direct path using a consistency model. Extensive experiments demonstrate TweezeEdit's superior performance in semantic preservation and target alignment, outperforming existing methods. Remarkably, it requires only 12 steps (1.6 seconds per edit), underscoring its potential for real-time applications. The appendix is available in the extended version. Jianda Mao, Kaibo Wang, Kani Chen |
AAAI | 2 |
| 2026 | Hamiltonian monte carlo based neural process for few-shot knowledge graph completion
Kaibo Wang, Jinguang Chen |
Inf. Sci. | 1 |
| 2026 | Multi-Stage Robust Federated Learning: Addressing Label Noise under Data Heterogeneity and ImbalanceabstractFederated Learning (FL) enables collaborative model training while preserving data privacy, but the presence of noisy labels in local datasets remains a significant challenge, particularly under heterogeneous noise conditions and class imbalance. In this work, we introduce a novel Multi-Stage Robust Federated Learning (MRFL) framework to address these issues. In the warm-up noise detection stage, MRFL computes per-class average losses on each client and employs a Gaussian mixture model to accurately identify clients with substantial label noise. In the subsequent noise-robust training stage, a robust loss function and noise solver are designed to distinguish clean from noisy samples, while semi-supervised learning is used to recover valuable information from tail classes. Moreover, a robust weighted aggregation strategy is adopted to mitigate the adverse effects of noisy clients. Extensive experiments on CIFAR-10/100-LT and ICH datasets demonstrate that MRFL outperforms state-of-the-art methods in federated noisy label learning scenarios characterized by data heterogeneity and imbalance. Kaibo Wang, Anqi Zhang 0001, Tangyou Liu, Wenqian Zhang 0003, Guanglin Zhang |
ACM Trans. Intell. Syst. Technol. | 1 |
| 2025 | Towards a Golden Classifier-Free Guidance Path via Foresight Fixed Point IterationsabstractClassifier-Free Guidance (CFG) is an essential component of text-to-image diffusion models, and understanding and advancing its operational mechanisms remains a central focus of research. Existing approaches stem from divergent theoretical interpretations, thereby limiting the design space and obscuring key design choices. To address this, we propose a unified perspective that reframes conditional guidance as fixed point iterations, seeking to identify a golden path where latents produce consistent outputs under both conditional and unconditional generation. We demonstrate that CFG and its variants constitute a special case of single-step short-interval iteration, which is theoretically proven to exhibit inefficiency. To this end, we introduce Foresight Guidance (FSG), which prioritizes solving longer-interval subproblems in early diffusion stages with increased iterations. Extensive experiments across diverse datasets and model architectures validate the superiority of FSG over state-of-the-art methods in both image quality and computational efficiency. Our work offers novel perspectives for conditional guidance and unlocks the potential of adaptive design. Kaibo Wang, Jianda Mao |
NeurIPS | 1 |
| 2025 | Ship re-identification in foggy weather: A two-branch network with dynamic feature enhancement and dual attention
Wei Sun 0012, Kaibo Wang |
Eng. Appl. Artif. Intell. | 5 |
| 2024 | DiffHammer: Rethinking the Robustness of Diffusion-Based Adversarial PurificationabstractDiffusion-based purification has demonstrated impressive robustness as an adversarial defense. However, concerns exist about whether this robustness arises from insufficient evaluation. Our research shows that EOT-based attacks face gradient dilemmas due to global gradient averaging, resulting in ineffective evaluations. Additionally, 1-evaluation underestimates resubmit risks in stochastic defenses. To address these issues, we propose an effective and efficient attack named DiffHammer. This method bypasses the gradient dilemma through selective attacks on vulnerable purifications, incorporating $N$-evaluation into loops and using gradient grafting for comprehensive and efficient evaluations. Our experiments validate that DiffHammer achieves effective results within 10-30 iterations, outperforming other methods. This calls into question the reliability of diffusion-based purification after mitigating the gradient dilemma and scrutinizing its resubmit risk. Kaibo Wang, Xiaowen Fu |
NeurIPS | 1 |
| 2024 | μSlope: High Compression and Fast Search on Semi-Structured Logs
Devin Gibson, Kirk Rodrigues, Yu Luo 0006, Kaibo Wang, Yupeng Fu, Ding Yuan 0004 |
OSDI | 6 |
| 2024 | Coupled Epidemic-Information Propagation With Stranding Mechanism on Multiplex Metapopulation NetworksabstractAcknowledging the significance of information propagation and individual adaptive behavior has been regarded as an indispensable prerequisite for a complete understanding of epidemic spreading. Recent studies have widely considered the metapopulation model, where epidemics spread over a single layer of physical networks via individual mobility. However, these advances neglected the interventions of accompanied information and individual behavior response related to epidemics. In this article, we develop a coupled epidemic-information propagation model on multiplex metapopulation networks leveraging the microscopic Markov chain (MMC) approach, aiming to explore the spatiotemporal characteristics of epidemic spreading process. Taking the individual adaptive behavior into account, the stranding mechanism based on infection level and medical resources is introduced to capture the population size dynamics during individual mobility among different patches. Theoretical epidemic threshold is analytically derived under the improved framework. Extensive numerical simulations are performed to validate our theoretical analysis and further examine the impacts of information propagation and spreading parameters on epidemic threshold and steady-state prevalence. Our results indicate that both the scale of information diffusion and the specific configuration of spreading parameters can significantly suppress the epidemic prevalence. These findings shed a novel light on theoretical research and decision-making of coupled epidemic-information process in the spatiotemporal perspective. Xuming An 0002, Chen Zhang 0007, Lin Hou 0003, Kaibo Wang |
IEEE Trans. Comput. Soc. Syst. | 4 |
| 2023 | Parallel crosschecking neural network based fault-tolerant flight parameter estimation and faulty sensor identification
Wanyong Zou, Ban Wang, Kaibo Wang, Shuhui Bu, He Shen 0001 |
Eng. Appl. Artif. Intell. | 4 |
| 2021 | FakeWake: Understanding and Mitigating Fake Wake-up Words of Voice AssistantsabstractIn the area of Internet of Things (IoT), voice assistants have become an important interface to operate smart speakers, smartphones, and even automobiles. To save power and protect user privacy, voice assistants send commands to the cloud only if a small set of preregistered wake-up words are detected. However, voice assistants are shown to be vulnerable to the FakeWake phenomena, whereby they are inadvertently triggered by innocent-sounding fuzzy words. In this paper, we present a systematic investigation of the FakeWake phenomena from three aspects. To start with, we design the first fuzzy word generator to automatically and efficiently produce fuzzy words instead of searching through a swarm of audio materials.We manage to generate 965 fuzzy words covering 8 most popular English and Chinese smart speakers. To explain the causes underlying the FakeWake phenomena, we construct an interpretable tree-based decision model, which reveals phonetic features that contribute to false acceptance of fuzzy words by wake-up word detectors. Finally, we propose remedies to mitigate the effect of FakeWake. The results show that the strengthened models are not only resilient to fuzzy words but also achieve better overall performance on original training datasets. Yanjiao Chen, Yijie Bai, Richard Mitev, Kaibo Wang, Ahmad-Reza Sadeghi, Wenyuan Xu 0001 |
CCS | 4 |
| 2021 | Joint distribution adaptation network with adversarial learning for rolling bearing fault diagnosis
Hongkai Jiang, Kaibo Wang, Zeyu Pei |
Knowl. Based Syst. | 3 |
| 2017 | A distributed in-memory key-value store system on heterogeneous CPU-GPU cluster
Kai Zhang 0006, Kaibo Wang, Yuan Yuan 0014, Lei Guo 0004, Rubao Li, Xiaodong Zhang 0001, Bingsheng He, Bei Hua |
VLDB J. | 2 |
| 2016 | Spark-GPU: An accelerated in-memory data processing engine on clustersabstractApache Spark is an in-memory data processing system that supports both SQL queries and advanced analytics over large data sets. In this paper, we present our design and implementation of Spark-GPU that enables Spark to utilize GPU's massively parallel processing ability to achieve both high performance and high throughput. Spark-GPU transforms a general-purpose data processing system into a GPU-supported system by addressing several real-world technical challenges including minimizing internal and external data transfers, preparing a suitable data format and a batching mode for efficient GPU execution, and determining the suitability of workloads for GPU with a task scheduling capability between CPU and GPU. We have comprehensively evaluated Spark-GPU with a set of representative analytical workloads to show its effectiveness. Our results show that Spark-GPU improves the performance of machine learning workloads by up to 16.13x and the performance of SQL queries by up to 4.83x. Yuan Yuan 0014, Meisam Fathi Salmi, Yin Huai, Kaibo Wang, Rubao Lee, Xiaodong Zhang 0001 |
IEEE BigData | 4 |
| 2016 | BCC: Reducing False Aborts in Optimistic Concurrency Control with Low Cost for In-Memory DatabasesabstractThe Optimistic Concurrency Control (OCC) method has been commonly used for in-memory databases to ensure transaction serializability --- a transaction will be aborted if its read set has been changed during execution. This simple criterion to abort transactions causes a large proportion of false positives, leading to excessive transaction aborts. Transactions aborted false-positively (i.e. false aborts) waste system resources and can significantly degrade system throughput (as much as 3.68x based on our experiments) when data contention is intensive. Modern in-memory databases run on systems with increasingly parallel hardware and handle workloads with growing concurrency. They must efficiently deal with data contention in the presence of greater concurrency by minimizing false aborts. This paper presents a new concurrency control method named Balanced Concurrency Control (BCC) which aborts transactions more carefully than OCC does. BCC detects data dependency patterns which can more reliably indicate unserializable transactions than the criterion used in OCC. The paper studies the design options and implementation techniques that can effectively detect data contention by identifying dependency patterns with low overhead. To test the performance of BCC, we have implemented it in Silo and compared its performance against that of the vanilla Silo system with OCC and two-phase locking (2PL). Our extensive experiments with TPC-W-like, TPC-C-like and YCSB workloads demonstrate that when data contention is intensive, BCC can increase transaction throughput by more than 3x versus OCC and more than 2x versus 2PL; meanwhile, BCC has comparable performance with OCC for workloads with low data contention. Yuan Yuan 0014, Kaibo Wang, Rubao Lee, Xiaoning Ding, Spyros Blanas, Xiaodong Zhang 0001 |
Proc. VLDB Endow. | 2 |
| 2016 | A Spatial Calibration Model for Nanotube Film Quality PredictionabstractA carbon nanotube (CNT) film, which is drawn from a CNT array, is a spatially distributed thin film with unique and appealing properties. Novel devices have been developed based on CNT films. The anisotropy of a CNT film, which is a spatially distributed quality index, is difficult to measure in practice due to metrology and cost constraints. As the anisotropy is highly correlated with the height of the CNT array and the height can be measured in a much easier and more cost-effective way, we propose a spatial model for predicting the anisotropy using the height. The model takes the spatially distributed two-dimensional (2-D) height as an input and provides a predicted anisotropy distribution in a 2-D space. If the anisotropy measures are obtained, the model can provide a more accurate prediction. The performance of the proposed model is verified by both a simulation study and real data samples. Note to Practitioners-Timely and accurate measurement of key product features is essential in scale-up nanomanufacturing processes. Even though a fast growth of metrology technology has been seen in recent years, some variables of nanoscale products are still hard to measure, either too costly or too time consuming, in highspeed large-scale production. However, physical mechanisms may suggest that a hard-to-measure variable may be correlated with another easy-to-measure variable. In such a case, a spatial calibration model could be constructed, based on which the prediction of the hard-to-measure variable is achievable given measures of the easyto-measure variable. Such a calibration model provides an effective alternative to physical metrology tools in large-scale nanomanufacturing processes in which metrology technology is not fully ready yet. Su Wu, Kaibo Wang, Xinwei Deng |
IEEE Trans Autom. Sci. Eng. | 3 |
| 2015 | Hetero-DB: Next Generation High-Performance Database Systems by Best Utilizing Heterogeneous Computing and Storage Resources
Kai Zhang 0006, Feng Chen 0005, Xiaoning Ding, Yin Huai, Rubao Lee, Kaibo Wang, Yuan Yuan 0014, Xiaodong Zhang 0001 |
J. Comput. Sci. Technol. | 7 |
| 2015 | Mega-KV: A Case for GPUs to Maximize the Throughput of In-Memory Key-Value StoresabstractIn-memory key-value stores play a critical role in data processing to provide high throughput and low latency data accesses. In-memory key-value stores have several unique properties that include (1) data intensive operations demanding high memory bandwidth for fast data accesses, (2) high data parallelism and simple computing operations demanding many slim parallel computing units, and (3) a large working set. As data volume continues to increase, our experiments show that conventional and general-purpose multicore systems are increasingly mismatched to the special properties of key-value stores because they do not provide massive data parallelism and high memory bandwidth; the powerful but the limited number of computing cores do not satisfy the demand of the unique data processing task; and the cache hierarchy may not well benefit to the large working set. In this paper, we make a strong case for GPUs to serve as special-purpose devices to greatly accelerate the operations of in-memory key-value stores. Specifically, we present the design and implementation of Mega-KV, a GPU-based in-memory key-value store system that achieves high performance and high throughput. Effectively utilizing the high memory bandwidth and latency hiding capability of GPUs, Mega-KV provides fast data accesses and significantly boosts overall performance. Running on a commodity PC installed with two CPUs and two GPUs, Mega-KV can process up to 160+ million key-value operations per second, which is 1.4-2.8 times as fast as the state-of-the-art key-value store system on a conventional CPU-based platform. Kai Zhang 0006, Kaibo Wang, Yuan Yuan 0014, Lei Guo 0004, Rubao Lee, Xiaodong Zhang 0001 |
Proc. VLDB Endow. | 2 |
| 2015 | A Run-to-Run Profile Control Algorithm for Improving the Flatness of Nano-Scale ProductsabstractIn scaling-up Carbon Nanotubes (CNTs) array manufacturing, the uniformity of CNTs' height, or flatness of array, is critical for the yield of nanodevices fabricated from CNTs, and thus needs to be properly controlled. However, since the flatness of the CNTs array is better characterized by a profile, the conventional run-to-run (R2R) controllers that are designed for a single or multiple quality indicators are not effective in controlling the CNTs array manufacturing process. Therefore, in this work, we first develop a statistical model to characterize the variation of the flatness profile, and then derive a novel R2R profile controller based on a state-space model and Kalman filter to improve the flatness of CNTs array. The performance of the proposed R2R control algorithm is studied and compared with existing controller via simulation studies. Su Wu, Kaibo Wang |
IEEE Trans Autom. Sci. Eng. | 3 |
| 2014 | GDM: device memory management for gpgpu computingabstractGPGPUs are evolving from dedicated accelerators towards mainstream commodity computing resources. During the transition, the lack of system management of device memory space on GPGPUs has become a major hurdle. In existing GPGPU systems, device memory space is still managed explicitly by individual applications, which not only increases the burden of programmers but can also cause application crashes, hangs, or low performance. Kaibo Wang, Xiaoning Ding, Rubao Lee, Shinpei Kato, Xiaodong Zhang 0001 |
SIGMETRICS | 1 |
| 2014 | Concurrent Analytical Query Processing with GPUsabstractIn current databases, GPUs are used as dedicated accelerators to process each individual query. Sharing GPUs among concurrent queries is not supported, causing serious resource underutilization. Based on the profiling of an open-source GPU query engine running commonly used single-query data warehousing workloads, we observe that the utilization of main GPU resources is only up to 25%. The underutilization leads to low system throughput. To address the problem, this paper proposes concurrent query execution as an effective solution. To efficiently share GPUs among concurrent queries for high throughput, the major challenge is to provide software support to control and resolve resource contention incurred by the sharing. Our solution relies on GPU query scheduling and device memory swapping policies to address this challenge. We have implemented a prototype system and evaluated it intensively. The experiment results confirm the effectiveness and performance advantage of our approach. By executing multiple GPU queries concurrently, system throughput can be improved by up to 55% compared with dedicated processing. Kaibo Wang, Kai Zhang 0006, Yuan Yuan 0014, Rubao Lee, Xiaoning Ding, Xiaodong Zhang 0001 |
Proc. VLDB Endow. | 1 |
| 2012 | BWS: balanced work stealing for time-sharing multicoresabstractRunning multithreaded programs in multicore systems has become a common practice for many application domains. Work stealing is a widely-adopted and effective approach for managing and scheduling the concurrent tasks of such programs. Existing work-stealing schedulers, however, are not effective when multiple applications time-share a single multicore---their management of steal-attempting threads often causes unbalanced system effects that hurt both workload throughput and fairness. Xiaoning Ding, Kaibo Wang, Phillip B. Gibbons, Xiaodong Zhang 0001 |
EuroSys | 2 |
| 2012 | Accelerating Pathology Image Data Cross-Comparison on CPU-GPU Hybrid SystemsabstractAs an important application of spatial databases in pathology imaging analysis, cross-comparing the spatial boundaries of a huge amount of segmented micro-anatomic objects demands extremely data- and compute-intensive operations, requiring high throughput at an affordable cost. However, the performance of spatial database systems has not been satisfactory since their implementations of spatial operations cannot fully utilize the power of modern parallel hardware. In this paper, we provide a customized software solution that exploits GPUs and multi-core CPUs to accelerate spatial cross-comparison in a cost-effective way. Our solution consists of an efficient GPU algorithm and a pipelined system framework with task migration support. Extensive experiments with real-world data sets demonstrate the effectiveness of our solution, which improves the performance of spatial cross-comparison by over 18 times compared with a parallelized spatial database approach. Kaibo Wang, Yin Huai, Rubao Lee, Fusheng Wang 0001, Xiaodong Zhang 0001, Joel H. Saltz |
Proc. VLDB Endow. | 1 |
| 2011 | SRM-buffer: an OS buffer management technique to prevent last level cache from thrashing in multicoresabstractBuffer caches in operating systems keep active file blocks in memory to reduce disk accesses. Related studies have been focused on how to minimize buffer misses and the caused performance degradation. However, the side effects and performance implications of accessing the data in buffer caches (i.e. buffer cache hits) have not been paid attention. In this paper, we show that accessing buffer caches can cause serious performance degradation on multicores, particularly with shared last level caches (LLCs). There are two reasons for this problem. First, data in files normally have weaker localities than data objects in virtual memory spaces. Second, due to the shared structure of LLCs on multicore processors, an application accessing the data in a buffer cache may flush the to-be-reused data of its co-running applications from the shared LLC and significantly slow down these applications. Xiaoning Ding, Kaibo Wang, Xiaodong Zhang 0001 |
EuroSys | 2 |
| 2011 | ULCC: a user-level facility for optimizing shared cache performance on multicoresabstractScientific applications face serious performance challenges on multicore processors, one of which is caused by access contention in last level shared caches from multiple running threads. The contention increases the number of long latency memory accesses, and consequently increases application execution times. Optimizing shared cache performance is critical to reduce significantly execution times of multi-threaded programs on multicores. However, there are two unique problems to be solved before implementing cache optimization techniques on multicores at the user level. First, available cache space for each running thread in a last level cache is difficult to predict due to access contention in the shared space, which makes cache conscious algorithms for single cores ineffective on multicores. Second, at the user level, programmers are not able to allocate cache space at will to running threads in the shared cache, thus data sets with strong locality may not be allocated with sufficient cache space, and cache pollution can easily happen. To address these two critical issues, we have designed ULCC (User Level Cache Control), a software runtime library that enables programmers to explicitly manage and optimize last level cache usage by allocating proper cache space for different data sets of different threads. We have implemented ULCC at the user level based on a page-coloring technique for last level cache usage management. By means of multiple case studies on an Intel multicore processor, we show that with ULCC, scientific applications can achieve significant performance improvements by fully exploiting the benefit of cache optimization algorithms and by partitioning the cache space accordingly to protect frequently reused data sets and to avoid cache pollution. Our experiments with various applications show that ULCC can significantly improve application performance by nearly 40%. Xiaoning Ding, Kaibo Wang, Xiaodong Zhang 0001 |
PPoPP | 2 |
| 2009 | An Adaptive Resource Monitoring Method for Distributed Heterogeneous Computing EnvironmentabstractResource performance monitoring is among the most active research topics in distributed computing. In this paper, we propose an adaptive resource monitoring method for applications in heterogeneous computing environment. According to the operating environment of distributed heterogeneous system and the changes of system resource workload, the method combines periodic pull mode with event-driven push mode to adaptively publish and retrieve system resource information. Preliminary experiments reveal that, by using our adaptive monitoring method, the efficiency of system monitoring is improved over that accrued by using regular monitoring approaches. Gang Yang 0008, Kaibo Wang, Xingshe Zhou 0001 |
ISPA | 2 |
| 2007 | A Review of Reliability Research on NanotechnologyabstractNano-reliability measures the ability of a nano-scaled product to perform its intended functionality. At the nano scale, the physical, chemical, and biological properties of materials differ in fundamental, valuable ways from the properties of individual atoms, molecules, or bulk matter. Conventional reliability theories need to be restudied to be applied to nano-engineering. Research on nano-reliability is extremely important due to the fact that nano-structure components account for a high proportion of costs, and serve critical roles in newly designed products. This review introduces the concepts of reliability to nano-technology; and presents the current work on identifying various physical failure mechanisms of nano-structured materials, and devices during fabrication process, and operation. Modeling techniques of degradation, reliability functions, and failure rates of nano-systems are also reviewed in this work. Shuen-Lin Jeng, Jye-Chyi Lu, Kaibo Wang |
IEEE Trans. Reliab. | 3 |
| 2006 | Monitoring Multivariate Processes Using an Adaptive T2 ChartabstractIn recent years, there has been an increasing demand for quality control of multivariate dynamic systems. Conventional directionally invariant charts fail to make use of versatile shift patterns in a multivariate process and are sensitive to general failures only. Directionally variant charts, however, are designed for specific and constant shifts and are not suitable for processes with dynamic and unknown failures. This paper proposes an adaptive T2scheme, which can successfully capture the unknown shift patterns of a multivariate system via an exponentially weighted moving average (EWMA) forecasting procedure. The adaptive scheme preserves the optimality of a directionally variant chart, while provides a scalable extension to Hotelling's T2procedure. The smoothing parameter of the proposed scheme can be tuned for desired shift sizes. Significant improvement of sensitivity over an intended range is demonstrated by Monte Carlo simulation. Kaibo Wang, Fugee Tsung |
SMC | 1 |
| 2001 | Portrait video phoneabstractThe rapid development of wired and wireless networks tremendouslyfacilitates communications between people. However, most of thecurrent wireless networks still work in low bandwidths, and mobiledevices still suffer from weak computational power, short batterylifetime and limited display capability. We developed a very lowbit-rate bi-level video coding technique, which can be used invideo communications almost anywhere, anytime on any device. Thespirit of this method is that rather than giving highest priorityto the basic colors of an image as in conventional DCT-basedcompression methods, we give preference to the outline features ofscenes when we have limited bandwidths. These features can berepresented by bi-level image sequences that are converted fromgray-scale image sequences. By analyzing the temporal correlationbetween successive frames and flexibilities in the scenepresentation using bi-level images, we achieve very high ratioswith our bi-level video compression scheme. Experiments show thatin low bandwidths, our method provides clearer shape, smoothermotion, shorter initial latency and much cheaper computational costthan do DCT-based methods. Our method is especially suitable forsmall mobile devices such as handheld PCs, palm-size PCs and mobilephones that possess small display screens and light computationalpower, and work in low bandwidth wireless networks. We have builtPC and Pocket PC versions of bi-level video phone systems, whichtypically provide QCIF-size video with a frame rate of 5-15 fps fora 9.6 Kbps bandwidth. Jiang Li 0008, Keman Yu, Harry Shum, Jizheng Xu, Hanning Zhou, King To Ng, Kaibo Wang |
ACM Multimedia | 9 |
| 2001 | Portrait video phoneabstractAs the Internet and wirless networks are developed rapidly, the demand of communicating anywhere, anytime on any device emerges. However, most of the current wireless networks still work in low bandwidths, and mobile devices still suffer from weak computational power, short battery lifetime and limited display capability. We developed portrait video phone systems that can run on Pcs and Pocket Pcs at very low bit rates through the Internet. The core technology that portrait video phones employ is the so-called portrait video (or bi-level video) codec. Portrait video codec first converts a full-color video into a black/white image sequence and then compresses it into a black/white portrait-like video. Portrait video processes clearer shape, smoother motion, shorter initial latency, and cheaper computational cost than MPEG2, MPEG4 and H.263 for low bandwidths. Typically the portrait video phone provides QCIF-size video with a frame rate of 5-15 fps for a 9.6 Kbps video bandwidth. The portrait video is so small that it can even be transmitted through an HTTP proxy as text. Experiments show that the portrait video phones work well on ordinary GSM wireless telecommunication networks. Jiang Li 0008, Keman Yu, Hanning Zhou, Jizheng Xu, King To Ng, Kaibo Wang, Harry Shum |
ACM Multimedia | 8 |