EDBT 2026 Demo / reviewers in the wild / expert
Xiaonan Zhao
dblp:15/3558
· DBLP profile ↗
24ranked-venue papers
3as first author
11since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 9 · 1 first-author · 2 since 2021Systems, architecture and hardware · 6 · 1 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | wdCP: Windowed Incremental Checkpointing for Efficient and Bounded LLM RecoveryabstractCheckpointing is essential for fault tolerance in large-scale LLM training, yet periodic full-state checkpoints bring heavy I/O overhead and training stalls. Prior work suggests that differential checkpointing is ineffective for LLMs, since most parameters updates every iteration, leading to dense updates. This paper revisit this assumption and observe that parameter updates are naturally generated inside optimizer execution and exhibit significant temporal and layer-wise heterogeneity. Guided by this, we present wdCP, a lightweight runtime that captures optimizer-level parameter deltas and asynchronously persists them using a windowed buffering mechanism. wdCP further introduces lightweight anchor snapshots to bound recovery cost. We implement wdCP and evaluate it on several representative models. Results show that wdCP introduces less than 5% training overhead while achieving up to 69.2× reduction in checkpoint size and enabling fast, bounded recovery. Wendi Cheng, Xiao Zhang 0014, Xiaonan Zhao, Xiaoling Shu, Jinjiang Wang, Shujie Han 0001 |
CF | 3 |
| 2025 | BVLSM: Write-Efficient LSM-Tree Storage via WAL-Time Key-Value SeparationabstractModern data-intensive applications increasingly store and process big-value items, such as multimedia objects and machine learning embeddings, which exacerbate storage inefficiencies in Log-Structured Merge-Tree (LSM)-based key-value stores. This paper presents BVLSM, a Write-Ahead Log (WAL)time key-value separation mechanism designed to address three key challenges in LSM-Tree storage systems: write amplification, poor memory utilization, and I/O jitter under big-value workloads. Unlike state-of-the-art approaches that delay key-value separation until the flush stage, leading to redundant data in MemTables and repeated writes. BVLSM proactively decouples keys and values during the WAL phase. The MemTable stores only lightweight metadata, allowing multi-queue parallel store for big value. Benchmark results demonstrate that BVLSM significantly outperforms both RocksDB and BlobDB across various write-intensive workloads. Specifically, in asynchronous WAL mode, its 64KB write throughput exceeds that of RocksDB by a factor of 7.6 and that of BlobDB by a factor of 1.9. Wendi Cheng, Jiahe Wei, Xueqiang Shan, Weikai Liu, Xiaonan Zhao, Xiao Zhang 0014 |
HPCC | 6 |
| 2025 | Large FoV Direction-Finding System Based on a Small-Aperture Phase Mode Antenna and its Applications on UAVsabstractUnmanned aerial vehicles (UAVs), serving as small mobile platforms, are well-suited for emitter detection and localization tasks in dynamic environments. However, implementing large antenna array apertures on the compact platform of UAVs is impractical. This article presents a small-aperture phase mode antenna (PMA) direction-finding (DF) system for UAV applications. Initially, a small-aperture, high port-isolation dual-port PMA is introduced. Through miniaturization and decoupling, a conventional slot antenna pair is transformed into the final dual-port PMA. The dual-port PMA features a compact aperture of$0.4~\lambda \times 0.47~\lambda $(where$\lambda $is the wavelength at 2 GHz), with port isolation reaching up to 40 dB. Next, we investigate the integration of the dual-port PMA with a novel amplitude-only DF (AODF) method, termed the single PMA multiple patterns amplitude comparison (SPMAMP-AC) method. This method utilizes multiple patterns to construct a DF function cluster, providing higher DF accuracy over a broader Field of View (FoV) compared to the traditional AODF method. The measured results indicate maximum DF errors of 2.4°, 4.3°, 4.9°, and 7.9° for FoVs of$60{^{\circ }}~(\theta \in $[−30°, 30°]),$140{^{\circ }}~(\theta \in $[−70°, 70°]),$160{^{\circ }}~(\theta \in $[−80°, 80°]), and$180{^{\circ }}~(\theta \in $[−90°, 90°]) in the yoz plane, respectively. Xudong Tang, Xiaonan Zhao, Han Zhou 0012, Junping Geng, Jingzheng Lu, Xuepeng Li, Enyu Li, Heci Liu, Chong He, Ronghong Jin, Guolin Tong |
IEEE Internet Things J. | 2 |
| 2025 | Scheduling virtual machines and containers: A comparative review of techniques, performance, and future trends
Jiameng Zhang, Ruofei Wu, Taoyu Zhong, Shujie Han 0001, Xiao Zhang 0014, Xiaonan Zhao |
J. Syst. Archit. | 10 |
| 2024 | TraceGen: A Block-level Storage System Performance Evaluation Tool for Analyzing and Generating I/O TracesabstractPerformance measurement is essential for detecting potential performance issues and guiding optimization efforts. However, acquiring I/O traces of real applications can be costly in production environments. Also, existing performance measurement tools, such as FIO and Iometer, often oversimplify real-world application characteristics. In this paper, we introduce TraceGen, a block-level performance measurement tool for storage systems that consists of a trace analyzer and a trace generator. The trace analyzer produces two categories of traces: (i) new traces with specified characteristics designed to accurately simulate a range of applications, and (ii) extended traces that maintain similar workload characteristics to the input traces, thereby improving measurement accuracy during trace replay. We evaluate TraceGen using traces from an enterprise production environment and demonstrate its capability to generate new traces with an error margin of less than 1%. Jiahe Wei, Huiru Xie, Jinjiang Wang, Xiaonan Zhao, Shujie Han 0001, Xiao Zhang 0014 |
HPCC | 5 |
| 2023 | Magnetic coupling governed pinning directions in magnetic tunnel junctions under magnetic field annealing with zero magnetic field cooling
Shaohua Yan, Shiyang Lu, Xiaonan Zhao, Runrun Hao, Zitong Zhou, Kun Zhang 0030, Shishen Yan, Qunwen Leng |
Sci. China Inf. Sci. | 5 |
| 2022 | Visual Representation Learning with Self-Supervised Attention for Low-Label High-Data RegimeabstractSelf-supervision has shown outstanding results for natural language processing, and more recently, for image recognition. Simultaneously, vision transformers and its variants have emerged as a promising and scalable alternative to convolutions on various computer vision tasks. In this paper, we are the first to question if self-supervised vision transformers (SSL-ViTs) can be adapted to two important computer vision tasks in the low-label, high-data regime: few-shot image classification and zero-shot image retrieval. The motivation is to reduce the number of manual annotations required to train a visual embedder, and to produce generalizable and semantically meaningful embeddings. For few-shot image classification we train SSL-ViTs without any supervision, on external data, and use this trained embedder to adapt quickly to novel classes with limited number of labels. For zero-shot image retrieval, we use SSL-ViTs pre-trained on a large dataset without any labels and fine-tune them with several metric learning objectives. Our self-supervised attention representations outperforms the state-of-the-art on several public benchmarks for both tasks, namely miniImageNet and CUB200 for few-shot image classification by up-to 6%-10%, and Stanford Online Products, Cars196 and CUB200 for zero-shot image retrieval by up-to 4%-11%. Code is available at https://github.com/AutoVision-cloud/SSL-ViT-lowlabel-highdata. Prarthana Bhattacharyya, Chenge Li, Xiaonan Zhao, István Fehérvári, Jason Sun |
ICASSP | 3 |
| 2022 | SeeTek: Very Large-Scale Open-set Logo Recognition with Text-Aware Metric LearningabstractRecent advances in deep learning and computer vision have set new state of the art in logo recognition [2], [9], [36]. Logo recognition has mostly been approached as a closed-set object recognition problem and more recently as an open-set retrieval problem. Current approaches suffer from distinguishing visually similar logos, especially in open-set retrieval for very large-scale applications with thousands of brands. To address the problem, we propose a multi-task learning architecture of deep metric learning and scene text recognition. We use brand names as weak labels and enforce the model to simultaneously extract distinct visual features as well as predict brand name text. To achieve it, we collected a dataset with 3 Million logos cropped from Amazon Product Catalog images across nearly 8K brands, named PL8K. Our experiments show that adding the task of text recognition during training boosts the model’s retrieval performance both on our PL8K dataset and on five other public logo datasets. Chenge Li, István Fehérvári, Xiaonan Zhao, Ives Macêdo, Srikar Appalaraju |
WACV | 3 |
| 2022 | An Adaptive Elastic Multi-model Big Data Analysis and Information Extraction SystemabstractAbstract With the diverse applications to industry and domain-specific context, multi-source information extraction on semi-structured and unstructured data, as well as across data models, is becoming more common. However, multi-model information extraction often requires the deployment of multiple data model management, storage, and analysis subsystems on the cloud, many subsystems are not high-resource utilization at the same time, and the resource waste phenomenon is often serious. Therefore, an adaptive scalable multi-model big data analysis and information extraction system is designed and implemented in this paper, which can support data maintenance and cross-model query of relational, graph, document, key and other data models, and can provide efficient cross-model information extraction. On this basis, we can achieve the system resource allocation on demand and fast scaling mechanism, according to the real-time requirements of multi-model big data analysis, and dynamic adjustment of each subsystem resource allocation. Therefore, our solution not only guarantees multi-model query and information extraction performance and quality of service, but also significantly reduces the total consumption of system resources and cost. Qiang Yin 0004, Sheng Du, Jianquan Leng, Yinhao Hong, Feng Zhang 0007, Yunpeng Chai, Xiao Zhang 0014, Xiaonan Zhao, Wei Lu 0015 |
Data Sci. Eng. | 10 |
| 2021 | WOBTree: a write-optimized B+-tree for non-volatile memory
Zhanhuai Li, Xiao Zhang 0014, Xiaonan Zhao, Song Jiang 0001 |
Frontiers Comput. Sci. | 4 |
| 2021 | Directional Modulation Design Under a Given Symbol-Independent Magnitude Constraint for Secure IoT NetworksabstractDirectional modulation (DM) is an important technology for physical layer security in wireless communications. Recently, a symbol-independent magnitude constraint for all antennas was proposed in DM design to reduce the design complexity of its analogue implementation. However, a limitation of the method is that it can only set the magnitude to a certain value, and all the antenna coefficients have the same magnitude. In this article, a more flexible solution is provided and the challenge of the design is the nonconvex constraint enforcing an arbitrary symbol-independent magnitude for all coefficients. To solve the problem, a convex iterative method is proposed, based on which the magnitudes of weight coefficients for all antennas can be chosen by designers in advance according to the specific requirements, allowing more freedom in the design process, which is the major difference between the previously proposed design and the newly proposed one. Two design examples are provided to demonstrate the effectiveness of the proposed design. One is is a general example, where coefficient magnitudes for different symbols are the same for the same antenna, but different for different antennas; the other one is a special case where magnitudes for all antennas are the same. Bo Zhang 0033, Wei Liu 0001, Qiang Li 0019, Yang Li 0045, Xiaonan Zhao, Cuiping Zhang, Cheng Wang 0018 |
IEEE Internet Things J. | 5 |
| 2020 | Design of and research on industrial measuring devices based on Internet of Things technology
Yang Li 0045, Licheng Yang 0001, Cuiping Zhang, Bo Zhang 0033, Xiaonan Zhao |
Ad Hoc Networks | 6 |
| 2020 | A study of a one-turn circular patch antenna array and the influence of the human body on the characteristics of the antenna
Yang Li 0045, Licheng Yang 0001, Xiaonan Zhao, Xin Zhang 0042 |
Ad Hoc Networks | 4 |
| 2020 | Directional modulation design under maximum and minimum magnitude constraints for weight coefficients
Bo Zhang 0033, Wei Liu 0001, Yang Li 0045, Xiaonan Zhao, Cheng Wang 0018 |
Ad Hoc Networks | 4 |
| 2020 | Symbol-independent weight magnitude design for antenna array based directional modulation
Bo Zhang 0033, Wei Liu 0001, Yang Li 0045, Xiaonan Zhao, Cuiping Zhang, Cheng Wang 0018 |
Ad Hoc Networks | 4 |
| 2020 | Design of the sleeping aid system based on face recognition
Xiaonan Zhao, Yang Li 0045 |
Ad Hoc Networks | 1 |
| 2020 | Consortium Blockchain-Based Secure Software Defined Vehicular Network
Hao Wu 0005, Xiaonan Zhao |
Mob. Networks Appl. | 3 |
| 2018 | OC-Cache: An Open-channel SSD Based Cache for Multi-Tenant SystemsabstractIn a multi-tenant cloud environment, tenants are usually hosted by virtual machines. Cloud providers deploy multiple virtual machines on a physical server to better utilize physical resources including CPU, memory, and storage devices. SSDs are often used as an I/O cache shared among the tenants for large storage systems using hard disk drives (HDDs) as their main storage devices, which can receive much of SSD's performance benefit and HDD's cost advantage. A key challenge in the use of the shared cache is to ensure strong performance isolation and maintain its high utilization at the same time. However, conventional SSD cache management approaches cannot effectively address this challenge. In this paper, we propose OC-Cache, an open-channel SSD cache framework which utilizes SSD'd internal parallelism to adaptively allocate cache to tenants for both good performance isolation and high SSD utilization. In particular, OC-Cache uses a tenant's miss ratio curve to determine the amount of cache space allocation and where the allocation is (in dedicated or shared SSD channels) and dynamically manages cache space according to the workload characteristics. Experiments show that OC-Cache significantly reduces interference among tenants, and maintains high utilization of the SSD cache. Zhanhuai Li, Xiao Zhang 0014, Xiaonan Zhao, Xingsheng Zhao, Song Jiang 0001 |
IPCCC | 4 |
| 2018 | FSObserver: A Performance Measurement and Monitoring Tool for Distributed Storage Systems
Xiao Zhang 0014, Lanxin Kong, Shunyi Zhu, Zhanhuai Li, Xiaonan Zhao |
NPC | 5 |
| 2010 | FDTM: Block Level Data Migration Policy in Tiered Storage System
Xiaonan Zhao, Zhanhuai Li, Leijie Zeng |
NPC | 1 |
| 2008 | Image spam hunterabstractSpammers are constantly creating sophisticated new weapons in their arms race with anti-spam technology, the latest of which is image-based spam. The newest image-based spam uses simple image processing technologies to vary the content of individual messages, e.g. by changing foreground colors, backgrounds, font types, or even rotating and adding artifacts to the images. Thus, they pose great challenges to conventional spam filters. In this paper, we propose a system using a probabilistic boosting tree to determine whether an incoming image is a spam or not based on global image features, i.e. color and gradient orientation histograms. The system identifies spam without the need for OCR and is robust in the face of the kinds of variation found in current spam images. Evaluation results show the system correctly classifies 90% of spam images while mislabeling only 0.86% of non-spam images as spam. Yan Gao 0003, Ming Yang 0007, Xiaonan Zhao, Bryan Pardo, Ying Wu 0001, Thrasyvoulos N. Pappas, Alok N. Choudhary |
ICASSP | 3 |
| 2008 | Structural texture similarity metrics for retrieval applicationsabstractTraditional image similarity metrics compare two images on a point-by-point basis. On the other hand, structural similarity metrics (SSIM) attempt to base image similarity on "structural" information. We evaluate the performance of SSIM metrics in the context of texture similarity, and propose new metrics that incorporate the best features of SSIM and eliminate the most serious drawbacks. We show that the proposed new texture similarity metrics outperform SSIM and its variations, as well as PSNR and other traditional metrics. We demonstrate the advantages of the new metrics on a carefully selected set of 39 texture pairs and comparisons with informal subjective test results. Xiaonan Zhao, Matthew G. Reyes, Thrasyvoulos N. Pappas, David L. Neuhoff |
ICIP | 1 |
| 2008 | Structural Similarity Quality Metrics in a Coding Context: Exploring the Space of Realistic DistortionsabstractPerceptual image quality metrics have explicitly accounted for human visual system (HVS) sensitivity to subband noise by estimating just noticeable distortion (JND) thresholds. A recently proposed class of quality metrics, known as structural similarity metrics (SSIM), models perception implicitly by taking into account the fact that the HVS is adapted for extracting structural information from images. We evaluate SSIM metrics and compare their performance to traditional approaches in the context of realistic distortions that arise from compression and error concealment in video compression/transmission applications. In order to better explore this space of distortions, we propose models for simulating typical distortions encountered in such applications. We compare specific SSIM implementations both in the image space and the wavelet domain; these include the complex wavelet SSIM (CWSSIM), a translation-insensitive SSIM implementation. We also propose a perceptually weighted multiscale variant of CWSSIM, which introduces a viewing distance dependence and provides a natural way to unify the structural similarity approach with the traditional JND-based perceptual approaches. Alan C. Brooks, Xiaonan Zhao, Thrasyvoulos N. Pappas |
IEEE Trans. Image Process. | 2 |
| 2007 | Lossy Compression of Bilevel Images Based on Markov Random FieldsabstractA new method for lossy compression of bilevel images based on Markov random fields (MRFs) is proposed. It preserves key structural information about the image, and then reconstructs the smoothest image that is consistent with this information. The smoother the original image, the lower the required bit rate, and conversely, the lower the bit rate, the smoother the approximation provided by the decoded image. The main idea is that as long as the key structural information is preserved, then any smooth contours consistent with this information will provide an acceptable reconstructed image. The use of MRFs in the decoding stage is the key to efficient compression. Experimental results demonstrate that the new technique outperforms existing lossy compression techniques, and provides substantially lower rates than lossless techniques (JBIG) with little loss in perceived image quality. Matthew G. Reyes, Xiaonan Zhao, David L. Neuhoff, Thrasyvoulos N. Pappas |
ICIP (2) | 2 |