Yiming Qiao

dblp:257/5648 · DBLP profile ↗
← Back
8ranked-venue papers
5as first author
6since 2021 · last 2026
0000-0001-8023-0276ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 4 · 3 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Artificial intelligence and machine learning · 1Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021Computer networks · 1 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Low-Complexity Overlapped Subarrays Based Hybrid Precoding for Beam Squint Mitigation in Massive MIMO Systems
Yiming Qiao, Bowen Zhong, Zhongxiang Wei, Zhixiang Xu
WCNC1
2026 Robust Predicate Transfer with Dynamic Execution
Yiming Qiao, Peter Boncz, Huanchen Zhang
Proc. VLDB Endow.1
2025 A Loss Weighting Algorithm Based on In-batch Positive Passage Rankings for Dense Retrievers
abstract
In the domain of dense retrieval, training with hard negatives is widely used. However, the existence of a large number of false negatives has given rise to some hard-to-train queries, which often have a significant impact on the loss at the same time. The training process overemphasizes these queries that account for a disproportionately high proportion in the training loss while neglecting the positive examples with moderate rankings, thus resulting in suboptimal model outcomes. We theoretically analyze the excessive influence of hard-to-train queries and introduce an in-batch loss weighting algorithm that adaptively assigns weights based on the in-batch ranking position of its positive document. Specifically, our method will calculate the in-batch ranking positions of the similarity between positive examples and queries, construct query weights based on the Gaussian distribution, and balance the attention in the training process. This strategy mitigates the effects of false negatives while retaining the benefits of hard negatives. Experiments on the MS-MARCO dataset show statistically significant improvements: a 1.3% increase in MRR@10, higher Recall@1000, and consistent NDCG@10 gains on the TREC test set across diverse hard negative sources. The relevant code has been open-sourced on the GitHub repository: https://github.com/TRcreeper/IBRWLoss.
Yiming Qiao
SMC5
2025 Data Chunk Compaction in Vectorized Execution
abstract
Modern analytical database management systems often adopt vectorized query execution engines that process columnar data in batches (i.e., data chunks) to minimize the interpretation overhead and improve CPU parallelism. However, certain database operators, especially hash joins, can drastically reduce the number of valid entries in a data chunk, resulting in numerous small chunks in an execution pipeline. These small chunks cannot fully enjoy the benefits of vectorized query execution, causing significant performance degradation. The key research question is when and how to compact these small data chunks during query execution. In this paper, we first model the chunk compaction problem and analyze the trade-offs between different compaction strategies. We then propose a learning-based algorithm that can adjust the compaction threshold dynamically at run time. To answer the ''how'' question, we propose a compaction method for the hash join operator, called logical compaction, that minimizes data movements when compacting data chunks. We implemented the proposed techniques in the state-of-the-art DuckDB and observed up to 63% speedup when evaluated using the Join Order Benchmark, TPC-H, and TPC-DS.
Yiming Qiao, Huanchen Zhang
Proc. ACM Manag. Data1
2024 A hardware acceleration Framework of Reconfigurable Edge Modules for Convolutional Neural Networks
abstract
Edge computing exploits node devices situated in close proximity to terminals to deliver distributed computing services directly to users, with FPGAs serving as the predominant platform. With the amplification in volume and intricacy of deep learning models, effectively deploying these models on FPGA devices introduces substantial challenges. To counter this predicament, this paper introduces an FPGA-based hardware acceleration framework for reconfigurable edge devices tailored for Convolutional Neural Networks (CNNs). Initially, a method for dynamically quantizing network models is devised to markedly curtail model memory consumption. Subsequently, a meticulously engineered hardware acceleration unit is formulated to attain augmented computing parallelism via meticulous temporal redesign. Finally, model deployment on FPGA devices for inference verification is showcased utilizing fully connected and convolutional neural networks as exemplars. On the MNIST dataset, the FPGA inference unit attains remarkable accuracy and computational efficiency in comparison to CPUs and GPUs
Yiming Qiao, Zidi Jia, Shixiang Li, Lei Ren 0001
IECON1
2024 Blitzcrank: Fast Semantic Compression for In-memory Online Transaction Processing
abstract
We present Blitzcrank, a high-speed semantic compressor designed for OLTP databases. Previous solutions are inadequate for compressing row-stores: they suffer from either low compression factor due to a coarse compression granularity or suboptimal performance due to the inefficiency in handling dynamic data sets. To solve these problems, we first propose novel semantic models that support fast inferences and dynamic value set for both discrete and continuous data types. We then introduce a new entropy encoding algorithm, called delayed coding, that achieves significant improvement in the decoding speed compared to modern arithmetic coding implementations. We evaluate Blitzcrank in both standalone microbenchmarks and a multicore in-memory row-store using the TPC-C benchmark. Our results show that Blitzcrank achieves a sub-microsecond latency for decompressing a random tuple while obtaining high compression factors. This leads to an 85% memory reduction in the TPC-C evaluation with a moderate (19%) throughput degradation. For data sets larger than the available physical memory, Blitzcrank help the database sustain a high throughput for more transactions before the I/O overhead dominates.
Yiming Qiao, Huanchen Zhang
Proc. VLDB Endow.1
2020 DSPNet: A Lightweight Dilated Convolution Neural Networks for Spectral Deconvolution With Self-Paced Learning
abstract
In the fields of industry research, infrared spectrometers are widely used in diverse applications. However, the spectrum often suffers from band overlap and random noise due to the distortion caused by the point spread function, especially for aging instruments. The problem of reconstructing the clear spectrum from the degraded spectrum is called spectrum deconvolution. Traditional partial differential equation (PDE) methods rely on distribution assumptions in the reconstructed process. This restriction makes PDE methods sensitive to tackle complex instrumental broadening effect in the dispersive IR spectrometers. Also, we need to spend much time setting the parameters of PDE models manually. These problems intuitively degrade the performances of PDE methods. In this article, we propose an end-to-end neural network framework for spectral deconvolution problem. The novelty of this article lies in its strong robustness from dilated deconvolution and self-paced learning procedure to challenge the complicated degraded spectra. Actually, the deconvolution problem is tailored to a dense prediction problem in this article. Inspired by the extensive use and excellent effects of dilated convolutions in dense prediction, a lightweight dilated convolution module is given to detect the overlaps of degraded spectra. Experimental results demonstrate that the proposed solution has an outstanding performance against many other approaches. Such improvements have the potential to facilitate industrial applications and further exploration of an unknown chemical mixture. Our framework has a good performance on feature extracting and spectrum reconstruction, even in the case of low signal-to-noise ratio.
Hu Zhu, Yiming Qiao, Guoxia Xu, Lizhen Deng, Yu-Feng Yu 0001
IEEE Trans. Ind. Informatics2
2019 Fast Classification Algorithms via Distributed Accelerated Alternating Direction Method of Multipliers
abstract
Distributed machine learning has gained lots of attention due to the rapid growth of data. In this paper, we focus regularized empirical risk minimization problems, and propose two novel Distributed Accelerated Alternating Direction Method of Multipliers (D-A2DM2) algorithms for distributed classification. Based on the framework of Alternating Direction Method of Multipliers (ADMM), we decentralize the distributed classification problem as a global consensus optimization problem with a series of sub-problems. In D-A2DM2, we exploit ADMM with variance reduction for sub-problem optimization in parallel. Taking global update and local update into consideration respectively, we propose two acceleration mechanisms in the framework of D-A2DM2. In particular, inspired by Nesterov's accelerated gradient descent, we utilize it for global update to further improve time efficiency. Moreover, we also introduce Nesterov's acceleration for local update, and develop the corrected local update and symmetric dual update to accelerate the convergence with only a little change in the computational effort. Theoretically, D-A2DM2 has a linear convergence rate. Empirically, experimental results show that D-A2DM2 converge faster than existing distributed ADMM-based classification, and could be a highly efficient algorithm for practical use.
Huihui Wang 0001, Shunmei Meng, Yiming Qiao, Jing Zhang 0015
ICDM3