Yi Qiao

dblp:67/2922 · DBLP profile ↗
← Back
19ranked-venue papers
9as first author
9since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 7 · 4 first-author · 1 since 2021Computer networks · 4 · 2 first-author · 1 since 2021Artificial intelligence and machine learning · 3 · 2 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
7 papers
Distributed systems · 33% Storage systems · 26% Cloud and datacenter computing · 26%
Computer networks
4 papers
Software-defined and programmable networks · 58% Physical-layer communications · 26% Content delivery and video streaming · 8%
Artificial intelligence
1 paper
Language models and text generation · 77% Knowledge representation and reasoning · 23%
Databases, data mining, and information retrieval
1 paper
Recommender systems · 100%
Interdisciplinary, comprehensive, and emerging computing
1 paper
Bioinformatics and computational biology · 100%
Software engineering, system software, and programming languages
1 paper
Operating systems · 100%

Topics — the 29 heaviest of 31, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Language models and text generation
large language model
1.012026
Behavior Tokens Speak Louder: Disentangled Explainable Recommendation with Behavior Vocabulary · AAAI 2026
Recommender systems
explainable recommendation
1.012026
Behavior Tokens Speak Louder: Disentangled Explainable Recommendation with Behavior Vocabulary · AAAI 2026
Physical-layer communications › channel coding › error control coding
forward error correction
0.812024
NetFEC: In-network FEC Encoding Acceleration for Latency-sensitive Multimedia Applications · INFOCOM 2024
Software-defined and programmable networks
programmable data plane
0.812024
NetFEC: In-network FEC Encoding Acceleration for Latency-sensitive Multimedia Applications · INFOCOM 2024
Software-defined and programmable networks › programmable data plane
programmable switching ASIC
0.812024
NetFEC: In-network FEC Encoding Acceleration for Latency-sensitive Multimedia Applications · INFOCOM 2024
Bioinformatics and computational biology
sequence analysis
0.712023
quickBAM: a parallelized BAM file access API for high-throughput sequence analysis informatics · Bioinform. 2023
Storage systems › repair
data reconstruction
0.612022
NetEC: Accelerating Erasure Coding Reconstruction With In-Network Aggregation · IEEE Trans. Parallel Distributed Syst. 2022
Storage systems › storage reliability
erasure coding
0.612022
NetEC: Accelerating Erasure Coding Reconstruction With In-Network Aggregation · IEEE Trans. Parallel Distributed Syst. 2022
Distributed systems › data aggregation
in-network aggregation
0.612022
NetEC: Accelerating Erasure Coding Reconstruction With In-Network Aggregation · IEEE Trans. Parallel Distributed Syst. 2022
Operating systems › kernel
kernel design
0.412019
Transkernel: Bridging Monolithic Kernels to Peripheral Cores · USENIX ATC 2019
Cloud and datacenter computing › cloud networking
cloud gateway
0.412019
Tripod: Towards a Scalable, Efficient and Resilient Cloud Gateway · IEEE J. Sel. Areas Commun. 2019
Cloud and datacenter computing
cloud infrastructure
0.412019
Tripod: Towards a Scalable, Efficient and Resilient Cloud Gateway · IEEE J. Sel. Areas Commun. 2019
Cloud and datacenter computing › virtualization › network virtualization
network function virtualization
0.412019
Tripod: Towards a Scalable, Efficient and Resilient Cloud Gateway · IEEE J. Sel. Areas Commun. 2019
Distributed systems
state management
0.412019
Tripod: Towards a Scalable, Efficient and Resilient Cloud Gateway · IEEE J. Sel. Areas Commun. 2019
Distributed systems
peer-to-peer systems
0.232008
Designing less-structured P2P systems for the expected high churn · IEEE/ACM Trans. Netw. 2008
Improving peer-to-peer performance through server-side scheduling · ACM Trans. Comput. Syst. 2008
Structured and Unstructured Overlays under the Microscope: A Measurement-based View of Two P2P Systems That People Use · USENIX ATC, General Track 2006
High-performance computing
parallel i/o
0.212023
quickBAM: a parallelized BAM file access API for high-throughput sequence analysis informatics · Bioinform. 2023
Edge and fog computing
offloading
0.212022
NetEC: Accelerating Erasure Coding Reconstruction With In-Network Aggregation · IEEE Trans. Parallel Distributed Syst. 2022
Software-defined and programmable networks › programmable data plane
programmable switch
0.212022
NetEC: Accelerating Erasure Coding Reconstruction With In-Network Aggregation · IEEE Trans. Parallel Distributed Syst. 2022
Processor architecture and microarchitecture › multicore design
heterogeneous cores
0.112019
Transkernel: Bridging Monolithic Kernels to Peripheral Cores · USENIX ATC 2019
Parallel and multicore computing
load balancing
0.112019
Tripod: Towards a Scalable, Efficient and Resilient Cloud Gateway · IEEE J. Sel. Areas Commun. 2019
Distributed systems › peer-to-peer systems
overlay networks
0.122008
Structured and Unstructured Overlays under the Microscope: A Measurement-based View of Two P2P Systems That People Use · USENIX ATC, General Track 2006
Designing less-structured P2P systems for the expected high churn · IEEE/ACM Trans. Netw. 2008
Distributed systems › peer-to-peer systems › churn
churn resilience
0.112008
Designing less-structured P2P systems for the expected high churn · IEEE/ACM Trans. Netw. 2008
Distributed systems › resource sharing
data sharing
0.112008
Improving peer-to-peer performance through server-side scheduling · ACM Trans. Comput. Syst. 2008
Performance modeling and evaluation
scheduling policy
0.112008
Improving peer-to-peer performance through server-side scheduling · ACM Trans. Comput. Syst. 2008
Performance modeling and evaluation › scheduling policy
shortest remaining processing time
0.112008
Improving peer-to-peer performance through server-side scheduling · ACM Trans. Comput. Syst. 2008
Network measurement and analytics
traffic characterization
0.012004
An Empirical Study of the Multiscale Predictability of Network Traffic · HPDC 2004
Performance modeling and evaluation › simulation › discrete-event simulation
trace-driven simulation
0.012008
Improving peer-to-peer performance through server-side scheduling · ACM Trans. Comput. Syst. 2008
Network measurement and analytics › internet measurement
peer-to-peer network measurement
0.012006
Structured and Unstructured Overlays under the Microscope: A Measurement-based View of Two P2P Systems That People Use · USENIX ATC, General Track 2006
Network measurement and analytics
traffic prediction
0.012004
An Empirical Study of the Multiscale Predictability of Network Traffic · HPDC 2004

Methods — techniques the papers use, named apart from their topics

vector-quantized autoencoding · 2.0semantic alignment regularization · 2.0multi-level semantic supervision · 2.0parallelization · 1.3on-switch TCP proxy · 1.1galois field offloading · 1.1ghost packet mechanism · 0.8forward error correction · 0.8traffic processing · 0.4state management service · 0.4trace-driven simulation · 0.1request service time estimation · 0.1multiscale analysis · 0.0empirical study · 0.0
YearPublicationVenuePosition
2026 Behavior Tokens Speak Louder: Disentangled Explainable Recommendation with Behavior Vocabulary
abstract
Recent advances in explainable recommendation have explored the integration of language models to analyze natural language rationales for user–item interactions. Despite their potential, existing methods often rely on ID-based representations that obscure semantic meaning and impose structural constraints on language models, thereby limiting their applicability in open-ended scenarios. These challenges are intensified by the complex nature of real-world interactions, where diverse user intents are entangled and collaborative signals rarely align with linguistic semantics. To overcome these limitations, we propose BEAT, a unified and transferable framework that tokenizes user and item behaviors into discrete, interpretable sequences. We construct a behavior vocabulary via a vector-quantized autoencoding process that disentangles macro-level interests and micro-level intentions from graph-based representations. We then introduce multi-level semantic supervision to bridge the gap between behavioral signals and language space. A semantic alignment regularization mechanism is designed to embed behavior tokens directly into the input space of frozen language models. Experiments on three public datasets show that BEAT improves zero-shot recommendation performance while generating coherent and informative explanations. Further analysis demonstrates that our behavior tokens capture fine-grained semantics and offer a plug-and-play interface for integrating complex behavior patterns into large language models.
Xinshun Feng, Mingzhe Liu 0002, Yi Qiao, Tongyu Zhu, Leilei Sun
AAAI3
2026 Real-time rehabilitation assessment and corrective guidance driven by dual regulation pose analysis
Yi Qiao, Zilong Wang 0029
Appl. Intell.1
2026 Correction: Real-time rehabilitation assessment and corrective guidance driven by dual regulation pose analysis
Yi Qiao, Zilong Wang 0029
Appl. Intell.1
2025 Domain-aware Node Representation Learning for Graph Out-of-Distribution Generalization
abstract
Graph Neural Networks (GNNs) have demonstrated impressive success across diverse fields when data satisfies in-distribution (ID) assumption. Nevertheless, GNN performance significantly declines in cases of distribution shifts between training and testing graph data. This degradation primarily stems from spurious correlations between irrelevant domain information and target labels in out-of-distribution (OOD) scenarios. Thus, maximizing the utilization of domain information becomes imperative. In light of this, we propose a novel approach named Domain-aware Node Representation Learning (DNRL), comprehensively incorporates domain information to bolster generalization capability. Specifically, DNRL selectively interpolates nodes with the same label but different domains, extending training data into unseen domains and alleviating the effects caused by domain-related spurious correlations. Futhermore, by introducing a domain-aware contrastive learning strategy, our method implicitly decouples domain information from node information to learn domain-independent node representations. Extensive experiments on graph out-of-distribution benchmarks demonstrate that DNRL can achieve effective OOD generalization performance across diverse domains.
Yi Qiao, Yang Liu 0200, Qing He 0003, Xiang Ao 0001
ICASSP1
2025 The application of compressed sensing on tumor mutation burden calculation from overlapped pooling sequencing data
abstract
BACKGROUND: Tumor Mutation Burden (TMB) is commonly characterized as the number of non-synonymous somatic SNVs per megabase within the gene region identified through whole exon sequencing or targeted sequencing in a tumor sample. It has been statistically demonstrated that TMB was related to the ability of neoantigen production and used to predict the efficacy of immunotherapy for various types of cancers. However, screening for TMB in patients poses challenges due to the extensive labor and financial resources required for the preparation of large quantities of parallel sequencing libraries. RESULTS: In this study, we employed compressed sensing (CS) to calculate TMB from overlapped pooling sequencing data, aiming to reduce the sequencing cost by minimizing the number of library builds. Over 90% SNPs could still be detected without a significant loss of mutation information even when the data is pooled from ten different samples. Based on this, the orthogonal matching pursuit (OMP) algorithm and the basic pursuit (BP) algorithm were used to reconstruct TMB from pooling sequencing data. The performance of these two algorithms was evaluated. The BP algorithm consistently performed well across all cases, albeit necessitating extended computational time. The OMP algorithm has been proved to be suitable for scenarios where the original matrix was sparse but it showed low overall performance. Based on an accurate calculation of TMB, we determined that the number of sequencing runs could be reduced to 0.6 times the total number of samples, resulting in a 40% reduction in sequencing cost. CONCLUSIONS: In conclusion, we calculated TMB from overlapped pooling sequencing data utilizing compressed sensing strategy to reduce sequencing cost. Our findings confirm that the SNP calling from ten samples' pooling sequencing data is feasible. Additionally, we performed an assessment of the reconstruction efficiency of both the BP model and the OMP model.
Yi Qiao, Rongming An, Xuan Pan, Jing Tu
BMC Bioinform.2
2024 NetFEC: In-network FEC Encoding Acceleration for Latency-sensitive Multimedia Applications
abstract
In face of packet loss, latency-sensitive multimedia applications cannot afford re-transmission because loss detection and re-transmission could lead to extra latency or otherwise compromised media quality. Alternatively, forward error correction (FEC) ensures reliability by adding redundancy and it is able to achieve lower latency at the cost of bandwidth and computational overheads. We propose to re-locate FEC encoding to hardware that better suits the computational pattern of FEC encoding than CPUs. In this paper, we present NetFEC, an in-network acceleration system that offloads the entire FEC encoding process on the emergent programmable switching ASICs, eliminating all CPU involvement. We design the ghost packet mechanism so that NetFEC can be compatible with important media transport functionalities, including congestion control, pacing and statistics. We integrate NetFEC with WebRTC and conduct extensive experiments with real hardwares. Our evaluations demonstrate that NetFEC is able to eliminate server CPU burden and adds negligible overheads.
Yi Qiao, Han Zhang 0009, Jilong Wang 0001
INFOCOM1
2023 Exploration of whole genome amplification generated chimeric sequences in long-read sequencing data
abstract
MOTIVATION: Multiple displacement amplification (MDA) has become the most commonly used method of whole genome amplification, generating a vast amount of DNA with higher molecular weight and greater genome coverage. Coupling with long-read sequencing, it is possible to sequence the amplicons of over 20 kb in length. However, the formation of chimeric sequences (chimeras, expressed as structural errors in sequencing data) in MDA seriously interferes with the bioinformatics analysis but its influence on long-read sequencing data is unknown. RESULTS: We sequenced the phi29 DNA polymerase-mediated MDA amplicons on the PacBio platform and analyzed chimeras within the generated data. The 3rd-ChimeraMiner has been constructed as a pipeline for recognizing and restoring chimeras into the original structures in long-read sequencing data, improving the efficiency of using TGS data. Five long-read datasets and one high-fidelity long-read dataset with various amplification folds were analyzed. The result reveals that the mis-priming events in amplification are more frequently occurring than widely perceived, and the propor tion gradually accumulates from 42% to over 78% as the amplification continues. In total, 99.92% of recognized chimeric sequences were demonstrated to be artifacts, whose structures were wrongly formed in MDA instead of existing in original genomes. By restoring chimeras to their original structures, the vast majority of supplementary alignments that introduce false-positive structural variants are recycled, removing 97% of inversions on average and contributing to the analysis of structural variation in MDA-amplified samples. The impact of chimeras in long-read sequencing data analysis should be emphasized, and the 3rd-ChimeraMiner can help to quantify and reduce the influence of chimeras. AVAILABILITY AND IMPLEMENTATION: The 3rd-ChimeraMiner is available on GitHub, https://github.com/dulunar/3rdChimeraMiner.
Yi Qiao, Pengfei An, Jiajian Luo, Changwei Bi, Musheng Li, Zuhong Lu, Jing Tu
Briefings Bioinform.2
2023 quickBAM: a parallelized BAM file access API for high-throughput sequence analysis informatics
abstract
MOTIVATION: In time-critical clinical settings, such as precision medicine, genomic data needs to be processed as fast as possible to arrive at data-informed treatment decisions in a timely fashion. While sequencing throughput has dramatically increased over the past decade, bioinformatics analysis throughput has not been able to keep up with the pace of computer hardware improvement, and consequently has now turned into the primary bottleneck. Modern computer hardware today is capable of much higher performance than current genomic informatics algorithms can typically utilize, therefore presenting opportunities for significant improvement of performance. Accessing the raw sequencing data from BAM files, e.g. is a necessary and time-consuming step in nearly all sequence analysis tools, however existing programming libraries for BAM access do not take full advantage of the parallel input/output capabilities of storage devices. RESULTS: In an effort to stimulate the development of a new generation of faster sequence analysis tools, we developed quickBAM, a software library to accelerate sequencing data access by exploiting the parallelism in commodity storage hardware currently widely available. We demonstrate that analysis software ported to quickBAM consistently outperforms their current versions, in some cases finishing an analysis in under 3 min while the original version took 1.5 h, using the same storage solution. AVAILABILITY AND IMPLEMENTATION: Open source and freely available at https://gitlab.com/yiq/quickbam/, we envision that quickBAM will enable a new generation of high-performance informatics tools, either directly boosting their performance if they are currently data-access bottlenecked, or allow data-access to keep up with further optimizations in algorithms and compute techniques.
Anders Pitman, Xiaomeng Huang, Gabor T. Marth, Yi Qiao
Bioinform.4
2022 NetEC: Accelerating Erasure Coding Reconstruction With In-Network Aggregation
abstract
In distributed storage systems, Erasure Coding (EC) is a crucial technology to enable high data availability. By downloading parity data from survived machines, EC can reconstruct lost data with much lower storage overheads than data replication. However, this reduction in storage cost comes at the expense of extra performance problems:low reconstruction rate,high degraded read latency, andhigh host CPU utilization. Our analysis shows that these performance problems are deeply rooted in thehost-basedEC processing. To resolve these problems, we present NetEC, an in-network accelerating framework that fully offloads EC to the new generation programmable switching ASICs. We propose Explicit Buffer Size Notification (EBSN) to constrain decoding buffer usage, and design an on-switch one-to-many TCP proxy to integrate EBSN with TCP. We also design two parallel Galois Field (GF) offloading methods—table lookup and bitmatrix methods—to maximize parsable bytes. We implement NetEC on programmable switches and integrate it with HDFS. Extensive evaluations show that NetEC improves the reconstruction rate by 2.7x-6.8x, reduces the degraded read latency significantly, and removes the host CPU overhead completely. We also emulate multi-rack scenarios and show that NetEC is able to support$\sim$∼GB/s reconstruction rate and tens of concurrent tasks.
Yi Qiao, Menghao Zhang 0001, Yu Zhou 0008, Han Zhang 0009, Mingwei Xu 0001, Jun Bi, Jilong Wang 0001
IEEE Trans. Parallel Distributed Syst.1
2019 Transkernel: Bridging Monolithic Kernels to Peripheral Cores
Yi Qiao, Felix Xiaozhu Lin
USENIX ATC3
2019 Tripod: Towards a Scalable, Efficient and Resilient Cloud Gateway
abstract
Cloud gateways are fundamental components of a cloud platform, where various network functions (e.g., L4/L7 load balancing, network address translation, stateful firewall, and SYN proxy) are deployed to process millions of connections and billions of packets. Providing high-performance and failure-resilient packet processing with a scalable traffic management mechanism is crucial to ensuring the quality of service of a cloud provider, and hence is of great importance. Many network functions nowadays are implemented in software with commodity servers for low cost and high flexibility. However, existing software-based network function frameworks oftentimes provide part of these features, while cannot satisfy all three requirements above simultaneously. To address these issues, in this paper, we introduce TRIPOD, a novel network function framework specialized for cloud gateways. Having identified the fundamental limitations of loosely coupling traffic, processing logic and state, TRIPOD jointly manages these three elements with the unique characteristics of cloud gateways, which is enabled by a simple, efficient traffic processing mechanism, and a high performance state management service. Adopting several effective techniques and optimizations, TRIPOD is able to achieve scalable traffic management (<;100 flow rules for even ~Tbps traffic), high performance (reducing 40% of latency compared with state of the art) and failure resilience (similar packet/connection loss rate compared to state of the art), with reasonable overheads (less than 10% of the workload traffic) even under an extremely heavy traffic, making it a good fit for cloud gateways.
Menghao Zhang 0001, Jun Bi, Kai Gao 0001, Yi Qiao, Zhaogeng Li, Hongxin Hu
IEEE J. Sel. Areas Commun.4
2008 Improving peer-to-peer performance through server-side scheduling
abstract
We show how to significantly improve the mean response time seen by both uploaders and downloaders in peer-to-peer data-sharing systems. Our work is motivated by the observation that response times are largely determined by the performance of the peers serving the requested objects, that is, by the peers in their capacity as servers. With this in mind, we take a close look at this server side of peers, characterizing its workload by collecting and examining an extensive set of traces. Using trace-driven simulation, we demonstrate the promise and potential problems with scheduling policies based on shortest-remaining-processing-time (SRPT), the algorithm known to be optimal for minimizing mean response time. The key challenge to using SRPT in this context is determining request service times. In addressing this challenge, we introduce two new estimators that enable predictive SRPT scheduling policies that closely approach the performance of ideal SRPT. We evaluate our approach through extensive single-server and system-level simulation coupled with real Internet deployment and experimentation.
Yi Qiao, Fabián E. Bustamante, Peter A. Dinda, Stefan Birrer
ACM Trans. Comput. Syst.1
2008 Designing less-structured P2P systems for the expected high churn
Fabián E. Bustamante, Yi Qiao
IEEE/ACM Trans. Netw.2
2006 Structured and Unstructured Overlays under the Microscope: A Measurement-based View of Two P2P Systems That People Use
Yi Qiao, Fabián E. Bustamante
USENIX ATC, General Track1
2005 Characterizing and Predicting TCP Throughput on the Wide Area Network
abstract
DualPats exploits the strong correlation between TCP throughput and flow size, and the statistical stability of Internet path characteristics to accurately predict the TCP throughput of large transfers using active probing. We propose additional mechanisms to explain the correlation, and then analyze why traditional TCP benchmarking fails to predict the throughput of large transfers well. We characterize stability and develop a dynamic sampling rate adjustment algorithm so that we probe a path based on its stability. Our analysis, design, and evaluation is based on a large-scale measurement study.
Yi Qiao, Peter A. Dinda, Fabián E. Bustamante
ICDCS2
2005 Effects and Implications of File Size/Service Time Correlation onWeb Server Scheduling Policies
abstract
Recently, size-based policies such as SRPT and FSP have been proposed for scheduling requests in Web servers. SRPT and FSP are superior to policies that ignore request size, such as PS, in both efficiency and fairness, given heavy-tailed service times. However, a central assumption that is usually made in implementing size-based policies in a Web server is that the service time of a request is strongly correlated with the size of the file it serves. By collecting Web server trace data taken from the logs of modified Apache Web servers, this paper reveals that the correlation between service time and file size can be quite low, and shows how the performance of SRPT and FSP can be dramatically affected by the weak correlation via trace-driven simulations. In response, we propose and evaluate domain-based scheduling, a simple technique that better estimates connection times by making use of the source IP address of the request. Domain-based scheduling improves SRPT and FSP performance on Web servers, bringing the performance benefits of these scheduling polices even to those regimes where the correlation between file size and service time is low.
Peter A. Dinda, Yi Qiao, Huanyuan Sheng
MASCOTS3
2005 Elders Know Best - Handling Churn in Less Structured P2P Systems
abstract
We address the problem of highly transient populations in unstructured and loosely-structured peer-to-peer systems. We propose a number of illustrative query-related strategies and organizational protocols that, by taking into consideration the expected session times of peers (their lifespans), yield systems with performance characteristics more resilient to the natural instability of their environments. We first demonstrate the benefits of lifespan-based organizational protocols in terms of end-application performance and in the context of dynamic and heterogeneous Internet environments. We do this using a number of currently adopted and proposed query-related strategies, including methods for query distribution, caching and replication. We then show, through trace-driven simulation and wide-area experimentation, the performance advantages of lifespan-based, query-related strategies when layered over currently employed and lifespan-based organizational protocols. While merely illustrative, the evaluated strategies and protocols clearly demonstrate the advantages of considering peers' session time in designing widely-deployed peer-to-peer systems.
Yi Qiao, Fabián E. Bustamante
Peer-to-Peer Computing1
2004 An Empirical Study of the Multiscale Predictability of Network Traffic
Yi Qiao, Jason A. Skicewicz, Peter A. Dinda
HPDC1
1995 A new stochastic projection-based image recovery method
abstract
The method of projections onto convex sets has been used effectively in the solution many important signal and image processing applications. This method however fails in circumstances when the desired constraints form non-convex sets. A new approach to image recovery based on a stochastic projections method is presented. At its core the method relies on the iteration of projections onto closed sets (not necessarily convex). A stochastic parameter is used to update the projections among the images in their neighborhood subsequent to each iteration. Computer simulation experiments are used for the comparison of the stochastic projections with several existing methods to signal recovery for phase-from-magnitude image restoration.
Dan Schonfeld, Yi Qiao
ICIP2