VLDB 2026 Research / reviewers in the wild / expert
Xihao Li
dblp:147/4915
· DBLP profile ↗
8ranked-venue papers
2as first author
8since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 4 since 2021Theory of computation · 2 · 1 first-author · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A novel two-sample Mendelian randomization framework integrating common and rare variants: application to assess the effect of HDL-C on preeclampsia riskabstractMendelian randomization (MR) has become an important technique for establishing causal relationships between risk factors and health outcomes. By using genetic variants as instrumental variables, it can mitigate bias due to confounding and reverse causation in observational studies. Current MR analyses have predominantly used common genetic variants as instruments, which represent only part of the genetic architecture of complex traits. Rare variants, which can have larger effect sizes and provide unique biological insights, have been understudied due to statistical and methodological challenges. We introduce MR-common and annotation-informed rare variants (MR-CARV), a novel framework integrating common and rare genetic variants in two-sample MR. This method leverages comprehensive genetic data made available by high-throughput sequencing technologies and large-scale consortia. Rare variants are aggregated into functional categories, such as gene-coding, gene-noncoding, and nongene regions, by leveraging variant annotations and biological impact as weights. The effects of rare variant sets are then estimated with STAARpipeline and combined with the estimated effects of common variants by the existing MR methods. Simulation studies demonstrate that MR-CARV maintains robust type I error and achieves higher statistical power, with up to a 66.3% relative increase compared with existing methods only based on common variants. Consistent with these findings, application to real data on high-density lipoprotein cholesterol (HDL-C) and preeclampsia showed that MR-CARV [inverse variance weighted (IVW)] yielded a more precise and statistically significant effect estimate (-0.020, SE = 0.0102, $P$ =.0470) than IVW using only common variants (-0.023, SE = 0.0123, $P$ =.0659). David M. Haas, C. Noel Bairey Merz, Tsegaselassie Workalemahu, Kelli Ryckman, Janet M. Catov, Lisa D. Levine, Alexa Freedman, George R. Saade, Jiaqi Hu 0004, Hongyu Zhao 0003, Xihao Li, Nianjun Liu |
Briefings Bioinform. | 13 |
| 2025 | Integrating Audio, Visual, and Semantic Information for Enhanced Multimodal Speaker Diarization on Multi-party ConversationabstractLuyao Cheng, Hui Wang, Chong Deng, Siqi Zheng, Yafeng Chen, Rongjie Huang, Qinglin Zhang, Qian Chen, Xihao Li, Wen Wang. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Luyao Cheng, Hui Wang 0030, Chong Deng, Yafeng Chen, Rongjie Huang 0001, Qian Chen 0003, Xihao Li, Wen Wang 0001 |
ACL (1) | 9 |
| 2025 | 3D-Speaker-Toolkit: An Open-Source Toolkit for Multimodal Speaker Verification and DiarizationabstractWe introduce 3D-Speaker-Toolkit, an open-source toolkit for multimodal speaker verification and diarization, designed for meeting the needs of academic researchers and industrial practitioners. The 3D-Speaker-Toolkit adeptly leverages the combined strengths of acoustic, semantic, and visual data, seamlessly fusing these modalities to offer robust speaker recognition capabilities. The acoustic module extracts speaker embeddings from acoustic features, employing both fully-supervised and self-supervised learning approaches. The semantic module leverages advanced language models to comprehend the substance and context of spoken language, thereby augmenting the system’s proficiency in distinguishing speakers through linguistic patterns. The visual module applies image processing technologies to scrutinize facial features, which bolsters the precision of speaker diarization in multi-speaker environments. Collectively, these modules empower the 3D-Speaker-Toolkit to achieve substantially improved accuracy and reliability in speaker-related tasks. With 3D-Speaker-Toolkit, we establish a new benchmark for multimodal speaker analysis. The toolkit also includes a handful of open-source state-of-the-art models and a large-scale dataset containing over 10,000 speakers. The toolkit is publicly available at https://github.com/modelscope/3D-Speaker. Yafeng Chen, Hui Wang 0030, Luyao Cheng, Tinglong Zhu, Rongjie Huang 0001, Chong Deng, Qian Chen 0003, Shiliang Zhang, Wen Wang 0001, Xihao Li |
ICASSP | 11 |
| 2025 | List Viterbi Algorithm Aided Low-Latency OSDabstractOrdered statistics decoding (OSD) requires Gaussian elimination (GE) to obtain the systematic generator matrix (SGM) for yielding codeword candidates, resulting in an uncompromised latency. Recently, low-latency OSD (LLOSD) has been proposed for BCH codes to avoid GE by computing the SGM of a Reed-Solomon (RS) code. Since BCH codes are binary subcodes of the RS codes, the BCH codeword candidates can be yielded with the RS SGM. To further facilitate the LLOSD, this paper proposes a local constraint-based LLOSD (LC-LLOSD). In particular, the RS SGM is converted into a binary BCH parity-check matrix, whose submatrix can be used to specify a trellis. The serial list Viterbi algorithm (SLVA) can be applied to generate extended test messages (TMs). Further, it can be facilitated by incorporating the TM generation scheme of the LLOSD. Since the SLVA generates the TMs in decreasing likelihood, the LCLLOSD can yield better decoding performance while re-encoding far less TMs compared to the LLOSD. Simulation results show the LC-LLOSD's complexity advantage over the LLOSD. Xihao Li, Li Chen 0011 |
ISIT | 1 |
| 2025 | A comprehensive comparison on clustering methods for multi-slice spatially resolved transcriptomics data analysisabstractSpatial transcriptomics (ST) data, by providing spatial information, enable simultaneous analysis of gene expression distributions and their spatial patterns within tissue. Clustering or spatial domain detection represents an essential methodology for ST data, facilitating the exploration of spatial organizations with shared gene expression or histological characteristics. Traditionally, clustering algorithms for ST have focused on individual tissue sections. However, the emergence of numerous contiguous tissue sections derived from the same or similar tissue specimens within or across individuals has led to the development of multi-slice clustering methods. In this study, we assess seven single-slice and four multi-slice clustering methods on two simulated datasets and four real datasets. Additionally, we investigate the effectiveness of preprocessing techniques, including spatial coordinate alignment (e.g. PASTE) and gene expression batch effect removal (e.g. Harmony), on clustering performance. Our study provides a comprehensive comparison of clustering methods for multi-slice ST data, serving as a practical guide for method selection in various scenarios. Caiwei Xiong, Muqing Zhou, Wenrong Wu, Xihao Li, Huaxiu Yao, Jiawen Chen 0002, Yun Li 0010 |
Briefings Bioinform. | 6 |
| 2025 | Efficient Ordered Statistics Decoding of BCH Codes Without Gaussian EliminationabstractOrdered statistics decoding (OSD) can achieve near maximum likelihood (ML) decoding performance for BCH codes. However, Gaussian elimination (GE) that delivers the systematic generator matrix of the code has an uncompromised latency. Addressing this challenge, this paper proposes a low-latency OSD (LLOSD) for BCH codes. Since BCH codes are binary subcodes of Reed-Solomon (RS) codes, codeword candidates can be produced using the RS systematic generator matrix, whose entries can be generated in parallel. By eliminating the non-binary codeword candidates and identifying the ML codeword, the LLOSD yields a lower latency as well as complexity than the OSD. It is shown that the LLOSD can be interpreted as generating the codeoword candidates through systematic encoding of a punctured BCH codeword, explaining its low-complexity feature. Moreover, the segmented variant is proposed to further facilitate the LLOSD. In order to decode long BCH codes, a hybrid soft decoding (HSD) is finally proposed. It integrates the LLOSD and the algebraic Chase decoding that can effectively provide extra TEPs for the LLOSD, enhancing the decoding performance. Both the complexity and performance of the proposed decoding are analyzed, demonstrating their advantage over the relevant state-of-the-art decoding. Lijia Yang, Xihao Li, Li Chen 0013, Huazi Zhang, Jiajie Tong |
IEEE Trans. Inf. Theory | 3 |
| 2024 | Order Skipping Ordered Statistics Decoding and its Performance AnalysisabstractThis paper proposes a reduced complexity ordered statistics decoding (OSD) algorithm for linear block codes, the namely order skipping (OS)-OSD algorithm. An approximated correlation distance lower bound (CDLB) is derived by utilizing likelihood of the received symbols over the least reliable positions (LRPs). It enables the assessment of whether the higher-order decoding can yield a more likely codeword estimation. If not, they can be skipped. Error-correction performance of the OSOSD is analyzed. In particular, the decoding error probability of OS-OSD with order one is theoretically characterized. Our simulation results verify that the OS-OSD can achieve a significant complexity reduction over the state-of-the-art OSD without compromising the decoding performance. Xihao Li, Li Chen 0013, Yuan Li 0034, Huazi Zhang |
ITW | 1 |
| 2022 | STAAR workflow: a cloud-based workflow for scalable and reproducible rare variant analysisabstractSUMMARY: We developed the variant-Set Test for Association using Annotation infoRmation (STAAR) workflow description language (WDL) workflow to facilitate the analysis of rare variants in whole genome sequencing association studies. The open-access STAAR workflow written in the WDL allows a user to perform rare variant testing for both gene-centric and genetic region approaches, enabling genome-wide, candidate and conditional analyses. It incorporates functional annotations into the workflow as introduced in the STAAR method in order to boost the rare variant analysis power. This tool was specifically developed and optimized to be implemented on cloud-based platforms such as BioData Catalyst Powered by Terra. It provides easy-to-use functionality for rare variant analysis that can be incorporated into an exhaustive whole genome sequencing analysis pipeline. AVAILABILITY AND IMPLEMENTATION: The workflow is freely available from https://dockstore.org/workflows/github.com/sheilagaynor/STAAR_workflow. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Sheila M. Gaynor, Kenneth E. Westerman, Lea L. Ackovic, Xihao Li, Zilin Li, Alisa Manning, Anthony A. Philippakis, Xihong Lin |
Bioinform. | 4 |