Xinmei Huang

dblp:99/5472 · DBLP profile ↗
← Back
10ranked-venue papers
2as first author
7since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 3 · 2 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 3 since 2021Theory of computation · 2 · 1 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Security and privacy · 1 · 1 first-author
YearPublicationVenuePosition
2026 Llmia: an Out-Of-The-Box Index Advisor Via in-Context Learning With Llms
abstract
Index recommendation is crucial for optimizing database performance. However, existing heuristic- and learning-based methods often rely on inefficient exhaustive search and estimated costs, leading to low efficiency (due to the vast search space) and unsatisfactory actual latency (due to inaccurate estimations). Inspired by the refinement strategies of experienced DBAs-who efficiently identify and iteratively refine indexes with database feedback-we present LLMIA, an out-of-the-box, tuning-free index advisor leveraging large language models (LLMs) through in-context learning for index recommendation. LLMIA injects database expertise into the LLM using a high-quality demonstration pool and comprehensive workload feature extraction, while iteratively incorporating database feedback to guide the index refinement. This design enables LLMIA to emulate the decision-making process of expert DBAs: efficiently recommending and refining indexes for various workloads within just a few interactions with the DBMS. We validate LLMIA with extensive experiments on five standard OLAP benchmarks (TPC-H with different scales, JOB, TPC-DS, SSB), where it consistently outperforms or matches 12 baselines by producing superior index recommendations with minimal database interactions. Additionally, LLMIA demonstrates robust generalization on two real-world commercial workloads, delivering high-quality recommendations without the need for additional adaptation or retraining, highlighting its out-of-the-box capability.
Xinxin Zhao, Xinmei Huang, Haoyang Li 0015, Jing Zhang 0001, Tieying Zhang, Jianjun Chen 0001, Cuiping Li 0001, Hong Chen 0001
ICDE2
2025 E2ETune: End-to-End Knob Tuning via Fine-tuned Generative Language Model
Xinmei Huang, Haoyang Li 0015, Jing Zhang 0001, Xinxin Zhao, Zhiming Yao, Yiyan Li, Tieying Zhang, Jianjun Chen 0001, Hong Chen 0001, Cuiping Li 0001
Proc. VLDB Endow.1
2025 OmniSQL: Synthesizing High-quality Text-to-SQL Data at Scale
abstract
Text-to-SQL, the task of translating natural language questions into SQL queries, plays a crucial role in enabling non-experts to interact with databases. While recent advancements in large language models (LLMs) have significantly enhanced text-to-SQL performance, existing approaches face notable limitations in real-world text-to-SQL applications. Prompting-based methods often depend on closed-source LLMs, which are expensive, raise privacy concerns, and lack customization. Fine-tuning-based methods, on the other hand, suffer from poor generalizability due to the limited coverage of publicly available training data. To overcome these challenges, we propose a novel and scalable text-to-SQL data synthesis framework for automatically synthesizing large-scale, high-quality, and diverse datasets without extensive human intervention. Using this framework, we introduce SynSQL-2.5M, the first million-scale text-to-SQL dataset, containing 2.5 million samples spanning over 16,000 synthetic databases. Each sample includes a database, SQL query, natural language question, and chain-of-thought (CoT) solution. Leveraging SynSQL-2.5M, we develop OmniSQL, a powerful open-source text-to-SQL model available in three sizes: 7B, 14B, and 32B. Extensive evaluations across nine datasets demonstrate that OmniSQL achieves state-of-the-art performance, matching or surpassing leading closed-source and open-source LLMs, including GPT-4o and DeepSeek-V3, despite its smaller size. We release all code, datasets, and models to support further research.
Haoyang Li 0015, Xinmei Huang, Jing Zhang 0001, Fuxin Jiang, Tieying Zhang, Jianjun Chen 0001, Hong Chen 0001, Cuiping Li 0001
Proc. VLDB Endow.4
2024 On the Trade-Off Between Communication Reliability and Latency in the Absence of Feedback
abstract
Reliability and latency are two key performance indicators of communications. This paper investigates the tradeoff between them over a random packet erasure channel in the absence of feedback. In contrast to the instant feedback case where guaranteed reliability with low latency can be achieved by simple automatic repeat query (ARQ), we show that in the absence of feedback, the reliability increases with the coding window size, at the cost of the degraded latency performance. Specifically, we propose a sliding window network coding (SWNC) scheme that works in the absence of feedback and achieves various reliability and latency tradeoff by adjusting the coding window size. The tradeoff between the reliability and latency are investigated by deriving the achievable performance of the proposed SWNC scheme as a function of the coding window size. We further show that the proposed scheme degenerates to the existing benchmark schemes that are superior in either latency or reliability, and it has much higher flexibility.
Zhicheng Zhu, Xiaoli Xu 0001, Yong Zeng 0001, Xinmei Huang
WCNC4
2024 Age-Optimal Packet Scheduling With Resource Constraint and Feedback Delay
abstract
This paper considers the design of age-optimal packet scheduling policies in networks with delayed feedback, long-term resource constraint and various packet arrival models. The problem is formulated as a constrained partial observable Markov decision process (CPOMDP) in general. We first derive the achievable age of information (AoI) by the random and determined transmission policies, which works in the absence of feedback. Then, we propose a greedy policy that traces the expected receiver AoI based on the delayed feedback, and selects the action that minimizes the expected immediate Lagrange cost defined in the formulated CPOMDP. The proposed policy outperforms the existing policies for a wide range of feedback delay, even in the absence of feedback, yet the implementation complexity is still low since the decision threshold is derived in the closed-form, as a function of the network parameters.
Yonghao Ji, Yuxiao Lu, Xiaoli Xu 0001, Xinmei Huang
IEEE Trans. Commun.4
2023 FC-KBQA: A Fine-to-Coarse Composition Framework for Knowledge Base Question Answering
abstract
Lingxi Zhang, Jing Zhang, Yanling Wang, Shulin Cao, Xinmei Huang, Cuiping Li, Hong Chen, Juanzi Li. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Lingxi Zhang, Jing Zhang 0001, Shulin Cao, Xinmei Huang, Cuiping Li 0001, Hong Chen 0001, Juan-Zi Li
ACL (1)5
2021 Hulls of Generalized Reed-Solomon Codes via Goppa Codes and Their Applications to Quantum Codes
abstract
A Goppa code over \Bbb Fqmis a well-known subclass of algebraic error-correcting code. If m=1, then it is a generalized Reed-Solomon(GRS) code and its dual code is called a GRS code via a Goppa code. In this paper, we give a necessary and sufficient condition that the dual codes of GRS codes via (expurgated) Goppa codes are also GRS codes via Goppa codes. Under the above condition, we show that the hulls of GRS codes via Goppa codes are still GRS codes via Goppa codes. As an application, we characterize LCD GRS codes and self-dual GRS codes under the above condition. Some numerical examples are also presented to illustrate our main results. Moreover, we also apply our result to entanglement-assisted quantum error correcting codes (EAQECCs) and obtain two new families of MDS EAQECCs with arbitrary parameters.
Yanyan Gao 0003, Qin Yue 0001, Xinmei Huang, Jun Zhang 0031
IEEE Trans. Inf. Theory3
2020 Binary primitive LCD BCH codes
Xinmei Huang, Qin Yue 0001, Yansheng Wu, Xiaoping Shi 0002, Jerod Michel
Des. Codes Cryptogr.1
2013 Check Node Reliability-Based Scheduling for BP Decoding of Non-Binary LDPC Codes
abstract
Scheduling strategy is considered an important aspect of belief-propagation (BP) decoding of low-density parity-check (LDPC) codes because it affects the decoder's convergence rate, decoding complexity and error-correction performance. In this paper, we propose two new scheduling strategies for the BP decoding of non-binary LDPC (NB-LDPC) codes. Both the strategies are devised based on the concept of check node reliability and employ a heuristically defined threshold which can adapt to the communication channel variations. As the scheduling strategies only update a subset of the check nodes in each iteration, they result in reduced iteration cost. Furthermore, since the BP performs suboptimally for finite-length LDPC codes, especially for short-length LDPC codes, by enhancing the message propagation over the Tanner Graphs of short-length NB-LDPC codes, the new scheduling strategies can even improve the error-correction performances of BP decoding. Simulation results demonstrate that the new scheduling strategies provide good performance/complexity tradeoffs.
Guojun Han, Yong Liang Guan 0001, Xinmei Huang
IEEE Trans. Commun.3
2009 A Note on Limited-Trial Chase-Like Algorithms Achieving Bounded-Distance Decoding
abstract
For the decoding of a binary linear block code of minimal Hamming distancedover additive white Gaussian noise (AWGN) channels, a soft-decision decoder achieves bounded-distance (BD) decoding if its squared error-correction radius is equal tod. A Chase-like algorithm outputs the best (most likely) codeword in a list of candidates generated by a conventional algebraic binary decoder in a few trials. It is of interest to design Chase-like algorithms that achieve BD decoding with as least trials as possible. In this paper, we show that Chase-like algorithms can achieve BD decoding with onlyO(d1/2+epsiv) trials for any given positive numberepsiv.
Yuansheng Tang, Xinmei Huang
IEEE Trans. Inf. Theory2