Zhichao Geng

dblp:157/4479 · DBLP profile ↗
← Back
9ranked-venue papers
4as first author
7since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 5 · 2 first-author · 4 since 2021Theory of computation · 4 · 2 first-author · 2 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Why Advanced Encoders Lag on Sparse Retrieval? The Answer and an Approach to Bridging Vocabulary Gaps
abstract
While advanced foundation models like ModernBERT significantly outperform older architectures in dense retrieval, they surprisingly lag behind the aging BERT-base baseline in learned sparse retrieval (LSR). We identify the root cause as the Vocabulary Gap : modern tokenizers utilize raw, case-sensitive vocabularies designed for lossless reconstruction, which map single semantic units to redundant surface forms, wasting model capacity on morphological noise and hindering lexical matching. We formalize this intuition through a theoretical framework, demonstrating that appropriate vocabulary coarse-graining can tighten the generalization bounds by reducing complexity of the hypothesis class, provided that semantic integrity is preserved. To resolve this, we propose Vocabulary Transfer (VT), a model-agnostic framework that migrates advanced encoders to sparse-friendly, normalized vocabularies with minimal computational cost. VT utilizes a novel Semantic Initialization via spatial topology to preserve geometric structure and an Activation Potential Calibration (APC) mechanism to align pre-trained manifolds with sparsity constraints, preventing the dead neuron and dense collapse observed in standard fine-tuning. Empirically, VT is universally effective: it enables ModernBERT to achieve state-of-the-art performance on the BEIR benchmark (52.4 nDCG, a +4.7 improvement), resuscitates failing models like RoBERTa-large, and generalizes seamlessly to inference-free architectures and specialized domains. These results confirm that the performance lag is not an architectural deficiency but a solvable vocabulary mismatch. We've released our code and models. https://anonymous.4open.science/r/vocab-transfer/. All details included.
Zhichao Geng, Yang Yang 0222
SIGIR1
2026 Surface-Form Neural Sparse Retrieval: Robust Fuzzy Matching for Industrial Music Search
abstract
Music search at the scale of Amazon Music presents a unique challenge: queries frequently deviate from indexed metadata due to misspellings, transpositions, and phonetic variations, yet the retrieval system must operate under strict millisecond-level latency constraints. Our existing learning-to-retrieve system, the High Confidence Index (HCI), learns query-entity associations from customer behavior, relying on continual "exploration" to choose candidates. Traditional n-gram matching enables this exploration but suffers from poor semantic robustness and high noise, limiting the system's ability to learn from long-tail queries. In this work, we present a robust neural sparse retrieval system designed to maximize exploration efficiency. We adapt a state-of-the-art inference-free sparse retrieval architecture to the music domain, combining it with an effective domain-specific granular subword tokenization strategy. Our approach utilizes short-length token constraints (max 3 chars) to enforce the learning of surface-form robustness over lexical memorization. By pre-computing the neural embeddings and term expansions during the offline indexing phase, online processing is reduced to minimal tokenization and IDF weighting, achieving effectively zero latency overhead for query encoding. Evaluations on a 6M-document production corpus show an aggregate 91.4% recall@10 (vs. 57.7% for trigrams) at comparable throughput. Simulation of the HCI feedback loop demonstrates improved exploration efficiency, with +0.8% higher stabilized recall than production trigrams. Ablation studies indicate that our sparse training methodology drives the performance gains, while domain-specific pretraining provides a cost-effective alternative to large-scale general-purpose pretraining.
Paul Greyson, Zhichao Geng, Yang Yang 0222
SIGIR2
2026 Multiple parallel-batch machines scheduling with additive resource assignment and machine available times
Zhichao Geng, Renxia Chen
Discret. Appl. Math.3
2025 Exploring ℓ0 Sparsification for Inference-free Sparse Retrievers
abstract
With increasing demands for efficiency, information retrieval has developed a branch of sparse retrieval, further advancing towards inference-free retrieval where the documents are encoded during indexing time and there is no model-inference for queries. Existing sparse retrieval models rely on FLOPS regularization for sparsification, while this mechanism was originally designed for Siamese encoders, it is considered to be suboptimal in inference-free scenarios which is asymmetric. Previous attempts to adapt FLOPS for inference-free scenarios have been limited to rule-based methods, leaving the potential of sparsification approaches for inference-free retrieval models largely unexplored. In this paper, we explore ℓ0 inspired sparsification manner for inference-free retrievers. Through comprehensive out-of-domain evaluation on the BEIR benchmark, our method achieves state-of-the-art performance among inference-free sparse retrieval models and is comparable to leading Siamese sparse retrieval models. Furthermore, we provide insights into the trade-off between retrieval effectiveness and computational efficiency, demonstrating practical value for real-world applications.
Xinjie Shen, Zhichao Geng, Yang Yang 0222
SIGIR2
2024 CPT: a pre-trained unbalanced transformer for both Chinese language understanding and generation
Yunfan Shao, Zhichao Geng, Yitao Liu, Junqi Dai, Hang Yan 0001, Li Zhe, Hujun Bao, Xipeng Qiu
Sci. China Inf. Sci.2
2023 Bicriteria scheduling on an unbounded parallel-batch machine for minimizing makespan and maximum cost
Shuguang Li 0003, Zhichao Geng
Inf. Process. Lett.2
2022 Improving Abstractive Dialogue Summarization with Speaker-Aware Supervised Contrastive Learning
abstract
Pre-trained models have brought remarkable success on the text summarization task. For dialogue summarization, the subdomain of text summarization, utterances are concatenated to flat text before being processed. As a result, existing summarization systems based on pre-trained models are unable to recognize the unique format of the speaker-utterance pair well in the dialogue. To investigate this issue, we conduct probing tests and manual analysis, and find that the powerful pre-trained model can not identify different speakers well in the conversation, which leads to various factual errors. Moreover, we propose three speaker-aware supervised contrastive learning (SCL) tasks: Token-level SCL, Turn-level SCL, and Global-level SCL. Comprehensive experiments demonstrate that our methods achieve significant performance improvement on two mainstream dialogue summarization datasets. According to detailed human evaluations, pre-trained models equipped with SCL tasks effectively generate summaries with better factual consistency.
Zhichao Geng, Ming Zhong 0005, Zhangyue Yin, Xipeng Qiu, Xuanjing Huang 0001
COLING1
2015 A note on unbounded parallel-batch scheduling
Zhichao Geng, Jinjiang Yuan
Inf. Process. Lett.1
2015 Pareto optimization scheduling of family jobs on a p-batch machine to minimize makespan and maximum lateness
Zhichao Geng, Jinjiang Yuan
Theor. Comput. Sci.1