Feilong Bao

dblp:136/5318 · DBLP profile ↗
← Back
9ranked-venue papers in the field
0as first author
7since 2021 · last 2026
0000-0001-7312-1629ORCID · corroborated

Domains — venue-derived; a paper can count in several

Information Retrieval & Web Search · 4Knowledge Engineering, Semantic Web & Information Systems · 3Other / Interdisciplinary · 2
YearPublicationVenuePosition
2026 Construction and Evaluation of Large Language Models for the Mongolian Medicine Diagnostic and Treatment System
Jixieqi Bai, Feilong Bao, Hui Zhang 0031, Aruukhan Bai
KSEM (6)2
2026 TM-Bench: Benchmarking Large Language Models on Low-Resource Traditional Mongolian
abstract
Large language models (LLMs) have achieved remarkable success in high-resource languages, yet their performance on Traditional Mongolian remains highly limited. A primary bottleneck is the absence of a systematic evaluation framework, which precludes quantitative comparison and obscures directions for model optimization. In this paper, we introduce TM-Bench, the first comprehensive benchmark for LLMs on Traditional Mongolian. TM-Bench adopts a hybrid construction strategy consisting of human-verified Translation-based Adaptation, Expert-Original Authoring, and Semi-automated Synthesis. It comprises 18,357 instances spanning five tasks across both natural language understanding and generation to evaluate models' reasoning, knowledge application, and linguistic proficiency. We conduct systematic evaluations across representative model families. The results show that on understanding tasks, model performance lags significantly behind high-resource languages, with only a few models performing slightly above the random baseline. For generation tasks, both automatic metrics and double-blind human evaluations reveal severe semantic collapse, failing to generate coherent text and often producing unreadable gibberish. These findings underscore the critical role of TM-Bench as a foundational infrastructure for evaluating LLMs in Traditional Mongolian and catalyzing future model optimization. Our benchmark and code are available at https://github.com/gao1948083886/TM-Bench.
Zhenjie Gao, Feilong Bao, Aruukhan Bai, Ruichen Hou, Xieqi Ji, Dabalgan Wang, Hugjil Ming
SIGIR2
2026 Selective Distillation for Continual Named Entity Recognition with Memory Replay
Weihua Wang 0006, Feilong Bao
SIGIR3
2026 How to teach and forget: Towards cross-modal semantic consistency for entity alignment
Cunda Wang, Chenglong Miao, Po Hu 0001, Weihua Wang 0006, Feilong Bao
Inf. Process. Manag.5
2025 Zero-Shot Speech Recognition from Text-Only Data through Synthesized Spectrogram Refinement Using Style Truncation and Contextual Alignment Loss
abstract
Utilizing pseudo speech-label pairs synthesized via Text-to-Speech (TTS) systems as supplementary training data for automatic speech recognition (ASR) has shown significant benefits. However, the mismatch between synthesized and real speech make them unsuitable for direct use in zero-shot speech recognition tasks. In this paper, we propose a generative-adversarial model with a style truncation strategy that enhances mel-spectrograms by projecting style vectors into high-density regions. We also introduce a Contextual alignment loss function to align synthesized and real mel-spectrograms by computing global distances between high-dimensional features. Experimental results on zero-shot ASR tasks using the SLURP and LibriSpeech datasets show that our method achieves WER reductions of 4.3%, 11.1% (clean), and 6.9% (other) compared to the baseline mel-spectrograms enhanced model.
Yonghe Wang, Zhenjie Gao, Feilong Bao
MMAsia4
2025 Domain disentanglement and fusion based on hyperbolic neural networks for zero-shot sketch-based image retrieval
Xiangdong Su, Yonghe Wang, Feilong Bao, Guanglai Gao
Inf. Process. Manag.5
2021 Panoptic-DLA: Document Layout Analysis of Historical Newspapers Based on Proposal-Free Panoptic Segmentation Model
Feilong Bao, Guanglai Gao
KSEM2
2019 An Automatic Spelling Correction Method for Classical Mongolian
Feilong Bao, Guanglai Gao, Weihua Wang 0006, Hui Zhang 0031
KSEM (2)2
2017 Segmentation-Free Printed Traditional Mongolian OCR Using Sequence to Sequence with Attention Model
abstract
Mongolian Optical Character Recognition (OCR) systems are required for printed document digitization and Mongolian cultural resources utilization. Existing Mongolian OCR systems are based on segmentation. But, the Mongolian segmentation is more difficult than other languages. So, these methods are highly costly and error suffering. In this study, a segmentation-free based traditional Mongolian word recognition method is proposed. Specifically, we formalize the OCR task as a sequence to sequence mapping problem, in which the input Mongolian word image and the output textual string are treated as a sequence of image frames and a sequence of letters, respectively. A sequence to sequence with attention model is adopted to solve this problem. Experimental results on a dataset show the effectiveness of the proposed method.
Hui Zhang 0031, Hongxi Wei, Feilong Bao, Guanglai Gao
ICDAR3