Jiu Sha

dblp:270/5958 · DBLP profile ↗
← Back
4ranked-venue papers
3as first author
4since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 4 · 3 first-author · 4 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
2 papers
Vision and language · 50% Language models and text generation · 32% Image recognition and object detection · 18%

Topics — the 5 heaviest of 6, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › Image recognition and object detection › text recognition
optical character recognition
1.012026
Beyond Atomic Characters: Glyph-Aware Sub-character Alignment for Low-Resource Multilingual OCR · ACL (1) 2026
Natural language and speech › Language models and text generation › evaluation of language models › multilingual evaluation
low-resource language evaluation
0.912025
TVQACML: Benchmarking Text-Centric Visual Question Answering in Multilingual Chinese Minority Languages · EMNLP 2025
Natural language and speech › Language models and text generation › evaluation of language models
multilingual evaluation
0.912025
TVQACML: Benchmarking Text-Centric Visual Question Answering in Multilingual Chinese Minority Languages · EMNLP 2025
Computer vision › Vision and language › visual question answering
text-based visual question answering
0.912025
TVQACML: Benchmarking Text-Centric Visual Question Answering in Multilingual Chinese Minority Languages · EMNLP 2025
Computer vision › Vision and language
visual question answering
0.912025
TVQACML: Benchmarking Text-Centric Visual Question Answering in Multilingual Chinese Minority Languages · EMNLP 2025

Methods — techniques the papers use, named apart from their topics

reverse synthesis · 1.0glyph-aware adapter · 1.0curriculum learning · 1.0instruction tuning · 0.9benchmark construction · 0.9
YearPublicationVenuePosition
2026 Beyond Atomic Characters: Glyph-Aware Sub-character Alignment for Low-Resource Multilingual OCR
abstract
Low-resource multilingual OCR faces a dual challenge: complex script structures and severe data scarcity.In such settings, existing OCR models often struggle, as coarse visual representations combined with weak linguistic priors lead to frequent errors among visually similar characters.To address this, we present BASA (Beyond Atomic Sub-character Alignment), a OCR framework built upon highresolution visual and language backbones with a novel glyph-aware interface.The core technical contribution is the Glyph-Aware Finegrained Adapter (GAFA).Unlike standard linear projectors, GAFA employs learnable glyph prototypes to actively align sub-character structural primitives (e.g., strokes and radicals) with visual features, explicitly resolving topological ambiguities during vision-language alignment.To complement this, we introduce a twostage curriculum learning strategy supported by a Glyph-Aware Reverse Synthesis pipeline, which generates large-scale multilingual training corpora with automatic, zero-cost component labels.Furthermore, we construct BASA-Bench, a representative benchmark spanning 11 languages with diverse script structures and 23 authentic scenarios.Experiments demonstrate that BASA achieves consistent improvements over strong OCR baselines, particularly on scripts with complex compositions.Our model and benchmark will be available at https://github.com/NcutLLM/BASA.
Mengxiao Zhu 0004, Jiu Sha
ACL (1)3
2025 VEEF-Multi-LLM: Effective Vocabulary Expansion and Parameter Efficient Finetuning Towards Multilingual Large Language Models
abstract
Large Language Models(LLMs) have brought significant transformations to various aspects of human life and productivity. However, the heavy reliance on vast amounts of data in developing these models has resulted in a notable disadvantage for low-resource languages, such as Nuosu and others, which lack large datasets. Moreover, many LLMs exhibit significant performance discrepancies between high-and lowresource languages, thereby restricting equitable access to technological advances for all linguistic communities. To address these challenges, this paper propose a low-resource multilingual large language model, termed VEEF-Multi-LLM, constructed through effective vocabulary expansion and parameter-efficient fine-tuning. We introduce a series of innovative methods to address challenges in low-resource languages. First, we adopt Byte-level Byte-Pair Encoding to expand the vocabulary for broader multilingual support. We separate input and output embedding weights to boost performance, and apply RoPE for long-context handling, as well as RMSNorm for efficient training. To generate high-quality supervised fine-tuning (SFT) data, we use self-training and selective translation, and refine the resulting dataset with the assistance of native speakers to ensure cultural and linguistic accuracy. Our model, VEEF-Multi-LLM-8B, is trained on 600 billion tokens across 50 natural and 16 programming languages. Experimental results show that the model excels in multilingual instruction-following tasks, particularly in translation, outperforming competing models in benchmarks such as XCOPA and XStoryCloze. Although it lags slightly behind English-centric models in some tasks (e.g., m-MMLU), it prioritizes safety, reliability, and inclusivity, making it valuable for diverse linguistic communities. We open-source our models on GitHub and Huggingface.
Jiu Sha, Mengxiao Zhu 0004, Chong Feng 0001, Yuming Shang
COLING1
2025 TVQACML: Benchmarking Text-Centric Visual Question Answering in Multilingual Chinese Minority Languages
abstract
Text-Centric Visual Question Answering (TEC-VQA) serves as a key benchmark for evaluating AI's ability to reason over text-rich visual scenes.However, most existing TEC-VQA datasets focus on high-resource languages and are susceptible to benchmark contamination due to overlap with pretraining corpora of large models.These limitations severely hinder progress in low-resource language scenarios and compromise the reliability of current evaluations.To address both the underrepresentation of low-resource languages and the contamination issue, we propose TVQACML, the first large-scale TEC-VQA benchmark for multilingual Chinese minority languages, constructed through a scalable, reproducible pipeline.It comprises 8,000 real-world images and 32,000 high-quality QA pairs across eight languages and 30 application scenarios.We conduct comprehensive benchmarking of open-source, closed-source, and text-centric MLLMs, revealing substantial performance gaps from human accuracy, especially in scenetext and document understanding tasks.Furthermore, instruction tuning with TVQACML yields consistent performance gains, in some cases surpassing leading closed models demonstrating the dataset's utility for model alignment.We also introduce a lightweight, extensible evaluation metric for robust multilingual, multi-format answer assessment.The code and dataset for TVQACML are available at https://github.com/Shajiu/TVQACML.
Jiu Sha, Mengxiao Zhu 0004, Chong Feng 0001, Jialedongzhu
EMNLP1
2025 Tibetan-LLaMA 2: Large Language Model for Tibetan
abstract
Large language models (LLMs), such as ChatGPT and LLama, have shown remarkable capability in a wide range of natural language tasks. However, the current LLMs are mainly concentrated in resource-rich languages, such as English and Chinese. For low-resource language such as Tibetan, research and applications related to LLMs are still in their infancy. To address the existing gap, we present a method to enhance LLaMA with the ability to understand and generate Tibetan text, as well as to follow instructions. This is achieved by creating large-scale unsupervised pre-training and supervised fine-tuning datasets, which mitigate the limited availability of Tibetan data. Additionally, we expand LLaMA’s vocabulary by incorporating Tibetan tokens through Unigram tokenization, thereby improving both its encoding efficiency and semantic understanding of Tibetan. Furthermore, we conduct secondary pre-training and fine-tune the model using the constructed datasets, thereby enhancing its capability to interpret and execute instructions effectively. To verify the effectiveness of the model, we establish ten evaluation benchmarks for Tibetan. The experimental results indicate that the proposed model significantly enhances the LLaMA’s proficiency in understanding and generating Tibetan content. To promote further research, we release our model and inference resources at https://github.com/Shajiu/Tibetan-LLaMA-2 .
Jiu Sha, Mengxiao Zhu 0004, Chong Feng 0001, Jizhuoma Ci
ACM Trans. Asian Low Resour. Lang. Inf. Process.1