VLDB 2026 Research / reviewers in the wild / expert
Mayire Ibrayim
dblp:135/9689 · also Mayire Ibrahim
· DBLP profile ↗
19ranked-venue papers
4as first author
18since 2021 · last 2026
0000-0002-8766-0647ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 12 · 2 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 3 first-author · 6 since 2021Databases, data management, data science and information retrieval · 4 · 1 first-author · 4 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021Security and privacy · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Visual Information Facilitation Scene Text Retrieval
Mayire Ibrayim, Hailong Luo, Pengyang Li |
ICPR (9) | 1 |
| 2026 | PrQAC : Prompting LLaMA3 with question-aware image captions and answer candidates for knowledge-based VQA
Peichao Jiang, Mayire Ibrayim, Linying Wang |
Inf. Process. Manag. | 2 |
| 2026 | Adaptive dynamic graph interaction network for fine-grained image-text retrieval
Pengyang Li, Mayire Ibrayim, Alkut Mardan, Peichao Jiang |
Knowl. Based Syst. | 2 |
| 2026 | DocSR-DETR: Enhancing Document Layout Analysis with emphasis on background regions of document images
Mayire Ibrayim, Linying Wang |
Pattern Recognit. | 2 |
| 2025 | An Optional 2D Feature Scene Text Recognition Network Based on Transformer
Mayire Ibrayim, Jianjun Kang, Zhicheng Bao |
ICIG (3) | 1 |
| 2025 | Multi-modal End-to-End Text Spotting Networks: Interactive Enhancements Between Visual and Semantic Features
Mayire Ibrayim, Yefei Qian, Zhicheng Bao |
ICIG (3) | 1 |
| 2025 | Arbitrary-Shaped Text Detection with Hierarchical Feature Refinement and Nonlinear-Enhanced TransformersabstractDetecting arbitrary-shaped scene text is a challenging task due to the irregularities in font, size, color, orientation, and shape, which often lead to detection errors. In this paper, we propose a novel boundary-learning-based, coarse-to-fine text detection network that combines the strengths of CNNs and Transformers. Specifically, we introduce two key components: the Hierarchical Feature Refinement Network (HFRNet) and the Nonlinear-Enhanced Transformer Module (NETM). HFRNet enhances multi-scale feature extraction, improving the detection of text at various scales through dynamic convolution kernel sampling and attention mechanisms. This enables better adaptation to spatial scale variations and geometric deformations. NETM, on the other hand, leverages multi-head self-attention and nonlinear feature mapping to improve the representation of complex sequential data, allowing for more accurate text boundary detection in a coarse-to-fine manner. Our method, integrating HFRNet and NETM, achieves state-of-the-art performance on benchmark scene text detection datasets, including Total-Text and CTW1500. Zhicheng Bao, Mayire Ibrayim |
IJCNN | 2 |
| 2025 | TriBiaNet: Hierarchical-Biaffine Multimodal Fusion for Key Information Extraction in Visually Rich Documents
Linying Wang, Mayire Ibrayim, Peichao Jiang |
PRCV (7) | 2 |
| 2025 | Object-Centric Transformer Framework for Fine-Grained Image-Text Retrieval with Global Consistency *abstractCross-modal image-text retrieval enables efficient heterogeneous modality interaction via vision-language semantic alignment, advancing multimodal intelligence applications. However, traditional cross-modal retrieval methods typically rely on pre-trained feature extractors, whose performance limitations hinder the accuracy of image-text retrieval. In this paper, we propose an end-to-end Transformer-based cross-modal image-text retrieval framework, OCGC, which extracts visual and textual features through Transformer models. The framework employs a Global-Attentive Slot Fusion module to aggregate object-centric visual features to address feature redundancy, combined with a Global Semantic Consistency Alignment module to enhance cross-modal feature alignment between aggregated image features and text embeddings. Experimental results demonstrate that the proposed model performs exceptionally well on two benchmark datasets, outperforming the state-of-the-art methods by over 6% Rsum on the Flickr30K dataset. Sitong Shen, Mayire Ibrayim, Peichao Jiang |
SMC | 2 |
| 2025 | TCaEx:Targeted Caption as External Knowledge for knowledge-based visual question answering
Peichao Jiang, Mayire Ibrayim, Sitong Shen |
Image Vis. Comput. | 2 |
| 2025 | Enhanced spatial and interaction channel feature network for skin lesion segmentation
Rumeng Wang, Mayire Ibrayim, Askar Hamdulla |
Multim. Tools Appl. | 2 |
| 2025 | Contextual xLSTM-based multimodal fusion for conversational emotion recognition
Yupeng Qi, Mayire Ibrayim, Turdi Tohti |
Pattern Anal. Appl. | 2 |
| 2024 | Attention-based Dual-Branch Network for Micro-Expression Recognition with Global-Local Feature FusionabstractMicro-expression(ME) is an uncontrollable muscle movement that appears on the face when people try to hide or inhibit their real emotions, which has the problems of short duration, small movement amplitude and uneven distribution. In order to solve the problem of localization and asymmetry in the distribution of ME features, this paper proposes a new two-branch attention network to recognize MEs, which utilizes the attention mechanism to capture global and local ME features. The network is mainly divided into three parts: data preprocessing, ME feature learning and feature fusion classification. First, the data preprocessing first extracts the ME optical flow features and then divides them into four regions, which are used as inputs to the two-branch network respectively. Second, the two-branch network uses the Inception-MSFE global multi-scale feature extraction network incorporating the attention mechanism (CBAM) and the Swin Transformer-based local feature extraction network for feature learning, respectively. Finally, MEs are predicted by fusing global and local features of MEs. Experimental validation is carried out on three datasets, CASME II, SAMM, and SMIC, which proves that Acc, UAR, and UF1 are 0.797, 0.702, and 0.698 on the SAMM dataset, respectively; Acc, UAR, and UF1 are 0.734, 0.723, and 0.729 on SMIC, respectively; and on the CASME II dataset, Acc, UAR and UF1 are 0.865, 0.872, and 0.889, respectively, which are competitive with other state-of-the-art methods. Yupeng Qi, Mayire Ibrayim, Askar Hamdulla |
IJCB | 2 |
| 2024 | Doc-DINO: A Transformer Model for Complex Logical Document Layout Analysis
Mayire Ibrayim, Askar Hamdulla, Hailong Luo, Chunhu Zhang |
ICDAR (4) | 2 |
| 2024 | A Real-Time Scene Uyghur Text Detection Network Based on Feature Complementation
Mayire Ibrayim, Askar Hamdulla, Jianjun Kang, Chunhu Zhang |
ICDAR (5) | 1 |
| 2024 | More and Less: Enhancing Abundance and Refining Redundancy for Text-Prior-Guided Scene Text Image Super-Resolution
Yihong Luo, Mayire Ibrayim, Askar Hamdulla |
ICDAR (5) | 3 |
| 2024 | G-YOLOv5: A Face Mask Detector That Balances Effectiveness and Real-TimeabstractIn crowded scenarios, face mask detection algorithms still suffer from the target misdetection and omission, and the difficulty of reconciling real time and effectiveness. To address these issues, we present an G-YOLOv5 face mask detection algorithm. First, we adopt a weighted bidirectional feature pyramid network as the feature fusion network to enhance multiscale feature fusion; second, we utilize the WIoU loss function to strengthen the impact of good anchor frames while weakening the detrimental impact of poor quality anchor frames. Finally, considering the real-time and effectiveness issues of detector, we designed the Ghostv2-C3 module instead of the C3 module of the backbone to improve the inference speed of the model. The results of the experiment demonstrates that compared with other detectors, our detector plays a positive role in solving the target misdetection and omission and balancing the real-time nature of the model. Rong Ye, Mayire Ibrayim, Askar Hamdulla |
IJCNN | 2 |
| 2024 | Visual and semantic guided scene text retrieval
Hailong Luo, Mayire Ibrayim, Askar Hamdulla |
J. Supercomput. | 2 |
| 2020 | A benchmark for unconstrained online handwritten Uyghur word recognition
Wujiahemaiti Simayi, Mayire Ibrayim, Xu-Yao Zhang, Cheng-Lin Liu 0001, Askar Hamdulla |
Int. J. Document Anal. Recognit. | 2 |