VLDB 2026 Research / reviewers in the wild / expert
Yang Xue 0001
dblp:25/6299-1
· DBLP profile ↗
16ranked-venue papers
3as first author
13since 2021 · last 2026
0000-0002-1947-4957ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 12 · 2 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 1 first-author · 6 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Transformer-based dynamic cell bounding box refinement for end-to-end Table Structure Recognition
Yang Xue 0001, Haosheng Cai, Zhuoming Li |
Pattern Recognit. Lett. | 1 |
| 2025 | G2LFormer: Global-to-Local Query Enhancement for Robust Table Structure RecognitionabstractTable structure recognition (TSR), the task of extracting logical and physical structures from table images, is critical for document understanding. Current end-to-end image-to-text methods typically employ a top-down strategy where physical structure prediction depends on the logical decoder's output sequence. However, this process often suffers from training instability and misalignment between predicted bounding boxes and ground-truth cell positions. To address this issue, we propose G2LFormer, a novel transformer-based framework that employs a ''Global-to-Local'' query enhancement strategy. Specifically, G2LFormer introduces a Vision-guided Query Enhancer to integrate both textual and visual modalities, significantly improving the overall query representation capability and boosting prediction accuracy. Additionally, we design a Multi-scale Manhattan Vision-guider that leverages a spatial attenuation matrix to guide each query towards its corresponding cell location, effectively balancing local and global information for more precise bounding box generation. Extensive experiments on benchmark datasets demonstrate G2LFormer's superior performance, while ablation studies confirming the significant contribution of each proposed module in achieving state-of-the-art results. The source code and model have been released at: https://github.com/Hzbupahaozi/G2LFormer. Haosheng Cai, Yang Xue 0001 |
ACM Multimedia | 2 |
| 2025 | Bi-VLDoc: bidirectional vision-language modeling for visually-rich document understanding
Chuwei Luo, Guozhi Tang, Qi Zheng 0002, Cong Yao, Yang Xue 0001, Luo Si |
Int. J. Document Anal. Recognit. | 7 |
| 2024 | DocNLC: A Document Image Enhancement Framework with Normalized and Latent Contrastive Representation for Multiple DegradationsabstractDocument Image Enhancement (DIE) remains challenging due to the prevalence of multiple degradations in document images captured by cameras. In this paper, we respond an interesting question: can the performance of pre-trained models and downstream DIE models be improved if they are bootstrapped using different degradation types of the same semantic samples and their high-dimensional features with ambiguous inter-class distance? To this end, we propose an effective contrastive learning paradigm for DIE — a Document image enhancement framework with Normalization and Latent Contrast (DocNLC). While existing DIE methods focus on eliminating one type of degradation, DocNLC considers the relationship between different types of degradation while utilizing both direct and latent contrasts to constrain content consistency, thus achieving a unified treatment of multiple types of degradation. Specifically, we devise a latent contrastive learning module to enforce explicit decorrelation of the normalized representations of different degradation types and to minimize the redundancy between them. Comprehensive experiments show that our method outperforms state-of-the-art DIE models in both pre-training and fine-tuning stages on four publicly available independent datasets. In addition, we discuss the potential benefits of DocNLC for downstream tasks. Our code is released at https://github.com/RylonW/DocNLC Ruilu Wang, Yang Xue 0001 |
AAAI | 2 |
| 2024 | GARDEN: Generative Prior Guided Network for Scene Text Image Super-Resolution
Yuxin Kong, Weihong Ma, Yang Xue 0001 |
ICDAR (5) | 4 |
| 2023 | Scene Table Structure Recognition with Segmentation and Key Point Collaboration
Zhuoming Li, Fan Peng, Yang Xue 0001, Ni Hao |
ICDAR (2) | 3 |
| 2023 | GridFormer: Towards Accurate Table Structure Recognition via Grid PredictionabstractAll tables can be represented as grids. Based on this observation, we propose GridFormer, a novel approach for interpreting unconstrained table structures by predicting the vertex and edge of a grid. First, we propose a flexible table representation in the form of an M X N grid. In this representation, the vertexes and edges of the grid store the localization and adjacency information of the table. Then, we introduce a DETR-style table structure recognizer to efficiently predict this multi-objective information of the grid in a single shot. Specifically, given a set of learned row and column queries, the recognizer directly outputs the vertexes and edges information of the corresponding rows and columns. Extensive experiments on five challenging benchmarks which include wired, wireless, multi-merge-cell, oriented, and distorted tables demonstrate the competitive performance of our model over other methods. Pengyuan Lv, Weihong Ma, Hongyi Wang 0008, Yuechen Yu, Chengquan Zhang, Yang Xue 0001, Jingdong Wang 0001 |
ACM Multimedia | 7 |
| 2023 | Scene table structure recognition with segmentation collaboration and alignment
Hongyi Wang 0008, Yang Xue 0001, Jiaxin Zhang 0003 |
Pattern Recognit. Lett. | 2 |
| 2022 | CMT-Co: Contrastive Learning with Character Movement Task for Handwritten Text Recognition
Jiapeng Wang 0003, Yujin Ren, Yang Xue 0001 |
ACCV (7) | 5 |
| 2022 | ChaCo: Character Contrastive Learning for Handwritten Text Recognition
Jiapeng Wang 0003, Canjie Luo, Yang Xue 0001 |
ICFHR | 6 |
| 2022 | TextSRNet: Scene Text Super-Resolution Based on Contour Prior and Atrous ConvolutionabstractLow-resolution (LR) scene text images are often encountered in real application scenarios. It is difficult to recognize these scene texts due to the loss of their character details. Super-resolution (SR) is an intuitive alternative to solve this problem. Unlike generic SR, which mainly focuses on image quality, text SR focuses on the accuracy of the downstream recognition tasks. Although several scene text SR methods have been recently proposed, they ignore retaining character contours that are meaningful for character readability and poorly perform on long texts. Therefore, we proposes a new scene text SR network called TextSRNet. In order to get fine character details, we adopt the segmentation maps of scene text images as the prior knowledge of character contours and embed it into the proposed TextSRNet. In addition, we incorporate Sobel loss to enhance the shape boundaries of characters. To improve the robustness of the model for long texts, we adopt a text atrous spatial pyramid pooling module. The proposed TextSRNet is trained and evaluated on TextZoom, a real-world scene text SR dataset. Extensive experiments demonstrate that our TextSRNet outperforms existing methods in terms of both recognition accuracy and image quality. Jizhao Ma, Jiaxin Zhang 0003, Yang Xue 0001, Mengchao He |
ICPR | 5 |
| 2021 | DSCNN: Dimension Separable Convolutional Neural Networks for Character Recognition Based on Inertial Sensor Signal
Fan Peng, Zhendong Zhuang, Yang Xue 0001 |
ICDAR (3) | 3 |
| 2021 | A Novel Unsupervised domain adaptation method for inertia-Trajectory translation of in-air handwriting
Songbin Xu, Yang Xue 0001, Xin Zhang 0013 |
Pattern Recognit. | 2 |
| 2016 | Ensemble Manifold Rank Preserving for Acceleration-Based Human Activity RecognitionabstractWith the rapid development of mobile devices and pervasive computing technologies, acceleration-based human activity recognition, a difficult yet essential problem in mobile apps, has received intensive attention recently. Different acceleration signals for representing different activities or even a same activity have different attributes, which causes troubles in normalizing the signals. We thus cannot directly compare these signals with each other, because they are embedded in a nonmetric space. Therefore, we present a nonmetric scheme that retains discriminative and robust frequency domain information by developing a novel ensemble manifold rank preserving (EMRP) algorithm. EMRP simultaneously considers three aspects: 1) it encodes the local geometry using the ranking order information of intraclass samples distributed on local patches; 2) it keeps the discriminative information by maximizing the margin between samples of different classes; and 3) it finds the optimal linear combination of the alignment matrices to approximate the intrinsic manifold lied in the data. Experiments are conducted on the South China University of Technology naturalistic 3-D acceleration-based activity dataset and the naturalistic mobile-devices based human activity dataset to demonstrate the robustness and effectiveness of the new nonmetric scheme for acceleration-based human activity recognition. Dapeng Tao, Yuan Yuan 0001, Yang Xue 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2010 | A New Rotation Feature for Single Tri-axial Accelerometer Based 3D Spatial Handwritten Digit RecognitionabstractA new rotation feature extracted from tri-axial acceleration signals for 3D spatial handwritten digit recognition is proposed. The feature can effectively express the clockwise and anti-clockwise direction changes of the users' movement while writing in a 3D space. Based on the rotation feature, an algorithm for 3D spatial handwritten digit recognition is presented. First, the rotation feature of the handwritten digit is extracted and coded. Then, the normalized edit distance between the digit and class model is computed. Finally, classification is performed using Support Vector Machine (SVM). The proposed approach outperforms time-domain features with a 22.12% accuracy improvement, peak-valley features with a 12.03% accuracy improvement, and FFT features with a 3.24% accuracy improvement, respectively. Experimental results show that the proposed approach is effective. Yang Xue 0001 |
ICPR | 1 |
| 2010 | A naturalistic 3D acceleration-based activity dataset & benchmark evaluationsabstractIn this paper, a naturalistic 3D acceleration-based activity dataset, the SCUT-NAA dataset, is created to assist researchers in the field of acceleration-based activity recognition and to provide a standard dataset for comparing and evaluating the performance of different algorithms. The SCUT-NAA dataset is the first publicly available 3D acceleration-based activity dataset and contains 1278 samples from 44 subjects (34 males and 10 females) collected in naturalistic settings with only one tri-axial accelerometer located alternatively on the waist belt, in the trousers pocket, and in the shirt pocket. Each subject was asked to perform ten activities. Benchmark evaluations of the dataset are provided based on FFT coefficients, DCT coefficients, time-domain features, and AR coefficients for the different accelerometer locations. Yang Xue 0001 |
SMC | 1 |