VLDB 2026 Research / reviewers in the wild / expert
Liangcai Gao
dblp:23/7062
· DBLP profile ↗
77ranked-venue papers
8as first author
35since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 54 · 6 first-author · 27 since 2021Databases, data management, data science and information retrieval · 42 · 7 first-author · 13 since 2021Graphics, computer vision, multimedia, augmented reality and games · 18 · 1 first-author · 12 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | GlyphSR: A Simple Glyph-Aware Framework for Scene Text Image Super-ResolutionabstractThe goal of scene text image super-resolution (STISR) is to enhance the clarity of text within line images, thereby improving readability and enabling more accurate text recognition. However, existing STISR methods often rely heavily on Text Prior (TP) derived from trained recognizers, which can be unreliable and may lead to incorrect glyph restoration. Text images contain two crucial types of information: semantic content from word meanings and structural details from glyphs. When semantic information is unreliable, accurate perception of glyph structures becomes essential. This paper introduces GlyphSR, a novel STISR framework that addresses three key challenges: precise extraction, effective learning, and optimal utilization of glyph structural information. GlyphSR incorporates the Glyph Extraction Module (GEM), a training-free approach leveraging the Segment Anything Model (SAM) to accurately extract character-level glyphs. The Glyph Perception Module (GPM) models and learns glyph structures through segmentation and classification tasks, while the Glyph Fusion Module (GFM) integrates glyph information to enhance overall STISR model performance. Extensive experiments on the TextZoom dataset demonstrate that GlyphSR achieves a new state-of-the-art performance. Baole Wei, Liangcai Gao, Zhi Tang 0001 |
AAAI | 3 |
| 2025 | TAMER: Tree-Aware Transformer for Handwritten Mathematical Expression RecognitionabstractHandwritten Mathematical Expression Recognition (HMER) has extensive applications in automated grading and office automation. However, existing sequence-based decoding methods, which directly predict LaTeX sequences, struggle to understand and model the inherent tree structure of LaTeX and often fail to ensure syntactic correctness in the decoded results. To address these challenges, we propose a novel model named TAMER (Tree-Aware Transformer) for handwritten mathematical expression recognition. TAMER introduces an innovative Tree-aware Module while maintaining the flexibility and efficient training of Transformer. TAMER combines the advantages of both sequence decoding and tree decoding models by jointly optimizing sequence prediction and tree structure prediction tasks, which enhances the model's understanding and generalization of complex mathematical expression structures. During inference, TAMER employs a Tree Structure Prediction Scoring Mechanism to improve the structural validity of the generated LaTeX sequences. Experimental results on CROHME datasets demonstrate that TAMER outperforms traditional sequence decoding and tree decoding models, especially in handling complex mathematical structures, achieving state-of-the-art (SOTA) performance. Wenqi Zhao, Yu Li 0023, Xingjian Hu, Liangcai Gao |
AAAI | 5 |
| 2025 | LogicPro: Improving Complex Logical Reasoning via Program-Guided LearningabstractJin Jiang, Yuchen Yan, Yang Liu, Jianing Wang, Shuai Peng, Xunliang Cai, Yixin Cao, Mengdi Zhang, Liangcai Gao. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Shuai Peng, Liangcai Gao |
ACL (1) | 9 |
| 2025 | Do Large Language Models excel in Complex Logical Reasoning with Formal Language?abstractLarge Language Models (LLMs) have been shown to achieve breakthrough performance on complex logical reasoning tasks.Nevertheless, most existing research focuses on employing formal language to guide LLMs to derive reliable reasoning paths, while systematic evaluations of these capabilities are still limited.In this paper, we aim to conduct a comprehensive evaluation of LLMs across various logical reasoning problems utilizing formal languages.From the perspective of three dimensions, i.e., spectrum of LLMs, taxonomy of tasks, and format of trajectories, our key findings are: 1) Thinking models significantly outperform Instruct models, especially when formal language is employed; 2) All LLMs exhibit limitations in inductive reasoning capability, irrespective of whether they use a formal language; 3) Data with PoT format achieves the best generalization performance across other languages.Additionally, we also curate the formal-relative training data to further enhance the small language models, and the experimental results 1 indicate that a simple rejected fine-tuning method can better enable LLMs to generalize across formal languages and achieve the best overall performance. Liangcai Gao |
EMNLP | 7 |
| 2025 | DocVideoQA: Towards Comprehensive Understanding of Document-Centric Videos through Question AnsweringabstractRemote work and online courses have become important methods of knowledge dissemination, leading to a large number of document-based instructional videos. Unlike traditional video datasets, these videos mainly feature rich-text images and audio that are densely packed with information closely tied to the visual content, requiring advanced multimodal understanding capabilities. However, this domain remains underexplored due to dataset availability and its inherent complexity. In this paper, we introduce the DocVideoQA task and dataset for the first time, comprising 1,454 videos across 23 categories with a total duration of about 828 hours. The dataset is annotated with 154K question-answer pairs generated manually and via GPT, assessing models’ comprehension, temporal awareness, and modality integration capabilities. Initially, we establish a baseline using open-source MLLMs. Recognizing the challenges in modality comprehension for document-centric videos, we present DV-LLaMA, a robust video MLLM baseline. Our method enhances unimodal feature extraction with diverse instruction-tuning data and employs contrastive learning to strengthen modality integration. Through fine-tuning, the LLM is equipped with audio-visual capabilities, leading to significant improvements in document-centric video understanding. Extensive testing on the DocVideoQA dataset shows that DV-LLaMA significantly outperforms existing models. We’ll release the code and dataset to facilitate future research. Liangcai Gao |
ICASSP | 3 |
| 2025 | ICDAR 2025 Competition on Understanding Chinese College Entrance Exam Papers
Wenhao Yu 0013, Tianrui Zong, He Zhang 0034, Guanhao Wu, Ruohua Xu, Qinqin Yan, Liangcai Gao |
ICDAR (5) | 9 |
| 2025 | SketchRef: a Multi-Task Evaluation Benchmark for Sketch SynthesisabstractSketching is a powerful artistic technique for capturing essential visual information about real-world objects and has increasingly attracted attention in image synthesis research. However, the field lacks a unified benchmark to evaluate the performance of various synthesis methods. To address this, we propose SketchRef, the first comprehensive multi-task evaluation benchmark for sketch synthesis. SketchRef fully leverages the shared characteristics between sketches and reference photos. It introduces two primary tasks: category prediction and structural consistency estimation, the latter being largely overlooked in previous studies. These tasks are further divided into five subtasks across four domains: animals, common things, human body, and faces. Recognizing the inherent trade-off between recognizability and simplicity in sketches, we are the first to quantify this balance by introducing a recognizability calculation method constrained by simplicity, mRS, ensuring fair and meaningful evaluations. To validate our approach, we collected 7,920 responses from art enthusiasts, confirming the effectiveness of our proposed evaluation metrics. Additionally, we evaluate the performance of existing sketch synthesis methods on our benchmark, highlighting their strengths and weaknesses. We hope this study establishes a standardized benchmark and offers valuable insights for advancing sketch synthesis algorithms. Xingyue Lin, Xingjian Hu, Shuai Peng, Liangcai Gao |
ICME | 5 |
| 2025 | Vote & Mix: Plug-and-Play Token Reduction for Efficient Vision TransformerabstractDespite the remarkable success of Vision Transformers (ViTs) in various visual tasks, they are often hindered by substantial computational cost. In this work, we introduce Vote&Mix (VoMix), a plug-and-play and parameter-free token reduction method, which can be readily applied to off-the-shelf ViT models without any training. VoMix tackles the computational redundancy of ViTs by identifying tokens with high homogeneity through a layer-wise token similarity voting mechanism. Subsequently, the selected tokens are mixed into the retained set, thereby preserving visual information. Experiments demonstrate VoMix significantly improves the speed-accuracy tradeoff of ViTs on both images and videos. Without any training, VoMix achieves a 2× increase in throughput of existing ViT-H on ImageNet-1K and a 2.4× increase in throughput of existing ViT-L on Kinetics-400 video dataset, with a mere 0.3% drop in top-1 accuracy. Shuai Peng, Di Fu, Baole Wei, Liangcai Gao, Zhi Tang 0001 |
ICME | 4 |
| 2025 | Uni-MuMER: Unified Multi-Task Fine-Tuning of Vision-Language Model for Handwritten Mathematical Expression RecognitionabstractHandwritten Mathematical Expression Recognition (HMER) remains a persistent challenge in Optical Character Recognition (OCR) due to the inherent freedom of symbol layouts and variability in handwriting styles.
Prior methods have faced performance bottlenecks by proposing isolated architectural modifications, making them difficult to integrate coherently into a unified framework.
Meanwhile, recent advances in pretrained vision-language models (VLMs) have demonstrated strong cross-task generalization, offering a promising foundation for developing unified solutions.
In this paper, we introduce Uni-MuMER, which fully fine-tunes a VLM for the HMER task without modifying its architecture, effectively injecting domain-specific knowledge into a generalist framework.
Our method integrates three data-driven tasks: Tree-Aware Chain-of-Thought (Tree-CoT) for structured spatial reasoning, Error-Driven Learning (EDL) for reducing confusion among visually similar
characters, and Symbol Counting (SC) for improving recognition consistency in long expressions.
Experiments on the CROHME and HME100K datasets show that Uni-MuMER achieves super state-of-the-art performance,
outperforming the best lightweight specialized model SSAN by 16.31\% and the top-performing VLM Gemini2.5-flash by 24.42\% under zero-shot setting.
Our datasets, models, and code are open-sourced at: https://github.com/BFlameSwift/Uni-MuMER Shuai Peng, Baole Wei, Liangcai Gao |
NeurIPS | 7 |
| 2025 | Enhancing Graph Representation Learning with Localized Topological FeaturesabstractRepresentation learning on graphs is a fundamental problem that can be crucial in various tasks. Graph neural networks, the dominant approach for graph representation learning, are limited in their representation power. Therefore, it can be beneficial to explicitly extract and incorporate high-order topological and geometric information into these models. In this paper, we propose a principled approach to extract the rich connectivity information of graphs based on the theory of persistent homology. Our method utilizes the topological features to enhance the representation learning of graph neural networks and achieve state-of-the-art performance on various node classification and link prediction benchmarks. We also explore the option of end-to-end learning of the topological features, i.e., treating topological computation as a differentiable operator during learning. Our theoretical analysis and empirical study provide insights and potential guidelines for employing topological features in graph learning tasks. Zuoyu Yan, Qi Zhao 0007, Ze Ye, Tengfei Ma 0001, Liangcai Gao, Zhi Tang 0001, Yusu Wang 0001, Chao Chen 0012 |
J. Mach. Learn. Res. | 5 |
| 2024 | Maskstr: Guide Scene Text Recognition Models with MaskingabstractText recognition in information loss scenarios like blurriness, occlusion, and perspective distortion is challenging in real-world applications. To enhance robustness, some studies use extra unlabeled data for encoder pretraining. Others focus on improving decoder context reasoning. However, pretraining methods require abundant unlabeled data and high computing resources, while decoder-based approaches risk over-correction. In this paper, we propose MaskSTR, a dual-branch training framework for STR models, using patch masking to simulate information loss. MaskSTR guides visual representation learning, improving robustness to information loss conditions without extra data or training stages. Furthermore, we introduce Block Masking, a novel and straightforward mask generation method, for further performance enhancement. Experiments demonstrate MaskSTR’s effectiveness across CTC, attention, and Transformer decoding methods, achieving significant performance gains and setting new state-of-the-art results. Baole Wei, Minghang He, Liangcai Gao, Duoyou Zhou, Xiang Bai, Zhi Tang 0001 |
ICASSP | 3 |
| 2024 | Recognition-Guided Diffusion Model for Scene Text Image Super-ResolutionabstractScene Text Image Super-Resolution (STISR) aims to enhance the resolution and legibility of text within low-resolution (LR) images, consequently elevating recognition accuracy in Scene Text Recognition (STR). Previous methods predominantly employ discriminative Convolutional Neural Networks (CNNs) augmented with diverse forms of text guidance to address this issue. Nevertheless, they remain deficient when confronted with severely blurred images, due to their insufficient generation capability when little structural or semantic information can be extracted from original images. Therefore, we introduce RGDiffSR, a Recognition-Guided Diffusion model for scene text image Super-Resolution, which exhibits great generative diversity and fidelity even in challenging scenarios. Moreover, we propose a Recognition-Guided Denoising Network, to guide the diffusion model generating LR-consistent results through succinct semantic guidance. Experiments on the TextZoom dataset demonstrate the superiority of RGDiffSR over prior state-of-the-art methods in both text recognition accuracy and image fidelity. Liangcai Gao, Zhi Tang 0001, Baole Wei |
ICASSP | 2 |
| 2024 | SegHist: A General Segmentation-Based Framework for Chinese Historical Document Text Line Detection
Xingjian Hu, Baole Wei, Liangcai Gao |
ICDAR (3) | 3 |
| 2024 | DocTabQA: Answering Questions from Long Documents Using Tables
Liangcai Gao |
ICDAR (1) | 4 |
| 2024 | ICAL: Implicit Character-Aided Learning for Enhanced Handwritten Mathematical Expression Recognition
Liangcai Gao, Wenqi Zhao |
ICDAR (5) | 2 |
| 2024 | ODTr: Transformer Integrating OCR Auxiliary Map and Image Depth Information for Document Image Unwarping
Xiangyu Xie, Liangcai Gao |
ICPR (31) | 3 |
| 2024 | ICPR 2024 Competition on Multi-line Mathematical Expressions Recognition
Tianrui Zong, Wei Luo 0001, Wenhao Yu 0013, Borui Cai, He Zhang 0034, Fenghong Liu 0002, Qinqin Yan, Liangcai Gao |
ICPR (34) | 9 |
| 2024 | An Efficient Subgraph GNN with Provable Substructure Counting PowerabstractWe investigate the enhancement of graph neural networks' (GNNs) representation power through their ability in substructure counting. Recent advances have seen the adoption of subgraph GNNs, which partition an input graph into numerous subgraphs, subsequently applying GNNs to each to augment the graph's overall representation. Despite their ability to identify various substructures, subgraph GNNs are hindered by significant computational and memory costs. In this paper, we tackle a critical question: Is it possible for GNNs to count substructures both efficiently and provably? Our approach begins with a theoretical demonstration that the distance to rooted nodes in subgraphs is key to boosting the counting power of subgraph GNNs. To avoid the need for repetitively applying GNN across all subgraphs, we introduce precomputed structural embeddings that encapsulate this crucial distance information. Experiments validate that our proposed model retains the counting power of subgraph GNNs while achieving significantly faster performance. Zuoyu Yan, Junru Zhou, Liangcai Gao, Zhi Tang 0001, Muhan Zhang |
KDD | 3 |
| 2024 | Enhancing Transformer-Based Table Structure Recognition for Long Tables
Wenqi Zhao, Liangcai Gao |
PRCV (7) | 3 |
| 2024 | Table image dewarping with key element segmentation
Zhi Tang 0001, Liangcai Gao |
Int. J. Document Anal. Recognit. | 3 |
| 2023 | Improving Table Structure Recognition with Visual-Alignment Sequential Coordinate ModelingabstractTable structure recognition aims to extract the logical and physical structure of unstructured table images into a machine-readable format. The latest end-to-end image-to-text approaches simultaneously predict the two structures by two decoders, where the prediction of the physical structure (the bounding boxes of the cells) is based on the representation of the logical structure. However, the previous methods struggle with imprecise bounding boxes as the logical representation lacks local visual information. To address this issue, we propose an end-to-end sequential modeling framework for table structure recognition called VAST. It contains a novel coordinate sequence decoder triggered by the representation of the non-empty cell from the logical structure decoder. In the coordinate sequence decoder, we model the bounding box coordinates as a language sequence, where the left, top, right and bottom coordinates are decoded sequentially to leverage the inter-coordinate dependency. Furthermore, we propose an auxiliary visual-alignment loss to enforce the logical representation of the non-empty cells to contain more local visual details, which helps produce better cell bounding boxes. Extensive experiments demonstrate that our proposed method can achieve state-of-the-art results in both logical and physical structure recognition. The ablation study also validates that the proposed coordinate sequence decoder and the visual-alignment loss are the keys to the success of our method. Yongshuai Huang, Ning Lu 0003, Dapeng Chen, Zecheng Xie, Shenggao Zhu, Liangcai Gao, Wei Peng 0011 |
CVPR | 7 |
| 2023 | An Iterative Graph Learning Convolution Network for Key Information Extraction Based on the Document Inductive Bias
Jiyao Deng, Zhi Tang 0001, Liangcai Gao |
ICDAR (3) | 5 |
| 2023 | A Character-Level Document Key Information Extraction Method with Contrastive Learning
Jiyao Deng, Liangcai Gao |
ICDAR (3) | 3 |
| 2022 | MIR: A Benchmark for Molecular Image Retrival with a Cross-modal Pretraining FrameworkabstractMolecular image retrieval is one of the crucial steps in automatic mining and utilization of biochemistry-related literatures, which is also a relatively open and challenging task in cross fields of biochemistry and artificial intelligence. The challenges come from two aspects: 1) there is a lack of open datasets and evaluation criteria for molecular image retrieval. 2) Common retrieval methods always ignore that molecular image retrieval has cross-modal information of both images and SMILES texts. To address the first challenge, we firstly construct a new molecular image retrieval benchmark, named MIR, including 130770 molecular images, labeled structural similarity, and reasonable evaluation metrics. Faced with the second challenge, we propose an effective cross-modal pre-training framework for molecular image retrieval following CLIP. Experimental results reflect the effectiveness of our proposed benchmark MIR and cross-modal pre-training framework. Baole Wei, Ruiqi Jia, Shihan Fu, Xiaoqing Lyu, Liangcai Gao, Zhi Tang 0001 |
BIBM | 5 |
| 2022 | CoMER: Modeling Coverage for Transformer-Based Handwritten Mathematical Expression Recognition
Wenqi Zhao, Liangcai Gao |
ECCV (28) | 2 |
| 2022 | Cycle Representation Learning for Inductive Relation PredictionabstractIn recent years, algebraic topology and its modern development, the theory of persistent homology, has shown great potential in graph representation learning. In this paper, based on the mathematics of algebraic topology, we propose a novel solution for inductive relation prediction, an important learning task for knowledge graph completion. To predict the relation between two entities, one can use the existence of rules, namely a sequence of relations. Previous works view rules as paths and primarily focus on the searching of paths between entities. The space of rules is huge, and one has to sacrifice either efficiency or accuracy. In this paper, we consider rules as cycles and show that the space of cycles has a unique structure based on the mathematics of algebraic topology. By exploring the linear structure of the cycle space, we can improve the searching efficiency of rules. We propose to collect cycle bases that span the space of cycles. We build a novel GNN framework on the collected cycles to learn the representations of cycles, and to predict the existence/non-existence of a relation. Our method achieves state-of-the-art performance on benchmarks. Zuoyu Yan, Tengfei Ma 0001, Liangcai Gao, Zhi Tang 0001, Chao Chen 0012 |
ICML | 3 |
| 2022 | Compute Like Humans: Interpretable Step-by-step Symbolic Computation with Deep Neural NetworkabstractNeural network capability in symbolic computation has emerged in much recent work. However, symbolic computation is always treated as an end-to-end blackbox prediction task, where human-like symbolic deductive logic is missing. In this paper, we argue that any complex symbolic computation can be broken down to a sequence of finite Fundamental Computation Transformations (FCT), which are grounded as certain mathematical expression computation transformations. The entire computation sequence represents a full human understandable symbolic deduction process. Instead of studying on different end-to-end neural network applications, this paper focuses on approximating FCT which further build up symbolic deductive logic. To better mimic symbolic computations with math expression transformations, we propose a novel tree representation learning architecture GATE (Graph Aggregation Transformer Encoder) for math expressions. We generate a large-scale math expression transformation dataset for training purpose and collect a real-world dataset for validation. Experiments demonstrate the feasibility of producing step-by-step human-like symbolic deduction sequences with the proposed approach, which outperforms other neural network approaches and heuristic approaches. Shuai Peng, Di Fu, Yijun Liang, Gu Xu, Liangcai Gao, Zhi Tang 0001 |
KDD | 6 |
| 2022 | Neural Approximation of Graph Topological FeaturesabstractTopological features based on persistent homology capture high-order structural information so as to augment graph neural network methods. However, computing extended persistent homology summaries remains slow for large and dense graphs and can be a serious bottleneck for the learning pipeline. Inspired by recent success in neural algorithmic reasoning, we propose a novel graph neural network to estimate extended persistence diagrams (EPDs) on graphs efficiently. Our model is built on algorithmic insights, and benefits from better supervision and closer alignment with the EPD computation algorithm. We validate our method with convincing empirical results on approximating EPDs and downstream graph representation learning tasks. Our method is also efficient; on large and dense graphs, we accelerate the computation by nearly 100 times. Zuoyu Yan, Tengfei Ma 0001, Liangcai Gao, Zhi Tang 0001, Yusu Wang 0001, Chao Chen 0012 |
NeurIPS | 3 |
| 2021 | An Infantile Hemangioma Dataset IH-2021 and a Deep Learning based Recognition Method on itabstractInfantile hemangioma(IH) is one of the most common skin and soft tissue tumors in children. From the appearance, it is very easy to be confused with vascular malformation. As a result, the misdiagnosis between them often occurs. To help doctors screen and identify hemangioma patients, we employed image recognition technology based on deep learning to recognize photos of patients. At first, we process and sort out the existing medical image to establish and release an infantile hemangioma dataset named IH-2021. Then we do classification experiments and obtain the baseline results on it. However, the number of medical images is usually rather small, directly leading to the overfitting of deep learning methods. To further improve the performance of image classification, we constructed a new neural network for image classification of infantile hemangioma, in which we introduced data augmentation approaches based on generative adversarial network and active learning. It has achieved better results on IH-2021 compared with other state-of-the-art models for this task. Liangcai Gao, Li Li 0054, Zuoyu Yan |
BIBM | 2 |
| 2021 | Rethinking Table Structure Recognition Using Sequence Labeling Methods
Yilun Huang 0001, Lemeng Pan, Yongshuai Huang, Zhi Tang 0001, Liangcai Gao |
ICDAR (2) | 8 |
| 2021 | Image to LaTeX with Graph Neural Network for Mathematical Formula Recognition
Shuai Peng, Liangcai Gao, Zhi Tang 0001 |
ICDAR (2) | 2 |
| 2021 | Formula Citation Graph Based Mathematical Information Retrieval
Liangcai Gao, Zhuoren Jiang, Zhi Tang 0001 |
ICDAR (1) | 2 |
| 2021 | Handwritten Mathematical Expression Recognition with Bidirectionally Trained Transformer
Wenqi Zhao, Liangcai Gao, Zuoyu Yan, Shuai Peng, Ziyin Zhang |
ICDAR (2) | 2 |
| 2021 | NTable: A Dataset for Camera-Based Table Detection
Liangcai Gao, Yilun Huang 0001 |
ICDAR (2) | 2 |
| 2021 | Link Prediction with Persistent Homology: An Interactive ViewabstractLink prediction is an important learning task for graph-structured data. In this paper, we propose a novel topological approach to characterize interactions between two nodes. Our topological feature, based on the extended persistent homology, encodes rich structural information regarding the multi-hop paths connecting nodes. Based on this feature, we propose a graph neural network method that outperforms state-of-the-arts on different benchmarks. As another contribution, we propose a novel algorithm to more efficiently compute the extended persistence diagrams for graphs. This algorithm can be generally applied to accelerate many other topological methods for graph learning tasks. Zuoyu Yan, Tengfei Ma 0001, Liangcai Gao, Zhi Tang 0001, Chao Chen 0012 |
ICML | 3 |
| 2020 | Automatic Generation of Headlines for Online Math QuestionsabstractMathematical equations are an important part of dissemination and communication of scientific information. Students, however, often feel challenged in reading and understanding math content and equations. With the development of the Web, students are posting their math questions online. Nevertheless, constructing a concise math headline that gives a good description of the posted detailed math question is nontrivial. In this study, we explore a novel summarization task denoted as geNerating A concise Math hEadline from a detailed math question (NAME). Compared to conventional summarization tasks, this task has two extra and essential constraints: 1) Detailed math questions consist of text and math equations which require a unified framework to jointly model textual and mathematical information; 2) Unlike text, math equations contain semantic and structural features, and both of them should be captured together. To address these issues, we propose MathSum, a novel summarization model which utilizes a pointer mechanism combined with a multi-head attention mechanism for mathematical representation augmentation. The pointer mechanism can either copy textual tokens or math tokens from source questions in order to generate math headlines. The multi-head attention mechanism is designed to enrich the representation of math equations by modeling and integrating both its semantic and structural features. For evaluation, we collect and make available two sets of real-world detailed math questions along with human-written math headlines, namely EXEQ-300k and OFEQ-10k. Experimental results demonstrate that our model (MathSum) significantly outperforms state-of-the-art models for both the EXEQ-300k and OFEQ-10k datasets. Dafang He, Zhuoren Jiang, Liangcai Gao, Zhi Tang 0001, C. Lee Giles |
AAAI | 4 |
| 2020 | ConvMath: A Convolutional Sequence Network for Mathematical Expression RecognitionabstractDespite the recent advances in optical character recognition (OCR), mathematical expressions still face a great challenge to recognize due to their two-dimensional graphical layout. In this paper, we propose a convolutional sequence modeling network, ConvMath, which converts the mathematical expression description in an image into a LaTeX sequence in an end-to-end way. The network combines an image encoder for feature extraction and a convolutional decoder for sequence generation. Compared with other Long Short Term Memory(LSTM) based encoder-decoder models, ConvMath is entirely based on convolution, thus it is easy to perform parallel computation. Besides, the network adopts multi-layer attention mechanism in the decoder, which allows the model to align output symbols with source feature vectors automatically, and alleviates the problem of lacking coverage while training the model. The performance of ConvMath is evaluated on an open dataset named IM2LATEX-100K, including 103556 samples. The experimental results demonstrate that the proposed network achieves state-of-the-art accuracy and much better efficiency than previous methods. Zuoyu Yan, Xiaode Zhang, Liangcai Gao, Zhi Tang 0001 |
ICPR | 3 |
| 2019 | A non-local based segmentation method for Pelvic MR ImagesabstractA number of people are infected with functional disorders of the pelvic floor such as pelvic organ prolapse and defecation dysfunction. For radiologists, to discern anatomical structures from tomographic images is an important way to find such diseases. However, it is a challenge to recognize the anatomical structures from pelvic MR (magnetic resonance) images, since the same pelvic organs usually are with various shapes and different organs sometimes are with similar appearance in MR images. To address this problem, we introduce a non-local module in the deep neural networks so that both the local features and the global features could be fused to segment organs. Moreover, for the very small quantities of the open pelvic MR images, we have built up a fine-grained MR image segmentation dataset annotated by professional medical researchers, and proposed a data augmentation strategy that is especially suitable for our dataset. In the primary experiments on the dataset, the proposed method in this paper is compared with state-of-the-art image segmentation methods and achieves better performance. Zuoyu Yan, Liangcai Gao, Zhi Tang 0001 |
BIBM | 2 |
| 2019 | ICDAR 2019 Competition on Table Detection and Recognition (cTDaR)abstractThe cTDaR competition aims at benchmarking state-of-the-art table detection (TRACK A) and table recognition (TRACK B) methods. In particular, we wish to investigate and compare general methods that can reliably and robustly identify the table regions within a document image on the one hand, and the table structure on the other hand. Due to the presence of hand-drawn tables and handwritten text, the methods must be robust against various noise conditions, interfering annotations, and variations of the tables. Two new challenging datasets were created to test the behaviour of state-of-the-art table detection and recognition systems on real world data. One dataset consists of modern documents, while the other consists of archival documents with presence of hand-drawn tables and handwritten text. The evaluation scheme is adapted from the ICDAR 2013 Table competition. We received results of Track A from 11 teams and results of Track B from 2 teams. Results for Track A are very good for the top participants. The winner and his runner-up are very close while using very different approaches. Track B was more challenging and only one participant was able to produce good results. Liangcai Gao, Yilun Huang 0001, Hervé Déjean, Jean-Luc Meunier, Qinqin Yan, Florian Kleber, Eva Maria Lang |
ICDAR | 1 |
| 2019 | A YOLO-Based Table Detection MethodabstractDue to various table layouts and styles, table detection is always a difficult task in the field of document analysis. Inspired by the great progress of deep learning based methods on object detection, in this paper, we present a YOLO-based method for this task. Considering the large difference between document objects and natural objects, we introduce some adaptive adjustments to YOLOv3, including an anchor optimization strategy and two post processing methods. For anchor optimization, we use k-means clustering to find anchors which are more suitable for tables rather than natural objects and make it easier for our model to find exact positions of tables. In post-processing process, the extra whitespaces and noisy page objects (e.g. page headers, page footers) are removed from the predicted results, so that our model can get more accurate table margins and higher IoU scores. The proposed method is evaluated on two datasets from ICDAR 2013 Table Competition and ICDAR 2017 Page Object Detection (POD) Competition and achieves state-of-the-art performance. Yilun Huang 0001, Qinqin Yan, Liangcai Gao, Zhi Tang 0001 |
ICDAR | 6 |
| 2019 | A GAN-Based Feature Generator for Table DetectionabstractTable detection is of great significance for the documents analysis and recognition. Although many methods have been proposed and great progress have been made, it is still a great challenge to recognize the less-ruled tables due to the lack of table line features. In this paper, we propose a novel network to generate the layout features for table text to improve the performance of less-ruled table recognition. This feature generator model is similar to the Generative Adversarial Networks (GAN). We force the feature generator model to extract similar features for both ruling tables and less-ruled tables. It can be added into some common object detection and semantic segmentation models such as Mask R-CNN, U-Net. Extensive experiments are conducted on the dataset of ICDAR2017 Page Object Detection Competition dataset and a closed dataset full of the less-ruled tables and non-ruled tables. The primary experimental results show that the proposed GAN-based feature generator is very helpful for less-ruled table detection. Liangcai Gao, Zhi Tang 0001, Qinqin Yan, Yilun Huang 0001 |
ICDAR | 2 |
| 2019 | Weakly supervised precise segmentation for historical document imagesabstractWith the passing of history, precious cultural heritage was left behind to tell ancient stories, especially those in the form of written documents. In this paper, a weakly supervised segmentation system with recognition-guided information on attention area, is proposed for high-precision historical document segmentation under strict intersection-over-union (IoU) requirements. We formulate the character segmentation problem from Bayesian decision theory perspective and propose boundary box segmentation (BBS), recognition-guided BBS (Rg-BBS), and recognition-guided attention BBS (Rg-ABBS), progressively, to search for the segmentation path. Furthermore, a novel judgment gate mechanism is proposed to train a high-performance character recognizer in an incremental weakly supervised learning manner. The proposed Rg-ABBS method is shown to substantially reduce time consumption while maintaining sufficiently high precision of the segmentation result by incorporating both character recognition knowledge and line-level annotation. Experiments show that the proposed Rg-ABBS system significantly outperforms traditional segmentation methods as well as deep-learning-based instance segmentation and detection methods under strict IoU requirements. Zecheng Xie, Yaoxiong Huang, Liangcai Gao, Xiaode Zhang |
Neurocomputing | 6 |
| 2018 | A Benchmark for Automatic Acral Melanoma Preliminary Screening
Liangcai Gao, Zhi Tang 0001, Menglong Ran, Zhimiao Lin |
BIBM | 2 |
| 2018 | Mathematics Content Understanding for Cyberlearning via Formula Evolution MapabstractAlthough the scientific digital library is growing at a rapid pace, scholars/students often find reading Science, Technology, Engineering, and Mathematics (STEM) literature daunting, especially for the math-content/formula. In this paper, we propose a novel problem, "mathematics content understanding", for cyberlearning and cyberreading. To address this problem, we create a Formula Evolution Map (FEM) offline and implement a novel online learning/reading environment, PDF Reader with Math-Assistant (PRMA), which incorporates innovative math-scaffolding methods. The proposed algorithm/system can auto-characterize student emerging math-information need while reading a paper and enable students to readily explore the formula evolution trajectory in FEM. Based on a math-information need, PRMA utilizes innovative joint embedding, formula evolution mining, and heterogeneous graph mining algorithms to recommend high quality Open Educational Resources (OERs), e.g., video, Wikipedia page, or slides, to help students better understand the math-content in the paper. Evaluation and exit surveys show that the PRMA system and the proposed formula understanding algorithm can effectively assist master and PhD students better understand the complex math-content in the class readings. Zhuoren Jiang, Liangcai Gao, Zheng Gao 0001, Zhi Tang 0001, Xiaozhong Liu 0001 |
CIKM | 2 |
| 2018 | A Sequence Labeling Based Approach for Character Segmentation of Historical DocumentsabstractAs an important prerequisite step of historical document image analysis, character segmentation is fundamental but challenging. In this paper, we propose a novel approach for the handwritten character segmentation of historical documents by treating it as a sequence labeling problem. In more detail, the proposed model first segments document image into lines, then each column in the line image is given a label to indicate it is a segmentation position or not. The segmentation labeling is achieved by a neural model, which combines a CNN for feature extraction, a LSTM for sequence modeling and a CRF for sequence labeling. The performance of our methods has been evaluated on a 300-page dataset including 96,479 characters. The experimental results demonstrate that the proposed methods achieve superior or highly competitive performance compared with other methods. Liangcai Gao, Xiaode Zhang, Zhi Tang 0001, Yaoxiong Huang |
DAS | 1 |
| 2018 | APNet: Semantic Segmentation for Pelvic MR Image
Ting-Ting Liang, Mengyan Sun, Liangcai Gao, Jingjing Lu, Satoshi Tsutsui |
PRCV (2) | 3 |
| 2018 | Cross-language Citation Recommendation via Hierarchical Representation Learning on Heterogeneous GraphabstractWhile the volume of scholarly publications has increased at a frenetic pace, accessing and consuming the useful candidate papers, in very large digital libraries, is becoming an essential and challenging task for scholars. Unfortunately, because of language barrier, some scientists (especially the junior ones or graduate students who do not master other languages) cannot efficiently locate the publications hosted in a foreign language repository. In this study, we propose a novel solution, cross-language citation recommendation via Hierarchical Representation Learning on Heterogeneous Graph (HRLHG), to address this new problem. HRLHG can learn a representation function by mapping the publications, from multilingual repositories, to a low-dimensional joint embedding space from various kinds of vertexes and relations on a heterogeneous graph. By leveraging both global (task specific) plus local (task independent) information as well as a novel supervised hierarchical random walk algorithm, the proposed method can optimize the publication representations by maximizing the likelihood of locating the important cross-language neighborhoods on the graph. Experiment results show that the proposed method can not only outperform state-of-the-art baseline models, but also improve the interpretability of the representation model for cross-language citation recommendation task. Zhuoren Jiang, Liangcai Gao, Yao Lu 0007, Xiaozhong Liu 0001 |
SIGIR | 3 |
| 2017 | Citation Metadata Extraction via Deep Neural Network-based Segment Sequence LabelingabstractCitation metadata extraction plays an important role in academic information retrieval and knowledge management. Current works on this task generally use rule-based, template-based or learning-based approaches but these methods usually either rely on handcrafted features or are limited with domains. Recently, neural networks have shown strong ability in addressing sequence labeling tasks. Liangcai Gao, Zhuoren Jiang, Runtao Liu, Zhi Tang 0001 |
CIKM | 2 |
| 2017 | ICDAR2017 Competition on Page Object DetectionabstractThis paper presents the results of ICDAR2017 Competition on Page Object Detection (POD). POD is to detect page objects (tables, mathematical equations, graphics, figures, etc.) from document images. This competition makes use of a dataset consists of 2,000 document page images This dataset contains abundant page objects with various types and layouts. During the competition, we received 13 different teams' registrations and finally 8 of them submitted their results. All teams used deep learning as the basic method, then combined different traditional features or methods to improve the detection performance. The team NLPR-PAL achieved the averaged F1 of 0.898 and mAPs of 0.805 in the detection of all page objects under the IOU threshold 0.8. In this overview paper, we summarize the task design, dataset, results, and the approaches used by those teams of this competitions. Liangcai Gao, Xiaohan Yi, Zhuoren Jiang, Leipeng Hao, Zhi Tang 0001 |
ICDAR | 1 |
| 2017 | A Deep Learning-Based Formula Detection Method for PDF DocumentsabstractIn practice, PDF files may be generated by different tools and their character information quality could be different. As a result, the approaches to detecting formulae from PDF documents usually have much different performance on different PDF files. To address this problem, in this paper we combine and refine the Convolutional Neural Network (CNN) and Recurrent Neural Network (RNN) model to detect formulae according to both their character and vision features. Based on the characteristic of PDF documents, we propose a series of strategies to train and optimize deep networks, such as the implicit class down-sampling strategy which can reduce the unbalancedness between formulae and other page elements (e.g., text paragraphs, tables, figures, etc.). The region proposal method is also redesigned to generate moderate formula candidates through combining the bottom-up and top-down layout analysis. The experimental results show that the combination of CNN and RNN can increase the robustness of our proposed detection method. Furthermore, the proposed method outperforms the existing formula detection methods on both a ground-truth dataset and a larger self-built dataset, which would be released and available for research purposes. Liangcai Gao, Xiaohan Yi, Zhuoren Jiang, Zuoyu Yan, Zhi Tang 0001 |
ICDAR | 1 |
| 2017 | CNN Based Page Object Detection in Document ImagesabstractThis electronic document is a "live" template. The various components of your paper [title, text, heads, etc.] are Abstract-Object detection in natural scenes has been widely researched in the past decade, and many deep learning based methods have achieved good performance on this task. This paper focuses on how to transfer and refine those object detection approaches from natural scene images to documents images, and proposes a deep learning-based page object (e.g., tables, formulae, figures) detection method. On the basis of traditional Convolutional Neural Network (CNN) based object detection methods, we redesign the region proposal method, the training strategy, the network structure and replace the Non-Maximum Suppression (NMS) with a dynamic programming algorithm. The experimental results show that it is essential to adjust some modules of the natural scene object detection approaches in order to better process the document images. The proposed method also achieved better performance compared with existing page object detection methods. Xiaohan Yi, Liangcai Gao, Xiaode Zhang, Runtao Liu, Zhuoren Jiang |
ICDAR | 2 |
| 2017 | A Symbol Dominance Based Formulae Recognition Approach for PDF DocumentsabstractWith more and more scientific documents becoming available in PDF format, recognition of formulae in these PDF documents is of great significance. In this paper, we propose a symbol dominance based formulae recognition approach to recovering formulae structures by using the rich information extracted directly from PDF files. The hierarchical structure of formula is represented by relationship tree, and the tree is built recursively based on symbol dominance, which considers both the spatial layout of symbols and the typesetting conventions of mathematics. In addition, we propose a special character recognition method to identify the formula characters with multiple components or variable unicode. Repeatable and comparable experiments have been done over two large datasets, IM2LATEX-100K and PDFME-10K. Experimental results demonstrate that our method is more adaptive and practical for PDF documents compared with other two existing available formulae recognition systems, INFTY and WYGIWYS. Xiaode Zhang, Liangcai Gao, Runtao Liu, Zhuoren Jiang, Zhi Tang 0001 |
ICDAR | 2 |
| 2017 | Automatic Document Metadata Extraction Based on Deep Networks
Runtao Liu, Liangcai Gao, Zhuoren Jiang, Zhi Tang 0001 |
NLPCC | 2 |
| 2016 | A Table Detection Method for PDF Documents Based on Convolutional Neural NetworksabstractBecause of the better performance of deep learning on many computer vision tasks, researchers in the area of document analysis and recognition begin to adopt this technique into their work. In this paper, we propose a novel method for table detection in PDF documents based on convolutional neutral networks, one of the most popular deep learning models. In the proposed method, some table-like areas are selected first by some loose rules, and then the convolutional networks are built and refined to determine whether the selected areas are tables or not. Besides, the visual features of table areas are directly extracted and utilized through the convolutional networks, while the non-visual information (e.g. characters, rendering instructions) contained in original PDF documents is also taken into consideration to help achieve better recognition results. The primary experimental results show that the approach is effective in table detection. Leipeng Hao, Liangcai Gao, Xiaohan Yi, Zhi Tang 0001 |
DAS | 2 |
| 2016 | Community-based Cyberreading for Information UnderstandingabstractAlthough the content in scientific publications is increasingly challenging, it is necessary to investigate another important problem, that of scientific information understanding. For this proposed problem, we investigate novel methods to assist scholars (readers) to better understand scientific publications by enabling physical and virtual collaboration. For physical collaboration, an algorithm will group readers together based on their profiles and reading behavior, and will enable the cyberreading collaboration within a online reading group. For virtual collaboration, instead of pushing readers to communicate with others, we cluster readers based on their estimated information needs. For each cluster, a learning to rank model will be generated to recommend readers' communitized resources (i.e., videos, slides, and wikis) to help them understand the target publication. Zhuoren Jiang, Xiaozhong Liu 0001, Liangcai Gao, Zhi Tang 0001 |
SIGIR | 3 |
| 2015 | Chronological Citation Recommendation with Information-Need ShiftingabstractAs the volume of publications has increased dramatically, an urgent need has developed to assist researchers in locating high-quality, candidate-cited papers from a research repository. Traditional scholarly-recommendation approaches ignore the chronological nature of citation recommendations. In this study, we propose a novel method called "Chronological Citation Recommendation" which assumes initial user information needs could shift while users are searching for papers in different time slices. We model the information-need shifts with two-level modeling: dynamic time-related ranking feature construction and dynamic evolving feature weight training. In more detail, we employed a supervised document influence model to characterize the content "time-varying" dynamics and constructed a novel heterogeneous graph that encapsulates dynamic topic-based information, time-decay paper/topic citation information, and word-based information. We applied multiple meta-paths for different ranking hypotheses which carried different types of information for citation recommendation in various time slices, along with information-need shifting. We also used multiple learning-to-rank models to optimize the feature weights for different time slices to generate the final "Chronological Citation Recommendation" rankings. The use of Chronological Citation Recommendation suggests time-series ranking lists based on initial user textual information need and characterizes the information-need shifting. Experiments on the ACM corpus show that Chronological Citation Recommendation can significantly enhance citation recommendation performance. Zhuoren Jiang, Xiaozhong Liu 0001, Liangcai Gao |
CIKM | 3 |
| 2015 | Classification of forms with similar layouts based on Mixed Gaussian Weighted MaskabstractAs an essential step of form processing, form classification has attracted much attention from researchers. However, for the forms with similar layout, most of the previous classification methods still suffer from two issues: huge variation among areas of user-filled-in data and insufficient discriminative identifiers in areas of preprinted data. In this paper, we propose a novel Mixed Gaussian Weighted Mask (MGWM) based method to identify forms with similar layouts by leveraging the multiple information extracted from areas of user-filled-in data, areas of preprinted data and dithering data of a form. The proposed method utilizes a combination of three Gaussian weighted masks to mitigate the impact of noise from areas of user-filled-in data, layout consistency and position dithering among form images respectively. Experimental results show that the proposed method achieves more than 85% classification accuracy on a number of forms and outperforms the state-of-the-art form classification method. Simeng Wang, Liangcai Gao, Yuehan Wang |
ICDAR | 2 |
| 2015 | A User-Oriented Special Topic Generation System for Digital NewspaperabstractWith the coming of digital newspaper, user-oriented special topic generation becomes extremely urgent to satisfy the users’ requirements both functionally and emotionally. We propose an applicable automatic special topic generation system for digital newspapers based on users’ interests. Firstly, extract subject heading vector of the topic of interest by filtering out function words, localizing Latent Dirichlet Allocation (LDA) and training the LDA model. Secondly, remove semantically repetitive vector component by constructing a synonymy word map. Lastly, organize and refine the special topic according to the similarity between the candidate news and the topic, and the density of topic-related terms. The experimental results show that the system has both simple operation and high accuracy, and it is stable enough to be applied for user-oriented special topic generation in practical applications. Zhi Tang 0001, Jianbo Xu, Liangcai Gao |
NLPCC | 5 |
| 2015 | Scientific Information Understanding via Open Educational Resources (OER)abstractScientific publication retrieval/recommendation has been investigated in the past decade. However, to the best of our knowledge, few efforts have been made to help junior scholars and graduate students to understand and consume the essence of those scientific readings. This paper proposes a novel learning/reading environment, OER-based Collaborative PDF Reader (OCPR), that incorporates innovative scaffolding methods that can: 1. auto-characterize student emerging information need while reading a paper; and 2. enable students to readily access open educational resources (OER) based on their information need. By using metasearch methods, we pre-indexed 1,112,718 OERs, including presentation videos, slides, algorithm source code, or Wikipedia pages, for 41,378 STEM publications. Based on the computational information need, we use text mining and heterogeneous graph mining algorithms to recommend high quality OERs to help students better understand the scientific content in the paper. Evaluation results and exit surveys for an information retrieval course show that the OCPR system alone with the recommended OERs can effectively assist graduate students better understand the complex STEM publications. For instance, 78.42% of participants believe the OCPR system and recommended OERs can provide precise and useful information they need, while 78.43% of them believe the recommended OERs are close to exactly what they need when reading the paper. From OER ranking viewpoint, MRR, MAP and NDCG results prove that learning to rank and cold start solutions can efficiently integrate different text and graph ranking features. Xiaozhong Liu 0001, Zhuoren Jiang, Liangcai Gao |
SIGIR | 3 |
| 2014 | Plane Geometry Figure Retrieval with Bag of ShapesabstractDigital education is serving an increasingly important function in most educational institutions, thus resulting in the production of a large number of digital documents online for education purposes. However, convenient ways to retrieve mathematic geometry questions are lacking because current retrieval systems largely rely on keywords instead of geometry figure images. This study focuses on plane geometry figure (PGF) image retrieval with the aim of retrieving relevant geometry images that contain more structural information than a question text stem. To fully use geometrical properties, a Bag-of-shapes (BoS) method is proposed to build the feature descriptor of an image. The BoS method contains either basic geometric primitives or dual-primitive structures along with several specific geometrical features for shape description. Based on the BoS feature descriptor, we apply cosine similarity with group feature weight as vector similarity measure for ranking to achieve high efficiency. For a PGF image query, the retrieval results are provided in an appropriate ranking order, which has high visual similarity with respect to human perception. Retrieval experiments and evaluation results show the effectiveness and efficiency of the proposed BoS shape descriptor. Lu Liu 0018, Xiaoqing Lu, Jingwei Qu, Liangcai Gao, Zhi Tang 0001 |
Document Analysis Systems | 5 |
| 2014 | Ground-Truth and Performance Evaluation for Page Layout Analysis of Born-Digital DocumentsabstractIn this paper, a new dataset is proposed for page layout analysis of born-digital documents. By extracting uniformly the document contents, an XML based data format is designed in terms of raw data and structure data. Utilizing a self-developed ground-truthing tool, a public dataset is constructed from diverse styles of document resources. With consideration of physical segmentation and logical labeling, automatic performance evaluation methods are adjusted to cope with different scenarios. The applications of the proposed dataset have shown that it is suitable for evaluating various layout analysis tasks. Zhi Tang 0001, Canhui Xu 0002, Liangcai Gao |
Document Analysis Systems | 4 |
| 2014 | Plane Geometry Figure Retrieval Based on Bilayer Geometric Attributed Graph MatchingabstractWith the development of computer-aided education and digital library, there have emerged large numbers of digital documents online for education purposes. However, it is far from convenient to retrieve mathematic geometry questions because current retrieval systems largely rely on keywords instead of geometry figure images. We focus on plane geometry figure (PGF) image retrieval aiming at retrieving relevant geometry images that hold more similar geometric attributes and structure properties than a question text stem. Motivated by Attribute Graph (AG), and aiming to catch more delicate local geometric attributes and overall structure layout, we propose a Bilayer Geometric Attributed Graph (Bilayer-GAG) matching method to retrieve the relevant PGF images. The root node of Bilayer-GAG catches the spatial relationships among its children - the graph elements of the second layer, the second layer contains curvilinear geometric primitives and linear nested AGs that consist of nodes and edges with geometric signatures. Then we calculate the overall matching cost in three perspectives and finally retrieve top-k relevant Bilayer-GAGs. For a PGF image query, the retrieval results are shown in an appropriate ranking order, which has high visual similarity with respect to human perception. Retrieval experiments results show the effectiveness and efficiency of the proposed Bilayer-GAG. Lu Liu 0018, Xiaoqing Lu, Songping Fu, Jingwei Qu, Liangcai Gao, Zhi Tang 0001 |
ICPR | 5 |
| 2014 | A mathematics retrieval system for formulae in layout presentationsabstractThe semantics of mathematical formulae depend on their spatial structure, and they usually exist in layout presentations such as PDF, LaTeX, and Presentation MathML, which challenges previous text index and retrieval methods. This paper proposes an innovative mathematics retrieval system along with the novel algorithms, which enables efficient formula index and retrieval from both webpages and PDF documents. Unlike prior studies, which require users to manually input formula markup language as query, the new system enables users to "copy" formula queries directly from PDF documents. Furthermore, by using a novel indexing and matching model, the system is aimed at searching for similar mathematical formulae based on both textual and spatial similarities. A hierarchical generalization technique is proposed to generate sub-trees from the semi-operator tree of formulae and support substructure match and fuzzy match. Experiments based on massive Wikipedia and CiteSeer repositories show that the new system along with novel algorithms, comparing with two representative mathematics retrieval systems, provides more efficient mathematical formula index and retrieval, while simplifying user query input for PDF documents. Xiaoyan Lin, Liangcai Gao, Zhi Tang 0001, Yingnan Xiao, Xiaozhong Liu 0001 |
SIGIR | 2 |
| 2014 | Mathematical formula identification and performance evaluation in PDF documents
Xiaoyan Lin, Liangcai Gao, Zhi Tang 0001, Josef B. Baker, Volker Sorge |
Int. J. Document Anal. Recognit. | 2 |
| 2014 | Automatic comic page segmentation based on polygon detection
Luyuan Li, Yongtao Wang, Zhi Tang 0001, Liangcai Gao |
Multim. Tools Appl. | 4 |
| 2013 | Unsupervised Speech Text Localization in Comic ImagesabstractLocalizing speech texts in comic images is a crucial step for catering the growing needs of reading comics on mobile devices. For example, automatically reading speech texts while adding sound effects alongside can not only render comic contents vividly but also help visually impaired readers. Unlike conventional text localization methods, we present an effective unsupervised speech text localization method in this paper that is free of training data. The proposed method consists of two major stages: (1) based on the concurrence of characters, the first stage of our method is to generate some of the character strings (a row or column of characters that align horizontally or vertically) from the comic images while the fonts and gaps of the adjacent characters within the character string are also obtained, (2) in the second stage, the obtained fonts and gaps of adjacent characters are used to detect rest of the character strings within the comic image via Bayesian classifier. The proposed method is tested on a dataset consists of 1000 comic images from ten printed comic series and provide satisfactory results. Luyuan Li, Yongtao Wang, Zhi Tang 0001, Xiaoqing Lu, Liangcai Gao |
ICDAR | 5 |
| 2013 | A Text Line Detection Method for Mathematical Formula RecognitionabstractText line detection is a prerequisite procedure of mathematical formula recognition, however, many incorrectly segmented text lines are often produced due to the two-dimensional structures of mathematics when using existing segmentation methods such as Projection Profiles Cutting or white space analysis. In consequence, mathematical formula recognition is adversely affected by these incorrectly detected text lines, with errors propagating through further processes. Aimed at mathematical formula recognition, we propose a text line detection method to produce reliable line segmentation. Based on the results produced by PPC, a learning based merging strategy is presented to combine incorrectly split text lines. In the merging strategy, the features of layout and text for a text line and those between successive lines are utilised to detect the incorrectly split text lines. Experimental results show that the proposed approach obtains good performance in detecting text lines from mathematical documents. Furthermore, the error rate in mathematical formula identification is reduced significantly through adopting the proposed text line detection method. Xiaoyan Lin, Liangcai Gao, Zhi Tang 0001, Josef B. Baker, Mohamed A. Alkalai, Volker Sorge |
ICDAR | 2 |
| 2013 | Stroke-Based Character Segmentation of Low-Quality Images on Ancient Chinese TabletabstractAncient Chinese tablets are invaluable in terms of historical and aesthetic value. Automatic character segmentation of images from degraded tablets poses a challenging problem. Therefore, this paper proposes a new character segmentation method that utilizes an enhanced stroke filter and an energy propagation process based on local layout information. A ground-truth dataset was established to evaluate the accuracy of the algorithm adopted by the proposed segmentation method. Experimental results indicate that the proposed method can effectively extract characters from low-quality ancient Chinese tablet images. Xiaoqing Lu, Zhi Tang 0001, Liangcai Gao |
ICDAR | 4 |
| 2012 | Performance Evaluation of Mathematical Formula IdentificationabstractThis paper presents a performance evaluation system for mathematical formula identification. First, a ground-truth dataset is constructed to facilitate the performance comparison of different mathematical formula identification algorithms. Statistics analysis of the dataset shows the diversities of the dataset to reflect the real-world documents. Second, a performance evaluation metric for mathematical formula identification is proposed, including the error type definitions and the scenario-adjustable scoring. The proposed metric enables in-depth analysis of mathematical formula identification systems in different scenarios. Finally, based on the proposed evaluation metric, a tool is developed to automatically evaluate mathematical formula identification results. It is worth noting that the ground-truth dataset and the evaluation tool are freely available for academic purpose. Xiaoyan Lin, Liangcai Gao, Zhi Tang 0001 |
Document Analysis Systems | 2 |
| 2012 | A graph-based method of newspaper article reconstruction
Liangcai Gao, Zhi Tang 0001, Xiaoyan Lin, Yongtao Wang |
ICPR | 1 |
| 2011 | A Table Detection Method for Multipage PDF Documents via Visual Seperators and Tabular StructuresabstractTable detection is always an important task of document analysis and recognition. In this paper, we propose a novel and effective table detection method via visual separators and geometric content layout information, targeting at PDF documents. The visual separators refer to not only the graphic ruling lines but also the white spaces to handle tables with or without ruling lines. Furthermore, we detect page columns in order to assist table region delimitation in complex layout pages. Evaluations of our algorithm on an e-Book dataset and a scientific document dataset show competitive performance. It is noteworthy that the proposed method has been successfully incorporated into a commercial software package for large-scale Chinese e-Book production. Liangcai Gao, Ruiheng Qiu, Zhi Tang 0001 |
ICDAR | 2 |
| 2011 | Metadata Extraction System for Chinese BooksabstractExtracting metadata from academic papers has attracted much attention from researchers in past years. But how to extract metadata automatically from books is still seldom discussed. In this paper, we address this task on Chinese books and present a system to extract metadata from the title page of a book. This system consists of three components: metadata segmentation, metadata labeling, and post-processing. Different strategies are adopted in the system to identify different metadata types, and a variety of information sources, including geometric layout, linguistic, semantic content and header-footer, are used to accommodate the wide range of metadata layouts. Experimental results on real-world data have demonstrated the effectiveness of the proposed system. Liangcai Gao, Yingmin Tang, Zhi Tang 0001 |
ICDAR | 1 |
| 2011 | Mathematical Formula Identification in PDF DocumentsabstractRecognizing mathematical expressions in PDF documents is a new and important field in document analysis. It is quite different from extracting mathematical expressions in image-based documents. In this paper, we propose a novel method by combining rule-based and learning-based methods to detect both isolated and embedded mathematical expressions in PDF documents. Moreover, various features of formulas, including geometric layout, character and context content, are used to adapt to a wide range of formula types. Experimental results show satisfactory performance of the proposed method. Furthermore, the method has been successfully incorporated into a commercial software package for large-scale Chinese e-Book production. Xiaoyan Lin, Liangcai Gao, Zhi Tang 0001 |
ICDAR | 2 |
| 2011 | An Efficient Pre-processing Method to Identify Logical Components from PDF Documents
Ying Liu 0006, Liangcai Gao |
PAKDD (1) | 3 |
| 2010 | MC-JBIG2: an improved algorithm for Chinese textual image compression
Kui Hu, Zhi Tang 0001, Liangcai Gao, Yadong Mu |
Int. J. Document Anal. Recognit. | 3 |
| 2009 | Analysis of Book Documents' Table of Content Based on ClusteringabstractTable of contents (TOC) recognition has attracted a great deal of attention in recent years. After reviewing the merits and drawbacks of the existing TOC recognition methods, we have observed that book documents are multi-page documents with intrinsic local format consistency. Based on this finding we introduce an automatic TOC analysis method through clustering. This method first detects the decorative elements in TOC pages. Then it learns a layout model used in the TOC pages through clustering. Finally, it generates TOC entries and extracts their hierarchical structure under the guidance of the model. More specifically, broken lines are taken into account in the method. Experimental results show that this method achieves high accuracy and efficiency. In addition, this method has been successfully applied in a commercial e-book production software package. Liangcai Gao, Zhi Tang 0001, Yimin Chu |
ICDAR | 1 |
| 2008 | Comprehensive Global Typography Extraction System for Electronic Book DocumentsabstractBook documents usually have consistent typographies throughout the whole book, including headers, footers, columns, text line directions, and fonts used in the each level of headings. Such document-level typography information is of great value for downstream document processing applications. This paper presents a document analysis system that can extract a comprehensive set of typographies used in book documents. The system consists of several components: recognition of fonts used in the body text and chapter headings; detection of page body area, headers and footers; detection of columns, text line direction and line spacing of body text. Page-association is employed in the system. The preliminary experimental results demonstrate the effectiveness of the system. Liangcai Gao, Zhi Tang 0001, Ruiheng Qiu |
Document Analysis Systems | 1 |