VLDB 2026 Research / reviewers in the wild / expert
Zhi Tang 0001
dblp:16/4222-1
· DBLP profile ↗
121ranked-venue papers
0as first author
27since 2021 · last 2025
0000-0002-6021-8357ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 65 · 17 since 2021Graphics, computer vision, multimedia, augmented reality and games · 41 · 10 since 2021Databases, data management, data science and information retrieval · 39 · 6 since 2021Applied, interdisciplinary, general and emerging computing · 17 · 4 since 2021Security and privacy · 2Human-computer interaction and ubiquitous computing · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | GlyphSR: A Simple Glyph-Aware Framework for Scene Text Image Super-ResolutionabstractThe goal of scene text image super-resolution (STISR) is to enhance the clarity of text within line images, thereby improving readability and enabling more accurate text recognition. However, existing STISR methods often rely heavily on Text Prior (TP) derived from trained recognizers, which can be unreliable and may lead to incorrect glyph restoration. Text images contain two crucial types of information: semantic content from word meanings and structural details from glyphs. When semantic information is unreliable, accurate perception of glyph structures becomes essential. This paper introduces GlyphSR, a novel STISR framework that addresses three key challenges: precise extraction, effective learning, and optimal utilization of glyph structural information. GlyphSR incorporates the Glyph Extraction Module (GEM), a training-free approach leveraging the Segment Anything Model (SAM) to accurately extract character-level glyphs. The Glyph Perception Module (GPM) models and learns glyph structures through segmentation and classification tasks, while the Glyph Fusion Module (GFM) integrates glyph information to enhance overall STISR model performance. Extensive experiments on the TextZoom dataset demonstrate that GlyphSR achieves a new state-of-the-art performance. Baole Wei, Liangcai Gao, Zhi Tang 0001 |
AAAI | 4 |
| 2025 | Vote & Mix: Plug-and-Play Token Reduction for Efficient Vision TransformerabstractDespite the remarkable success of Vision Transformers (ViTs) in various visual tasks, they are often hindered by substantial computational cost. In this work, we introduce Vote&Mix (VoMix), a plug-and-play and parameter-free token reduction method, which can be readily applied to off-the-shelf ViT models without any training. VoMix tackles the computational redundancy of ViTs by identifying tokens with high homogeneity through a layer-wise token similarity voting mechanism. Subsequently, the selected tokens are mixed into the retained set, thereby preserving visual information. Experiments demonstrate VoMix significantly improves the speed-accuracy tradeoff of ViTs on both images and videos. Without any training, VoMix achieves a 2× increase in throughput of existing ViT-H on ImageNet-1K and a 2.4× increase in throughput of existing ViT-L on Kinetics-400 video dataset, with a mere 0.3% drop in top-1 accuracy. Shuai Peng, Di Fu, Baole Wei, Liangcai Gao, Zhi Tang 0001 |
ICME | 5 |
| 2025 | Enhancing Graph Representation Learning with Localized Topological FeaturesabstractRepresentation learning on graphs is a fundamental problem that can be crucial in various tasks. Graph neural networks, the dominant approach for graph representation learning, are limited in their representation power. Therefore, it can be beneficial to explicitly extract and incorporate high-order topological and geometric information into these models. In this paper, we propose a principled approach to extract the rich connectivity information of graphs based on the theory of persistent homology. Our method utilizes the topological features to enhance the representation learning of graph neural networks and achieve state-of-the-art performance on various node classification and link prediction benchmarks. We also explore the option of end-to-end learning of the topological features, i.e., treating topological computation as a differentiable operator during learning. Our theoretical analysis and empirical study provide insights and potential guidelines for employing topological features in graph learning tasks. Zuoyu Yan, Qi Zhao 0007, Ze Ye, Tengfei Ma 0001, Liangcai Gao, Zhi Tang 0001, Yusu Wang 0001, Chao Chen 0012 |
J. Mach. Learn. Res. | 6 |
| 2024 | Maskstr: Guide Scene Text Recognition Models with MaskingabstractText recognition in information loss scenarios like blurriness, occlusion, and perspective distortion is challenging in real-world applications. To enhance robustness, some studies use extra unlabeled data for encoder pretraining. Others focus on improving decoder context reasoning. However, pretraining methods require abundant unlabeled data and high computing resources, while decoder-based approaches risk over-correction. In this paper, we propose MaskSTR, a dual-branch training framework for STR models, using patch masking to simulate information loss. MaskSTR guides visual representation learning, improving robustness to information loss conditions without extra data or training stages. Furthermore, we introduce Block Masking, a novel and straightforward mask generation method, for further performance enhancement. Experiments demonstrate MaskSTR’s effectiveness across CTC, attention, and Transformer decoding methods, achieving significant performance gains and setting new state-of-the-art results. Baole Wei, Minghang He, Liangcai Gao, Duoyou Zhou, Xiang Bai, Zhi Tang 0001 |
ICASSP | 6 |
| 2024 | Recognition-Guided Diffusion Model for Scene Text Image Super-ResolutionabstractScene Text Image Super-Resolution (STISR) aims to enhance the resolution and legibility of text within low-resolution (LR) images, consequently elevating recognition accuracy in Scene Text Recognition (STR). Previous methods predominantly employ discriminative Convolutional Neural Networks (CNNs) augmented with diverse forms of text guidance to address this issue. Nevertheless, they remain deficient when confronted with severely blurred images, due to their insufficient generation capability when little structural or semantic information can be extracted from original images. Therefore, we introduce RGDiffSR, a Recognition-Guided Diffusion model for scene text image Super-Resolution, which exhibits great generative diversity and fidelity even in challenging scenarios. Moreover, we propose a Recognition-Guided Denoising Network, to guide the diffusion model generating LR-consistent results through succinct semantic guidance. Experiments on the TextZoom dataset demonstrate the superiority of RGDiffSR over prior state-of-the-art methods in both text recognition accuracy and image fidelity. Liangcai Gao, Zhi Tang 0001, Baole Wei |
ICASSP | 3 |
| 2024 | An Efficient Subgraph GNN with Provable Substructure Counting PowerabstractWe investigate the enhancement of graph neural networks' (GNNs) representation power through their ability in substructure counting. Recent advances have seen the adoption of subgraph GNNs, which partition an input graph into numerous subgraphs, subsequently applying GNNs to each to augment the graph's overall representation. Despite their ability to identify various substructures, subgraph GNNs are hindered by significant computational and memory costs. In this paper, we tackle a critical question: Is it possible for GNNs to count substructures both efficiently and provably? Our approach begins with a theoretical demonstration that the distance to rooted nodes in subgraphs is key to boosting the counting power of subgraph GNNs. To avoid the need for repetitively applying GNN across all subgraphs, we introduce precomputed structural embeddings that encapsulate this crucial distance information. Experiments validate that our proposed model retains the counting power of subgraph GNNs while achieving significantly faster performance. Zuoyu Yan, Junru Zhou, Liangcai Gao, Zhi Tang 0001, Muhan Zhang |
KDD | 4 |
| 2024 | Training-Free Open-Ended Object Detection and Segmentation via Attention as PromptsabstractExisting perception models achieve great success by learning from large amounts of labeled data, but they still struggle with open-world scenarios. To alleviate this issue, researchers introduce open-set perception tasks to detect or segment unseen objects in the training set. However, these models require predefined object categories as inputs during inference, which are not available in real-world scenarios. Recently, researchers pose a new and more practical problem, i.e., open-ended object detection, which discovers unseen objects without any object categories as inputs. In this paper, we present VL-SAM, a training-free framework that combines the generalized object recognition model (i.e., Vision-Language Model) with the generalized object localization model (i.e., Segment-Anything Model), to address the open-ended object detection and segmentation task. Without additional training, we connect these two generalized models with attention maps as the prompts. Specifically, we design an attention map generation module by employing head aggregation and a regularized attention flow to aggregate and propagate attention maps across all heads and layers in VLM, yielding high-quality attention maps. Then, we iteratively sample positive and negative points from the attention maps with a prompt generation module and send the sampled points to SAM to segment corresponding objects. Experimental results on the long-tail instance segmentation dataset (LVIS) show that our method surpasses the previous open-ended method on the object detection task and can provide additional instance segmentation masks. Besides, VL-SAM achieves favorable performance on the corner case object detection dataset (CODA), demonstrating the effectiveness of VL-SAM in real-world applications. Moreover, VL-SAM exhibits good model generalization that can incorporate various VLMs and SAMs. Yongtao Wang, Zhi Tang 0001 |
NeurIPS | 3 |
| 2024 | Table image dewarping with key element segmentation
Zhi Tang 0001, Liangcai Gao |
Int. J. Document Anal. Recognit. | 2 |
| 2023 | Graph-Graph Context Dependency Attention for Graph Edit DistanceabstractAs a popular similarity measurement, Graph Edit Distance (GED) can be applied to various downstream tasks. Traditional GED solvers have difficulty in generalization and time efficiency, so deep graph similarity learning models have emerged as a promising trend. However, existing methods lack distinguishing embeddings for cross-dependencies between graphs and intra-graph representations. In this paper, we propose a deep network architecture GED-CDA by introducing a graph-graph context dependency attention module to enhance embeddings. Our graph-graph context dependency attention module consists of a cross-attention layer and a self-attention layer, which can both learn the interactions between graph pairs and capture the connections within graphs. In addition, we also introduce a spectral encoding strategy to bring topological information of graphs into node representations. Extensive experimental results on three real-world graph datasets demonstrate the effectiveness of our proposed GED-CDA in the GED task. Ruiqi Jia, Xianbing Feng, Xiaoqing Lyu, Zhi Tang 0001 |
ICASSP | 4 |
| 2023 | An Iterative Graph Learning Convolution Network for Key Information Extraction Based on the Document Inductive Bias
Jiyao Deng, Zhi Tang 0001, Liangcai Gao |
ICDAR (3) | 4 |
| 2022 | CMUA-Watermark: A Cross-Model Universal Adversarial Watermark for Combating DeepfakesabstractMalicious applications of deepfakes (i.e., technologies generating target facial attributes or entire faces from facial images) have posed a huge threat to individuals' reputation and security. To mitigate these threats, recent studies have proposed adversarial watermarks to combat deepfake models, leading them to generate distorted outputs. Despite achieving impressive results, these adversarial watermarks have low image-level and model-level transferability, meaning that they can protect only one facial image from one specific deepfake model. To address these issues, we propose a novel solution that can generate a Cross-Model Universal Adversarial Watermark (CMUA-Watermark), protecting a large number of facial images from multiple deepfake models. Specifically, we begin by proposing a cross-model universal attack pipeline that attacks multiple deepfake models iteratively. Then, we design a two-level perturbation fusion strategy to alleviate the conflict between the adversarial watermarks generated by different facial images and models. Moreover, we address the key problem in cross-model optimization with a heuristic approach to automatically find the suitable attack step sizes for different models, further weakening the model-level conflict. Finally, we introduce a more reasonable and comprehensive evaluation method to fully test the proposed method and compare it with existing ones. Extensive experimental results demonstrate that the proposed CMUA-Watermark can effectively distort the fake facial images generated by multiple deepfake models while achieving a better performance than existing methods. Our code is available at https://github.com/VDIGPKU/CMUA-Watermark. Yongtao Wang, Zhaoyu Chen 0001, Zhi Tang 0001, Jingdong Chen, Weisi Lin, Kai-Kuang Ma |
AAAI | 6 |
| 2022 | Molecular Formula Image Segmentation with Shape Constraint Loss and Data AugmentationabstractThe increasing demand for molecular formula image data leads to formidable pressure for researchers. Most existing image segmentation approaches can not be directly utilized for molecules, and how to improve the coverage fineness and generate a large amount of labeled training data is worthy of further exploration. To this end, we establish a deep learning based molecular formula image segmentation model (DL-MFS). Specifically, we design a shape constraint loss (SCL) function to refine the detection frame position and propose a rule-based molecular formula image data augmentation method for solving the bottleneck problem that the lack of training data. Experimental results demonstrate the effectiveness of the proposed segmentation model. Ruiqi Jia, Baole Wei, Guanren Qiao, Xiaoqing Lyu, Zhi Tang 0001 |
BIBM | 7 |
| 2022 | MIR: A Benchmark for Molecular Image Retrival with a Cross-modal Pretraining FrameworkabstractMolecular image retrieval is one of the crucial steps in automatic mining and utilization of biochemistry-related literatures, which is also a relatively open and challenging task in cross fields of biochemistry and artificial intelligence. The challenges come from two aspects: 1) there is a lack of open datasets and evaluation criteria for molecular image retrieval. 2) Common retrieval methods always ignore that molecular image retrieval has cross-modal information of both images and SMILES texts. To address the first challenge, we firstly construct a new molecular image retrieval benchmark, named MIR, including 130770 molecular images, labeled structural similarity, and reasonable evaluation metrics. Faced with the second challenge, we propose an effective cross-modal pre-training framework for molecular image retrieval following CLIP. Experimental results reflect the effectiveness of our proposed benchmark MIR and cross-modal pre-training framework. Baole Wei, Ruiqi Jia, Shihan Fu, Xiaoqing Lyu, Liangcai Gao, Zhi Tang 0001 |
BIBM | 6 |
| 2022 | Cycle Representation Learning for Inductive Relation PredictionabstractIn recent years, algebraic topology and its modern development, the theory of persistent homology, has shown great potential in graph representation learning. In this paper, based on the mathematics of algebraic topology, we propose a novel solution for inductive relation prediction, an important learning task for knowledge graph completion. To predict the relation between two entities, one can use the existence of rules, namely a sequence of relations. Previous works view rules as paths and primarily focus on the searching of paths between entities. The space of rules is huge, and one has to sacrifice either efficiency or accuracy. In this paper, we consider rules as cycles and show that the space of cycles has a unique structure based on the mathematics of algebraic topology. By exploring the linear structure of the cycle space, we can improve the searching efficiency of rules. We propose to collect cycle bases that span the space of cycles. We build a novel GNN framework on the collected cycles to learn the representations of cycles, and to predict the existence/non-existence of a relation. Our method achieves state-of-the-art performance on benchmarks. Zuoyu Yan, Tengfei Ma 0001, Liangcai Gao, Zhi Tang 0001, Chao Chen 0012 |
ICML | 4 |
| 2022 | Compute Like Humans: Interpretable Step-by-step Symbolic Computation with Deep Neural NetworkabstractNeural network capability in symbolic computation has emerged in much recent work. However, symbolic computation is always treated as an end-to-end blackbox prediction task, where human-like symbolic deductive logic is missing. In this paper, we argue that any complex symbolic computation can be broken down to a sequence of finite Fundamental Computation Transformations (FCT), which are grounded as certain mathematical expression computation transformations. The entire computation sequence represents a full human understandable symbolic deduction process. Instead of studying on different end-to-end neural network applications, this paper focuses on approximating FCT which further build up symbolic deductive logic. To better mimic symbolic computations with math expression transformations, we propose a novel tree representation learning architecture GATE (Graph Aggregation Transformer Encoder) for math expressions. We generate a large-scale math expression transformation dataset for training purpose and collect a real-world dataset for validation. Experiments demonstrate the feasibility of producing step-by-step human-like symbolic deduction sequences with the proposed approach, which outperforms other neural network approaches and heuristic approaches. Shuai Peng, Di Fu, Yijun Liang, Gu Xu, Liangcai Gao, Zhi Tang 0001 |
KDD | 7 |
| 2022 | BEVFusion: A Simple and Robust LiDAR-Camera Fusion FrameworkabstractFusing the camera and LiDAR information has become a de-facto standard for 3D object detection tasks. Current methods rely on point clouds from the LiDAR sensor as queries to leverage the feature from the image space. However, people discovered that this underlying assumption makes the current fusion framework infeasible to produce any prediction when there is a LiDAR malfunction, regardless of minor or major. This fundamentally limits the deployment capability to realistic autonomous driving scenarios. In contrast, we propose a surprisingly simple yet novel fusion framework, dubbed BEVFusion, whose camera stream does not depend on the input of LiDAR data, thus addressing the downside of previous methods. We empirically show that our framework surpasses the state-of-the-art methods under the normal training settings. Under the robustness training settings that simulate various LiDAR malfunctions, our framework significantly surpasses the state-of-the-art methods by 15.7% to 28.9% mAP. To the best of our knowledge, we are the first to handle realistic LiDAR malfunction and can be deployed to realistic scenarios without any post-processing procedure. Tingting Liang, Hongwei Xie, Kaicheng Yu, Zhongyu Xia, Yongtao Wang, Zhi Tang 0001 |
NeurIPS | 9 |
| 2022 | Neural Approximation of Graph Topological FeaturesabstractTopological features based on persistent homology capture high-order structural information so as to augment graph neural network methods. However, computing extended persistent homology summaries remains slow for large and dense graphs and can be a serious bottleneck for the learning pipeline. Inspired by recent success in neural algorithmic reasoning, we propose a novel graph neural network to estimate extended persistence diagrams (EPDs) on graphs efficiently. Our model is built on algorithmic insights, and benefits from better supervision and closer alignment with the EPD computation algorithm. We validate our method with convincing empirical results on approximating EPDs and downstream graph representation learning tasks. Our method is also efficient; on large and dense graphs, we accelerate the computation by nearly 100 times. Zuoyu Yan, Tengfei Ma 0001, Liangcai Gao, Zhi Tang 0001, Yusu Wang 0001, Chao Chen 0012 |
NeurIPS | 4 |
| 2022 | CBNet: A Composite Backbone Network Architecture for Object Detectionabstracttop-performing object detectors depend heavily on backbone networks, whose advances bring consistent performance gains through exploring more effective network structures. In this paper, we propose a novel and flexible backbone framework, namely CBNet, to construct high-performance detectors using existing open-source pre-trained backbones under the pre-training fine-tuning paradigm. In particular, CBNet architecture groups multiple identical backbones, which are connected through composite connections. Specifically, it integrates the high- and low-level features of multiple identical backbone networks and gradually expands the receptive field to more effectively perform object detection. We also propose a better training strategy with auxiliary supervision for CBNet-based detectors. CBNet has strong generalization capabilities for different backbones and head designs of the detector architecture. Without additional pre-training of the composite backbone, CBNet can be adapted to various backbones (i.e., CNN-based vs. Transformer-based) and head designs of most mainstream detectors (i.e., one-stage vs. two-stage, anchor-based vs. anchor-free-based). Experiments provide strong evidence that, compared with simply increasing the depth and width of the network, CBNet introduces a more efficient, effective, and resource-friendly way to build high-performance backbone networks. Particularly, our CB-Swin-L achieves 59.4% box AP and 51.6% mask AP on COCO test-dev under the single-model and single-scale testing protocol, which are significantly better than the state-of-the-art results (i.e., 57.7% box AP and 50.2% mask AP) achieved by Swin-L, while reducing the training time by 6×. With multi-scale testing, we push the current best single model result to a new record of 60.1% box AP and 52.3% mask AP without using extra training data. Code is available at https://github.com/VDIGPKU/CBNetV2. Ting-Ting Liang, Xiaojie Chu, Yongtao Wang, Zhi Tang 0001, Jingdong Chen, Haibin Ling |
IEEE Trans. Image Process. | 5 |
| 2021 | Attention-enhanced Graph Cross-convolution for Protein-Ligand Binding Affinity PredictionabstractThe binding affinity between drugs and proteins is a substantial part of the drug discovery process. Graph neural networks (GNNs) have shown great promise in graph-related structure by learning the representations of graphs, which are suitable for tasks such as binding affinity prediction. However, most of the existing GNN architectures only pay attention to the information flow on a single graph, while the interaction between two graphs is unconcerned. In this paper, we propose an attention-enhanced graph cross-convolution network (GCAT) to explore binding affinity on pure 3D atomistic geometry. It consists of two components: cross-convolution and self-attention pooling. Specifically, cross-convolution performs an aggregate-update mechanism to simulate the interaction between the protein and the drug, then self-attention pooling is adopted to capture global interactions and get graph-level representations. Extensive experiments conducted on the PDB-Bind dataset demonstrate the effectiveness of our GCAT. Xianbing Feng, Jingwei Qu, Tianle Wang 0005, Bei Wang 0004, Xiaoqing Lyu, Zhi Tang 0001 |
BIBM | 6 |
| 2021 | Understanding Multivariate Drug-Target-DiseaseInterdependence via Event-GraphabstractDrug repurposing aims at identifying new indications for approved drugs that are outside the scope of the original indications. Understanding the acting mechanism among drugs, protein targets, and diseases, especially the interdependent and indecomposable relationships, is a critical step. However, most existing methods rely on pairwise relationships. To model the biological interactions between the three types of entities, which are likely ignored by the pairwise paradigm, we propose an end-to-end Event-Graph Neural Network (EGNN) to predict multivariate relationships of drugs, targets, and diseases for drug repurposing. Specifically, we introduce the event to describe the interdependence of drug-target-disease as a complete semantic unit and design the Event-Graph to model the multivariate relationships. To predict the potential relationships, we perform the representation learning on the Event-Graph by a bidirectional aggregating operation and an event-level attention mechanism. Experimental results on real-world datasets demonstrate the effectiveness and promising performance of EGNN compared with several competitive methods. Jingwei Qu, Bei Wang 0004, Zhixun Li, Xiaoqing Lyu, Zhi Tang 0001 |
BIBM | 5 |
| 2021 | OPANAS: One-Shot Path Aggregation Network Architecture Search for Object DetectionabstractRecently, neural architecture search (NAS) has been exploited to design feature pyramid networks (FPNs) and achieved promising results for visual object detection. Encouraged by the success, we propose a novel One-Shot Path Aggregation Network Architecture Search (OPANAS) algorithm, which significantly improves both searching efficiency and detection accuracy. Specifically, we first introduce six heterogeneous information paths to build our search space, namely top-down, bottom-up, fusing-splitting, scale-equalizing, skip-connect and none. Second, we propose a novel search space of FPNs, in which each FPN candidate is represented by a densely-connected directed acyclic graph (each node is a feature pyramid and each edge is one of the six heterogeneous information paths). Third, we propose an efficient one-shot search method to find the optimal path aggregation architecture; specifically, we first train a super-net and then find the optimal candidate with an evolutionary algorithm. Experimental results demonstrate the efficacy of the proposed OPANAS for object detection: (1) OPANAS is more efficient than state-of-the-art methods (e.g., NAS-FPN and Auto-FPN) at significantly smaller searching cost (e.g., only 4 GPU days on MS-COCO); (2) the optimal architecture found by OPANAS significantly improves main-stream detectors including RetinaNet, Faster R-CNN and Cascade R-CNN, by 2.3∼3.2 % mAP compared to their FPN counterparts; and (3) a new state-of-the-art accuracy-speed trade-off (52.2 % mAP at 7.6 FPS) is achieved at smaller training costs than comparable recent arts. Code will be released at https://github.com/VDIGPKU/OPANAS. Tingting Liang, Yongtao Wang, Zhi Tang 0001, Guosheng Hu, Haibin Ling |
CVPR | 3 |
| 2021 | Rethinking Table Structure Recognition Using Sequence Labeling Methods
Yilun Huang 0001, Lemeng Pan, Yongshuai Huang, Zhi Tang 0001, Liangcai Gao |
ICDAR (2) | 7 |
| 2021 | Image to LaTeX with Graph Neural Network for Mathematical Formula Recognition
Shuai Peng, Liangcai Gao, Zhi Tang 0001 |
ICDAR (2) | 4 |
| 2021 | Formula Citation Graph Based Mathematical Information Retrieval
Liangcai Gao, Zhuoren Jiang, Zhi Tang 0001 |
ICDAR (1) | 4 |
| 2021 | Rpattack: Refined Patch Attack on General Object DetectorsabstractNowadays, general object detectors like YOLO and Faster R-CNN as well as their variants are widely exploited in many applications. Many works have revealed that these detectors are extremely vulnerable to adversarial patch attacks. The perturbed regions generated by previous patch-based attack works on object detectors are very large which are not necessary for attacking and perceptible for human eyes. To generate much less but more efficient perturbation, we propose a novel patch-based method for attacking general object detectors. Firstly, we propose a patch selection and refining scheme to find the pixels which have the greatest importance for attack and remove the inconsequential perturbations gradually. Then, for a stable ensemble attack, we balance the gradients of detectors to avoid over-optimizing one of them during the training phase. Our RPAttack can achieve an amazing missed detection rate of 100% for both Yolo v4 and Faster R-CNN while only modifies 0.32% pixels on VOC 2007 test set. Our code is available at https://github.com/VDIGPKU/RPAttack. Yongtao Wang, Zhaoyu Chen 0001, Zhi Tang 0001, Kai-Kuang Ma |
ICME | 4 |
| 2021 | Link Prediction with Persistent Homology: An Interactive ViewabstractLink prediction is an important learning task for graph-structured data. In this paper, we propose a novel topological approach to characterize interactions between two nodes. Our topological feature, based on the extended persistent homology, encodes rich structural information regarding the multi-hop paths connecting nodes. Based on this feature, we propose a graph neural network method that outperforms state-of-the-arts on different benchmarks. As another contribution, we propose a novel algorithm to more efficiently compute the extended persistence diagrams for graphs. This algorithm can be generally applied to accelerate many other topological methods for graph learning tasks. Zuoyu Yan, Tengfei Ma 0001, Liangcai Gao, Zhi Tang 0001, Chao Chen 0012 |
ICML | 4 |
| 2021 | Adaptive Edge Attention for Graph Matching with OutliersabstractGraph matching aims at establishing correspondence between node sets of given graphs while keeping the consistency between their edge sets. However, outliers in practical scenarios and equivalent learning of edge representations in deep learning methods are still challenging. To address these issues, we present an Edge Attention-adaptive Graph Matching (EAGM) network and a novel description of edge features. EAGM transforms the matching relation between two graphs into a node and edge classification problem over their assignment graph. To explore the potential of edges, EAGM learns edge attention on the assignment graph to 1) reveal the impact of each edge on graph matching, as well as 2) adjust the learning of edge representations adaptively. To alleviate issues caused by the outliers, we describe an edge by aggregating the semantic information over the space spanned by the edge. Such rich information provides clear distinctions between different edges (e.g., inlier-inlier edges vs. inlier-outlier edges), which further distinguishes outliers in the view of their associated edges. Extensive experiments demonstrate that EAGM achieves promising matching quality compared with state-of-the-arts, on cases both with and without outliers. Our source code along with the experiments is available at https://github.com/bestwei/EAGM. Jingwei Qu, Haibin Ling, Xiaoqing Lyu, Zhi Tang 0001 |
IJCAI | 5 |
| 2020 | CBNet: A Novel Composite Backbone Network Architecture for Object DetectionabstractIn existing CNN based detectors, the backbone network is a very important component for basic feature1 extraction, and the performance of the detectors highly depends on it. In this paper, we aim to achieve better detection performance by building a more powerful backbone from existing ones like ResNet and ResNeXt. Specifically, we propose a novel strategy for assembling multiple identical backbones by composite connections between the adjacent backbones, to form a more powerful backbone named Composite Backbone Network (CBNet). In this way, CBNet iteratively feeds the output features of the previous backbone, namely high-level features, as part of input features to the succeeding backbone, in a stage-by-stage fashion, and finally the feature maps of the last backbone (named Lead Backbone) are used for object detection. We show that CBNet can be very easily integrated into most state-of-the-art detectors and significantly improve their performances. For example, it boosts the mAP of FPN, Mask R-CNN and Cascade R-CNN on the COCO dataset by about 1.5 to 3.0 points. Moreover, experimental results show that the instance segmentation results can be improved as well. Specifically, by simply integrating the proposed CBNet into the baseline detector Cascade Mask R-CNN, we achieve a new state-of-the-art result on COCO dataset (mAP of 53.3) with a single model, which demonstrates great effectiveness of the proposed CBNet architecture. Code will be made available at https://github.com/PKUbahuangliuhe/CBNet. Yongtao Wang, Siwei Wang 0001, Tingting Liang, Qijie Zhao, Zhi Tang 0001, Haibin Ling |
AAAI | 6 |
| 2020 | Automatic Generation of Headlines for Online Math QuestionsabstractMathematical equations are an important part of dissemination and communication of scientific information. Students, however, often feel challenged in reading and understanding math content and equations. With the development of the Web, students are posting their math questions online. Nevertheless, constructing a concise math headline that gives a good description of the posted detailed math question is nontrivial. In this study, we explore a novel summarization task denoted as geNerating A concise Math hEadline from a detailed math question (NAME). Compared to conventional summarization tasks, this task has two extra and essential constraints: 1) Detailed math questions consist of text and math equations which require a unified framework to jointly model textual and mathematical information; 2) Unlike text, math equations contain semantic and structural features, and both of them should be captured together. To address these issues, we propose MathSum, a novel summarization model which utilizes a pointer mechanism combined with a multi-head attention mechanism for mathematical representation augmentation. The pointer mechanism can either copy textual tokens or math tokens from source questions in order to generate math headlines. The multi-head attention mechanism is designed to enrich the representation of math equations by modeling and integrating both its semantic and structural features. For evaluation, we collect and make available two sets of real-world detailed math questions along with human-written math headlines, namely EXEQ-300k and OFEQ-10k. Experimental results demonstrate that our model (MathSum) significantly outperforms state-of-the-art models for both the EXEQ-300k and OFEQ-10k datasets. Dafang He, Zhuoren Jiang, Liangcai Gao, Zhi Tang 0001, C. Lee Giles |
AAAI | 5 |
| 2020 | An Enhanced LRMC Method for Drug Repositioning via GCN-based HIN EmbeddingabstractDrug repositioning has received ever-increasing attention in the field of drug discovery over the last few years. However, the high efficient prediction methods taking full advantage of heterogeneous information networks (HINs) still deserves further research. To this end, this paper proposes an approach for drug repositioning via integrating HINs embedding and link prediction for more potential drug-target interactions. To utilize multiple side information, we introduce a graph convolutional network (GCN) based embedding method for HINs. The obtained drug-related and target-related information is adopted to improve the low-rank matrix completion (LRMC) model. Moreover, a regulation for alleviating the noise of negative samples is designed to enhance the optimization of LRMC. The experiments conducted on the comparative database demonstrate that the proposed method is more effective than the existing approaches in the prediction of drug repositioning. Xiaoqing Lyu, Bei Wang 0004, Zhi Tang 0001, Zhenming Liu |
BIBM | 5 |
| 2020 | Dual Autoencoder Network with Swap Reconstruction for Cold-Start RecommendationabstractCold-start is a long-standing and challenging problem in recommendation systems. To tackle this issue, many cross-domain recommendation approaches are proposed. However, most of them follow a two-stage embedding-and-mapping paradigm, which is hard to be optimized. Besides, they ignore the structure information of the user-item interaction graph, resulting in that the embedding is insufficient to capture the latent collaborative filtering effect. In this paper, we propose a Dual Autoencoder Network (DAN), which implements cross-domain recommendations to cold-start users in an end-to-end manner. The graph convolutional network (GCN) based encoder in DAN explicitly captures high-order collaborative information in user-item interaction graphs. The two-branched decoder is proposed for fully exploiting the data across domains, and therefore the elaborate reconstruction constraints are obtained under a domain swapping strategy. Experiments on two pairs of real-world cross-domain datasets demonstrate that DAN outperforms existing state-of-the-art methods. Bei Wang 0004, Xiaoqing Lyu, Zhi Tang 0001 |
CIKM | 5 |
| 2020 | Follow The Curve: Arbitrarily Oriented Scene Text Detection Using Key Points Spotting And Curve PredictionabstractDetecting arbitrarily oriented text in natural images is still a challenging and unsolved problem in multimedia. In this work, we propose an efficient and accurate scene text detector. The detector first detects key points which are carefully designed and semantically meaningful. Then the key points are learnt to be associated together to form a hexagon for each text instance. Starting from a key point, the detector then predicts curves alongside the border of the text region. Simple heuristic post-processing followed by the predicted curve resulting in more accurate text region prediction. The predicted key points are used as anchors to correct errors from the search process. The proposed method is efficient since it is a single stage key point detection method with simple post-processing. It is also effective and shows state-of-the-art or comparable performance on several benchmark datasets. Dafang He, Xiao Yang 0004, Zhi Tang 0001, Daniel Kifer, C. Lee Giles |
ICME | 4 |
| 2020 | Topology Discriminative Pooling via Graph Kernel Derived Nodes and Edges Co-EmbeddingabstractAs a vital component of graph neural networks, graph pooling remains largely an open problem. Most existing methods are induced by empirical insights, while ignore the effect of topology information for graph coarsening guidance. In this paper, we propose a Topology Discriminative Pooling (TD-Pool) approach, which derives graph pooling architectures from graph kernels for semantic and structural feature co-embedding. Concretely, TD-Pool contains two graph modules, namely recurrent edge module and convolutional node module, for respectively learning topological/semantic embeddings on edges/nodes. The edge module combines Weisfeiler-Lehman graph kernel with neural networks, which captures topology features in a data-driven manner. It is deployed on the line graphs, whose nodes correspond to the original graph edges for semantic-topology decoupling. The node module learns semantic features with graph convolution in the spectral domain, which acts on the original graphs nodes. Extensive experiments demonstrate the effectiveness of our TD-Pool in graph-level tasks. Xiaoqing Lyu, Zhi Tang 0001 |
ICME | 4 |
| 2020 | Dual Loss for Manga Character Recognition with Imbalanced Training DataabstractManga character recognition is a key technology for manga character retrieval and verification. This task is very challenging since the manga character images have a long-tailed distribution and large quality variations. Training models with cross-entropy softmax loss on such imbalanced data would introduce biases to feature and class weight norms. To handle this problem, we propose a novel dual loss which is the sum of two losses: dual ring loss and dual adaptive re-weighting loss. Dual ring loss combines weight and feature soft normalization and serves as a regularization term to softmax loss. Dual adaptive re-weighting loss re-weights softmax loss according to the norm of both feature and class weight. With the proposed losses, we have achieved encouraging results on the Manga109 benchmark. Specifically, compared with the baseline softmax loss, our method improves the character retrieval mAP from 35.72% to 38.88% and the character verification accuracy from 87.00% to 88.50%. Yonggang Li 0001, Yafeng Zhou, Yongtao Wang, Xiaoran Qin, Zhi Tang 0001 |
ICPR | 5 |
| 2020 | MixTConv: Mixed Temporal Convolutional Kernels for Efficient Action RecognitionabstractTo efficiently extract spatiotemporal features of video for action recognition, most state-of-the-art methods integrate 1D temporal convolutional filters into 2D CNN backbones. However, they all exploit 1D temporal convolutional filters of fixed kernel size (i.e., 3) in their network building block, thus have suboptimal temporal modeling capability to handle both long-term and short-term actions. To address this problem, we first investigate the impacts of different kernel sizes for the 1D temporal convolutional filters. Then, we propose a simple yet efficient operation called Mixed Temporal Convolution (MixTConv), which consists of multiple depthwise 1D convolutional filters with different kernel sizes. By plugging MixTConv into the conventional 2D CNN backbone ResNet-50, we further propose an efficient and effective network architecture named MSTNet for action recognition, and achieve state-of-the-art results on multiple large-scale benchmarks. Kaiyu Shan, Yongtao Wang, Zhi Tang 0001, Yangyan Li |
ICPR | 3 |
| 2020 | GSTO: Gated Scale-Transfer Operation for Multi-Scale Feature Learning in Semantic SegmentationabstractExisting CNN-based methods for semantic segmentation heavily depend on multi-scale features to meet the requirements of both semantic comprehension and detail preservation. State-of-the-art segmentation networks widely exploit conventional scale-transfer operations, i.e., up-sampling and down-sampling to learn multi-scale features. In this work, we find that these operations lead to scale-confused features and suboptimal performance because they are spatial-invariant and directly transit all feature information cross scales without spatial selection. To address this issue, we propose the Gated Scale-Transfer Operation (GSTO) to properly transit spatial-filtered features to another scale. Specifically, GSTO can work either with or without extra supervision. Unsupervised GSTO is learned from the feature itself while the supervised one is guided by the supervised probability matrix. Both forms of GSTO are lightweight and plug-and-play, which can be flexibly integrated into networks or modules for learning better multi-scale features. In particular, by plugging GSTO into HRNet, we get a more powerful backbone (namely GSTO-HRNet) for pixel labeling, and it achieves new state-of-the-art results on multiple benchmarks for semantic segmentation including Cityscapes, LIP, and Pascal Context, with a negligible extra computational cost. Moreover, experiment results demonstrate that GSTO can also significantly boost the performance of multi-scale feature aggregation modules like PPM and ASPP. Zhuoying Wang, Yongtao Wang, Zhi Tang 0001, Yangyan Li, Haibin Ling, Weisi Lin |
ICPR | 3 |
| 2020 | ConvMath: A Convolutional Sequence Network for Mathematical Expression RecognitionabstractDespite the recent advances in optical character recognition (OCR), mathematical expressions still face a great challenge to recognize due to their two-dimensional graphical layout. In this paper, we propose a convolutional sequence modeling network, ConvMath, which converts the mathematical expression description in an image into a LaTeX sequence in an end-to-end way. The network combines an image encoder for feature extraction and a convolutional decoder for sequence generation. Compared with other Long Short Term Memory(LSTM) based encoder-decoder models, ConvMath is entirely based on convolution, thus it is easy to perform parallel computation. Besides, the network adopts multi-layer attention mechanism in the decoder, which allows the model to align output symbols with source feature vectors automatically, and alleviates the problem of lacking coverage while training the model. The performance of ConvMath is evaluated on an open dataset named IM2LATEX-100K, including 103556 samples. The experimental results demonstrate that the proposed network achieves state-of-the-art accuracy and much better efficiency than previous methods. Zuoyu Yan, Xiaode Zhang, Liangcai Gao, Zhi Tang 0001 |
ICPR | 5 |
| 2020 | Towards Accurate Panel Detection in Manga: A Combined Effort of CNN and Heuristics
Yafeng Zhou, Yongtao Wang, Zheqi He, Zhi Tang 0001, Ching Y. Suen |
MMM (1) | 4 |
| 2020 | A quadrilateral scene text detector with two-stage network architecture
Siwei Wang 0006, Zheqi He, Yongtao Wang, Zhi Tang 0001 |
Pattern Recognit. | 5 |
| 2020 | Scalable Triangle Discovery Algorithm for Large Scale-Free Network with Limited Internal MemoryabstractIn complex network, the links among the nodes fabricate thousands of polygons, where triangles are the most foundamental ones. As more links are added to the network, more polygons are split to be triangles. Discovering the triangles can help understand the current status of the network. Due to the limitations of internal memory, discovering all the triangles in a large-scale real-world network, e.g., online social network (OSN), is challenging and has intrigued researchers for years. Currently, parallel computation frameworks, e.g., MapReduce, have been employed to address this issue. The last reducer is required to load all the neighbors of each node into internal memory, which causes “the curse of the last reducer.” The scale-free characteristic in real-world networks further exacerbate the applicability of the parallel triangle discovery algorithms. The current efforts on this issue mainly focus on dividing a large network into several small subnetworks and later employ internal memory algorithms. In this paper, we address this issue from a different perspective by employing the power-law degree distribution in scale-free networks. We theoretically and empirically prove that one node in scale-free network has a limited number of neighbors with a higher degree. Employing this finding, an novel parallel triangle discovery algorithm called degree partition (DePart) is designed. Since only the portion of neighbors with higher degrees are loaded into internal memory in the reducers, DePart solves “the curse of the last reducer” problem and works efficiently in clusters with limited internal memory. We develop DePart using MapReduce framework and examine the performance of DePart on five large real-world networks. The empirical testing results demonstrate that DePart has a favorable tradeoff between internal memory and total disk space and outperforms the current state-of-the-art MapReduce-based triangle discovery algorithms. Xun Liang 0001, Zhi Tang 0001 |
IEEE Trans. Big Data | 3 |
| 2020 | Convex Edges in Social NetworksabstractConvex edges are located on convex surfaces of hidden geometry spaces, resulting in shorter geometrical distance than observed. Convex edges reveal the close relationships between two connected nodes and have a key role in many research fields and applications, including social, trust, and protein-protein interaction networks. We first quantitatively define convex edges using hidden geometry space defined by Ricci curvature and show these to be capable of predicting highly interacted edges in social networks. From empirical studies of social networks, we find that convex edges often occur on two connected nodes with similar popularity and allow for frenemy (or versatile node) identification, and that almost all nodes maintain an upper bound of convex edges. These findings present further evidence of the limited attention in an attention economy from a different perspective. Using three well-known synthetic networks, we show that convex edges can reflect the formation of a network from scatter (concave) to aggregation (convex). Taken together, our results show that convex edges are a novel mechanism to investigate social networks from different disciplines and offer insights into many network phenomena. Xun Liang 0001, Zhi Tang 0001 |
IEEE Trans. Comput. Soc. Syst. | 4 |
| 2019 | M2Det: A Single-Shot Object Detector Based on Multi-Level Feature Pyramid NetworkabstractFeature pyramids are widely exploited by both the state-of-the-art one-stage object detectors (e.g., DSSD, RetinaNet, RefineDet) and the two-stage object detectors (e.g., Mask RCNN, DetNet) to alleviate the problem arising from scale variation across object instances. Although these object detectors with feature pyramids achieve encouraging results, they have some limitations due to that they only simply construct the feature pyramid according to the inherent multiscale, pyramidal architecture of the backbones which are originally designed for object classification task. Newly, in this work, we present Multi-Level Feature Pyramid Network (MLFPN) to construct more effective feature pyramids for detecting objects of different scales. First, we fuse multi-level features (i.e. multiple layers) extracted by backbone as the base feature. Second, we feed the base feature into a block of alternating joint Thinned U-shape Modules and Feature Fusion Modules and exploit the decoder layers of each Ushape module as the features for detecting objects. Finally, we gather up the decoder layers with equivalent scales (sizes) to construct a feature pyramid for object detection, in which every feature map consists of the layers (features) from multiple levels. To evaluate the effectiveness of the proposed MLFPN, we design and train a powerful end-to-end one-stage object detector we call M2Det by integrating it into the architecture of SSD, and achieve better detection performance than state-of-the-art one-stage detectors. Specifically, on MSCOCO benchmark, M2Det achieves AP of 41.0 at speed of 11.8 FPS with single-scale inference strategy and AP of 44.2 with multi-scale inference strategy, which are the new stateof-the-art results among one-stage detectors. The code will be made available on https://github.com/qijiezhao/M2Det. Qijie Zhao, Yongtao Wang, Zhi Tang 0001, Haibin Ling |
AAAI | 4 |
| 2019 | Chemical Similarity Based on Map Edit DistanceabstractQuantifying the similarity of compounds is a particularly important problem, as it is able to give rise to many new possibilities for understanding the relationship between molecules. Here we develop a novel algorithm for computing the Molecular similarity. Experimental results on a broad range of real world graphs and tasks show that the proposed similarity measure is very effective. Xiaoqing Lyu, Zhi Tang 0001 |
BIBM | 3 |
| 2019 | GNDD: A Graph Neural Network-Based Method for Drug-Disease Association PredictionabstractPotential drug-disease association prediction is important to facilitate drug discovery. However, most of existing drug-disease association prediction approaches rely on assembling multiple drug (disease)-related biological information, which is usually not comprehensively available, and they always fail to explore the latent information in drug-disease network. To tackle these challenges, we propose a graph neural network-based method for drug-disease association prediction, dubbed GNDD, with capturing the complex information between drugs and diseases dispense with any side information. Specifically, GNDD introduces the idea of collaborative filtering in recommendation system to avoid the dependency on multi-data. Furthermore, an embedding propagation strategy is exploited to model the high-order relationships in drug-disease network. We conduct experiments on the Comparative Toxicogenomics Database, demonstrating the effectiveness of our method in drug-disease association prediction. Bei Wang 0004, Xiaoqing Lyu, Jingwei Qu, Zehua Pan, Zhi Tang 0001 |
BIBM | 6 |
| 2019 | Comparison of Molecule Graph Representation with Similarity ConsistencyabstractMany research tasks in the bioinformatics area, such as bioactivity prediction, de novo molecular design, and synthesis prediction, increasingly utilize graph-similarity-based analysis, which relies on selecting graph representation methods. However, most studies focus on the improvement of the final results of the tasks, and less attention has been paid on the differences of the distribution of the result of graph representation, (e.g. the embedded vectors) and how to evaluate them. In this paper, we propose a metric, mean degree of consistency (MDC), to evaluate the different approaches of graph representation with the help of retrieval results. To implement MDC, we introduce a serialized matching matrix and an optimization method based on partial sequence matching. We evaluate the efficiency and reliability of MDC with a series of experiments on graph-similarity-based retrievals for molecule graphs (MGs). Our result is helpful to the graph researchers for selecting graph representation methods. Bei Wang 0004, Xiaoqing Lyu, Zhi Tang 0001 |
BIBM | 3 |
| 2019 | A non-local based segmentation method for Pelvic MR ImagesabstractA number of people are infected with functional disorders of the pelvic floor such as pelvic organ prolapse and defecation dysfunction. For radiologists, to discern anatomical structures from tomographic images is an important way to find such diseases. However, it is a challenge to recognize the anatomical structures from pelvic MR (magnetic resonance) images, since the same pelvic organs usually are with various shapes and different organs sometimes are with similar appearance in MR images. To address this problem, we introduce a non-local module in the deep neural networks so that both the local features and the global features could be fused to segment organs. Moreover, for the very small quantities of the open pelvic MR images, we have built up a fine-grained MR image segmentation dataset annotated by professional medical researchers, and proposed a data augmentation strategy that is especially suitable for our dataset. In the primary experiments on the dataset, the proposed method in this paper is compared with state-of-the-art image segmentation methods and achieves better performance. Zuoyu Yan, Liangcai Gao, Zhi Tang 0001 |
BIBM | 3 |
| 2019 | Molecular Graph Generation with Deep Reinforced Multitask Network and Adversarial Imitation LearningabstractMolecular graph generation aims to design molecules with desired biochemical properties, which is promising in drug discovery. Existing methods typically combine deep reinforcement models with adversarial training. However, the reinforced rewards in molecule generation are delayed and sparse while adversarial training suffers mode collapse issue. Moreover, they optimize multiple properties with a linear combination manner, where the models get distracted by potentially conflicting objectives. To tackle the above challenges, we propose a Deep Reinforced framework with Adversarial Imitation and Multitask learning (DR-AIM). First, the reinforced agent generates discrete molecular graphs via deep Q-learning, where the trajectories of high rewards are cached as experts. Then policies are extracted directly from expert trajectories by adversarial imitation learning, in which the discriminator delivers behavior distribution signals to the agent as dense rewards. Second, we propose to realize multi-goal molecule generation as a multitask learning process, where different property optimizations are treated as different tasks to be trained jointly. Extensive experiments demonstrate the effectiveness of DR-AIM in molecule generation. Xiaoqing Lyu, Zhi Tang 0001, Zhenming Liu |
BIBM | 4 |
| 2019 | A YOLO-Based Table Detection MethodabstractDue to various table layouts and styles, table detection is always a difficult task in the field of document analysis. Inspired by the great progress of deep learning based methods on object detection, in this paper, we present a YOLO-based method for this task. Considering the large difference between document objects and natural objects, we introduce some adaptive adjustments to YOLOv3, including an anchor optimization strategy and two post processing methods. For anchor optimization, we use k-means clustering to find anchors which are more suitable for tables rather than natural objects and make it easier for our model to find exact positions of tables. In post-processing process, the extra whitespaces and noisy page objects (e.g. page headers, page footers) are removed from the predicted results, so that our model can get more accurate table margins and higher IoU scores. The proposed method is evaluated on two datasets from ICDAR 2013 Table Competition and ICDAR 2017 Page Object Detection (POD) Competition and achieves state-of-the-art performance. Yilun Huang 0001, Qinqin Yan, Liangcai Gao, Zhi Tang 0001 |
ICDAR | 7 |
| 2019 | A GAN-Based Feature Generator for Table DetectionabstractTable detection is of great significance for the documents analysis and recognition. Although many methods have been proposed and great progress have been made, it is still a great challenge to recognize the less-ruled tables due to the lack of table line features. In this paper, we propose a novel network to generate the layout features for table text to improve the performance of less-ruled table recognition. This feature generator model is similar to the Generative Adversarial Networks (GAN). We force the feature generator model to extract similar features for both ruling tables and less-ruled tables. It can be added into some common object detection and semantic segmentation models such as Mask R-CNN, U-Net. Extensive experiments are conducted on the dataset of ICDAR2017 Page Object Detection Competition dataset and a closed dataset full of the less-ruled tables and non-ruled tables. The primary experimental results show that the proposed GAN-based feature generator is very helpful for less-ruled table detection. Liangcai Gao, Zhi Tang 0001, Qinqin Yan, Yilun Huang 0001 |
ICDAR | 3 |
| 2019 | Scene Text Recognition via Gated Cascade AttentionabstractScene text recognition is very challenging due to the complex background, low resolution, perspective distortion and curved placement, etc. Most of the state-of-the-art methods adopt the attention-based encoder-decoder framework, and usually get suboptimal recognition performance for challenging text images due to the misalignment between attention region and target character region. In this paper, a novel module, named Gated Cascade Attention Module (GCAM), is proposed to increase the alignment precision of attention in a cascade way. Moreover, a channel and spatial attention module is introduced into the encoder to extract more discriminative features for text recognition. By assembling these two modules, a novel scene text recognizer is developed, and extensive experiments demonstrate it can achieve state-of-the-art results on multiple benchmarks of regular and irregular text images. Siwei Wang 0006, Yongtao Wang, Xiaoran Qin, Qijie Zhao, Zhi Tang 0001 |
ICME | 5 |
| 2019 | Regularizing Variational Autoencoders for Molecular Graph Generation
Xiaoqing Lyu, Keqi Hu, Zhi Tang 0001 |
ICONIP (5) | 5 |
| 2019 | TGG: Transferable Graph Generation for Zero-shot and Few-shot LearningabstractZero-shot and few-shot learning aim to improve generalization to unseen concepts, which are promising in many realistic scenarios. Due to the lack of data in unseen domain, relation modeling between seen and unseen domains is vital for knowledge transfer in these tasks. Most existing methods capture seen-unseen relationimplicitly via semantic embedding or feature generation, resulting in inadequate use of relation and some issues remain (e.g. domain shift). To tackle these challenges, we propose a Transferable Graph Generation (TGG ) approach, in which the relation is modeled and utilizedexplicitly via graph generation. Specifically, our proposed TGG contains two main components: (1) Graph generation for relation modeling. Anattention-based aggregate network and arelation kernel are proposed, which generate instance-level graph based on a class-level prototype graph and visual features. Proximity information aggregating is guided by a multi-head graph attention mechanism, where seen and unseen features synthesized by GAN are revised as node embeddings. The relation kernel further generates edges with GCN and graph kernel method, to capture instance-level topological structure while tackling data imbalance and noise. (2) Relation propagation for relation utilization. Adual relation propagation approach is proposed, where relations captured by the generated graph are separately propagated from the seen and unseen subgraphs. The two propagations learn from each other in a dual learning fashion, which performs as an adaptation way for mitigating domain shift. All components are jointly optimized with a meta-learning strategy, and our TGG acts as an end-to-end framework unifying conventional zero-shot, generalized zero-shot and few-shot learning. Extensive experiments demonstrate that it consistently surpasses existing methods of the above three fields by a significant margin. Xiaoqing Lyu, Zhi Tang 0001 |
ACM Multimedia | 3 |
| 2019 | Automatic Chinese Spelling Checking and Correction Based on Character-Based Pre-trained Contextual Representations
Haihua Xie, Aolin Li, Yabo Li, Zhiyou Chen, Xiaoqing Lyu, Zhi Tang 0001 |
NLPCC (2) | 7 |
| 2019 | Synthesizing data for text recognition with style transfer
Siwei Wang 0006, Yongtao Wang, Zhi Tang 0001 |
Multim. Tools Appl. | 4 |
| 2018 | Comprehensive Feature Enhancement Module for Single-Shot Object Detector
Qijie Zhao, Yongtao Wang, Zhi Tang 0001 |
ACCV (5) | 4 |
| 2018 | A method for improving the reliability of causal inference from large-scale data in biomedicine
Yitao Liu, Xiaoqing Lyu, Haihua Xie, Xiaotong Yan, Bei Wang 0004, Zhi Tang 0001 |
BIBM | 6 |
| 2018 | Understanding Markush Structures in Chemistry Documents With Deep Learning
Penghui Sun, Xiaoqing Lyu, Bei Wang 0004, Xiaohan Yi, Zhi Tang 0001 |
BIBM | 6 |
| 2018 | A Benchmark for Automatic Acral Melanoma Preliminary Screening
Liangcai Gao, Zhi Tang 0001, Menglong Ran, Zhimiao Lin |
BIBM | 3 |
| 2018 | Mathematics Content Understanding for Cyberlearning via Formula Evolution MapabstractAlthough the scientific digital library is growing at a rapid pace, scholars/students often find reading Science, Technology, Engineering, and Mathematics (STEM) literature daunting, especially for the math-content/formula. In this paper, we propose a novel problem, "mathematics content understanding", for cyberlearning and cyberreading. To address this problem, we create a Formula Evolution Map (FEM) offline and implement a novel online learning/reading environment, PDF Reader with Math-Assistant (PRMA), which incorporates innovative math-scaffolding methods. The proposed algorithm/system can auto-characterize student emerging math-information need while reading a paper and enable students to readily explore the formula evolution trajectory in FEM. Based on a math-information need, PRMA utilizes innovative joint embedding, formula evolution mining, and heterogeneous graph mining algorithms to recommend high quality Open Educational Resources (OERs), e.g., video, Wikipedia page, or slides, to help students better understand the math-content in the paper. Evaluation and exit surveys show that the PRMA system and the proposed formula understanding algorithm can effectively assist master and PhD students better understand the complex math-content in the class readings. Zhuoren Jiang, Liangcai Gao, Zheng Gao 0001, Zhi Tang 0001, Xiaozhong Liu 0001 |
CIKM | 5 |
| 2018 | A Sequence Labeling Based Approach for Character Segmentation of Historical DocumentsabstractAs an important prerequisite step of historical document image analysis, character segmentation is fundamental but challenging. In this paper, we propose a novel approach for the handwritten character segmentation of historical documents by treating it as a sequence labeling problem. In more detail, the proposed model first segments document image into lines, then each column in the line image is given a label to indicate it is a segmentation position or not. The segmentation labeling is achieved by a neural model, which combines a CNN for feature extraction, a LSTM for sequence modeling and a CRF for sequence labeling. The performance of our methods has been evaluated on a 300-page dataset including 96,479 characters. The experimental results demonstrate that the proposed methods achieve superior or highly competitive performance compared with other methods. Liangcai Gao, Xiaode Zhang, Zhi Tang 0001, Yaoxiong Huang |
DAS | 3 |
| 2018 | A Free-Sketch Recognition Method for Chemical Structural FormulaabstractChemical Structural Formula(CSF) recognition plays an important role in the molecular design and component retrieval. However, sketch-based CSF recognition remains an obstacle in current retrieval systems. This paper introduces a system for sketch CSF recognition on smart mobile devices. A dual-mode-based method is proposed to distinguish the gestures for character inputs and non-character inputs instead of the ordinary segmentation approaches. An attribute graph model is established to describe effectively all necessary information of a sketched CSF. Chemical knowledge is adopted to refine the candidates of structure relationship among elements. The experiments results demonstrate that the proposed method outperforms the existing methods for free-sketch CSFs on effectiveness and flexibility. Penghui Sun, Xiaoqing Lyu, Bei Wang 0004, Jingwei Qu, Zhi Tang 0001 |
DAS | 6 |
| 2018 | An End-to-End Quadrilateral Regression Network for Comic Panel ExtractionabstractComic panel extraction, i.e., decomposing a comic page image into panels, has become a fundamental technique for meeting many practical needs of mobile comic reading such as comic content adaptation and comic animating. Most of existing approaches are based on handcrafted low-level visual patterns and heuristics rules, thus having limited ability to deal with irregular comic panels. Only one existing method is based on deep learning and achieves better experimental results, but its architecture is redundant and its time efficiency is not good. To address these problems, we propose an end-to-end, two-stage quadrilateral regressing network architecture for comic panel detection, which inherits the architecture of Faster R-CNN. At the first stage, we propose a quadrilateral region proposal network for generating panel proposals, based on a newly proposed quadrilateral regression method. At the second stage, we classify the proposals and refine their shapes with the proposed quadrilateral regression method again. Extensive experimental results demonstrate that the proposed method significantly outperforms the existing comic panel detection methods on multiple datasets by F1-score and page accuracy. Zheqi He, Yafeng Zhou, Yongtao Wang, Siwei Wang 0006, Xiaoqing Lu, Zhi Tang 0001 |
ACM Multimedia | 6 |
| 2017 | SReN: Shape Regression Network for Comic Storyboard ExtractionabstractThe goal of storyboard extraction is to decompose the comic image into several storyboards(or frames), which is the fundamental step of comic image understanding and producing digital comic documents suitable for mobile reading. Most of existing approaches are based on hand crafted low-level visual patters like edge segments and line segments, which do not capture high-level vision. To overcome shortcomings of the existing approaches, we propose a novel architecture based on deep convolutional neural network, namely Shape Regression Network(SReN), to detect storyboards within comic images. Firstly, we use Fast R-CNN to generate rectangle bounding boxes as storyboard proposals. Then we train a deep neural network to predict quadrangles for these propos- als. Unlike existing object detection methods which only output rectangle bounding boxes, SReN can produce more precise quadrangle bounding boxes. Experimental results, evaluating on 7382 comic pages, demonstrate that SReN outperforms the state-of-the-art methods by more than 10% in terms of F1-score and page correction rate. Zheqi He, Yafeng Zhou, Yongtao Wang, Zhi Tang 0001 |
AAAI | 4 |
| 2017 | Citation Metadata Extraction via Deep Neural Network-based Segment Sequence LabelingabstractCitation metadata extraction plays an important role in academic information retrieval and knowledge management. Current works on this task generally use rule-based, template-based or learning-based approaches but these methods usually either rely on handcrafted features or are limited with domains. Recently, neural networks have shown strong ability in addressing sequence labeling tasks. Liangcai Gao, Zhuoren Jiang, Runtao Liu, Zhi Tang 0001 |
CIKM | 5 |
| 2017 | Mutual Enhancement for Detection of Multiple Logos in Sports VideosabstractDetecting logo frequency and duration in sports videos provides sponsors an effective way to evaluate their advertising efforts. However, general-purposed object detection methods cannot address all the challenges in sports videos. In this paper, we propose a mutual-enhanced approach that can improve the detection of a logo through the information obtained from other simultaneously occurred logos. In a Fast-RCNN-based framework, we first introduce a homogeneity-enhanced re-ranking method by analyzing the characteristics of homogeneous logos in each frame, including type repetition, color consistency, and mutual exclusion. Different from conventional enhance mechanism that improves the weak proposals with the dominant proposals, our mutual method can also enhance the relatively significant proposals with weak proposals. Mutual enhancement is also included in our frame propagation mechanism that improves logo detection by utilizing the continuity of logos across frames. We use a tennis video dataset and an associated logo collection for detection evaluation. Experiments show that the proposed method outperforms existing methods with a higher accuracy. Xiaoqing Lu, Chengcui Zhang, Yongtao Wang, Zhi Tang 0001 |
ICCV | 5 |
| 2017 | ICDAR2017 Competition on Page Object DetectionabstractThis paper presents the results of ICDAR2017 Competition on Page Object Detection (POD). POD is to detect page objects (tables, mathematical equations, graphics, figures, etc.) from document images. This competition makes use of a dataset consists of 2,000 document page images This dataset contains abundant page objects with various types and layouts. During the competition, we received 13 different teams' registrations and finally 8 of them submitted their results. All teams used deep learning as the basic method, then combined different traditional features or methods to improve the detection performance. The team NLPR-PAL achieved the averaged F1 of 0.898 and mAPs of 0.805 in the detection of all page objects under the IOU threshold 0.8. In this overview paper, we summarize the task design, dataset, results, and the approaches used by those teams of this competitions. Liangcai Gao, Xiaohan Yi, Zhuoren Jiang, Leipeng Hao, Zhi Tang 0001 |
ICDAR | 5 |
| 2017 | A Deep Learning-Based Formula Detection Method for PDF DocumentsabstractIn practice, PDF files may be generated by different tools and their character information quality could be different. As a result, the approaches to detecting formulae from PDF documents usually have much different performance on different PDF files. To address this problem, in this paper we combine and refine the Convolutional Neural Network (CNN) and Recurrent Neural Network (RNN) model to detect formulae according to both their character and vision features. Based on the characteristic of PDF documents, we propose a series of strategies to train and optimize deep networks, such as the implicit class down-sampling strategy which can reduce the unbalancedness between formulae and other page elements (e.g., text paragraphs, tables, figures, etc.). The region proposal method is also redesigned to generate moderate formula candidates through combining the bottom-up and top-down layout analysis. The experimental results show that the combination of CNN and RNN can increase the robustness of our proposed detection method. Furthermore, the proposed method outperforms the existing formula detection methods on both a ground-truth dataset and a larger self-built dataset, which would be released and available for research purposes. Liangcai Gao, Xiaohan Yi, Zhuoren Jiang, Zuoyu Yan, Zhi Tang 0001 |
ICDAR | 6 |
| 2017 | Geometric Object 3D Reconstruction from Single Line Drawing Image with Bottom-Up and Top-Down Classification and Sketch GenerationabstractSince geometric objects are usually illustrated as 2D line drawings with the loss of depth information, it is difficult for people to fully understand the geometric structure of them. Most of the existing 3D reconstruction methods require the input to be perfect sketch of the line drawing, therefore they are not suitable when only line drawing image is offered. To address this problem, we propose a robust reconstruction method that can directly conduct reconstruction from the input single line drawing image. Our main contribution is that we introduce a bottom-up and top-down scheme to the task of geometric object 3D reconstruction. First, for the sub-task of geometric object classification, we perform a bottom-up process of geometric features extraction together with a top-down process of additional feature searching to classify the geometric object. Second, for the sub-task of sketch completion, we complete the sketch extracted from the geometric object's contour based on the class of the geometric object, in a top-down manner. Extensive experimental results demonstrate that our approach can improve in both accuracy and efficiency compared with the existing methods. Yongtao Wang, Yafeng Zhou, Zheqi He, Zhi Tang 0001 |
ICDAR | 5 |
| 2017 | A Faster R-CNN Based Method for Comic Characters Face DetectionabstractFace detection of comic characters is a necessary step in most applications, such as comic character retrieval, automatic character classification and comic analysis. However, the existing methods were developed for simple cartoon images or small size comic datasets, and detection performance remains to be improved. In this paper, we propose a Faster R-CNN based method for face detection of comic characters. Our contribution is twofold. First, for the binary classification task of face detection, we empirically find that the sigmoid classifier shows a slightly better performance than the softmax classifier. Second, we build two comic datasets, JC2463 and AEC912, consisting of 3375 comic pages in total for characters face detection evaluation. Experimental results have demonstrated that the proposed method not only performs better than existing methods, but also works for comic images with different drawing styles. Xiaoran Qin, Yafeng Zhou, Zheqi He, Yongtao Wang, Zhi Tang 0001 |
ICDAR | 5 |
| 2017 | A Symbol Dominance Based Formulae Recognition Approach for PDF DocumentsabstractWith more and more scientific documents becoming available in PDF format, recognition of formulae in these PDF documents is of great significance. In this paper, we propose a symbol dominance based formulae recognition approach to recovering formulae structures by using the rich information extracted directly from PDF files. The hierarchical structure of formula is represented by relationship tree, and the tree is built recursively based on symbol dominance, which considers both the spatial layout of symbols and the typesetting conventions of mathematics. In addition, we propose a special character recognition method to identify the formula characters with multiple components or variable unicode. Repeatable and comparable experiments have been done over two large datasets, IM2LATEX-100K and PDFME-10K. Experimental results demonstrate that our method is more adaptive and practical for PDF documents compared with other two existing available formulae recognition systems, INFTY and WYGIWYS. Xiaode Zhang, Liangcai Gao, Runtao Liu, Zhuoren Jiang, Zhi Tang 0001 |
ICDAR | 6 |
| 2017 | Deep-learning-based object-level contour detection with CCG and CRF optimizationabstractContour detection is a fundamental problem in computer vision. However, there is still a considerable disparity between detection results and actual contours. To detect object-level contours on the basis of comprehensive analysis of potential edges, we present a deep-learning-based approach with a conditional random fields (CRF) model. We obtain the initial edgemap with a VGGNet-based model, and establish a contour correlation graph (CCG) to describe the potential edges and their relationships. Then a CRF model is adopted to fulfill the optimization on the basis of the analysis of object-level contour characteristics and to predict the validity of candidate contours. The experiment results on various datasets demonstrate that the proposed method outperforms the existing methods by a noticeable margin. Songping Fu, Xiaoqing Lu, Chengcui Zhang, Zhi Tang 0001 |
ICME | 5 |
| 2017 | Automatic Document Metadata Extraction Based on Deep Networks
Runtao Liu, Liangcai Gao, Zhuoren Jiang, Zhi Tang 0001 |
NLPCC | 5 |
| 2016 | Analysis of Stroke Intersection for Overlapping PGF ElementsabstractQuery-by-figure is an effective retrieval approach for educational documents. However, complex geometric diagrams in the field of mathematics education remain as obstacles in current retrieval systems. This study aims to explore a query method for plane geometric figures (PGFs) via sketched figures on smart mobile devices. We adopt an undirected graph model to describe PGFs and a divide-and-conquer strategy to analyze the relationships among strokes. Our main contribution is the detailed analysis of the stroke intersection that frequently occurs in PGFs. Numerous accurate elements obtained through overlapping analysis are then selected to construct strong descriptors for PGFs. Only the compressed query features, instead of a query figure, are transmitted to an image-based retrieval system located on a remote server, where the sketched PGF is finally recognized with low delay response. The experiments show that the proposed method achieves high efficiency and provides users with good interactive experience. Xiaoqing Lu, Jingwei Qu, Zhi Tang 0001 |
DAS | 4 |
| 2016 | A Table Detection Method for PDF Documents Based on Convolutional Neural NetworksabstractBecause of the better performance of deep learning on many computer vision tasks, researchers in the area of document analysis and recognition begin to adopt this technique into their work. In this paper, we propose a novel method for table detection in PDF documents based on convolutional neutral networks, one of the most popular deep learning models. In the proposed method, some table-like areas are selected first by some loose rules, and then the convolutional networks are built and refined to determine whether the selected areas are tables or not. Besides, the visual features of table areas are directly extracted and utilized through the convolutional networks, while the non-visual information (e.g. characters, rendering instructions) contained in original PDF documents is also taken into consideration to help achieve better recognition results. The primary experimental results show that the approach is effective in table detection. Leipeng Hao, Liangcai Gao, Xiaohan Yi, Zhi Tang 0001 |
DAS | 4 |
| 2016 | Improving PGF retrieval effectiveness with active learningabstractMultimedia education is playing a significant and increasing role for education purposes, thus leading to a large number of electronic documents. Plane geometry figures (PGFs), as important components of these documents, are regarded as very helpful information to most retrieval systems in the field of mathematics education. However, the burdensome work of annotation has become one of the chief obstacles to improve the efficiency of retrieval systems. In this paper, we introduce an active learning-based frame to select candidate instances for training the classifiers in retrieval systems, which are an emerging non-text-based information systems. In addition, an enhanced uncertainty measure and the selection of specific features of PGFs are proposed for our active learning algorithm. Comparative experiment results indicate that the proposed method effectively improves the performance of the PGF retrieval system and reduces the burdensome annotation workload. Jingwei Qu, Xiaoqing Lu, Songping Fu, Zhi Tang 0001 |
ICPR | 4 |
| 2016 | Context-aware Geometric Object Reconstruction for Mobile EducationabstractThe solid geometric objects in the educational geometric books are usually illustrated as 2D line drawings accompanied with description text. In this paper, we present a method to recover the geometric objects from 2D to 3D. Unlike the previous methods, we not only use the geometric information from the line drawing itself, but also the textual information extracted from its context. The essential of our method is a cost function to mix the two types of information, and we optimize the cost function to identify the geometric object and recover its 3D information. Our method can recover various types of solid geometric objects including straight-edge manifolds and curved objects such as cone, cylinder and sphere. We show that our method performs significantly better compared to the previous ones. Jinxin Zheng, Yongtao Wang, Zhi Tang 0001 |
ACM Multimedia | 3 |
| 2016 | An R-CNN Based Method to Localize Speech Balloons in Comics
Yongtao Wang, Xicheng Liu, Zhi Tang 0001 |
MMM (1) | 3 |
| 2016 | Community-based Cyberreading for Information UnderstandingabstractAlthough the content in scientific publications is increasingly challenging, it is necessary to investigate another important problem, that of scientific information understanding. For this proposed problem, we investigate novel methods to assist scholars (readers) to better understand scientific publications by enabling physical and virtual collaboration. For physical collaboration, an algorithm will group readers together based on their profiles and reading behavior, and will enable the cyberreading collaboration within a online reading group. For virtual collaboration, instead of pushing readers to communicate with others, we cluster readers based on their estimated information needs. For each cluster, a learning to rank model will be generated to recommend readers' communitized resources (i.e., videos, slides, and wikis) to help them understand the target publication. Zhuoren Jiang, Xiaozhong Liu 0001, Liangcai Gao, Zhi Tang 0001 |
SIGIR | 4 |
| 2016 | Comic storyboard extraction via edge segment analysis
Yongtao Wang, Yafeng Zhou, Zhi Tang 0001 |
Multim. Tools Appl. | 4 |
| 2016 | Recovering solid geometric object from single line drawing image
Jinxin Zheng, Yongtao Wang, Zhi Tang 0001 |
Multim. Tools Appl. | 3 |
| 2016 | Improving retrieval of plane geometry figure with learning to rank
Lu Liu 0018, Xiaoqing Lu, Yongtao Wang, Zhi Tang 0001 |
Pattern Recognit. Lett. | 5 |
| 2015 | A fast and robust ellipse detector based on top-down least-square fitting
Yongtao Wang, Zheqi He, Xicheng Liu, Zhi Tang 0001, Luyuan Li |
BMVC | 4 |
| 2015 | A Novel Method for Chinese Named Entity Recognition Based on Character Vector
Zhi Tang 0001, Xiao-Jun Huang, Jia-Le Ma |
CollaborateCom | 3 |
| 2015 | A clump splitting based method to localize speech balloons in comicsabstractComic books enjoy great popularity across the world and people tend to read comics on digital devices nowadays. Localizing speech balloons is a necessary step in the process of digitalizing comic books, not only because they are one of the basic elements constituting a comic page but also because they are helpful to the localization of other comic elements. However, only a few studies have been done in this direction. In this paper, we propose a clump splitting based speech balloon localization method which can detect both closed and unclosed speech balloons. The contour generated by our method is accurate at pixel level and our method works for comic books of different drawing styles. Experimental results have demonstrated that our speech balloon localization method achieves better performance than existing methods on the common evaluation measures of recall, precision and F1score. Xicheng Liu, Yongtao Wang, Zhi Tang 0001 |
ICDAR | 3 |
| 2015 | Overlapped-triangle analysis with hierarchical ranking of dominanceabstractPlane geometric figures (PGFs) are essential diagrams that regularly appear in mathematics documents. To understand PGFs profoundly and intuitively, it is necessary to decompose them into visual elements rather than traditional line segments. In this paper, we present a method for detection and analysis of overlapped triangles on the basis of dominance rank. Overlapped triangles lead to a large quantity of redundant sub-triangles or incident triangles. The most important and representative triangles must be detected and adopted to reduce redundancy. Diverse types of relationships among potential triangles present at least three obstacles: (1) how to precisely detect all potential triangles, (2) how to select predominant triangles, and (3) how to verify that the combination of these predominant triangles can reconstruct the original figure. Based on Gestalt theory, we propose a method for selecting predominant triangles, including selection of the main convex triangle and composition of interior auxiliary triangles. Experiments show that our algorithm can detect most of the potential triangles and can present reasonable decomposition solutions using only few predominant elements. Xiaoqing Lu, Lu Liu 0018, Zhi Tang 0001, Haibin Ling |
ICDAR | 3 |
| 2015 | Comic frame extraction via line segments combinationabstractAutomatic comic frame extraction is the core technique for comic content adaptation. Typical algorithms segment a comic page into a set of frames using connected components or division lines, but they cannot produce frames without blank margins and there is still room for improvement for complex comic images. We present a method that identifies frame polygons via connected component labeling and line segments combination. We analyze lines within each component and preliminarily judge the frame type, and then we optimize an energy-like score function constrained by several rules to choose frames. Experimental results indicate an increase of both accuracy and F-score compared with previous methods. Yongtao Wang, Yafeng Zhou, Zhi Tang 0001 |
ICDAR | 3 |
| 2015 | A User-Oriented Special Topic Generation System for Digital NewspaperabstractWith the coming of digital newspaper, user-oriented special topic generation becomes extremely urgent to satisfy the users’ requirements both functionally and emotionally. We propose an applicable automatic special topic generation system for digital newspapers based on users’ interests. Firstly, extract subject heading vector of the topic of interest by filtering out function words, localizing Latent Dirichlet Allocation (LDA) and training the LDA model. Secondly, remove semantically repetitive vector component by constructing a synonymy word map. Lastly, organize and refine the special topic according to the similarity between the candidate news and the topic, and the density of topic-related terms. The experimental results show that the system has both simple operation and high accuracy, and it is stable enough to be applied for user-oriented special topic generation in practical applications. Zhi Tang 0001, Jianbo Xu, Liangcai Gao |
NLPCC | 3 |
| 2015 | A tree conditional random field model for panel detection in comic images
Luyuan Li, Yongtao Wang, Ching Y. Suen, Zhi Tang 0001 |
Pattern Recognit. | 4 |
| 2014 | Plane Geometry Figure Retrieval with Bag of ShapesabstractDigital education is serving an increasingly important function in most educational institutions, thus resulting in the production of a large number of digital documents online for education purposes. However, convenient ways to retrieve mathematic geometry questions are lacking because current retrieval systems largely rely on keywords instead of geometry figure images. This study focuses on plane geometry figure (PGF) image retrieval with the aim of retrieving relevant geometry images that contain more structural information than a question text stem. To fully use geometrical properties, a Bag-of-shapes (BoS) method is proposed to build the feature descriptor of an image. The BoS method contains either basic geometric primitives or dual-primitive structures along with several specific geometrical features for shape description. Based on the BoS feature descriptor, we apply cosine similarity with group feature weight as vector similarity measure for ranking to achieve high efficiency. For a PGF image query, the retrieval results are provided in an appropriate ranking order, which has high visual similarity with respect to human perception. Retrieval experiments and evaluation results show the effectiveness and efficiency of the proposed BoS shape descriptor. Lu Liu 0018, Xiaoqing Lu, Jingwei Qu, Liangcai Gao, Zhi Tang 0001 |
Document Analysis Systems | 6 |
| 2014 | Ground-Truth and Performance Evaluation for Page Layout Analysis of Born-Digital DocumentsabstractIn this paper, a new dataset is proposed for page layout analysis of born-digital documents. By extracting uniformly the document contents, an XML based data format is designed in terms of raw data and structure data. Utilizing a self-developed ground-truthing tool, a public dataset is constructed from diverse styles of document resources. With consideration of physical segmentation and logical labeling, automatic performance evaluation methods are adjusted to cope with different scenarios. The applications of the proposed dataset have shown that it is suitable for evaluating various layout analysis tasks. Zhi Tang 0001, Canhui Xu 0002, Liangcai Gao |
Document Analysis Systems | 2 |
| 2014 | Logical Labeling of Fixed Layout PDF Documents Using Multiple ContextsabstractThe task of logical structure recovery is known to be of crucial importance, yet remains unsolved not only for image based document but also for born-digital document system. In this work, the modeling of contextual information based on 2D Conditional Random Fields is proposed to learn page structure for born-digital fixed-layout documents. Heuristic prior knowledge of Portable Document Format (PDF) content and layout are interpreted to construct neighborhood graphs and various pair wise clique templates for the modeling of multiple contexts. By integrating local and contextual observations obtained from PDF attributes, the ambiguities of semantic labels are better resolved. Experimental comparisons for six types of clique templates has demonstrated the benefits of contextual information in logical labeling of 16 finely defined categories. Zhi Tang 0001, Canhui Xu 0002, Yongtao Wang |
Document Analysis Systems | 2 |
| 2014 | Plane Geometry Figure Retrieval Based on Bilayer Geometric Attributed Graph MatchingabstractWith the development of computer-aided education and digital library, there have emerged large numbers of digital documents online for education purposes. However, it is far from convenient to retrieve mathematic geometry questions because current retrieval systems largely rely on keywords instead of geometry figure images. We focus on plane geometry figure (PGF) image retrieval aiming at retrieving relevant geometry images that hold more similar geometric attributes and structure properties than a question text stem. Motivated by Attribute Graph (AG), and aiming to catch more delicate local geometric attributes and overall structure layout, we propose a Bilayer Geometric Attributed Graph (Bilayer-GAG) matching method to retrieve the relevant PGF images. The root node of Bilayer-GAG catches the spatial relationships among its children - the graph elements of the second layer, the second layer contains curvilinear geometric primitives and linear nested AGs that consist of nodes and edges with geometric signatures. Then we calculate the overall matching cost in three perspectives and finally retrieve top-k relevant Bilayer-GAGs. For a PGF image query, the retrieval results are shown in an appropriate ranking order, which has high visual similarity with respect to human perception. Retrieval experiments results show the effectiveness and efficiency of the proposed Bilayer-GAG. Lu Liu 0018, Xiaoqing Lu, Songping Fu, Jingwei Qu, Liangcai Gao, Zhi Tang 0001 |
ICPR | 6 |
| 2014 | A Method of Density Analysis for Chinese Characters
Jingwei Qu, Xiaoqing Lu, Lu Liu 0018, Zhi Tang 0001, Yongtao Wang |
NLPCC | 4 |
| 2014 | A mathematics retrieval system for formulae in layout presentationsabstractThe semantics of mathematical formulae depend on their spatial structure, and they usually exist in layout presentations such as PDF, LaTeX, and Presentation MathML, which challenges previous text index and retrieval methods. This paper proposes an innovative mathematics retrieval system along with the novel algorithms, which enables efficient formula index and retrieval from both webpages and PDF documents. Unlike prior studies, which require users to manually input formula markup language as query, the new system enables users to "copy" formula queries directly from PDF documents. Furthermore, by using a novel indexing and matching model, the system is aimed at searching for similar mathematical formulae based on both textual and spatial similarities. A hierarchical generalization technique is proposed to generate sub-trees from the semi-operator tree of formulae and support substructure match and fuzzy match. Experiments based on massive Wikipedia and CiteSeer repositories show that the new system along with novel algorithms, comparing with two representative mathematics retrieval systems, provides more efficient mathematical formula index and retrieval, while simplifying user query input for PDF documents. Xiaoyan Lin, Liangcai Gao, Zhi Tang 0001, Yingnan Xiao, Xiaozhong Liu 0001 |
SIGIR | 4 |
| 2014 | Mathematical formula identification and performance evaluation in PDF documents
Xiaoyan Lin, Liangcai Gao, Zhi Tang 0001, Josef B. Baker, Volker Sorge |
Int. J. Document Anal. Recognit. | 3 |
| 2014 | Automatic comic page segmentation based on polygon detection
Luyuan Li, Yongtao Wang, Zhi Tang 0001, Liangcai Gao |
Multim. Tools Appl. | 3 |
| 2013 | Detection of Overlapped Quadrangles in Plane Geometric FiguresabstractDigital plane geometric figures (PGFs) are important resources of digital education, especially in mathematical pedagogy. The related applications, such as recognition and retrieval system on geometric figure images, have not been fully exploited. Special quadrangles, including rectangles, parallelograms, and trapezoids, are important components of PGFs, and extracting these special quadrangles is a prerequisite task. In this paper, we focus on the detection of overlapped quadrangles in PGFs using the proposed one-pass detection algorithm based on the geometric symmetrical property of involved shapes. We also introduce a method to refine the detected geometric primitives by significance analysis in order to optimize the shape descriptor. Experimental results show that the proposed method for special quadrangle detection is accurate and efficient, and is ready for further use in PGF retrieval systems. Xiaoqing Lu, Haibin Ling, Lu Liu 0018, Tianxiao Feng, Zhi Tang 0001 |
ICDAR | 6 |
| 2013 | Unsupervised Speech Text Localization in Comic ImagesabstractLocalizing speech texts in comic images is a crucial step for catering the growing needs of reading comics on mobile devices. For example, automatically reading speech texts while adding sound effects alongside can not only render comic contents vividly but also help visually impaired readers. Unlike conventional text localization methods, we present an effective unsupervised speech text localization method in this paper that is free of training data. The proposed method consists of two major stages: (1) based on the concurrence of characters, the first stage of our method is to generate some of the character strings (a row or column of characters that align horizontally or vertically) from the comic images while the fonts and gaps of the adjacent characters within the character string are also obtained, (2) in the second stage, the obtained fonts and gaps of adjacent characters are used to detect rest of the character strings within the comic image via Bayesian classifier. The proposed method is tested on a dataset consists of 1000 comic images from ten printed comic series and provide satisfactory results. Luyuan Li, Yongtao Wang, Zhi Tang 0001, Xiaoqing Lu, Liangcai Gao |
ICDAR | 3 |
| 2013 | A Text Line Detection Method for Mathematical Formula RecognitionabstractText line detection is a prerequisite procedure of mathematical formula recognition, however, many incorrectly segmented text lines are often produced due to the two-dimensional structures of mathematics when using existing segmentation methods such as Projection Profiles Cutting or white space analysis. In consequence, mathematical formula recognition is adversely affected by these incorrectly detected text lines, with errors propagating through further processes. Aimed at mathematical formula recognition, we propose a text line detection method to produce reliable line segmentation. Based on the results produced by PPC, a learning based merging strategy is presented to combine incorrectly split text lines. In the merging strategy, the features of layout and text for a text line and those between successive lines are utilised to detect the incorrectly split text lines. Experimental results show that the proposed approach obtains good performance in detecting text lines from mathematical documents. Furthermore, the error rate in mathematical formula identification is reduced significantly through adopting the proposed text line detection method. Xiaoyan Lin, Liangcai Gao, Zhi Tang 0001, Josef B. Baker, Mohamed A. Alkalai, Volker Sorge |
ICDAR | 3 |
| 2013 | Stroke-Based Character Segmentation of Low-Quality Images on Ancient Chinese TabletabstractAncient Chinese tablets are invaluable in terms of historical and aesthetic value. Automatic character segmentation of images from degraded tablets poses a challenging problem. Therefore, this paper proposes a new character segmentation method that utilizes an enhanced stroke filter and an energy propagation process based on local layout information. A ground-truth dataset was established to evaluate the accuracy of the algorithm adopted by the proposed segmentation method. Experimental results indicate that the proposed method can effectively extract characters from low-quality ancient Chinese tablet images. Xiaoqing Lu, Zhi Tang 0001, Liangcai Gao |
ICDAR | 2 |
| 2013 | Structure-Based Web Access Method for Ancient Chinese Characters
Xiaoqing Lu, Yingmin Tang, Zhi Tang 0001, Yujun Gao |
NLPCC | 3 |
| 2012 | Table Header Detection and ClassificationabstractIn digital libraries, a table, as a specific document component as well as a condensed way to present structured and relational data, contains rich information and often the only source of .that information. In order to explore, retrieve, and reuse that data, tables should be identified and the data extracted. Table recognition is an old field of research. However, due to the diversity of table styles, the results are still far from satisfactory, and not a single algorithm performs well on all different types of tables. In this paper, we randomly take samples from the CiteSeerX to investigate diverse table styles for automatic table extraction. We find that table headers are one of the main characteristics of complex table styles. We identify a set of features that can be used to segregate headers from tabular data and build a classifier to detect table headers. Our empirical evaluation on PDF documents shows that using a Random Forest classifier achieves an accuracy of 92%. Prasenjit Mitra 0001, Zhi Tang 0001, C. Lee Giles |
AAAI | 3 |
| 2012 | Dataset, Ground-Truth and Performance Metrics for Table Detection EvaluationabstractTable detection is an important task in the field of document analysis. It has been extensively studied since a couple of decades. Various kinds of document mediums are involved, from scanned images to web pages, from plain texts to PDF files. Numerous algorithms published bring up a challenging issue: how to evaluate algorithms in different context. Currently, most work on table detection conducts experiments on their in-house dataset. Even the few sources of online datasets are targeted at image documents only. Moreover, Precision and recall measurement are usual practice in order to account performance based on human evaluation. In this paper, we provide a dataset that is representative, large and most importantly, publicly available. The compatible format of the ground truth makes evaluation independent of document medium. We also propose a set of new measures, implement them, and open the source code. Finally, three existing table detection algorithms are evaluated to demonstrate the reliability of the dataset and metrics. Zhi Tang 0001, Ruiheng Qiu, Ying Liu 0006 |
Document Analysis Systems | 3 |
| 2012 | Performance Evaluation of Mathematical Formula IdentificationabstractThis paper presents a performance evaluation system for mathematical formula identification. First, a ground-truth dataset is constructed to facilitate the performance comparison of different mathematical formula identification algorithms. Statistics analysis of the dataset shows the diversities of the dataset to reflect the real-world documents. Second, a performance evaluation metric for mathematical formula identification is proposed, including the error type definitions and the scenario-adjustable scoring. The proposed metric enables in-depth analysis of mathematical formula identification systems in different scenarios. Finally, based on the proposed evaluation metric, a tool is developed to automatically evaluate mathematical formula identification results. It is worth noting that the ground-truth dataset and the evaluation tool are freely available for academic purpose. Xiaoyan Lin, Liangcai Gao, Zhi Tang 0001 |
Document Analysis Systems | 3 |
| 2012 | A graph-based method of newspaper article reconstruction
Liangcai Gao, Zhi Tang 0001, Xiaoyan Lin, Yongtao Wang |
ICPR | 2 |
| 2012 | Envelope extraction for composite shapes for shape retrieval
Jianguo Song, Xiaoqing Lu, Haibin Ling, Zhi Tang 0001 |
ICPR | 5 |
| 2011 | An extended rights expression model supporting dynamic digital rights managementabstractLicense often plays the role of an agreement between the user and the copyright owner. Rights expression language should provide a description of interactive mechanism and license negotiation to avoid the ambiguous transformation between REL and authorization protocols. By extending the ODRL to establish a digital rights expression model of appending, updating and aggregating license, the interactive license management mechanism can be described. In this paper, the supposed scheme can support fine-grained authorization, progressive or updating authorization and combination authorization to achieve dynamic rights management of digital content. Zhi Tang 0001, Yinyan Yu |
IAS | 2 |
| 2011 | A decentralized authorization scheme for DRM in P2P file-sharing systemsabstractTo enable secure content trade between peers with good scalability and efficiency in Peer-to-Peer (P2P) file-sharing systems, we propose a decentralized authorization scheme for Digital Rights Management (DRM). Based on proxy re-encryption mechanism, our scheme deploys authorization functions on semi-trusted authorization proxy peers who re-encrypt the ciphertext of content keys and issue licenses to content user peers. A trusted server ensures that content users are properly charged for authorization. Our scheme has no alteration to the structure of P2P network and can be implemented with a wide range of proxy re-encryption algorithms. Compared with authorization schemes in existing DRM solutions, our scheme relieves the server from intense authorization processing, and achieves reliable decentralized authorization with good scalability and efficiency. Qin Qiu, Zhi Tang 0001, Yinyan Yu |
CCNC | 2 |
| 2011 | An efficient key management scheme for segment-based document protectionabstractHierarchical cryptographic key management of access control can be modeled as a partially ordered set in which a high security class can derive its descendant encryption keys, but not vice versa. In this paper, we propose a practical key management scheme for our segment-based document which is a novel XML-based document format for web publishing, called CEBX. The proposed scheme is not only efficient on key generation and key derivation, but also secure against collusion attack, reverse attack and key modification attack. Excellent dynamics are also supported in our scheme and updates of nodes or edges are locally affected in the hierarchy. Nodes of the hierarchy can define passwords to protect the segments of CEBX document, supporting more flexible business strategies in a digital rights management system. Zhi Tang 0001, Yinyan Yu |
CCNC | 2 |
| 2011 | A Table Detection Method for Multipage PDF Documents via Visual Seperators and Tabular StructuresabstractTable detection is always an important task of document analysis and recognition. In this paper, we propose a novel and effective table detection method via visual separators and geometric content layout information, targeting at PDF documents. The visual separators refer to not only the graphic ruling lines but also the white spaces to handle tables with or without ruling lines. Furthermore, we detect page columns in order to assist table region delimitation in complex layout pages. Evaluations of our algorithm on an e-Book dataset and a scientific document dataset show competitive performance. It is noteworthy that the proposed method has been successfully incorporated into a commercial software package for large-scale Chinese e-Book production. Liangcai Gao, Ruiheng Qiu, Zhi Tang 0001 |
ICDAR | 6 |
| 2011 | Metadata Extraction System for Chinese BooksabstractExtracting metadata from academic papers has attracted much attention from researchers in past years. But how to extract metadata automatically from books is still seldom discussed. In this paper, we address this task on Chinese books and present a system to extract metadata from the title page of a book. This system consists of three components: metadata segmentation, metadata labeling, and post-processing. Different strategies are adopted in the system to identify different metadata types, and a variety of information sources, including geometric layout, linguistic, semantic content and header-footer, are used to accommodate the wide range of metadata layouts. Experimental results on real-world data have demonstrated the effectiveness of the proposed system. Liangcai Gao, Yingmin Tang, Zhi Tang 0001 |
ICDAR | 4 |
| 2011 | Mathematical Formula Identification in PDF DocumentsabstractRecognizing mathematical expressions in PDF documents is a new and important field in document analysis. It is quite different from extracting mathematical expressions in image-based documents. In this paper, we propose a novel method by combining rule-based and learning-based methods to detect both isolated and embedded mathematical expressions in PDF documents. Moreover, various features of formulas, including geometric layout, character and context content, are used to adapt to a wide range of formula types. Experimental results show satisfactory performance of the proposed method. Furthermore, the method has been successfully incorporated into a commercial software package for large-scale Chinese e-Book production. Xiaoyan Lin, Liangcai Gao, Zhi Tang 0001 |
ICDAR | 3 |
| 2010 | Copy protection for emailabstractEmail has been one of the most common medium for online communication. Directly proportional to this situation is the significant financial loss or privacy disclosure resulting from the unprotected emails. We need a secure email system to protect email wherever it goes. However, most existing secure email solutions can only achieve security during email transmission. After an encrypted email is received and downloaded on the recipient's computer, the email will lose protection. In this paper, we propose a copy protection method for email based on "client-to-client" architecture. The foundation of our method is a password based key negotiation protocol between two clients. Moreover, device-bound technology and commutative symmetric encryption are applied in our method to achieve persistent copy protection. Our method balances flexibility and security by allowing users to receive and use email on any of their devices. Qin Qiu, Zhi Tang 0001 |
ICEC | 3 |
| 2010 | MC-JBIG2: an improved algorithm for Chinese textual image compression
Kui Hu, Zhi Tang 0001, Liangcai Gao, Yadong Mu |
Int. J. Document Anal. Recognit. | 2 |
| 2009 | An Efficient Contents Sharing Method for DRMabstractThis paper presents a rights management method for the multiple devices of end-user. In recent years, most existing contents sharing solutions are based on Authorized Domain where legally acquired copyrighted content can seamlessly move from device to device. However, Authorized Domain systems are relatively complex and will usually increase the usage cost. In this paper, we propose an efficient contents sharing method for digital rights management system without authorized domain. In particular, we design a novel license acquisition and sharing mechanism based on ergodic encryption and machine authentication techniques. The mechanism works well in a loosely connected network environment. Our design supports efficient and flexible device alterations without expensive revocation mechanism. In addition, a novel pre-experiencing function is supported conveniently in our approach. Zhi Tang 0001, Yinyan Yu |
CCNC | 2 |
| 2009 | Analysis of Book Documents' Table of Content Based on ClusteringabstractTable of contents (TOC) recognition has attracted a great deal of attention in recent years. After reviewing the merits and drawbacks of the existing TOC recognition methods, we have observed that book documents are multi-page documents with intrinsic local format consistency. Based on this finding we introduce an automatic TOC analysis method through clustering. This method first detects the decorative elements in TOC pages. Then it learns a layout model used in the TOC pages through clustering. Finally, it generates TOC entries and extracts their hierarchical structure under the guidance of the model. More specifically, broken lines are taken into account in the method. Experimental results show that this method achieves high accuracy and efficiency. In addition, this method has been successfully applied in a commercial e-book production software package. Liangcai Gao, Zhi Tang 0001, Yimin Chu |
ICDAR | 2 |
| 2008 | Comprehensive Global Typography Extraction System for Electronic Book DocumentsabstractBook documents usually have consistent typographies throughout the whole book, including headers, footers, columns, text line directions, and fonts used in the each level of headings. Such document-level typography information is of great value for downstream document processing applications. This paper presents a document analysis system that can extract a comprehensive set of typographies used in book documents. The system consists of several components: recognition of fonts used in the body text and chapter headings; detection of page body area, headers and footers; detection of columns, text line direction and line spacing of body text. Page-association is employed in the system. The preliminary experimental results demonstrate the effectiveness of the system. Liangcai Gao, Zhi Tang 0001, Ruiheng Qiu |
Document Analysis Systems | 2 |
| 2008 | Improved quantization index modulation watermarking robust against amplitude scaling and constant change distortionsabstractThe original quantization index modulation watermarking is largely vulnerable to valumetric scaling and constant change attacks. To overcome these two drawbacks, the normalized dither modulation (NDM) is presented in this paper. The main idea of it is to construct a gain-invariant vector with zero mean for quantization. Performance analysis shows that NDM is theoretically invariant to valumetric scaling and constant change and achieves similar performance to dither modulation (DM) in other aspects. Some useful strategies are further provided to improve the performance of NDM. Experiments on real data (images) demonstrate that not only the proposed method is extremely robust to amplitude scaling and constant change, and outperform the original DM subject to several other typical attacks. Xinshan Zhu, Zhi Tang 0001 |
ICIP | 2 |
| 2008 | Improved quantization index modulation watermarking robust against amplitude scaling distortionsabstractThe original quantization index modulation (QIM) watermarking is largely vulnerable to valumetric scaling. To solve this problem, a gain-invariant vector for quantization is constructed through dividing the host signal by a statistical feature extracted from the target content. With this idea, the improved dither modulation (IM-DM) and spread transform dither modulation (IM-STDM) are developed in this paper. We present the strategies on the choice of the introduced feature and analyze the performance of two improved QIM schemes theoretically. It is shown that the improved QIM is theoretically invariant to valumetric scaling, but becomes sensitive to constant change. Experiments on real data (images) demonstrate that the proposed methods are extremely robust to amplitude scaling and possess the similar performance to the original QIM subject to several other typical attacks. Xinshan Zhu, Zhi Tang 0001 |
ICME | 2 |
| 2007 | The valuation of china venture capital guiding fund policy based on options modelabstractIn this paper, we analyze the China venture capital guiding fund policy based on options model. Options theory determines the present value of a future uncertainty interest. Following this principle, we propose the methods to valuate the policy with both Monte Carlo simulation and numerical analysis technique on Black-Scholes method. Our work is a remarkable step towards the quantitative analysis of public policies using options theory. Kui Hu, Zhi Tang 0001, Xun Liang 0001 |
SMC | 2 |
| 2006 | A Novel Multibit Watermarking Scheme Combining Spread Spectrum and Quantization
Xinshan Zhu, Zhi Tang 0001, Liesen Yang |
IWDW | 2 |