Tong Lu 0002

dblp:33/4058-2 · DBLP profile ↗
← Back
181ranked-venue papers
4as first author
88since 2021 · last 2026
0000-0002-7051-5347ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 119 · 2 first-author · 63 since 2021Graphics, computer vision, multimedia, augmented reality and games · 103 · 3 first-author · 45 since 2021Databases, data management, data science and information retrieval · 17 · 3 since 2021Computer networks · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
YearPublicationVenuePosition
2026 MMFuser: Multimodal Multi-layer Feature Fuser for Fine-Grained Vision-Language Understanding
Yangzhou Liu, Guangchen Shi, Yong Fa, Song Mei, Tong Lu 0002
ICPR (5)10
2026 Video Mamba Suite: State Space Model as a Versatile Alternative for Video Understanding
Guo Chen 0006, Yifei Huang 0002, Jilan Xu, Baoqi Pei, Jiahao Wang 0005, Zhe Chen 0017, Tong Lu 0002, Limin Wang 0002
Int. J. Comput. Vis.8
2026 Feature matters: Revisiting channel attention for Temporal Action Detection
Guo Chen 0006, Yin-Dong Zheng, Jiahao Wang 0005, Tong Lu 0002
Pattern Recognit.5
2026 Diffusion models with spatial control and attention fusion for incremental few-shot semantic segmentation
Guangchen Shi, Yirui Wu, Palaiahnakote Shivakumara, Shirong Zou, Tong Lu 0002
Pattern Recognit.7
2025 Docopilot: Improving Multimodal Models for Document-Level Understanding
abstract
Despite significant progress in multimodal large language models (MLLMs), their performance on complex, multi-page document comprehension remains inadequate, largely due to the lack of high-quality, document-level datasets. While current retrieval-augmented generation (RAG) methods offer partial solutions, they suffer from issues, such as fragmented retrieval contexts, multi-stage error accumulation, and extra time costs of retrieval. In this work, we present a high-quality document-level dataset, Doc-750K, designed to support in-depth understanding of multimodal documents. This dataset includes diverse document structures, extensive cross-page dependencies, and real question-answer pairs derived from the original documents. Building on the dataset, we develop a native multimodal model—Docopilot, which can accurately handle document-level dependencies without relying on RAG. Experiments demonstrate that Docopilot achieves superior coherence, accuracy, and efficiency in document understanding tasks and multi-turn interactions, setting a new baseline for document-level multimodal understanding. Data, code, and models are released at https://github.com/OpenGVLab/Docopilot.
Yuchen Duan, Zhe Chen 0017, Yusong Hu, Weiyun Wang, Shenglong Ye, Botian Shi, Lewei Lu, Qibin Hou, Tong Lu 0002, Hongsheng Li 0001, Jifeng Dai, Wenhai Wang
CVPR9
2025 CorrMoE: Mixture of Experts with De-Stylization Learning for Cross-Scene and Cross-Domain Correspondence Pruning
abstract
Establishing reliable correspondences between image pairs is a fundamental task in computer vision, underpinning applications such as 3D reconstruction and visual localization. Although recent methods have made progress in pruning outliers from dense correspondence sets, they often hypothesize consistent visual domains and overlook the challenges posed by diverse scene structures. In this paper, we propose CorrMoE, a novel correspondence pruning framework that enhances robustness under cross-domain and cross-scene variations. To address domain shift, we introduce a De-stylization Dual Branch, performing style mixing on both implicit and explicit graph features to mitigate the adverse influence of domain-specific representations. For scene diversity, we design a Bi-Fusion Mixture of Experts module that adaptively integrates multi-perspective features through linear-complexity attention and dynamic expert routing. Extensive experiments on benchmark datasets demonstrate that CorrMoE achieves superior accuracy and generalization compared to state-of-the-art methods. The code and pre-trained models are available at https://github.com/peiwenxia/CorrMoE.
Peiwen Xia, Tangfei Liao, Danhuai Zhao, Jianjun Ke, Kaihao Zhang, Tong Lu 0002, Tao Wang 0052
ECAI7
2025 LLFA: Fusing Global Illumination and Local Priors for Low-Light Face Image Enhancement with Adaptor
abstract
Low-light image enhancement problem has been widely studied. However, most existing methods do not perform well on low-light face images due to no specific facial characteristic considerations. We first create large-scale low-light face datasets with synthesized and real-world images to address the absence of suitable datasets. Our experiments show that existing LLIE and face restoration methods are limited in enhancing low-light face images. To overcome these challenges, we propose a novel framework, the Low-Light Face Adaptor (LLFA), featuring an auxiliary encoder and an adaptor module. The encoder captures global illumination information, while the adaptor module adaptively fuses this information with high-quality priors. We also introduce a joint learning strategy that optimizes the model by simultaneously learning face priors and the enhancement process. Comprehensive experiments demonstrate that LLFA significantly outperforms state-of-the-art methods.
Ziqian Shao, Tao Wang 0052, Kaihao Zhang, Danhuai Zhao, Tong Lu 0002
ICASSP5
2025 Conditional Convolutions for End-to-End Single-Stage Video Text Detection
abstract
We propose a simple yet effective single-stage video text detection framework, termed CVTD (Conditional convolutions for Video Text Detection), which, to the best of our knowledge, is the first end-to-end single-stage video text detection framework.Most existing video text detection methods adopt text tracking to enhance text detection performance, but treat text detection and tracking as two separate tasks. In contrast, we propose to solve text detection and tracking in an end-to-end, unified way. Instead of using an additional tracking module, we employ dynamic instance-aware conditional convolution (CondConv) to implicitly model the temporal variation of video text. Each CondConv represents a text instance over time and is responsible to predict corresponding text mask. It is propagated and updated frame-by-frame to perform text detection and tracking in a seamless manner.CVTD enjoys two advantages: (1) Text detection and tracking are integrated in a single-stage network, eliminating the need for additional text tracking process. (2) The CondConv transferred across frames can compactly encodes temporal context features of text, leading to enhanced text detection performance and real-time inference speed. Experiments on multiple video text benchmarks demonstrate the superiority of our method in terms of both accuracy and efficiency.
Xiaoge Song, Danhuai Zhao, Tong Lu 0002
ICASSP5
2025 MOERL: When Mixture-Of-Experts Meet Reinforcement Learning for Adverse Weather Image Restoration
Tao Wang 0052, Peiwen Xia, Peng-Tao Jiang, Zhe Kong, Kaihao Zhang, Tong Lu 0002, Wenhan Luo
ICCV7
2025 Personality Trait Prediction from Twitter Data Using Text and Image Features
Kunal Biswas, Palaiahnakote Shivakumara, Umapada Pal 0001, Daniel P. Lopresti, Tong Lu 0002
ICDAR (1)5
2025 CG-Bench: Clue-grounded Question Answering Benchmark for Long Video Understanding
abstract
The existing video understanding benchmarks for multimodal large language models (MLLMs) mainly focus on short videos. The few benchmarks for long video understanding often rely on multiple-choice questions (MCQs). Due to the limitations of MCQ evaluations and the advanced reasoning abilities of MLLMs, models can often answer correctly by combining short video insights with elimination, without truly understanding the content. To bridge this gap, we introduce CG-Bench, a benchmark for clue-grounded question answering in long videos. CG-Bench emphasizes the model's ability to retrieve relevant clues, enhancing evaluation credibility. It includes 1,219 manually curated videos organized into 14 primary, 171 secondary, and 638 tertiary categories, making it the largest benchmark for long video analysis. The dataset features 12,129 QA pairs in three question types: perception, reasoning, and hallucination. To address the limitations of MCQ-based evaluation, we develop two novel clue-based methods: clue-grounded white box and black box evaluations, assessing whether models generate answers based on accurate video understanding. We evaluated multiple closed-source and open-source MLLMs on CG-Bench. The results show that current models struggle significantly with long videos compared to short ones, and there is a notable gap between open-source and commercial models. We hope CG-Bench will drive the development of more reliable and capable MLLMs for long video comprehension.
Guo Chen 0006, Yifei Huang 0002, Baoqi Pei, Jilan Xu, Yuping He, Tong Lu 0002, Yali Wang 0001, Limin Wang 0002
ICLR7
2025 Vision-RWKV: Efficient and Scalable Visual Perception with RWKV-Like Architectures
abstract
Transformers have revolutionized computer vision and natural language processing, but their high computational complexity limits their application in high-resolution image processing and long-context analysis. This paper introduces Vision-RWKV (VRWKV), a model that builds upon the RWKV architecture from the NLP field with key modifications tailored specifically for vision tasks. Similar to the Vision Transformer (ViT), our model demonstrates robust global processing capabilities, efficiently handles sparse inputs like masked images, and can scale up to accommodate both large-scale parameters and extensive datasets. Its distinctive advantage is its reduced spatial aggregation complexity, enabling seamless processing of high-resolution images without the need for window operations. Our evaluations demonstrate that VRWKV surpasses ViT's performance in image classification and has significantly faster speeds and lower memory usage processing high-resolution inputs. In dense prediction tasks, it outperforms window-based models, maintaining comparable speeds. These results highlight VRWKV's potential as a more efficient alternative for visual perception tasks. Code and models are available at~\url{https://github.com/OpenGVLab/Vision-RWKV}.
Yuchen Duan, Weiyun Wang, Zhe Chen 0017, Xizhou Zhu, Lewei Lu, Tong Lu 0002, Yu Qiao 0001, Hongsheng Li 0001, Jifeng Dai, Wenhai Wang
ICLR6
2025 Egocentric Object-Interaction Anticipation with Retentive and Predictive Learning
abstract
Egocentric object-interaction anticipation is critical for applications like augmented reality and robotics, but existing methods struggle with misaligned egocentric encoding, insufficient supervision, and underutilized historical context. These limitations stem from a lack of focus on retention, i.e., retaining long-term object-centric interactions, and prediction, i.e., future-centric encoding and future uncertainty modeling. We introduce EgoAnticipator, a novel Retentive and Predictive Learning framework that addresses these challenges. Our approach combines retentive pre-training for domain-specific encoding, predictive pre-training for future uncertainty modeling, and mirror distillation to transfer future-informed knowledge. Additionally, we propose long-term memory prompting to integrate historical interaction cues. We evaluate the effectiveness of our framework using the Ego4D short-term object interaction anticipation benchmark, covering both STAv1 and STAv2. Extensive experiments demonstrate that our framework outperforms existing methods, while ablation studies highlight the effectiveness of each design inside our retentive and predictive learning framework.
Guo Chen 0006, Yifei Huang 0002, Yin-Dong Zheng, Jiahao Wang 0005, Tong Lu 0002
IJCAI6
2025 Eagle 2.5: Boosting Long-Context Post-Training for Frontier Vision-Language Models
abstract
We introduce Eagle2.5, a frontier vision-language model (VLM) for long-context multimodal learning. Our work addresses the challenges in long video comprehension and high-resolution image understanding, introducing a generalist framework for both tasks. The proposed training framework incorporates Automatic Degrade Sampling and Image Area Preservation, two techniques that preserve contextual integrity and visual details. The framework also includes numerous efficiency optimizations in the pipeline for long-context data training. Finally, we propose Eagle-Video-110K, a novel dataset that integrates both story-level and clip-level annotations, facilitating long-video understanding. Eagle2.5 demonstrates substantial improvements on long-context multimodal benchmarks, providing a robust solution to the limitations of existing VLMs. Notably, our best model Eagle2.5-8B achieves 72.4\% on Video-MME with 512 input frames, matching the results of top-tier commercial model such as GPT-4o and large-scale open-source models like Qwen2.5-VL-72B and InternVL2.5-78B.
Guo Chen 0006, Jindong Jiang, Lidong Lu, De-An Huang, Wonmin Byeon, Matthieu Le, Max Ehrlich, Tong Lu 0002, Limin Wang 0002, Bryan Catanzaro, Jan Kautz, Andrew Tao, Zhiding Yu, Guilin Liu
NeurIPS11
2025 EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs
abstract
Transferring and integrating knowledge across first-person (egocentric) and third-person (exocentric) viewpoints is intrinsic to human intelligence, enabling humans to learn from others and convey insights from their own experiences. Despite rapid progress in multimodal large language models (MLLMs), their ability to perform such cross-view reasoning remains unexplored. To address this, we introduce EgoExoBench, the first benchmark for egocentric exocentric video understanding and reasoning. Built from publicly available datasets, EgoExoBench comprises over 7300 question–answer pairs spanning eleven sub-tasks organized into three core challenges: semantic alignment, viewpoint association, and temporal reasoning. We evaluate 13 state-of-the-art MLLMs and find that while these models excel on single-view tasks, they struggle to align semantics across perspectives, accurately associate views, and infer temporal dynamics in the ego-exo context. We hope EgoExoBench can serve as a valuable resource for research on embodied agents and intelligent assistants seeking human-like cross-view intelligence.
Yuping He, Yifei Huang 0002, Guo Chen 0006, Baoqi Pei, Jilan Xu, Tong Lu 0002, Jiangmiao Pang
NeurIPS6
2025 Guiding Audio-Visual Question Answering with Collective Question Reasoning
abstract
Abstract Audio-Visual Question Answering (AVQA) requires the model to answer questions with complex dynamic audio-visual information. Prior works on this task mainly consider only using single question-answer pairs during training, overlooking the rich semantic associations between questions. In this work, we propose a novel Collective Question-Guided Network (CoQo), which accepts multiple question-answer pairs as input and leverages the reasoning over these questions to assist the model training process. The core module is the proposed Question Guided Transformer (QGT), which uses collective question reasoning to perform question-guided feature extraction. Since multiple question-answer pairs are not always available, especially during inference, our QGT uses a set of learnable tokens to learn the collective information from multiple questions during training. At inference time, these learnable tokens bring additional reasoning information even when only one question is used as input. We employ QGT in both spatial and temporal dimensions to extract question-related features effectively and efficiently. To better capture detailed audio-visual associations, we train the model in a finer level by distinguishing feature pairs of different questions within the same video. Extensive experiments demonstrate that our method can achieve state-of-the-art performance on three AVQA datasets while reducing training time significantly. We also observe strong performances of our method on three VQA benchmarks. Detailed ablation studies further confirm the effectiveness of our proposed collective question reasoning scheme, both quantitatively and qualitatively.
Baoqi Pei, Yifei Huang 0002, Guo Chen 0006, Jilan Xu, Yali Wang 0001, Limin Wang 0002, Tong Lu 0002, Yu Qiao 0001, Fei Wu 0001
Int. J. Comput. Vis.7
2025 Lightweight Hybrid Device Identification for IoT Applications
abstract
The rapid proliferation of Internet of Things (IoT) devices has increased the variety of devices and data traffic, making data management and analysis more complex. This complexity has raised the demand for efficient device identification methods to ensure the smooth operation of the network. Conventional identification methods rely on Machine Learning (ML) and Deep Learning (DL), which either suffer from unstable feature engineering or rely on large labeled datasets with confined representation. To overcome these shortcomings, generic hybrid representations of raw traffic are essential for precise device identification. Additionally, existing work mainly investigated device identification in clouds, incurring high network latency and computation costs. A few studies have identified IoT devices in edge, but such methods used simple neural networks, resulting in incomplete representation and redundant operations. Comprehensive representations typically require complex models, but the limited resources at the edge are insufficient to execute these models. Therefore, this paper proposes a lightweight hybrid device identification (LHDI) approach, which achieves efficient device identification in resource-constrained edge nodes. First, we adopt the unsupervised pre-training to enhance the characterization of network packets. Second, we devise LHDI by integrating bidirectional long short-term memory (Bi-LSTM) and Transformerbased blocks in a parallel configuration. Third, a pruning framework is introduced to automatically reduce Transformer parameters using structured sparsity methods without retraining. By reducing redundant neural network parameters, the proposed lightweight model facilitates effective device identification in edge, without losing representation capabilities. Experimental results demonstrate that our methods deliver high accuracy with low cost compared to others.
Wei Liu 0004, Tong Lu 0002, Chao Cai 0001, Menglan Hu, Kai Peng 0001, Zehui Xiong
IEEE Internet Things J.3
2025 BEVFormer: Learning Bird's-Eye-View Representation From LiDAR-Camera via Spatiotemporal Transformers
abstract
Multi-modality fusion strategy is currently the de-facto most competitive solution for 3D perception tasks. In this work, we present a new framework termed BEVFormer, which learns unified BEV representations from multi-modality data with spatiotemporal transformers to support multiple autonomous driving perception tasks. In a nutshell, BEVFormer exploits both spatial and temporal information by interacting with spatial and temporal space through predefined grid-shaped BEV queries. To aggregate spatial information, we design spatial cross-attention that each BEV query extracts the spatial features from both point cloud and camera input, thus completing multi-modality information fusion under BEV space. For temporal information, we propose temporal self-attention to fuse the history BEV information recurrently. By comparing with other fusion paradigms, we demonstrate that the fusion method proposed in this work is both succinct and effective. Our approach achieves the new state-of-the-art 74.1% in terms of NDS metric on the nuScenes test set. In addition, we extend BEVFormer to encompass a wide range of autonomous driving tasks, including object tracking, vectorized mapping, occupancy prediction, and end-to-end autonomous driving, achieving outstanding results across these tasks. The code is released at https://github.com/fundamentalvision/BEVFormer.
Wenhai Wang, Hongyang Li 0001, Enze Xie, Chonghao Sima, Tong Lu 0002, Yu Qiao 0001, Jifeng Dai
IEEE Trans. Pattern Anal. Mach. Intell.6
2025 LLDiffusion: Learning degradation representations in diffusion models for low-light image enhancement
Tao Wang 0052, Kaihao Zhang, Yong Zhang 0034, Wenhan Luo, Björn Stenger, Tong Lu 0002, Tae-Kyun Kim 0001, Wei Liu 0005
Pattern Recognit.6
2024 Intern VL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks
abstract
The exponential growth of large language models (LLMs) has opened up numerous possibilities for multi-modal AGI systems. However, the progress in vision and vision-language foundation models, which are also critical elements of multi-modal AGI, has not kept pace with LLMs. In this work, we design a large-scale vision-language foun-dation model (Intern VL), which scales up the vision foun-dation model to 6 billion parameters and progressively aligns it with the LLM, using web-scale image-text data from various sources. This model can be broadly applied to and achieve state-of-the-art performance on 32 generic visual-linguistic benchmarks including visual perception tasks such as image-level or pixel-level recognition, vision-language tasks such as zero-shot image/video classification, zero-shot image/video-text retrieval, and link with LLMs to create multi-modal dialogue systems. It has powerful visual capabilities and can be a good alternative to the ViT-22B. We hope that our research could contribute to the development of multi-modal large models.
Zhe Chen 0017, Jiannan Wu, Wenhai Wang, Weijie Su 0002, Guo Chen 0006, Sen Xing, Muyan Zhong, Xizhou Zhu, Lewei Lu, Bin Li 0025, Ping Luo 0002, Tong Lu 0002, Yu Qiao 0001, Jifeng Dai
CVPR13
2024 Is Ego Status All You Need for Open-Loop End-to-End Autonomous Driving?
abstract
End-to-end autonomous driving recently emerged as a promising research direction to target autonomy from a full-stack perspective. Along this line, many of the latest works follow an open-loop evaluation setting on nuScenes to study the planning behavior. In this paper, we delve deeper into the problem by conducting thorough analyses and demystifying more devils in the details. We initially observed that the nuScenes dataset, characterized by relatively simple driving scenarios, leads to an under-utilization of perception information in end-to-end models incorporating ego status, such as the ego vehicle's velocity. These models tend to rely predominantly on the ego vehicle's status for future path planning. Beyond the limitations of the dataset, we also note that current metrics do not comprehensively assess the planning quality, leading to potentially biased conclusions drawn from existing benchmarks. To address this issue, we introduce a new metric to evaluate whether the predicted trajectories adhere to the road. We further propose a simple baseline able to achieve competitive results without relying on perception annotations. Given the current limitations on the benchmark and metrics, we suggest the community reassess relevant prevailing research and be cautious about whether the continued pursuit of state-of-the-art would yield convincing and universal conclusions. Code and models are available at https://github.com/NVlabs/BEV-Planner.
Zhiding Yu, Shiyi Lan, Jiahan Li, Jan Kautz, Tong Lu 0002, José M. Álvarez 0004
CVPR6
2024 Efficient Deformable ConvNets: Rethinking Dynamic and Sparse Operator for Vision Applications
abstract
We introduce Deformable Convolution v4 (DCNv4), a highly efficient and effective operator designed for a broad spectrum of vision applications. DCNv4 addresses the limitations of its predecessor, DCNv3, with two key enhancements: 1. removing softmax normalization in spatial aggregation to enhance its dynamic property and expressive power and 2. optimizing memory access to minimize redundant operations for speedup. These improvements result in a significantly faster convergence compared to DCNv3 and a substantial increase in processing speed, with DCNv4 achieving more than three times the forward speed. DCNv4 demonstrates exceptional performance across various tasks, including image classification, instance and semantic segmentation, and notably, image generation. When integrated into generative models like U-Net in the latent diffusion model, DCNv4 outperforms its baseline, underscoring its possibility to enhance generative models. In practical applications, replacing DCNv3 with DCNv4 in the InternImage model to create FlashInternImage results in up to 80% speed increase and further performance improvement without further modifications. The advancements in speed and efficiency of DCNv4, combined with its robust performance across diverse vision tasks, show its potential as a foundational building block for future vision models.
Yuwen Xiong, Yuntao Chen, Feng Wang 0015, Xizhou Zhu, Jiapeng Luo, Wenhai Wang, Tong Lu 0002, Hongsheng Li 0001, Yu Qiao 0001, Lewei Lu, Jie Zhou 0001, Jifeng Dai
CVPR8
2024 CorrAdaptor: Adaptive Local Context Learning for Correspondence Pruning
abstract
In the fields of computer vision and robotics, accurate pixel-level correspondences are essential for enabling advanced tasks such as structure-from-motion and simultaneous localization and mapping. Recent correspondence pruning methods usually focus on learning local consistency through k-nearest neighbors, which makes it difficult to capture robust context for each correspondence. We propose CorrAdaptor, a novel architecture that introduces a dual-branch structure capable of adaptively adjusting local contexts through both explicit and implicit local graph learning. Specifically, the explicit branch uses KNN-based graphs tailored for initial neighborhood identification, while the implicit branch leverages a learnable matrix to softly assign neighbors and adaptively expand the local context scope, significantly enhancing the model’s robustness and adaptability to complex image variations. Moreover, we design a motion injection module to integrate motion consistency into the network to suppress the impact of outliers and refine local context learning, resulting in substantial performance improvements. The experimental results on extensive correspondence-based tasks indicate that our CorrAdaptor achieves state-of-the-art performance both qualitatively and quantitatively.
Yuping He, Tangfei Liao, Xiaoqiu Xu, Tao Wang 0052, Tong Lu 0002
ECAI8
2024 The All-Seeing Project: Towards Panoptic Visual Recognition and Understanding of the Open World
abstract
We present the All-Seeing (AS) project: a large-scale dataset and model for recognizing and understanding everything in the open world. Using a scalable data engine that incorporates human feedback and efficient models in the loop, we create a new dataset (AS-1B) with over 1.2 billion regions annotated with semantic tags, question-answering pairs, and detailed captions. It covers a wide range of 3.5 million common and rare concepts in the real world and has 132.2 billion tokens that describe the concepts and their attributes. Leveraging this new dataset, we develop the All-Seeing model (ASM), a unified framework for panoptic visual recognition and understanding. The model is trained with open-ended language prompts and locations, which allows it to generalize to various vision and language tasks with remarkable zero-shot performance, including both region- and image-level retrieval, region recognition, captioning, and question-answering. We hope that this project can serve as a foundation for vision-language artificial general intelligence research. Code is available at https://github.com/OpenGVLab/all-seeing.
Weiyun Wang, Min Shi 0004, Qingyun Li, Wenhai Wang, Zhenhang Huang, Linjie Xing, Zhe Chen 0017, Hao Li 0069, Xizhou Zhu, Zhiguo Cao 0001, Tong Lu 0002, Jifeng Dai, Yu Qiao 0001
ICLR12
2024 Few-shot Semantic Segmentation via Perceptual Attention and Spatial Control
abstract
Few-shot semantic segmentation (FSS) aims to locate pixels of unseen classes with clues from a few labeled samples. Recently, thanks to profound prior knowledge, diffusion models have been expanded to achieve FSS tasks. However, due to probabilistic noising and denoising processes, it is difficult for them to maintain spatial relationships between inputs and outputs, leading to inaccurate segmentation masks. To address this issue, we propose a Diffusion-based Segmentation network (DiffSeg), which decouples probabilistic denoising and segmentation processes. Specifically, DiffSeg leverages attention maps extracted from a pretrained diffusion model as support-query interaction information to guide segmentation, which mitigates the impact of probabilistic processes while benefiting from rich prior knowledge of diffusion models. In the segmentation stage, we present a Perceptual Attention Module (PAM), where two cross-attention mechanisms capture semantic information of support-query interaction and spatial information produced by the pretrained diffusion model. Furthermore, a self-attention mechanism within PAM ensures a balanced dependence for segmentation, thus preventing inconsistencies between the aforementioned semantic and spatial information. Additionally, considering the uncertainty inherent in the generation process of diffusion models, we equip DiffSeg with a Spatial Control Module (SCM), which models spatial structural information of query images to control boundaries of attention maps, thus aligning the spatial location between knowledge representation and query images. Experiments on PASCAL-5i and COCO datasets show that DiffSeg achieves new state-of-the-art performance with remarkable advantages.
Guangchen Shi, Yirui Wu, Danhuai Zhao, Tong Lu 0002
ACM Multimedia6
2024 VisionLLM v2: An End-to-End Generalist Multimodal Large Language Model for Hundreds of Vision-Language Tasks
abstract
We present VisionLLM v2, an end-to-end generalist multimodal large model (MLLM) that unifies visual perception, understanding, and generation within a single framework. Unlike traditional MLLMs limited to text output, VisionLLM v2 significantly broadens its application scope. It excels not only in conventional visual question answering (VQA) but also in open-ended, cross-domain vision tasks such as object localization, pose estimation, and image generation and editing. To this end, we propose a new information transmission mechanism termed ``super link'', as a medium to connect MLLM with task-specific decoders. It not only allows flexible transmission of task information and gradient feedback between the MLLM and multiple downstream decoders but also effectively resolves training conflicts in multi-tasking scenarios. In addition, to support the diverse range of tasks, we carefully collected and combed training data from hundreds of public vision and vision-language tasks. In this way, our model can be joint-trained end-to-end on hundreds of vision language tasks and generalize to these tasks using a set of shared parameters through different user prompts, achieving performance comparable to task-specific models. We believe VisionLLM v2 will offer a new perspective on the generalization of MLLMs.
Jiannan Wu, Muyan Zhong, Sen Xing, Zeqiang Lai, Zhaoyang Liu 0001, Zhe Chen 0017, Wenhai Wang, Xizhou Zhu, Lewei Lu, Tong Lu 0002, Ping Luo 0002, Yu Qiao 0001, Jifeng Dai
NeurIPS10
2024 How far are we to GPT-4V? Closing the gap to commercial multimodal models with open-source suites
Zhe Chen 0017, Weiyun Wang, Hao Tian 0006, Shenglong Ye, Zhangwei Gao, Erfei Cui, Wenwen Tong, Kongzhi Hu, Jiapeng Luo, Zheng Ma 0012, Jiaqi Wang 0003, Xiaoyi Dong, Hang Yan 0001, Hewei Guo, Conghui He, Botian Shi, Zhenjiang Jin, Bin Wang 0065, Xingjian Wei, Wei Li 0320, Wenjian Zhang, Bo Zhang 0069, Pinlong Cai, Licheng Wen, Xiangchao Yan, Min Dou, Lewei Lu, Xizhou Zhu, Tong Lu 0002, Dahua Lin, Yu Qiao 0001, Jifeng Dai, Wenhai Wang
Sci. China Inf. Sci.31
2024 MMInstruct: a high-quality multi-modal instruction tuning dataset with extensive diversity
Yangzhou Liu, Zhangwei Gao, Weiyun Wang, Zhe Chen 0017, Wenhai Wang, Hao Tian 0006, Lewei Lu, Xizhou Zhu, Tong Lu 0002, Yu Qiao 0001, Jifeng Dai
Sci. China Inf. Sci.10
2024 A new deep CNN for 3D text localization in the wild through shadow removal
Palaiahnakote Shivakumara, Ayan Banerjee 0002, Lokesh Nandanwar, Umapada Pal 0001, Apostolos Antonacopoulos, Tong Lu 0002, Michael Blumenstein
Comput. Vis. Image Underst.6
2024 A robust script independent handwriting system for gender identification
abstract
Gender identification at the word level in a multi-script environment is challenging due to variations posed by free-style handwriting of individuals and geographical differences in writing styles. This paper presents a new approach, Multi-Orientation-Scale Gabor Response Fusion (MOSGF), for gender identification at the word level using handwritten text. Our method has two steps: (i) word segmentation from unconstrained lines and (ii) gender identification at the word level. In the first step, the method explores the number of zero crossing points and gradient information for word segmentation from handwritten text lines. In the second step, employs Gabor responses at different orientations and scales to detect fine details in female and male handwriting. For each Gabor response, the proposed model estimates the correlation between average templates of all Gabor responses and the individual Gabor response to extract global consistency in writing. To strengthen correlation features, the proposed method uses the Mahalanobis distance measure, which extracts local similarity. Further, the proposed approach fuses correlation coefficient and distance-based features in a novel way. The fused features are then fed to a Neural Network (NN) for gender identification. Experiments on our dataset, which comprises Roman (English), Chinese, Farsi (Persian), Arabic, and Indian scripts, and a benchmark dataset, namely, IAM which includes English text, KHATT which includes Arabic, and QUWI which includes both English and Arabic, show that the proposed system outperforms the existing methods in terms of word segmentation and gender identification.
Palaiahnakote Shivakumara, Maryam Asadzadeh Kaljahi, Swati Kanchan, Umapada Pal 0001, Daniel P. Lopresti, Tong Lu 0002
Expert Syst. Appl.6
2024 GridFormer: Residual Dense Transformer with Grid Structure for Image Restoration in Adverse Weather Conditions
Tao Wang 0052, Kaihao Zhang, Ziqian Shao, Wenhan Luo, Björn Stenger, Tong Lu 0002, Tae-Kyun Kim 0001, Wei Liu 0005, Hongdong Li
Int. J. Comput. Vis.6
2024 Restoring vision in hazy weather with hierarchical contrastive learning
Tao Wang 0052, Guangpin Tao, Wanglong Lu, Kaihao Zhang, Wenhan Luo, Xiaoqin Zhang 0002, Tong Lu 0002
Pattern Recognit.7
2024 Revisiting of AlphaStar
abstract
Research onStarCraftII (SC2) is considered important due to its similarity to real-life tasks and its potential to inspire game artificial intelligence design. However, the complexity of SC2 presents considerable challenges. In 2019, DeepMind proposed AlphaStar (AS), an agent that achieved Grandmaster level in SC2. Nevertheless, the reasons for AS's success remain unclear. In this article, we revisit AS by analyzing its technical details, implementation codes, and replays. We also propose the open-sourced mini-scaled AS's new versions to do ablation studies. We classify SC2 problems by difficulty level and suggest a research path for tackling them. We then identify several limitations of AS, such as its lack of strategic view, reasoning, scouting, changes in tactics, and planning. Our article also presents the first analysis of AS's replays. In conclusion, we emphasize that there is still a long way to solve the final SC2 problem.
Ruo-Ze Liu, Yanjie Shen, Yang Yu 0001, Tong Lu 0002
IEEE Trans. Games4
2024 Feature Selection Based on Intrusive Outliers Rather Than All Instances
abstract
Feature selection (FS) has recently attracted considerable attention in many fields. Highly-overlapping classes and skewed distributions of data within classes have been found in various classification tasks. Most existing FS methods are all instance-based, which ignores the significant differences in characteristics between the particular outliers and the main body of the class, causing confusion for classifiers. In this paper, we propose a novel supervised FS method, Intrusive Outliers-based Feature Selection (IOFS), to find out what kind of outliers lead to misclassification and exploit the characteristics of such outliers. In order to accurately identify the intrusive outliers (IOs), we provide a density-mean center algorithm to obtain the appropriate representative of a class. A special distance threshold is given to obtain the candidate for IOs. Combining with several metrics, mathematical formulations are provided to evaluate the overlapping degree of the intrusive class pairs. Features with high overlapping degrees are assigned to low rankings in IOFS method. An extension of IOFS based on a small number of extreme IOs, called E-IOFS, is also proposed. Three theoretical proofs are provided for the essential theoretical basis of IOFS. Experiments comparing against various state-of-the-art methods on eleven benchmark datasets show that IOFS is rational and effective, especially on the datasets with higher overlapping classes. And E-IOFS almost always outperforms IOFS.
Lixin Yuan, Cheng Mei, Wenhai Wang, Tong Lu 0002
IEEE Trans. Image Process.4
2024 TTS: Hilbert Transform-Based Generative Adversarial Network for Tattoo and Scene Text Spotting
abstract
Text spotting in natural scenes is of increasing interest and significance due to its critical role in several applications, such as visual question answering, named entity recognition and event rumor detection on social media. One of the newly emerging challenging problems is Tattoo Text Spotting (TTS) in images for assisting forensic teams and for person identification. Unlike the generally simpler scene text addressed by current state-of-the-art methods, tattoo text is typically characterized by the presence of decorative backgrounds, calligraphic handwriting and several distortions due to the deformable nature of the skin. This paper describes the first approach to address TTS in a real-world application context by designing an end-to-end text spotting method employing a Hilbert transform-based Generative Adversarial Network (GAN). To reduce the complexity of the TTS task, the proposed approach first detects fine details in the image using the Hilbert transform and the Optimum Phase Congruency (OPC). To overcome the challenges of only having a relatively small number of training samples, a GAN is then used for generating suitable text samples and descriptors for text spotting (i.e., both detection and recognition). The superior performance of the proposed TTS approach, for both tattoo and general scene text, over the state-of-the-art methods is demonstrated on a new TTS-specific dataset (publicly available) as well as on the existing benchmark natural scene text datasets: Total-Text, CTW1500 and ICDAR 2015.
Ayan Banerjee 0002, Palaiahnakote Shivakumara, Umapada Pal 0001, Apostolos Antonacopoulos, Tong Lu 0002, Josep Lladós 0001
IEEE Trans. Multim.5
2024 A Conformable Moments-Based Deep Learning System for Forged Handwriting Detection
abstract
Detecting forged handwriting is important in a wide variety of machine learning applications, and it is challenging when the input images are degraded with noise and blur. This article presents a new model based on conformable moments (CMs) and deep ensemble neural networks (DENNs) for forged handwriting detection in noisy and blurry environments. Since CMs involve fractional calculus with the ability to model nonlinearities and geometrical moments as well as preserving spatial relationships between pixels, fine details in images are preserved. This motivates us to introduce a DENN classifier, which integrates stenographic kernels and spatial features to classify input images as normal (original, clean images), altered (handwriting changed through copy-paste and insertion operations), noisy (added noise to original image), blurred (added blur to original image), altered-noise (noise is added to the altered image), and altered-blurred (blur is added to the altered image). To evaluate our model, we use a newly introduced dataset, which comprises handwritten words altered at the character level, as well as several standard datasets, namely ACPR 2019, ICPR 2018-FDC, and the IMEI dataset. The first two of these datasets include handwriting samples that are altered at the character and word levels, and the third dataset comprises forged International Mobile Equipment Identity (IMEI) numbers. Experimental results demonstrate that the proposed method outperforms the existing methods in terms of classification rate.
Lokesh Nandanwar, Palaiahnakote Shivakumara, Hamid Abdullah Jalab, Rabha W. Ibrahim, Ramachandra Raghavendra, Umapada Pal 0001, Tong Lu 0002, Michael Blumenstein
IEEE Trans. Neural Networks Learn. Syst.7
2023 Ultra-High-Definition Low-Light Image Enhancement: A Benchmark and Transformer-Based Method
abstract
As the quality of optical sensors improves, there is a need for processing large-scale images. In particular, the ability of devices to capture ultra-high definition (UHD) images and video places new demands on the image processing pipeline. In this paper, we consider the task of low-light image enhancement (LLIE) and introduce a large-scale database consisting of images at 4K and 8K resolution. We conduct systematic benchmarking studies and provide a comparison of current LLIE algorithms. As a second contribution, we introduce LLFormer, a transformer-based low-light enhancement method. The core components of LLFormer are the axis-based multi-head self-attention and cross-layer attention fusion block, which significantly reduces the linear complexity. Extensive experiments on the new dataset and existing public datasets show that LLFormer outperforms state-of-the-art methods. We also show that employing existing LLIE methods trained on our benchmark as a pre-processing step significantly improves the performance of downstream tasks, e.g., face detection in low-light conditions. The source code and pre-trained models are available at https://github.com/TaoWangzj/LLFormer.
Tao Wang 0052, Kaihao Zhang, Tianrun Shen, Wenhan Luo, Björn Stenger, Tong Lu 0002
AAAI6
2023 InternImage: Exploring Large-Scale Vision Foundation Models with Deformable Convolutions
abstract
Compared to the great progress of large-scale vision transformers (ViTs) in recent years, large-scale models based on convolutional neural networks (CNNs) are still in an early state. This work presents a new large-scale CNN-based foundation model, termed InternImage, which can obtain the gain from increasing parameters and training data like ViTs. Different from the recent CNNs that focus on large dense kernels, InternImage takes deformable convolution as the core operator, so that our model not only has the large effective receptive field required for downstream tasks such as detection and segmentation, but also has the adaptive spatial aggregation conditioned by input and task information. As a result, the proposed InternImage reduces the strict inductive bias of traditional CNNs and makes it possible to learn stronger and more robust patterns with large-scale parameters from massive data like ViTs. The effectiveness of our model is proven on challenging benchmarks including ImageNet, COCO, andADE20K. It is worth mentioning that InternImage-H achieved a new record 65.4 mAP on COCO test-dev and 62.9 mIoU on ADE20K, outperforming current leading CNNs and ViTs.
Wenhai Wang, Jifeng Dai, Zhe Chen 0017, Zhenhang Huang, Xizhou Zhu, Xiaowei Hu 0001, Tong Lu 0002, Lewei Lu, Hongsheng Li 0001, Xiaogang Wang 0001, Yu Qiao 0001
CVPR8
2023 DDP: Diffusion Model for Dense Visual Prediction
abstract
We propose a simple, efficient, yet powerful framework for dense visual predictions based on the conditional diffusion pipeline. Our approach follows a "noise-to-map" generative paradigm for prediction by progressively removing noise from a random Gaussian distribution, guided by the image. The method, called DDP, efficiently extends the denoising diffusion process into the modern perception pipeline. Without task-specific design and architecture customization, DDP is easy to generalize to most dense prediction tasks, e.g., semantic segmentation and depth estimation. In addition, DDP shows attractive properties such as dynamic inference and uncertainty awareness, in contrast to previous single-step discriminative methods. We show top results on three representative tasks with six diverse benchmarks, without tricks, DDP achieves state-of-the-art or competitive performance on each task compared to the specialist counterparts. For example, semantic segmentation (83.9 mIoU on Cityscapes), BEV map segmentation (70.6 mIoU on nuScenes), and depth estimation (0.05 REL on KITTI). We hope that our approach will serve as a solid baseline and facilitate future research.
Yuanfeng Ji, Zhe Chen 0017, Enze Xie, Lanqing Hong, Xihui Liu, Zhaoqiang Liu, Tong Lu 0002, Zhenguo Li, Ping Luo 0002
ICCV7
2023 FB-BEV: BEV Representation from Forward-Backward View Transformations
abstract
View Transformation Module (VTM), where transformations happen between multi-view image features and Bird-Eye-View (BEV) representation, is a crucial step in camera-based BEV perception systems. Currently, the two most prominent VTM paradigms are forward projection and backward projection. Forward projection, represented by Lift-Splat-Shoot, leads to sparsely projected BEV features without post-processing. Backward projection, with BEV-Former being an example, tends to generate false-positive BEV features from incorrect projections due to the lack of utilization on depth. To address the above limitations, we propose a novel forward-backward view transformation module. Our approach compensates for the deficiencies in both existing methods, allowing them to enhance each other to obtain higher quality BEV representations mutually. We instantiate the proposed module with FB-BEV, which achieves a new state-of-the-art result of 62.4% NDS on the nuScenes test set. Code and models are available at https://github.com/NVlabs/FB-BEV.
Zhiding Yu, Wenhai Wang, Anima Anandkumar, Tong Lu 0002, José M. Álvarez 0004
ICCV5
2023 Memory-and-Anticipation Transformer for Online Action Understanding
abstract
Most existing forecasting systems are memory-based methods, which attempt to mimic human forecasting ability by employing various memory mechanisms and have progressed in temporal modeling for memory dependency. Nevertheless, an obvious weakness of this paradigm is that it can only model limited historical dependence and can not transcend the past. In this paper, we rethink the temporal dependence of event evolution and propose a novel memory-anticipation-based paradigm to model an entire temporal structure, including the past, present, and future. Based on this idea, we present Memory-and-Anticipation Transformer (MAT), a memory-anticipation-based approach, to address the online action detection and anticipation tasks. In addition, owing to the inherent superiority of MAT, it can process online action detection and anticipation tasks in a unified manner. The proposed MAT model is tested on four challenging benchmarks TVSeries, THUMOS’14, HDD, and EPIC-Kitchens-100, for online action detection and anticipation tasks, and it significantly outperforms all existing methods. Code is available at https://github.com/Echo0125/Memory-and-Anticipation-Transformer.
Jiahao Wang 0005, Guo Chen 0006, Yifei Huang 0002, Limin Wang 0002, Tong Lu 0002
ICCV5
2023 ICDAR 2023 Competition on Born Digital Video Text Question Answering
Zhibo Yang 0003, Xiaoge Song, Sibo Song, Tong Lu 0002, Xiang Bai, Cheng-Lin Liu 0001, Fei Huang 0002, Cong Yao
ICDAR (2)4
2023 Vision Transformer Adapter for Dense Predictions
Zhe Chen 0017, Yuchen Duan, Wenhai Wang, Junjun He, Tong Lu 0002, Jifeng Dai, Yu Qiao 0001
ICLR5
2023 ELAN: Enhancing Temporal Action Detection with Location Awareness
abstract
Current query-based temporal action detection methods lack multiple levels of location awareness, leading to performance degradation. In this paper, we present a novel query-based method called Enhanced Location-Aware Network (ELAN) for temporal action detection. ELAN adopts a lightweight convolution-based encoder, termed Temporal Location-Aware (TLA) encoder, to model temporal continuous location-aware context. Moreover, ELAN can re-aware the location-related context inside and between queries through our proposed Instance Location-Aware (ILA) decoder. As a result, ELAN can learn strong position discrimination of actions and effectively eliminates the ambiguity caused by sparse action decoding, yielding significant improvement in detection performance. ELAN achieves state-of-the-art performance on two temporal action detection benchmarks, including THUMOS-14 and ActivityNet-1.3.
Guo Chen 0006, Yin-Dong Zheng, Zhe Chen 0017, Jiahao Wang 0005, Tong Lu 0002
ICME5
2023 MRSN: Multi-Relation Support Network for Video Action Detection
abstract
Action detection is a challenging video understanding task, requiring modeling spatio-temporal and interaction relations. Current methods usually model actor-actor and actor-context relations separately, ignoring their complementarity and mutual support. To solve this problem, we propose a novel network called Multi-Relation Support Network (MRSN). In MRSN, Actor-Context Relation Encoder (ACRE) and Actor-Actor Relation Encoder (AARE) model the actor-context and actor-actor relation separately. Then Relation Support Encoder (RSE) computes the supports between the two relations and performs relation-level interactions. Finally, Relation Consensus Module (RCM) enhances two relations with the long-term relations from the Long-term Relation Bank (LRB) and yields a consensus. Our experiments demonstrate that modeling relations separately and performing relation-level interactions can achieve and outperformer state-of-the-art results on two challenging video datasets: AVA and UCF101-24.
Yin-Dong Zheng, Guo Chen 0006, Minglei Yuan, Tong Lu 0002
ICME4
2023 Graph Propagation Transformer for Graph Representation Learning
abstract
This paper presents a novel transformer architecture for graph representation learning. The core insight of our method is to fully consider the information propagation among nodes and edges in a graph when building the attention module in the transformer blocks. Specifically, we propose a new attention mechanism called Graph Propagation Attention (GPA). It explicitly passes the information among nodes and edges in three ways, i.e. node-to-node, node-to-edge, and edge-to-node, which is essential for learning graph-structured data. On this basis, we design an effective transformer architecture named Graph Propagation Transformer (GPTrans) to further help learn graph data. We verify the performance of GPTrans in a wide range of graph learning experiments on several benchmark datasets. These results show that our method outperforms many state-of-the-art transformer-based graph models with better performance. The code will be released at https://github.com/czczup/GPTrans.
Zhe Chen 0017, Tao Wang 0052, Tianrun Shen, Tong Lu 0002, Qiuying Peng
IJCAI5
2023 VisionLLM: Large Language Model is also an Open-Ended Decoder for Vision-Centric Tasks
abstract
Large language models (LLMs) have notably accelerated progress towards artificial general intelligence (AGI), with their impressive zero-shot capacity for user-tailored tasks, endowing them with immense potential across a range of applications. However, in the field of computer vision, despite the availability of numerous powerful vision foundation models (VFMs), they are still restricted to tasks in a pre-defined form, struggling to match the open-ended task capabilities of LLMs. In this work, we present an LLM-based framework for vision-centric tasks, termed VisionLLM. This framework provides a unified perspective for vision and language tasks by treating images as a foreign language and aligning vision-centric tasks with language tasks that can be flexibly defined and managed using language instructions. An LLM-based decoder can then make appropriate predictions based on these instructions for open-ended tasks. Extensive experiments show that the proposed VisionLLM can achieve different levels of task customization through language instructions, from fine-grained object-level to coarse-grained task-level customization, all with good results. It's noteworthy that, with a generalist LLM-based framework, our model can achieve over 60% mAP on COCO, on par with detection-specific models. We hope this model can set a new baseline for generalist vision and language models. The code shall be released.
Wenhai Wang, Zhe Chen 0017, Xiaokang Chen, Jiannan Wu, Xizhou Zhu, Ping Luo 0002, Tong Lu 0002, Jie Zhou 0001, Yu Qiao 0001, Jifeng Dai
NeurIPS8
2023 BasicTAD: An astounding RGB-Only baseline for temporal action detection
Min Yang 0011, Guo Chen 0006, Yin-Dong Zheng, Tong Lu 0002, Limin Wang 0002
Comput. Vis. Image Underst.4
2023 Classification of aesthetic natural scene images using statistical and semantic features
Kunal Biswas, Palaiahnakote Shivakumara, Umapada Pal 0001, Tong Lu 0002, Michael Blumenstein, Josep Lladós 0001
Multim. Tools Appl.4
2023 Writer age estimation through handwriting
Zhiheng Huang, Palaiahnakote Shivakumara, Maryam Asadzadeh Kaljahi, Ahlad Kumar, Umapada Pal 0001, Tong Lu 0002, Michael Blumenstein
Multim. Tools Appl.6
2023 Refine-Net: Normal Refinement Neural Network for Noisy Point Clouds
abstract
Point normal, as an intrinsic geometric property of 3D objects, not only serves conventional geometric tasks such as surface consolidation and reconstruction, but also facilitates cutting-edge learning-based techniques for shape analysis and generation. In this paper, we propose a normal refinement network, called Refine-Net, to predict accurate normals for noisy point clouds. Traditional normal estimation wisdom heavily depends on priors such as surface shapes or noise distributions, while learning-based solutions settle for single types of hand-crafted features. Differently, our network is designed to refine the initial normal of each point by extracting additional information from multiple feature representations. To this end, several feature modules are developed and incorporated into Refine-Net by a novel connection module. Besides the overall network architecture of Refine-Net, we propose a new multi-scale fitting patch selection scheme for the initial normal estimation, by absorbing geometry domain knowledge. Also, Refine-Net is a generic normal estimation framework: 1) point normals obtained from other methods can be further refined, and 2) any feature module related to the surface geometric structures can be potentially integrated into the framework. Qualitative and quantitative evaluations demonstrate the clear superiority of Refine-Net over the state-of-the-arts on both synthetic and real-scanned datasets.
Honghua Chen, Yingkui Zhang, Mingqiang Wei, Haoran Xie 0001, Jun Wang 0039, Tong Lu 0002, Harry Qin, Xiao-Ping Zhang 0002
IEEE Trans. Pattern Anal. Mach. Intell.7
2023 A New Language-Independent Deep CNN for Scene Text Detection and Style Transfer in Social Media Images
abstract
Due to the adverse effect of quality caused by different social media and arbitrary languages in natural scenes, detecting text from social media images and transferring its style is challenging. This paper presents a novel end-to-end model for text detection and text style transfer in social media images. The key notion of the proposed work is to find dominant information, such as fine details in the degraded images (social media images), and then restore the structure of character information. Therefore, we first introduce a novel idea of extracting gradients from the frequency domain of the input image to reduce the adverse effect of different social media, which outputs text candidate points. The text candidates are further connected into components and used for text detection via a UNet++ like network with an EfficientNet backbone (EffiUNet++). Then, to deal with the style transfer issue, we devise a generative model, which comprises a target encoder and style parameter networks (TESP-Net) to generate the target characters by leveraging the recognition results from the first stage. Specifically, a series of residual mapping and a position attention module are devised to improve the shape and structure of generated characters. The whole model is trained end-to-end so as to optimize the performance. Experiments on our social media dataset, benchmark datasets of natural scene text detection and text style transfer show that the proposed model outperforms the existing text detection and style transfer methods in multilingual and cross-language scenario.
Palaiahnakote Shivakumara, Ayan Banerjee 0002, Umapada Pal 0001, Lokesh Nandanwar, Tong Lu 0002, Cheng-Lin Liu 0001
IEEE Trans. Image Process.5
2023 Digital Twin of Intelligent Small Surface Defect Detection with Cyber-manufacturing Systems
abstract
With the remarkable technological development in cyber-physical systems, industry 4.0 has evolved by use of a significant concept named digital twin (DT). However, it is still difficult to construct a relationship between twin simulation and a real scenario considering dynamic variations, especially when dealing with small surface defect detection tasks with high performance and computation resource requirements. In this article, we aim to construct cyber-manufacturing systems to achieve a DT solution for small surface defect detection task. Focusing on DT-based solution, the proposed system consists of an Edge–Cloud architecture and a surface defect detection algorithm. Considering dynamic characteristics and real-time response requirement, Edge–Cloud architecture is built to achieve smart manufacturing by efficiently collecting, processing, analyzing, and storing data produced by factory. A deep learning–based algorithm is then constructed to detect surface defeats based on multi-modal data, i.e., imaging and depth data. Experiments show the proposed algorithm could achieve high accuracy and recall in small defeat detection task, thus constructing DT in cyber-manufacturing.
Yirui Wu, Guoqiang Yang, Tong Lu 0002, Shaohua Wan 0001
ACM Trans. Internet Techn.4
2022 Towards Ultra-Resolution Neural Style Transfer via Thumbnail Instance Normalization
abstract
We present an extremely simple Ultra-Resolution Style Transfer framework, termed URST, to flexibly process arbitrary high-resolution images (e.g., 10000x10000 pixels) style transfer for the first time. Most of the existing state-of-the-art methods would fall short due to massive memory cost and small stroke size when processing ultra-high resolution images. URST completely avoids the memory problem caused by ultra-high resolution images by (1) dividing the image into small patches and (2) performing patch-wise style transfer with a novel Thumbnail Instance Normalization (TIN). Specifically, TIN can extract thumbnail features' normalization statistics and apply them to small patches, ensuring the style consistency among different patches. Overall, the URST framework has three merits compared to prior arts. (1) We divide input image into small patches and adopt TIN, successfully transferring image style with arbitrary high-resolution. (2) Experiments show that our URST surpasses existing SOTA methods on ultra-high resolution images benefiting from the effectiveness of the proposed stroke perceptual loss in enlarging the stroke size. (3) Our URST can be easily plugged into most existing style transfer methods and directly improve their performance even without training. Code is available at https://git.io/URST.
Zhe Chen 0017, Wenhai Wang, Enze Xie, Tong Lu 0002, Ping Luo 0002
AAAI4
2022 DCAN: Improving Temporal Action Detection via Dual Context Aggregation
abstract
Temporal action detection aims to locate the boundaries of action in the video. The current method based on boundary matching enumerates and calculates all possible boundary matchings to generate proposals. However, these methods neglect the long-range context aggregation in boundary prediction. At the same time, due to the similar semantics of adjacent matchings, local semantic aggregation of densely-generated matchings cannot improve semantic richness and discrimination. In this paper, we propose the end-to-end proposal generation method named Dual Context Aggregation Network (DCAN) to aggregate context on two levels, namely, boundary level and proposal level, for generating high-quality action proposals, thereby improving the performance of temporal action detection. Specifically, we design the Multi-Path Temporal Context Aggregation (MTCA) to achieve smooth context aggregation on boundary level and precise evaluation of boundaries. For matching evaluation, Coarse-to-fine Matching (CFM) is designed to aggregate context on the proposal level and refine the matching map from coarse to fine. We conduct extensive experiments on ActivityNet v1.3 and THUMOS-14. DCAN obtains an average mAP of 35.39% on ActivityNet v1.3 and reaches mAP 54.14% at [email protected] on THUMOS-14, which demonstrates DCAN can generate high-quality proposals and achieve state-of-the-art performance. We release the code at https://github.com/cg1177/DCAN.
Guo Chen 0006, Yin-Dong Zheng, Limin Wang 0002, Tong Lu 0002
AAAI4
2022 Panoptic SegFormer: Delving Deeper into Panoptic Segmentation with Transformers
abstract
Panoptic segmentation involves a combination of joint semantic segmentation and instance segmentation, where image contents are divided into two types: things and stuff. We present Panoptic SegFormer, a general framework for panoptic segmentation with transformers. It contains three innovative components: an efficient deeply-supervised mask decoder, a query decoupling strategy, and an improved postprocessing method. We also use Deformable DETR to efficiently process multiscale features, which is a fast and efficient version of DETR. Specifically, we supervise the attention modules in the mask decoder in a layer-wise manner. This deep supervision strategy lets the attention modules quickly focus on meaningful semantic regions. It improves performance and reduces the number of required training epochs by half compared to Deformable DETR. Our query decoupling strategy decouples the responsibilities of the query set and avoids mutual interference between things and stuff. In addition, our post-processing strategy improves performance without additional costs by jointly considering classification and segmentation qualities to resolve conflicting mask overlaps. Our approach increases the accuracy 6.2% PQ over the baseline DETR model. Panoptic SegFormer achieves state-of-the-art results on COCO testdev with 56.2% PQ. It also shows stronger zero-shot robustness over existing methods.
Wenhai Wang, Enze Xie, Zhiding Yu, Anima Anandkumar, José M. Álvarez 0004, Ping Luo 0002, Tong Lu 0002
CVPR8
2022 BEVFormer: Learning Bird's-Eye-View Representation from Multi-camera Images via Spatiotemporal Transformers
Wenhai Wang, Hongyang Li 0001, Enze Xie, Chonghao Sima, Tong Lu 0002, Yu Qiao 0001, Jifeng Dai
ECCV (9)6
2022 SeedFormer: Patch Seeds Based Point Cloud Completion with Upsample Transformer
Yun Cao 0002, Wenqing Chu, Tong Lu 0002, Ying Tai, Chengjie Wang 0001
ECCV (3)5
2022 Uncertainty-Based Network for Few-Shot Image Classification
abstract
The transductive inference is an effective technique in the few-shot learning task, where query sets update prototypes to improve themselves. However, these methods optimize the model by considering only the classification scores of the query instances as confidence while ignoring the uncertainty of these classification scores. In this paper, we propose a novel method called Uncertainty-Based Network, which models the uncertainty of classification results with the help of mutual information. Specifically, we first data augment and classify the query instance and calculate the mutual information of these classification scores. Then, mutual information is used as uncertainty to assign weights to classification scores, and the iterative update strategy based on classification scores and uncertainties assigns the optimal weights to query instances in prototype optimization. Extensive results on four benchmarks show that Uncertainty-Based Network achieves comparable performance in classification accuracy compared to state-of-the-art methods.
Minglei Yuan, Chunhao Cai, Yin-Dong Zheng, Tao Wang 0052, Tong Lu 0002, Wenbin Li 0006
ICME6
2022 Incremental Few-Shot Semantic Segmentation via Embedding Adaptive-Update and Hyper-class Representation
abstract
Incremental few-shot semantic segmentation (IFSS) targets at incrementally expanding model's capacity to segment new class of images supervised by only a few samples. However, features learned on old classes could significantly drift, causing catastrophic forgetting. Moreover, few samples for pixel-level segmentation on new classes lead to notorious overfitting issues in each learning session. In this paper, we explicitly represent class-based knowledge for semantic segmentation as a category embedding and a hyper-class embedding, where the former describes exclusive semantical properties, and the latter expresses hyper-class knowledge as class-shared semantic properties. Aiming to solve IFSS problems, we present EHNet, i.e., Embedding adaptive-update and Hyper-class representation Network from two aspects. First, we propose an embedding adaptive-update strategy to avoid feature drift, which maintains old knowledge by hyper-class representation, and adaptively update category embeddings with a class-attention scheme to involve new classes learned in individual sessions. Second, to resist overfitting issues caused by few training samples, a hyper-class embedding is learned by clustering all category embeddings for initialization and aligned with category embedding of the new class for enhancement, where learned knowledge assists to learn new knowledge, thus alleviating performance dependence on training data scale. Significantly, these two designs provide representation capability for classes with sufficient semantics and limited biases, enabling to perform segmentation tasks requiring high semantic dependence. Experiments on PASCAL-5i and COCO datasets show that EHNet achieves new state-of-the-art performance with remarkable advantages.
Guangchen Shi, Yirui Wu, Jun Liu 0036, Shaohua Wan 0001, Wenhai Wang, Tong Lu 0002
ACM Multimedia6
2022 PVT v2: Improved baselines with Pyramid Vision Transformer
abstract
Transformers have recently lead to encouraging progress in computer vision. In this work, we present new baselines by improving the original Pyramid Vision Transformer (PVT v1) by adding three designs: (i) a linear complexity attention layer, (ii) an overlapping patch embedding, and (iii) a convolutional feed-forward network. With these modifications, PVT v2 reduces the computational complexity of PVT v1 to linearity and provides significant improvements on fundamental vision tasks such as classification, detection, and segmentation. In particular, PVT v2 achieves comparable or better performance than recent work such as the Swin transformer. We hope this work will facilitate state-of-the-art transformer research in computer vision. Code is available at https://github.com/whai362/PVT .
Wenhai Wang, Enze Xie, Xiang Li 0028, Deng-Ping Fan, Kaitao Song, Ding Liang, Tong Lu 0002, Ping Luo 0002, Ling Shao 0001
Comput. Vis. Media7
2022 Text line segmentation from struck-out handwritten document images
Palaiahnakote Shivakumara, Tanmay Jain, Umapada Pal 0001, Nitish Surana, Apostolos Antonacopoulos, Tong Lu 0002
Expert Syst. Appl.6
2022 A Knowledge Enforcement Network-Based Approach for Classifying a Photographer's Images
abstract
Classification of photos captured by different photographers is an important and challenging problem in knowledge-based and image processing. Monitoring and authenticating images uploaded on social media are essential, and verifying the source is one key piece of evidence. We present a novel framework for classifying photos of different photographers based on the combination of local features and deep learning models. The proposed work uses focused and defocused information in the input images to extract contextual information. The model estimates the weighted gradient and calculates entropy to strengthen context features. The focused and defocused information is fused to estimate cross-covariance and define a linear relationship between them. This relationship results in a feature matrix fed to Knowledge Enforcement Network (KEN) for obtaining representative features. Due to the strong discriminative ability of deep learning models, we employ the lightweight and accurate MobileNetV2. The output of KEN and MobileNetV2 is sent to a classifier for photographer classification. Experimental results of the proposed model on our dataset of 46 photographer classes (46234 images) and publicly available datasets of 41 photographer classes (218303 images) show that the method outperforms the existing techniques by 5%–10% on average. The dataset created for the experimental purpose will be made available upon publication.
Palaiahnakote Shivakumara, Pinaki Nath Chowdhury, Umapada Pal 0001, David S. Doermann, Ramachandra Raghavendra, Tong Lu 0002, Michael Blumenstein
Int. J. Pattern Recognit. Artif. Intell.6
2022 On Efficient Reinforcement Learning for Full-length Game of StarCraft II
abstract
StarCraft II (SC2) poses a grand challenge for reinforcement learning (RL), of which the main difficulties include huge state space, varying action space, and a long time horizon. In this work, we investigate a set of RL techniques for the full-length game of StarCraft II. We investigate a hierarchical RL approach, where the hierarchy involves two. One is the extracted macro-actions from experts’ demonstration trajectories to reduce the action space in an order of magnitude. The other is a hierarchical architecture of neural networks, which is modular and facilitates scale. We investigate a curriculum transfer training procedure that trains the agent from the simplest level to the hardest level. We train the agent on a single machine with 4 GPUs and 48 CPU threads. On a 64x64 map and using restrictive units, we achieve a win rate of 99% against the difficulty level-1 built-in AI. Through the curriculum transfer learning algorithm and a mixture of combat models, we achieve a 93% win rate against the most difficult non-cheating level built-in AI (level-7). In this extended version of the paper, we improve our architecture to train the agent against the most difficult cheating level AIs (level-8, level-9, and level-10). We also test our method on different maps to evaluate the extensibility of our approach. By a final 3-layer hierarchical architecture and applying significant tricks to train SC2 agents, we increase the win rate against the level-8, level-9, and level-10 to 96%, 97%, and 94%, respectively. Our codes and models are all open-sourced now at https://github.com/liuruoze/HierNet-SC2. To provide a baseline referring the AlphaStar for our work as well as the research and open-source community, we reproduce a scaled-down version of it, mini-AlphaStar (mAS). The latest version of mAS is 1.07, which can be trained using supervised learning and reinforcement learning on the raw action space which has 564 actions. It is designed to run training on a single common machine, by making the hyper-parameters adjustable and some settings simplified. We then can compare our work with mAS using the same computing resources and training time. By experiment results, we show that our method is more effective when using limited resources. The inference and training codes of mini-AlphaStar are all open-sourced at https://github.com/liuruoze/mini-AlphaStar. We hope our study could shed some light on the future research of efficient reinforcement learning on SC2 and other large-scale games.
Ruo-Ze Liu, Zhen-Jia Pang, Zhou-Yu Meng, Wenhai Wang, Yang Yu 0001, Tong Lu 0002
J. Artif. Intell. Res.6
2022 Fuzzy and genetic algorithm based approach for classification of personality traits oriented social media images
Kunal Biswas, Palaiahnakote Shivakumara, Umapada Pal 0001, Tapabrata Chakraborti, Tong Lu 0002, Mohamad Nizam Ayub
Knowl. Based Syst.5
2022 A new deep model for family and non-family photo identification
Tapan Karnik, Palaiahnakote Shivakumara, Pinaki Nath Chowdhury, Umapada Pal 0001, Tong Lu 0002, Nor Badrul Anuar
Multim. Tools Appl.5
2022 PAN++: Towards Efficient and Accurate End-to-End Spotting of Arbitrarily-Shaped Text
abstract
Scene text detection and recognition have been well explored in the past few years. Despite the progress, efficient and accurate end-to-end spotting of arbitrarily-shaped text remains challenging. In this work, we propose an end-to-end text spotting framework, termed PAN++, which can efficiently detect and recognize text of arbitrary shapes in natural scenes. PAN++ is based on the kernel representation that reformulates a text line as a text kernel (central region) surrounded by peripheral pixels. By systematically comparing with existing scene text representations, we show that our kernel representation can not only describe arbitrarily-shaped text but also well distinguish adjacent text. Moreover, as a pixel-based representation, the kernel representation can be predicted by a single fully convolutional network, which is very friendly to real-time applications. Taking the advantages of the kernel representation, we design a series of components as follows: 1) a computationally efficient feature enhancement network composed of stacked Feature Pyramid Enhancement Modules (FPEMs); 2) a lightweight detection head cooperating with Pixel Aggregation (PA); and 3) an efficient attention-based recognition head with Masked RoI. Benefiting from the kernel representation and the tailored components, our method achieves high inference speed while maintaining competitive accuracy. Extensive experiments show the superiority of our method. For example, the proposed PAN++ achieves an end-to-end text spotting F-measure of 64.9 at 29.2 FPS on the Total-Text dataset, which significantly outperforms the previous best method. Code will be available at: git.io/PAN.
Wenhai Wang, Enze Xie, Xiang Li 0041, Xuebo Liu 0001, Ding Liang, Zhibo Yang 0003, Tong Lu 0002, Chunhua Shen
IEEE Trans. Pattern Anal. Mach. Intell.7
2022 A novel forget-update module for few-shot domain generalization
Minglei Yuan, Chunhao Cai, Tong Lu 0002, Yirui Wu
Pattern Recognit.3
2022 Oil palm tree counting in drone images
Pinaki Nath Chowdhury, Palaiahnakote Shivakumara, Lokesh Nandanwar, Faizal Samiron, Umapada Pal 0001, Tong Lu 0002
Pattern Recognit. Lett.6
2022 A new method for detection and prediction of occluded text in natural scene images
Ayush Mittal, Palaiahnakote Shivakumara, Umapada Pal 0001, Tong Lu 0002, Michael Blumenstein
Signal Process. Image Commun.4
2022 Efficient Reinforcement Learning for StarCraft by Abstract Forward Models and Transfer Learning
abstract
Injecting human knowledge is an effective way to accelerate reinforcement learning (RL). However, these methods are underexplored. This article presents our discovery that an abstract forward model [thought-game (TG)] combined with transfer learning is an effective way. We takeStarCraft IIas our study environment. With the help of a designed TG, the agent can learn a 99% win-rate on a 64×64 map against the Level-7 built-in AI, using only 1.08 h in a single commercial machine. We also show that the TG method is not as restrictive as it was thought to be. It can work with roughly designed TGs, and can also be useful when the environment changes. Comparing with previous model-based RL, we show TG is more effective. We also present a TG hypothesis that gives the influence of different fidelity levels of TG. For real games that have unequal state and action spaces, we proposed a novel XfrNet of which usefulness is validated while achieving a 90% win-rate against the cheating Level-10 AI. We argue that the TG method might shed light on further studies of efficient RL with human knowledge.
Ruo-Ze Liu, Xiaozhong Ji, Yang Yu 0001, Zhen-Jia Pang, Zitai Xiao, Yuzhou Wu, Tong Lu 0002
IEEE Trans. Games8
2022 An Episodic Learning Network for Text Detection on Human Bodies in Sports Images
abstract
Due to the proliferation of sports-related multimedia content on the WWW, effective visual search and retrieval present interesting research challenges. These are caused by poor image quality, a wide range of possible camera points of view, pose variations on the part of athletes engaged in playing a sport, deformations of text appearing on sports person’s clothing and uniforms in motion, occlusions caused by other objects, etc. To address these challenges, this paper presents a new method for detecting text on human bodies in sports images. Unlike most existing methods, which attempt to exploit locations of a player’s torso, face, and skin, we propose an end-to-end episodic learning approach that employs inductive learning criteria for detecting clothing regions in an image, which are, in turn, then used for text detection. Our method integrates a Residual Network (ResNet) and Pyramidal Pooling Module (PPM) for generating a spatial attention map. The Progressive Scalable Expansion Algorithm (PSE) is adapted for text detection from these regions. Experimental results on our own dataset as well as several benchmarks (like RBNR and MMM which contain images of runners in marathons, and Re-ID which is a person re-identification dataset) demonstrate that the proposed method outperforms existing methods in terms of precision and F1-score. We also present results for sports images chosen from natural scene text detection datasets such as CTW1500 and MS-COCO to show the proposed method is effective and reliable across a range of inputs.
Pinaki Nath Chowdhury, Palaiahnakote Shivakumara, Ramachandra Raghavendra, Sauradip Nag, Umapada Pal 0001, Tong Lu 0002, Daniel P. Lopresti
IEEE Trans. Circuits Syst. Video Technol.6
2022 A New Deep Wavefront Based Model for Text Localization in 3D Video
abstract
With the evolution of electronic devices, such as 3D cameras, addressing the challenges of text localization in 3D video (e.g., for indexing) is increasingly drawing the attention of the multimedia and video processing community. Existing methods focus on 2D video and their performance in the presence of the challenges in 3D video, such as shadow areas associated with text and irregularly sized and shaped text, degrades. This paper proposes the first approach that successfully addresses the challenges of 3D video in addition to those of 2D. It employs a number of innovations, among which, the first is the Generalized Gradient Vector Flow (GGVF) for dominant points detection. The second is the Wavefront concept for text candidate point detection from those dominant points. In addition, an Adaptive B-Spline Polygon Curve Network (ABS-Net) is proposed for accurate text localization in 3D videos by constructing tight fitting bounding polygons using text candidate points. Extensive experiments on custom (3D video) and standard datasets (2D video and scene text) show that the proposed method is practical and useful, and overall outperforms existing state-of-the-art methods.
Lokesh Nandanwar, Palaiahnakote Shivakumara, Ramachandra Raghavendra, Tong Lu 0002, Umapada Pal 0001, Apostolos Antonacopoulos, Yue Lu 0001
IEEE Trans. Circuits Syst. Video Technol.4
2021 Frequency Consistent Adaptation for Real World Super Resolution
abstract
Recent deep-learning based Super-Resolution (SR) methods have achieved remarkable performance on images with known degradation. However, these methods always fail in real-world scene, since the Low-Resolution (LR) images after the ideal degradation (e.g., bicubic down-sampling) deviate from real source domain. The domain gap between the LR images and the real-world images can be observed clearly on frequency density, which inspires us to explicitly narrow the undesired gap caused by incorrect degradation. From this point of view, we design a novel Frequency Consistent Adaptation (FCA) that ensures the frequency domain consistency when applying existing SR methods to the real scene. We estimate degradation kernels from unsupervised images and generate the corresponding LR images. To provide useful gradient information for kernel estimation, we propose Frequency Density Comparator (FDC) by distinguishing the frequency density of images on different scales. Based on the domain-consistent LR-HR pairs, we train easy-implemented Convolutional Neural Network (CNN) SR models. Extensive experiments show that the proposed FCA improves the performance of the SR model under real-world setting achieving state-of-the-art results with high fidelity and plausible perception, thus providing a novel effective framework for real-world SR application.
Xiaozhong Ji, Guangpin Tao, Yun Cao 0002, Ying Tai, Tong Lu 0002, Chengjie Wang 0001, Feiyue Huang
AAAI5
2021 TAM: Temporal Adaptive Module for Video Recognition
abstract
Video data is with complex temporal dynamics due to various factors such as camera motion, speed variation, and different activities. To effectively capture this diverse motion pattern, this paper presents a new temporal adaptive module (TAM) to generate video-specific temporal kernels based on its own feature map. TAM proposes a unique two-level adaptive modeling scheme by decoupling the dynamic kernel into a location sensitive importance map and a location invariant aggregation weight. The importance map is learned in a local temporal window to capture short-term information, while the aggregation weight is generated from a global view with a focus on long-term structure. TAM is a modular block and could be integrated into 2D CNNs to yield a powerful video architecture (TANet) with a very small extra computational cost. The extensive experiments on Kinetics-400 and Something-Something datasets demonstrate that our TAM outperforms other temporal modeling methods consistently, and achieves the state-of-the-art performance under the similar complexity. The code is available at https://github.com/liu-zhy/temporal-adaptive-module.
Zhaoyang Liu 0001, Limin Wang 0002, Wayne Wu, Chen Qian 0006, Tong Lu 0002
ICCV5
2021 Pyramid Vision Transformer: A Versatile Backbone for Dense Prediction without Convolutions
abstract
Although convolutional neural networks (CNNs) have achieved great success in computer vision, this work investigates a simpler, convolution-free backbone network use-fid for many dense prediction tasks. Unlike the recently-proposed Vision Transformer (ViT) that was designed for image classification specifically, we introduce the Pyramid Vision Transformer (PVT), which overcomes the difficulties of porting Transformer to various dense prediction tasks. PVT has several merits compared to current state of the arts. (1) Different from ViT that typically yields low-resolution outputs and incurs high computational and memory costs, PVT not only can be trained on dense partitions of an image to achieve high output resolution, which is important for dense prediction, but also uses a progressive shrinking pyramid to reduce the computations of large feature maps. (2) PVT inherits the advantages of both CNN and Transformer, making it a unified backbone for various vision tasks without convolutions, where it can be used as a direct replacement for CNN backbones. (3) We validate PVT through extensive experiments, showing that it boosts the performance of many downstream tasks, including object detection, instance and semantic segmentation. For example, with a comparable number of parameters, PVT+RetinaNet achieves 40.4 AP on the COCO dataset, surpassing ResNet50+RetinNet (36.3 AP) by 4.1 absolute AP (see Figure 2). We hope that PVT could, serre as an alternative and useful backbone for pixel-level predictions and facilitate future research.
Wenhai Wang, Enze Xie, Xiang Li 0028, Deng-Ping Fan, Kaitao Song, Ding Liang, Tong Lu 0002, Ping Luo 0002, Ling Shao 0001
ICCV7
2021 Adaptive Graph Convolution for Point Cloud Analysis
abstract
Convolution on 3D point clouds that generalized from 2D grid-like domains is widely researched yet far from perfect. The standard convolution characterises feature correspondences indistinguishably among 3D points, presenting an intrinsic limitation of poor distinctive feature learning. In this paper, we propose Adaptive Graph Convolution (AdaptConv) which generates adaptive kernels for points according to their dynamically learned features. Compared with using a fixed/isotropic kernel, AdaptConv improves the flexibility of point cloud convolutions, effectively and precisely capturing the diverse relations between points from different semantic parts. Unlike popular attentional weight schemes, the proposed AdaptConv implements the adaptiveness inside the convolution operation instead of simply assigning different weights to the neighboring points. Extensive qualitative and quantitative evaluations show that our method outperforms state-of-the-art point cloud classification and segmentation approaches on several benchmark datasets. Our code is available at https://github.com/hrzhou2/AdaptConv-master.
Yidan Feng, Mingsheng Fang, Mingqiang Wei, Harry Qin, Tong Lu 0002
ICCV6
2021 DCINN: Deformable Convolution and Inception Based Neural Network for Tattoo Text Detection Through Skin Region
Tamal Chowdhury, Palaiahnakote Shivakumara, Umapada Pal 0001, Tong Lu 0002, Ramachandra Raghavendra, Sukalpa Chanda
ICDAR (2)4
2021 ARNet: Active-Reference Network for Few-Shot Image Semantic Segmentation
abstract
To make predictions on unseen classes, few-shot segmentation becomes a research focus recently. However, most methods build on pixel-level annotation requiring quantity of manual work. Moreover, inherent information on same-category objects to guide segmentation could have large diversity in feature representation due to differences in size, appearance, layout, and so on. To tackle these problems, we present an active-reference network (ARNet) for few-shot segmentation. The proposed active-reference mechanism not only supports accurately cooccurrent objects in either support or query images, but also relaxes high constraint on pixel-level labeling, allowing for weakly boundary labeling. To extract more intrinsic feature representation, a category-modulation module (CMM) is further applied to fuse features extracted from multiple support images, thus forgetting useless and enhancing contributive information. Experiments on PASCAL-5idataset show the proposed method achieves a m-IOU score of 56.5% for 1-shot and 59.8% for 5-shot segmentation, being 0.5% and 1.3% higher than current state-of-the-art method.
Guangchen Shi, Yirui Wu, Palaiahnakote Shivakumara, Umapada Pal 0001, Tong Lu 0002
ICME5
2021 Spectrum-to-Kernel Translation for Accurate Blind Image Super-Resolution
abstract
Deep-learning based Super-Resolution (SR) methods have exhibited promising performance under non-blind setting where blur kernel is known; however, blur kernels of Low-Resolution (LR) images in different practical applications are usually unknown. It may lead to a significant performance drop when degradation process of training images deviates from that of real images. In this paper, we propose a novel blind SR framework to super-resolve LR images degraded by arbitrary blur kernel with accurate kernel estimation in frequency domain. To our best knowledge, this is the first deep learning method which conducts blur kernel estimation in frequency domain. Specifically, we first demonstrate that feature representation in frequency domain is more conducive for blur kernel reconstruction than in spatial domain. Next, we present a Spectrum-to-Kernel (S$2$K) network to estimate general blur kernels in diverse forms. We use a conditional GAN (CGAN) combined with SR-oriented optimization target to learn the end-to-end translation from degraded images' spectra to unknown kernels. Extensive experiments on both synthetic and real-world images demonstrate that our proposed method sufficiently reduces blur kernel estimation error, thus enables the off-the-shelf non-blind SR methods to work under blind setting effectively, and achieves superior performance over state-of-the-art blind SR methods, averagely by 1.39dB, 0.48dB (Gaussian kernels) and 6.15dB, 4.57dB (motion kernels) for scales $2\times$ and $4\times$ respectively.
Guangpin Tao, Xiaozhong Ji, Wenzhuo Wang, Shuo Chen 0003, Chuming Lin, Yun Cao 0002, Tong Lu 0002, Donghao Luo 0001, Ying Tai
NeurIPS7
2021 DCT-phase statistics for forged IMEI numbers and air ticket detection
Lokesh Nandanwar, Palaiahnakote Shivakumara, Swati Kanchan, V. Basavaraja, D. S. Guru, Umapada Pal 0001, Tong Lu 0002, Michael Blumenstein
Expert Syst. Appl.7
2021 Improved Ring Radius Transform-Based Reconstruction for Video Character Recognition
abstract
Character shape reconstruction in video is challenging due to low contrast, complex backgrounds and arbitrary orientation of characters. This work proposes an Improved Ring Radius Transform (IRRT) for reconstructing impaired characters through medial axis prediction. At first, the technique proposes a novel idea based on the Tangent Vector (TV) concept that identifies each actual pair of end pixels caused by gaps in impaired character components. Next, the actual direction to predict medial axis pixels using IRRT for each pair of end pixels is proposed with a new normal vector concept. The process of prediction repeats iteratively to find all the medial axis pixels for every gap in question. Further, medial axis pixels with their radii are used to reconstruct the shapes of impaired characters. The proposed technique is tested on benchmark datasets consisting of video, natural scenes, objects and multi-lingual data to demonstrate that it reconstructs shapes well, even for heterogeneous data. Comparative studies with different binarization and character recognition methods show that the proposed technique is effective, useful and outperforms existing methods.
Zhiheng Huang, Palaiahnakote Shivakumara, Tong Lu 0002, Umapada Pal 0001, Michael Blumenstein, Bhaarat Chetty, G. Hemantha Kumar 0001
Int. J. Pattern Recognit. Artif. Intell.3
2021 A New Hybrid Method for Caption and Scene Text Classification in Action Video Images
abstract
Achieving a better recognition rate for text in action video images is challenging due to multiple types of text with unpredictable actions in the background. In this paper, we propose a new method for the classification of caption (which is edited text) and scene text (text that is a part of the video) in video images. This work considers five action classes, namely, Yoga, Concert, Teleshopping, Craft, and Recipes, where it is expected that both types of text play a vital role in understanding the video content. The proposed method introduces a new fusion criterion based on Discrete Cosine Transform (DCT) and Fourier coefficients to obtain the reconstructed images for caption and scene text. The fusion criterion involves computing the variances for coefficients of corresponding pixels of DCT and Fourier images, and the same variances are considered as the respective weights. This step results in Reconstructed image-1. Inspired by the special property of Chebyshev-Harmonic-Fourier-Moments (CHFM) that has the ability to reconstruct a redundancy-free image, we explore CHFM for obtaining the Reconstructed image-2. The reconstructed images along with the input image are passed to a Deep Convolutional Neural Network (DCNN) for classification of caption/scene text. Experimental results on five action classes and a comparative study with the existing methods demonstrate that the proposed method is effective. In addition, the recognition results of the before and after the classification obtained from different methods show that the recognition performance improves significantly after classification, compared to before classification.
Lokesh Nandanwar, Palaiahnakote Shivakumara, Umapada Pal 0001, Tong Lu 0002, Michael Blumenstein
Int. J. Pattern Recognit. Artif. Intell.4
2021 A New Method for Detecting Altered Text in Document Images
abstract
As more and more office documents are captured, stored, and shared in digital format, and as image editing software are becoming increasingly more powerful, there is a growing concern about document authenticity. To prevent illicit activities, this paper presents a new method for detecting altered text in document images. The proposed method explores the relationship between positive and negative coefficients of DCT to extract the effect of distortions caused by tampering by fusing reconstructed images of respective positive and negative coefficients, which results in Positive-Negative DCT coefficients Fusion (PNDF). To take advantage of spatial information, we propose to fuse R, G, and B color channels of input images, which results in RGBF (RGB Fusion). Next, the same fusion operation is used for fusing PNDF and RGBF, which results in a fused image for the original input one. We compute a histogram to extract features from the fused image, which results in a feature vector. The feature vector is then fed to a deep neural network for classifying altered text images. The proposed method is tested on our own dataset and the standard datasets from the ICPR 2018 Fraud Contest, Altered Handwriting (AH), and faked IMEI number images. The results show that the proposed method is effective and the proposed method outperforms the existing methods irrespective of image type.
Lokesh Nandanwar, Palaiahnakote Shivakumara, Umapada Pal 0001, Tong Lu 0002, Daniel P. Lopresti, Bhagesh Seraogi, Bidyut B. Chaudhuri
Int. J. Pattern Recognit. Artif. Intell.4
2021 A new context-based feature for classification of emotions in photographs
Divya Krishnani, Palaiahnakote Shivakumara, Tong Lu 0002, Umapada Pal 0001, Daniel P. Lopresti, G. Hemantha Kumar 0001
Multim. Tools Appl.3
2021 A new DCT-PCM method for license plate number detection in drone images
Hamam Mokayed, Palaiahnakote Shivakumara, Hock Woon Hon, Mohan Kankanhalli, Tong Lu 0002, Umapada Pal 0001
Pattern Recognit. Lett.5
2021 Arbitrarily-Oriented Text Detection in Low Light Natural Scene Images
abstract
Text detection in low light natural scene images is challenging due to poor image quality and low contrast. Unlike most existing methods that focus on well-lit (normally daylight) images, the proposed method considers much darker natural scene images. For this task, our method first integrates spatial and frequency domain features through fusion to enhance fine details in the image. Next, we use Maximally Stable Extremal Regions (MSER) for detecting text candidates from the enhanced images. We then introduce Cloud of Line Distribution (COLD) features, which capture the distribution of pixels of text candidates in the polar domain. The extracted features are sent to a Convolution Neural Network (CNN) to correct the bounding boxes for arbitrarily oriented text lines by removing false positives. Experiments are conducted on a dataset of low light images to evaluate the proposed enhancement step. The results show our approach is more effective compared to existing methods in terms of standard quality measures, namely, BRISQE, NIQE and PIQE. In addition, experimental results on a variety of standard benchmark datasets, namely, ICDAR 2013, ICDAR 2015, SVT, Total-Text, ICDAR 2017-MLT and CTW1500, show that the proposed approach not only produces better results for low light images, at the same time it is also competitive for daylight images.
Minglong Xue, Palaiahnakote Shivakumara, Tong Lu 0002, Umapada Pal 0001, Daniel P. Lopresti, Zhibo Yang 0003
IEEE Trans. Multim.5
2021 A New Foreground-Background based Method for Behavior-Oriented Social Media Image Classification
abstract
Due to various applications, research on personal traits using information on social media has become an important area. In this paper, a new method for the classification of behavior-oriented social images uploaded on various social media platforms is presented. The proposed method introduces a multimodality concept using skin of different parts of human body and background information, such as indoor and outdoor environments. For each image, the proposed method detects skin candidate components based on R, G, B color spaces and entropy features. The iterative mutual nearest neighbor approach is proposed to detect accurate skin candidate components, which result in foreground components. Next, the proposed method detects the remaining part (other than skin components) as background components based on structure tensor of R, G, B color spaces, and Maximally Stable Extremal Regions (MSER ) concept in the wavelet domain. We then explore Hanman Transform for extracting context features from foreground and background components through clustering and fusion operation. These features are then fed to an SVM classifier for the classification of behavior-oriented images. Comprehensive experiments on 10-class datasets of Normal Behavior-Oriented Social media Image (NBSI) and Abnormal Behavior-Oriented Social media Image (ABSI) show that the proposed method is effective and outperforms the existing methods in terms of average classification rate. Also, the results on the benchmark dataset of five classes of personality traits and two classes of emotions of different facial expressions (FERPlus dataset) demonstrated the robustness of the proposed method over the existing methods.
Lokesh Nandanwar, Palaiahnakote Shivakumara, Divya Krishnani, Ramachandra Raghavendra, Tong Lu 0002, Umapada Pal 0001, Mohan Kankanhalli
ACM Trans. Multim. Comput. Commun. Appl.5
2020 TEINet: Towards an Efficient Architecture for Video Recognition
abstract
Efficiency is an important issue in designing video architectures for action recognition. 3D CNNs have witnessed remarkable progress in action recognition from videos. However, compared with their 2D counterparts, 3D convolutions often introduce a large amount of parameters and cause high computational cost. To relieve this problem, we propose an efficient temporal module, termed as Temporal Enhancement-and-Interaction (TEI Module), which could be plugged into the existing 2D CNNs (denoted by TEINet). The TEI module presents a different paradigm to learn temporal features by decoupling the modeling of channel correlation and temporal interaction. First, it contains a Motion Enhanced Module (MEM) which is to enhance the motion-related features while suppress irrelevant information (e.g., background). Then, it introduces a Temporal Interaction Module (TIM) which supplements the temporal contextual information in a channel-wise manner. This two-stage modeling scheme is not only able to capture temporal structure flexibly and effectively, but also efficient for model inference. We conduct extensive experiments to verify the effectiveness of TEINet on several benchmarks (e.g., Something-Something V1&V2, Kinetics, UCF101 and HMDB51). Our proposed TEINet can achieve a good recognition accuracy on these datasets but still preserve a high efficiency.
Zhaoyang Liu 0001, Donghao Luo 0001, Yabiao Wang, Limin Wang 0002, Ying Tai, Chengjie Wang 0001, Feiyue Huang, Tong Lu 0002
AAAI9
2020 A New Context-Based Method for Restoring Occluded Text in Natural Scene Images
Ayush Mittal, Palaiahnakote Shivakumara, Umapada Pal 0001, Tong Lu 0002, Michael Blumenstein, Daniel P. Lopresti
DAS4
2020 A New Common Points Detection Method for Classification of 2D and 3D Texts in Video/Scene Images
Lokesh Nandanwar, Palaiahnakote Shivakumara, Ahlad Kumar, Tong Lu 0002, Umapada Pal 0001, Daniel P. Lopresti
DAS4
2020 AE TextSpotter: Learning Visual and Linguistic Representation for Ambiguous Text Spotting
Wenhai Wang, Xuebo Liu 0001, Xiaozhong Ji, Enze Xie, Ding Liang, Zhibo Yang 0003, Tong Lu 0002, Chunhua Shen, Ping Luo 0002
ECCV (14)7
2020 Dynamic Low-Light Image Enhancement for Object Detection via End-to-End Training
abstract
Object detection based on convolutional neural networks is a hot research topic in computer vision. The illumination component in the image has a great impact on object detection, and it will cause a sharp decline in detection performance under low-light conditions. Using low-light image enhancement technique as a pre-processing mechanism can improve image quality and obtain better detection results. However, due to the complexity of low-light environments, the existing enhancement methods may have negative effects on some samples. Therefore, it is difficult to improve the overall detection performance in low-light conditions. In this paper, our goal is to use image enhancement to improve object detection performance rather than perceptual quality for humans. We propose a novel framework that combines low-light enhancement and object detection for end-to-end training. The framework can dynamically select different enhancement subnetworks for each sample to improve the performance of the detector. Our proposed method consists of two stage: the enhancement stage and the detection stage. The enhancement stage dynamically enhances the low-light images under the supervision of several enhancement methods and output corresponding weights. During the detection stage, the weights offers information on object classification to generate high-quality region proposals and in turn result in accurate detection. Our experiments present promising results, which show that the proposed method can significantly improve the detection performance in low-light environment.
Tong Lu 0002, Yirui Wu
ICPR2
2020 Multi-scale Relational Reasoning with Regional Attention for Visual Question Answering
abstract
One of the main challenges of visual question answering (VQA) lies in properly reasoning relations among visual regions involved in the question. In this paper, we propose a novel neural network to perform question-guided relational reasoning in multi-scales for visual question answering, in which each region of image is enhanced by regional attention. Specifically, we present regional attention module, which consists of a soft attention module and a hard attention module, to select informative regions of the image according to informative evaluations implemented by question-guided soft attention. Combinations of different informative regions are then concatenated with question embedding in different scales to capture relational information. Relational reasoning module can extract question-based relational information among regions, in which multi-scale mechanism gives it the ability to model scaled relationships with diversity making it sensitive to numbers. We conduct experiments to show that our proposed architecture is effective and achieves a new state-of-the-art on VQA v2.
Yun-Tao Ma, Tong Lu 0002, Yirui Wu
ICPR2
2020 Chebyshev-Harmonic-Fourier-Moments and Deep CNNs for Detecting Forged Handwriting
abstract
Recently developed sophisticated image processing techniques and tools have made easier the creation of high-quality forgeries of handwritten documents including financial and property records. To detect such forgeries of handwritten documents, this paper presents a new method by exploring the combination of Chebyshev-Harmonic-Fourier-Moments (CHFM) and deep Convolutional Neural Networks (D-CNNs). Unlike existing methods work based on abrupt changes due to distortion created by forgery operation, the proposed method works based on inconsistencies and irregular changes created by forgery operations. Inspired by the special properties of CHFM, such as its reconstruction ability by removing redundant information, the proposed method explores CHFM to obtain reconstructed images for the color components of the Original, Forged Noisy and Blurred classes. Motivated by the strong discriminative power of deep CNNs, for the reconstructed images of respective color components, the proposed method used deep CNNs for forged handwriting detection. Experimental results on our dataset and benchmark datasets (namely, ACPR 2019, ICPR 2018 FCD and IMEI datasets) show that the proposed method outperforms existing methods in terms of classification rate.
Lokesh Nandanwar, Palaiahnakote Shivakumara, Sayani Kundu, Umapada Pal 0001, Tong Lu 0002, Daniel P. Lopresti
ICPR5
2020 Local Gradient Difference Features for Classification of 2D-3D Natural Scene Text Images
abstract
Methods developed for normal 2D text detection do not work well for text that is rendered using decorative, 3D effects, etc. This paper proposes a new method for classification of 2D and 3D natural scene text images so that an appropriate recognition method can be chosen accordingly based on the classification results for better performance. The proposed method explores local gradient differences for obtaining candidate pixels, which represent a stroke. To study the spatial distribution of candidate pixels, we propose a measure, called COLD, which is denser for pixels toward the center of strokes and scattered for non-stroke pixels. This observation leads us to introduce mass features for extracting the regular spatial pattern of COLD, which indicates a 2D text image. The extracted features are fed into a Neural Network (NN) for classification. The proposed method is tested on (i) a new dataset introduced in this work (ii) a second dataset assembled from standard natural scene datasets (iii) Non-Text Image datasets which does not contain text, rather it contains objects. Experimental results of the proposed method on images with text and non-text show that the proposed method is independent of text. The proposed approach improves text detection and recognition performance significantly after classification.
Lokesh Nandanwar, Palaiahnakote Shivakumara, Ramachandra Raghavendra, Tong Lu 0002, Umapada Pal 0001, Daniel P. Lopresti, Nor Badrul Anuar
ICPR4
2020 Context-Aware Residual Network with Promotion Gates for Single Image Super-Resolution
Xiaozhong Ji, Yirui Wu, Tong Lu 0002
MMM (2)3
2020 TK-Text: Multi-shaped Scene Text Detection via Instance Segmentation
Xiaoge Song, Yirui Wu, Wenhai Wang, Tong Lu 0002
MMM (2)4
2020 Rotation invariant angle-density based features for an ice image classification system
Shengkai Yue, Minglei Yuan, Tong Lu 0002, Palaiahnakote Shivakumara, Michael Blumenstein, G. Hemantha Kumar 0001
Expert Syst. Appl.3
2020 Forged text detection in video, scene, and document images
abstract
Rapid advances in artificial intelligence have made it possible to produce forgeries good enough to fool an average user. As a result, there is growing interest in developing robust methods to counter such forgeries. This study presents a new Fourier spectrum‐based method for detecting forged text in video images. The authors' premise is that brightness distribution and the spectrum shape exhibit irregular patterns (inconsistencies) for forged text, while appearing more regular for original text. The method divides the spectrum of an input image into sectors and tracks to highlight these effects. Specifically, positive and negative coefficients for sectors and tracks are extracted to quantify the brightness distribution. Variations in the shape of the spectrum are analysed by determining the angular relationship between the principal axes and the sectors/tracks of the spectrum. Next, it combines these two features to detect forged text in the images of IMEI (International Mobile Equipment Identity) numbers and document. For evaluation, the following datasets are used: own video dataset and standard datasets, namely, IMEI number, ICPR 2018 Fraud Document Contest, and a natural scene text dataset. Experimental results show that the proposed method outperforms existing methods in terms of average classification rate and F ‐score.
Lokesh Nandanwar, Palaiahnakote Shivakumara, Prabir Mondal, Raghunandan K. Srinivas, Umapada Pal 0001, Tong Lu 0002, Daniel P. Lopresti
IET Image Process.6
2020 A new augmentation-based method for text detection in night and day license plate images
Pinaki Nath Chowdhury, Palaiahnakote Shivakumara, Umapada Pal 0001, Tong Lu 0002, Michael Blumenstein
Multim. Tools Appl.4
2020 A new unified method for detecting text from marathon runners and sports players in video (PR-D-19-01078R2)
Sauradip Nag, Palaiahnakote Shivakumara, Umapada Pal 0001, Tong Lu 0002, Michael Blumenstein
Pattern Recognit.4
2020 Graph attention network for detecting license plates in crowded street scenes
Pinaki Nath Chowdhury, Palaiahnakote Shivakumara, Swati Kanchan, Ramachandra Raghavendra, Umapada Pal 0001, Tong Lu 0002, Daniel P. Lopresti
Pattern Recognit. Lett.6
2020 Delaunay triangulation based text detection from multi-view images of natural scene
Soumyadip Roy, Palaiahnakote Shivakumara, Umapada Pal 0001, Tong Lu 0002, G. Hemantha Kumar 0001
Pattern Recognit. Lett.4
2020 A new Fractal Series Expansion based enhancement model for license plate recognition
Pinaki Nath Chowdhury, Palaiahnakote Shivakumara, Hamid Abdullah Jalab, Rabha W. Ibrahim, Umapada Pal 0001, Tong Lu 0002
Signal Process. Image Commun.6
2020 Dynamic Sampling Networks for Efficient Action Recognition in Videos
abstract
The existing action recognition methods are mainly based on clip-level classifiers such as two-stream CNNs or 3D CNNs, which are trained from the randomly selected clips and applied to densely sampled clips during testing. However, this standard setting might be suboptimal for training classifiers and also requires huge computational overhead when deployed in practice. To address these issues, we propose a new framework for action recognition in videos, called Dynamic Sampling Networks (DSN), by designing a dynamic sampling module to improve the discriminative power of learned clip-level classifiers and as well increase the inference efficiency during testing. Specifically, DSN is composed of a sampling module and a classification module, whose objective is to learn a sampling policy to on-the-fly select which clips to keep and train a clip-level classifier to perform action recognition based on these selected clips, respectively. In particular, given an input video, we train an observation network in an associative reinforcement learning setting to maximize the rewards of the selected clips with a correct prediction. We perform extensive experiments to study different aspects of the DSN framework on four action recognition datasets: UCF101, HMDB51, THUMOS14, and ActivityNet v1.3. The experimental results demonstrate that DSN is able to greatly improve the inference efficiency by only using less than half of the clips, which can still obtain a slightly better or comparable recognition accuracy to the state-of-the-art approaches.
Yin-Dong Zheng, Zhaoyang Liu 0001, Tong Lu 0002, Limin Wang 0002
IEEE Trans. Image Process.3
2019 On Reinforcement Learning for Full-Length Game of StarCraft
abstract
StarCraft II poses a grand challenge for reinforcement learning. The main difficulties include huge state space, varying action space, long horizon, etc. In this paper, we investigate a set of techniques of reinforcement learning for the full-length game of StarCraft II. We investigate a hierarchical approach, where the hierarchy involves two levels of abstraction. One is the macro-actions extracted from expert’s demonstration trajectories, which can reduce the action space in an order of magnitude yet remain effective. The other is a two-layer hierarchical architecture, which is modular and easy to scale. We also investigate a curriculum transfer learning approach that trains the agent from the simplest opponent to harder ones. On a 64×64 map and using restrictive units, we train the agent on a single machine with 4 GPUs and 48 CPU threads. We achieve a winning rate of more than 99% against the difficulty level-1 built-in AI. Through the curriculum transfer learning algorithm and a mixture of combat model, we can achieve over 93% winning rate against the most difficult noncheating built-in AI (level-7) within days. We hope this study could shed some light on the future research of large-scale reinforcement learning.
Zhen-Jia Pang, Ruo-Ze Liu, Zhou-Yu Meng, Yi Zhang 0108, Yang Yu 0001, Tong Lu 0002
AAAI6
2019 Shape Robust Text Detection With Progressive Scale Expansion Network
abstract
Scene text detection has witnessed rapid progress especially with the recent development of convolutional neural networks. However, there still exists two challenges which prevent the algorithm into industry applications. On the one hand, most of the state-of-art algorithms require quadrangle bounding box which is in-accurate to locate the texts with arbitrary shape. On the other hand, two text instances which are close to each other may lead to a false detection which covers both instances. Traditionally, the segmentation-based approach can relieve the first problem but usually fail to solve the second challenge. To address these two challenges, in this paper, we propose a novel Progressive Scale Expansion Network (PSENet), which can precisely detect text instances with arbitrary shapes. More specifically, PSENet generates the different scale of kernels for each text instance, and gradually expands the minimal scale kernel to the text instance with the complete shape. Due to the fact that there are large geometrical margins among the minimal scale kernels, our method is effective to split the close text instances, making it easier to use segmentation-based methods to detect arbitrary-shaped text instances. Extensive experiments on CTW1500, Total-Text, ICDAR 2015 and ICDAR 2017 MLT validate the effectiveness of PSENet. Notably, on CTW1500, a dataset full of long curve texts, PSENet achieves a F-measure of 74.3% at 27 FPS, and our best F-measure (82.2%) outperforms state-of-art algorithms by 6.6%. The code will be released in the future.
Wenhai Wang, Enze Xie, Xiang Li 0041, Wenbo Hou, Tong Lu 0002, Gang Yu 0002, Shuai Shao 0005
CVPR5
2019 Efficient and Accurate Arbitrary-Shaped Text Detection With Pixel Aggregation Network
abstract
Scene text detection, an important step of scene text reading systems, has witnessed rapid development with convolutional neural networks. Nonetheless, two main challenges still exist and hamper its deployment to real-world applications. The first problem is the trade-off between speed and accuracy. The second one is to model the arbitrary-shaped text instance. Recently, some methods have been proposed to tackle arbitrary-shaped text detection, but they rarely take the speed of the entire pipeline into consideration, which may fall short in practical applications. In this paper, we propose an efficient and accurate arbitrary-shaped text detector, termed Pixel Aggregation Network (PAN), which is equipped with a low computational-cost segmentation head and a learnable post-processing. More specifically, the segmentation head is made up of Feature Pyramid Enhancement Module (FPEM) and Feature Fusion Module (FFM). FPEM is a cascadable U-shaped module, which can introduce multi-level information to guide the better segmentation. FFM can gather the features given by the FPEMs of different depths into a final feature for segmentation. The learnable post-processing is implemented by Pixel Aggregation (PA), which can precisely aggregate text pixels by predicted similarity vectors. Experiments on several standard benchmarks validate the superiority of the proposed PAN. It is worth noting that our method can achieve a competitive F-measure of 79.9% at 84.2 FPS on CTW1500.
Wenhai Wang, Enze Xie, Xiaoge Song, Yuhang Zang, Wenjia Wang 0009, Tong Lu 0002, Gang Yu 0002, Chunhua Shen
ICCV6
2019 Age Estimation using Disconnectedness Features in Handwriting
abstract
Real-time applications of handwriting analysis have increased drastically in the fields of forensic and information security because of accurate cues. One of such applications is human age estimation based on handwriting for the purpose of immigrant checking. In this paper, we have proposed a new method for age estimation using handwriting analysis using Hu invariant moments and disconnectedness features. To make the proposed method robust to both ruled and un-ruled documents, we propose to explore intersection point detection in Canny edge images of each input document, which results in text components. For each text component pair, we propose Hu invariant moments for extracting disconnectedness features, which in fact measure multi-shape components based on distance, shape and mutual position analysis of components. Furthermore, iterative k-means clustering is proposed for the classification of different age groups. Experimental results on our dataset and some standard datasets, namely, IAM and KHATT, show that the proposed method is effective and outperforms the state-of-the-art methods.
V. Basavaraja, Palaiahnakote Shivakumara, D. S. Guru, Umapada Pal 0001, Tong Lu 0002, Michael Blumenstein
ICDAR5
2019 CRNN Based Jersey-Bib Number/Text Recognition in Sports and Marathon Images
abstract
The primary challenge in tracing the participants in sports and marathon video or images is to detect and localize the jersey/Bib number that may present in different regions of their outfit captured in cluttered environment conditions. In this work, we proposed a new framework based on detecting the human body parts such that both Jersey Bib number and text is localized reliably. To achieve this, the proposed method first detects and localize the human in a given image using Single Shot Multibox Detector (SSD). In the next step, different human body parts namely, Torso, Left Thigh, Right Thigh, that generally contain a Bib number or text region is automatically extracted. These detected individual parts are processed individually to detect the Jersey Bib number/text using a deep CNN network based on the 2-channel architecture based on the novel adaptive weighting loss function. Finally, the detected text is cropped out and fed to a CNN-RNN based deep model abbreviated as CRNN for recognizing jersey/Bib/text. Extensive experiments are carried out on the four different datasets including both bench-marking dataset and a new dataset. The performance of the proposed method is compared with the state-of-the-art methods on all four datasets that indicates the improved performance of the proposed method on all four datasets.
Sauradip Nag, Ramachandra Raghavendra, Palaiahnakote Shivakumara, Umapada Pal 0001, Tong Lu 0002, Mohan Kankanhalli
ICDAR5
2019 A Text-Context-Aware CNN Network for Multi-oriented and Multi-language Scene Text Detection
abstract
The existing deep learning based state-of-theart scene text detection methods treat scene texts a type of general objects, or segment text regions directly. The latter category achieves remarkable detection results on arbitrary orientation and large aspect ratios of scene texts based on instance segmentation algorithms. However, due to the lack of context information with consideration of scene text unique characteristics, directly applying instance segmentation to text detection task is prone to result in low accuracy, especially producing false positive detection results. To ease this problem, we propose a novel text-context-aware scene text detection CNN structure, which appropriately encodes channel and spatial attention information to construct context-aware and discriminative feature map for multi-oriented and multi-language text detection tasks. With high representation ability of text context-aware feature map, the proposed instance segmentation based method can not only robustly detect multi-oriented and multi-language text from natural scene images, but also produce better text detection results by greatly reducing false positives. Experiments on ICDAR2015 and ICDAR2017-MLT datasets show that the proposed method has achieved superior performances in precision, recall and F-measure than most of the existing studies.
Minglong Xue, Tong Lu 0002, Yirui Wu, Palaiahnakote Shivakumara
ICDAR3
2019 Hierarchical Bayesian Network Based Incremental Model for Flood Prediction
Yirui Wu, Weigang Xu, Qinghan Yu, Jun Feng 0001, Tong Lu 0002
MMM (1)5
2019 An Automatic System for Generating Artificial Fake Character Images
Yisheng Yue, Palaiahnakote Shivakumara, Yirui Wu, Tong Lu 0002, Umapada Pal 0001
MMM (2)5
2019 An automatic zone detection system for safe landing of UAVs
Maryam Asadzadeh Kaljahi, Palaiahnakote Shivakumara, Mohd Yamani Idna Bin Idris, Mohammad Hossein Anisi, Tong Lu 0002, Michael Blumenstein, Noorzaily Mohamed Noor
Expert Syst. Appl.5
2019 A novel character segmentation-reconstruction approach for license plate recognition
Vijeta Khare, Palaiahnakote Shivakumara, Chee Seng Chan, Tong Lu 0002, Kim Meng Liang, Hock Woon Hon, Michael Blumenstein
Expert Syst. Appl.4
2019 Fractional means based method for multi-oriented keyword spotting in video/scene/license plate images
Palaiahnakote Shivakumara, Sangheeta Roy, Hamid Abdullah Jalab, Rabha W. Ibrahim, Umapada Pal 0001, Tong Lu 0002, Vijeta Khare, Ainuddin Wahid Abdul Wahab
Expert Syst. Appl.6
2019 Curved text detection in blurred/non-blurred video/scene images
Minglong Xue, Palaiahnakote Shivakumara, Tong Lu 0002, Umapada Pal 0001
Multim. Tools Appl.4
2019 Multi-Script-Oriented Text Detection and Recognition in Video/Scene/Born Digital Images
abstract
Achieving good text detection and recognition results for multi-script-oriented images is a challenging task. First, we explore bit plane slicing in order to utilize the advantage of the most significant bit information to identify text components. A new iterative nearest neighbor symmetry is then proposed based on shapes of convex and concave deficiencies of text components in bit planes to identify candidate planes. Further, we introduce a new concept called mutual nearest neighbor pair components based on gradient direction to identify representative pairs of texts in each candidate bit plane. The representative pairs are used to restore words with the help of edge image of the input one, which results in text detection results (words). Second, we propose a new idea by fixing window for character components of arbitrary oriented words based on angular relationship between sub-bands and a fused band. For each window, we extract features in contourlet wavelet domain to detect characters with the help of an SVM classifier. Further, we propose to explore HMM for recognizing characters and words of any orientation using the same feature vector. The proposed method is evaluated on standard databases such as ICDAR, YVT video, ICDAR, SVT, MSRA scene data, ICDAR born digital data, and multi-lingual data to show its superiority to the state of the art methods.
Raghunandan K. Srinivas, Palaiahnakote Shivakumara, Sangheeta Roy, G. Hemantha Kumar 0001, Umapada Pal 0001, Tong Lu 0002
IEEE Trans. Circuits Syst. Video Technol.6
2018 New COLD Feature Based Handwriting Analysis for Enthnicity/Nationality Identification
abstract
Identifying crime for forensic investigating teams when crimes involve people of different nationals is challenging. This paper proposes a new method for ethnicity (nationality) identification based on Cloud of Line Distribution (COLD) features of handwriting components. The proposed method, at first, uses tangent angle of the contour pixels in each row and the mean of intensity values of each row for segmenting text lines. For segmented text lines, we use tangent angle and direction of base lines to remove rule lines in the image. We use polygonal approximation for finding dominant points for contours of edge components. Then the proposed method connects the nearest dominant points of every dominant point, which results in line segments of dominant point pairs. For each line segment, the proposed method estimates angle and length, which gives a point in polar domain. For all the line segments, the proposed method generates dense points in polar domain, which results in COLD distribution. As character component shapes change, according to nationals, the shape of the distribution changes. This observation is extracted based on distance from pixels of distribution to Principal Axis of the distribution. Then the features are subjected to an SVM classifier for identifying nationals. Experiments are conducted on a complex dataset, which show the proposed method is effective and outperforms the existing method.
Sauradip Nag, Palaiahnakote Shivakumara, Yirui Wu, Umapada Pal 0001, Tong Lu 0002
ICFHR5
2018 Adaptive Multi-Gradient Kernels for Handwritting Based Gender Identification
abstract
Handwriting based Gender identification is challenging due to unconstrained handwriting and individual differences in writing. To solve this problem, we propose a new adaptive multi-gradient of Sobel kernels for extracting Adaptive Multi-Gradient Features (AMGF). For extracted text lines, the proposed method finds dominant pixels based on directional symmetry of text pixels given by AMGF. We perform histogram operation for adaptive multi-gradient values extracted corresponding to dominant pixels. The gradient values that give the highest peak in respective histograms is chosen as features. This results in feature vector having four AMGF values. The same vector are generated for successive text lines in each image to study either consistency, which is expected for females or inconsistency, which is expected for males in writing styles. The correlation is estimated based on feature vectors of the first and the successive text lines until converging or diverging criteria is met. If convergence happens, the input document is considered as female else is considered as male. The method is tested on our own dataset, which includes large variations and standard datasets, namely, QUWI, IAM-1+IAM-2 and KHATT, to demonstrate the effectiveness of the proposed method. Experimental results show that the proposed method outperforms the existing methods.
B. J. Navya, Palaiahnakote Shivakumara, G. C. Swetha, Sangheeta Roy, D. S. Guru, Umapada Pal 0001, Tong Lu 0002
ICFHR7
2018 A New RGB Based Fusion for Forged IMEI Number Detection in Mobile Images
abstract
As technology advances to make living comfortable for people, at the same time, different crimes also increase. One such sensitive crime is creating fake International Mobile Equipment Identity (IMEI) for smart mobile devices. In this paper, we present a new fusion based method using R, G and B color components for detecting forged IMEI numbers. To the best of our knowledge, this is the first work for forged IMEI number detection in mobile images. The proposed method first finds variances for R, G and B images of a forged input image to study local changes. The variances are used to derive weights for respective color components. The same weights are convolved with respective pixel values of R, G and B components, which results in the fused image. For the fused image, the proposed method extracts features based on sparsity, the number of connected components, and the average intensity values for edge components in respective R, G and B components, which gives six features. The proposed method finds absolute difference between fused and input images, which gives feature vector containing six difference values. The proposed method constructs templates based on samples chosen randomly. Feature vectors are compared with the templates for detecting forged IMEI numbers. Experiments are conducted on our own dataset and standard datasets to evaluate the proposed method. Furthermore, comparative studies with the related existing methods show that the proposed method outperforms the existing methods.
Palaiahnakote Shivakumara, V. Basavaraja, Harsha S. Gowda, D. S. Guru, Umapada Pal 0001, Tong Lu 0002
ICFHR6
2018 End-To-End Chromosome Karyotyping with Data Augmentation Using GAN
abstract
Classifying human chromosomes from input cell images, i.e., karyotyping, requires domain expertise and quantity of manual effort to perform. In this paper, we propose an end-to-end chromosome karyotyping method, which can automatically detect, segment and classify chromosomes from cell images. During detection, we explore Extremal Regions (ER) to obtain chromosome candidates in input images. During segmentation, we segment overlapping chromosome candidates by approximating chromosome shapes with eclipses. In classification, we first propose Multiple Distribution Generative Advertising Network (MD-GAN) to effectively cover diverse data modes and generate more labeled samples for data augmentation. Then, we finetune pre-trained convolutional neural network (CNN) to classify chromosomes with samples generated by MD-GAN. We demonstrate the accuracy of the proposed end-to-end method in detecting, segmenting and classifying by experiments on a self-collected dataset. Experiments also prove data augmentation with MD-GAN could improve classification performance of CNN.
Yirui Wu, Yisheng Yue, Tong Lu 0002
ICIP5
2018 Weighted-Gradient Features for Handwritten Line Segmentation
abstract
Text line segmentation from handwritten documents is challenging when a document image contains severe touching. In this paper, we propose a new idea based on Weighted-Gradient Features (WGF) for segmenting text lines. The proposed method finds the number of zero crossing points for every row of Canny edge image of the input one, which is considered as the weights of respective rows. The weights are then multiplied with gradient values of respective rows of the image to widen the gap between pixels in the middle portion of text and the other portions. Next, k-means clustering is performed on WGF to classify middle and other pixels of text. The method performs morphological operation to obtain word components as patches for the result of clustering. The patches in both the clusters are matched to find common patch areas, which helps in reducing touching effect. Then the proposed method checks linearity and non-linearity iteratively based on patch direction to segment text lines. The method is tested on our own and standard datasets, namely, Alaei, ICDAR 2013 robust competition on handwriting context and ICDAR 2015-HTR, to evaluate the performance. Further, the method is compared with the state of art methods to show its effectiveness and usefulness.
Vijeta Khare, Palaiahnakote Shivakumara, B. J. Navya, G. C. Swetha, D. S. Guru, Umapada Pal 0001, Tong Lu 0002
ICPR7
2018 Multi-Gradient Directional Features for Gender Identification
abstract
Gender identification based on handwriting analysis has received a special attention to researchers in the field of document image analysis as it is useful for several real-time applications like forensic, population counting, etc. In this paper, we explore Multi-Gradient Directional (MGD) features, which provide direction of dominant pixels obtained by Canny edge image, and gradient direction symmetry. The proposed method further performs histogram operation for gradient angle information of dominant pixels of respective multi-gradient directional images to select angles, which contribute to the highest peak. This results in feature vectors. The process of feature vector formation continues for the segmented first, second, and third text lines in each image by male or female. Next, correlation is estimated for the vector of the first line with successive lines until converging or diverging criteria is met. If the convergence happens, a document is considered as by female, else is considered as by male. The method is tested on our own dataset, which includes images of different scripts, writers, papers, pens, and ages, and the standard database QUWI which includes Arabic and English texts, to demonstrate the efficiency of the proposed method. Comparative studies with the state of the art methods show that the proposed method is effective and useful.
B. J. Navya, G. C. Swetha, Palaiahnakote Shivakumara, Sangheeta Roy, D. S. Guru, Umapada Pal 0001, Tong Lu 0002
ICPR7
2018 Em-SLAM: a Fast and Robust Monocular SLAM Method for Embedded Systems
abstract
Simultaneous Localization and Mapping (SLAM) is difficult to deploy in the embedded systems due to its high computation cost and stable input requirements. Building on excellent algorithms of recent years, we present Em-SLAM, a monocular SLAM method which is fast and robust in the embedded system. We present Em-SLAM in three stages comprising initial pose estimation, iterative pose optimization and correspondences, and mapping with nearest frame queue. During the first stage, we perform stable initial pose estimation based on the matched ORB features extracted around the selected key points. Regarding initial pose and corresponding key points as input, the second stage of Em-SLAM iteratively optimizes these inputs values by tracking key points in the new frames. At the last stage, we firstly determine keyframes with the help of the proposed nearest frame queue and then design a greedy search algorithm to find matched ORB features between keyframes, which are adopted for compact and robust map reconstruction. Due to the special designs for the embedded systems, Em-SLAM demonstrates a high accurate and fast performance on the embedded system for all SLAM tasks: tracking, mapping and loop closing. We evaluate Em-SLAM on he most popular datasets by comparing with one latest SLAM method.
Yirui Wu, Zhikai Li, Palaiahnakote Shivakumara, Tong Lu 0002
ICPR4
2018 Context-Aware Attention LSTM Network for Flood Prediction
abstract
To minimize the negative impacts brought by floods, researchers from pattern recognition community utilize artificial intelligence based methods to solve the problem of flood prediction. Inspired by the significant power of Long Short-Term Memory (LSTM) networks in modeling the dynamics and dependencies of sequential data, we intend to utilize LSTM networks to predict sequential flow rate values based on a set of collected flood factors. Since not all factors are informative for flood prediction and the irrelevant factors often bring a lot of noise, we need to pay more attention to the informative ones. However, original LSTM doesn't have strong attention capability. Hence we propose an context-aware attention LSTM (CA-LSTM) network for flood prediction, which is capable to selectively focus on informative factors. During training, the local context-aware attention model is constructed by learning probability distributions between flow rate and hidden output of each LSTM cell. During testing, the learned local attention model assign weights to adjust relations between input factors and predictions at all steps of LSTM network. We conduct experiments on a flood dataset with several comparative methods to demonstrate high accuracy of the proposed method and the effectiveness of the proposed context-aware attention model.
Yirui Wu, Zhaoyang Liu 0001, Weigang Xu, Jun Feng 0001, Palaiahnakote Shivakumara, Tong Lu 0002
ICPR6
2018 Fourier Transform based Features for Clean and Polluted Water Image Classification
abstract
Water image classification is challenging because water images of ocean or river share the same properties with images of polluted water such as fungus, waste and rubbish. In this paper, we present a method for classifying clean and polluted water images. The proposed method explores Fourier transform based features for extracting texture properties of clean and polluted water images. Fourier spectrum of each input image is divided into several sub-regions based on angle and spatial information. For each region over the spectrum, the proposed method extracts mean and variance features using intensity values, which results in a feature matrix. The feature matrix is then passed to an SVM classifier for the classification of clean and polluted water images. Experimental results on classes of clean and polluted water images show that the proposed method is effective. Furthermore, a comparative study with the state-of-the-art method shows that the proposed method outperforms the existing method in terms of classification rate, recall, precision and F-measure.
Xuerong Wu, Palaiahnakote Shivakumara, Hualu Zhang, Tong Lu 0002, Umapada Pal 0001, Michael Blumenstein
ICPR6
2018 Local and Global Bayesian Network based Model for Flood Prediction
abstract
To minimize the negative impacts brought by floods, researchers from pattern recognition community pay special attention to the problem of flood prediction by involving technologies of machine learning. In this paper, we propose to construct hierarchical Bayesian network to predict floods for small rivers, which appropriately embed hydrology expert knowledge for high rationality and robustness. We present the construction of the hierarchical Bayesian network in two stages comprising local and global network construction. During the local network construction, we firstly divide the river watershed into small local regions. Following the idea of a famous hydrology model - the Xinanjiang model, we establish the entities and connections of the local Bayesian network to represent the variables and physical processes of the Xinanjiang model, respectively. During the global network construction, intermediate variables for local regions, computed by the local Bayesian network, are coupled to offer an estimation for time-varying values of flow rate by proper inferences of the global network. At last, we propose to improve the output of Bayesian network by utilizing former flow rate values. We demonstrate the accuracy and robustness of the proposed method by conducting experiments on a collected dataset with several comparative methods.
Yirui Wu, Weigang Xu, Jun Feng 0001, Palaiahnakote Shivakumara, Tong Lu 0002
ICPR5
2018 Mixed Link Networks
abstract
On the basis of the analysis by revealing the equivalence of modern networks, we find that both ResNet and DenseNet are essentially derived from the same "dense topology", yet they only differ in the form of connection: addition (dubbed "inner link") vs. concatenation (dubbed "outer link"). However, both forms of connections have the superiority and insufficiency. To combine their advantages and avoid certain limitations on representation learning, we present a highly efficient and modularized Mixed Link Network (MixNet) which is equipped with flexible inner link and outer link modules. Consequently, ResNet, DenseNet and Dual Path Network (DPN) can be regarded as a special case of MixNet, respectively. Furthermore, we demonstrate that MixNets can achieve superior efficiency in parameter over the state-of-the-art architectures on many competitive datasets like CIFAR-10/100, SVHN and ImageNet.
Wenhai Wang, Xiang Li 0041, Tong Lu 0002, Jian Yang 0003
IJCAI3
2018 Cloud of Line Distribution and Random Forest Based Text Detection from Natural/Video Scene Images
Wenhai Wang, Yirui Wu, Palaiahnakote Shivakumara, Tong Lu 0002
MMM (2)4
2018 A Novel 3D Human Action Recognition Framework for Video Content Analysis
Lianglei Wei, Yirui Wu, Wenhai Wang, Tong Lu 0002
MMM (1)4
2018 Rough-fuzzy based scene categorization for text detection and recognition in video
Sangheeta Roy, Palaiahnakote Shivakumara, Namita Jain, Vijeta Khare, Anjan Dutta 0001, Umapada Pal 0001, Tong Lu 0002
Pattern Recognit.7
2018 Riesz Fractional Based Model for Enhancing License Plate Detection and Recognition
abstract
One of the major causes of poor results in license plate recognition is low quality of images affected by multiple factors, such as severe illumination condition, complex background, different weather conditions, night light, and perspective distortions. In this paper, we propose a new mathematical model based on Riesz fractional operator for enhancing details of edge information in license plate images to improve the performances of text detection and recognition methods. The proposed model performs convolution operation of the Riesz fractional derivative over each input image by enhancing the edge strength in it. To test the performance of the proposed model, we conduct experiments on benchmark license plate image databases, namely, UCSD and ICDAR 2015-SR competition text image databases. Experimental results on enhancement show that the proposed model outperforms the existing baseline enhancement techniques in terms of quality measures. Furthermore, experimental results on text detection and recognition show that text detection and recognition rates are improved significantly after enhancement compared with before enhancement.
Raghunandan K. Srinivas, Palaiahnakote Shivakumara, Hamid Abdullah Jalab, Rabha W. Ibrahim, G. Hemantha Kumar 0001, Umapada Pal 0001, Tong Lu 0002
IEEE Trans. Circuits Syst. Video Technol.7
2017 New Fuzzy-Mass Based Features for Video Image Type Categorization
abstract
Due to the large variety of video type collections, it becomes difficult to achieve good text detection and recognition accuracy. We propose a new fuzzy-mass based method for classifying (categorizing) text frames from different types of video. For each frame of a video type, we formulate Fuzzy logic to identify straight and curved edge components from edge images. We then estimate mass locally and globally by drawing consecutive ellipses over edge images with respect to straight and curved edge components. Further, we extract features based on spatial proximity between centroid of classified straight/curved edge components and that of the whole image. This results local features. Next, the features are extracted for the whole image without ellipse drawing, which results in global features. The combination of both local and global features is then fed to an SVM classifier for video type classification. Experimental results on the proposed and existing classification methods show that the proposed classification outperforms the stat of art methods. Furthermore, experiments on before and after classification with several text detection and binarization methods show that the proposed classification is significant in improving text detection and recognition performance.
Sangheeta Roy, Palaiahnakote Shivakumara, Namita Jain, Vijeta Khare, Umapada Pal 0001, Tong Lu 0002
ICDAR6
2017 Temporal Integration for Word-Wise Caption and Scene Text Identification
abstract
Generally video consists of edited text (i.e., caption text) and natural text (i.e., scene text), and these two texts differ from one another in nature as well as characteristics. Such different behaviors of caption and scene texts lead to poor accuracy for text recognition in video. In this paper, we explore wavelet decomposition and temporal coherency for the classification of caption and scene text. We propose wavelet of high frequency sub-bands to separate text candidates that are represented by high frequency coefficients in an input word. The proposed method studies the distribution of text candidates over word images based on the fact that the standard deviation of text candidates is high at the first zone, low at the middle zone and high at the third zone. This is extracted by mapping standard deviation values to 8 equal sized bins formed based on the range of standard deviation values. The correlation among bins at the first and second levels of wavelets is explored to differentiate caption and scene text and for determining the number of temporal frames to be analyzed. The properties of caption and scene texts are validated with the chosen temporal frames to find the stable property for classification. Experimental results on three standard datasets (ICDAR 2015, YVT and License Plate Video) show that the proposed method outperforms the existing methods in terms of classification rate and improves recognition rate significantly based on classification results.
Sangheeta Roy, Palaiahnakote Shivakumara, Umapada Pal 0001, Tong Lu 0002, Ainuddin Wahid Abdul Wahab
ICDAR4
2017 Fourier-Residual for Printer Identification
abstract
Printer identification is challenging due to advanced software technologies in the field of forgery detection. This paper presents a new idea of using the Fourier transform residual for the identification of documents printed by different printers. The proposed approach first convolves a Laplacian mask with a Fourier transform in the frequency domain to smoothen the edges. Next, we apply an inverse Fourier transform to reconstruct images from smoothed information (RFL). Similarly, the proposed approach reconstructs images using gray information of the input image (RFG). Then the residual is calculated by subtracting RFG from RFL. The set of statistical features, texture and spatial features are extracted from residual images for printer identification. Experimental results with the existing method on our dataset and a standard dataset show that the proposed approach outperforms the existing approach on both the datasets in terms of classification rate, recall, precision and F-measure.
Palaiahnakote Shivakumara, Tong Lu 0002, M. Basavanna, Umapada Pal 0001, Michael Blumenstein
ICDAR3
2017 A Robust Symmetry-Based Method for Scene/Video Text Detection through Neural Network
abstract
Text detection in video/scene images has gained a significant attention in the field of image processing and document analysis due to the inherent challenges caused by variations in contrast, orientation, background, text type, font type, non-uniform illumination and so on. In this paper, we propose a novel text detection method to explore symmetry property and appearance features of text for improved accuracy and robustness. First, the proposed method explores Extremal Regions (ER) for detecting text candidates in images. Then we propose a novel feature named as Multi-domain Strokes Symmetry Histogram (MSSH) for each text candidate, which describes the inherent symmetry property of stroke pixel pairs in gray, gradient and frequency domains. Furthermore, deep convolutional features are extracted to describe the appearance for each text candidate. We further fuse them by Auto-Encoder network to define a more discriminative text descriptor for classification. Finally, the proposed method constructs text lines based on the classification results. We demonstrate the effectiveness and robustness detection results of our proposed method by testing on four different benchmark databases.
Yirui Wu, Wenhai Wang, Palaiahnakote Shivakumara, Tong Lu 0002
ICDAR4
2017 Deep-dense Conditional Random Fields for Object Co-segmentation
abstract
We address the problem of object co-segmentation in images. Object co-segmentation aims to segment common objects in images and has promising applications in AI agents. We solve it by proposing a co-occurrence map, which measures how likely an image region belongs to an object and also appears in other images. The co-occurrence map of an image is calculated by combining two parts: objectness scores of image regions and similarity evidences from object proposals across images. We introduce a deep-dense conditional random field framework to infer co-occurrence maps. Both similarity metric and objectness measure are learned end-to-end in a single deep network. We evaluate our method on two benchmarks and achieve competitive performance.
Ze-Huan Yuan, Tong Lu 0002, Yirui Wu
IJCAI2
2017 Robust Scene Text Detection for Multi-script Languages Using Deep Learning
Ruo-Ze Liu, Xin Sun 0009, Hailiang Xu, Palaiahnakote Shivakumara, Feng Su, Tong Lu 0002, Ruoyu Yang
MMM (1)6
2017 Visual Robotic Object Grasping Through Combining RGB-D Data and 3D Meshes
Yiyang Zhou, Wenhai Wang, Wenjie Guan, Yirui Wu, Heng Lai, Tong Lu 0002, Min Cai
MMM (1)6
2017 Script independent approach for multi-oriented text detection in scene image
Sounak Dey, Palaiahnakote Shivakumara, Raghunandan K. Srinivas, Umapada Pal 0001, Tong Lu 0002, G. Hemantha Kumar 0001, Chee Seng Chan
Neurocomputing5
2017 A new multi-modal approach to bib number/text detection and recognition in Marathon images
Palaiahnakote Shivakumara, Ramachandra Raghavendra, Longfei Qin, Kiran B. Raja, Tong Lu 0002, Umapada Pal 0001
Pattern Recognit.5
2017 Fractals based multi-oriented text detection system for recognition in mobile video images
Palaiahnakote Shivakumara, Liang Wu 0009, Tong Lu 0002, Chew Lim Tan, Michael Blumenstein, Basavaraj S. Anami
Pattern Recognit.3
2017 Learning discriminated and correlated patches for multi-view object detection using sparse coding
Ze-Huan Yuan, Tong Lu 0002, Chew Lim Tan
Pattern Recognit.2
2017 FreeScup: A Novel Platform for Assisting Sculpture Pose Design
abstract
Sculpture design is challenging due to its inherent difficulty in characterizing artworks quantitatively; thus, few works have been done to assist sculpture design in the past decades in the multimedia community. We have cooperated with several sculptors on analyzing styles of different artists consisting of Giacometti, Augeuste Rodin, Henry Moore, and Marino Marini from which we find pose editing plays an important role in sculpture design. Motivated by this, we present a novel platform that allows sculptors to edit virtual three-dimensional (3-D) sculptures by a free way. The proposed platform consists of three modules, namely,sculpture initialization,sculptor-sculpture mapping, andinteractive pose editing. In sculpture initialization, a virtual 3-D sculpture is first incrementally reconstructed from multiview images. Then, we define Laplace operator and its corresponding spectrum to describe the geometry information of the reconstructed sculpture. During sculptor–sculpture mapping, we apply spectral analysis on the low-frequency parts of the spectrum to search for candidate editing points on the surface of the sculpture. Next, body actions of the sculptor are captured by Kinect and further mapped onto editing points as a predefined configuration set. Finally, during interactive pose editing, a real-time Kinect-driven sculpture pose editing scheme is presented, which not only preserves geometry features of the sculpture but also allows instant changes of sculpture poses. We demonstrate that our platform successfully assists sculptors on real-time pose editing by comparing its performance with those of the existing sculpture assisting methods.
Yirui Wu, Tong Lu 0002, Ze-Huan Yuan
IEEE Trans. Multim.2
2016 New Sharpness Features for Image Type Classification Based on Textual Information
abstract
Achieving good recognition results from a single method for text lines in video/natural scene images captured by high resolution cameras or low resolution mobile cameras, and images in web pages, is often hard. In this paper, we propose new sharpness based features of textual portion of each input text line image using HSI color space for the classification of an input image into one of the four classes (video, scene, mobile or born digital). This helps in choosing an appropriate method based on the class type of the input text for its improved recognition rate. For a given input text line image, the proposed method obtains H, S and I images. Then Canny edge images are obtained for H, S and I spaces, which results in text candidates. We perform sliding window operation over the text candidate image of each text line of each color space to estimate new sharpness by calculating stroke width and gradient information. The sharpness values of the text lines of the three color spaces are then fed to k-means clustering with maximum, minimum and average guesses, which results in three respective clusters. The mean of each cluster for respective color spaces outputs a feature vector having nine feature values for image classification with the help of an SVM classifier. Experimental results on standard datasets, namely, ICDAR 2013, ICDAR 2015 video, ICDAR 2015 natural scene data, ICDAR 2013 born digital data and the images captured by a mobile camera (our own data) show that the proposed classification method helps in improving recognition results.
Raghunandan K. Srinivas, Palaiahnakote Shivakumara, G. Hemantha Kumar 0001, Umapada Pal 0001, Tong Lu 0002
DAS5
2016 Fourier Coefficients for Fraud Handwritten Document Classification through Age Analysis
abstract
As new digital technologies emerge to improve living style, at the same time, it also lead to increase crimes. Unlike existing approaches that use content of handwriting for fraud/forged document identification, in this paper we propose a novel approach that explores the quality of handwritten documents by considering both foreground and background information to identify whether it is old or new. The proposed approach works based on the fact that if a fraud document is created with some gaps after the original one, the fraud document happened to be a new one and the original happened to be an old one in this work. To identify whether a given handwritten document is old or new with gaps, we propose to divide Fourier coefficients of the input image into positive and negative coefficient images, and then reconstruct respective images to conquer two reconstructed ones. The contrast of the reconstructed images obtained before and after divide-conquer is studied to analyze the ages of the document based on image quality. The proposed approach finds a unique relationship between reconstructed images, obtained before and after divide-conquer, to identify the input image as old or new. To evaluate the proposed approach, we conduct experiments on our own handwritten dataset and a standard database, namely, Google-LIFE magazine. Comparative studies with the existing approaches show that the proposed approach outperforms the existing approaches in terms of classification rate.
Raghunandan K. Srinivas, Palaiahnakote Shivakumara, B. J. Navya, G. Pooja, Navya Prakash, G. Hemantha Kumar 0001, Umapada Pal 0001, Tong Lu 0002
ICFHR8
2016 New Tampered Features for Scene and Caption Text Classification in Video Frame
abstract
The presence of both caption/graphics/superimposed and scene texts in video frames is the major cause for the poor accuracy of text recognition methods. This paper proposes an approach for identifying tampered information by analyzing the spatial distribution of DCT coefficients in a new way for classifying caption and scene text. Since caption text is edited/superimposed, which results in artificially created texts comparing to scene texts that exist naturally in frames. We exploit this fact to identify the presence of caption and scene texts in video frames based on the advantage of DCT coefficients. The proposed method analyzes the distributions of both zero and non-zero coefficients (only positive values) locally by moving a window, and studies histogram operations over each input text line image. This generates line graphs for respective zero and non-zero coefficient coordinates. We further study the behavior of text lines, namely, linearity and smoothness based on centroid location analysis, and the principal axis direction of each text line for classification. Experimental results on standard datasets, namely, ICDAR 2013 video, 2015 video, YVT video and our own data, show that the performances of text recognition methods are improved significantly after-classification compared to before-classification.
Sangheeta Roy, Palaiahnakote Shivakumara, Umapada Pal 0001, Tong Lu 0002, Chew Lim Tan
ICFHR4
2016 A quad tree based method for blurred and non-blurred video text frames classification through quality metrics
abstract
Blur is a common artifact in video, which adds more complexity to text detection and recognition. To achieve good accuracies for text detection and recognition, this paper suggests a new method for classifying blurred and non-blurred frames in video. We explore quality metrics, namely, BRISQUE, NRIQA, GPC and SI, in a new way for classification. We estimate the values of these metrics with the help of predefined samples called reference values. To widen the difference between metric values for better classification, we introduce scaling factors as a non-linear sigmoidal function, which considers the metric of each current frame and its reference and results in templates. Based on the characteristics of metrics, the proposed method finds a relationship between the metrics to derive rules for classification. To classify the frame containing local blur, we explore quad tree division with classification rules which divide non-blurred blocks to identify local blur. We use standard databases, namely, ICDAR 2013, ICDAR 2015 and YVT videos for experimentation, and evaluate the proposed method in terms of text detection and recognition rates given by text detection and binarization methods before and after classification.
Vijeta Khare, Palaiahnakote Shivakumara, Ahlad Kumar, Chee Seng Chan, Tong Lu 0002, Michael Blumenstein
ICPR5
2016 Video scene text frames categorization for text detection and recognition
abstract
Developing a unified text detection and recognition method is hard for different video types due to varying characteristics in video. This paper proposes a new method for categorizing different types of video text frames, namely, videos containing advertisement, signboard, license plate, front page of book or magazine, street view, and video of general items, for better text detection and recognition rate. We propose symmetry features using gradient vector flow for Canny and Sobel edge images of each input frame to identify candidate edge components. Then for a candidate edge component image, we extract both global and local features using colors from different channels in a new way. Besides, the proposed method extracts statistical and structural features from the spatial distribution of candidate pixels in a multi-scale environment. Lastly, the extracted features are fed to a logistic classifier for categorization. The features extracted locally and globally are tested both separately and altogether in terms of confusion matrix. The performance of the proposed categorization method is evaluated through several text detection and recognition experiments before and after categorization. We noted that the proposed categorization method is very useful in improving text detection and recognition performance.
Longfei Qin, Palaiahnakote Shivakumara, Tong Lu 0002, Umapada Pal 0001, Chew Lim Tan
ICPR3
2016 EvaToon: A novel graph matching system for evaluating cartoon drawings
abstract
Imitation cartoon drawing is an important skill for cartoonists, requiring quantity of efforts on practising and guidance. In this paper, we propose EvaToon, an imitated drawing evaluate system, which automatically assigns judging scores and marks improper drawing regions. With our system, cartoonists can practise and get guidance by themselves. We have cooperated with several experts on developing such an evaluation system. Based on their guide, we present EvaToon in two stages comprising cartoon drawings analyzing and similarity evaluating. During analyzing, we first locate contour pixels with high curvature as interest points and then extract multi-scale features around interest points to hierarchically describe shape. During evaluating, we first match interest points between original and imitated drawing based on distance of features. After matching, we construct a regression tree to map high dimensional difference of matching features to scores and marks based on quantity of manually evaluated training examples. Finally, our system matches an input imitated drawing with the original one and predicts its scores automatically. We demonstrate the accuracy of our EvaToon system in matching and predicting and prove the capability of describing shape of our proposed features by experiments on a collected dataset of imitated drawings.
Yirui Wu, Xianli Zhou, Tong Lu 0002, Guo Mei, Linbi Sun
ICPR3
2016 Modeling spatial layout for scene image understanding via a novel multiscale sum-product network
Ze-Huan Yuan, Limin Wang 0002, Tong Lu 0002, Palaiahnakote Shivakumara, Chew Lim Tan
Expert Syst. Appl.4
2016 Weakly-supervised region annotation for understanding scene images
Tong Lu 0002, Palaiahnakote Shivakumara, Chew Lim Tan
Multim. Tools Appl.2
2016 Fractional poisson enhancement model for text detection and recognition in video frames
Sangheeta Roy, Palaiahnakote Shivakumara, Hamid Abdullah Jalab, Rabha W. Ibrahim, Umapada Pal 0001, Tong Lu 0002
Pattern Recognit.6
2016 A new method for multi-oriented graphics-scene-3D text classification in video
Jiamin Xu, Palaiahnakote Shivakumara, Tong Lu 0002, Chew Lim Tan, Seiichi Uchida
Pattern Recognit.3
2016 Contour Restoration of Text Components for Recognition in Video/Scene Images
abstract
Text recognition in video/natural scene images has gained significant attention in the field of image processing in many computer vision applications, which is much more challenging than recognition in plain background images. In this paper, we aim to restore complete character contours in video/scene images from gray values, in contrast to the conventional techniques that consider edge images/binary information as inputs for text detection and recognition. We explore and utilize the strengths of zero crossing points given by the Laplacian to identify stroke candidate pixels (SPC). For each SPC pair, we propose new symmetry features based on gradient magnitude and Fourier phase angles to identify probable stroke candidate pairs (PSCP). The same symmetry properties are proposed at the PSCP level to choose seed stroke candidate pairs (SSCP). Finally, an iterative algorithm is proposed for SSCP to restore complete character contours. Experimental results on benchmark databases, namely, the ICDAR family of video and natural scenes, Street View Data, and MSRA data sets, show that the proposed technique outperforms the existing techniques in terms of both quality measures and recognition rate. We also show that character contour restoration is effective for text detection in video and natural scene images.
Yirui Wu, Palaiahnakote Shivakumara, Tong Lu 0002, Chew Lim Tan, Michael Blumenstein, G. Hemantha Kumar 0001
IEEE Trans. Image Process.3
2015 A new method based on bag of filters for character recognition in scene images by learning
abstract
Achieving a good recognition rate for scene characters is a big challenge due to non-uniform illumination effects, perspective distortions, multiple colors or contrasts, different fonts and their various sizes, background or orientation variations, etc. Unlike the existing recognition methods that use binary information or the features extracted from different domains, the proposed method explores gray information in the form of a filter bank to extract the discriminative power for all the 62 scene character classes. We propose a sliding window (patch) operation over a character image for learning the global features, which represent the structures of character images of all the classes by reconstructing a filter bank from the original data. We introduce shareable constrains to activate class-specific filters from the filter bank. Further, we propose constraints by studying the nearest neighbor patches and exemplar selection to maximize the gap between inter-classes and minimize the gap between intra-classes. The method is evaluated and compared with several existing recognition methods in terms of character recognition rate. Experimental results show that the proposed method outperforms the existing methods.
Qisu Li, Tong Lu 0002, Palaiahnakote Shivakumara, Umapada Pal 0001, Chew Lim Tan
ICDAR2
2015 A new wavelet-Laplacian method for arbitrarily-oriented character segmentation in video text lines
abstract
Character segmentation is an important topic to improve the overall performance of text recognition methods due to low resolution, complex background and lots of visual variations in video. This paper presents a novel idea for segmenting characters from arbitrarily-oriented text lines based on wavelet and Laplacian combination. Firstly, we explore wavelet which decomposes a given input image into sub-levels like a pyramid structure for segmenting words based on the fact that as decomposition level increases, the gap between characters decreases due to the reduction in the size of the input image, which results in a single component for each word. Secondly, for each segmented word, we propose Laplacian wavelet combination in a new way to extract text candidates. Thirdly, we propose horizontal and vertical sampling for character segmentation from words. The proposed method is tested on curved, non-horizontal and horizontal text lines of video and the ICDAR 2005 natural scene dataset to evaluate its performance. A comparative study with an existing method shows that the proposed method outperforms it in terms of precision and f-measure.
Guozhu Liang, Palaiahnakote Shivakumara, Tong Lu 0002, Chew Lim Tan
ICDAR3
2015 FreeScup: A novel platform for assisting sculpture pose design
abstract
Sculpture design is challenging due to its inherent difficulty in characterizing an artwork quantitatively, and few works have been done to assist sculpture design. We present a novel platform to help sculptors in two stages, comprising automatic sculpture reconstruction and free spectral-based sculpture pose editing. During sculpture reconstruction, we co-segment a sculpture from real scene images of different views through a two-label MRF framework, aiming at performing sculpture reconstruction efficiently. During sculpture pose editing, we automatically extract candidate editing points on the sculpture by searching in the spectrums of Laplacian operator. After manually mapping body joints of a sculptor to particular editing points, we further construct a global Laplacian-based linear system by adopting the spectrums of Laplacian operator and using Kinect captured body motions for real time pose editing. The constructed system thus allows the sculptor to freely edit different kinds of sculpture artworks through Kinect. Experimental results demonstrate that our platform successfully assists sculptors in real-time pose editing.
Yirui Wu, Tong Lu 0002, Ze-Huan Yuan
ICME2
2015 HIRM: A handle-independent reduced model for incremental mesh editing
Yirui Wu, Oscar Kin-Chung Au, Chiew-Lan Tai, Tong Lu 0002
Comput. Aided Geom. Des.4
2015 New Gradient-Spatial-Structural Features for video script identification
Palaiahnakote Shivakumara, Ze-Huan Yuan, Danni Zhao, Tong Lu 0002, Chew Lim Tan
Comput. Vis. Image Underst.4
2015 Bayesian classifier for multi-oriented video text recognition system
Sangheeta Roy, Palaiahnakote Shivakumara, Partha Pratim Roy 0001, Umapada Pal 0001, Chew Lim Tan, Tong Lu 0002
Expert Syst. Appl.6
2015 A new ring radius transform-based thinning method for multi-oriented video characters
Yirui Wu, Palaiahnakote Shivakumara, Tong Lu 0002, Umapada Pal 0001
Int. J. Document Anal. Recognit.4
2015 Character shape restoration system through medial axis points in video
Shangxuan Tian, Palaiahnakote Shivakumara, Trung Quy Phan, Tong Lu 0002, Chew Lim Tan
Neurocomputing4
2015 Content-oriented multimedia document understanding through cross-media correlation
Tong Lu 0002, Yukang Jin, Feng Su, Palaiahnakote Shivakumara, Chew Lim Tan
Multim. Tools Appl.1
2015 Multi-Spectral Fusion Based Approach for Arbitrarily Oriented Scene Text Detection in Video Images
abstract
Scene text detection from video as well as natural scene images is challenging due to the variations in background, contrast, text type, font type, font size, and so on. Besides, arbitrary orientations of texts with multi-scripts add more complexity to the problem. The proposed approach introduces a new idea of convolving Laplacian with wavelet sub-bands at different levels in the frequency domain for enhancing low resolution text pixels. Then, the results obtained from different sub-bands (spectral) are fused for detecting candidate text pixels. We explore maxima stable extreme regions along with stroke width transform for detecting candidate text regions. Text alignment is done based on the distance between the nearest neighbor clusters of candidate text regions. In addition, the approach presents a new symmetry driven nearest neighbor for restoring full text lines. We conduct experiments on our collected video data as well as several benchmark data sets, such as ICDAR 2011, ICDAR 2013, and MSRA-TD500 to evaluate the proposed method. The proposed approach is compared with the state-of-the-art methods to show its superiority to the existing methods.
Guozhu Liang, Palaiahnakote Shivakumara, Tong Lu 0002, Chew Lim Tan
IEEE Trans. Image Process.3
2015 A New Technique for Multi-Oriented Scene Text Line Detection and Tracking in Video
abstract
Text detection and tracking in video is challenging due to contrast, resolution and background variations, and different orientations and text movements. In addition, the presence of both caption and scene texts in video aggravates the problem because these two text types differ in characteristics significantly . This paper proposes a new technique for detecting and tracking video texts of any orientation by using spatial and temporal information, respectively. The technique explores gradient directional symmetry at component level for smoothing edge components before text detection. Spatial information is preserved by forming Delaunay triangulation in a novel way at this level, which results in text candidates. Text characteristics are then proposed in a different way for eliminating false text candidates , which results in potential text candidates. Then grouping is proposed for combining potential text candidates regardless of orientation based on the nearest neighbor criterion. To tackle the problems of multi-font and multi-sized texts, we propose multi-scale integration by a pyramid structure, which helps in extracting full text lines. Then, the detected text lines are tracked in video by matching the subgraphs of triangulation. Experimental results for text detection and tracking on our video dataset, the benchmark video datasets, and the natural scene image benchmark datasets show that the proposed method is superior to the state-of-the-art methods in terms of recall, precision , and F-measure.
Liang Wu 0009, Palaiahnakote Shivakumara, Tong Lu 0002, Chew Lim Tan
IEEE Trans. Multim.3
2014 Text Detection Using Delaunay Triangulation in Video Sequence
abstract
Text detection and tracking in video sequence is gaining interest due to the challenges posed by low resolution and complex background. This paper proposes a new method for text detection by estimating trajectories between the corners of texts in video sequence over time. Each trajectory is considered as one node to form a graph for all trajectories and Delaunay triangulation is used to obtain edges to connect nodes of the graph. In order to identify the edges that represent text regions, we propose four pruning criteria based on spatial proximity, motion coherence, local appearance and canny rate. This results in several sub-graphs. Then we use depth first search to collect corner points, which essentially represent text candidates. False positives are eliminated using heuristics and missing trajectories will be obtained by tracking the corners in temporal frames. We test the method on different videos and evaluate the method in terms of recall, precision, f-measure with existing results. Experimental result shows that the proposed method is superior to existing method.
Liang Wu 0009, Palaiahnakote Shivakumara, Tong Lu 0002, Chew Lim Tan
Document Analysis Systems3
2014 A Novel Topic-Level Random Walk Framework for Scene Image Co-segmentation
Ze-Huan Yuan, Tong Lu 0002, Palaiahnakote Shivakumara
ECCV (1)2
2014 Optical flow based dynamic curved video text detection
abstract
Text detection in video is a challenging problem as it is useful in several real time applications in the field of video indexing and retrieval. Unlike existing methods that generally focus on horizontal caption or graphics text, the proposed method focuses on detecting dynamic curved text in video. The method explores the characteristics of the optical flow of text, namely, constant velocity, uniform magnitude distribution and unique angle distribution, to identify text candidates with the help of k-means clustering algorithm. We propose an iterative procedure which finds the standard deviation of text candidates between the first and its successive frames, and it terminates when there is a sudden decrease in the standard deviation values. The proposed method eliminates false text candidates based on the characteristics of optical flow at component level while retaining the potential text candidates. Then, direction guided boundary growing is proposed to traverse curved text lines in video. Furthermore, the characteristics of optical flow of text are utilized at block level to eliminate false positives. Experiments are conducted with various videos, including video with static text, static and dynamic text, and dynamic text only, to evaluate the proposed method. The results are benchmarked with the existing methods to verify the superiority of our method over the existing methods in terms of recall, precision, F-measure and average processing time.
Palaiahnakote Shivakumara, Mohamed Lubani, Koksheik Wong, Tong Lu 0002
ICIP4
2014 Anomaly Detection through Spatio-temporal Context Modeling in Crowded Scenes
abstract
A novel statistical framework for modeling the intrinsic structure of crowded scenes and detecting abnormal activities is presented in this paper. The proposed framework essentially turns the anomaly detection process into two parts, namely, motion pattern representation and crowded context modeling. During the first stage, we averagely divide the spatiotemporal volume into atomic blocks. Considering the fact that mutual interference of several human body parts potentially happen in the same block, we propose an atomic motion pattern representation using the Gaussian Mixture Model (GMM) to distinguish the motions inside each block in a refined way. Usual motion patterns can thus be defined as a certain type of steady motion activities appearing at specific scene positions. During the second stage, we further use the Markov Random Field (MRF) model to characterize the joint label distributions over all the adjacent local motion patterns inside the same crowded scene, aiming at modeling the severely occluded situations in a crowded scene accurately. By combining the determinations from the two stages, a weighted scheme is proposed to automatically detect anomaly events from crowded scenes. The experimental results on several different outdoor and indoor crowded scenes illustrate the effectiveness of the proposed algorithm.
Tong Lu 0002, Liang Wu 0009, Xiaolin Ma, Palaiahnakote Shivakumara, Chew Lim Tan
ICPR1
2014 Graphics and Scene Text Classification in Video
abstract
Achieving good accuracy for text detection and recognition is a challenging and interesting problem in the field of video document analysis because of the presences of both graphics text that has good clarity and scene text that is unpredictable in video frames. Therefore, in this paper, we present a novel method for classifying graphics texts and scene texts by exploiting temporal information and finding the relationship between them in video. The method proposes an iterative procedure to identify Probable Graphics Text Candidates (PGTC) and Probable Scene Text Candidates (PSTC) in video based on the fact that graphics texts in general do not have large movements especially compared to scene texts which are usually embedded on background. In addition to PGTC and PSTC, the iterative process automatically identifies the number of video frames with the help of a converging criterion. The method further explores the symmetry between intra and inter character components to identify graphics text candidates and scene text candidates. Boundary growing method is employed to restore the complete text line. For each segmented text line, we finally introduce Eigen value analysis to classify graphics and scene text lines based on the distribution of respective Eigen values. Experimental results with the existing methods show that the proposed method is effective and useful to improve the accuracy of text detection and recognition.
Jiamin Xu, Palaiahnakote Shivakumara, Tong Lu 0002, Trung Quy Phan, Chew Lim Tan
ICPR3
2014 2D and 3D Video Scene Text Classification
abstract
Text detection and recognition is a challenging problem in document analysis due to the presence of the unpredictable nature of video texts, such as the variations of orientation, font and size, illumination effects, and even different 2D/3D text shadows. In this paper, we propose a novel horizontal and vertical symmetry feature by calculating the gradient direction and the gradient magnitude of each text candidate, which results in Potential Text Candidates (PTCs) after applying the k-means clustering algorithm on the gradient image of each input frame. To verify PTCs, we explore temporal information of video by proposing an iterative process that continuously verifies the PTCs of the first frame and the successive frames, until the process meets the converging criterion. This outputs Stable Potential Text Candidates (SPTCs). For each SPTC, the method obtains text representatives with the help of the edge image of the input frame. Then for each text representative, we divide it into four quadrants and check a new Mutual Nearest Neighbor Symmetry (MNNS) based on the dominant stroke width distances of the four quadrants. A voting method is finally proposed to classify each text block as either 2D or 3D by counting the text representatives that satisfy MNNS. Experimental results on classifying 2D and 3D text images are promising, and the results are further validated by text detection and recognition before classification and after classification with the exiting methods, respectively.
Jiamin Xu, Palaiahnakote Shivakumara, Tong Lu 0002, Chew Lim Tan
ICPR3
2014 Spectral 3D mesh segmentation with a novel single segmentation field
Tong Lu 0002, Oscar Kin-Chung Au, Chiew-Lan Tai
Graph. Model.2
2013 Recognition of Video Text through Temporal Integration
abstract
This paper presents a method for temporal integration, which can be used to improve the recognition accuracy of video texts. Given a word detected in a video frame, we use a combination of Stroke Width Transform and SIFT (Scale Invariant Feature Transform) to track it both backward and forward in time. The text instances within the word's frame span are then extracted and aligned at pixel level. In the second step, we integrate these instances into a text probability map. By thresholding this map, we obtain an initial binarization of the word. In the final step, the shapes of the characters are refined using the intensity values. This helps to preserve the distinctive character features (e.g., sharp edges and holes), which are useful for OCR engines to distinguish between the different character classes. Experiments on English and German videos show that the proposed method outperforms existing ones in terms of recognition accuracy.
Trung Quy Phan, Palaiahnakote Shivakumara, Tong Lu 0002, Chew Lim Tan
ICDAR3
2013 Online stroke segmentation by quick penalty-based dynamic programming
abstract
A stroke segmentation method named quick penalty‐based dynamic programming is proposed for splitting a sketchy stroke into several regular primitive shapes, such as line segments and elliptical arcs. The authors extend the dynamic programming framework with a customisable penalty function, which measures the correctness of splitting a stroke at a particular point. With the help of the penalty function, the proposed dynamic programming framework can finish the stroke segmentation process without any prior knowledge of the number and/or the type of segments contained in the sketchy stroke. Its response time is sufficiently short for online applications, even for long strokes. Experiments show that the proposed method is robust for strokes with arbitrary shape and size.
Wenyin Liu, Tong Lu 0002, Yajie Yu, Shuang Liang 0001, Rui Zhang 0031
IET Comput. Vis.2
2011 Multiclass object detection by combining local appearances and context
abstract
In this paper, we present a novel approach for multiclass object detection by combining local appearances and contextual constraints. We first construct a multiclass Hough forest of local patches, which can well deal with multiclass object deformations and local appearance variations, due to randomization and discrimination of the forest. Then, in the object hypothesis space, a new multiclass context model is proposed to capture relative location constraints, disambiguating appearance inputs in multiclass object detection. Finally, multiclass objects are detected with a greedy search algorithm efficiently. Experimental evaluations on two image data sets show that the combination of local appearances and context achieves state-of-the-art performance in multiclass object detection.
Limin Wang 0002, Yirui Wu, Tong Lu 0002
ACM Multimedia3
2010 3D Model Comparison through Kernel Density Matching
abstract
A novel 3D shape matching method is proposed in this paper. We first extract angular and distance feature pairs from pre-processed 3D models, then estimate their kernel densities after quantifying the feature pairs into a fixed number of bins. During 3D matching, we adopt the KL-divergence as a distance of 3D comparison. Experimental results show that our method is effective to match similar 3D shapes, and robust to model deformations or rotation transformations.
Tong Lu 0002, Rongjun Gao, Wenyin Liu
ICPR2
2009 A Novel Knowledge-Based System for Interpreting Complex Engineering Drawings: Theory, Representation, and Implementation
abstract
We present a novel knowledge-based system to automatically convert real-life engineering drawings to content-oriented high-level descriptions. The proposed method essentially turns the complex interpretation process into two parts: knowledge representation and knowledge-based interpretation. We propose a new hierarchical descriptor-based knowledge representation method to organize the various types of engineering objects and their complex high-level relations. The descriptors are defined using an Extended Backus Naur Form (EBNF), facilitating modification and maintenance. When interpreting a set of related engineering drawings, the knowledge-based interpretation system first constructs an EBNF-tree from the knowledge representation file, then searches for potential engineering objects guided by a depth-first order of the nodes in the EBNF-tree. Experimental results and comparisons with other interpretation systems demonstrate that our knowledge-based system is accurate and robust for high-level interpretation of complex real-life engineering projects.
Tong Lu 0002, Chiew-Lan Tai, Huafei Yang, Shijie Cai
IEEE Trans. Pattern Anal. Mach. Intell.1
2005 A new recognition model for electronic architectural drawings
Tong Lu 0002, Chiew-Lan Tai, Feng Su, Shijie Cai
Comput. Aided Des.1