Min Cai

dblp:87/5549 · DBLP profile ↗
← Back
36ranked-venue papers
11as first author
10since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 10 · 3 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 10 · 2 first-author · 1 since 2021Databases, data management, data science and information retrieval · 6 · 3 first-author · 1 since 2021Artificial intelligence and machine learning · 5 · 2 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 3 · 2 first-author · 3 since 2021Security and privacy · 2 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 1 since 2021Software engineering, systems software and programming languages · 1
YearPublicationVenuePosition
2026 How transparency impacts trust in teleoperated autonomous robots under uncertainty
Min Cai, Kuangyong Gao, Ziling Ji, Xueqi Xu
Int. J. Hum. Comput. Stud.1
2025 Strategist: Self-improvement of LLM Decision Making via Bi-Level Tree Search
abstract
Traditional reinforcement learning and planning require a lot of data and training to develop effective strategies. On the other hand, large language models (LLMs) can generalize well and perform tasks without prior training but struggle with complex planning and decision-making. We introduce **STRATEGIST**, a new approach that combines the strengths of both methods. It uses LLMs to generate and update high-level strategies in text form, while a Monte Carlo Tree Search (MCTS) algorithm refines and executes them. STRATEGIST is a general framework that optimizes strategies through self-play simulations without requiring any training data. We test STRATEGIST in competitive, multi-turn games with partial information, such as **Game of Pure Strategy (GOPS)** and **The Resistance: Avalon**, a multi-agent hidden-identity discussion game. Our results show that STRATEGIST-based agents outperform traditional reinforcement learning models, other LLM-based methods, and existing LLM agents while achieving performance levels comparable to human players.
Jonathan Light, Min Cai, Weiqin Chen 0003, Guanzhi Wang, Xiusi Chen, Wei Cheng 0002, Yisong Yue, Ziniu Hu
ICLR2
2025 Prompt-Based Relation Extraction By Reasoning with Contextual Knowledge
Xinhe Zhang, Min Cai, Yuanfeng Song
PAKDD (5)3
2024 Recognizing Textual Entailment by Hierarchical Crowdsourcing with Diverse Labor Costs
abstract
With the rapid advancement of supervised learning and the rise of large language models, the demand for high-quality labeled datasets has surged. However, crowdsourced datasets often suffer from label noise. To address this challenge, we present a hierarchical crowdsourcing framework with diverse labor costs that aims to mitigate label noise. Our framework models crowdsourcing workers with varying labor costs and leverages hierarchical crowdsourcing under limited budget constraints to enhance data quality. We establish a loop for label selection and checking, strategically selecting checkers from different cost levels, including perfect workers (Oracles) and regular experts, to optimize the checking process. Additionally, we tackle an NP-hard problem in label selection. Experimental evaluation on a real-world dataset for the Recognizing Textual Entailment (RTE) task demonstrates a significant improvement in labeled dataset quality, leading to state-of-the-art performance in downstream tasks.
Junyu Yang, Wenxi Huang, Min Cai, Chen Zhang 0013, Kaishun Wu
CSCWD4
2024 NFP-UNet: Deep Learning Estimation of Placeable Areas for 2D Irregular Packing
Min Cai, Zixin Gong
PRCV (4)1
2023 Configurable and High-Level Pipelined Lattice-Based Post Quantum Cryptography Hardware Accelerator Design
abstract
Number Theoretic Transform (NTT) and Secure Hash Algorithm 3 (SHA3), are the two main operators in the lattice-based Post-Quantum Cryptography (PQC) algorithms. Lattice-based PQC algorithms have different parameter settings, e.g., the length and modulus of NTT polynomials and the different hash functions. Motivated by the demands for more versatile NTT and SHA3 hardware accelerators, we implement the NTT and SHA3 designs that can accommodate to different parameters at run-time. Furthermore, to reduce the running cycles of the whole NTT operation and whole SHA3 operation including data transferring and calculation, we propose a pipelined architecture to optimize the gap between data transfer and calculation process in high-level. The designed configurable accelerators can be embedded in SoC to accelerate different lattice-based PQC algorithms efficiently. The experimental results show that our high-level pipelined and configurable NTT and SHA3 designs have good area-time efficiency. In specific, for the NTT design, our architecture is 4.1 times more area-time efficient compared with the state-of-the-art. For SHA3, our architecture is 1.4 times more area-time efficient over the existing configurable SHA3 designs.
Jianan Mu, Huajie Tan, Min Cai, Jing Ye 0001, Huawei Li 0001, Xiaowei Li 0001
ATS4
2023 Integrated Framework of Kansei Engineering and Kano Model Applied to Service Design
abstract
To provide high-quality service, it is necessary to emphasize customers’ emotional needs. It is necessary to determine the priority strategy for service attributes owing to limited resources and different relative levels of influence on customer perception. Therefore, this study proposed a service design framework based on the Kansei Engineering (KE) and Kano models to design services oriented toward customer perception. In the modeling analysis stage, the partial least squares algorithm and decision tree mining were used to construct a relationship model between service perception and attribute, and use the intention model. Finally, the in-flight service of a Chinese airline was considered as an application case study. The results obtained thus revealed a relationship between perception and service attributes, and their influence on customers’ intentions. These validated the proposed method and provided new concepts for service design.
Min Cai, Miaohuan Wu, Qianqian Wang 0020, Zhongliang Zhang 0001, Ziling Ji
Int. J. Hum. Comput. Interact.1
2023 A Prefetch-Adaptive Intelligent Cache Replacement Policy Based on Machine Learning
Huijing Yang, Min Cai, Zhi Cai
J. Comput. Sci. Technol.3
2023 A prefetch control strategy based on improved hill-climbing method in asymmetric multi-core architecture
abstract
Abstract Cache prefetching is a traditional way to reduce memory access latency. In multi-core systems, aggressive prefetching may harm the system. In the past, prefetching throttling strategies usually set thresholds through certain factors. When the threshold is exceeded, prefetch throttling strategies will control the aggressive prefetcher. However, these strategies usually work well in homogeneous multi-core systems and do not work well in heterogeneous multi-core systems. This paper considers the performance difference between cores under the asymmetric multi-core architecture. Through the improved hill-climbing method, the aggressiveness of prefetching for different cores is controlled, and the IPC of the core is improved. Through experiments, it is found that compared with the previous strategy, the average performance of big core is improved by more than 3%, and the average performance of little cores is improved by more than 24%.
Yixiang Xu, Han Kong, Min Cai
J. Supercomput.4
2023 WSMP: a warp scheduling strategy based on MFQ and PPF
abstract
Abstract Normally, threads in a warp do not severely interfere with each other. However, the scheduler must wait until all the threads within complete before scheduling the next warp, resulting in memory divergence. The crux of the problem is scheduling the warp in a more reasonable order. Therefore, we propose a new warp scheduling strategy called WSMP, which is based on multi-level feedback queue (MFQ) and perceptron-based prefetch filtering (PPF). All the warps are sorted beforehand according to the latency tolerance of the warps and pushed into a certain queue in MFQ. We also remold PPF to enhance the modified underlying prefetcher. We are able to strike a balance between cache hit rate and prefetch coverage then. We verify its feasibility using GPGPU-Sim, along with exclusive GPGPU workload. The results show that compared to the baseline, WSMP improves IPC by 26.45% and reduces L2 cache miss rate by 9.54% on average.
Li'ang Zhao, Min Cai, Huijing Yang
J. Supercomput.3
2020 Module dividing for brain functional networks by employing betweenness efficiency
Min Cai, Xuelian Ming, Yin Cao, Ling Zou 0002, Shuihua Wang
Multim. Tools Appl.2
2020 Rich club characteristics of dynamic brain functional networks in resting state
Min Cai, Yin Cao, Ling Zou 0002, Shuihua Wang
Multim. Tools Appl.3
2019 A Deep Neural Framework for Sales Forecasting in E-Commerce
abstract
Product sales forecasting plays a fundamental role in enhancing timeliness of product delivery in E-Commerce. Among many heterogeneous features relevant to sales forecasting, promotion campaigns held in E-Commerce and competing relation between substitutable products would greatly complicate the matter. Unfortunately, these factors are usually overlooked in the existing literature, since the conventional time series analysis based techniques mainly consider the sales records alone. In this paper, we propose a novel deep neural framework for sales forecasting in E-Commerce, named DSF. In DSF, sales forecasting is formulated as a sequence-to-sequence learning problem where the sales is estimated in a recurrent fashion. On top of the decoder, we introduce a sales residual network to explicitly model the impact of competing relation when a promotion campaign is launched for a target item or some substitutable counterparts. Extensive experiments are conducted over two real-world datasets in different domains from Taobao E-Commerce platform. Our results demonstrate that the proposed DSF obtains substantial performance gain over the traditional baselines and up-to-date deep learning alternatives in terms of forecasting accuracy. Further comparison shows that DSF has also surpassed the deep learning based solution currently depolyed in Taobao platform.
Min Cai, Yunwei Qi, Yuming Deng
CIKM4
2017 Visual Robotic Object Grasping Through Combining RGB-D Data and 3D Meshes
Yiyang Zhou, Wenhai Wang, Wenjie Guan, Yirui Wu, Heng Lai, Tong Lu 0002, Min Cai
MMM (1)7
2016 Energy and Performance Efficient Underloading Detection Algorithm of Virtual Machines in Cloud Data Centers
abstract
The problem of high energy consumption is becoming more and more serious due to the construction of large-scale cloud data centers. Detecting and switching underloading hosts into sleep mode is the key approach to reduce energy consumption. In this paper, we propose Minimum Utilization with Threshold algorithm to reduce energy consumption and VLA violation. We use the simulator CloudSim in our experiment with our algorithm and have a significant reduction on both energy consumption and SLA violation based on our data.
Lifu Zhou, Xiaoting Hao, Min Cai, Xingtian Ren
CLUSTER4
2014 XvMotion: Unified Virtual Machine Migration over Long Distance
Ali José Mashtizadeh, Min Cai, Gabriel Tarasuk-Levin, Ricardo Koller, Tal Garfinkel, Sreekanth Setty
USENIX ATC2
2012 Solving Parameter Selection Problem of Helper Thread Prefetching via Realtime Hardware Performance Monitoring
abstract
Helper thread prefetching have the potential of improving the performance of irregular data intensive applications, but the prefetching effect depends on how efficiently and swiftly the control parameters can be selected. The parameter selection and optimization was done by executing the application exhaustively in prior works. In this study, we propose a helper thread prefetching control framework, which adjusts the control parameters of helper thread automatically, called HPCF. We present the idea, initial design and implementation of HPCF. In particular, we establish a dynamic control model of helper thread prefetching and develop a two-level parameter selection algorithm. We evaluate the proposed HPCF framework on commodity multi-core platforms by using selected benchmarks which come from SPEC2006, Olden and SSCA2. Results show that our approach performs almost equal to our prior static Skip Helper Thread prefetching scheme, while the parameter selection was done by executing the application only once. And it achieves up to 33.3%, 18.2% and 18.6% performance improvement for MST, MCF, and SSCA2 benchmarks, respectively.
Zhimin Gu, Yan Huang 0011, Min Cai, Xiaohan Hu
PDCAT4
2011 The Design and Evolution of Live Storage Migration in VMware ESX
Ali José Mashtizadeh, Emré Celebi, Tal Garfinkel, Min Cai
USENIX ATC4
2009 A closed-form reduction of multi-class cost-sensitive learning to weighted multi-class learning
Fen Xia, Fuxin Li, Min Cai, Daniel Dajun Zeng
Pattern Recognit.5
2008 GossipTrust for Fast Reputation Aggregation in Peer-to-Peer Networks
abstract
In peer-to-peer (P2P) networks, reputation aggregation and ranking are the most time-consuming and space-demanding operations. This paper proposes a new gossip protocol for fast score aggregation. We developed a Bloom filter architecture for efficient score ranking. These techniques do not require any secure hashing or fast lookup mechanism, thus are applicable to both unstructured and structured P2P networks. We report the design principles and performance results of a simulated GossipTrust reputation system. Randomized gossiping with effective use of power nodes enables light-weight aggregation and fast dissemination of global scores in O(log2n) time steps, where n is the P2P network size. The Gossip-based protocol is designed to tolerate dynamic peer joining and departure, as well as to avoid possible peer collusions. The scheme has a considerably low gossiping message overhead, i.e. O(n log2n) messages for n nodes. Bloom filters demand at most 512 KB memory per node for a 10,000-node network. We evaluate the performance of GossipTrust with distributed P2P file-sharing and parameter-sweeping applications. The simulation results demonstrate that GossipTrust has small aggregation time, low memory demand, and high ranking accuracy. These results suggest promising advantages of using the GossipTrust system for trusted P2P applications.
Runfang Zhou, Kai Hwang 0001, Min Cai
IEEE Trans. Knowl. Data Eng.3
2007 Distributed Aggregation Algorithms with Load-Balancing for Scalable Grid Resource Monitoring
abstract
Scalable resource monitoring and discovery are essential to the planet-scale infrastructures such as grids and PlanetLab. This paper proposes a scalable grid monitoring architecture that builds distributed aggregation trees (DAT) on a structured P2P network like Chord. By leveraging Chord topology and routing mechanisms, the DAT trees are implicitly constructed from native Chord routing paths without membership maintenance. To balance the DAT trees, we propose a balanced routing algorithm on Chord that dynamically selects the parent of a node from its finger nodes by its distance to the root. This paper shows that this balanced routing algorithm enables the construction of almost completely balanced DATs, when nodes are evenly distributed in the Chord identifier space. We have evaluated the performance and scalability of a DAT prototype implementation with up to 8192 nodes. Our experimental results show that the balanced DAT scheme scales well to a large number of nodes and corresponding aggregation trees. Without maintaining explicit parent-child membership, it has very low overhead during node arrival and departure. We demonstrate that the DAT scheme performs well in grid resource monitoring.
Min Cai, Kai Hwang 0001
IPDPS1
2007 WormShield: Fast Worm Signature Generation with Distributed Fingerprint Aggregation
abstract
Fast and accurate generation of worm signatures is essential to contain zero-day worms at the Internet scale. Recent work has shown that signature generation can be automated by analyzing the repetition of worm substrings (that is, fingerprints) and their address dispersion. However, at the early stage of a worm outbreak, individual edge networks are often short of enough worm exploits for generating accurate signatures. This paper presents both theoretical and experimental results on a collaborative worm signature generation system (WormShield) that employs distributed fingerprint filtering and aggregation over multiple edge networks. By analyzing real-life Internet traces, we discovered that fingerprints in background traffic exhibit a Zipf-like distribution. Due to this property, a distributed fingerprint filtering reduces the amount of aggregation traffic significantly. WormShield monitors utilize a new distributed aggregation tree (DAT) to compute global fingerprint statistics in a scalable and load-balanced fashion. We simulated a spectrum of scanning worms including CodeRed and Slammer by using realistic Internet configurations of about 100,000 edge networks. On average, 256 collaborative monitors generate the signature of CodeRedl-v2 135 times faster than using the same number of isolated monitors. In addition to speed gains, we observed less than 100 false signatures out of 18.7-Gbyte Internet traces, yielding a very low false-positive rate. Each monitor only generates about 0.6 kilobit per second of aggregation traffic, which is 0.003 percent of the 18 megabits per second link traffic sniffed. These results demonstrate that the WormShield system offers distinct advantages in speed gains, signature accuracy, and scalability for large-scale worm containment.
Min Cai, Kai Hwang 0001, Jianping Pan 0001, Christos Papadopoulos
IEEE Trans. Dependable Secur. Comput.1
2007 Hybrid Intrusion Detection with Weighted Signature Generation over Anomalous Internet Episodes
abstract
This paper reports the design principles and evaluation results of a new experimental hybrid intrusion detection system (HIDS). This hybrid system combines the advantages of low false-positive rate of signature-based intrusion detection system (IDS) and the ability of anomaly detection system (ADS) to detect novel unknown attacks. By mining anomalous traffic episodes from Internet connections, we build an ADS that detects anomalies beyond the capabilities of signature-based SNORT or Bro systems. A weighted signature generation scheme is developed to integrate ADS with SNORT by extracting signatures from anomalies detected. HIDS extracts signatures from the output of ADS and adds them into the SNORT signature database for fast and accurate intrusion detection. By testing our HIDS scheme over real-life Internet trace data mixed with 10 days of Massachusetts Institute of Technology/Lincoln Laboratory (MIT/LL) attack data set, our experimental results show a 60 percent detection rate of the HIDS, compared with 30 percent and 22 percent in using the SNORT and Bro systems, respectively. This sharp increase in detection rate is obtained with less than 3 percent false alarms. The signatures generated by ADS upgrade the SNORT performance by 33 percent. The HIDS approach proves the vitality of detecting intrusions and anomalies, simultaneously, by automated data mining and signature generation over Internet connection episodes
Kai Hwang 0001, Min Cai
IEEE Trans. Dependable Secur. Comput.2
2006 Applying Peer-to-Peer Techniques to Grid Replica Location Services
Ann L. Chervenak, Min Cai
J. Grid Comput.2
2005 A comprehensive method for multilingual video text detection, localization, and extraction
abstract
Text in video is a very compact and accurate clue for video indexing and summarization. Most video text detection and extraction methods hold assumptions on text color, background contrast, and font style. Moreover, few methods can handle multilingual text well since different languages may have quite different appearances. This paper performs a detailed analysis of multilingual text characteristics, including English and Chinese. Based on the analysis, we propose a comprehensive, efficient video text detection, localization, and extraction method, which emphasizes the multilingual capability over the whole processing. The proposed method is also robust to various background complexities and text appearances. The text detection is carried out by edge detection, local thresholding, and hysteresis edge recovery. The coarse-to-fine localization scheme is then performed to identify text regions accurately. The text extraction consists of adaptive thresholding, dam point labeling, and inward filling. Experimental results on a large number of video images and comparisons with other methods are reported in detail.
Michael R. Lyu, Jiqiang Song, Min Cai
IEEE Trans. Circuits Syst. Video Technol.3
2004 Stochastic power control algorithm in CDMA systems
abstract
Transmitter power control has been proven to be an efficient method to solve co-channel interference in CDMA system, and to increase the capacity of systems. A novel distributed power control algorithm based on estimates of channel gains is presented in this paper. The proposed algorithm is robust and does not require the exact knowledge of channel gains. In the proposed scheme, power control problem is converted into the zero point problem of a certain function. It is proved that the algorithm converges to the average optimal solution by stochastic approximation approach.
Min Cai, Wei Wang 0036, Huanshui Zhang
ICARCV1
2004 A Peer-to-Peer Replica Location Service Based on a Distributed Hash Table
abstract
A Replica Location Service (RLS) allows registration and discovery of data replicas. In earlier work, we proposed an RLS framework and described the performance and scalability of an RLS implementation in Globus Toolkit Version 3.0. In this paper, we present a Peer-to-Peer Replica Location Service (P-RLS) with properties of self-organization, fault-tolerance and improved scalability. P-RLS uses the Chord algorithm to self-organize PRLS servers and exploits the Chord overlay network to replicate P-RLS mappings adaptively. Our performance measurements demonstrate that update and query latencies increase at a logarithmic rate with the size of the P-RLS network, while the overhead of maintaining the P-RLS network is reasonable. Our simulation results for adaptive replication demonstrate that as the number of replicas per mapping increases, the mappings are more evenly distributed among P-RLS nodes. We introduce a predecessor replication scheme and show it reduces query hotspots of popular mappings by distributing queries among nodes.
Min Cai, Ann L. Chervenak, Martin R. Frank
SC1
2004 RDFPeers: a scalable distributed RDF repository based on a structured peer-to-peer network
abstract
Centralized Resource Description Framework (RDF) repositories have limitations both in their failure tolerance and in their scalability. Existing Peer-to-Peer (P2P) RDF repositories either cannot guarantee to find query results, even if these results exist in the network, or require up-front definition of RDF schemas and designation of super peers. We present a scalable distributed RDF repository (RDFPeers) that stores each triple at three places in a multi-attribute addressable network by applying globally known hash functions to its subject predicate and object. Thus all nodes know which node is responsible for storing triple values they are looking for and both exact-match and range queries can be efficiently routed to those nodes. RDFPeers has no single point of failure nor elevated peers and does not require the prior definition of RDF schemas. Queries are guaranteed to find matched triples in the network if the triples exist. In RDFPeers both the number of neighbors per node and the number of routing hops for inserting RDF triples and for resolving most queries are logarithmic to the number of nodes in the network. We further performed experiments that show that the triple-storing load in RDFPeers differs by less than an order of magnitude between the most and the least loaded nodes for real-world RDF data.
Min Cai, Martin R. Frank
WWW1
2004 MAAN: A Multi-Attribute Addressable Network for Grid Information Services
Min Cai, Martin R. Frank, Pedro A. Szekely
J. Grid Comput.1
2004 A subscribable peer-to-peer RDF repository for distributed metadata management
Min Cai, Martin R. Frank, Baoshi Yan, Robert M. MacGregor
J. Web Semant.1
2003 A robust statistic method for classifying color polarity of video text
abstract
Video text extraction and recognition are prerequisite tasks for video indexing and retrieval. Color polarity classification of video text is very important to these tasks. Most existing text extraction methods assume that the text color is always light (or dark). Obviously, this assumption restricts the application of these methods to some specific domains. Only a few methods can detect the color polarity in the condition that the background is clear. However, many real video texts have various appearances and complex backgrounds that existing methods cannot handle. The paper proposes a statistical color polarity classification method that is robust to various background complexities, font styles, stroke widths, and languages. We discover the intrinsic relationships between text edges and background edges, and then develop an efficient measurement to detect the color polarity. The experimental results show that the proposed method achieves a much higher accuracy, 98.5%, than existing methods.
Jiqiang Song, Min Cai, Michael R. Lyu
ICASSP (3)2
2003 A robust statistic method for classifying color polarity of video text
abstract
Video text extraction and recognition are prerequisite tasks for video indexing and retrieval. Color polarity classification of video text is very important to these tasks. Most existing text extraction methods assume that the text color is always light (or dark). Obviously, this assumption restricts the application of these methods to some specific domains. Only a few methods can detect the color polarity on condition that the background is clear. However, many real video texts have various appearances and complex backgrounds that existing methods cannot handle. This paper proposes a statistic color polarity classification method that is robust to various background complexities, font styles, stroke widths, and languages. We discover the intrinsic relationships between text edges and background edges, and then develop an efficient measurement to detect the color polarity. The experimental results show that the proposed method achieves a much higher accuracy, 98.5%, than existing methods.
Jiqiang Song, Min Cai, Michael R. Lyu
ICME2
2003 PVCAIS: a personal videoconference archive indexing system
abstract
Wilh the rapid deployment of videoconference, the fast-accumulated personal videoconference archives need to be effectively indexed. This paper proposes a well-designed indexing system - PVCAIS that integrates many multimedia-indexing techniques to manage personal videoconference archives. Firstly, the contents of video, audio, text and whiteboard communications are stored after removing the redundancies. Next, more information, e.g., participants, title, keywords and slides, is extracted by face detection and recognition, speech recognition, OCR, automatic title generation, keyword selection, etc. Then, an XML index file containing the summary of the videoconference is also generated. The whole indexing process is automatic except that the face of a new contact needs to be interactively identified for only once. Finally, we demonstrate a graphical user interface which allows the user to search and browse the indexed videoconference archives conveniently.
Jiqiang Song, Michael R. Lyu, Jenq-Neng Hwang, Min Cai
ICME4
2003 Proteus: A System for Dynamically Composing and Intelligently Executing Web Services
Shahram Ghandeharizadeh, Craig A. Knoblock, Christos Papadopoulos, Cyrus Shahabi, Esam Alwagait, José Luis Ambite, Min Cai, Ching-Chien Chen, Parikshit Pol, Rolfe R. Schmidt, Saihong Song, Snehal Thakkar, Runfang Zhou
ICWS7
2002 A Comparison of Alternative Encoding Mechanisms for Web Services
Min Cai, Shahram Ghandeharizadeh, Rolfe R. Schmidt, Saihong Song
DEXA1
2002 A new approach for video text detection
abstract
Text detection is fundamental to video information retrieval and indexing. Existing methods cannot handle well those texts with different contrast or embedded in a complex background. To handle these difficulties, this paper proposes an efficient text detection approach, which is based on invariant features, such as edge strength, edge density, and horizontal distribution. First, it applies edge detection and uses a low threshold to filter out definitely non-text edges. Then, a local threshold is selected to both keep low-contrast text and simplify complex background of high-contrast text. Next, two text-area enhancement operators are proposed to highlight those areas with either high edge strength or high edge density. Finally, coarse-to-fine detection locates text regions efficiently. Experimental results show that this approach is robust for contrast, font-size, font-color, language, and background complexity.
Michael R. Lyu, Jiqiang Song, Min Cai
ICIP (1)3