EDBT 2026 Demo / reviewers in the wild / expert
Tin Kam Ho
dblp:22/6176
· DBLP profile ↗
51ranked-venue papers
25as first author
1since 2021 · last 2022
0000-0001-9635-1769ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 40 · 24 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 16 · 8 first-author · 1 since 2021Databases, data management, data science and information retrieval · 10 · 6 first-authorComputer networks · 3Applied, interdisciplinary, general and emerging computing · 2Systems, architecture and hardware · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
6 papers |
Information extraction and text analysis · 41% Question answering and dialogue systems · 27% Learning paradigms · 25% | |
| Computer networks
1 paper |
Network management and operations · 77% Routing and switching · 23% |
Topics — the 15 heaviest of 16, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Information extraction and text analysis
text classification |
0.8 | 2 | 2020 | Iterative Data Programming for Expanding Text Classification Corpora · AAAI 2020 Bootstrapping Conversational Agents with Weak Supervision · AAAI 2019 |
Machine learning › Learning paradigms
weakly supervised learning |
0.8 | 2 | 2020 | Iterative Data Programming for Expanding Text Classification Corpora · AAAI 2020 Bootstrapping Conversational Agents with Weak Supervision · AAAI 2019 |
Natural language and speech › Information extraction and text analysis › data annotation
data programming |
0.5 | 2 | 2020 | Iterative Data Programming for Expanding Text Classification Corpora · AAAI 2020 Bootstrapping Conversational Agents with Weak Supervision · AAAI 2019 |
Natural language and speech › Question answering and dialogue systems
intent detection |
0.5 | 2 | 2020 | Bootstrapping Conversational Agents with Weak Supervision · AAAI 2019 Iterative Data Programming for Expanding Text Classification Corpora · AAAI 2020 |
Natural language and speech › Question answering and dialogue systems
conversational agents |
0.4 | 1 | 2019 | Bootstrapping Conversational Agents with Weak Supervision · AAAI 2019 |
Network management and operations › fault management
fault diagnosis |
0.1 | 1 | 2009 | An Online Mechanism for BGP Instability Detection and Analysis · IEEE Trans. Computers 2009 |
Machine learning › Trustworthy machine learning
interpretability |
0.1 | 2 | 2002 | Complexity Measures of Supervised Classification Problems · IEEE Trans. Pattern Anal. Mach. Intell. 2002 The Random Subspace Method for Constructing Decision Forests · IEEE Trans. Pattern Anal. Mach. Intell. 1998 |
Machine learning › Kernel, tree and ensemble methods
ensemble learning |
0.0 | 2 | 1998 | The Random Subspace Method for Constructing Decision Forests · IEEE Trans. Pattern Anal. Mach. Intell. 1998 Decision Combination in Multiple Classifier Systems · IEEE Trans. Pattern Anal. Mach. Intell. 1994 |
Computer vision › Image recognition and object detection
character recognition |
0.0 | 2 | 1997 | Large-Scale Simulation Studies in Image Pattern Recognition · IEEE Trans. Pattern Anal. Mach. Intell. 1997 Decision Combination in Multiple Classifier Systems · IEEE Trans. Pattern Anal. Mach. Intell. 1994 |
Routing and switching › inter-domain routing
BGP |
0.0 | 1 | 2009 | An Online Mechanism for BGP Instability Detection and Analysis · IEEE Trans. Computers 2009 |
Machine learning › Kernel, tree and ensemble methods › decision tree
decision tree classifiers |
0.0 | 1 | 1998 | The Random Subspace Method for Constructing Decision Forests · IEEE Trans. Pattern Anal. Mach. Intell. 1998 |
Machine learning › Kernel, tree and ensemble methods › ensemble learning
random subspace method |
0.0 | 1 | 1998 | The Random Subspace Method for Constructing Decision Forests · IEEE Trans. Pattern Anal. Mach. Intell. 1998 |
Machine learning › Probabilistic and Bayesian machine learning › statistical inference
bayes risk estimation |
0.0 | 1 | 1997 | Large-Scale Simulation Studies in Image Pattern Recognition · IEEE Trans. Pattern Anal. Mach. Intell. 1997 |
Machine learning › Kernel, tree and ensemble methods › ensemble learning
classifier ensemble |
0.0 | 1 | 1994 | Decision Combination in Multiple Classifier Systems · IEEE Trans. Pattern Anal. Mach. Intell. 1994 |
Machine learning › Learning theory › classification
classifier evaluation |
0.0 | 1 | 1997 | Large-Scale Simulation Studies in Image Pattern Recognition · IEEE Trans. Pattern Anal. Mach. Intell. 1997 |
Methods — techniques the papers use, named apart from their topics
labeling functions · 0.4iterative data programming · 0.4ensemble learning · 0.4weak supervision · 0.4search-label-propagate · 0.4data programming · 0.4statistical pattern recognition · 0.1adaptive segmentation · 0.1random labeling · 0.0geometric complexity measures · 0.0pseudorandom feature subspace sampling · 0.0bagging · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2022 | Complexity of Representations in Deep LearningabstractDeep neural networks use multiple layers of functions to map an object represented by an input vector progressively to different representations, and with sufficient training, eventually to a single score for each class that is the output of the final decision function. Ideally, in this output space, the objects of different classes achieve maximum separation. Motivated by the need to better understand the inner working of a deep neural network, we analyze the effectiveness of the learned representations in separating the classes from a data complexity perspective. Using a simple complexity measure, a popular benchmarking task, and a well-known architecture design, we show how the data complexity evolves through the network, how it changes during training, and how it is impacted by the network design and the availability of training samples. We discuss the implications of the observations and the potentials for further studies. Tin Kam Ho |
ICPR | 1 |
| 2020 | Iterative Data Programming for Expanding Text Classification CorporaabstractReal-world text classification tasks often require many labeled training examples that are expensive to obtain. Recent advancements in machine teaching, specifically the data programming paradigm, facilitate the creation of training data sets quickly via a general framework for building weak models, also known as labeling functions, and denoising them through ensemble learning techniques. We present a fast, simple data programming method for augmenting text data sets by generating neighborhood-based weak models with minimal supervision. Furthermore, our method employs an iterative procedure to identify sparsely distributed examples from large volumes of unlabeled data. The iterative data programming techniques improve newer weak models as more labeled data is confirmed with human-in-loop. We show empirical results on sentence classification tasks, including those from a task of improving intent recognition in conversational agents. Neil Mallinar, Abhishek Shah, Tin Kam Ho, Rajendra Ugrani |
AAAI | 3 |
| 2020 | Special issue on recent advances in statistical, structural and syntactic pattern recognition
Xiao Bai 0001, Edwin R. Hancock, Richard C. Wilson 0001, Tin Kam Ho |
Pattern Recognit. Lett. | 4 |
| 2019 | Bootstrapping Conversational Agents with Weak SupervisionabstractMany conversational agents in the market today follow a standard bot development framework which requires training intent classifiers to recognize user input. The need to create a proper set of training examples is often the bottleneck in the development process. In many occasions agent developers have access to historical chat logs that can provide a good quantity as well as coverage of training examples. However, the cost of labeling them with tens to hundreds of intents often prohibits taking full advantage of these chat logs. In this paper, we present a framework called search, label, and propagate (SLP) for bootstrapping intents from existing chat logs using weak supervision. The framework reduces hours to days of labeling effort down to minutes of work by using a search engine to find examples, then relies on a data programming approach to automatically expand the labels. We report on a user study that shows positive user feedback for this new approach to build conversational agents, and demonstrates the effectiveness of using data programming for autolabeling. While the system is developed for training conversational agents, the framework has broader application in significantly reducing labeling effort for training text classifiers. Neil Mallinar, Abhishek Shah, Rajendra Ugrani, Manikandan Gurusankar, Tin Kam Ho, Qingzi Vera Liao, Rachel K. E. Bellamy, Robert Yates, Chris Desmarais, Blake McGregor |
AAAI | 6 |
| 2018 | Classifier Recommendation Using Data Complexity MeasuresabstractApplication of machine learning to new and unfamiliar domains calls for increasing automation in choosing a learning algorithm suitable for the data arising from each domain. Meta-learning could address this need since it has been largely used in the last years to support the recommendation of the most suitable algorithms for a new dataset. The use of complexity measures could increase the systematic comprehension over the meta-models and also allow to differentiate the performance of a set of techniques taking into account the overlap between classes imposed by feature values, the separability and distribution of the data points. In this paper we compare the effectiveness of several standard regression models in predicting the accuracies of classifiers for classification problems from the OpenML repository. We show that the models can predict the classifiers' accuracies with low mean-squared-error and identify the best classifier for a problem that results in statistically significant improvements over a randomly chosen classifier or a fixed classifier believed to be good on average. Luís Paulo F. Garcia, Ana Carolina Lorena, Marcílio Carlos Pereira de Souto, Tin Kam Ho |
ICPR | 4 |
| 2015 | Design of the 2015 ChaLearn AutoML challengeabstractChaLearn is organizing the Automatic Machine Learning (AutoML) contest for IJCNN 2015, which challenges participants to solve classification and regression problems without any human intervention. Participants' code is automatically run on the contest servers to train and test learning machines. However, there is no obligation to submit code; half of the prizes can be won by submitting prediction results only. Datasets of progressively increasing difficulty are introduced throughout the six rounds of the challenge. (Participants can enter the competition in any round.) The rounds alternate phases in which learners are tested on datasets participants have not seen, and phases in which participants have limited time to tweak their algorithms on those datasets to improve performance. This challenge will push the state of the art in fully automatic machine learning on a wide range of real-world problems. The platform will remain available beyond the termination of the challenge. Isabelle Guyon, Kristin P. Bennett, Gavin C. Cawley, Hugo Jair Escalante, Sergio Escalera, Tin Kam Ho, Núria Macià, Bisakha Ray, Mehreen Saeed, Alexander R. Statnikov, Evelyne Viegas |
IJCNN | 6 |
| 2015 | Mitigating Mimicry Attacks Against the Session Initiation ProtocolabstractThe U.S. National Academies of Science's Board on Science, Technology and Economic Policy estimates that the Internet and voice-over-IP (VoIP) communications infrastructure generates 10% of U.S. economic growth. As market forces move increasingly towards Internet and VoIP communications, there is proportional increase in telephony denial of service (TDoS) attacks. Like denial of service (DoS) attacks, TDoS attacks seek to disrupt business and commerce by directing a flood of anomalous traffic towards key communication servers. In this work, we focus on a new class of anomalous traffic that exhibits a mimicry TDoS attack. Such an attack can be launched by crafting malformed messages with small changes from normal ones. We show that such malicious messages easily bypass intrusion detection systems (IDS) and degrade the goodput of the server drastically by forcing it to parse the message looking for the needed token. Our approach is not to parse at all; instead, we use multiple classifier systems (MCS) to exploit the strength of multiple learners to predict the true class of a message with high probability (98.50% ≤ p ≤ 99.12%). We proceed systematically by first formulating an optimization problem of picking the minimum number of classifiers such that their combination yields the optimal classification performance. Next, we analytically bound the maximum performance of such a system and empirically demonstrate that it is possible to attain close to the maximum theoretical performance across varied datasets. Finally, guided by our analysis we construct an MCS appliance that demonstrates superior classification accuracy with O(1) runtime complexity across varied datasets. Samuel Marchal, Anil Mehta, Vijay K. Gurbani, Radu State, Tin Kam Ho, Flavia Sancier-Barbosa |
IEEE Trans. Netw. Serv. Manag. | 5 |
| 2014 | Building Optimal Radio-Frequency Signal MapsabstractA popular way for using radio-frequency (RF) signals (e.g. WiFi) to position people or device indoors is by matching received radio signal strength (RSS) to fingerprints that are spatial signatures of such measures. Traditionally such signal maps are built by manual collection of repeated measurements at predefined locations following a spatial sampling scheme. Recently, such labor intensive processes are being replaced by robot-based automation or crowd-sourced simultaneous localization and mapping (SLAM). These new approaches produce time-stamped trajectories along with time-stamped RSS as the human or robot moves freely about the building. However, they require an additional procedure to segment the continuous RF samples into fingerprint cells to produce a robust signal map. In this paper, we explore several strategies for building optimal signal maps from RSS collected along robotic or pedestrian trajectories. We compare two clustering algorithms with a baseline strategy that divides the trajectories into a hierarchy of fixed-size grids. We study the trade-off between the spatial extent of the fingerprint cells and the differentiability of the RSS distribution in each cell, as well as their impact on localization accuracy and on fingerprint storage. We experimented with traces collected by an autonomous robot exploring a large multi-floor office building. Piotr Mirowski, Tin Kam Ho, Phil Whiting |
ICPR | 2 |
| 2014 | Pose Invariant Activity Classification for Multi-floor Indoor LocalizationabstractSmartphone based indoor localization caught massive interest of the localization community in recent years. Combining pedestrian dead reckoning obtained using the phone's inertial sensors with the Graph SLAM (Simultaneous Localization and Mapping) algorithm is one of the most effective approaches to reconstruct the entire pedestrian trajectory given a set of visited landmarks during movement. A key to Graph SLAM-based localization is the detection of reliable landmarks, which are typically identified using visual cues or via NFC tags or QR codes. Alternatively, human activity can be classified to detect organic landmarks such as visits to stairs and elevators while in movement. We provide a novel human activity classification framework that is invariant to the pose of the smartphone. Pose invariant features allow robust observation no matter how a user puts the phone in the pocket. In addition, activity classification obtained by an SVM (Support Vector Machine) is used in a Bayesian framework with an HMM (Hidden Markov Model) that improves the activity inference based on temporal smoothness. Furthermore, the HMM jointly infers activity and floor information, thus providing multi-floor indoor localization. Our experiments show that the proposed framework detects landmarks accurately and enables multi-floor indoor localization from the pocket using Graph SLAM. Saehoon Yi, Piotr Mirowski, Tin Kam Ho, Vladimir Pavlovic 0001 |
ICPR | 3 |
| 2014 | Motion feature filtering for event detection in crowded scenes
Lawrence O'Gorman, Tin Kam Ho |
Pattern Recognit. Lett. | 3 |
| 2013 | SignalSLAM: Simultaneous localization and mapping with mixed WiFi, Bluetooth, LTE and magnetic signalsabstractIndoor localization typically relies on measuring a collection of RF signals, such as Received Signal Strength (RSS) from WiFi, in conjunction with spatial maps of signal fingerprints. A new technology for localization could arise with the use of 4G LTE telephony small cells, with limited range but with rich signal strength information, namely Reference Signal Received Power (RSRP). In this paper, we propose to combine an ensemble of available sources of RF signals to build multi-modal signal maps that can be used for localization or for network deployment optimization. We primarily rely on Simultaneous Localization and Mapping (SLAM), which provides a solution to the challenge of building a map of observations without knowing the location of the observer. SLAM has recently been extended to incorporate signal strength from WiFi in the so-called WiFi-SLAM. In parallel to WiFi-SLAM, other localization algorithms have been developed that exploit the inertial motion sensors and a known map of either WiFi RSS or of magnetic field magnitude. In our study, we use all the measurements that can be acquired by an off-the-shelf smartphone and crowd-source the data collection from several experimenters walking freely through a building, collecting time-stamped WiFi and Bluetooth RSS, 4G LTE RSRP, magnetic field magnitude, GPS reference points when outdoors, Near-Field Communication (NFC) readings at specific landmarks and pedestrian dead reckoning based on inertial data. We resolve the location of all the users using a modified version of Graph-SLAM optimization of the users poses with a collection of absolute location and pairwise constraints that incorporates multi-modal signal similarity. We demonstrate that we can recover the user positions and thus simultaneously generate dense signal maps for each WiFi access point and 4G LTE small cell, “from the pocket”. Finally, we demonstrate the localization performance using selected single modalities, such as only WiFi and the WiFi signal maps that we generated. Piotr Mirowski, Tin Kam Ho, Saehoon Yi, Michael MacDonald |
IPIN | 2 |
| 2013 | Learner excellence biased by data set selection: A case for data characterisation and artificial data sets
Núria Macià, Ester Bernadó-Mansilla, Albert Orriols-Puig, Tin Kam Ho |
Pattern Recognit. | 4 |
| 2012 | On using multiple classifier systems for Session Initiation Protocol (SIP) anomaly detectionabstractThe Session Initiation Protocol (SIP) is an important multimedia session establishment protocol used on the Internet. It is a text-based protocol, which is complex to parse due to the wide variability in representing the information elements. Building a parser for SIP may appear straight-forward; however, writing an efficient, robust, and scalable parser that is immune to low-effort attacks using malformed messages is surprisingly difficult. To mitigate this, self-learning systems based on Euclidean distance classifiers have been proposed to determine whether a message is well-formed or not. The efficacy of such machine learning algorithms must be studied on varied data sets before they can be successfully used. Our previous work has shown that Euclidean distance-based classifiers and standard classifiers used for self-learning problems are unable to detect malformed self-similar SIP messages (i.e., invalid SIP messages that differ by only a few bytes from normal SIP messages). This paper proposes using multiple classifier systems to detect malformed self-similar SIP messages. Our results show that a judiciously constructed multiple classifier system yields classification performance as high as 97.56% of the messages being classified correctly. We further show that for self-similar SIP messages, feature reduction measures based on the first moment are insufficient for improving classification accuracy. Anil Mehta, Neda Hantehzadeh, Vijay K. Gurbani, Tin Kam Ho, Flavia Sander |
ICC | 4 |
| 2012 | Improving cross-validation based classifier selection using meta-learning
Jesse H. Krijthe, Tin Kam Ho, Marco Loog |
ICPR | 2 |
| 2011 | On the inefficacy of Euclidean classifiers for detecting self-similar Session Initiation Protocol (SIP) messagesabstractThe Session Initiation Protocol (SIP) is an important multimedia session establishment protocol used on the Internet. Due to the nature and deployment realities of the protocol (ASCII message representation, most deployments over UDP, limited use of message encryption), it becomes relatively easy to attack the protocol at the message level. To mitigate this, self-learning systems have been proposed to counteract new threats. However the efficacy of existing machine learning algorithms must be studied on varied data sets before they can be successfully used. Existing literature indicates that Euclidean distance based classifiers work well to detect anomalous messages. Our work suggests that such classifiers do not produce adequate results for well-crafted malicious messages that differ very slightly from normal messages. To demonstrate this, we gather SIP traffic and minimally perturb it using 13 generic transforms to create malicious SIP messages. We use the Levenshtein distance, L, as a measure of similarity between normal and malicious SIP messages. We subject our dataset — consisting of malicious and normal SIP messages — to Euclidean distance-based classifiers as well as four standard classifiers. Our results show vast differences for Euclidean distance-based classifiers on our dataset than reported in current literature. We further see that the standard classifiers are better able to classify an anomalous message when L is small. Anil Mehta, Neda Hantehzadeh, Vijay K. Gurbani, Tin Kam Ho, Jun Koshiko, R. Viswanathan 0002 |
Integrated Network Management | 4 |
| 2011 | KL-divergence kernel regression for non-Gaussian fingerprint based localizationabstractVarious methods have been developed for indoor localization using WLAN signals. Algorithms that fingerprint the Received Signal Strength Indication (RSSI) of WiFi for different locations can achieve tracking accuracies of the order of a few meters. RSSI fingerprinting suffers though from two main limitations: first, as the signal environment changes, so does the fingerprint database, which requires regular updates; second, it has been reported that, in practice, certain devices record more complex (e.g bimodal) distributions of WiFi signals, precluding algorithms based on the mean RSSI. In this article, we propose a simple methodology that takes into account the full distribution for computing similarities among fingerprints using Kullback-Leibler divergence, and that performs localization through kernel regression. Our method provides a natural way of smoothing over time and trajectories. Moreover, we propose unsupervised KL-divergence-based recalibration of the training fingerprints. Finally, we apply our method to work with histograms of WiFi connections to access points, ignoring RSSI distributions, and thus removing the need for recalibration. We demonstrate that our results outperform nearest neighbors or Kalman and Particle Filters, achieving up to 1m accuracy in office environments. We also show that our method generalizes to non-Gaussian RSSI distributions. Piotr Mirowski, Harald Steck, Phil Whiting, Ravishankar Palaniappan, Michael MacDonald, Tin Kam Ho |
IPIN | 6 |
| 2009 | An Online Mechanism for BGP Instability Detection and AnalysisabstractThe importance of border gateway protocol (BGP) as the primary interautonomous system (AS) routing protocol that maintains the connectivity of the Internet imposes stringent stability requirements on its route selection process. Accidental and malicious activities such as misconfigurations, failures, and worm attacks can induce severe BGP instabilities leading to data loss, extensive delays, and loss of connectivity. In this work, we propose an online instability detection architecture that can be implemented by individual routers. We use statistical pattern recognition techniques for detecting the instabilities, and the algorithm is evaluated using real Internet data for a diverse set of events including misconfiguration, node failures, and several worm attacks. The proposed scheme is based on adaptive segmentation of feature traces extracted from BGP update messages and exploiting the temporal and spatial correlations in the traces for robust detection of the instability events. Furthermore, we use route change information to pinpoint the culprit ASes where the instabilities have originated. Shivani Deshpande, Marina Thottan, Tin Kam Ho, Biplab Sikdar 0001 |
IEEE Trans. Computers | 3 |
| 2008 | Mobile User Profile Acquisition through Network Observables and Explicit User QueriesabstractThis paper describes a novel approach for gathering profile information about mobile phone users. The focus is on information that can be used to enhance targeting of advertisements. (The ads might be delivered into the mobile phones, or to other devices such as the user's IPTV.) Unlike previous approaches, we use a two-tiered approach for learning end-user habits and preferences. In this approach the first tier involves statistical learning from network observable data (in the current paper, primarily logs of cell towers visited), and the second tier involves explicit queries to the user (in the current paper, to ask, e.g., what kinds of activities the user does in a given region that he frequents). The user might be willing to answer occasional queries of this sort through offers of service discounts, or to be able to receive more relevant ads. The paper focuses on two key aspects of our approach, which correspond to how the two tiers are instantiated in the current version of the prototype system that we have developed at Bell Labs. The first concerns the statistical techniques used to determine information about regions visited, along with the frequency of visits, typical durations, and typical visit times. These techniques were developed based on a training set consisting of logs of 6 users with mobile devices over a period of several months. The techniques address issues that arise when a given small region is serviced by multiple cell towers (in which case oscillations between cell towers can be confused with movement between regions). The second key aspect concerns optimizing the order in which queries are presented to users, in a context where different query answers have different value for the advertising process. (The values of answers might be influenced by the mix of advertising campaigns from which ads are to be matched against users.) Optimization is NP-complete in a relatively general context. We develop a polynomial time algorithm which yields optimal sequences for the case where the family of queries to be asked satisfies a tree-based property. This is extended to create a heuristic polynomial time algorithm for the general case. Nilton Bila, Robert Dinoff, Tin Kam Ho, Richard Hull 0001, Bharat Kumar, Paulo Santos 0002 |
MDM | 4 |
| 2006 | A Statistical Approach to Anomaly Detection in Interdomain RoutingabstractA number of events such as hurricanes, earthquakes, power outages can cause large-scale failures in the Internet. These in turn cause anomalies in the interdomain routing process. The policy-based nature of border gateway protocol (BGP) further aggravates the effect of these anomalies causing severe, long lasting route fluctuations. In this work we propose an architecture for anomaly detection that can be implemented on individual routers. We use statistical pattern recognition techniques for extracting meaningful features from the BGP update message data. A time-series segmentation algorithm is then carried out on the feature traces to detect the onset of an instability event The performance of the proposed algorithm is evaluated using real Internet trace data. We show that instabilities triggered by events like router mis-configurations, infrastructure failures and worm attacks can be detected with a false alarm rate as low as 0.0083 alarms per hour. We also show that our learning based mechanism is highly robust as compared to methods like exponentially weighted moving average (EWMA) based detection. Shivani Deshpande, Marina Thottan, Tin Kam Ho, Biplab Sikdar 0001 |
BROADNETS | 3 |
| 2006 | Exploratory Analysis System for Semi-structured Engineering Logs
Michael Flaster, Bruce Hillyer, Tin Kam Ho |
Document Analysis Systems | 3 |
| 2006 | Recent submissions in linear dimensionality reduction and face recognition
Robert P. W. Duin, Marco Loog, Tin Kam Ho |
Pattern Recognit. Lett. | 3 |
| 2005 | Article retraction: Least-squares fitting for deformable superquadric model based on orthogonal distance
Tin Kam Ho |
Pattern Recognit. Lett. | 1 |
| 2005 | Editorial
Tin Kam Ho |
Pattern Recognit. Lett. | 1 |
| 2005 | Domain of competence of XCS classifier system in complexity measurement spaceabstractThe XCS classifier system has recently shown a high degree of competence on a variety of data mining problems, but to what kind of problems XCS is well and poorly suited is seldom understood, especially for real-world classification problems. The major inconvenience has been attributed to the difficulty of determining the intrinsic characteristics of real-world classification problems. This paper investigates the domain of competence of XCS by means of a methodology that characterizes the complexity of a classification problem by a set of geometrical descriptors. In a study of 392 classification problems along with their complexity characterization, we are able to identify difficult and easy domains for XCS. We focus on XCS with hyperrectangle codification, which has been predominantly used for real-attributed domains. The results show high correlations between XCS's performance and measures of length of class boundaries, compactness of classes, and nonlinearities of decision boundaries. We also compare the relative performance of XCS with other traditional classifier schemes. Besides confirming the high degree of competence of XCS in these problems, we are able to relate the behavior of the different classifier schemes to the geometrical complexity of the problem. Moreover, the results highlight certain regions of the complexity measurement space where a classifier scheme excels, establishing a first step toward determining the best classifier scheme for a given classification problem. Ester Bernadó-Mansilla, Tin Kam Ho |
IEEE Trans. Evol. Comput. | 2 |
| 2002 | A Data Complexity Analysis of Comparative Advantages of Decision Forest Constructors
Tin Kam Ho |
Pattern Anal. Appl. | 1 |
| 2002 | Complexity Measures of Supervised Classification ProblemsabstractWe studied a number of measures that characterize the difficulty of a classification problem, focusing on the geometrical complexity of the class boundary. We compared a set of real-world problems to random labelings of points and found that real problems contain structures in this measurement space that are significantly different from the random sets. Distributions of problems in this space show that there exist at least two independent factors affecting a problem's difficulty. We suggest using this space to describe a classifier's domain of competence. This can guide static and dynamic selection of classifiers for specific problems as well as subproblems formed by confinement, projection, and transformations of the feature vectors. Tin Kam Ho, Mitra Basu |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2001 | Exploration of Contextual Constraints for Character Pre-ClassificationabstractWe present strategies and results for identifying the symbol type (lower-case, upper-case, digit, and punctuation or special symbols) of every character in a text document by using various kinds of information from neighboring characters. In the expectation of reasonable word and character segmentation for shape clustering, we designed several type recognition methods that depend on cluster n-grams, shape codes, and within word context. On an ASCII test corpus of 925 articles that simulates perfect image-level processing, these methods achieve a substantial improvement over default assignment of all characters to lower case. Tin Kam Ho, George Nagy |
ICDAR | 1 |
| 2000 | Measuring the Complexity of Classification ProblemsabstractWe study a number of measures that characterize the difficulty of a classification problem. We compare a set of real world problems to random combinations of points in this measurement space and found that real problems contain structures that are significantly different from the random sets. Distribution of problems in this space reveals that there exist at least two independent factors affecting a problem's difficulty, and that they have notable joint effects. We suggest using this space to describe a classifier domain of competence. This can guide static and dynamic selection of classifiers for specific problems as well as sub-problems formed by confinement, projections, and transformations of the feature vectors. Tin Kam Ho, Mitra Basu |
ICPR | 1 |
| 2000 | OCR with No Shape TrainingabstractWe present a document-specific OCR system and apply it to a corpus of fixed business letters. Unsupervised classification of the segmented character bitmaps on each page, using a "clump" metric, typically yields several hundred clusters with highly skewed populations. Letter identities are assigned to each cluster by maximizing matches with a lexicon of English words. We found that for 2/3 of the pages, we can identify almost 80% of the words included in the lexicon, without any shape training. Residual errors are caused by mis-segmentation including missed lines and punctuation. This research differs from earlier attempts to apply cipher decoding to OCR in: (1) using real data; (2) a more appropriate clustering algorithm; and (3) decoding a many-to-many instead of a one-to-one mapping between clusters and letters. Tin Kam Ho, George Nagy |
ICPR | 1 |
| 2000 | Stop word location and identification for adaptive text recognition
Tin Kam Ho |
Int. J. Document Anal. Recognit. | 1 |
| 1999 | Fast Identification of Stop Words for Font Learning and Keyword SpottingabstractA recently proposed adaptive strategy for text recognition uses a linguistic fact that over half of the words on a typical English page are among 150 common stop words. The small lexicon permits word-shape based recognition that yields word identities from which character prototypes can be extracted. This paper describes a fast procedure for locating the best candidates for those stop words. The procedure uses width statistics of individual words and their immediate neighbors. In an experiment using 400 page images, the method removed 63% of the words from consideration. The stop/nonstop word discrimination also assists keyword spotting for information retrieval. Tin Kam Ho |
ICDAR | 1 |
| 1999 | The learning behavior of single neuron classifiers on linearly separable or nonseparable inputabstractDetermining linear separability is an important way of understanding structures present in data. We explore the behavior of several classical descent procedures for determining linear separability and training linear classifiers in the presence of linearly nonseparable input. We compare the adaptive procedures to linear programming methods using many pairwise discrimination problems from a public database. We found that the adaptive procedures have serious implementation problems which make them less preferable than linear programming. Mitra Basu, Tin Kam Ho |
IJCNN | 2 |
| 1998 | C4.5 decision forestsabstractMuch of previous attention on decision trees focuses on the splitting criteria and optimization of tree sizes. The dilemma between overfitting and achieving maximum accuracy is seldom resolved. We propose a method to construct a decision tree based classifier that maintains highest accuracy on training data and improves on generalization accuracy as it grows in complexity. Trees are generated using the well-known C4.5 algorithm, and the classifier consists of multiple trees constructed in pseudo-randomly selected subspaces of the given feature space. We compare the method to single-tree classifiers and other forest construction methods by experiments on four public data sets, where the method's superiority is demonstrated. A measure is given to describe the similarity between trees in a forest, and is related to the combined classification accuracy. Tin Kam Ho |
ICPR | 1 |
| 1998 | Bootstrapping text recognition from stop wordsabstractRecognition of arbitrary noisy English text has been difficult because of problems in character segmentation and multi-font symbol classification. Both segmentation and recognition can be easier with more knowledge of the dominant font used in a given text page. This has led to some recent studies that show promising methods for extracting character prototypes from a text image provided that truth is given for part of the image. In this paper we investigate the feasibility of such a strategy without dependence on ground truth. We replace the needed truth by results of direct recognition of some frequently occurring words. The method makes use of the observation that over half of the words in a typical English text passage are contained in a very small lexicon. Tin Kam Ho |
ICPR | 1 |
| 1998 | Pattern Classification with Compact Distribution Maps
Tin Kam Ho, Henry S. Baird |
Comput. Vis. Image Underst. | 1 |
| 1998 | The Random Subspace Method for Constructing Decision ForestsabstractMuch of previous attention on decision trees focuses on the splitting criteria and optimization of tree sizes. The dilemma between overfitting and achieving maximum accuracy is seldom resolved. A method to construct a decision tree based classifier is proposed that maintains highest accuracy on training data and improves on generalization accuracy as it grows in complexity. The classifier consists of multiple trees constructed systematically by pseudorandomly selecting subsets of components of the feature vector, that is, trees constructed in randomly chosen subspaces. The subspace method is compared to single-tree classifiers and other forest construction methods by experiments on publicly available datasets, where the method's superiority is demonstrated. We also discuss independence between trees in a forest and relate that to the combined classification accuracy. Tin Kam Ho |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 1997 | Enhancing Degraded Document Images via Bitmap Clustering and AveragingabstractProper display and accurate recognition of document images are often hampered by degradations caused by poor scanning or transmission conditions. The authors propose a method to enhance such degraded document images for better display quality and recognition accuracy. The essence of the method is in finding and averaging bitmaps of the same symbol that are scattered across a text page. Outline descriptions of the symbols are then obtained that can be rendered at arbitrary solution. The paper describes details of the algorithm and an experiment to demonstrate its capabilities using fax images. John D. Hobby, Tin Kam Ho |
ICDAR | 2 |
| 1997 | Large-Scale Simulation Studies in Image Pattern RecognitionabstractMany obstacles to progress in image pattern recognition result from the fact that per-class distributions are often too irregular to be well-approximated by simple analytical functions. Simulation studies offer one way to circumvent these obstacles. We present three closely related studies of machine-printed character recognition that rely on synthetic data generated pseudo-randomly in accordance with an explicit stochastic model of document image degradations. The unusually large scale of experiments - involving several million samples that makes this methodology possible have allowed us to compute sharp estimates of the intrinsic difficulty (Bayes risk) of concrete image recognition problems, as well as the asymptotic accuracy and domain of competency of classifiers. Tin Kam Ho, Henry S. Baird |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 1996 | Adaptive Coordination of Multiple ClassifiersabstractIntroduction The advantages of using multiple classifiers in parallel in a document analysis system have been realized in recent years. 3 It has become well known that using combined decisions often results in improved accuracy. However, since using multiple classifiers incurs significant additional costs, it is less than ideal in a production environment demanding high speed. The combination strategies would be more interesting if they can adapt to certain conditions of the input, so that additional classifiers would be invoked only if necessary. To design adaptive coordination strategies, it is important to analyze the classifiers' performance in response to different conditions of the input. Among the characteristics of the input that could affect classification accuracy, image quality is often believed to be the most important. In practice, however, analyses have been hampered by the lack of a systematic way to measure and describe image quality. In addi Tin Kam Ho |
DAS | 1 |
| 1996 | Commercial Exploitation of OCR and Document Analysis Systems Technology: Market and User Requirements: das'96 Working Group Report
Stefan Knerr, Tin Kam Ho |
DAS | 2 |
| 1996 | Building projectable classifiers of arbitrary complexityabstractConventional methods for classifier design often suffer from having two conflicting goals-to develop arbitrarily complex decision boundaries to suit a given problem, and at the same time to constrain the complexity of those boundaries to avoid overfitting given training data. A recent analysis reveals that the conflict is resolvable by building classifiers based on projectable elements, which are weak discriminators that perform equally well for both training and testing data. Based on this analysis, we present a method that constructs a classifier up to arbitrary complexity while presenting generalization accuracy. Tin Kam Ho, Eugene M. Kleinberg |
ICPR | 1 |
| 1995 | Random decision forestsabstractDecision trees are attractive classifiers due to their high execution speed. But trees derived with traditional methods often cannot be grown to arbitrary complexity for possible loss of generalization accuracy on unseen data. The limitation on complexity usually means suboptimal accuracy on training data. Following the principles of stochastic modeling, we propose a method to construct tree-based classifiers whose capacity can be arbitrarily expanded for increases in accuracy for both training and unseen data. The essence of the method is to build multiple trees in randomly selected subspaces of the feature space. Trees in, different subspaces generalize their classification in complementary ways, and their combined classification can be monotonically improved. The validity of the method is demonstrated through experiments on the recognition of handwritten digits. Tin Kam Ho |
ICDAR | 1 |
| 1994 | Estimating the intrinsic difficulty of a recognition problemabstractDescribes an experiment in estimating the Bayes error of an image classification problem: a difficult, practically important, two-class character recognition problem. The Bayes error gives the "intrinsic difficulty" of the problem since it is the minimum error achievable by any classification method. Since for many realistically complex problems, deriving this analytically appears to be hopeless, the authors approach the task empirically. The authors proceed first by expressing the problem precisely in terms of ideal prototype images and an image defect model, and then by carrying out the estimation on pseudorandomly simulated data. Arriving at sharp estimates seems inevitably to require both large sample sizes-in the authors' trial, over a million images-and careful statistical extrapolation. The study of the data reveals many interesting statistics, which allow the prediction of the worst-case time/space requirements for any given classifier performance, expressed as a combination of error and reject rates. Tin Kam Ho, Henry S. Baird |
ICPR (2) | 1 |
| 1994 | Decision Combination in Multiple Classifier SystemsabstractA multiple classifier system is a powerful solution to difficult pattern recognition problems involving large class sets and noisy input because it allows simultaneous use of arbitrary feature descriptors and classification procedures. Decisions by the classifiers can be represented as rankings of classifiers and different instances of a problem. The rankings can be combined by methods that either reduce or rerank a given set of classes. An intersection method and union method are proposed for class set reduction. Three methods based on the highest rank, the Borda count, and logistic regression are proposed for class set reranking. These methods have been tested in applications of degraded machine-printed characters and works from large lexicons, resulting in substantial improvement in overall correctness.> Tin Kam Ho, Jonathan J. Hull, Sargur N. Srihari |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 1993 | Recognition of handwritten digits by combining independent learning vector quantizationsabstractClassifiers derived by learning vector quantization (LVQ) have well-defined decision regions that can be combined to construct a more accurate classifier. A given point may be included in a number of decision regions associated with different LVQ classifiers. The relative densities of classes in each region can be combined to obtain a final classification. The method allows useful inferences from small training sets, which is needed for problems involving large variations within each class. In an application of this method to the recognition of handwritten digits, it is shown that the classifier can be improved almost monotonically without suffering from over-adaptation to the training data.> Tin Kam Ho |
ICDAR | 1 |
| 1993 | Perfect metricsabstractThe authors describe an experiment in the construction of perfect metrics for minimum-distance classification of character images. A perfect metric is one that, with high probability, is zero for correct classifications and non-zero for incorrect classifications. They promise excellent reject behavior in addition to good rank ordering. The approach is to infer from the training data faithful but concise representations of the empirical class-conditional distributions. In doing this, the authors have abandoned many visual simplifying assumptions about the distributions, e.g., that they are simply-connected, unimodal, convex, or parametric (e.g., Gaussian). The method requires unusually large and representative training sets, which we provide through pseudorandom generation of training samples using a realistic model of printing and imaging distortions. The authors illustrate the method on a challenging recognition problem: 3755 character classes of machine-print Chinese, in four typefaces, over a range of text sizes. In a test on over three million images, the perfect-metric classifier achieved better than 99% top-choice accuracy. In addition, it is shown that it is superior to a conventional parametric classifier.> Tin Kam Ho, Henry S. Baird |
ICDAR | 1 |
| 1992 | On multiple classifier systems for pattern recognitionabstractDifficult pattern recognition problems involving large class sets and noisy input can be solved by a multiple classifier system, which allows simultaneous use of arbitrary feature descriptors and classification procedures. Independent decisions by each classifier can be combined by methods of the highest rank, Borda count, and logistic regression, resulting in substantial improvement in overall correctness.> Tin Kam Ho, Jonathan J. Hull, Sargur N. Srihari |
ICPR (2) | 1 |
| 1992 | World image matching as a technique for degraded text recognitionabstractA technique is presented that determines equivalences between word images in a passage of text. A clustering procedure is applied to group visually similar words. Initial hypotheses for the identities of words are then generated by matching the word groups to language statistics that predict the frequency at which certain words will occur. This is followed by a recognition step that assigns identifications to the images in the clusters. This paper concentrates on the clustering algorithm. A clustering technique is presented and its performance on a running text of 1062 word images is determined. It is shown that the clustering algorithm can correctly locate groups of short function words with better than a 95 percent correct rate.> Jonathan J. Hull, Siamak Khoubyari, Tin Kam Ho |
ICPR (2) | 3 |
| 1992 | A hypothesis testing approach to word recognition using dynamic feature selectionabstractA top-down approach to word recognition is proposed. Discussions are presented on dynamically selecting the most effective feature combinations, which are applied to discriminate between a limited set of word hypotheses.> Tin Kam Ho, Jonathan J. Hull, Sargur N. Srihari |
ICPR (2) | 2 |
| 1992 | A computational model for recognition of multifont word images
Tin Kam Ho, Jonathan J. Hull, Sargur N. Srihari |
Mach. Vis. Appl. | 1 |
| 1992 | A word shape analysis approach to lexicon based word recognition
Tin Kam Ho, Jonathan J. Hull, Sargur N. Srihari |
Pattern Recognit. Lett. | 1 |