EDBT 2026 Demo / reviewers in the wild / expert
Frederick H. Lochovsky
dblp:l/FHLochovsky
· DBLP profile ↗
41ranked-venue papers
4as first author
0since 2021 · last 2013
0000-0001-7470-4172ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 29 · 3 first-authorArtificial intelligence and machine learning · 9Human-computer interaction and ubiquitous computing · 3Applied, interdisciplinary, general and emerging computing · 2Systems, architecture and hardware · 1Security and privacy · 1 · 1 first-authorSoftware engineering, systems software and programming languages · 1Graphics, computer vision, multimedia, augmented reality and games · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Databases, data mining, and information retrieval
13 papers |
Data integration and cleaning · 71% Information retrieval · 15% Knowledge graphs · 6% | |
| Artificial intelligence
5 papers |
Kernel, tree and ensemble methods · 34% Optimization for machine learning · 32% Learning paradigms · 16% |
Topics — the 30 heaviest of 45, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Data integration and cleaning
data extraction |
0.3 | 3 | 2012 | Combining Tag and Value Similarity for Data Extraction and Alignment · IEEE Trans. Knowl. Data Eng. 2012 ODE: Ontology-assisted data extraction · ACM Trans. Database Syst. 2009 Data extraction and label assignment for web databases · WWW 2003 |
Data integration and cleaning › entity resolution
record matching |
0.3 | 2 | 2012 | Combining Tag and Value Similarity for Data Extraction and Alignment · IEEE Trans. Knowl. Data Eng. 2012 Record Matching over Query Results from Multiple Web Databases · IEEE Trans. Knowl. Data Eng. 2010 |
Data integration and cleaning › data extraction
web data extraction |
0.1 | 4 | 2012 | Combining Tag and Value Similarity for Data Extraction and Alignment · IEEE Trans. Knowl. Data Eng. 2012 Data extraction and label assignment for web databases · WWW 2003 Record Matching over Query Results from Multiple Web Databases · IEEE Trans. Knowl. Data Eng. 2010 |
Data integration and cleaning › schema matching
attribute correspondence |
0.1 | 1 | 2012 | Combining Tag and Value Similarity for Data Extraction and Alignment · IEEE Trans. Knowl. Data Eng. 2012 |
Machine learning › Optimization for machine learning
hyperparameter optimization |
0.1 | 2 | 2007 | A kernel path algorithm for support vector machines · ICML 2007 Two-dimensional solution path for support vector regression · ICML 2006 |
Machine learning › Kernel, tree and ensemble methods
support vector machine |
0.1 | 2 | 2007 | A kernel path algorithm for support vector machines · ICML 2007 Two-dimensional solution path for support vector regression · ICML 2006 |
Machine learning › Optimization for machine learning
solution path |
0.1 | 2 | 2006 | Two-dimensional solution path for support vector regression · ICML 2006 Solution Path for Semi-Supervised Classification with Manifold Regularization · ICDM 2006 |
Data integration and cleaning
schema matching |
0.1 | 2 | 2006 | Holistic Query Interface Matching using Parallel Schema Matching · ICDE 2006 Instance-based Schema Matching for Web Databases by Domain-specific Query Probing · VLDB 2004 |
Data integration and cleaning › entity resolution
duplicate detection |
0.1 | 1 | 2010 | Record Matching over Query Results from Multiple Web Databases · IEEE Trans. Knowl. Data Eng. 2010 |
Information retrieval › query understanding › query classification
query intent classification |
0.1 | 1 | 2009 | Understanding user's query intent with wikipedia · WWW 2009 |
Information retrieval
query understanding |
0.1 | 1 | 2009 | Understanding user's query intent with wikipedia · WWW 2009 |
Machine learning › Kernel, tree and ensemble methods
kernel selection |
0.1 | 1 | 2007 | A kernel path algorithm for support vector machines · ICML 2007 |
Machine learning › Learning paradigms › semi-supervised learning › graph-based semi-supervised learning
manifold regularization |
0.1 | 1 | 2006 | Solution Path for Semi-Supervised Classification with Manifold Regularization · ICDM 2006 |
Machine learning › Learning paradigms
semi-supervised learning |
0.1 | 1 | 2006 | Solution Path for Semi-Supervised Classification with Manifold Regularization · ICDM 2006 |
Machine learning › Kernel, tree and ensemble methods › support vector machine
support vector regression |
0.1 | 1 | 2006 | Two-dimensional solution path for support vector regression · ICML 2006 |
Data integration and cleaning › schema matching
query interface matching |
0.1 | 1 | 2006 | Holistic Query Interface Matching using Parallel Schema Matching · ICDE 2006 |
Machine learning › Probabilistic and Bayesian machine learning › structured models
latent variable model |
0.1 | 1 | 2005 | A Bernoulli Relational Model for Nonlinear Embedding · ICDM 2005 |
Machine learning › Representation and self-supervised learning › representation learning › dimensionality reduction › manifold learning › neighbor embedding
stochastic neighbor embedding |
0.1 | 1 | 2005 | A Bernoulli Relational Model for Nonlinear Embedding · ICDM 2005 |
Data mining
dimensionality reduction |
0.1 | 1 | 2005 | A Bernoulli Relational Model for Nonlinear Embedding · ICDM 2005 |
Information retrieval
hidden web database |
0.0 | 1 | 2003 | Data extraction and label assignment for web databases · WWW 2003 |
Data integration and cleaning › data extraction › web data extraction
wrapper generation |
0.0 | 1 | 2003 | Data extraction and label assignment for web databases · WWW 2003 |
Data models and query languages
object-oriented data model |
0.0 | 2 | 1998 | ADOME: An Advanced Object Modeling Environment · IEEE Trans. Knowl. Data Eng. 1998 A Data Model and Semantics of Objects with Dynamic Roles · ICDE 1997 |
Knowledge, reasoning and agents › Knowledge representation and reasoning › knowledge-based systems › rule-based systems
production rules |
0.0 | 1 | 1998 | ADOME: An Advanced Object Modeling Environment · IEEE Trans. Knowl. Data Eng. 1998 |
Knowledge, reasoning and agents › Knowledge representation and reasoning › knowledge-based systems › rule-based systems
rule-based reasoning |
0.0 | 1 | 1998 | ADOME: An Advanced Object Modeling Environment · IEEE Trans. Knowl. Data Eng. 1998 |
Data models and query languages
object model |
0.0 | 1 | 1998 | ADOME: An Advanced Object Modeling Environment · IEEE Trans. Knowl. Data Eng. 1998 |
Database system architecture and tuning
web databases |
0.0 | 1 | 2004 | Instance-based Schema Matching for Web Databases by Domain-specific Query Probing · VLDB 2004 |
Information retrieval
document retrieval |
0.0 | 1 | 1990 | HYTREM - A Hybrid Text-Retrieval Machine for Large Databases · IEEE Trans. Computers 1990 |
Information retrieval › document retrieval › text search
hardware systems for text retrieval |
0.0 | 1 | 1990 | HYTREM - A Hybrid Text-Retrieval Machine for Large Databases · IEEE Trans. Computers 1990 |
Programming languages and type systems
language semantics |
0.0 | 1 | 1997 | A Data Model and Semantics of Objects with Dynamic Roles · ICDE 1997 |
Distributed systems › distributed coordination › multi-agent systems
distributed problem solving |
0.0 | 1 | 1986 | Supporting Distributed Office Problem Solving in Organizations · ACM Trans. Inf. Syst. 1986 |
Methods — techniques the papers use, named apart from their topics
support vector machine · 0.2value similarity · 0.1tag similarity · 0.1weighted component similarity · 0.1iterative classification · 0.1wikipedia-based concept mapping · 0.1record segmentation · 0.1ontology matching · 0.1machine learning · 0.1solution path · 0.1breakpoint approximation · 0.1regularization path · 0.1piecewise linear path · 0.1manifold regularization · 0.1count-based greedy algorithm · 0.1relational model · 0.1gaussian process latent variable model · 0.1role facilities · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2013 | Understanding query interfaces by statistical parsingabstractUsers submit queries to an online database via its query interface. Query interface parsing, which is important for many applications, understands the query capabilities of a query interface. Since most query interfaces are organized hierarchically, we present a novel query interface parsing method, StatParser (Statistical Parser), to automatically extract the hierarchical query capabilities of query interfaces. StatParser automatically learns from a set of parsed query interfaces and parses new query interfaces. StatParser starts from a small grammar and enhances the grammar with a set of probabilities learned from parsed query interfaces under the maximum-entropy principle. Given a new query interface, the probability-enhanced grammar identifies the parse tree with the largest global probability to be the query capabilities of the query interface. Experimental results show that StatParser very accurately extracts the query capabilities and can effectively overcome the problems of existing query interface parsers. Weifeng Su, Hejun Wu, Frederick H. Lochovsky, Hongmin Cai |
ACM Trans. Web | 5 |
| 2012 | Combining Tag and Value Similarity for Data Extraction and AlignmentabstractWeb databases generate query result pages based on a user's query. Automatically extracting the data from these query result pages is very important for many applications, such as data integration, which need to cooperate with multiple web databases. We present a novel data extraction and alignment method called CTVS that combines both tag and value similarity. CTVS automatically extracts data from query result pages by first identifying and segmenting the query result records (QRRs) in the query result pages and then aligning the segmented QRRs into a table, in which the data values from the same attribute are put into the same column. Specifically, we propose new techniques to handle the case when the QRRs are not contiguous, which may be due to the presence of auxiliary information, such as a comment, recommendation or advertisement, and for handling any nested structure that may exist in the QRRs. We also design a new record alignment algorithm that aligns the attributes in a record, first pairwise and then holistically, by combining the tag and data value similarity information. Experimental results show that CTVS achieves high precision and outperforms existing state-of-the-art data extraction methods. Weifeng Su, Jiying Wang, Frederick H. Lochovsky |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2012 | Solution Path for Manifold Regularized Semisupervised ClassificationabstractTraditional learning algorithms use only labeled data for training. However, labeled examples are often difficult or time consuming to obtain since they require substantial human labeling efforts. On the other hand, unlabeled data are often relatively easy to collect. Semisupervised learning addresses this problem by using large quantities of unlabeled data with labeled data to build better learning algorithms. In this paper, we use the manifold regularization approach to formulate the semisupervised learning problem where a regularization framework which balances a tradeoff between loss and penalty is established. We investigate different implementations of the loss function and identify the methods which have the least computational expense. The regularization hyperparameter, which determines the balance between loss and penalty, is crucial to model selection. Accordingly, we derive an algorithm that can fit the entire path of solutions for every value of the hyperparameter. Its computational complexity after preprocessing is quadratic only in the number of labeled examples rather than the total number of labeled and unlabeled examples. Gang Wang 0004, Fei Wang 0001, Dit-Yan Yeung, Frederick H. Lochovsky |
IEEE Trans. Syst. Man Cybern. Part B | 5 |
| 2010 | A regularization framework for multiclass classification: A deterministic annealing approach
Zhihua Zhang 0004, Gang Wang 0004, Dit-Yan Yeung, Guang Dai, Frederick H. Lochovsky |
Pattern Recognit. | 5 |
| 2010 | Record Matching over Query Results from Multiple Web DatabasesabstractRecord matching, which identifies the records that represent the same real-world entity, is an important step for data integration. Most state-of-the-art record matching methods are supervised, which requires the user to provide training data. These methods are not applicable for the Web database scenario, where the records to match are query results dynamically generated on-the-fly. Such records are query-dependent and a prelearned method using training examples from previous query results may fail on the results of a new query. To address the problem of record matching in the Web database scenario, we present an unsupervised, online record matching method, UDD, which, for a given query, can effectively identify duplicates from the query result records of multiple Web databases. After removal of the same-source duplicates, the ¿presumed¿ nonduplicate records from the same source can be used as training examples alleviating the burden of users having to manually label training examples. Starting from the nonduplicate set, we use two cooperating classifiers, a weighted component similarity summing classifier and an SVM classifier, to iteratively identify duplicates in the query results from multiple Web databases. Experimental results show that UDD works well for the Web database scenario where existing supervised methods do not apply. Weifeng Su, Jiying Wang, Frederick H. Lochovsky |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2009 | Understanding user's query intent with wikipediaabstractUnderstanding the intent behind a user's query can help search engine to automatically route the query to some corresponding vertical search engines to obtain particularly relevant contents, thus, greatly improving user satisfaction. There are three major challenges to the query intent classification problem: (1) Intent representation; (2) Domain coverage and (3) Semantic interpretation. Current approaches to predict the user's intent mainly utilize machine learning techniques. However, it is difficult and often requires many human efforts to meet all these challenges by the statistical machine learning approaches. In this paper, we propose a general methodology to the problem of query intent classification. With very little human effort, our method can discover large quantities of intent concepts by leveraging Wikipedia, one of the best human knowledge base. The Wikipedia concepts are used as the intent representation space, thus, each intent domain is represented as a set of Wikipedia articles and categories. The intent of any input query is identified through mapping the query into the Wikipedia representation space. Compared with previous approaches, our proposed method can achieve much better coverage to classify queries in an intent domain even through the number of seed intent examples is very small. Moreover, the method is very general and can be easily applied to various intent domains. We demonstrate the effectiveness of this method in three different applications, i.e., travel, job, and person name. In each of the three cases, only a couple of seed intent queries are provided. We perform the quantitative evaluations in comparison with two baseline methods, and the experimental results shows that our method significantly outperforms other methods in each intent domain. Jian Hu 0001, Gang Wang 0004, Frederick H. Lochovsky, Jian-Tao Sun, Zheng Chen 0001 |
WWW | 3 |
| 2009 | ODE: Ontology-assisted data extractionabstractOnline databases respond to a user query with result records encoded in HTML files. Data extraction, which is important for many applications, extracts the records from the HTML files automatically. We present a novel data extraction method, ODE (Ontology-assisted Data Extraction), which automatically extracts the query result records from the HTML pages. ODE first constructs an ontology for a domain according to information matching between the query interfaces and query result pages from different Web sites within the same domain. Then, the constructed domain ontology is used during data extraction to identify the query result section in a query result page and to align and label the data values in the extracted records. The ontology-assisted data extraction method is fully automatic and overcomes many of the deficiencies of current automatic data extraction methods. Experimental results show that ODE is extremely accurate for identifying the query result section in an HTML page, segmenting the query result section into query result records, and aligning and labeling the data values in the query result records. Weifeng Su, Jiying Wang, Frederick H. Lochovsky |
ACM Trans. Database Syst. | 3 |
| 2008 | A New Solution Path Algorithm in Support Vector RegressionabstractIn this paper, regularization path algorithms were proposed as a novel approach to the model selection problem by exploring the path of possibly all solutions with respect to some regularization hyperparameter in an efficient way. This approach was later extended to a support vector regression (SVR) model called epsilon-SVR. However, the method requires that the error parameter epsilon be set a priori. This is only possible if the desired accuracy of the approximation can be specified in advance. In this paper, we analyze the solution space for epsilon-SVR and propose a new solution path algorithm, called epsilon-path algorithm, which traces the solution path with respect to the hyperparameter epsilon rather than lambda. Although both two solution path algorithms possess the desirable piecewise linearity property, our epsilon-path algorithm overcomes some limitations of the original lambda-path algorithm and has more advantages. It is thus more appealing for practical use. Gang Wang 0004, Dit-Yan Yeung, Frederick H. Lochovsky |
IEEE Trans. Neural Networks | 3 |
| 2007 | A kernel path algorithm for support vector machinesabstractThe choice of the kernel function which determines the mapping between the input space and the feature space is of crucial importance to kernel methods. The past few years have seen many efforts in learning either the kernel function or the kernel matrix. In this paper, we address this model selection issue by learning the hyperparameter of the kernel function for a support vector machine (SVM). We trace the solution path with respect to the kernel hyperparameter without having to train the model multiple times. Given a kernel hyperparameter value and the optimal solution obtained for that value, we find that the solutions of the neighborhood hyperparameters can be calculated exactly. However, the solution path does not exhibit piecewise linearity and extends nonlinearly. As a result, the breakpoints cannot be computed in advance. We propose a method to approximate the breakpoints. Our method is both efficient and general in the sense that it can be applied to many kernel functions in common use. Gang Wang 0004, Dit-Yan Yeung, Frederick H. Lochovsky |
ICML | 3 |
| 2006 | Query result ranking over e-commerce web databasesabstractTo deal with the problem of too many results returned from an E-commerce Web database in response to a user query, this paper proposes a novel approach to rank the query results. Based on the user query, we speculate how much the user cares about each attribute and assign a corresponding weight to it. Then, for each tuple in the query result, each attribute value is assigned a score according to its "desirableness" to the user. These attribute value scores are combined according to the attribute weights to get a final ranking score for each tuple. Tuples with the top ranking scores are presented to the user first. Our ranking method is domain independent and requires no user feedback. Experimental results demonstrate that this ranking method can effectively capture a user's preferences. Weifeng Su, Jiying Wang, Frederick H. Lochovsky |
CIKM | 4 |
| 2006 | Holistic Schema Matching for Web Query Interfaces
Weifeng Su, Jiying Wang, Frederick H. Lochovsky |
EDBT | 3 |
| 2006 | Holistic Query Interface Matching using Parallel Schema MatchingabstractUsing query interfaces of different Web databases, we propose a new complex schema matching approach, Parallel Schema Matching (PSM). A parallel schema is formed by comparing two individual schemas and deleting common attributes. The attribute matching can be discovered from the attribute-occurrence patterns if many parallel schemas are available. A count-based greedy algorithm identifies which attributes are more likely to be matched. Experiments show that PSM can identify both simple matching and complex matching accurately and efficiently. Weifeng Su, Jiying Wang, Frederick H. Lochovsky |
ICDE | 3 |
| 2006 | Solution Path for Semi-Supervised Classification with Manifold RegularizationabstractWith very low extra computational cost, the entire solution path can be computed for various learning algorithms like support vector classification (SVC) and support vector regression (SVR). In this paper, we extend this promising approach to semi-supervised learning algorithms. In particular, we consider finding the solution path for the Laplacian support vector machine (LapSVM) which is a semi-supervised classification model based on manifold regularization. One advantage of the this algorithm is that the coefficient path is piecewise linear with respect to the regularization parameter, hence its computational complexity is quadratic in the number of labeled examples. Gang Wang 0004, Dit-Yan Yeung, Frederick H. Lochovsky |
ICDM | 4 |
| 2006 | Two-dimensional solution path for support vector regressionabstractRecently, a very appealing approach was proposed to compute the entire solution path for support vector classification (SVC) with very low extra computational cost. This approach was later extended to a support vector regression (SVR) model called ε-SVR. However, the method requires that the error parameter ε be set a priori, which is only possible if the desired accuracy of the approximation can be specified in advance. In this paper, we show that the solution path for ε-SVR is also piecewise linear with respect to ε. We further propose an efficient algorithm for exploring the two-dimensional solution space defined by the regularization and error parameters. As opposed to the algorithm for SVC, our proposed algorithm for ε-SVR initializes the number of support vectors to zero and then increases it gradually as the algorithm proceeds. As such, a good regression function possessing the sparseness property can be obtained after only a few iterations. Gang Wang 0004, Dit-Yan Yeung, Frederick H. Lochovsky |
ICML | 3 |
| 2006 | Automatic Hierarchical Classification of Structured Deep Web Databases
Weifeng Su, Jiying Wang, Frederick H. Lochovsky |
WISE | 3 |
| 2005 | Annealed Discriminant Analysis
Gang Wang 0004, Zhihua Zhang 0004, Frederick H. Lochovsky |
ECML | 3 |
| 2005 | A Bernoulli Relational Model for Nonlinear EmbeddingabstractThe notion of relations is extremely important in mathematics. In this paper, we use relations to describe the embedding problem and propose a novel stochastic relational model for nonlinear embedding. Given some relation among points in a high-dimensional space, we start from preserving the same relation in a low embedded space and model the relation as probabilistic distributions over these two spaces, respectively. We illustrate that the stochastic neighbor embedding and the Gaussian process latent variable model can be derived from our relational model. Moreover we devise a new stochastic embedding model and refer to it as Bernoulli relational embedding (BRE). BRE's ability in nonlinear dimensionality reduction is illustrated on a set of synthetic data and collections of bitmaps of handwritten digits and face images. Gang Wang 0004, Zhihua Zhang 0004, Frederick H. Lochovsky |
ICDM | 4 |
| 2005 | Latent Process Model for Manifold LearningabstractIn this paper, we propose a novel stochastic framework for unsupervised manifold learning. The latent variables are introduced, and the latent processes are assumed to characterize the pairwise relations of points over a high dimensional and a low dimensional space. The elements in the embedding space are obtained by minimizing the divergence between the latent processes over the two spaces. Different priors of the latent variables, such as Gaussian and multinominal, are examined. The Kullback-Leibler divergence and the Bhattachartyya distance are investigated. The latent process model incorporates some existing embedding methods and gives a clear view on the properties of each method. The embedding ability of this latent process model is illustrated on a collection of bitmaps of handwritten digits and on a set of synthetic data Gang Wang 0004, Weifeng Su, Xiangye Xiao, Frederick H. Lochovsky |
ICTAI | 4 |
| 2004 | Feature selection with conditional mutual information maximin in text categorizationabstractFeature selection is an important component of text categorization. This technique can both increase a classifier's computation speed, and reduce the overfitting problem. Several feature selection methods, such as information gain and mutual information, have been widely used. Although they greatly improve the classifier's performance, they have a common drawback, which is that they do not consider the mutual relationships among the features. In this situation, where one feature's predictive power is weakened by others, and where the selected features tend to bias towards major categories, such selection methods are not very effective. In this paper, we propose a novel feature selection method for text categorization called conditional mutual information maximin (CMIM). It can select a set of individually discriminating and weakly dependent features. The experimental results show that CMIM can perform much better than traditional feature selection methods. Gang Wang 0004, Frederick H. Lochovsky |
CIKM | 2 |
| 2004 | Instance-based Schema Matching for Web Databases by Domain-specific Query Probing
Jiying Wang, Ji-Rong Wen, Frederick H. Lochovsky, Wei-Ying Ma |
VLDB | 3 |
| 2004 | AC-Tree: An Adaptive Structural Join Index
Kaiyang Liu, Frederick H. Lochovsky |
WISE | 2 |
| 2004 | Ranking Search Results by Web Quality Dimensions
Joshua C. C. Pun, Frederick H. Lochovsky |
J. Web Eng. | 2 |
| 2003 | Efficient Computation of Aggregate Structural JoinsabstractAlthough several XML query-processing techniques have been proposed to do structural joins, no work has been done to handle aggregate structural join, which computes the aggregate value for two given XML element sets. One important kind of aggregate structural join is to determine the exact selectivity of a structural join, which proves to be crucial for overall path expression optimization and on-line systems. Related work has been done to estimate the selectivity of a path expression by applying summary techniques for XML data based on either tree simplification or a Markov model. However, because of either space restrictions or performance issues (for example, previous approaches may need to store the whole XML tree to answer an aggregate structural join), they are not applicable for aggregate structural joins. In this paper, we consider a modified aggregate B-tree, which we call a XA-tree, to index the aggregate attribute values of XML elements. Since it is based on the B-tree, the XA-tree supports query and update efficiently and places little additional burden on database administrators. Two basic algorithms are proposed to compute aggregate structural joins efficiently by utilizing the XA-tree. Furthermore, a heuristic is also proposed to further improve the performance significantly. Extensive experiments have confirmed that our approaches greatly outperform competitors by an order of magnitude. Kaiyang Liu, Frederick H. Lochovsky |
WISE | 2 |
| 2003 | Data extraction and label assignment for web databasesabstractMany tools have been developed to help users query, extract and integrate data from web pages generated dynamically from databases, i.e., from the Hidden Web. A key prerequisite for such tools is to obtain the schema of the attributes of the retrieved data. In this paper, we describe a system called, DeLa, which reconstructs (part of) a "hidden" back-end web database. It does this by sending queries through HTML forms, automatically generating regular expression wrappers to extract data objects from the result pages and restoring the retrieved data into an annotated (labelled) table. The whole process needs no human involvement and proves to be fast (less than one minute for wrapper induction for each site) and accurate (over 90% correctness for data extraction and around 80% correctness for label assignment). Jiying Wang, Frederick H. Lochovsky |
WWW | 2 |
| 2002 | Data-rich Section Extraction from HTML pagesabstractWe propose a novel algorithm, DSE (data-rich subtree extraction) to recognize and extract the data-rich section of an HTML page. We apply the DSE algorithm as a pre-processing "clean-up" step for two typical Web information retrieval problems: topic distillation and Web information extraction. Our experiments show that, for the test data sets used, the DSE algorithm can correctly identify data-rich sections of HTML pages with 100% accuracy. Therefore, it can effectively reduce the root set size for the topic distillation problem thereby improving the precision and accuracy of the IETS algorithm. Furthermore, when applied to the Web information extraction problem using the IEPAD algorithm, it can decrease the number of patterns discovered by this algorithm, thus shortening its time cost to generalize a wrapper for HTML pages. Jiying Wang, Frederick H. Lochovsky |
WISE | 2 |
| 1999 | Guest Editors' Introduction
James Geller, Frederick H. Lochovsky |
Int. J. Cooperative Inf. Syst. | 2 |
| 1998 | ADOME: An Advanced Object Modeling EnvironmentabstractADOME, which stands for ADvanced Object Modeling Environment, is an approach to integrating data and knowledge management based on object oriented technology. Next generation information systems will require more flexible data modeling capabilities than those provided by current object oriented DBMSs. In particular, integration of data and knowledge management capabilities will become increasingly important. In this context, ADOME provides versatile role facilities that serve as "dynamic binders" between data objects and production rules, thereby facilitating flexible data and knowledge management integration. A prototype that implements this mechanism and the associated operators has been constructed on top of a commercial object oriented DBMS and a rule base system. Qing Li 0001, Frederick H. Lochovsky |
IEEE Trans. Knowl. Data Eng. | 2 |
| 1997 | A Data Model and Semantics of Objects with Dynamic RolesabstractAlthough the concept of roles is becoming a popular research issue in object oriented databases and has been proven to be useful for dynamic and evolving applications, it has only been described conceptually in most of the previous work. Moreover, the important issues such as the semantics of roles (e.g., message passing) are seldom discussed. Furthermore, none of the previous work has investigated the idea of role player qualification, which models the fact that not every object is qualified to play a particular role. We present a data model and the semantics of roles. We discuss each of the above issues and illustrate the ideas with examples. From these examples, we can easily see that the problems we discussed are fundamental and indeed exist in many complex applications. Raymond K. Wong 0001, H. Lewis Chau, Frederick H. Lochovsky |
ICDE | 3 |
| 1996 | The Roles and Views of Multimedia Objects
Raymond K. Wong 0001, H. Lewis Chau, Frederick H. Lochovsky |
MMM | 3 |
| 1992 | Knowledge Communication in Intelligent Information SystemsabstractAn intelligent information system may be composed of hundreds or thousands of entities each of which may possess some part of the overall knowledge of an organization. In such an environment there is a need for these entities to communicate in order to share knowledge and to cooperate in accomplishing organizational activities. In this paper we propose an architecture for modelling intelligent information systems and discuss how this architecture supports the communication of knowledge in such information systems. Carson C. Woo, Frederick H. Lochovsky |
Int. J. Cooperative Inf. Syst. | 2 |
| 1990 | User's command line reference behaviour: Locality versus recency
Alison Lee, Frederick H. Lochovsky |
INTERACT | 2 |
| 1990 | HYTREM - A Hybrid Text-Retrieval Machine for Large DatabasesabstractThe design of a text-retrieval machine, called HYTREM (hybrid text-retrieval machine), for the support of large unformatted text databases is described. A signature file is used as an access method to reduce the amount of data that need to be searched directly. Therefore, HYTREM consists of two major subsystems: a signature processor and a text processor. The signature processor is based on a world-parallel, bit-serial organization which is faster, more efficient, and more flexible than a word-serial, bit-parallel organization proposed by S.R. Ahuja and C.S. Roberts (1980). The text processor, called ALTEP (associative linear text processor), is a linear array of logic cells capable of matching regular expressions at a much higher speed than that of previous designs. Since both the signature processor and ALTEP are highly parallel processors, a high-speed multiple-response resolver is provided to facilitate data transfer between the processors and the controllers over a single common bus. Issues about th design of a cost-effective mass-storage system are also discussed. Performance and implementation issues for HYTREM are discussed.> Dik Lun Lee, Frederick H. Lochovsky |
IEEE Trans. Computers | 2 |
| 1989 | Interactive Specification and Integration of User Views Using Forms
Jürgen Diet, Frederick H. Lochovsky |
ER | 2 |
| 1987 | Role-Based Security in Data Base Management Systems
Frederick H. Lochovsky, Carson C. Woo |
DBSec | 1 |
| 1986 | Supporting Distributed Office Problem Solving in OrganizationsabstractTo improve the effectiveness of office workers in their decision making, office systems have been built to support (rather than replace) their judgment. However, these systems model office work in a centralized environment, and/or they can only support a single office worker. Office work that is divided into specialized domains handled by different office workers (where cooperation is needed in order to accomplish the work) is not supported. In this paper, we will present a model that supports office problem solving in a logically distributed environment. (In some systems, information is geographically distributed for performance purposes rather than for conceptual need. The term, logically , is therefore used to indicate the logical need of organizing information without having to worry about the physical location of the information.) In particular, cooperative tools that can be used to support office workers during the process of their problem solving is discussed. Carson C. Woo, Frederick H. Lochovsky |
ACM Trans. Inf. Syst. | 2 |
| 1984 | Logical Routing Specification in Office Information SystemsabstractA message management system is an office information system for managing structured messages, integrating the facilities of computer-based message systems and database management systems, and adding to them the capability of "intelligent" handling of messages.This allows the office information system to support messages that can use information about themselves (such as structure and content) or about the system to effect their own processing.Logical routing of messages in an office information system is a function that can benefit from such intelligent processing.A framework and language are introduced for the specification of logical routing for messages in an office information system.By associating routing specifications with message types, the system assumes the responsibility both for evaluating the current message instance state to yield the next destination for the instance and for forwarding the instance.The user is freed from the need to direct explicitly each instance of a message type.The routing specifications are based on a variety of criteria, including message instance state and system characteristics.A routing specification language is described, with examples, and an implementation for a distributed workstation environment is outlined. Murray S. Mazer, Frederick H. Lochovsky |
ACM Trans. Inf. Syst. | 2 |
| 1983 | Enhancing the usability of an Office Information System through direct manipulationabstractIn Office Information Systems, the primary focus has been to integrate facilities for the communication and management of information. However, the human factors aspects of the design of office systems are equally important considerations if such office systems are to gain widespread acceptance and use. The application of design techniques from Human Factors can help enhance the usability of an office system. In this paper, we describe the user interface of an office system developed by adapting such design techniques. Alison Lee, Frederick H. Lochovsky |
CHI | 2 |
| 1983 | On evaluating interactive query languages
Frederick H. Lochovsky, Dennis Tsichritzis |
Inf. Sci. | 1 |
| 1982 | An Interactive Query Language for External Data Bases
Frederick H. Lochovsky, Dennis Tsichritzis |
VLDB | 1 |
| 1979 | A Graphical Database Design Aid using the Entity-Relationship Model
Edward P. F. Chan, Frederick H. Lochovsky |
ER | 2 |
| 1977 | User Performance Considerations in DBMS SelectionabstractUser performance as well as system related considerations should be part of the DBMS selection process. However, appropriate procedures and measures for determining user performance characteristics of a DBMS are lacking. This paper describes a methodology and proposes a set of measures for determining the user performance level of the data model and data language of a DBMS. The methodology was applied to three DBMS's having different data models and data languages and the results of its application are discussed. Frederick H. Lochovsky, Dennis Tsichritzis |
SIGMOD Conference | 1 |