EDBT 2026 Demo / reviewers in the wild / expert
Gultekin Özsoyoglu
dblp:o/GultekinOzsoyoglu
· DBLP profile ↗
86ranked-venue papers
28as first author
0since 2021 · last 2016
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 57 · 21 first-authorArtificial intelligence and machine learning · 14 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 12Software engineering, systems software and programming languages · 6 · 4 first-authorSecurity and privacy · 4 · 2 first-authorSystems, architecture and hardware · 3Graphics, computer vision, multimedia, augmented reality and games · 3Human-computer interaction and ubiquitous computing · 2Theory of computation · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Databases, data mining, and information retrieval
40 papers |
Data models and query languages · 33% Query processing and optimization · 28% Information retrieval · 23% | |
| Interdisciplinary, comprehensive, and emerging computing
5 papers |
Bioinformatics and computational biology · 100% |
Topics — the 30 heaviest of 82, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Bioinformatics and computational biology › systems biology
metabolic network analysis |
0.3 | 2 | 2014 | MIRA: mutual information-based reporter algorithm for metabolic networks · Bioinform. 2014 Mining biological networks for unknown pathways · Bioinform. 2007 |
Bioinformatics and computational biology › biological database
pathway database |
0.1 | 2 | 2008 | PathCase: pathways database system · Bioinform. 2008 Pathways Database System: An Integrated System for Biological Pathways · Bioinform. 2003 |
Information retrieval › document retrieval
digital library search |
0.1 | 1 | 2009 | Context-based literature digital collection search · VLDB J. 2009 |
Information retrieval › document retrieval
literature search |
0.1 | 1 | 2009 | Context-based literature digital collection search · VLDB J. 2009 |
Bioinformatics and computational biology › protein function prediction
network-based function prediction |
0.1 | 1 | 2008 | Protein Function Prediction Based on Patterns in Biological Networks · RECOMB 2008 |
Bioinformatics and computational biology › systems bioinformatics › pathway analysis
pathway visualization |
0.1 | 1 | 2008 | PathCase: pathways database system · Bioinform. 2008 |
Bioinformatics and computational biology
protein function prediction |
0.1 | 1 | 2008 | Protein Function Prediction Based on Patterns in Biological Networks · RECOMB 2008 |
Data models and query languages › query language
visual query language |
0.1 | 8 | 2002 | A Graphical Query Language: VISUAL and Its Query Processing · IEEE Trans. Knowl. Data Eng. 2002 VISUAL: A Graphical Icon-Based Query Language · ICDE 1996 Towards a Unified Visual Database Access · SIGMOD Conference 1993 |
Bioinformatics and computational biology › systems bioinformatics › pathway analysis
pathway discovery |
0.1 | 1 | 2007 | Mining biological networks for unknown pathways · Bioinform. 2007 |
Bioinformatics and computational biology
gene expression analysis |
0.1 | 1 | 2014 | MIRA: mutual information-based reporter algorithm for metabolic networks · Bioinform. 2014 |
Data models and query languages
graph query language |
0.0 | 2 | 1999 | Querying Multimedia Presentations Based on Content · IEEE Trans. Knowl. Data Eng. 1999 A Graph Query Language and Its Query Processing · ICDE 1999 |
Query processing and optimization › top-k query processing
rank-aware query processing |
0.0 | 1 | 2004 | Querying web metadata: Native score management and text support in databases · ACM Trans. Database Syst. 2004 |
Bioinformatics and computational biology › data integration
biological data integration |
0.0 | 1 | 2003 | Pathways Database System: An Integrated System for Biological Pathways · Bioinform. 2003 |
Data models and query languages
object-relational database |
0.0 | 1 | 2002 | Sideway Value Algebra for Object-Relational Databases · VLDB 2002 |
Graph data management
graph query processing |
0.0 | 1 | 1999 | A Graph Query Language and Its Query Processing · ICDE 1999 |
Query processing and optimization › online query processing
time-constrained query processing |
0.0 | 2 | 1995 | Time-Constrained Query Processing in CASE-DB · IEEE Trans. Knowl. Data Eng. 1995 Processing Real-Time, Non-Aggregate Queries with Time-Constraints in CASE-DB · ICDE 1992 |
Multimedia systems and quality of experience
multimedia presentation |
0.0 | 2 | 2004 | On Automated Lesson Construction from Electronic Textbooks · IEEE Trans. Knowl. Data Eng. 2004 Querying Multimedia Presentations Based on Content · IEEE Trans. Knowl. Data Eng. 1999 |
Multimedia systems and quality of experience › multimedia delivery
multimedia server |
0.0 | 1 | 1998 | Delivering Presentations from Multimedia Servers · VLDB J. 1998 |
Database system architecture and tuning
real-time database |
0.0 | 2 | 1997 | A database server architecture for agile manufacturing · ICRA 1997 Real-Time Databases: Are they Real? · SIGMOD Conference 1989 |
Privacy and data protection
statistical database privacy |
0.0 | 6 | 1989 | On the Cell Suppression by Merging Technique in the Lattice Model of Summary Tables · S&P 1989 Information Loss in the Lattice Model of Summary Tables due to Cell Suppression · ICDE 1986 Rounding and Inference Controlin Conceptual Models for Statistical Databases · S&P 1985 |
Query processing and optimization
cardinality estimation |
0.0 | 2 | 1993 | Processing Time-Constrained Aggregate Queries in CASE-DB · ACM Trans. Database Syst. 1993 Statistical Estimators for Aggregate Relational Algebra Queries · ACM Trans. Database Syst. 1991 |
Robotics › Motion planning and robot control › robot control
robotic workcell control |
0.0 | 1 | 1997 | Advances in agile manufacturing · ICRA 1997 |
Embedded and real-time systems
real-time control |
0.0 | 1 | 1997 | Advances in agile manufacturing · ICRA 1997 |
Distributed and cloud data management
distributed query processing |
0.0 | 1 | 1996 | VISUAL: A Graphical Icon-Based Query Language · ICDE 1996 |
Query processing and optimization
parallel query processing |
0.0 | 1 | 1996 | VISUAL: A Graphical Icon-Based Query Language · ICDE 1996 |
Privacy and data protection › statistical database privacy
inference control |
0.0 | 6 | 1987 | Data Dependencies and Inference Control in Multilevel Relational Database Systems · S&P 1987 Rounding and Inference Controlin Conceptual Models for Statistical Databases · S&P 1985 Enhancing the Security of Statistical Databases with a Question-Answering System and a Kernel Design · IEEE Trans. Software Eng. 1982 |
Query processing and optimization
approximate query processing |
0.0 | 3 | 1995 | Statistical Estimators for Relational Algebra Expressions · PODS 1988 Synthetic Query Response Construction in Scientific Databases with Time Constraints and Incomplete Information · ICDE 1987 Time-Constrained Query Processing in CASE-DB · IEEE Trans. Knowl. Data Eng. 1995 |
Query processing and optimization
iterative query processing |
0.0 | 1 | 1995 | Time-Constrained Query Processing in CASE-DB · IEEE Trans. Knowl. Data Eng. 1995 |
Transaction processing and concurrency control
real-time transaction processing |
0.0 | 1 | 1995 | Temporal and Real-Time Databases: A Survey · IEEE Trans. Knowl. Data Eng. 1995 |
Data models and query languages
temporal data model |
0.0 | 1 | 1995 | Temporal and Real-Time Databases: A Survey · IEEE Trans. Knowl. Data Eng. 1995 |
Methods — techniques the papers use, named apart from their topics
mutual information · 0.2multivariate scoring · 0.2heuristic evaluation · 0.1axiomatization · 0.1web-based querying · 0.1pattern mining · 0.1client-server architecture · 0.1biological network analysis · 0.1XML-based data dissemination · 0.1graph pattern matching · 0.1gene ontology · 0.1frequent pattern mining · 0.1graph calculus · 0.0sideway-value algebra · 0.0SQL extension · 0.0simulation · 0.0object-oriented software design · 0.0object algebra · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2016 | Emotion -and area-driven topic shift analysis in social media discussionsabstractInternet-based social media platforms allow individuals to discuss/comment on the “topic” of an article in an interactive manner. The topic of a comment/reply in these discussions occasionally shifts, sometimes drastically and abruptly, other times slightly, away from the topic of the article. In this paper we study the phenomena of topic shifts in article-originated social media comments, and identify quantitatively the effects on topic shifts of comments (i) emotion levels (of various emotion dimensions), (ii) topic areas, and (iii) the structure of the discussion tree. We show that, with a better understanding of the topic shift phenomena in comments, automated systems can easily be built to personalize and cater to the comment-browsing and comment-viewing needs of different users. Kamil Topal, Mehmet Koyutürk, Gultekin Özsoyoglu |
ASONAM | 3 |
| 2016 | Movie review analysis: Emotion analysis of IMDb movie reviewsabstractMovie ratings and reviews at sites such as IMDb or Amazon are commonly used by moviegoers to decide which movie to watch or buy next. Currently, moviegoers base their decisions as to which movie to watch by looking at the ratings of movies as well as reading some of the reviews at IMDb or Amazon. This paper argues that there is a better way: reviewers movie scores and reviews can be analyzed with respect to their emotion content, aggregated and projected onto a movie, resulting in an emotion map for a movie. One can then make a decision on which movie to watch next by selecting those movies having emotion maps with certain emotion map patterns desirable for him/her. This paper is a first step towards the above-listed scenario. Kamil Topal, Gultekin Özsoyoglu |
ASONAM | 2 |
| 2015 | A method for imputation of semantic class in diagnostic radiology textabstractDiagnostic medicine produces large volumes of free-text reports used primarily for communication between medical professionals. Secondary use of these reports requires extraction of structured information from the free text. State-of-the-art computational natural language processing techniques can make partial identification of semantics in text, but the diverse terminology used in medical settings makes training classifiers for every lexicon a laborious task. We present statistics of semantics from a large-scale machine-annotated corpus of 83,452 chest x-ray reports. We show that the distribution of semantics is consistent with Zipfian distributions observed in other natural language corpora, and we quantify the semantic focus imparted by limiting a study by body area and modality. We demonstrate that within our semantically focused corpus, pairwise co-occurrence statistics can be used to accurately impute the semantic class for frequently occurring unknown entities, thereby reducing the number of semantically unclassified phrases by up to 25%. Finally, we show that our imputation approach is consistent across multiple reconstructions of the underlying text data. Eamon Johnson, W. Christopher Baughman, Gultekin Özsoyoglu |
BIBM | 3 |
| 2015 | A distributional approach to summarization of radiology reportsabstractDiagnostic radiology reports contain a summary written by a radiologist for communication with primary care providers. Information not included in the summary may be lost in the communication, resulting in substandard patient care. We introduce a notion of salience based on distributional characteristics of term usage in a corpus of radiology reports and present an algorithm for generation of suggestions for inclusion of additional information in report summaries. We evaluate our method on a corpus of 98,913 reports and show our method suggests additions to 11-28% of reports, depending on body location and imaging modality. Eamon Johnson, W. Christopher Baughman, Gultekin Özsoyoglu |
BIBM | 3 |
| 2015 | MIRA: mutual information-based reporter algorithm for metabolic networksabstractdoi: 10.1093/bioinformatics/btu290 Bioinformatics (2014) 30(12), i175–i184 The authors of the above article would like the following to be noted. Results reported for the reporter algorithm (RA) in the Results section are raw z-scores per metabolite, normalized by subtracting the sample mean and dividing by the sample standard deviation. Using the background correction as described in the original scoring scheme (Ideker et al., 2002; Patil and Nielsen, 2005) removes the hub metabolite bias we claim for the RA (for both normalization by k and k). This affects only our comparisons with RA, and MIRA’s results are not changed. The background correction improves the empirical significance of the RA as shown in Figure 5A and C. After the background correction, RA is among the top three permutations. MIRA still performs better but with a smaller margin than shown in Figure 5. We have updated the lists of reporter metabolites (for RA) in the supplementary file. A. Ercüment Çiçek, Kathryn Roeder, Gultekin Özsoyoglu |
Bioinform. | 3 |
| 2014 | MIRA: mutual information-based reporter algorithm for metabolic networksabstractMOTIVATION: Discovering the transcriptional regulatory architecture of the metabolism has been an important topic to understand the implications of transcriptional fluctuations on metabolism. The reporter algorithm (RA) was proposed to determine the hot spots in metabolic networks, around which transcriptional regulation is focused owing to a disease or a genetic perturbation. Using a z-score-based scoring scheme, RA calculates the average statistical change in the expression levels of genes that are neighbors to a target metabolite in the metabolic network. The RA approach has been used in numerous studies to analyze cellular responses to the downstream genetic changes. In this article, we propose a mutual information-based multivariate reporter algorithm (MIRA) with the goal of eliminating the following problems in detecting reporter metabolites: (i) conventional statistical methods suffer from small sample sizes, (ii) as z-score ranges from minus to plus infinity, calculating average scores can lead to canceling out opposite effects and (iii) analyzing genes one by one, then aggregating results can lead to information loss. MIRA is a multivariate and combinatorial algorithm that calculates the aggregate transcriptional response around a metabolite using mutual information. We show that MIRA's results are biologically sound, empirically significant and more reliable than RA. RESULTS: We apply MIRA to gene expression analysis of six knockout strains of Escherichia coli and show that MIRA captures the underlying metabolic dynamics of the switch from aerobic to anaerobic respiration. We also apply MIRA to an Autism Spectrum Disorder gene expression dataset. Results indicate that MIRA reports metabolites that highly overlap with recently found metabolic biomarkers in the autism literature. Overall, MIRA is a promising algorithm for detecting metabolic drug targets and understanding the relation between gene expression and metabolic activity. AVAILABILITY AND IMPLEMENTATION: The code is implemented in C# language using .NET framework. Project is available upon request. A. Ercüment Çiçek, Kathryn Roeder, Gultekin Özsoyoglu |
Bioinform. | 3 |
| 2013 | Locating basic bio-entities in genome-scale reconstructed metabolic networksabstractThe numbers and use of Genome-Scale Reconstructed Metabolic Networks (GSRMN) have been increasing in recent years. Comparing and identifying matching metabolites, reactions, and compartments in GSRMNs can be difficult due to inconsistent naming in GSRMNs. In this paper, we propose metabolite & reaction identification techniques for GSRMNs (by matching metabolites & reactions to corresponding metabolites & reactions in different models. We employ a variety of techniques that include approximate string matching, similarity score functions and filtering techniques, all enhanced by a set of rules based on the underlying metabolic biochemistry. The proposed techniques are evaluated by an empirical study on four pairs of GSRMNs, and significant accuracy gains are achieved using the proposed metabolite & reaction identification techniques Xinjian Qi, Gultekin Özsoyoglu |
BIBM | 2 |
| 2013 | ADEMA: An Algorithm to Determine Expected Metabolite Level Alterations Using Mutual InformationabstractMetabolomics is a relatively new "omics" platform, which analyzes a discrete set of metabolites detected in bio-fluids or tissue samples of organisms. It has been used in a diverse array of studies to detect biomarkers and to determine activity rates for pathways based on changes due to disease or drugs. Recent improvements in analytical methodology and large sample throughput allow for creation of large datasets of metabolites that reflect changes in metabolic dynamics due to disease or a perturbation in the metabolic network. However, current methods of comprehensive analyses of large metabolic datasets (metabolomics) are limited, unlike other "omics" approaches where complex techniques for analyzing coexpression/coregulation of multiple variables are applied. This paper discusses the shortcomings of current metabolomics data analysis techniques, and proposes a new multivariate technique (ADEMA) based on mutual information to identify expected metabolite level changes with respect to a specific condition. We show that ADEMA better predicts De Novo Lipogenesis pathway metabolite level changes in samples with Cystic Fibrosis (CF) than prediction based on the significance of individual metabolite level changes. We also applied ADEMA's classification scheme on three different cohorts of CF and wildtype mice. ADEMA was able to predict whether an unknown mouse has a CF or a wildtype genotype with 1.0, 0.84, and 0.9 accuracy for each respective dataset. ADEMA results had up to 31% higher accuracy as compared to other classification algorithms. In conclusion, ADEMA advances the state-of-the-art in metabolomics analysis, by providing accurate and interpretable classification results. A. Ercüment Çiçek, Ilya Bederman, Leigh Henderson, Mitchell L. Drumm, Gultekin Özsoyoglu |
PLoS Comput. Biol. | 5 |
| 2009 | Progressive Evaluation of XML Queries for Online Aggregation and Progress Indicator
Zhewei Jiang, Wen-Chi Hou, Gultekin Özsoyoglu |
DEXA | 4 |
| 2009 | Corrigendum for Elliott, B. et al., 'PathCase pathways database system', Bioinformatics 2008, 24(21) 2526-2533abstractContact: [email protected] Brendan Elliott, Mustafa Kirac, Ali Cakmak 0001, Gökhan Yavas, Stephen Mayes, En Cheng, Gultekin Özsoyoglu, Z. Meral Özsoyoglu |
Bioinform. | 9 |
| 2009 | Context-based literature digital collection search
Nattakarn Ratprasartporn, Jonathan Po, Ali Cakmak 0001, Sulieman Bani-Ahmad, Gultekin Özsoyoglu |
VLDB J. | 5 |
| 2008 | Taxonomy-superimposed graph miningabstractNew graph structures where node labels are members of hierarchically organized ontologies or taxonomies have become commonplace in different domains, e.g., life sciences. It is a challenging task to mine for frequent patterns in this new graph model which we call taxonomy-superimposed graphs, as there may be many patterns that are implied by the generalization/specialization hierarchy of the associated node label taxonomy. Hence, standard graph mining techniques are not directly applicable. In this paper, we present Taxogram, a taxonomy-superimposed graph mining algorithm that can efficiently discover frequent graph structures in a database of taxonomy-superimposed graphs. Taxogram has two advantages: (i) It performs a subgraph isomorphism test once per class of patterns which are structurally isomorphic, but have different labels, and (ii) it reconciles standard graph mining methods with taxonomy-based graph mining and takes advantage of well-studied methods in the literature. Taxogram has three stages: (a) relabeling nodes in the input database, (b) mining pattern classes/families and constructing associated occurrence indices, and (c) computing patterns and eliminating useless (i.e., over-generalized) patterns by post-processing occurrence indices. Experimental results show that Taxogram is significantly more efficient and more scalable compared to other alternative approaches. 1. Ali Cakmak 0001, Gultekin Özsoyoglu |
EDBT | 2 |
| 2008 | Protein Function Prediction Based on Patterns in Biological Networks
Mustafa Kirac, Gultekin Özsoyoglu |
RECOMB | 2 |
| 2008 | PathCase: pathways database systemabstractMOTIVATION: As the blueprints of cellular actions, biological pathways characterize the roles of genomic entities in various cellular mechanisms, and as such, their availability, manipulation and queriability over the web is important to facilitate ongoing biological research. RESULTS: In this article, we present the new features of PathCase, a system to store, query, visualize and analyze metabolic pathways at different levels of genetic, molecular, biochemical and organismal detail. The new features include: (i) a web-based system with a new architecture, containing a server-side and a client-side, and promoting scalability, and flexible and easy adaptation of different pathway databases, (ii) an interactive client-side visualization tool for metabolic pathways, with powerful visualization capabilities, and with integrated gene and organism viewers, (iii) two distinct querying capabilities: an advanced querying interface for computer savvy users, and built-in queries for ease of use, that can be issued directly from pathway visualizations and (iv) a pathway functionality analysis tool. PathCase is now available for three different datasets, namely, KEGG pathways data, sample pathways from the literature and BioCyc pathways for humans. AVAILABILITY: Available online at http://nashua.case.edu/pathways Brendan Elliott, Mustafa Kirac, Ali Cakmak 0001, Gökhan Yavas, Stephen Mayes, En Cheng, Gultekin Özsoyoglu, Z. Meral Özsoyoglu |
Bioinform. | 9 |
| 2008 | Discovering gene annotations in biomedical text databasesabstractBACKGROUND: Genes and gene products are frequently annotated with Gene Ontology concepts based on the evidence provided in genomics articles. Manually locating and curating information about a genomic entity from the biomedical literature requires vast amounts of human effort. Hence, there is clearly a need forautomated computational tools to annotate the genes and gene products with Gene Ontology concepts by computationally capturing the related knowledge embedded in textual data. RESULTS: In this article, we present an automated genomic entity annotation system, GEANN, which extracts information about the characteristics of genes and gene products in article abstracts from PubMed, and translates the discoveredknowledge into Gene Ontology (GO) concepts, a widely-used standardized vocabulary of genomic traits. GEANN utilizes textual "extraction patterns", and a semantic matching framework to locate phrases matching to a pattern and produce Gene Ontology annotations for genes and gene products. In our experiments, GEANN has reached to the precision level of 78% at therecall level of 61%. On a select set of Gene Ontology concepts, GEANN either outperforms or is comparable to two other automated annotation studies. Use of WordNet for semantic pattern matching improves the precision and recall by 24% and 15%, respectively, and the improvement due to semantic pattern matching becomes more apparent as the Gene Ontology terms become more general. CONCLUSION: GEANN is useful for two distinct purposes: (i) automating the annotation of genomic entities with Gene Ontology concepts, and (ii) providing existing annotations with additional "evidence articles" from the literature. The use of textual extraction patterns that are constructed based on the existing annotations achieve high precision. The semantic pattern matching framework provides a more flexible pattern matching scheme with respect to "exactmatching" with the advantage of locating approximate pattern occurrences with similar semantics. Relatively low recall performance of our pattern-based approach may be enhanced either by employing a probabilistic annotation framework based on the annotation neighbourhoods in textual data, or, alternatively, the statistical enrichment threshold may be adjusted to lower values for applications that put more value on achieving higher recall values. Ali Cakmak 0001, Gultekin Özsoyoglu |
BMC Bioinform. | 2 |
| 2007 | Gene Ontology-Based Annotation Analysis and Categorization of Metabolic PathwaysabstractFunctional characterizations of pathways provide new opportunities in defining, understanding, and comparing existing biological pathways, and in helping discover new ones in different organisms. In this paper, we present and evaluate computational techniques for categorizing pathways, based upon the Gene Ontology (GO) annotations of enzymes within metabolic pathways. Our approach is to use the notion of functionality templates, GO-functional graphs of pathways. Pathway categorization is then achieved through learning models built on different characteristics of functionality templates. We have experimentally evaluated the accuracy of automated pathway categorization with respect to different learning models and their parameters. Using KEGG metabolic pathways, the pathway categorization tool reaches to 90% and higher accuracy. Ali Cakmak 0001, Mustafa Kirac, Marc R. Reynolds, Z. Meral Özsoyoglu, Gultekin Özsoyoglu |
SSDBM | 5 |
| 2007 | Mining biological networks for unknown pathwaysabstractMOTIVATION: Biological pathways provide significant insights on the interaction mechanisms of molecules. Presently, many essential pathways still remain unknown or incomplete for newly sequenced organisms. Moreover, experimental validation of enormous numbers of possible pathway candidates in a wet-lab environment is time- and effort-extensive. Thus, there is a need for comparative genomics tools that help scientists predict pathways in an organism's biological network. RESULTS: In this article, we propose a technique to discover unknown pathways in organisms. Our approach makes in-depth use of Gene Ontology (GO)-based functionalities of enzymes involved in metabolic pathways as follows: i. Model each pathway as a biological functionality graph of enzyme GO functions, which we call pathway functionality template. ii. Locate frequent pathway functionality patterns so as to infer previously unknown pathways through pattern matching in metabolic networks of organisms. We have experimentally evaluated the accuracy of the presented technique for 30 bacterial organisms to predict around 1500 organism-specific versions of 50 reference pathways. Using cross-validation strategy on known pathways, we have been able to infer pathways with 86% precision and 72% recall for enzymes (i.e. nodes). The accuracy of the predicted enzyme relationships has been measured at 85% precision with 64% recall. AVAILABILITY: Code upon request. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Ali Cakmak 0001, Gultekin Özsoyoglu |
Bioinform. | 2 |
| 2006 | Task-Oriented Integrated Use of Biological Web Data SourcesabstractBiological Web data sources have now become essential information sources for researchers. However, their use is tedious, labor-intensive, repetitive, and possibly involve the integration of data from multiple Web data sources. In this paper, as a first step towards the full integration of Web data sources, we propose a framework that allows an integrated use of biological sources in a task-oriented manner. We define and experimentally evaluate a toolkit-based framework for semi-automatically constructing an integrated (software) system that automates and optimizes the execution of a biology-related computational task at hand. To test and refine the principles of the framework, we build and evaluate "pathway-infer" as a benchmark integrated system Mustafa Kirac, Ali Cakmak 0001, Gultekin Özsoyoglu |
SSDBM | 3 |
| 2006 | On Data and Visualization Models for Signaling PathwaysabstractSignaling pathways are chains of interacting proteins, through which the cell converts a (usually) extracellular signal into a biological response. The number of known signaling pathways in the biological literature and on the Web has been increasing at a very high rate, thus demanding a need for efficient ways of storing, visualizing, querying, and mining signaling pathways. In this paper, first we briefly compare the data modeling and visualization capabilities of existing signaling pathways systems. Then, we present a signaling pathway data model and its visualization that subsumes the existing models. Our model visualizes a signaling pathway (a) as a nested graph, (b) with explicit location information (e.g., cell, tissue, organelle, nucleus, etc.), and (c) in four abstraction levels, namely, the levels of molecule-to-molecule signaling steps, collapsed sub-pathways, molecule-to-pathway connections, and pathway-to-pathway connections. We model (1) the effects of specific signaling steps, (2) state changes of signaling molecules, (3) various (extensible) structural/physical changes of signaling molecules such as complex formation, dissociation, assembly, oligomerization, di-/trimerization, cleavage and degradation, (4) condensation/hydrolysis signaling steps, and (5) exchanges and translocations as signaling steps. The visualization model gracefully models incomplete information and hierarchical levels of signaling molecules. Finally, we introduce a completely new visualization dimension for pathways, namely, gene ontology (GO)-based functional visualizations of pathways. We believe that functional visualizations of pathways provides new opportunities in understanding, defining and comparing existing pathways, and in helping discover new ones Nattakarn Ratprasartporn, Ali Cakmak 0001, Gultekin Özsoyoglu |
SSDBM | 3 |
| 2004 | Metadata-based modeling of information resources on the WebabstractAbstract This paper deals with the problem of modeling Web information resources using expert knowledge and personalized user information for improved Web searching capabilities. We propose a “Web information space” model, which is composed of Web‐based information resources (HTML/XML [Hypertext Markup Language/Extensible Markup Language] documents on the Web), expert advice repositories (domain‐expert‐specified metadata for information resources), and personalized information about users (captured as user profiles that indicate users' preferences about experts as well as users' knowledge about topics). Expert advice, the heart of the Web information space model, is specified using topics and relationships among topics (called metalinks), along the lines of the recently proposed topic maps. Topics and metalinks constitute metadata that describe the contents of the underlying HTML/XML Web resources. The metadata specification process is semiautomated, and it exploits XML DTDs (Document Type Definition) to allow domain‐expert guided mapping of DTD elements to topics and metalinks. The expert advice is stored in an object‐relational database management system (DBMS). To demonstrate the practicality and usability of the proposed Web information space model, we created a prototype expert advice repository of more than one million topics/metalinks for DBLP (Database and Logic Programming) Bibliography data set. We also present a query interface that provides sophisticated querying facilities for DBLP Bibliography resources using the expert advice repository. Selma Ayse Özel, Ismail Sengör Altingövde, Özgür Ulusoy, Gultekin Özsoyoglu, Z. Meral Özsoyoglu |
J. Assoc. Inf. Sci. Technol. | 4 |
| 2004 | On Automated Lesson Construction from Electronic TextbooksabstractAn electronic book may be viewed as an application with a multimedia database. We define an electronic textbook as an electronic book that is used in conjunction with instructional resources such as lectures. We propose an electronic textbook data model with topics, topic sources, metalinks (relationships among topics), and instructional modules, which are multimedia presentations possibly capturing real-life lectures of instructors. Using the data model, the system provides users a topic-guided multimedia lesson construction. We concentrate, in detail, on the use of one metalink type in lesson construction, namely, prerequisite dependencies, and provide a sound and complete axiomatization of prerequisite dependencies. We present a simple automated way of constructing lessons for users where the user lists a set of topic names (s)he is interested in, and the system automatically constructs and delivers the "best" user-tailored lesson as a multimedia presentation, where "best" is characterized in terms of both topic closures with respect to prerequisite dependencies and what the user knows about topics. We model and present sample lesson construction requests for users, discuss their complexity, and give algorithms that evaluate such requests. For expensive lesson construction requests, we list heuristics and empirically evaluate their performance. We also discuss the worst-case performance guarantees of lesson request algorithms. Gultekin Özsoyoglu, Nevzat Hurkan Balkir, Z. Meral Özsoyoglu, Graham Cormode |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2004 | Querying web metadata: Native score management and text support in databasesabstractIn this article, we discuss the issues involved in adding a native score management system to object-relational databases, to be used in querying Web metadata (that describes the semantic content of Web resources). The Web metadata model is based on topics (representing entities), relationships among topics (called metalinks ), and importance scores (sideway values) of topics and metalinks. We extend database relations with scoring functions and importance scores. We add to SQL score-management clauses with well-defined semantics, and propose the sideway-value algebra (SVA), to evaluate the extended SQL queries. SQL extensions and the SVA algebra are illustrated through two Web resources, namely, the DBLP Bibliography and the SIGMOD Anthology.SQL extensions include clauses for propagating input tuple importance scores to output tuples during query processing, clauses that specify query stopping conditions, threshold predicates (a type of approximate similarity predicates for text comparisons), and user-defined-function-based predicates. The propagated importance scores are then used to rank and return a small number of output tuples. The query stopping conditions are propagated to SVA operators during query processing. We show that our SQL extensions are well-defined, meaning that, given a database and a query Q, under any query processing scheme, the output tuples of Q and their importance scores stay the same.To process the SQL extensions, we discuss two sideway value algebra operators, namely, sideway value algebra join and topic closure, give their implementation algorithms, and report their experimental evaluations. Gultekin Özsoyoglu, Ismail Sengör Altingövde, Abdullah Al-Hamdani, Selma Ayse Özel, Özgür Ulusoy, Z. Meral Özsoyoglu |
ACM Trans. Database Syst. | 1 |
| 2003 | Anti-Tamper Databases: Querying Encrypted Databases
Gultekin Özsoyoglu, David A. Singer, Sun S. Chung |
DBSec | 1 |
| 2003 | Selecting Topics for Web Resource Discovery: Efficiency Issues in a Database Approach
Abdullah Al-Hamdani, Gultekin Özsoyoglu |
DEXA | 2 |
| 2003 | XML Restructuring and Integration for Tabular Data
Z. Meral Özsoyoglu, Gultekin Özsoyoglu |
DEXA | 3 |
| 2003 | Pathways Database System: An Integrated System for Biological PathwaysabstractMOTIVATION: During the next phase of the Human Genome Project, research will focus on functional studies of attributing functions to genes, their regulatory elements, and other DNA sequences. To facilitate the use of genomic information in such studies, a new modeling perspective is needed to examine and study genome sequences in the context of many kinds of biological information. Pathways are the logical format for modeling and presenting such information in a manner that is familiar to biological researchers. RESULTS: In this paper we present an integrated system, called Pathways Database System, with a set of software tools for modeling, storing, analyzing, visualizing, and querying biological pathways data at different levels of genetic, molecular, biochemical and organismal detail. The novel features of the system include: (a) genomic information integrated with other biological data and presented from a pathway, rather than from the DNA sequence, perspective; (b) design for biologists who are possibly unfamiliar with genomics, but whose research is essential for annotating gene and genome sequences with biological functions; (c) database design, implementation and graphical tools which enable users to visualize pathways data in multiple abstraction levels, and to pose predetermined queries; and (d) an implementation that allows for web(XML)-based dissemination of query outputs (i.e. pathways data) to researchers in the community, giving them control on the use of pathways data. AVAILABILITY: Available on request from the authors. Larkshmi Krishnamurthy, Joseph H. Nadeau, Gultekin Özsoyoglu, Z. Meral Özsoyoglu, Greg Schaeffer, Murat Tasan, Wanhong Xu |
Bioinform. | 3 |
| 2002 | Sideway Value Algebra for Object-Relational Databases
Gultekin Özsoyoglu, Abdullah Al-Hamdani, Ismail Sengör Altingövde, Selma Ayse Özel, Özgür Ulusoy, Z. Meral Özsoyoglu |
VLDB | 1 |
| 2002 | A Graphical Query Language: VISUAL and Its Query ProcessingabstractThis paper describes VISUAL, a graphical icon-based query language with a user-friendly graphical user interface for scientific databases and its query processing techniques. VISUAL is suitable for domains where visualization of the relationships is important for the domain scientist to express queries. In VISUAL, graphical objects are not tied to the underlying formalism; instead, they represent the relationships of the application domain. VISUAL supports relational, nested, and object-oriented models naturally and has formal basis. For ease of understanding and for efficiency reasons, two VISUAL semantics are introduced, namely, the interpretation and execution semantics. Translations from VISUAL to the Object Query Language (for portability considerations) and to an object algebra (for query processing purposes) are presented. Concepts of external and internal queries are developed as modularization tools. Nevzat Hurkan Balkir, Gultekin Özsoyoglu, Z. Meral Özsoyoglu |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2001 | Topic-Centric Querying of Web Information Resources
Ismail Sengör Altingövde, Selma Ayse Özel, Özgür Ulusoy, Gultekin Özsoyoglu, Z. Meral Özsoyoglu |
DEXA | 4 |
| 2000 | Query Processing Techniques for Multimedia Presentations
Taekyong Lee, Nevzat Hurkan Balkir, Abdullah Al-Hamdani, Gultekin Özsoyoglu, Z. Meral Özsoyoglu |
Multim. Tools Appl. | 5 |
| 1999 | A Graph Query Language and Its Query ProcessingabstractMany new database applications involve querying of graph data. We present an object-oriented graph data model, and an OQL like graph query language, GOQL. The data model and the language are illustrated in the application domain of multimedia presentation graphs. We then discuss the query processing techniques for GOQL, more specifically, the translation of GOQL, into an operator-based language, called O-Algebra, extended with operators to deal with paths and sequences. We also discuss different approaches for efficient implementation of algebra operators for paths and sequences. Z. Meral Özsoyoglu, Gultekin Özsoyoglu |
ICDE | 3 |
| 1999 | Constraint-Based Automation of Multimedia Presentation Assembly
Veli Hakkoymaz, Joel Kraft, Gultekin Özsoyoglu |
Multim. Syst. | 3 |
| 1999 | Querying Multimedia Presentations Based on ContentabstractConsiders the problem of querying multimedia presentations based on content information. Multimedia presentations are modeled as presentation graphs, which are directed acyclic graphs that visually specify the presentations. We present a graph data model for the specification of multimedia presentations and discuss query languages as effective tools to query and manipulate multimedia presentation graphs with respect to content information. To query the information flow throughout a multimedia presentation, as well as in each individual multimedia stream, we use revised versions of temporal operators Next, Connected and Until, together with path formulas. These constructs allow us to specify and query paths along a presentation graph. We present an icon-based graphical query language, GVISUAL, that provides iconic representations for these constructs and a user-friendly graphical interface for query specification. We also present an OQL-like language, GOQL (Graph OQL), with similar constructs, that allows textual and more traditional specifications of graph queries. Finally, we introduce GCalculus (Graph Calculus), a calculus-based language that establishes the formal grounds for the use of temporal operators in path formulas and for querying presentation graphs with respect to content information. We also discuss GCalculus/S (GCalculus with Sets) which avoids highly complex query expressions by eliminating the universal path quantifier, the negation operator and the universal quantifier. GCalculus/S represents the formal basis for GVISUAL, i.e. GVISUAL uses the constructs of GCalculus/S directly. Taekyong Lee, Tolga Bozkaya, Nevzat Hurkan Balkir, Z. Meral Özsoyoglu, Gultekin Özsoyoglu |
IEEE Trans. Knowl. Data Eng. | 6 |
| 1998 | Selecting Actions to Trigger in Active Database Applications
Huang-Cheng Kuo, Gultekin Özsoyoglu |
DEXA | 2 |
| 1998 | Delivering Presentations from Multimedia Servers
Nevzat Hurkan Balkir, Gultekin Özsoyoglu |
VLDB J. | 2 |
| 1997 | Real-Time Transactions with Execution Histories: Priority Assignment and Load ControlabstractArticle Free Access Share on Real-time transactions with execution histories: priority assignment and load control Authors: Erdoğan Doğdu Department of Computer Engineering and Science, Case Western Reserve University, Cleveland, OH Department of Computer Engineering and Science, Case Western Reserve University, Cleveland, OHView Profile , Gültekin Özsoyoğlu Department of Computer Engineering and Science, Case Western Reserve University, Cleveland, OH Department of Computer Engineering and Science, Case Western Reserve University, Cleveland, OHView Profile Authors Info & Claims CIKM '97: Proceedings of the sixth international conference on Information and knowledge managementJanuary 1997 Pages 301–308https://doi.org/10.1145/266714.266915Online:01 January 1997Publication History 1citation258DownloadsMetricsTotal Citations1Total Downloads258Last 12 Months3Last 6 weeks0 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteeReaderPDF Erdogan Dogdu, Gultekin Özsoyoglu |
CIKM | 2 |
| 1997 | A flexible software architecture for agile manufacturingabstractThe flexibility required of an agile manufacturing system must be achieved largely through computer software. The system's control software must be adaptable to new products and to new system components without becoming unreliable or difficult to maintain. This requires designing the software specifically to facilitate future changes. As part of the Agile Manufacturing Project at Case Western Reserve University, we have developed a software architecture for control of an agile manufacturing workcell, and we have demonstrated its flexibility with rapid changeover and introduction of new products. In this paper, we describe the requirements for agile manufacturing software and how our software architecture addresses them. Yoohwan Kim, Ju-Yeon Jo, Virgilio B. Velasco Jr., Nicholas A. Barendt, Andy Podgurski, Gultekin Özsoyoglu, Francis L. Merat |
ICRA | 6 |
| 1997 | A database server architecture for agile manufacturingabstractAgile manufacturing systems can benefit significantly from a database support. This paper describes AMDS, an agile manufacturing database system designed for capturing and manipulating the operational data of a manufacturing cell. AMDS is a continuous data-gathering real-time DBMS, and it can be logged either locally or remotely and used for off-line analysis as well. The temporal operational data obtained is used for performance and reliability analysis, high-level summary report generation, real-time monitoring and active interaction. This paper gives an overview of AMDS and its real-time features. Sungkil Lee 0001, Huang-Cheng Kuo, Nevzat Hurkan Balkir, Gultekin Özsoyoglu |
ICRA | 4 |
| 1997 | Advances in agile manufacturingabstractAn agile workcell has been developed for light mechanical assembly in collaboration with industrial sponsors. The workcell includes multiple Adept robots, a Bosch conveyor system, multiple flexible parts feeders at each robot's workstation, CCD cameras for parts feeding and hardware registration, and a dual VMEbus control system. Our flexible pairs feeder design uses multiple conveyors to singulate the parts and machine vision to locate them. Specialized hardware is encapsulated on modular grippers and modular worktables which can be quickly interchanged for assembly of different products. Object-oriented software (C++) running under VxWorks, a real-time operating system, is used for workcell control. An agile software architecture was developed for rapid introduction of new assemblies through code re-use. A simulation of the workcell was developed so that controller software could be written and tested off-line, enabling the rapid introduction of new products. Francis L. Merat, Nicholas A. Barendt, Roger D. Quinn, Greg C. Causey, Wyatt S. Newman, Virgilio B. Velasco Jr., Andy Podgurski, Yoohwan Kim, Gultekin Özsoyoglu, Ju-Yeon Jo |
ICRA | 9 |
| 1997 | A Constraint-Driven Approach to Automate the Organization and Playout of Presentations in Multimedia Databases
Veli Hakkoymaz, Gultekin Özsoyoglu |
Multim. Tools Appl. | 2 |
| 1996 | Distributed Processing of Time-Constrained Queries in CASE-DBabstractWe develop and experimentally evaluate a real-time distributed relational prototype database system that permits the specification of time constraints for relational algebra queries and evaluates queries in parallel over distributed sites in a completely replicated data environment.We develop, implement and evaluate on multiple server sites iterative distributed query evaluation techniques.The risk of overspending the time constraint at each iteration is controlled using a probabilistic risk cent rol technique.The query execution plan is designed to restructure a query into an equivalent, esay-to-parallelize form, and distributed execution is achieved by multi-threaded processes and RPC calls. Sungkil Lee 0001, Gultekin Özsoyoglu |
CIKM | 2 |
| 1996 | VISUAL: A Graphical Icon-Based Query LanguageabstractVISUAL is a graphical icon-based query language designed for scientific databases where visualization of the relationships are important for the domain scientist to express queries. Graphical objects are not tied to the underlying formalism; instead, they represent the relationships of the application domain. VISUAL supports relational, nested, and object-oriented models naturally and has formal basis. In addition to set and bag constructs for complex objects, sequences are also supported by the data model. Concepts of external and internal queries are developed as modularization tools. A new parallel/distributed query processing paradigm is presented. VISUAL query processing techniques are also discussed. Nevzat Hurkan Balkir, Eser Sükan, Gultekin Özsoyoglu, Z. Meral Özsoyoglu |
ICDE | 3 |
| 1996 | Automating the Assembly of Presentations from Multimedia DatabasesabstractA multimedia presentation refers to the presentation of multimedia data using output devices such as monitors for text and video, and speakers for audio. Each presentation consists of multimedia segments which are obtained from a multimedia data model. In this paper, we propose to express the semantic coherency of a multimedia presentation in terms of presentation inclusion and exclusion constraints that are incorporated into the multimedia data model. Thus, when a user specifies a set of segments for a presentation, the DBMS adds segments into and/or deletes segments from the set in order to satisfy the inclusion and exclusion constraints. To automate the assembly of a presentation with concurrent presentation streams, we also propose presentation organization constraints that are incorporated into the multimedia data model, independent of any presentation. We give two algorithms for automated presentation assembly and discuss their complexity. We discuss the satisfiability of inclusion and exclusion constraints when negation is allowed, and we briefly describe a prototype system that is being developed for automated presentation assembly. Gultekin Özsoyoglu, Veli Hakkoymaz, Joel Kraft |
ICDE | 1 |
| 1996 | A Scientific Multimedia Database System for Polymer Science ExperimentsabstractThis paper describes SciMMDB, a scientific multimedia database system designed for polymer science experiment data applications. SciMMDB maintains multimedia data types of video, test and picture (as well as other standard data types), and allows users to: form multimedia presentation streams about experiments; define a presentation which organizes a number of presentation streams for playout purposes; and control the playout of a presentation. To form video, text or picture streams, a presentation stream construction language is described. Queries in this language express spatial and temporal relationships involving content objects and efficient evaluation of such queries are facilitated by indexing structures. SciMMDB uses IB+trees for queries involving temporal intervals. In SciMMDB, presentation streams are formed into a presentation using the notion of presentation graphs. During the playout of a presentation, each presentation stream in the presentation is played out by a separate transaction, called an output agent. Such transactions communicate and synchronize with each other using the cooperative transaction model. Taekyong Lee, Tolga Bozkaya, Huang-Cheng Kuo, Gultekin Özsoyoglu, Z. Meral Özsoyoglu |
SSDBM | 4 |
| 1995 | Time-Constrained Query Processing in CASE-DBabstractCASE-DB is a real-time, single-user, relational prototype DBMS that permits the specification of strict time constraints for relational algebra queries. Given a time constrained nonaggregate relational algebra query and a "fragment chain" for each relation involved in the query, CASE-DB initially obtains a response to a modified version of the query and then uses an "iterative query evaluation" technique to successively improve and evaluate the modified version of the query, CASE-DB controls the risk of overspending the time quota at each step using a "risk control technique". Gultekin Özsoyoglu, Sujatha Guru Swamy, Kaizheng Du, Wen-Chi Hou |
IEEE Trans. Knowl. Data Eng. | 1 |
| 1995 | Guest Editors' Introduction to Special Section On Temporal and Real-Time Databases
Gultekin Özsoyoglu, Richard T. Snodgrass |
IEEE Trans. Knowl. Data Eng. | 1 |
| 1995 | Temporal and Real-Time Databases: A SurveyabstractA temporal database contains time-varying data. In a real-time database transactions have deadlines or timing constraints. In this paper we review the substantial research in these two previously separate areas. First we characterize the time domain; then we investigate temporal and real-time data models. We evaluate temporal and real-time query languages along several dimensions. We examine temporal and real-time DBMS implementation. Finally, we summarize major research accomplishments to date and list several unanswered research questions.> Gultekin Özsoyoglu, Richard T. Snodgrass |
IEEE Trans. Knowl. Data Eng. | 1 |
| 1994 | A Scientific Database System for Polymers and Materials Engineering NeedsabstractThis paper describes the environment, the issues and the progress made in developing a database management system for managing data generated by fracture study experiments in polymers and materials engineering. The experiment data produced by materials engineers is large and has several data types that include summary data, image data, spatial data, and temporal data obtained by applying differing amounts of multiaxial stresses to materials at different times. The types of queries on the experiment data include statistical queries, matching queries, mapping queries, and spatiotemporal queries. We discuss the data model, query language, and spatial data structure issues associated with our environment, and how they translate into specific research problems. We also briefly summarize the progress made in each area.> Gultekin Özsoyoglu, Z. Meral Özsoyoglu, Kumar V. Vadaparty |
SSDBM | 1 |
| 1994 | A Framework for Feature-Based Indexing for Spatial DatabasesabstractWe propose a data-driven method for efficient retrieval of objects as well as similar shapes with a given query object in the spatial database environment. The idea is to find some features from the image object in order to build an index search structure. For the sake of similarity matching among shapes, the features must be invariant to rotation, translation and scaling. We propose a set of generic features that are invariant to these transformations. Each feature in the feature vector is associated with a weight based on the application, which is used in the search process. Any multidimensional point access method can then be used to build an index. In this paper, a variant of the K-D-B tree is used to construct the index structure. Finally, we define a similarity measure to find objects similar to a given query object, and discuss how similarity queries can be processed using the index structure.> Nasser Yazdani, Z. Meral Özsoyoglu, Gultekin Özsoyoglu |
SSDBM | 3 |
| 1993 | Modifying Datbase Queries and Error Constraints
Kaizheng Du, Gultekin Özsoyoglu |
DEXA | 2 |
| 1993 | Towards a Unified Visual Database AccessabstractSince the development of QBE, over fifty visual query languages have been proposed to facilitate easy database access. Although these languages have introduced some very useful paradigms, a number of these have some severe limitations, such as: (a) not extending beyond the relational model (b) not considering negation and safety, formally (c) using ad hoc constructs, with no analysis of expressivity or complexity done, etc. Note that visual database access is an important issue being revisted, with the emergence of different flavors of object-oriented databases. We believe that there is a need for developing a unified visual query language. Kumar V. Vadaparty, Y. Alp Aslandogan, Gultekin Özsoyoglu |
SIGMOD Conference | 3 |
| 1993 | Information loss in three cell-level control techniques for summary tables
Gultekin Özsoyoglu, JiYoung Chung, Tzong-An Su |
Inf. Sci. | 1 |
| 1993 | Irreversibility problem in NL-class relations
Hsiu-Hsen Yao, Gultekin Özsoyoglu |
Inf. Sci. | 2 |
| 1993 | Incomplete Relational Database Models Based on IntervalsabstractTables are used to represent unknown relations. The three partial tuple types are defined in a table to specify incompleteness relationships among tuples of the same table. For tuples of different tables, the cases where incompleteness is introduced at the relation level, tuple level, or attribute value level are discussed. For each of the models, it is shown that query evaluation is sound in the Imielinski-Lipski sense. None of the models is complete in the Imielinski-Lipski sense. Two of the models in the family are compared with other approaches.> Adegbeniga Ola, Gultekin Özsoyoglu |
IEEE Trans. Knowl. Data Eng. | 2 |
| 1993 | Near-Optimum Storage Models for Nested Relations Based on Workload InformationabstractThe problem of choosing a storage model for a nested relation (i.e., a relation containing relations) is considered. A technique is introduced that uses the workload information of the database system under consideration to obtain a better storage model (i.e., one with a lower query cost) for a given nested relation. The nested relation scheme is first represented as a tree called the scheme tree. By using the workload information and by performing a series of merges in the nodes of the scheme tree, a near-optimum scheme tree is produced, and file organization types are assigned to each node (file) in the scheme tree. The authors' methodology is applied by using a specific nested relational algebra and three file organization types, namely, sequential, heap, and dense index files. The proposed methodology locates the optimum storage model and the optimum file organization techniques for the external and internal relations of the nested relations tested.> Gultekin Özsoyoglu, Aladdin Hafez |
IEEE Trans. Knowl. Data Eng. | 1 |
| 1993 | Processing Time-Constrained Aggregate Queries in CASE-DBabstractIn this paper, we present an algorithm to strictly control the time to process an estimator for an aggregate relational query. The algorithm implemented in a prototype database management system, called CASE-DB, iteratively samples from input relations, and evaluates the associated estimator until the time quota expires. In order to estimate the time cost of a query, CASE-DB uses adaptive time cost formulas. The formulas are adaptive in that the parameters of the formulas can be adjusted at runtime to better fit the characteristics of a query. To control the use of time quota, CASE-DB adopts the one-at-a-time-interval time control strategy to make a tradeoff between the risks of overspending and the overhead, finally, experimental evaluation of the methodology is presented. Wen-Chi Hou, Gultekin Özsoyoglu |
ACM Trans. Database Syst. | 2 |
| 1992 | Processing Real-Time, Non-Aggregate Queries with Time-Constraints in CASE-DBabstractThe problem of time-constrained query evaluation in a single-user database management system (DBMS) is considered. CASE-DB is a real-time, single user, relational prototype DBMS that uses the relational algebra as its query language. Given a nonaggregate query and a fragment chain for each input relation of the query. CASE-DB uses iterative query evaluation techniques to obtain a response first to a modified version of the query, and then to successively improved versions of the query. CASE-DB controls the risk of overspending the time quota at each step using a risk control technique. For periodically occurring queries, CASE-DB uses incremental query evaluation techniques that make sure that each operator in the query has at least one operand relation which contains the changes in the last period, and is expected to be very small compared to the actual database relation.> Gultekin Özsoyoglu, Kaizheng Du, Sujatha Guru Swamy, Wen-Chi Hou |
ICDE | 1 |
| 1992 | Human factors study of two screen-oriented query languages: STBE and QBE
Gultekin Özsoyoglu, W. A. Abdul-Qader |
Inf. Softw. Technol. | 1 |
| 1991 | On Estimating COUNT, SUM, and AVERAGE
Gultekin Özsoyoglu, Kaizheng Du, A. Tjahjana, Wen-Chi Hou, D. Y. Rowland |
DEXA | 1 |
| 1991 | Error-Constraint COUNT Query Evaluation in Relational Databases
Wen-Chi Hou, Gultekin Özsoyoglu, Erdogan Dogdu |
SIGMOD Conference | 2 |
| 1991 | Controlling FD and MVD Inferences in Multilevel Relational Database SystemsabstractThe authors investigate the inference problems due to functional dependencies (FD) and multivalued dependencies (MVD) in a multilevel relational database (MDB) with attribute and record classification schemes, respectively. The set of functional dependencies to be taken into account in order to prevent FD-compromises is determined. It is proven that incurring minimum information loss to prevent compromises is an NP-complete problem. An exact algorithm to adjust the attribute levels so that no compromise due to functional dependencies occurs is given. Some necessary and sufficient conditions for MVD-compromises are presented. The set of MVDs to be taken into account for controlling inferences is determined. An algorithm to prevent MVD-compromises in a relation with conflict-free MVDs is given.> Tzong-An Su, Gultekin Özsoyoglu |
IEEE Trans. Knowl. Data Eng. | 2 |
| 1991 | Statistical Estimators for Aggregate Relational Algebra Queriesabstractarticle Free Access Share on Statistical estimators for aggregate relational algebra queries Authors: Wen-Chi Hou Case Western Reserve Univ., Cleveland, OH Case Western Reserve Univ., Cleveland, OHView Profile , Gultekin Ozsoyoglu Case Western Reserve Univ., Cleveland, OH Case Western Reserve Univ., Cleveland, OHView Profile Authors Info & Claims ACM Transactions on Database SystemsVolume 16Issue 4pp 600–654https://doi.org/10.1145/115302.115300Published:01 December 1991Publication History 44citation559DownloadsMetricsTotal Citations44Total Downloads559Last 12 Months24Last 6 weeks2 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteeReaderPDF Wen-Chi Hou, Gultekin Özsoyoglu |
ACM Trans. Database Syst. | 2 |
| 1990 | Database Systems for Programmable Logic Controllers
Gultekin Özsoyoglu, Wen-Chi Hou, Adegbemiga Ola |
SSDBM | 1 |
| 1990 | On Inference Control in Semantic Data Models for Statistical Databases
Gultekin Özsoyoglu, Tzong-An Su |
J. Comput. Syst. Sci. | 1 |
| 1989 | Processing Aggregate Relational Queries with Hard Time Constraints
Wen-Chi Hou, Gultekin Özsoyoglu, Baldeo K. Taneja |
SIGMOD Conference | 2 |
| 1989 | Real-Time Databases: Are they Real?
Gultekin Özsoyoglu |
SIGMOD Conference | 1 |
| 1989 | On the Cell Suppression by Merging Technique in the Lattice Model of Summary TablesabstractThe authors investigate the suitability of the cell suppression by merging (CSM) technique as an SDB (statistical database) protection mechanism, and give various heuristic algorithms for the minimum information loss. They first revise the definition for the information loss when query probabilities are taken into account. This definition reflects the actual utilization of cells in the lattice. The authors then propose a heuristic approach to be used with the CSM technique. This approach tries to minimize the information loss by properly choosing the merging pairs. Experimental results show that, in most cases, the information loss is lower than that of the case in which the query probabilities are not considered. This indicates that the actual information loss under query probabilities is low when the CSM technique is combined with the heuristic approach. It is concluded that the CSM technique is therefore an effective protection mechanism for summary tables. It is also shown that the CSM technique is also applicable in the generalized lattice model as an effective security enforcement mechanism. The authors propose several heuristic approaches to minimize the information loss when this technique is applied.> Tzong-An Su, JiYoung Chung, Gultekin Özsoyoglu |
S&P | 3 |
| 1989 | A Family of Incomplete Relational Database Models
Adegbemiga Ola, Gultekin Özsoyoglu |
VLDB | 2 |
| 1989 | Query Processing Techniques in the Summary-Table-by-Example Database Query LanguageabstractSummary-Table-by-Example (STBE) is a graphical language suitable for statistical database applications. STBE queries have a hierarchical subquery structure and manipulate summary tables and relations with set-valued attributes. The hierarchical arrangement of STBE queries naturally implies a tuple-by-tuple subquery evaluation strategy (similar to the nested loops join implementation technique) which may not be the best query processing strategy. In this paper we discuss the query processing techniques used in STBE. We first convert an STBE query into an “extended” relational algebra (ERA) expression. Two transformations are introduced to remove the hierarchical arrangement of subqueries so that query optimization is possible. To solve the “empty partition” problem of aggregate function evaluation, directional join (one-sided outer-join) is utilized. We give the algebraic properties of the ERA operators to obtain an “improved” ERA expression. Finally we briefly discuss the generation of alternative implementations of a given ERA expression. STBE is implemented in a prototype statistical database management system. We discuss the STBE-related features of the implemented system. Gultekin Özsoyoglu, Victor Matos, Z. Meral Özsoyoglu |
ACM Trans. Database Syst. | 1 |
| 1989 | A Relational Calculus with Set Operators, Its Safety and Equivalent Graphical LanguagesabstractThe authors propose a relational calculus (RC/S) which uses set comparison and set manipulation operators to replace universal quantifiers and negations. It is argued that compared to the Codd relational calculus (RC), RC/S queries are easier to construct and comprehend. It is proved that the expressive power of RC is equivalent to the expressive power of RC/S, and algorithms for translating an RC query into an RC/S query and vice versa are given. A safe RC/S query is defined as one that has finite output and can be evaluated in finite time. Then a subset of RC/S queries, called RC/S* is defined, and it is proved that RC/S* is safe. RC/S* is compared to the existing largest safe subsets of RC, i.e. the evaluable formulas and the allowed formulas. Algorithms are given to transform any evaluable formula into an RC/S* query, and some RC/S* formulas that are not evaluable are given. RC/S* queries can be directly implemented using a graphical language similar to Query-by-Example (QBE). Two different graphical languages are described that are equivalent to the RC/S* in expressive power, and these languages are compared to QBE.> Gultekin Özsoyoglu, Huaqing Wang |
IEEE Trans. Software Eng. | 1 |
| 1989 | Time-by-Example Query Language for Historical DatabasesabstractThe authors propose a graphical query language, Time-by-Example (TBE), which has suitable constructs for interacting with historical relational databases in a natural way. TBE is user-friendly. It follows the graphical, two-dimensional approach of such previous languages as Query-by-Example (QBE), Aggregation-by-Example (ABE), and Summary-Table-by-Example (STBE). TBE also uses the hierarchical window (subquery) concept of ABE and STBE. TBE manipulates triple-valued (set-triple-valued) attributes and historical relations. Set-theoretic expressions are followed to deal with time intervals. The BNF specification for TBE is given.> Abdullah Uz Tansel, M. Erol Arkun, Gultekin Özsoyoglu |
IEEE Trans. Software Eng. | 3 |
| 1988 | Statistical Estimators for Relational Algebra ExpressionsabstractPresent database systems process all the data related to a query before giving out responses. As a result, the size of the data to be processed becomes excessive for real-time/time-constrained environments. A new methodology is needed to cut down systematically the time to process the data involved in processing the query. To this end, we propose to use data samples and construct an approximate synthetic response to a given query. Wen-Chi Hou, Gultekin Özsoyoglu, Baldeo K. Taneja |
PODS | 2 |
| 1988 | The Partial Normalized Storage Model of Nested Relations
Aladdin Hafez, Gultekin Özsoyoglu |
VLDB | 2 |
| 1987 | Synthetic Query Response Construction in Scientific Databases with Time Constraints and Incomplete InformationabstractFor some statistical and scientific database applications, data gathering and analysis capabilities can be greatly enhanced if there were real-time/online querying capabilities that return approximate responses to check the quality of the generated data. In some experiments and applications, the data gathered is so large that arguments have been raised for “processing the data on-the-fly” during the execution of an experiment or an application. With real-time/online “on-the-fly” data analysis (querying) capabilities, the size of the stored data in statistical and scientific databases can be kept small and, thus, manageable. Gultekin Özsoyoglu |
ICDE | 1 |
| 1987 | Data Dependencies and Inference Control in Multilevel Relational Database SystemsabstractWe investigate the inference problems due to functional dependencies (FD) and multi-valued dependencies (hND) in a multilevel relational database (MDB) with attribute and record classification schemes, respectively. For FDs, we show that, to prevent compromise, the security levels of attributes must be assigned by using the knowledge of functional dependencies. Under the assumption that all the attributes in the database have been assigned classification levels according to real world requirements, we first determine the set of functional dependencies to be taken into account. Then, we prove that changing the minimum number of attribute levels to prevent compromise is an NP-complete problem. However, assuming that the number of functional dependencies involved in inference is low, we give an exact algorithm to adjust the minimum number of attribute levels so that no compromise due to functional dependencies occurs. For NfVDs, we give a necessary and sufficient condition for compromise due to a single MVD, and then propose an algorithm to prevent single MVD inferences. Tzong-An Su, Gultekin Özsoyoglu |
S&P | 2 |
| 1987 | Extending Relational Algebra and Relational Calculus with Set-Valued Attributes and Aggregate FunctionsabstractIn commercial network database management systems, set-valued fields and aggregate functions are commonly supported. However, the relational database model, as defined by Codd, does not include set-valued attributes or aggregate functions. Recently, Klug extended the relational model by incorporating aggregate functions and by defining relational algebra and calculus languages. In this paper, relational algebra and relational calculus database query languages (as defined by Klug) are extended to manipulate set-valued attributes and to utilize aggregate functions. The expressive power of the extended languages is shown to be equivalent. We extend the relational algebra with three new operators, namely, pack, unpack, and aggregation-by-template. The extended languages form a theoretical framework for statistical database query languages. Gultekin Özsoyoglu, Z. Meral Özsoyoglu, Victor Matos |
ACM Trans. Database Syst. | 1 |
| 1986 | Information Loss in the Lattice Model of Summary Tables due to Cell SuppressionabstractA statistical database (SDB) is a database that is used mostly to provide simple summary statistics (e.g., SUM, COUNT, MAX, MEDIAN, etc.) about individuals in the database and that supports statistical data analysis. When SDB users infer protected information in the SDB from responses to queries, we say that the SDB is compromised. Summary tables are tabular representations of summary data. For a given aggregate function and a set of attributes to specify subsets of individuals in the SDB, all possible (primitive) summary tables form a lattice. The SDB security problem in the lattice model is defined as preventing the users to obtain the information that a table element (i.e., cell) is of size one. In this paper, to solve the SDB security problem in the lattice model, we generalize a technique called cell suppression by merging, and analyze its information loss. Gultekin Özsoyoglu, JiYoung Chung |
ICDE | 1 |
| 1985 | On Optimizing Summary-Table-by-Example QueriesabstractTable-by-Example (STBE) is a graphical language suitable for statistical database applications.STBE queries have a hierarchical subquery structure and manipulate relations with set-valued attributes and summary-tables.The hierarchical structure of queries in STBE naturally implies a tuple-bytuple query evaluation strategy which may not be optimum.This paper describes an algo rithmic approach to transform any STBE query into a relational algebra expression that utilizes an extended set of operators. Gultekin Özsoyoglu, Victor Matos |
PODS | 1 |
| 1985 | A Language and a Physical Organization Technique for Summary Tablesabstractarticle A language and a physical organization technique for summary tables Share on Authors: Gultekin Ozsoyoglu Department of Computer Engineering and Science and Center for Automation and Intelligent Systems, Case Western Reserve University, Cleveland, Ohio Department of Computer Engineering and Science and Center for Automation and Intelligent Systems, Case Western Reserve University, Cleveland, OhioView Profile , Z. Meral Ozsoyoglu View Profile , Francisco Mata Department of Computer Engineering and Science and Center for Automation and Intelligent Systems, Case Western Reserve University, Cleveland, Ohio Department of Computer Engineering and Science and Center for Automation and Intelligent Systems, Case Western Reserve University, Cleveland, OhioView Profile Authors Info & Claims ACM SIGMOD RecordVolume 14Issue 4May 1985 pp 3–16https://doi.org/10.1145/971699.318899Online:01 May 1985Publication History 34citation304DownloadsMetricsTotal Citations34Total Downloads304Last 12 Months5Last 6 weeks1 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteGet Access Gultekin Özsoyoglu, Z. Meral Özsoyoglu, Francisco Mata |
SIGMOD Conference | 1 |
| 1985 | Rounding and Inference Controlin Conceptual Models for Statistical DatabasesabstractA statistical database (SDB) is a database that is used to provide simple summary statistics (e.g., SUM, COUNT, MAX, MEDIAN, etc.) about populations stored in the database and that supports statistical data analysis. When SDB users infer protected information in the SDB from responses to queries, we say that the SDB is compromised. The security problem of SDB is to allow simple summary statistics about protected information in the SDB while preventing compromise. In this paper, we investigate the effectiveness of rounding in statistical databases as a protection technique for SUM and COUNT queries. We consider the generalization hiermchy of the Data Abstraction model and assume there are four different types of inference mechanisms (called R1-R4 range reductions) available for SDB users. For a two-level generalization hierarchy, we find (a) necessary conditions for compromise, and (b) a necessary and sufficient condition for eliminating R1-R4 range reductions. We then describe a procedure for choosing a round-ing base for a tree-organized generalization hierarchy that allows range reductions, but guarantees a minimum range size for protected values in the hierarchy. Gultekin Özsoyoglu, Tzong-An Su |
S&P | 1 |
| 1985 | Statistical Database Query LanguagesabstractDatabases that are mainly used for statistical analysis are called statistical databases (SDB). A statistical database management system (SDBMS) may be defined as a database management system that provides capabilities 1) to model, store, and manipulate data in a manner suitable for the needs of SDB users, and 2) to apply statistical data analysis techniques that range from simple summary statistics to advanced procedures. This paper surveys the existing and proposed SDB data definition and data manipulation (i.e., query) languages. Gultekin Özsoyoglu, Z. Meral Özsoyoglu |
IEEE Trans. Software Eng. | 1 |
| 1984 | Summary-Table-By-Example: A Database Query Language for Manipulating Summary DataabstractIn this paper we introduce the notion of summary table and a high level nonprocedural language, Summary-Table-By-Example (STBE) to manipulate summary data in databases. STBE is similar to Query-by-Example in that it uses graphical two-dimensional objects such as relations and summary tables in formulating a relational database query. STBE is an extension of the Aggregates-by-example database language. STBE may be used in general-purpose databases and statistical databases to extract and format summary data in a tabular form. It is believed to be user-friendly and sufficiently powerful to be used in application areas such as medical research, health planning, energy production and consumption, scientific experiments, political planning and office automation. STBE is relationally complete, i.e., its expressive power is at least that of the relational calculus extended to allow set-valued attributes and aggregate functions. Z. Meral Özsoyoglu, Gultekin Özsoyoglu |
ICDE | 2 |
| 1984 | Statistical Databases
Gultekin Özsoyoglu, Z. Meral Özsoyoglu |
VLDB | 1 |
| 1982 | Auditing and Inference Control in Statistical DatabasesabstractA statistical database (SDB) may be defined as an ordinary database with the capability of providing statistical information to user queries. The security problem for the SDB is to limit the use of the SDB so(that only statistical information is available and no sequence of queries is sufficient to infer protected information about any individual. When such information is obtained, the SDB is said to be compromised. Francis Y. L. Chin, Gultekin Özsoyoglu |
IEEE Trans. Software Eng. | 2 |
| 1982 | Enhancing the Security of Statistical Databases with a Question-Answering System and a Kernel DesignabstractThe security problem of a statistical database is to limit database use so that no private information is deducible. This paper discusses the advantages of using a Question-Answering System and a security kernel to enhance the security constraints at the conceptual model level. An SDB design with the goal of helping the DBA in specifying certain security contraints is proposed. Gultekin Özsoyoglu, Francis Y. L. Chin |
IEEE Trans. Software Eng. | 1 |
| 1981 | Statistical Database DesignabstractThe security problem of a statistical database is to limit the use of the database so that no sequence of statistical queries is sufficient to deduce confidential or private information. In this paper it is suggested that the problem be investigated at the conceptual data model level. The design of a statistical database should utilize a statistical security management facility to enforce the security constraints at the conceptual model level. Information revealed to users is well defined in the sense that it can at most be reduced to nondecomposable information involving a group of individuals. In addition, the design also takes into consideration means of storing the query information for auditing purposes, changes in the database, users' knowledge, and some security measures. Francis Y. L. Chin, Gultekin Özsoyoglu |
ACM Trans. Database Syst. | 2 |