VLDB 2026 Research / reviewers in the wild / expert
Hasan Davulcu
dblp:d/HasanDavulcu
· DBLP profile ↗
48ranked-venue papers
6as first author
6since 2021 · last 2025
0000-0001-5602-8270ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 32 · 6 first-author · 4 since 2021Artificial intelligence and machine learning · 23 · 1 first-author · 4 since 2021Human-computer interaction and ubiquitous computing · 11 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 1 since 2021Security and privacy · 2Software engineering, systems software and programming languages · 2Graphics, computer vision, multimedia, augmented reality and games · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | NARRA-SCALE: Scaling Users and Messaging Through Narrative Detection in Retweet NetworksabstractIn politically charged environments, understanding how ideological narratives emerge, spread, and shape user behavior on social media is critical for applications ranging from misinformation detection to enriching public discourse with verifiable truth. In this study, we present NARRA-SCALE, a framework that brings together network analysis, narrative detection, stance classification, and bipartite scaling to place users, communities, and messages along a single ideological dimension. As a case study, we apply NARRA-SCALE to a U.S. race relations related dataset chiefly polarized between “Black Lives Matter” and “All Lives Matter” supporters. We extract topic coded key phrases and named entities using frequency-based heuristics. Using key phrase co-occurrence relationships and latent representations of matching messages, we mine grouped (entities, issues/aspects, values) triplets characterizing key recurring narratives within the corpus. We use an LLM to summarize messages matching each triplet. Subsequently, a panel of experts label the stance information of the narratives on their key phrases, enabling weak supervision for training a high-accuracy stance detection model which achieves an 81 % F1 score on a held-out gold standard. Next, we construct a signed bipartite graph with colored edges (i.e., representing support versus opposition) between users and key terms mentioned in their messaging corresponding to debated core values, issues, and actors to co-scale their positions on a [-1,+1] range. Our method reaches 91 % agreement with user groups identified through community structure as well as with the “ideal points” of political elites and the general public on Twitter in the U.S. and five European countries. Yusuf Mücahit Çetinkaya, Anshul Trivedi, Vishnu Datta Yanamandala, Michael A. Cowan, Ismail Hakki Toroslu, Hasan Davulcu |
ICTAI | 6 |
| 2025 | Neighborhood search with heuristic-based feature selection for click-through rate prediction
Dogukan Aksu, Ismail Hakki Toroslu, Hasan Davulcu |
Eng. Appl. Artif. Intell. | 3 |
| 2024 | Masking the Bias: From Echo Chambers to Large Scale Aspect-Based Sentiment Analysis
Yeonjung Lee, Yusuf Mücahit Çetinkaya, Emre Külah, Ismail Hakki Toroslu, Hasan Davulcu |
ASONAM (2) | 5 |
| 2022 | Coherent Personalized Paragraph Generation for a Successful Landing PageabstractSocial media has become an important place for online marketing like never before. Businesses use various techniques to identify and reach potential customers across multiple platforms and deliver a message to grab their attention. A notable post could attract potential customers to the product landing page. However, the acquisition is only the beginning. The landing page should respond to the visitor's need for persuasion to increase conversion rates. Showing every visitor the same page is far from that goal. Even if the product meets everyone's needs, their priorities may differ. In this study, we propose a pipeline that includes gathering and identifying potential customers from Twitter, determining their priorities by understanding the context of their message, and creating a coherent paragraph that addresses the issue to display on the landing page. Yusuf Mücahit Çetinkaya, Ismail Hakki Toroslu, Hasan Davulcu |
ASONAM | 3 |
| 2022 | Effective Messaging on Social Media: What Makes Online Content Go Viral?abstractIn this paper, we propose and test three content-based hypotheses that significantly increase message virality. We measure virality as the retweet counts of messages in a pair of real-world Twitter datasets A large dataset - UK Brexit with 51 million tweets from 2.8 million users between June 1, 2015 and May 12, 2019 and a smaller dataset - Nord Stream 2 with 516,000 tweets from 250,000 users between October 1, 2019 and October 15, 2019. We hypothesize, test and conclude that messages incorporating “negativity bias”, “causal arguments” and “threats to personal or societal core values of target audiences” singularly and jointly increase message virality on social media. Maryam Mousavi, Hasan Davulcu, Mohsen Ahmadi, Robert Axelrod, Scott Atran |
WWW | 2 |
| 2021 | High-Order Structure Exploration on Massive Graphs: A Local Graph Clustering PerspectiveabstractModeling and exploring high-order connectivity patterns, also called network motifs, are essential for understanding the fundamental structures that control and mediate the behavior of many complex systems. For example, in social networks, triangles have been proven to play the fundamental role in understanding social network communities; in online transaction networks, detecting directed looped transactions helps identify money laundering activities; in personally identifiable information networks, the star-shaped structures may correspond to a set of synthetic identities. Despite the ubiquity of such high-order structures, many existing graph clustering methods are either not designed for the high-order connectivity patterns, or suffer from the prohibitive computational cost when modeling high-order structures in the large-scale networks. This article generalizes the challenges in multiple dimensions. First ( Model ), we introduce the notion of high-order conductance, and define the high-order diffusion core, which is based on a high-order random walk induced by the user-specified high-order network structure. Second ( Algorithm ), we propose a novel high-order structure-preserving graph clustering framework named HOSGRAP , which partitions the graph into structure-rich clusters in polylogarithmic time with respect to the number of edges in the graph. Third ( Generalization ), we generalize our proposed algorithm to solve the real-world problems on various types of graphs, such as signed graphs, bipartite graphs, and multi-partite graphs. Experimental results on both synthetic and real graphs demonstrate the effectiveness and efficiency of the proposed algorithms. Dawei Zhou 0003, Mehmet Yigit Yildirim, Scott Alcorn, Hanghang Tong, Hasan Davulcu, Jingrui He |
ACM Trans. Knowl. Discov. Data | 6 |
| 2020 | Graph Attention Auto-EncodersabstractAuto-encoders have emerged as a successful framework for unsupervised learning. However, conventional auto-encoders are incapable of utilizing explicit relations in structured data. To take advantage of relations in graph-structured data, several graph auto-encoders have recently been proposed, but they neglect to reconstruct either the graph structure or node attributes. In this paper, we present the graph attention auto-encoder (GATE), a neural network architecture for unsupervised representation learning on graph-structured data. Our architecture is able to reconstruct graph-structured inputs, including both node attributes and the graph structure, through stacked encoder/decoder layers equipped with self-attention mechanisms. In the encoder, by considering node attributes as initial node representations, each layer generates new representations of nodes by attending over their neighbors' representations. In the decoder, we attempt to reverse the encoding process to reconstruct node attributes. Moreover, node representations are regularized to reconstruct the graph structure. Our proposed architecture does not need to know the graph structure upfront, and thus it can be applied to inductive learning. Our experiments demonstrate competitive performance on several node classification benchmark datasets for transductive and inductive tasks, even exceeding the performance of supervised learning baselines in most cases. Amin Salehi, Hasan Davulcu |
ICTAI | 2 |
| 2018 | Detecting Antagonistic and Allied Communities on Social MediaabstractCommunity detection on social media has attracted considerable attention for many years. However, existing methods do not reveal the relations between communities. Communities can form alliances or engage in antagonisms due to various factors, e.g., shared or conflicting goals and values. Uncovering such relations can provide better insights to understand communities and the structure of social media. According to social science findings, the attitudes that members from different communities express towards each other are largely shaped by their community membership. Hence, we hypothesize that intercommunity attitudes expressed among users in social media have the potential to reflect their inter-community relations. Therefore, we first validate this hypothesis in the context of social media. Then, inspired by the hypothesis, we develop a framework to detect communities and their relations by jointly modeling users' attitudes and social interactions. We present experimental results using three real-world social media datasets to demonstrate the efficacy of our framework. Amin Salehi, Hasan Davulcu |
ASONAM | 2 |
| 2018 | Motif-Preserving Dynamic Local Graph CutabstractModeling and characterizing high-order connectivity patterns are essential for understanding many complex systems, ranging from social networks to collaboration networks, from finance to neuroscience. However, existing works on high-order graph clustering assume that the input networks are static. Consequently, they fail to explore the rich high-order connectivity patterns embedded in the network evolutions, which may play fundamental roles in real applications. For example, in financial fraud detection, detecting loops formed by sequenced transactions helps identify money laundering activities; in emerging trend detection, star-shaped structures showing in a short burst may indicate novel research topics in citation networks. In this paper, we bridge this gap by proposing a local graph clustering framework that captures structure-rich subgraphs, taking into consideration the information of high-order structures in temporal networks. In particular, our motif-preserving dynamic local graph cut framework (MOTLOC) is able to model various user-defined temporal network structures and find clusters with minimum conductance in a polylogarithmic time complexity. Extensive empirical evaluations on synthetic and real networks demonstrate the effectiveness and efficiency of our MOTLOC framework. Dawei Zhou 0003, Jingrui He, Hasan Davulcu, Ross Maciejewski |
IEEE BigData | 3 |
| 2018 | Triadic co-clustering of users, issues and sentiments in political tweets
Sefa Sahin Koc, Mert Ozer, Ismail Hakki Toroslu, Hasan Davulcu, Jeremy D. Jordan |
Expert Syst. Appl. | 4 |
| 2017 | A Local Algorithm for Structure-Preserving Graph CutabstractNowadays, large-scale graph data is being generated in a variety of real-world applications, from social networks to co-authorship networks, from protein-protein interaction networks to road traffic networks. Many existing works on graph mining focus on the vertices and edges, with the first-order Markov chain as the underlying model. They fail to explore the high-order network structures, which are of key importance in many high impact domains. For example, in bank customer personally identifiable information (PII) networks, the star structures often correspond to a set of synthetic identities; in financial transaction networks, the loop structures may indicate the existence of money laundering. In this paper, we focus on mining user-specified high-order network structures and aim to find a structure-rich subgraph which does not break many such structures by separating the subgraph from the rest. Dawei Zhou 0003, Mehmet Yigit Yildirim, Scott Alcorn, Hanghang Tong, Hasan Davulcu, Jingrui He |
KDD | 6 |
| 2017 | HiDDen: Hierarchical Dense Subgraph Detection with Application to Financial Fraud DetectionabstractDense subgraphs are fundamental patterns in graphs, and dense subgraph detection is often the key step of numerous graph mining applications. Most of the existing methods aim to find a single subgraph with a high density. However, dense subgraphs at different granularities could reveal more intriguing patterns in the underlying graph. In this paper, we propose to hierarchically detect dense subgraphs. The key idea of our method (HiDDen) is to envision the density of subgraphs as a relative measure to its background (i.e., the subgraph at the coarse granularity). Given that the hierarchical dense subgraph detection problem is essentially a nonconvex quadratic programming problem, we propose effective and efficient alternative projected gradient based algorithms to solve it. The experimental evaluations on real graphs demonstrate that (1) our proposed algorithms find subgraphs with an up to 40% higher density in almost every hierarchy; (2) the densities of different hierarchies exhibit a desirable variety across different granularities; (3) our projected gradient descent based algorithm scales linearly w.r.t the number of edges of the input graph; and (4) our methods are able to reveal interesting patterns in the underlying graphs (e.g., synthetic ID in financial fraud detection). Dawei Zhou 0003, Mehmet Yigit Yildirim, Scott Alcorn, Jingrui He, Hasan Davulcu, Hanghang Tong |
SDM | 6 |
| 2017 | Directional prediction of stock prices using breaking news on TwitterabstractStock market news and investing tips are popular topics in Twitter. In this paper, first we utilize a 5-year financial news corpus comprising over 50,000 articles collected from NASDAQ website matching 30 stock components in Dow Jones Index (DJI) to train a directional stock price prediction system based on news content. Next, we proceed to show that information in articles indicated by breaking Tweet volumes leads to a statistically significant boost in hourly directional prediction accuracies for the DJI stock prices mentioned in these articles. Secondly, we show that using document-level sentiment extraction does not yield a statistically significant boost in the directional predictive accuracies in the presence of other 1-gram keyword features. Thirdly we test the performance of the system on several time-frames and identify 4 hour time-frame for both the price charts and for Tweet breakout detection as the best time-frame. Finally, we develop a set of price momentum based trade exit rules to cut losing trades early and to allow winning trades run longer. We show that the Tweet volume breakout based trading system with price momentum based exit rules not only improve the winning accuracies and the return on investment, but they also lower the maximum drawdown and achieve highest overall return over maximum drawdown. Hana Alostad, Hasan Davulcu |
Web Intell. | 2 |
| 2016 | Co-clustering signed 3-partite graphsabstractIn this paper, we propose a new algorithm, called STRICLUSTER, to find tri-clusters from signed 3-partite graphs. The dataset contains three different types of nodes. Hyperedges connecting three nodes from three different partitions represent either positive or negative relations among those nodes. The aim of our algorithm is to find clusters with strong positive relations among its nodes. Moreover, negative relations up to a certain threshold is also allowed. Also, the clusters can have no overlapping hyperedges. We show the effectiveness of our algorithm via several experiments. Sefa Sahin Koc, Ismail Hakki Toroslu, Hasan Davulcu |
ASONAM | 3 |
| 2016 | Community detection in political Twitter networks using Nonnegative Matrix Factorization methodsabstractCommunity detection is a fundamental task in social network analysis. In this paper, first we develop an endorsement filtered user connectivity network by utilizing Heider's structural balance theory and certain Twitter triad patterns. Next, we develop three Nonnegative Matrix Factorization frameworks to investigate the contributions of different types of user connectivity and content information in community detection. We show that user content and endorsement filtered connectivity information are complementary to each other in clustering politically motivated users into pure political communities. Word usage is the strongest indicator of users' political orientation among all content categories. Incorporating user-word matrix and word similarity regularizer provides the missing link in connectivity-only methods which suffer from detection of artificially large number of clusters for Twitter networks. Mert Ozer, Nyunsu Kim, Hasan Davulcu |
ASONAM | 3 |
| 2016 | Predicting the Location and Time of Mobile Phone Users by Using Sequential Pattern Mining TechniquesabstractIn recent years, using cell phone log data to model human mobility patterns became an active research area. This problem is a challenging data mining problem due to huge size and non-uniformity of the log data, which introduces several granularity levels for the specification of temporal and spatial dimensions. This paper focuses on the prediction of the location of the next activity of the mobile phone users. There are several versions of this problem. In this work, we have concentrated on the following three problems: predicting the location and the time of the next user activity, predicting the location of the next activity of the user when the location of the user changes, and predicting both the location and the time of the activity of the user when the user's location changes. We have developed sequential pattern mining-based techniques for these three problems and validated the success of these methods with real data obtained from one of the largest mobile phone operators in Turkey. Our results are very encouraging, since we were able to obtain quite high accuracy results under small prediction sets. Mert Ozer, Ilkcan Keles, Ismail Hakki Toroslu, Pinar Karagöz, Hasan Davulcu |
Comput. J. | 5 |
| 2016 | A Generalized Hierarchical Multi-Latent Space Model for Heterogeneous LearningabstractIn many real world applications such as image annotation, gene function prediction, and insider threat detection, the data collected from heterogeneous sources often exhibit multiple types of heterogeneity, such as task heterogeneity, view heterogeneity, and label heterogeneity. To address this problem, we propose a Hierarchical Multi-Latent Space (HiMLS) learning framework to jointly model the triple types of heterogeneity. The basic idea is to learn a hierarchical multi-latent space by which we can simultaneously leverage the task relatedness, view consistency and the label correlations to improve the learning performance. We first propose a multi-latent space approach to model the complex heterogeneity, which is then used as a building block to stack up a multi-layer structure in order to learn the hierarchical multi-latent space. In such a way, we can gradually learn the more abstract concepts in the higher level. We present two instantiated models of the generalized framework using different divergence measures. The two-phase learning algorithms are used to train the multi-layer models. We drive the multiplicative update rules for pre-training and fine-tuning in each model, and prove the convergence and correctness of the update methods. The effectiveness of the proposed approach is verified on various data sets. Pei Yang 0001, Hasan Davulcu, Yada Zhu, Jingrui He |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2016 | Exploring Evolving Media Discourse Through Event CueingabstractOnline news, microblogs and other media documents all contain valuable insight regarding events and responses to events. Underlying these documents is the concept of framing, a process in which communicators act (consciously or unconsciously) to construct a point of view that encourages facts to be interpreted by others in a particular manner. As media discourse evolves, how topics and documents are framed can undergo change, shifting the discussion to different viewpoints or rhetoric. What causes these shifts can be difficult to determine directly; however, by linking secondary datasets and enabling visual exploration, we can enhance the hypothesis generation process. In this paper, we present a visual analytics framework for event cueing using media data. As discourse develops over time, our framework applies a time series intervention model which tests to see if the level of framing is different before or after a given date. If the model indicates that the times before and after are statistically significantly different, this cues an analyst to explore related datasets to help enhance their understanding of what (if any) events may have triggered these changes in discourse. Our framework consists of entity extraction and sentiment analysis as lenses for data exploration and uses two different models for intervention analysis. To demonstrate the usage of our framework, we present a case study on exploring potential relationships between climate change framing and conflicts in Africa. Yafeng Lu, Michael Steptoe, Sarah E. Burke, Jiun-Yi Tsai, Hasan Davulcu, Douglas C. Montgomery, Steven R. Corman, Ross Maciejewski |
IEEE Trans. Vis. Comput. Graph. | 6 |
| 2015 | A Dynamic Modularity Based Community Detection Algorithm for Large-scale Networks: DSLMabstractIn this work, a new fast dynamic community detection algorithm for large scale networks is presented. Most of the previous community detection algorithms are designed for static networks. However, large scale social networks are dynamic and evolve frequently over time. To quickly detect communities in dynamic large scale networks, we proposed dynamic modularity optimizer framework (DMO) that is constructed by modifying well-known static modularity based community detection algorithm. The proposed framework is tested using several different datasets. According to our results, community detection algorithms in the proposed framework perform better than static algorithms when large scale dynamic networks are considered. Riza Aktunc, Ismail Hakki Toroslu, Mert Ozer, Hasan Davulcu |
ASONAM | 4 |
| 2015 | Story Detection Using Generalized Concepts and RelationsabstractA major challenge in automated text analysis is that different words are used for related concepts. Analyzing text at the surface level would treat related concepts (i.e. actors, actions, targets, and victims) as different objects, potentially missing common narrative patterns. Shallow parsers reveal semantic roles of words leading to subject-verb-object triplets. We developed a novel algorithm to extract information from triplets by clustering them into generalized concepts by utilizing syntactic criteria based on common contexts and semantic corpus-based statistical criteria based on "contextual synonyms". We show that generalized concepts representation of text (1) overcomes surface level differences (which arise when different keywords are used for related concepts) without drift, (2) leads to a higher-level semantic network representation of related stories, and (3) when used as features, they yield a significant 36% boost in performance for the story detection task. Betul Ceran, Nitesh Kedia, Steven R. Corman, Hasan Davulcu |
ASONAM | 4 |
| 2015 | MUVIR: Multi-View Rare Category Detection
Dawei Zhou 0003, Jingrui He, K. Selçuk Candan, Hasan Davulcu |
IJCAI | 4 |
| 2014 | Rank-One Matrix Pursuit for Matrix CompletionabstractLow rank matrix completion has been applied successfully in a wide range of machine learning applications, such as collaborative filtering, image inpainting and Microarray data imputation. However, many existing algorithms are not scalable to large-scale problems, as they involve computing singular value decomposition. In this paper, we present an efficient and scalable algorithm for matrix completion. The key idea is to extend the well-known orthogonal matching pursuit from the vector case to the matrix case. In each iteration, we pursue a rank-one matrix basis generated by the top singular vector pair of the current approximation residual and update the weights for all rank-one matrices obtained up to the current iteration. We further propose a novel weight updating rule to reduce the time and storage complexity, making the proposed algorithm scalable to large matrices. We establish the linear convergence of the proposed algorithm. The fast convergence is achieved due to the proposed construction of matrix bases and the estimation of the weights. We empirically evaluate the proposed algorithm on many real-world large scale datasets. Results show that our algorithm is much more efficient than state-of-the-art matrix completion algorithms while achieving similar or better prediction performance. Zheng Wang 0011, Ming-Jun Lai, Zhaosong Lu, Wei Fan 0001, Hasan Davulcu, Jieping Ye |
ICML | 5 |
| 2013 | LookingGlass: a visual intelligence platform for tracking online social movementsabstractWe propose a multi-scale text mining methodology and develop a visual intelligence platform for tracking the diffusion of online social movements. The algorithms utilize large amounts of text collected from a wide variety of organizations' media outlets to discover their hotly debated topics, and their discriminative perspectives voiced by opposing camps organized into multiple scales. We utilize discriminating perspectives to classify and map individual Tweeter's message content to social movements based on the perspectives expressed in their weekly tweets. We developed a visual intelligence platform, named LookingGlass, to track the geographical footprint, shifting positions and flows of individuals, topics and perspectives between groups. Nyunsu Kim, Sedat Gokalp, Hasan Davulcu, Mark Woodward |
ASONAM | 3 |
| 2012 | A Semantic Triplet Based Story ClassifierabstractA story is defined as “an actor(s) taking action(s) that culminates in a resolution(s).” In this paper, we investigate the utility of standard keyword based features, statistical features based on shallow-parsing (such as density of POS tags and named entities), and a new set of semantic features to develop a story classifier. This classifier is trained to identify a paragraph as a “story,” if the paragraph contains mostly story(ies). Training data is a collection of expert-coded story and non-story paragraphs from RSS feeds from a list of extremist web sites. Our proposed semantic features are based on suitable aggregation and generalization of; triplets that can be extracted using a parser. Experimental results show that a model of statistical features alongside memory-based semantic linguistic features achieves the best accuracy with a Support Vector Machine (SVM) classifier. Betul Ceran, Ravi Karad, Ajay Mandvekar, Steven R. Corman, Hasan Davulcu |
ASONAM | 5 |
| 2012 | Perspective Analysis for Online DebatesabstractInternet and social media devices created a new public space for online debate on political and social topics. A debate is defined as a formal discussion on a set of related topics in a public meeting, in which opposing perspectives and arguments are put forward. In this paper, we develop automated perspective discovery techniques which would contribute to the understanding of features (i.e. social, political, cultural, religious beliefs, goals, and practices) shared by each side of the debate. Secondly, we show that, compared to a semi-automated process, our perspective discovery algorithms not only identify larger number of relevant features, but they also yield a higher accuracy scaling of moderate to extreme organizations on both sides of a debate. Sukru Tikves, Sedat Gokalp, M'hamed H. Temkit, Sujogya Banerjee, Jieping Ye, Hasan Davulcu |
ASONAM | 6 |
| 2010 | Clustering and mapping related news about violence events on their time-linesabstractKeeping track of news stories and events as they progress can be a tedious job, but as every day routine most of the web users read and follow many stories and events in news. If an analyst in her area has to follow and map all these according to the time-line they happen, the task quickly becomes overwhelming. We present an online tool which attempts to ease the analyst's task of finding all news articles about an event, and sorting and mapping them on a time-line. We implemented an incremental clustering algorithm working on real-time incoming news, experimenting with different feature sets, including named entities and sentence overlap methods. We evaluated these approaches using Document Understand Conference (DUC) datasets. Syed Toufeeq Ahmed, Sukru Tikves, Hasan Davulcu |
ISI | 3 |
| 2009 | Tracking terrorism news threads by extracting event signaturesabstractWith the humongous amount of news stories published daily and the range of ways (RSS feeds, blogs etc) to disseminate them, even an expert at tracking new developing stories can feel the information overload. At most times, when a user is reading a news story, she would like to know ldquowhat happened before this?ldquo or ldquohow things progressed after this incident?rdquo. In this paper, we present a novel real-time yet simple method to detect and track new events related to violence and terrorism in news streams through their life over a time line. We do this by first extracting signature of the event, at microscopic level rather than topic or macroscopic level, and then tracking and linking this event with mentions of same event signature in other incoming news articles. There by forming a thread that links all the news articles that describe this specific event, with no training data used or machine learning algorithms employed. We also present our experimental evaluations conducted with Document Understand Conference (DUC) datasets that validate our observations and methodology. Syed Toufeeq Ahmed, Ruchi Bhindwale, Hasan Davulcu |
ISI | 3 |
| 2008 | Extracting Protein-Protein Interactions from MEDLINE Using Syntactic RolesabstractWith rapid growth in genomics research in last decade, amount of information a biomedical researcher has to keep track of and understand has increased tremendously. We present a fully automated information extraction system to aid these researchers to identify and locate gene and protein interactions in biomedical text. Our extraction system handles complex sentences and extracts multiple and nested interactions specified in these sentences. Experimental evaluations with two other state of the art extraction systems indicate that the IntEx system achieves better performance without the labor intensive pattern engineering requirement. Syed Toufeeq Ahmed, Hasan Davulcu, Chitta Baral |
BIBM | 2 |
| 2008 | Mining Search-Phrase Definitions from Item DescriptionsabstractIn this paper, we develop a model for representing term dependence based on Markov Random Fields and present an approach based on Markov Chain Monte Carlo technique for generating phrase definitions. This approach can use a small corpus of keyword matching and a random sample of other product descriptions for an advertiser's search-phrase to effectively mine and rank alternative but highly relevantsearch-phrasedefinitions. These definitions, which are search-phrases themselves, can then be provided as alternative phrases to an advertiser. Hung V. Nguyen, Hasan Davulcu |
ICDE | 2 |
| 2007 | Fixing Weakly Annotated Web Data Using Relational Models
Fatih Gelgi, Srinivas Vadrevu, Hasan Davulcu |
ICWE | 3 |
| 2007 | Term Ranking for Clustering Web Search Results
Fatih Gelgi, Hasan Davulcu, Srinivas Vadrevu |
WebDB | 2 |
| 2007 | Baum-Welch Style EM Approach on Simple Bayesian Models forWeb Data AnnotationabstractIn this paper, our focus will be on weakly annotated data (WAD) which is typically generated by a (semi) automated information extraction system from the Web documents. The extracted information has a certain level of accuracy which can be surpassed by using statistical models that are capable of contextual reasoning such as Bayesian models. Our contribution is an EM algorithm that operates on simple Bayesian models to re-annotate WAD. EM estimates the parameters, i.e., the prior and conditional probabilities by iterating Bayesian model on the given Web data. In the expectation step, Bayesian classifier is trained from current annotations, and in the maximization step, the roles of all the labels are re-annotated to find the best fitting annotation with the current model then the probabilities are re-adjusted from the new annotations. Our experiments show that EM increases the Web data annotation accuracies up to 8%. We use Baum-Welch methodology in our EM approach. Fatih Gelgi, Hasan Davulcu |
Web Intelligence | 2 |
| 2007 | Scuba Diver: Subspace Clustering of Web Search Results
Fatih Gelgi, Srinivas Vadrevu, Hasan Davulcu |
WEBIST (2) | 3 |
| 2005 | Situation-Awareness for Adaptive Coordination in Service-Based SystemsabstractService-based systems have many applications, including collaborative research and development, e-business, health care, environmental control, military applications, and homeland security. Service coordination is required for these systems to coordinate distributed activities. To achieve adaptive service coordination under changing environment and workload, situation-awareness is needed. In this paper, a model is presented for situation-awareness (SAW) requirements in service-based systems. Based on this model, SAW agents are developed to incorporate situation-awareness and adaptive coordination in service-based systems. Stephen S. Yau, Dazhi Huang, Haishan Gong, Hasan Davulcu |
COMPSAC (1) | 4 |
| 2005 | Adaptable Situation-Aware Secure Service-Based (AS3) SystemsabstractService-oriented systems are distributed systems which have the major advantage of enabling rapid composition of distributed applications, regardless of the programming languages and platforms used in developing and running different components of the applications. In these systems, various capabilities are provided by different organizations as services interconnected by various types of networks. The services can be integrated following a specific workflow to achieve a mission goal for users. For large-scale service-based systems involving multiple organizations, high confidence and adaptability are of prime concern in order to ensure that users can use these systems anywhere, anytime with various devices, knowing that their confidentiality and privacy are well protected and the systems will adapt to satisfy their needs in various situations. Hence, these systems must be adaptable, situation-aware and secure. In this paper, an approach to rapid development of adaptable situation-aware secure service-based (AS/sup 3/) systems is presented. Our approach enables users to rapidly generate, discover, compose services into processes to achieve their goals based on the situation and adapt these processes when situation changes. Stephen S. Yau, Hasan Davulcu, Supratik Mukhopadhyay, Dazhi Huang, Yisheng Yao |
ISORC | 2 |
| 2005 | Boosting Item Keyword Search with Spreading ActivationabstractMost keyword search engines return directly matching keyword phrases. However, publishers cannot anticipate all possible ways in which users would search for the items in their documents. In fact, many times, there may be no direct keyword match between a keyword search phrase and descriptions of relevant items that are perfect matches for the search. We present an automated, high precision-based information retrieval solution to boost item find-ability by bridging the semantic gap between item information and popular keyword search phrases. Our solution achieves an average of 80% F-measure for various boosted matches for keyword search phrases in various categories. Dipti Aswath, Syed Toufeeq Ahmed, James D'cunha, Hasan Davulcu |
Web Intelligence | 4 |
| 2005 | Automated Metadata and Instance Extraction from News Web SitesabstractOver the past few years World Wide Web has established as a vital resource for news. With the continuous growth in the number of available news Web sites and the diversity in their presentation of content, there is an increasing need to organize the news related information on the Web and keep track of it. In this paper, we present automated techniques for extracting metadata instance information by organizing and mining a set of news Web sites. We develop algorithms that detect and utilize HTML regularities in the Web documents to turn them into hierarchical semantic structures encoded as XML. The tree-mining algorithms that we present identify key domain concepts and their taxonomical relationships. We also extract semi-structured concept instances annotated with their labels whenever they are available. We report experimental evaluation for the news domain to demonstrate the efficacy of our algorithms. Srinivas Vadrevu, Saravanakumar Nagarajan, Fatih Gelgi, Hasan Davulcu |
Web Intelligence | 4 |
| 2005 | Improving Web Data Annotations with Spreading Activation
Fatih Gelgi, Srinivas Vadrevu, Hasan Davulcu |
WISE | 3 |
| 2005 | Semantic Partitioning of Web Pages
Srinivas Vadrevu, Fatih Gelgi, Hasan Davulcu |
WISE | 3 |
| 2004 | WinAgent: a system for creating and executing personal information assistants using a web browserabstractWinAgent is a software system for creating and executing Personal Information Assistants (PIAs). These are software robots that can locate and extract targeted data buried deep within a web site. They do so by automatically navigating to relevant sites, locating the correct Web pages (which can be either directly accessed by traversing appropriate links or by filling out HTML forms), and extracting, structuring, and organizing data of interest from these pages into XML. The primary thrust of WinAgent technology effort was to make these tools easy-to-use by users who are not necessarily trained in computing. In particular users create and execute PIAs through a Web Browser. Nikeeta Julasana, Akshat Khandelwal, Anupama Lolage, Prabhdeep Singh, Priyanka Vasudevan, Hasan Davulcu, I. V. Ramakrishnan |
IUI | 6 |
| 2003 | Mining "Hidden Phrase" Definitions from the Web
Hung V. Nguyen, P. Velamuru, Deepak Kolippakkam, Hasan Davulcu, Huan Liu 0001, M. Ates |
APWeb | 4 |
| 2002 | Extraction Techniques for Mining Services from Web SourcesabstractThe Web has established itself as the dominant medium for doing electronic commerce. Consequently the number of service providers, both large and small, advertising their services on the web continues to proliferate. In this paper we describe new extraction algorithms for mining service directories from web pages. We develop a novel propagation technique for identifying and accumulating all of the attributes related to a service entity in a web page. We provide experimental results of the effectiveness of our extraction techniques by mining a database of veterinarian service providers from web sources. Hasan Davulcu, Saikat Mukherjee, I. V. Ramakrishnan |
ICDM | 1 |
| 2002 | A Clustering Technique for Mining Data from Text TablesabstractConsiderable quantities of valuable data about product information and financial statements is often available in sources and formats that are not amenable for querying using traditional database techniques.One such important source is text documents.In such documents these kinds of data often appear in tabular form.A data item in these text tables may span several words (e.g.product description).Furthermore items supposedly within the same column do not necessarily begin or end at the same position.Thus the absence of any regularity in column separators makes it difficult to automatically mine, i.e. extract data items from text tables.Nevertheless an interesting characteristic often exhibited by these tables is that intra-column items are "closer" to each other than inter-column items.We exploit this observation to develop a clustering-based technique to extract data items from these tables.In contrast to previous appproaches, a unique and important aspect of using clustering is that it makes the technique robust in the presence of misalignments.We provide a characterization theorem for text tables on which this technique will always produce a correct extraction.We discuss the design and implementation of a system for extracting tabular data based on this clustering technique.We present experimental evidence of its effectiveness and usability on real industrial data. Hasan Davulcu, Saikat Mukherjee, I. V. Ramakrishnan |
SDM | 1 |
| 2002 | YellowPager: a tool for ontology-based mining of service directories from web sourcesabstractThe web has established itself as the dominant medium for doing electronic commerce. Realizing that its global reach provides significant market and business opportunities, service providers, both large and small are advertising their services on the web. A number of them operate their own web sites promoting their services at length while others are merely listed in a referral site. Aggregating all of the providers into a queriable service directory makes it easy for customers to locate the one most suited for his/her needs.YellowPager is a tool for creating service directories by mining web sources. Service directories created by YellowPager have several merits compared to those generated by existing practices, which typically require participation by service providers (e.g. Verizon's SuperYellowPages.com). Firstly, the information content will be rich. Secondly since the process is automated and repeatable the content can always be kept current. Finally the same process can be readily adapted to different domains.YellowPager builds service directories by mining the web through a combination of keyword-based search engines,web agents, text classifiers and novel extraction algorithms.The extraction is driven by a services ontology consisting of a taxonomy of service concepts and their associated attributes (such as names and addresses) and type descriptions for the attributes. In addition the ontology also associates an extractor function with each attribute. Applying the function to a web page will identify all the occurrences of the attribute in that page.YellowPager's mining algorithm consists of a training step followed by classification and extraction steps. In the training step a classifier is trained to identify web pages relevant to the service of interest. The classification step proceeds by doing a search for the particular service of interest using a keyword based web search engine and retrieves all the matching web pages. From these pages the relevant ones are identified using the classifier. The final step is extraction of attribute values, associated with the service, from these pages. Each web page is parsed into a DOM tree and the extractor functions are applied. All of the attributes corresponding to a service provider are then correctly aggregated. This can pose difficulties especially in the presence of multiple service providers in a page. Using a novel concept of scoring and conflict resolution to prevent erroneous associations of attributes with service provider entities in the page, the algorithm aggregates all the attribute occurrences correctly. The extractor function may not be complete in the sense that it cannot always identify all the attributes in a page. By exploiting the regularity of the sequence in which attributes occurr in referral pages, the mining algorithm automatically learns generalized patterns to locate attributes that the extractor function misses. The distinguishing aspects of YellowPager's extraction algorithm are: (i) it is unsupervised, and (ii) the attribute values in the pages are extracted independent of any page-specific relationships that may exist among the markup tags.YellowPager has been used by a large pet food producer to build a directory of veterinarian service providers in the United States. The resulting database was found to be much larger and richer than that found in Vetquest, Vetworld, and the Super Yellow pages.YellowPager is implemented in JAVA and is interfaced to Rainbow, a library utility in C that is used for classification. The tool will demonstrate the creation of a service directory for any service domain by mining web sources. Prashant Choudhari, Hasan Davulcu, Abhishek Joglekar, Akshay More, Saikat Mukherjee, Supriya Patil, I. V. Ramakrishnan |
SIGIR | 2 |
| 2002 | CuTeX: a system for extracting data from text tablesabstractA wealth of information relevant for e-commerce often appears in text form. This includes specification and performance data sheets of products, financial statements, product offerings etc. Typically these types of product and financial data are published in tabular form. The only separators between items in the table are white spaces and line separators. We will refer to such tables as text tables. Due to the lack of structure in such tables, the information present is not readily queriable using traditional database query languages like SQL. One way to make it amenable to standard database querying techniques is to extract the data items in the tables and create a database out of the extracted data. But extraction from text tables poses difficulties due to the irregularity of the data in the column. Hasan Davulcu, Saikat Mukherjee, Arvind Seth, I. V. Ramakrishnan |
SIGIR | 1 |
| 2000 | Computational Aspects of Resilient Data Extraction from Semistructured SourcesabstractAutomatic data extraction from semistructured sources such as HTML pages is rapidly growing into a problem of significant importance, spurred by the growing popularity of the so called “shopbots” that enable end users to compare prices of goods and other services at various web sites without having to manually browse and fill out forms at each one of these sites. Hasan Davulcu, Guizhen Yang, Michael Kifer, I. V. Ramakrishnan |
PODS | 1 |
| 1999 | A Layered Architecture for Querying Dynamic Web ContentabstractThe design of webbases, database systems for supporting Web-based applications, is currently an active area of research. In this paper, we propose a 3-year architecture for designing and implementing webbases for querying dynamic Web content(i.e., data that can only be extracted by filling out multiple forms). The lowest layer, virtual physical layer, provides navigation independence by shielding the user from the complexities associated with retrieving data from raw Web sources. Next, the traditional logical layer supports site independence. The top layer is analogous to the external schema layer in traditional databases. Hasan Davulcu, Juliana Freire, Michael Kifer, I. V. Ramakrishnan |
SIGMOD Conference | 1 |
| 1998 | Logic Based Modeling and Analysis of WorkflowsabstractWC propose Concurrent Transaction Logic (C7X) as the language for specifying, analyzing, and scheduling of workflows.We show that both local and global properties of worktlows can be naturally represented as C7X formulas and reasoning can be done with the use of the proof theory and the semantics of this logic, We describe a transformation that leads to an eilicicnt algorithm for scheduling worldlows in the presencc of global temporal constraints, which leads to decision proccdurcs for dealing with several safety related properties such as whether every valid execution of the workflow satisfits a particular property or whether a worlcfiow execution is consistent with some given global constraints on the ordering of events in a workflow.We also provide tight complexity results on the running times of these algorithms. Hasan Davulcu, Michael Kifer, C. R. Ramakrishnan 0001, I. V. Ramakrishnan |
PODS | 1 |