VLDB 2026 Research / reviewers in the wild / expert
Anthony Tomasic
dblp:t/AnthonyTomasic
· DBLP profile ↗
50ranked-venue papers
13as first author
3since 2021 · last 2026
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 27 · 10 first-authorHuman-computer interaction and ubiquitous computing · 13 · 2 first-author · 1 since 2021Artificial intelligence and machine learning · 8 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-authorSystems, architecture and hardware · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Human-computer interaction and pervasive computing
8 papers |
Accessibility and assistive technology · 32% Human-AI interaction · 27% Design research and methods · 15% | |
| Databases, data mining, and information retrieval
16 papers |
Transaction processing and concurrency control · 46% Query processing and optimization · 23% Information retrieval · 18% | |
| Software engineering, system software, and programming languages
3 papers |
Program synthesis and code generation · 82% Compilers and program optimization · 12% Software maintenance and evolution · 7% | |
| Computer architecture, parallel and distributed computing, and storage systems
7 papers |
Cloud and datacenter computing · 50% Distributed systems · 30% Storage systems · 11% | |
| Network and information security
3 papers |
Privacy and data protection · 37% Web and mobile security · 37% Cryptographic primitives and cryptanalysis · 26% |
Topics — the 30 heaviest of 63, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Accessibility and assistive technology › assistive technology for visual impairment
assistive technology for blind and low-vision users |
1.0 | 1 | 2026 | Navigation and Interaction for Blind Users via a Cognitive Architecture · AAAI 2026 |
Human-AI interaction
conversational agents |
1.0 | 1 | 2026 | Navigation and Interaction for Blind Users via a Cognitive Architecture · AAAI 2026 |
Design research and methods
research through design |
0.6 | 1 | 2022 | Recentering Reframing as an RtD Contribution: The Case of Pivoting from Accessible Web Tables to a Conversational Internet · CHI 2022 |
Accessibility and assistive technology
screen reader accessibility |
0.6 | 1 | 2022 | Recentering Reframing as an RtD Contribution: The Case of Pivoting from Accessible Web Tables to a Conversational Internet · CHI 2022 |
Transaction processing and concurrency control
transaction scheduling |
0.3 | 1 | 2018 | Performance of OLTP via Intelligent Scheduling · ICDE 2018 |
Program synthesis and code generation
neural program synthesis |
0.3 | 1 | 2018 | Retrieval-Based Neural Code Generation · EMNLP 2018 |
Program synthesis and code generation › code generation with language models
retrieval-augmented code generation |
0.3 | 1 | 2018 | Retrieval-Based Neural Code Generation · EMNLP 2018 |
Human-AI interaction
mixed-initiative interaction |
0.2 | 2 | 2009 | User-created forms as an effective method of human-agent communication · CHI 2009 Vio: a mixed-initiative approach to learning and automating procedural update tasks · CHI 2007 |
Query processing and optimization
query result caching |
0.2 | 2 | 2008 | Scalable query result caching for web applications · Proc. VLDB Endow. 2008 Invalidation Clues for Database Scalability Services · ICDE 2007 |
Design research and methods › design practice
service design |
0.1 | 2 | 2011 | Understanding the space for co-design in riders' interactions with a transit service · CHI 2010 Field trial of Tiramisu: crowd-sourcing bus arrival times to spur co-design · CHI 2011 |
Ubiquitous computing and smart environments › urban computing
smart city |
0.1 | 1 | 2011 | Field trial of Tiramisu: crowd-sourcing bus arrival times to spur co-design · CHI 2011 |
Query processing and optimization › query rewriting
query transformation |
0.1 | 2 | 2009 | Holistic Query Transformations for Dynamic Web Applications · ICDE 2009 The Distributed Information Search Component (Disco) and the World Wide Web · SIGMOD Conference 1997 |
Collaborative and social computing › civic engagement
civic technology |
0.1 | 1 | 2010 | Understanding the space for co-design in riders' interactions with a transit service · CHI 2010 |
Human-AI interaction
human-agent interaction |
0.1 | 1 | 2009 | User-created forms as an effective method of human-agent communication · CHI 2009 |
Compilers and program optimization
program transformation |
0.1 | 1 | 2009 | Holistic Query Transformations for Dynamic Web Applications · ICDE 2009 |
Natural language and speech › Question answering and dialogue systems › conversational agents
conversational assistant |
0.1 | 1 | 2008 | RADAR: A Personal Assistant that Learns to Reduce Email Overload · AAAI 2008 |
Distributed systems › consistency models
cache consistency |
0.1 | 1 | 2008 | Scalable query result caching for web applications · Proc. VLDB Endow. 2008 |
Natural language and speech › Question answering and dialogue systems
intent detection |
0.1 | 1 | 2007 | Learning information intent via observation · WWW 2007 |
Interaction techniques and input
text entry |
0.1 | 1 | 2007 | Vio: a mixed-initiative approach to learning and automating procedural update tasks · CHI 2007 |
Web and mobile security › phishing detection
phishing email detection |
0.1 | 1 | 2007 | Learning to detect phishing emails · WWW 2007 |
Privacy and data protection
privacy-preserving data management |
0.1 | 1 | 2007 | Invalidation Clues for Database Scalability Services · ICDE 2007 |
Cryptographic primitives and cryptanalysis
encryption |
0.1 | 1 | 2006 | Simultaneous scalability and security for data-intensive web applications · SIGMOD Conference 2006 |
Information retrieval
distributed information retrieval |
0.1 | 4 | 1999 | GlOSS: Text-Source Discovery over the Internet · ACM Trans. Database Syst. 1999 Data Structures for Efficient Broker Implementation · ACM Trans. Inf. Syst. 1997 The Effectiveness of GlOSS for the Text Database Discovery Problem · SIGMOD Conference 1994 |
Software maintenance and evolution
software evolution |
0.1 | 1 | 2005 | Learning to Understand Web Site Update Requests · IJCAI 2005 |
Data integration and cleaning
heterogeneous data source integration |
0.0 | 2 | 1998 | Scaling Access to Heterogeneous Data Sources with DISCO · IEEE Trans. Knowl. Data Eng. 1998 Leveraging Mediator Cost Models with Heterogeneous Data Sources · ICDE 1998 |
Design research and methods
participatory design |
0.0 | 1 | 2011 | Field trial of Tiramisu: crowd-sourcing bus arrival times to spur co-design · CHI 2011 |
Information retrieval
retrieval models |
0.0 | 2 | 1999 | GlOSS: Text-Source Discovery over the Internet · ACM Trans. Database Syst. 1999 The Effectiveness of GlOSS for the Text Database Discovery Problem · SIGMOD Conference 1994 |
Information retrieval › indexing
inverted index |
0.0 | 3 | 1994 | Incremental Updates of Inverted Lists for Text Document Retrieval · SIGMOD Conference 1994 Query Processing and Inverted Indices in Shared-Nothing Document Information Retrieval Systems · VLDB J. 1993 Caching and Database Scaling in Distributed Shard-Nothing Information Retrieval Systems · SIGMOD Conference 1993 |
Query processing and optimization › query optimization
cost-based optimization |
0.0 | 2 | 1998 | Leveraging Mediator Cost Models with Heterogeneous Data Sources · ICDE 1998 The Distributed Information Search Component (Disco) and the World Wide Web · SIGMOD Conference 1997 |
Distributed and cloud data management
distributed query processing |
0.0 | 2 | 1998 | The Distributed Information Search Component (Disco) and the World Wide Web · SIGMOD Conference 1997 Leveraging Mediator Cost Models with Heterogeneous Data Sources · ICDE 1998 |
Methods — techniques the papers use, named apart from their topics
object detection · 1.0large language model · 1.0augmented reality · 1.0reflective case study · 0.6design experiment · 0.6transaction representation · 0.3subtree retrieval · 0.3neural encoder-decoder · 0.3intelligent scheduling · 0.3dynamic-programming sentence similarity · 0.3source-to-source compilation · 0.3latency hiding · 0.3machine learning · 0.2invalidation clues · 0.2encryption · 0.2longitudinal study · 0.2field experiment · 0.2publish-subscribe · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Navigation and Interaction for Blind Users via a Cognitive ArchitectureabstractNavigating new indoor spaces and interacting with the environment presents many challenges for people who are blind or have low vision (BLV). To address these challenges, we prototyped a smartphone-based conversational assistant that helps BLV people navigate and interact with their environment. The prototype utilizes a cognitive architecture to integrate three different technologies: (i) augmented-reality spatial anchors for high-precision localization and access to static information about the environment; (ii) real-time object/people detection for information about the environment and obstacle avoidance; and (iii) a conversational agent}that uses large language models (LLMs) for information extraction, conversational interaction, and turn-by-turn navigation. We assess the impact of different technologies on human performance by measuring user task time and errors. We found that conversational interaction holistically integrates the different technologies to deliver a better user experience while significantly reducing task completion time. Oscar J. Romero, Anthony Tomasic, Elizabeth J. Carter, John Zimmerman, Aaron Steinfeld |
AAAI | 2 |
| 2022 | Recentering Reframing as an RtD Contribution: The Case of Pivoting from Accessible Web Tables to a Conversational InternetabstractDesign produces valuable knowledge by offering new perspectives that reframe problematic situations. Research through Design (RtD) contributes new frames along with design work demonstrating a frame's value. Interestingly, RtD papers rarely describe how reframing happens. This gap in documentation unintentionally implies a romantic account of design, it implies that the first step of an RtD project is to have a brilliant idea. This is especially problematic in cases where the reframing causes a pivot that leads to a new research program. To help address this gap, we describe a case where through a series of three design experiments we experienced a research pivot. We describe how our work to improve web-table navigation for screen-reader users broke our frame. The break led to a new research program focused on constructing a conversational internet. This paper offers our case along with reflection on reporting design work that drives reframing. John Zimmerman, Aaron Steinfeld, Anthony Tomasic, Oscar J. Romero |
CHI | 3 |
| 2021 | A Task-Oriented Dialogue Architecture via Transformer Neural Language Models and Symbolic InjectionabstractRecently, transformer language models have been applied to build both task-and non-taskoriented dialogue systems.Although transformers perform well on most of the NLP tasks, they perform poorly on context retrieval and symbolic reasoning.Our work aims to address this limitation by embedding the model in an operational loop that blends both natural language generation and symbolic injection.We evaluated our system on the multi-domain DSTC8 data set and reported joint goal accuracy of 75.8% (ranked among the first half positions), intent accuracy of 97.4% (which is higher than the reported literature), and a 15% improvement for success rate compared to a baseline with no symbolic injection.These promising results suggest that transformer language models can not only generate proper system responses but also symbolic representations that can further be used to enhance the overall quality of the dialogue management as well as serving as scaffolding for complex conversational reasoning. Oscar J. Romero, Antian Wang, John Zimmerman, Aaron Steinfeld, Anthony Tomasic |
SIGDIAL | 5 |
| 2020 | A Long-Term Evaluation of Adaptive Interface Design for Mobile Transit InformationabstractPersonalization of user experience has a long history of success in the HCI community. More recently the community has focused on adaptive user interfaces, supported by machine learning, that reduce interaction efforts and improves user experience by collapsing transactions and pre-filtering results. However, generally, these more recent results have only been demonstrated in the laboratory environment. In this paper, we share the case of a deployed mobile transit app that adapts based on users’ previous usage. We examine the impact of adaptation, both good and bad, and user abandonment rates. We conducted an 18-month assessment where 2,616 participants (with and without vision impairments) were recruited and participated in an A/B study. Finally, we draw some insights on some unusual effects that appear over the long term. Oscar J. Romero, Alexander Haig, Lynn Kirabo, Qian Yang 0004, John Zimmerman, Anthony Tomasic, Aaron Steinfeld |
MobileHCI | 6 |
| 2018 | Contingent Responsiveness in Digital Storybooks: Effects on Children's Comprehension and the Role of Individual Differences in Attention
Cassondra M. Eng, Anthony Tomasic, Erik D. Thiessen |
CogSci | 2 |
| 2018 | Retrieval-Based Neural Code GenerationabstractIn models to generate program source code from natural language, representing this code in a tree structure has been a common approach.However, existing methods often fail to generate complex code correctly due to a lack of ability to memorize large and complex structures.We introduce RECODE, a method based on subtree retrieval that makes it possible to explicitly reference existing code examples within a neural code generation model.First, we retrieve sentences that are similar to input sentences using a dynamicprogramming-based sentence similarity scoring method.Next, we extract n-grams of action sequences that build the associated abstract syntax tree.Finally, we increase the probability of actions that cause the retrieved n-gram action subtree to be in the predicted code.We show that our approach improves the performance on two code generation tasks by up to +2.6 BLEU. 1 Shirley Anugrah Hayati, Raphaël Olivier, Pravalika Avvaru, Anthony Tomasic, Graham Neubig |
EMNLP | 5 |
| 2018 | Performance of OLTP via Intelligent SchedulingabstractCurrent architectures for main-memory online transaction processing (OLTP) database management systems (DBMS) typically use random scheduling to assign transactions to threads. This approach achieves uniform load across threads but it ignores the likelihood of conflicts between transactions. If the DBMS could estimate the potential for transaction conflict and then intelligently schedule transactions to avoid conflicts, then the system could improve its performance. Such estimation of transaction conflict, however, is non-trivial for several reasons. First, conflicts occur under complex conditions that are far removed in time from the scheduling decision. Second, transactions must be represented in a compact and efficient manner to allow for fast conflict detection. Third, given some evidence of potential conflict, the DBMS must schedule transactions in such a way that minimizes this conflict. In this paper, we systematically explore the design decisions for solving these problems. We then empirically measure the performance impact of different representations on a standard OLTP benchmark. Tieying Zhang, Anthony Tomasic, Yangjun Sheng, Andrew Pavlo |
ICDE | 2 |
| 2017 | Self-Driving Database Management Systems
Andrew Pavlo, Gustavo Angulo, Joy Arulraj, Haibin Lin, Jiexi Lin, Lin Ma 0006, Prashanth Menon, Todd C. Mowry, Matthew Perron, Ian Quah, Siddharth Santurkar, Anthony Tomasic, Skye Toor, Dana Van Aken, Ziqi Wang 0007, Yingjun Wu, Ran Xian, Tieying Zhang |
CIDR | 12 |
| 2016 | Planning Adaptive Mobile Experiences When WireframingabstractMachine learning improves mobile user experience. Interestingly, envisioning apps with adaptive interfaces that reduce navigation and selection effort is not standard UX practice. When implementing an adaptive UI for our mobile transit app, we encountered a number of problems. Our original design did not log necessary information nor did it induce users to provide good labels. On reflection, we realized UX designers should identify and refine UI adaptions when sketching wireframes. To advance on this insight, we reviewed the interfaces of popular apps and extracted six design patterns where UI adaptation can improve in-app navigation. Next, we designed an exemplar set of wireframes, illustrating how UX designers might annotate their interaction flows to communicate planned adaptation and note the information (logs and labels) needed to make the desired inferences. Qian Yang 0004, John Zimmerman, Aaron Steinfeld, Anthony Tomasic |
Conference on Designing Interactive Systems | 4 |
| 2016 | The utility of tables for screen reader usersabstractThe mere existence of digital documents on the web makes those documents more accessible to blind people than if they were not digitized. The web provides several opportunities, however, for web authors to describe what content looks like, rather than what it means, and this visual description inconveniences and confuses people using screen readers. We focus on the particular instance of this problem where information that is logically tabular is presented using repeated visual templates. We recruited a group of experienced screen reader users and observed them performing a series of tasks designed to elicit tabular behavior. We find that people who make use of tabular operations perform the tasks faster and more successfully than those who do not. We furthermore find that while many users discover the tables we introduced as part of the study, many others do not or do not choose to make use of the tables. We discuss some techniques for enhancing the discoverability of the introduced tables. Steven Gardiner, Anthony Tomasic, John Zimmerman |
CCNC | 2 |
| 2016 | Combining contribution interactions to increase coverage in mobile participatory sensing systemsabstractParticipatory sensing systems use people and their smartphones as a sensing infrastructure, and getting people to make contributions remains a critical challenge. Little work details how system designers should combine different interactions to increase coverage of service location. Tiramisu, a participatory sensing system, invites transit riders to crowdsource real-time arrival information by sharing location traces when they commute. We extended this system with a new feature that allows riders at stops to "spot" buses passing by. To better understand the impact of this new feature, we conducted an observational log analysis, examining changes in coverage and user behavior before and after the new feature. Following the addition of the spotting feature, participants' contributions increased coverage (the number of trips with real-time data) by 98%, and they used the app more than twice as much. The addition of the spotting feature was also followed by a significant increase of trace contributions. Yun Huang 0003, John Zimmerman, Anthony Tomasic, Aaron Steinfeld |
MobileHCI | 3 |
| 2015 | EnTable: Rewriting Web Data Sets as Accessible TablesabstractToday, many data-driven web pages present information in a way that is difficult for blind and low vision users to navigate and to understand. EnTable addresses this challenge. It re-writes confusing and complicated template-based data sets as accessible tables. EnTable allows blind and low vision users to submit requests for pages they wish to access. The system then employs sighted informants to markup the desired page with semantic information, allowing the page to be re-written using straightforward tags. Screen reader users who browse the web using the EnTable browser extension can report data sets that are confusing, and utilize data sets re-written with the tag based on their own requests or on the requests of other users. Steven Gardiner, Anthony Tomasic, John Zimmerman |
ASSETS | 2 |
| 2014 | Motivating contribution in a participatory sensing system via quid-pro-quoabstractParticipatory sensing systems (PSS) require frequent injection of information that has a short shelf-life. The use of crowds to gather information for PSS is therefore particularly challenging. In this study, we explore the impact of two policies on user contributions. A quid-pro-quo policy exchanges contributions from users for access to critical information in the system. A request policy simply reminds the user that information is needed to make the system function well. Prior research has shown that request for help in crowdsourced system is an effective mechanism to increase contributions. During a large-scale experimental study within a publicly deployed, crowdsourced, transit information system, we analyzed metrics associated with frequency of contribution and commitment to long-term use over a 10-month period. Our results confirmed that quid-pro-quo led to more contribution, but at a cost of faster departure from the study. When a participant was simply requested to contribute, but could still access community-generated data if they ignored a request, was largely ineffective and was statistically similar to the control condition where no request for contribution occurred. Thus crowdsource system designers should consider imposing quid-pro-quo type policies for PSS that concentrate on fewer users, but makes them more productive. Anthony Tomasic, John Zimmerman, Aaron Steinfeld, Yun Huang 0003 |
CSCW | 1 |
| 2014 | Citizen Motivation on the Go: The Role of Psychological EmpowermentabstractAlthough advances in technology now enable people to communicate ‘anytime, anyplace’, it is not clear how citizens can be motivated to actually do so. This paper evaluates the impact of three principles of psychological empowerment, namely perceived self-efficacy, sense of community and causal importance, on public transport passengers’ motivation to report issues and complaints while on the move. A week-long study with 65 participants revealed that self-efficacy and causal importance increased participation in short bursts and increased perceptions of service quality over longer periods. Finally, we discuss the implications of these findings for citizen participation projects and reflect on design opportunities for mobile technologies that motivate citizen participation. Jorge Gonçalves 0001, Vassilis Kostakos, Evangelos Karapanos, Mary Barreto, Tiago Camacho, Anthony Tomasic, John Zimmerman |
Interact. Comput. | 6 |
| 2013 | Energy efficient and accuracy aware (E2A2) location services via crowdsourcingabstractMany mobile applications rely on location information gained from location services on mobile devices. However, continuously tracking the device location with high accuracy drains the battery quickly. Furthermore, sensing the same location can be redundant when multiple devices are co-located. In this paper, we develop a crowdsourcing-based location service, E2A2 (energy efficient and accuracy aware), which places colo-cated devices into groups, and uses group location to represent individual device location. The E2A2 location service aims to reduce individual device battery consumption associated with location services while simultaneously maintaining high location accuracy for each device. Our experimental results from a prototype system show the effectiveness of our proposed solution with different mobility patterns. We also present results on the impact of different system parameters and the number of users in a group. Compared to running GPS location services on individual devices separately, our E2A2 service saves on average 33% battery consumption rate when 4 devices are co-located at walking speed and 26% battery consumption rate when 4 devices are colocated on the same bus while meeting the same accuracy requirements. Yun Huang 0003, Anthony Tomasic, Yufei An, Charles Garrod, Aaron Steinfeld |
WiMob | 2 |
| 2011 | Field trial of Tiramisu: crowd-sourcing bus arrival times to spur co-designabstractCrowd-sourcing social computing systems represent a new material for HCI designers. However, these systems are difficult to work with and to prototype, because they require a critical mass of participants to investigate social behavior. Service design is an emerging research area that focuses on how customers co-produce the services that they use, and thus it appears to be a great domain to apply this new material. To investigate this relationship, we developed Tiramisu, a transit information system where commuters share GPS traces and submit problem reports. Tiramisu processes incoming traces and generates real-time arrival time predictions for buses. We conducted a field trial with 28 participants. In this paper we report on the results and reflect on the use of field trials to evaluate crowd-sourcing prototypes and on how crowd sourcing can generate co-production between citizens and public services. John Zimmerman, Anthony Tomasic, Charles Garrod, Daisy Yoo, Chaya Hiruncharoenvate, Rafae Aziz, Nikhil Ravi Thiruvengadam, Yun Huang 0003, Aaron Steinfeld |
CHI | 2 |
| 2011 | Mixer: Mixed-Initiative Data Retrieval and Integration by Example
Steven Gardiner, Anthony Tomasic, John Zimmerman, Rafae Aziz, Kathryn Rivard |
INTERACT (1) | 2 |
| 2010 | Understanding the space for co-design in riders' interactions with a transit serviceabstractThe recent advances in web 2.0 technologies and the rapid adoption of smart phones raises many opportunities for public services to improve their services by engaging their users (who are also owners of the service) in co-design: a dialog where users help design the services they use. To investigate this opportunity, we began a service design project investigating how to create repeated information exchanges between riders and a transit agency in order to create a virtual "place" from which the dialog on services could take place. Through interviews with riders, a workshop with a transit agency, and speed dating of design concepts, we have developed a design direction. Specifically, we propose a service that combines vehicle location and "fullness" ratings provided by riders with dynamic route change information from the transit agency as a foundation for a dialog around riders conveying input for continuous service improvement. Daisy Yoo, John Zimmerman, Aaron Steinfeld, Anthony Tomasic |
CHI | 4 |
| 2009 | User-created forms as an effective method of human-agent communicationabstractA key challenge for mixed-initiative systems is to create a shared understanding of the task between human and agent. To address this challenge, we created a mixed-initiative interface called Mixer to aid administrators with automating tedious information-retrieval tasks. Users initiate communication with the agent by constructing a form, creating a structure to hold the information they require and to show context in order to interpret this information. They then populate the form with the desired results, demonstrating to the agent the steps required to retrieve the information. This method of form creation explicitly defines the shared understanding between human and agent. An evaluation of the interface shows that administrators can effectively create forms to communicate with the agent, that they are likely to accept this technology in their work environment, and that the agent's help can significantly reduce the time they spend on repeated information-retrieval tasks. John Zimmerman, Kathryn Rivard, Ian Hargraves, Anthony Tomasic, Ken Mohnkern |
CHI | 4 |
| 2009 | Holistic Query Transformations for Dynamic Web ApplicationsabstractA promising approach to scaling Web applications is to distribute the server infrastructure on which they run. This approach, unfortunately, can introduce latency between the application and database servers, which in turn increases the network latency of Web interactions for the clients (end users). In this paper we introduce the concept of source-to-source holistic transformations - transformations that seek to optimize both the application code and the database requests made by it, to reduce client latency. As examples of our concept, we propose and evaluate two source-to-source holistic transformations that focus on hiding the latencies of database queries. We argue that opportunities for applying these transformations will continue to exist in Web applications. We then present algorithms for automating these transformations in a source-to-source compiler. Finally, we evaluate the effect of these two transformations on three realistic Web benchmark applications, both in the traditional centralized setting and a distributed setting. Amit Manjhi, Charles Garrod, Bruce M. Maggs, Todd C. Mowry, Anthony Tomasic |
ICDE | 5 |
| 2008 | RADAR: A Personal Assistant that Learns to Reduce Email Overload
Michael Freed, Jaime G. Carbonell, Geoffrey J. Gordon, Jordan Hayes, Brad A. Myers, Daniel P. Siewiorek, Stephen F. Smith, Aaron Steinfeld, Anthony Tomasic |
AAAI | 9 |
| 2008 | Scalable query result caching for web applicationsabstractThe backend database system is often the performance bottleneck when running web applications. A common approach to scale the database component is query result caching, but it faces the challenge of maintaining a high cache hit rate while efficiently ensuring cache consistency as the database is updated. In this paper we introduce Ferdinand, the first proxy-based cooperative query result cache with fully distributed consistency management. To maintain a high cache hit rate, Ferdinand uses both a local query result cache on each proxy server and a distributed cache. Consistency management is implemented with a highly scalable publish/subscribe system. We implement a fully functioning Ferdinand prototype and evaluate its performance compared to several alternative query-caching approaches, showing that our high cache hit rate and consistency management are both critical for Ferdinand's performance gains over existing systems. Charles Garrod, Amit Manjhi, Anastasia Ailamaki, Bruce M. Maggs, Todd C. Mowry, Christopher Olston, Anthony Tomasic |
Proc. VLDB Endow. | 7 |
| 2007 | Vio: a mixed-initiative approach to learning and automating procedural update tasksabstractToday many workers spend too much of their time translating their co-workers' requests into structures that information systems can understand. This paper presents the novel interaction design and evaluation of VIO, an agent that helps workers trans late request. VIO monitors requests and makes suggestions to speed up the translation. VIO allows users to quickly correct agent errors. These corrections are used to improve agent performance as it learns to automate work. Our evaluations demonstrate that this type of agent can significantly reduce task completion time, freeing workers from mundane tasks. John Zimmerman, Anthony Tomasic, Isaac Simmons, Ian Hargraves, Ken Mohnkern, Jason Cornwell, Robert Martin McGuire |
CHI | 2 |
| 2007 | Invalidation Clues for Database Scalability ServicesabstractFor their scalability needs, data-intensive Web applications can use a database scalability service (DBSS), which caches applications' query results and answers queries on their behalf. One way for applications to address their security/privacy concerns when using a DBSS is to encrypt all data that passes through the DBSS. Doing so, however, causes the DBSS to invalidate large regions of its cache when data updates occur. To invalidate more precisely, the DBSS needs help in order to know which results to invalidate; such help inevitably reveals some properties about the data. In this paper, we present invalidation clues, a general technique that enables applications to reveal little data to the DBSS, yet limit the number of unnecessary invalidations. Compared with previous approaches, invalidation clues provide applications significantly improved tradeoffs between security/privacy and scalability. Our experiments using three Web application benchmarks, on a prototype DBSS we have built, confirm that invalidation clues are indeed a low-overhead, effective, and general technique for applications to balance their privacy and scalability needs. Amit Manjhi, Phillip B. Gibbons, Anastasia Ailamaki, Charles Garrod, Bruce M. Maggs, Todd C. Mowry, Christopher Olston, Anthony Tomasic |
ICDE | 8 |
| 2007 | Learning to detect phishing emailsabstractEach month, more attacks are launched with the aim of making web users believe that they are communicating with a trusted entity for the purpose of stealing account information, logon credentials, and identity information in general. This attack method, commonly known as "phishing," is most commonly initiated by sending out emails with links to spoofed websites that harvest information. We present a method for detecting these attacks, which in its most general form is an application of machine learning on a feature set designed to highlight user-targeted deception in electronic communication. This method is applicable, with slight modification, to detection of phishing websites, or the emails used to direct victims to these sites. We evaluate this method on a set of approximately 860 such phishing emails, and 6950 non-phishing emails, and correctly identify over 96% of the phishing emails while only mis-classifying on the order of 0.1% of the legitimate emails. We conclude with thoughts on the future for such techniques to specifically identify deception, specifically with respect to the evolutionary nature of the attacks and information available. Ian Fette, Norman M. Sadeh, Anthony Tomasic |
WWW | 3 |
| 2007 | Learning information intent via observationabstractUsers in an organization frequently request help by sending request messages to assistants that express information intent: an intention to update data in an information system. Human assistants spend a significant amount of time and effort processing these requests. For example, human resource assistants process requests to update personnel records, and executive assistants process requests to schedule conference rooms or to make travel reservations. To process the intent of a request, assistants read the request and then locate, complete, and submit a form that corresponds to the expressed intent. Automatically or semi-automatically processing the intent expressed in a request on behalf of an assistant would ease the mundane and repetitive nature of this kind of work.For a well-understood domain, a straightforward application of natural language processing techniques can be used to build an intelligent form interface to semi-automatically process information intent request messages. However, high performance parsers are based on machine learning algorithms that require a large corpus of examples that have been labeled by an expert. The generation of a labeled corpus of requests is a major barrier to the construction of a parser. In this paper, we investigate the construction of a natural language processing system and an intelligent form system that observes an assistant processing requests. The intelligent form system then generates a labeled training corpus by interpreting the observations. This paper reports on the measurement of the performance of the machine learning algorithms based on real data. The combination of observations, machine learning and interaction design produces an effective intelligent form interface based on natural language processing. Anthony Tomasic, Isaac Simmons, John Zimmerman |
WWW | 1 |
| 2006 | Processing information intent via weak labelingabstractNo abstract available. Anthony Tomasic, Isaac Simmons, John Zimmerman |
CIKM | 1 |
| 2006 | Linking messages and form requestsabstractLarge organizations with sophisticated infrastructures have large form-based systems that manage the interaction between the user community and the infrastructure. In many cases, when a user needs to complete a form to accomplish a task, the user e-mails a description of the task to the appropriate form expert. In many cases this description is incomplete and the expert engages in a clarification dialog to determine the details of the task. Since many tasks and descriptions are routine, this e-mail dialog can be replaced with an intelligent user interface. The interface proactively reads e-mail (or IM) messages and assists the user in completing the associated task without involving the expert. To ground our vision in a specific application, we have built an agent that functions as a webmaster assistant. For example, a user emails the request: "Change John Doe's home phone number to 800-555-1212" to the agent. The webmaster agent then replies with the biographical data form displaying information about John Doe with the new phone number pre-filled in the form. The user then simply approves the change.In this paper we describe a prototype website maintenance agent that (i) allows users to express the updates they want to make in human terms (free text input expression of intent), and (ii) allows users to quickly repair any inference errors the agent makes. In addition, we present the results of a proof of concept study that details how interacting with a webmaster agent that makes inference errors is both more efficient (faster) and more effective (errors made to site) than sending a request to a human webmaster. We conclude the paper with a discussion of the application of our work to any form-based system. Anthony Tomasic, John Zimmerman, Isaac Simmons |
IUI | 1 |
| 2006 | NER Systems that Suit User's Preferences: Adjusting the Recall-Precision Trade-off for Entity Extraction
Einat Minkov, Richard C. Wang, Anthony Tomasic, William W. Cohen |
HLT-NAACL | 3 |
| 2006 | Simultaneous scalability and security for data-intensive web applicationsabstractFor Web applications in which the database component is the bottleneck, scalability can be provided by a third-party Database Scalability Service Provider (DSSP) that caches application data and supplies query answers on behalf of the application. Cost-effective DSSPs will need to cache data from many applications, inevitably raising concerns about security. However, if all data passing through a DSSP is encrypted to enhance security, then data updates trigger invalidation of large regions of cache. Consequently, achieving good scalability becomes virtually impossible. There is a tradeoff between security and scalability, which requires careful consideration.In this paper we study the security-scalability tradeoff, both formally and empirically. We begin by providing a method for statically identifying segments of the database that can be encrypted without impacting scalability. Experiments over a prototype DSSP system show the effectiveness of our static analysis method--for all three realistic bench-mark applications that we study, our method enables a significant fraction of the database to be encrypted without impacting scalability. Moreover, most of the data that can be encrypted without impacting scalability is of the type that application designers will want to encrypt, all other things being equal. Based on our static analysis method, we propose a new scalability-conscious security design methodology that features: (a) compulsory encryption of highly sensitive data like credit card information, and (b) encryption of data for which encryption does not impair scalability. As a result, the security-scalability tradeoff needs to be considered only over data for which encryption impacts scalability, thus greatly simplifying the task of managing the tradeoff. Amit Manjhi, Anastasia Ailamaki, Bruce M. Maggs, Todd C. Mowry, Christopher Olston, Anthony Tomasic |
SIGMOD Conference | 6 |
| 2005 | Learning to Understand Web Site Update Requests
William W. Cohen, Einat Minkov, Anthony Tomasic |
IJCAI | 3 |
| 2002 | An Introduction to the e-XML Data Integration Suite
Georges Gardarin, Antoine Mensch, Anthony Tomasic |
EDBT | 3 |
| 2002 | Locating and accessing data repositories with WebSemantics
George A. Mihaila, Louiqa Raschid, Anthony Tomasic |
VLDB J. | 3 |
| 1999 | GlOSS: Text-Source Discovery over the InternetabstractThe dramatic growth of the Internet has created a new problem for users: location of the relevant sources of documents. This article presents a framework for (and experimentally analyzes a solution to) this problem, which we call the text-source discovery problem . Our approach consists of two phases. First, each text source exports its contents to a centralized service. Second, users present queries to the service, which returns an ordered list of promising text sources. This article describes GlOSS , Glossary of Servers Server, with two versions: bGlOSS , which provides a Boolean query retrieval model, and vGlOSS , which provides a vector-space retrieval model. We also present hGlOSS , which provides a decentralized version of the system. We extensively describe the methodology for measuring the retrieval effectiveness of these systems and provide experimental evidence, based on actual data, that all three systems are highly effective in determining promising text sources for a given query. Luis Gravano, Hector Garcia-Molina, Anthony Tomasic |
ACM Trans. Database Syst. | 3 |
| 1998 | Equal Time for Data on the Internet with WebSemantics
George A. Mihaila, Louiqa Raschid, Anthony Tomasic |
EDBT | 3 |
| 1998 | Partial Answers for Unavailable Data Sources
Philippe Bonnet, Anthony Tomasic |
FQAS | 2 |
| 1998 | Leveraging Mediator Cost Models with Heterogeneous Data SourcesabstractDistributed systems require declarative access to diverse information sources. One approach to solving this heterogeneous distributed database problem is based on mediator architectures. In these architectures, mediators accept queries from users, process them with respect to wrappers, and return answers. Wrappers provide access to underlying sources. To efficiently process queries, the mediator must optimize the plan used for processing the query. In classical databases, cost-estimate based query optimization is effective. In a heterogeneous distributed databases, cost-estimate based query optimization is difficult to achieve because the underlying data sources do not export cost information. This paper describes a new method that permits the wrapper programmer to export cost estimates. For the wrapper programmer to describe all cost estimates may be impossible due to lack of information or burdensome due to the amount of information. We ease this responsibility of the wrapper programmer by leveraging the generic cost model of the mediator with specific cost estimates from the wrappers. Hubert Naacke, Georges Gardarin, Anthony Tomasic |
ICDE | 3 |
| 1998 | Experiences in Federated Databases: From IRO-DB to MIRO-Web
Peter Fankhauser, Georges Gardarin, José Manuel Muñoz, Anthony Tomasic |
VLDB | 5 |
| 1998 | Dynamic Query Operator Scheduling for Wide-Area Remote Access
Laurent Amsaleg, Michael J. Franklin, Anthony Tomasic |
Distributed Parallel Databases | 3 |
| 1998 | Scaling Access to Heterogeneous Data Sources with DISCOabstractAccessing many data sources aggravates problems for users of heterogeneous distributed databases. Database administrators must deal with fragile mediators, that is, mediators with schemas and views that must be significantly changed to incorporate a new data source. When implementing translators of queries from mediators to data sources, database implementers must deal with data sources that do not support all the functionality required by mediators. Application programmers must deal with graceless failures for unavailable data sources. Queries simply return failure and no further information when data sources are unavailable for query processing. The Distributed Information Search COmponent (Disco) addresses these problems. Data modeling techniques manage the connections to data sources, and sources can be added transparently to the users and applications. The interface between mediators and data sources flexibly handles different query languages and different data source functionality. Query rewriting and optimization techniques rewrite queries so they are efficiently evaluated by sources. Query processing and evaluation semantics are developed to process queries over unavailable data sources. In this article, we describe: 1) the distributed mediator architecture of Disco; 2) the data model and its modeling of data source connections; 3) the interface to underlying data sources and the query rewriting process; and 4) query processing semantics. We describe several advantages of our system. Anthony Tomasic, Louiqa Raschid, Patrick Valduriez |
IEEE Trans. Knowl. Data Eng. | 1 |
| 1997 | The Distributed Information Search Component (Disco) and the World Wide WebabstractThe Distributed Information Search COmponent (DISCO) is a prototype heterogeneous distributed database that accesses underlying data sources. The DISCO prototype currently focuses on three central research problems in the context of these systems. First, since the capabilities of each data source is different, transforming queries into subqueries on data source is difficult. We call this problem the weak data source problem. Second, since each data source performs operations in a generally unique way, the cost for performing an operation may vary radically from one wrapper to another. We call this problem the radical cost problem. Finally, existing systems behave rudely when attempting to access an unavailable data source. We call this problem the ungraceful failure problem. Anthony Tomasic, Rémy Amouroux, Philippe Bonnet, Olga Kapitskaia, Hubert Naacke, Louiqa Raschid |
SIGMOD Conference | 1 |
| 1997 | Data Structures for Efficient Broker ImplementationabstractWith the profusion of text databases on the Internet, it is becoming increasingly hard to find the most useful databases for a given query. To attack this problem, several existing and proposed systems employ brokers to direct user queries, using a local database of summary information about the available databases. This summary information must effectively distinguish relevant databases and must be compact while allowing efficient access. We offer evidence that one broker, GlOSS , can be effective at locating databases of interest even in a system of hundreds of databased and can examine the performance of accessing the GlOSS summeries for two promising storage methods: the grid file and partitioned hashing. We show that both methods can be tuned to provide good performance for a particular workload (within a broad range of workloads), and we discuss the tradeoffs between the two data structures. As a side effect of our work, we show that grid files are more broadly applicable than previously thought; inparticular, we show that by varying the policies used to construct the grid file we can provide good performance for a wide range of workloads even when storing highly skewed data. Anthony Tomasic, Luis Gravano, Calvin Lue, Peter M. Schwarz, Laura M. Haas |
ACM Trans. Inf. Syst. | 1 |
| 1996 | Scaling Heterogeneous Databases and the Design of DiscoabstractAccess to large numbers of data sources introduces new problems for users of heterogeneous distributed databases. End users and application programmers must deal with unavailable data sources. Database administrators must deal with incorporating new sources into the model. Database implementers must deal with the translation of queries between query languages and schemas. The Distributed Information Search COmponent (Disco) addresses these problems. Query processing semantics are developed to process queries over data sources which do not return answers. Data modeling techniques manage connections to data sources. The component interface to data sources flexibly handles different query languages and translates queries. This paper describes (a) the distributed mediator architecture of Disco, (b) its query processing semantics, (C) the data model and its modeling of data source connections, and (d) the interface to underlying data sources. Anthony Tomasic, Louiqa Raschid, Patrick Valduriez |
ICDCS | 1 |
| 1996 | Performance Issues in Distributed Shared-Nothing Information-Retrieval Systems
Anthony Tomasic, Hector Garcia-Molina |
Inf. Process. Manag. | 1 |
| 1994 | Synthetic Workload Performance Analysis of Incremental Updates
Kurt A. Shoens, Anthony Tomasic, Hector Garcia-Molina |
SIGIR | 2 |
| 1994 | The Effectiveness of GlOSS for the Text Database Discovery ProblemabstractThe popularity of on-line document databases has led to a new problem: finding which text databases (out of many candidate choices) are the most relevant to a user. Identifying the relevant databases for a given query is the text database discovery problem. The first part of this paper presents a practical solution based on estimating the result size of a query and a database. The method is termed GlOSS—Glossary of Servers Server. The second part of this paper evaluates the effectiveness of GlOSS based on a trace of real user queries. In addition, we analyze the storage cost of our approach. Luis Gravano, Hector Garcia-Molina, Anthony Tomasic |
SIGMOD Conference | 3 |
| 1994 | Incremental Updates of Inverted Lists for Text Document RetrievalabstractWith the proliferation of the world's “information highways” a renewed interest in efficient document indexing techniques has come about. In this paper, the problem of incremental updates of inverted lists is addressed using a new dual-structure index. The index dynamically separates long and short inverted lists and optimizes retrieval, update, and storage of each type of list. To study the behavior of the index, a space of engineering trade-offs which range from optimizing update time to optimizing query performance is described. We quantitatively explore this space by using actual data and hardware in combination with a simulation of an information retrieval system. We then describe the best algorithm for a variety of criteria. Anthony Tomasic, Hector Garcia-Molina, Kurt A. Shoens |
SIGMOD Conference | 1 |
| 1993 | Caching and Database Scaling in Distributed Shard-Nothing Information Retrieval SystemsabstractA common class of existing information retrieval system provides access to abstracts. For example Stanford University, through its FOLIO system, provides access to the INSPECT database of abstracts of the literature on physics, computer science, electrical engineering, etc. In this paper this database is studied by using a trace-driven simulation. We focus on physical index design, inverted index caching, and database scaling in a distributed shared-nothing system. All three issues are shown to have a strong effect on response time and throughput. Database scaling is explored in two ways. One way assumes an “optimal” configuration for a single host and then linearly scales the database by duplicating the host architecture as needed. The second way determines the optimal number of hosts given a fixed database size. Anthony Tomasic, Hector Garcia-Molina |
SIGMOD Conference | 1 |
| 1993 | Query Processing and Inverted Indices in Shared-Nothing Document Information Retrieval Systems
Anthony Tomasic, Hector Garcia-Molina |
VLDB J. | 1 |
| 1988 | View Update Translation via Deduction and Annotation
Anthony Tomasic |
ICDT | 1 |