Anthony Tomasic

dblp:t/AnthonyTomasic · DBLP profile ↗
← Back
50ranked-venue papers
13as first author
3since 2021 · last 2026
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 27 · 10 first-authorHuman-computer interaction and ubiquitous computing · 13 · 2 first-author · 1 since 2021Artificial intelligence and machine learning · 8 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-authorSystems, architecture and hardware · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Human-computer interaction and pervasive computing
8 papers
Accessibility and assistive technology · 32% Human-AI interaction · 27% Design research and methods · 15%
Databases, data mining, and information retrieval
16 papers
Transaction processing and concurrency control · 46% Query processing and optimization · 23% Information retrieval · 18%
Software engineering, system software, and programming languages
3 papers
Program synthesis and code generation · 82% Compilers and program optimization · 12% Software maintenance and evolution · 7%
Computer architecture, parallel and distributed computing, and storage systems
7 papers
Cloud and datacenter computing · 50% Distributed systems · 30% Storage systems · 11%
Network and information security
3 papers
Privacy and data protection · 37% Web and mobile security · 37% Cryptographic primitives and cryptanalysis · 26%

Topics — the 30 heaviest of 63, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Accessibility and assistive technology › assistive technology for visual impairment
assistive technology for blind and low-vision users
1.012026
Navigation and Interaction for Blind Users via a Cognitive Architecture · AAAI 2026
Human-AI interaction
conversational agents
1.012026
Navigation and Interaction for Blind Users via a Cognitive Architecture · AAAI 2026
Design research and methods
research through design
0.612022
Recentering Reframing as an RtD Contribution: The Case of Pivoting from Accessible Web Tables to a Conversational Internet · CHI 2022
Accessibility and assistive technology
screen reader accessibility
0.612022
Recentering Reframing as an RtD Contribution: The Case of Pivoting from Accessible Web Tables to a Conversational Internet · CHI 2022
Transaction processing and concurrency control
transaction scheduling
0.312018
Performance of OLTP via Intelligent Scheduling · ICDE 2018
Program synthesis and code generation
neural program synthesis
0.312018
Retrieval-Based Neural Code Generation · EMNLP 2018
Program synthesis and code generation › code generation with language models
retrieval-augmented code generation
0.312018
Retrieval-Based Neural Code Generation · EMNLP 2018
Human-AI interaction
mixed-initiative interaction
0.222009
User-created forms as an effective method of human-agent communication · CHI 2009
Vio: a mixed-initiative approach to learning and automating procedural update tasks · CHI 2007
Query processing and optimization
query result caching
0.222008
Scalable query result caching for web applications · Proc. VLDB Endow. 2008
Invalidation Clues for Database Scalability Services · ICDE 2007
Design research and methods › design practice
service design
0.122011
Understanding the space for co-design in riders' interactions with a transit service · CHI 2010
Field trial of Tiramisu: crowd-sourcing bus arrival times to spur co-design · CHI 2011
Ubiquitous computing and smart environments › urban computing
smart city
0.112011
Field trial of Tiramisu: crowd-sourcing bus arrival times to spur co-design · CHI 2011
Query processing and optimization › query rewriting
query transformation
0.122009
Holistic Query Transformations for Dynamic Web Applications · ICDE 2009
The Distributed Information Search Component (Disco) and the World Wide Web · SIGMOD Conference 1997
Collaborative and social computing › civic engagement
civic technology
0.112010
Understanding the space for co-design in riders' interactions with a transit service · CHI 2010
Human-AI interaction
human-agent interaction
0.112009
User-created forms as an effective method of human-agent communication · CHI 2009
Compilers and program optimization
program transformation
0.112009
Holistic Query Transformations for Dynamic Web Applications · ICDE 2009
Natural language and speech › Question answering and dialogue systems › conversational agents
conversational assistant
0.112008
RADAR: A Personal Assistant that Learns to Reduce Email Overload · AAAI 2008
Distributed systems › consistency models
cache consistency
0.112008
Scalable query result caching for web applications · Proc. VLDB Endow. 2008
Natural language and speech › Question answering and dialogue systems
intent detection
0.112007
Learning information intent via observation · WWW 2007
Interaction techniques and input
text entry
0.112007
Vio: a mixed-initiative approach to learning and automating procedural update tasks · CHI 2007
Web and mobile security › phishing detection
phishing email detection
0.112007
Learning to detect phishing emails · WWW 2007
Privacy and data protection
privacy-preserving data management
0.112007
Invalidation Clues for Database Scalability Services · ICDE 2007
Cryptographic primitives and cryptanalysis
encryption
0.112006
Simultaneous scalability and security for data-intensive web applications · SIGMOD Conference 2006
Information retrieval
distributed information retrieval
0.141999
GlOSS: Text-Source Discovery over the Internet · ACM Trans. Database Syst. 1999
Data Structures for Efficient Broker Implementation · ACM Trans. Inf. Syst. 1997
The Effectiveness of GlOSS for the Text Database Discovery Problem · SIGMOD Conference 1994
Software maintenance and evolution
software evolution
0.112005
Learning to Understand Web Site Update Requests · IJCAI 2005
Data integration and cleaning
heterogeneous data source integration
0.021998
Scaling Access to Heterogeneous Data Sources with DISCO · IEEE Trans. Knowl. Data Eng. 1998
Leveraging Mediator Cost Models with Heterogeneous Data Sources · ICDE 1998
Design research and methods
participatory design
0.012011
Field trial of Tiramisu: crowd-sourcing bus arrival times to spur co-design · CHI 2011
Information retrieval
retrieval models
0.021999
GlOSS: Text-Source Discovery over the Internet · ACM Trans. Database Syst. 1999
The Effectiveness of GlOSS for the Text Database Discovery Problem · SIGMOD Conference 1994
Information retrieval › indexing
inverted index
0.031994
Incremental Updates of Inverted Lists for Text Document Retrieval · SIGMOD Conference 1994
Query Processing and Inverted Indices in Shared-Nothing Document Information Retrieval Systems · VLDB J. 1993
Caching and Database Scaling in Distributed Shard-Nothing Information Retrieval Systems · SIGMOD Conference 1993
Query processing and optimization › query optimization
cost-based optimization
0.021998
Leveraging Mediator Cost Models with Heterogeneous Data Sources · ICDE 1998
The Distributed Information Search Component (Disco) and the World Wide Web · SIGMOD Conference 1997
Distributed and cloud data management
distributed query processing
0.021998
The Distributed Information Search Component (Disco) and the World Wide Web · SIGMOD Conference 1997
Leveraging Mediator Cost Models with Heterogeneous Data Sources · ICDE 1998

Methods — techniques the papers use, named apart from their topics

object detection · 1.0large language model · 1.0augmented reality · 1.0reflective case study · 0.6design experiment · 0.6transaction representation · 0.3subtree retrieval · 0.3neural encoder-decoder · 0.3intelligent scheduling · 0.3dynamic-programming sentence similarity · 0.3source-to-source compilation · 0.3latency hiding · 0.3machine learning · 0.2invalidation clues · 0.2encryption · 0.2longitudinal study · 0.2field experiment · 0.2publish-subscribe · 0.2
YearPublicationVenuePosition
2026 Navigation and Interaction for Blind Users via a Cognitive Architecture
abstract
Navigating new indoor spaces and interacting with the environment presents many challenges for people who are blind or have low vision (BLV). To address these challenges, we prototyped a smartphone-based conversational assistant that helps BLV people navigate and interact with their environment. The prototype utilizes a cognitive architecture to integrate three different technologies: (i) augmented-reality spatial anchors for high-precision localization and access to static information about the environment; (ii) real-time object/people detection for information about the environment and obstacle avoidance; and (iii) a conversational agent}that uses large language models (LLMs) for information extraction, conversational interaction, and turn-by-turn navigation. We assess the impact of different technologies on human performance by measuring user task time and errors. We found that conversational interaction holistically integrates the different technologies to deliver a better user experience while significantly reducing task completion time.
Oscar J. Romero, Anthony Tomasic, Elizabeth J. Carter, John Zimmerman, Aaron Steinfeld
AAAI2
2022 Recentering Reframing as an RtD Contribution: The Case of Pivoting from Accessible Web Tables to a Conversational Internet
abstract
Design produces valuable knowledge by offering new perspectives that reframe problematic situations. Research through Design (RtD) contributes new frames along with design work demonstrating a frame's value. Interestingly, RtD papers rarely describe how reframing happens. This gap in documentation unintentionally implies a romantic account of design, it implies that the first step of an RtD project is to have a brilliant idea. This is especially problematic in cases where the reframing causes a pivot that leads to a new research program. To help address this gap, we describe a case where through a series of three design experiments we experienced a research pivot. We describe how our work to improve web-table navigation for screen-reader users broke our frame. The break led to a new research program focused on constructing a conversational internet. This paper offers our case along with reflection on reporting design work that drives reframing.
John Zimmerman, Aaron Steinfeld, Anthony Tomasic, Oscar J. Romero
CHI3
2021 A Task-Oriented Dialogue Architecture via Transformer Neural Language Models and Symbolic Injection
abstract
Recently, transformer language models have been applied to build both task-and non-taskoriented dialogue systems.Although transformers perform well on most of the NLP tasks, they perform poorly on context retrieval and symbolic reasoning.Our work aims to address this limitation by embedding the model in an operational loop that blends both natural language generation and symbolic injection.We evaluated our system on the multi-domain DSTC8 data set and reported joint goal accuracy of 75.8% (ranked among the first half positions), intent accuracy of 97.4% (which is higher than the reported literature), and a 15% improvement for success rate compared to a baseline with no symbolic injection.These promising results suggest that transformer language models can not only generate proper system responses but also symbolic representations that can further be used to enhance the overall quality of the dialogue management as well as serving as scaffolding for complex conversational reasoning.
Oscar J. Romero, Antian Wang, John Zimmerman, Aaron Steinfeld, Anthony Tomasic
SIGDIAL5
2020 A Long-Term Evaluation of Adaptive Interface Design for Mobile Transit Information
abstract
Personalization of user experience has a long history of success in the HCI community. More recently the community has focused on adaptive user interfaces, supported by machine learning, that reduce interaction efforts and improves user experience by collapsing transactions and pre-filtering results. However, generally, these more recent results have only been demonstrated in the laboratory environment. In this paper, we share the case of a deployed mobile transit app that adapts based on users’ previous usage. We examine the impact of adaptation, both good and bad, and user abandonment rates. We conducted an 18-month assessment where 2,616 participants (with and without vision impairments) were recruited and participated in an A/B study. Finally, we draw some insights on some unusual effects that appear over the long term.
Oscar J. Romero, Alexander Haig, Lynn Kirabo, Qian Yang 0004, John Zimmerman, Anthony Tomasic, Aaron Steinfeld
MobileHCI6
2018 Contingent Responsiveness in Digital Storybooks: Effects on Children's Comprehension and the Role of Individual Differences in Attention
Cassondra M. Eng, Anthony Tomasic, Erik D. Thiessen
CogSci2
2018 Retrieval-Based Neural Code Generation
abstract
In models to generate program source code from natural language, representing this code in a tree structure has been a common approach.However, existing methods often fail to generate complex code correctly due to a lack of ability to memorize large and complex structures.We introduce RECODE, a method based on subtree retrieval that makes it possible to explicitly reference existing code examples within a neural code generation model.First, we retrieve sentences that are similar to input sentences using a dynamicprogramming-based sentence similarity scoring method.Next, we extract n-grams of action sequences that build the associated abstract syntax tree.Finally, we increase the probability of actions that cause the retrieved n-gram action subtree to be in the predicted code.We show that our approach improves the performance on two code generation tasks by up to +2.6 BLEU. 1
Shirley Anugrah Hayati, Raphaël Olivier, Pravalika Avvaru, Anthony Tomasic, Graham Neubig
EMNLP5
2018 Performance of OLTP via Intelligent Scheduling
abstract
Current architectures for main-memory online transaction processing (OLTP) database management systems (DBMS) typically use random scheduling to assign transactions to threads. This approach achieves uniform load across threads but it ignores the likelihood of conflicts between transactions. If the DBMS could estimate the potential for transaction conflict and then intelligently schedule transactions to avoid conflicts, then the system could improve its performance. Such estimation of transaction conflict, however, is non-trivial for several reasons. First, conflicts occur under complex conditions that are far removed in time from the scheduling decision. Second, transactions must be represented in a compact and efficient manner to allow for fast conflict detection. Third, given some evidence of potential conflict, the DBMS must schedule transactions in such a way that minimizes this conflict. In this paper, we systematically explore the design decisions for solving these problems. We then empirically measure the performance impact of different representations on a standard OLTP benchmark.
Tieying Zhang, Anthony Tomasic, Yangjun Sheng, Andrew Pavlo
ICDE2
2017 Self-Driving Database Management Systems
Andrew Pavlo, Gustavo Angulo, Joy Arulraj, Haibin Lin, Jiexi Lin, Lin Ma 0006, Prashanth Menon, Todd C. Mowry, Matthew Perron, Ian Quah, Siddharth Santurkar, Anthony Tomasic, Skye Toor, Dana Van Aken, Ziqi Wang 0007, Yingjun Wu, Ran Xian, Tieying Zhang
CIDR12
2016 Planning Adaptive Mobile Experiences When Wireframing
abstract
Machine learning improves mobile user experience. Interestingly, envisioning apps with adaptive interfaces that reduce navigation and selection effort is not standard UX practice. When implementing an adaptive UI for our mobile transit app, we encountered a number of problems. Our original design did not log necessary information nor did it induce users to provide good labels. On reflection, we realized UX designers should identify and refine UI adaptions when sketching wireframes. To advance on this insight, we reviewed the interfaces of popular apps and extracted six design patterns where UI adaptation can improve in-app navigation. Next, we designed an exemplar set of wireframes, illustrating how UX designers might annotate their interaction flows to communicate planned adaptation and note the information (logs and labels) needed to make the desired inferences.
Qian Yang 0004, John Zimmerman, Aaron Steinfeld, Anthony Tomasic
Conference on Designing Interactive Systems4
2016 The utility of tables for screen reader users
abstract
The mere existence of digital documents on the web makes those documents more accessible to blind people than if they were not digitized. The web provides several opportunities, however, for web authors to describe what content looks like, rather than what it means, and this visual description inconveniences and confuses people using screen readers. We focus on the particular instance of this problem where information that is logically tabular is presented using repeated visual templates. We recruited a group of experienced screen reader users and observed them performing a series of tasks designed to elicit tabular behavior. We find that people who make use of tabular operations perform the tasks faster and more successfully than those who do not. We furthermore find that while many users discover the tables we introduced as part of the study, many others do not or do not choose to make use of the tables. We discuss some techniques for enhancing the discoverability of the introduced tables.
Steven Gardiner, Anthony Tomasic, John Zimmerman
CCNC2
2016 Combining contribution interactions to increase coverage in mobile participatory sensing systems
abstract
Participatory sensing systems use people and their smartphones as a sensing infrastructure, and getting people to make contributions remains a critical challenge. Little work details how system designers should combine different interactions to increase coverage of service location. Tiramisu, a participatory sensing system, invites transit riders to crowdsource real-time arrival information by sharing location traces when they commute. We extended this system with a new feature that allows riders at stops to "spot" buses passing by. To better understand the impact of this new feature, we conducted an observational log analysis, examining changes in coverage and user behavior before and after the new feature. Following the addition of the spotting feature, participants' contributions increased coverage (the number of trips with real-time data) by 98%, and they used the app more than twice as much. The addition of the spotting feature was also followed by a significant increase of trace contributions.
Yun Huang 0003, John Zimmerman, Anthony Tomasic, Aaron Steinfeld
MobileHCI3
2015 EnTable: Rewriting Web Data Sets as Accessible Tables
abstract
Today, many data-driven web pages present information in a way that is difficult for blind and low vision users to navigate and to understand. EnTable addresses this challenge. It re-writes confusing and complicated template-based data sets as accessible tables. EnTable allows blind and low vision users to submit requests for pages they wish to access. The system then employs sighted informants to markup the desired page with semantic information, allowing the page to be re-written using straightforward tags. Screen reader users who browse the web using the EnTable browser extension can report data sets that are confusing, and utilize data sets re-written with the tag based on their own requests or on the requests of other users.
Steven Gardiner, Anthony Tomasic, John Zimmerman
ASSETS2
2014 Motivating contribution in a participatory sensing system via quid-pro-quo
abstract
Participatory sensing systems (PSS) require frequent injection of information that has a short shelf-life. The use of crowds to gather information for PSS is therefore particularly challenging. In this study, we explore the impact of two policies on user contributions. A quid-pro-quo policy exchanges contributions from users for access to critical information in the system. A request policy simply reminds the user that information is needed to make the system function well. Prior research has shown that request for help in crowdsourced system is an effective mechanism to increase contributions. During a large-scale experimental study within a publicly deployed, crowdsourced, transit information system, we analyzed metrics associated with frequency of contribution and commitment to long-term use over a 10-month period. Our results confirmed that quid-pro-quo led to more contribution, but at a cost of faster departure from the study. When a participant was simply requested to contribute, but could still access community-generated data if they ignored a request, was largely ineffective and was statistically similar to the control condition where no request for contribution occurred. Thus crowdsource system designers should consider imposing quid-pro-quo type policies for PSS that concentrate on fewer users, but makes them more productive.
Anthony Tomasic, John Zimmerman, Aaron Steinfeld, Yun Huang 0003
CSCW1
2014 Citizen Motivation on the Go: The Role of Psychological Empowerment
abstract
Although advances in technology now enable people to communicate ‘anytime, anyplace’, it is not clear how citizens can be motivated to actually do so. This paper evaluates the impact of three principles of psychological empowerment, namely perceived self-efficacy, sense of community and causal importance, on public transport passengers’ motivation to report issues and complaints while on the move. A week-long study with 65 participants revealed that self-efficacy and causal importance increased participation in short bursts and increased perceptions of service quality over longer periods. Finally, we discuss the implications of these findings for citizen participation projects and reflect on design opportunities for mobile technologies that motivate citizen participation.
Jorge Gonçalves 0001, Vassilis Kostakos, Evangelos Karapanos, Mary Barreto, Tiago Camacho, Anthony Tomasic, John Zimmerman
Interact. Comput.6
2013 Energy efficient and accuracy aware (E2A2) location services via crowdsourcing
abstract
Many mobile applications rely on location information gained from location services on mobile devices. However, continuously tracking the device location with high accuracy drains the battery quickly. Furthermore, sensing the same location can be redundant when multiple devices are co-located. In this paper, we develop a crowdsourcing-based location service, E2A2 (energy efficient and accuracy aware), which places colo-cated devices into groups, and uses group location to represent individual device location. The E2A2 location service aims to reduce individual device battery consumption associated with location services while simultaneously maintaining high location accuracy for each device. Our experimental results from a prototype system show the effectiveness of our proposed solution with different mobility patterns. We also present results on the impact of different system parameters and the number of users in a group. Compared to running GPS location services on individual devices separately, our E2A2 service saves on average 33% battery consumption rate when 4 devices are co-located at walking speed and 26% battery consumption rate when 4 devices are colocated on the same bus while meeting the same accuracy requirements.
Yun Huang 0003, Anthony Tomasic, Yufei An, Charles Garrod, Aaron Steinfeld
WiMob2
2011 Field trial of Tiramisu: crowd-sourcing bus arrival times to spur co-design
abstract
Crowd-sourcing social computing systems represent a new material for HCI designers. However, these systems are difficult to work with and to prototype, because they require a critical mass of participants to investigate social behavior. Service design is an emerging research area that focuses on how customers co-produce the services that they use, and thus it appears to be a great domain to apply this new material. To investigate this relationship, we developed Tiramisu, a transit information system where commuters share GPS traces and submit problem reports. Tiramisu processes incoming traces and generates real-time arrival time predictions for buses. We conducted a field trial with 28 participants. In this paper we report on the results and reflect on the use of field trials to evaluate crowd-sourcing prototypes and on how crowd sourcing can generate co-production between citizens and public services.
John Zimmerman, Anthony Tomasic, Charles Garrod, Daisy Yoo, Chaya Hiruncharoenvate, Rafae Aziz, Nikhil Ravi Thiruvengadam, Yun Huang 0003, Aaron Steinfeld
CHI2
2011 Mixer: Mixed-Initiative Data Retrieval and Integration by Example
Steven Gardiner, Anthony Tomasic, John Zimmerman, Rafae Aziz, Kathryn Rivard
INTERACT (1)2
2010 Understanding the space for co-design in riders' interactions with a transit service
abstract
The recent advances in web 2.0 technologies and the rapid adoption of smart phones raises many opportunities for public services to improve their services by engaging their users (who are also owners of the service) in co-design: a dialog where users help design the services they use. To investigate this opportunity, we began a service design project investigating how to create repeated information exchanges between riders and a transit agency in order to create a virtual "place" from which the dialog on services could take place. Through interviews with riders, a workshop with a transit agency, and speed dating of design concepts, we have developed a design direction. Specifically, we propose a service that combines vehicle location and "fullness" ratings provided by riders with dynamic route change information from the transit agency as a foundation for a dialog around riders conveying input for continuous service improvement.
Daisy Yoo, John Zimmerman, Aaron Steinfeld, Anthony Tomasic
CHI4
2009 User-created forms as an effective method of human-agent communication
abstract
A key challenge for mixed-initiative systems is to create a shared understanding of the task between human and agent. To address this challenge, we created a mixed-initiative interface called Mixer to aid administrators with automating tedious information-retrieval tasks. Users initiate communication with the agent by constructing a form, creating a structure to hold the information they require and to show context in order to interpret this information. They then populate the form with the desired results, demonstrating to the agent the steps required to retrieve the information. This method of form creation explicitly defines the shared understanding between human and agent. An evaluation of the interface shows that administrators can effectively create forms to communicate with the agent, that they are likely to accept this technology in their work environment, and that the agent's help can significantly reduce the time they spend on repeated information-retrieval tasks.
John Zimmerman, Kathryn Rivard, Ian Hargraves, Anthony Tomasic, Ken Mohnkern
CHI4
2009 Holistic Query Transformations for Dynamic Web Applications
abstract
A promising approach to scaling Web applications is to distribute the server infrastructure on which they run. This approach, unfortunately, can introduce latency between the application and database servers, which in turn increases the network latency of Web interactions for the clients (end users). In this paper we introduce the concept of source-to-source holistic transformations - transformations that seek to optimize both the application code and the database requests made by it, to reduce client latency. As examples of our concept, we propose and evaluate two source-to-source holistic transformations that focus on hiding the latencies of database queries. We argue that opportunities for applying these transformations will continue to exist in Web applications. We then present algorithms for automating these transformations in a source-to-source compiler. Finally, we evaluate the effect of these two transformations on three realistic Web benchmark applications, both in the traditional centralized setting and a distributed setting.
Amit Manjhi, Charles Garrod, Bruce M. Maggs, Todd C. Mowry, Anthony Tomasic
ICDE5
2008 RADAR: A Personal Assistant that Learns to Reduce Email Overload
Michael Freed, Jaime G. Carbonell, Geoffrey J. Gordon, Jordan Hayes, Brad A. Myers, Daniel P. Siewiorek, Stephen F. Smith, Aaron Steinfeld, Anthony Tomasic
AAAI9
2008 Scalable query result caching for web applications
abstract
The backend database system is often the performance bottleneck when running web applications. A common approach to scale the database component is query result caching, but it faces the challenge of maintaining a high cache hit rate while efficiently ensuring cache consistency as the database is updated. In this paper we introduce Ferdinand, the first proxy-based cooperative query result cache with fully distributed consistency management. To maintain a high cache hit rate, Ferdinand uses both a local query result cache on each proxy server and a distributed cache. Consistency management is implemented with a highly scalable publish/subscribe system. We implement a fully functioning Ferdinand prototype and evaluate its performance compared to several alternative query-caching approaches, showing that our high cache hit rate and consistency management are both critical for Ferdinand's performance gains over existing systems.
Charles Garrod, Amit Manjhi, Anastasia Ailamaki, Bruce M. Maggs, Todd C. Mowry, Christopher Olston, Anthony Tomasic
Proc. VLDB Endow.7
2007 Vio: a mixed-initiative approach to learning and automating procedural update tasks
abstract
Today many workers spend too much of their time translating their co-workers' requests into structures that information systems can understand. This paper presents the novel interaction design and evaluation of VIO, an agent that helps workers trans late request. VIO monitors requests and makes suggestions to speed up the translation. VIO allows users to quickly correct agent errors. These corrections are used to improve agent performance as it learns to automate work. Our evaluations demonstrate that this type of agent can significantly reduce task completion time, freeing workers from mundane tasks.
John Zimmerman, Anthony Tomasic, Isaac Simmons, Ian Hargraves, Ken Mohnkern, Jason Cornwell, Robert Martin McGuire
CHI2
2007 Invalidation Clues for Database Scalability Services
abstract
For their scalability needs, data-intensive Web applications can use a database scalability service (DBSS), which caches applications' query results and answers queries on their behalf. One way for applications to address their security/privacy concerns when using a DBSS is to encrypt all data that passes through the DBSS. Doing so, however, causes the DBSS to invalidate large regions of its cache when data updates occur. To invalidate more precisely, the DBSS needs help in order to know which results to invalidate; such help inevitably reveals some properties about the data. In this paper, we present invalidation clues, a general technique that enables applications to reveal little data to the DBSS, yet limit the number of unnecessary invalidations. Compared with previous approaches, invalidation clues provide applications significantly improved tradeoffs between security/privacy and scalability. Our experiments using three Web application benchmarks, on a prototype DBSS we have built, confirm that invalidation clues are indeed a low-overhead, effective, and general technique for applications to balance their privacy and scalability needs.
Amit Manjhi, Phillip B. Gibbons, Anastasia Ailamaki, Charles Garrod, Bruce M. Maggs, Todd C. Mowry, Christopher Olston, Anthony Tomasic
ICDE8
2007 Learning to detect phishing emails
abstract
Each month, more attacks are launched with the aim of making web users believe that they are communicating with a trusted entity for the purpose of stealing account information, logon credentials, and identity information in general. This attack method, commonly known as "phishing," is most commonly initiated by sending out emails with links to spoofed websites that harvest information. We present a method for detecting these attacks, which in its most general form is an application of machine learning on a feature set designed to highlight user-targeted deception in electronic communication. This method is applicable, with slight modification, to detection of phishing websites, or the emails used to direct victims to these sites. We evaluate this method on a set of approximately 860 such phishing emails, and 6950 non-phishing emails, and correctly identify over 96% of the phishing emails while only mis-classifying on the order of 0.1% of the legitimate emails. We conclude with thoughts on the future for such techniques to specifically identify deception, specifically with respect to the evolutionary nature of the attacks and information available.
Ian Fette, Norman M. Sadeh, Anthony Tomasic
WWW3
2007 Learning information intent via observation
abstract
Users in an organization frequently request help by sending request messages to assistants that express information intent: an intention to update data in an information system. Human assistants spend a significant amount of time and effort processing these requests. For example, human resource assistants process requests to update personnel records, and executive assistants process requests to schedule conference rooms or to make travel reservations. To process the intent of a request, assistants read the request and then locate, complete, and submit a form that corresponds to the expressed intent. Automatically or semi-automatically processing the intent expressed in a request on behalf of an assistant would ease the mundane and repetitive nature of this kind of work.For a well-understood domain, a straightforward application of natural language processing techniques can be used to build an intelligent form interface to semi-automatically process information intent request messages. However, high performance parsers are based on machine learning algorithms that require a large corpus of examples that have been labeled by an expert. The generation of a labeled corpus of requests is a major barrier to the construction of a parser. In this paper, we investigate the construction of a natural language processing system and an intelligent form system that observes an assistant processing requests. The intelligent form system then generates a labeled training corpus by interpreting the observations. This paper reports on the measurement of the performance of the machine learning algorithms based on real data. The combination of observations, machine learning and interaction design produces an effective intelligent form interface based on natural language processing.
Anthony Tomasic, Isaac Simmons, John Zimmerman
WWW1
2006 Processing information intent via weak labeling
abstract
No abstract available.
Anthony Tomasic, Isaac Simmons, John Zimmerman
CIKM1
2006 Linking messages and form requests
abstract
Large organizations with sophisticated infrastructures have large form-based systems that manage the interaction between the user community and the infrastructure. In many cases, when a user needs to complete a form to accomplish a task, the user e-mails a description of the task to the appropriate form expert. In many cases this description is incomplete and the expert engages in a clarification dialog to determine the details of the task. Since many tasks and descriptions are routine, this e-mail dialog can be replaced with an intelligent user interface. The interface proactively reads e-mail (or IM) messages and assists the user in completing the associated task without involving the expert. To ground our vision in a specific application, we have built an agent that functions as a webmaster assistant. For example, a user emails the request: "Change John Doe's home phone number to 800-555-1212" to the agent. The webmaster agent then replies with the biographical data form displaying information about John Doe with the new phone number pre-filled in the form. The user then simply approves the change.In this paper we describe a prototype website maintenance agent that (i) allows users to express the updates they want to make in human terms (free text input expression of intent), and (ii) allows users to quickly repair any inference errors the agent makes. In addition, we present the results of a proof of concept study that details how interacting with a webmaster agent that makes inference errors is both more efficient (faster) and more effective (errors made to site) than sending a request to a human webmaster. We conclude the paper with a discussion of the application of our work to any form-based system.
Anthony Tomasic, John Zimmerman, Isaac Simmons
IUI1
2006 NER Systems that Suit User's Preferences: Adjusting the Recall-Precision Trade-off for Entity Extraction
Einat Minkov, Richard C. Wang, Anthony Tomasic, William W. Cohen
HLT-NAACL3
2006 Simultaneous scalability and security for data-intensive web applications
abstract
For Web applications in which the database component is the bottleneck, scalability can be provided by a third-party Database Scalability Service Provider (DSSP) that caches application data and supplies query answers on behalf of the application. Cost-effective DSSPs will need to cache data from many applications, inevitably raising concerns about security. However, if all data passing through a DSSP is encrypted to enhance security, then data updates trigger invalidation of large regions of cache. Consequently, achieving good scalability becomes virtually impossible. There is a tradeoff between security and scalability, which requires careful consideration.In this paper we study the security-scalability tradeoff, both formally and empirically. We begin by providing a method for statically identifying segments of the database that can be encrypted without impacting scalability. Experiments over a prototype DSSP system show the effectiveness of our static analysis method--for all three realistic bench-mark applications that we study, our method enables a significant fraction of the database to be encrypted without impacting scalability. Moreover, most of the data that can be encrypted without impacting scalability is of the type that application designers will want to encrypt, all other things being equal. Based on our static analysis method, we propose a new scalability-conscious security design methodology that features: (a) compulsory encryption of highly sensitive data like credit card information, and (b) encryption of data for which encryption does not impair scalability. As a result, the security-scalability tradeoff needs to be considered only over data for which encryption impacts scalability, thus greatly simplifying the task of managing the tradeoff.
Amit Manjhi, Anastasia Ailamaki, Bruce M. Maggs, Todd C. Mowry, Christopher Olston, Anthony Tomasic
SIGMOD Conference6
2005 Learning to Understand Web Site Update Requests
William W. Cohen, Einat Minkov, Anthony Tomasic
IJCAI3
2002 An Introduction to the e-XML Data Integration Suite
Georges Gardarin, Antoine Mensch, Anthony Tomasic
EDBT3
2002 Locating and accessing data repositories with WebSemantics
George A. Mihaila, Louiqa Raschid, Anthony Tomasic
VLDB J.3
1999 GlOSS: Text-Source Discovery over the Internet
abstract
The dramatic growth of the Internet has created a new problem for users: location of the relevant sources of documents. This article presents a framework for (and experimentally analyzes a solution to) this problem, which we call the text-source discovery problem . Our approach consists of two phases. First, each text source exports its contents to a centralized service. Second, users present queries to the service, which returns an ordered list of promising text sources. This article describes GlOSS , Glossary of Servers Server, with two versions: bGlOSS , which provides a Boolean query retrieval model, and vGlOSS , which provides a vector-space retrieval model. We also present hGlOSS , which provides a decentralized version of the system. We extensively describe the methodology for measuring the retrieval effectiveness of these systems and provide experimental evidence, based on actual data, that all three systems are highly effective in determining promising text sources for a given query.
Luis Gravano, Hector Garcia-Molina, Anthony Tomasic
ACM Trans. Database Syst.3
1998 Equal Time for Data on the Internet with WebSemantics
George A. Mihaila, Louiqa Raschid, Anthony Tomasic
EDBT3
1998 Partial Answers for Unavailable Data Sources
Philippe Bonnet, Anthony Tomasic
FQAS2
1998 Leveraging Mediator Cost Models with Heterogeneous Data Sources
abstract
Distributed systems require declarative access to diverse information sources. One approach to solving this heterogeneous distributed database problem is based on mediator architectures. In these architectures, mediators accept queries from users, process them with respect to wrappers, and return answers. Wrappers provide access to underlying sources. To efficiently process queries, the mediator must optimize the plan used for processing the query. In classical databases, cost-estimate based query optimization is effective. In a heterogeneous distributed databases, cost-estimate based query optimization is difficult to achieve because the underlying data sources do not export cost information. This paper describes a new method that permits the wrapper programmer to export cost estimates. For the wrapper programmer to describe all cost estimates may be impossible due to lack of information or burdensome due to the amount of information. We ease this responsibility of the wrapper programmer by leveraging the generic cost model of the mediator with specific cost estimates from the wrappers.
Hubert Naacke, Georges Gardarin, Anthony Tomasic
ICDE3
1998 Experiences in Federated Databases: From IRO-DB to MIRO-Web
Peter Fankhauser, Georges Gardarin, José Manuel Muñoz, Anthony Tomasic
VLDB5
1998 Dynamic Query Operator Scheduling for Wide-Area Remote Access
Laurent Amsaleg, Michael J. Franklin, Anthony Tomasic
Distributed Parallel Databases3
1998 Scaling Access to Heterogeneous Data Sources with DISCO
abstract
Accessing many data sources aggravates problems for users of heterogeneous distributed databases. Database administrators must deal with fragile mediators, that is, mediators with schemas and views that must be significantly changed to incorporate a new data source. When implementing translators of queries from mediators to data sources, database implementers must deal with data sources that do not support all the functionality required by mediators. Application programmers must deal with graceless failures for unavailable data sources. Queries simply return failure and no further information when data sources are unavailable for query processing. The Distributed Information Search COmponent (Disco) addresses these problems. Data modeling techniques manage the connections to data sources, and sources can be added transparently to the users and applications. The interface between mediators and data sources flexibly handles different query languages and different data source functionality. Query rewriting and optimization techniques rewrite queries so they are efficiently evaluated by sources. Query processing and evaluation semantics are developed to process queries over unavailable data sources. In this article, we describe: 1) the distributed mediator architecture of Disco; 2) the data model and its modeling of data source connections; 3) the interface to underlying data sources and the query rewriting process; and 4) query processing semantics. We describe several advantages of our system.
Anthony Tomasic, Louiqa Raschid, Patrick Valduriez
IEEE Trans. Knowl. Data Eng.1
1997 The Distributed Information Search Component (Disco) and the World Wide Web
abstract
The Distributed Information Search COmponent (DISCO) is a prototype heterogeneous distributed database that accesses underlying data sources. The DISCO prototype currently focuses on three central research problems in the context of these systems. First, since the capabilities of each data source is different, transforming queries into subqueries on data source is difficult. We call this problem the weak data source problem. Second, since each data source performs operations in a generally unique way, the cost for performing an operation may vary radically from one wrapper to another. We call this problem the radical cost problem. Finally, existing systems behave rudely when attempting to access an unavailable data source. We call this problem the ungraceful failure problem.
Anthony Tomasic, Rémy Amouroux, Philippe Bonnet, Olga Kapitskaia, Hubert Naacke, Louiqa Raschid
SIGMOD Conference1
1997 Data Structures for Efficient Broker Implementation
abstract
With the profusion of text databases on the Internet, it is becoming increasingly hard to find the most useful databases for a given query. To attack this problem, several existing and proposed systems employ brokers to direct user queries, using a local database of summary information about the available databases. This summary information must effectively distinguish relevant databases and must be compact while allowing efficient access. We offer evidence that one broker, GlOSS , can be effective at locating databases of interest even in a system of hundreds of databased and can examine the performance of accessing the GlOSS summeries for two promising storage methods: the grid file and partitioned hashing. We show that both methods can be tuned to provide good performance for a particular workload (within a broad range of workloads), and we discuss the tradeoffs between the two data structures. As a side effect of our work, we show that grid files are more broadly applicable than previously thought; inparticular, we show that by varying the policies used to construct the grid file we can provide good performance for a wide range of workloads even when storing highly skewed data.
Anthony Tomasic, Luis Gravano, Calvin Lue, Peter M. Schwarz, Laura M. Haas
ACM Trans. Inf. Syst.1
1996 Scaling Heterogeneous Databases and the Design of Disco
abstract
Access to large numbers of data sources introduces new problems for users of heterogeneous distributed databases. End users and application programmers must deal with unavailable data sources. Database administrators must deal with incorporating new sources into the model. Database implementers must deal with the translation of queries between query languages and schemas. The Distributed Information Search COmponent (Disco) addresses these problems. Query processing semantics are developed to process queries over data sources which do not return answers. Data modeling techniques manage connections to data sources. The component interface to data sources flexibly handles different query languages and translates queries. This paper describes (a) the distributed mediator architecture of Disco, (b) its query processing semantics, (C) the data model and its modeling of data source connections, and (d) the interface to underlying data sources.
Anthony Tomasic, Louiqa Raschid, Patrick Valduriez
ICDCS1
1996 Performance Issues in Distributed Shared-Nothing Information-Retrieval Systems
Anthony Tomasic, Hector Garcia-Molina
Inf. Process. Manag.1
1994 Synthetic Workload Performance Analysis of Incremental Updates
Kurt A. Shoens, Anthony Tomasic, Hector Garcia-Molina
SIGIR2
1994 The Effectiveness of GlOSS for the Text Database Discovery Problem
abstract
The popularity of on-line document databases has led to a new problem: finding which text databases (out of many candidate choices) are the most relevant to a user. Identifying the relevant databases for a given query is the text database discovery problem. The first part of this paper presents a practical solution based on estimating the result size of a query and a database. The method is termed GlOSS—Glossary of Servers Server. The second part of this paper evaluates the effectiveness of GlOSS based on a trace of real user queries. In addition, we analyze the storage cost of our approach.
Luis Gravano, Hector Garcia-Molina, Anthony Tomasic
SIGMOD Conference3
1994 Incremental Updates of Inverted Lists for Text Document Retrieval
abstract
With the proliferation of the world's “information highways” a renewed interest in efficient document indexing techniques has come about. In this paper, the problem of incremental updates of inverted lists is addressed using a new dual-structure index. The index dynamically separates long and short inverted lists and optimizes retrieval, update, and storage of each type of list. To study the behavior of the index, a space of engineering trade-offs which range from optimizing update time to optimizing query performance is described. We quantitatively explore this space by using actual data and hardware in combination with a simulation of an information retrieval system. We then describe the best algorithm for a variety of criteria.
Anthony Tomasic, Hector Garcia-Molina, Kurt A. Shoens
SIGMOD Conference1
1993 Caching and Database Scaling in Distributed Shard-Nothing Information Retrieval Systems
abstract
A common class of existing information retrieval system provides access to abstracts. For example Stanford University, through its FOLIO system, provides access to the INSPECT database of abstracts of the literature on physics, computer science, electrical engineering, etc. In this paper this database is studied by using a trace-driven simulation. We focus on physical index design, inverted index caching, and database scaling in a distributed shared-nothing system. All three issues are shown to have a strong effect on response time and throughput. Database scaling is explored in two ways. One way assumes an “optimal” configuration for a single host and then linearly scales the database by duplicating the host architecture as needed. The second way determines the optimal number of hosts given a fixed database size.
Anthony Tomasic, Hector Garcia-Molina
SIGMOD Conference1
1993 Query Processing and Inverted Indices in Shared-Nothing Document Information Retrieval Systems
Anthony Tomasic, Hector Garcia-Molina
VLDB J.1
1988 View Update Translation via Deduction and Annotation
Anthony Tomasic
ICDT1