Doreen Cheng

dblp:91/2230 · DBLP profile ↗
← Back
18ranked-venue papers
4as first author
0since 2021 · last 2016
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 5Systems, architecture and hardware · 5 · 4 first-authorHuman-computer interaction and ubiquitous computing · 4Databases, data management, data science and information retrieval · 2Software engineering, systems software and programming languages · 1Graphics, computer vision, multimedia, augmented reality and games · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
2 papers
Data integration and cleaning · 29% Data mining · 29% Machine learning and data management · 29%
Computer architecture, parallel and distributed computing, and storage systems
5 papers
Parallel and multicore computing · 34% High-performance computing · 29% Cloud and datacenter computing · 15%
Artificial intelligence
1 paper
Information extraction and text analysis · 100%

Topics — the 18 heaviest of 23, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning and data management
active learning
0.212014
Active Learning with Efficient Feature Weighting Methods for Improving Data Quality and Classification Accuracy · ACL (1) 2014
Data integration and cleaning › data quality
annotation quality
0.212014
Active Learning with Efficient Feature Weighting Methods for Improving Data Quality and Classification Accuracy · ACL (1) 2014
Data mining › predictive modeling
classification
0.212014
Active Learning with Efficient Feature Weighting Methods for Improving Data Quality and Classification Accuracy · ACL (1) 2014
Debugging and program repair › concurrent program debugging
parallel program debugging
0.011994
A portable debugger for parallel and distributed programs · SC 1994
Cloud and datacenter computing › cluster resource management and scheduling
cluster resource management
0.011994
NAS experiences with a prototype cluster of workstations · SC 1994
High-performance computing › cluster computing
network of workstations
0.011994
NAS experiences with a prototype cluster of workstations · SC 1994
Parallel and multicore computing
parallel programming models
0.021994
An evaluation of automatic and interactive parallel programming tools · SC 1991
A portable debugger for parallel and distributed programs · SC 1994
Distributed systems › distributed system architecture
heterogeneous distributed systems
0.011993
Heterogeneous distributed computing (Mini symposium) · SC 1993
Parallel and multicore computing
parallel programming environment
0.011991
An evaluation of automatic and interactive parallel programming tools · SC 1991
Debugging and program repair
fault localization
0.011989
HDB-a high level debugging · SC 1989
High-performance computing
scientific computing systems
0.021994
NAS experiences with a prototype cluster of workstations · SC 1994
HDB-a high level debugging · SC 1989
Performance modeling and evaluation
benchmarking
0.011994
NAS experiences with a prototype cluster of workstations · SC 1994
High-performance computing › scientific computing systems
computational fluid dynamics
0.011994
NAS experiences with a prototype cluster of workstations · SC 1994
Parallel and multicore computing › parallel programming models
message passing
0.011994
A portable debugger for parallel and distributed programs · SC 1994
Performance modeling and evaluation › benchmarking › parallel benchmark suites
NAS parallel benchmarks
0.011994
NAS experiences with a prototype cluster of workstations · SC 1994
Parallel and multicore computing › parallel computing › parallel software engineering
parallel application development
0.011991
An evaluation of automatic and interactive parallel programming tools · SC 1991
High-performance computing
supercomputing
0.011991
An evaluation of automatic and interactive parallel programming tools · SC 1991
Parallel and multicore computing › parallel computing
parallel program debugging
0.011989
HDB-a high level debugging · SC 1989

Methods — techniques the papers use, named apart from their topics

non-linear distribution spreading · 0.4feature weighting · 0.4active learning · 0.4protocol design · 0.0client-server model · 0.0performance evaluation · 0.0invariance assertions · 0.0checksum compression · 0.0interactive parallelization · 0.0automatic parallelization · 0.0
YearPublicationVenuePosition
2016 Clustering for Simultaneous Extraction of Aspects and Features from Reviews
abstract
Lu Chen, Justin Martineau, Doreen Cheng, Amit Sheth. Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2016.
Justin Martineau, Doreen Cheng, Amit P. Sheth
HLT-NAACL3
2015 Querying RDF data with text annotated graphs
abstract
Scientists and casual users need better ways to query RDF databases or Linked Open Data. Using the SPARQL query language requires not only mastering its syntax and semantics but also understanding the RDF data model, the ontology used, and URIs for entities of interest. Natural language query systems are a powerful approach, but current techniques are brittle in addressing the ambiguity and complexity of natural language and require expensive labor to supply the extensive domain knowledge they need. We introduce a compromise in which users give a graphical "skeleton" for a query and annotates it with freely chosen words, phrases and entity names. We describe a framework for interpreting these "schema-agnostic queries" over open domain RDF data that automatically translates them to SPARQL queries. The framework uses semantic textual similarity to find mapping candidates and uses statistical approaches to learn domain knowledge for disambiguation, thus avoiding expensive human efforts required by natural language interface systems. We demonstrate the feasibility of the approach with an implementation that performs well in an evaluation on DBpedia data.
Lushan Han, Tim Finin, Anupam Joshi, Doreen Cheng
SSDBM4
2014 Active Learning with Efficient Feature Weighting Methods for Improving Data Quality and Classification Accuracy
abstract
Many machine learning datasets are noisy with a substantial number of mislabeled instances. This noise yields sub-optimal classification performance. In this paper we study a large, low quality annotated dataset, created quickly and cheaply using Amazon Mechanical Turk to crowdsource annotations. We describe computationally cheap feature weighting techniques and a novel non-linear distribution spreading algorithm that can be used to iteratively and interactively correcting mislabeled instances to significantly improve annotation quality at low cost. Eight different emotion extraction experiments on Twitter data demonstrate that our approach is just as effective as more computationally expensive techniques. Our techniques save a considerable amount of time.
Justin Martineau, Doreen Cheng, Amit P. Sheth
ACL (1)3
2012 Situation-Aware on Mobile Phone Using Co-clustering: Algorithms and Extensions
Hyuk Cho, Deepthi Mandava, Qingzhong Liu, Lei Chen 0029, Sangoh Jeong, Doreen Cheng
IEA/AIE6
2012 Theme issue on adaptation and personalization for ubiquitous computing
Zhiwen Yu 0001, Doreen Cheng, Ismail Khalil, Judy Kay, Dominik Heckmann
Pers. Ubiquitous Comput.2
2009 SmartSearch: Situation-Aware Web Search on Mobile Devices
abstract
To provide consumers with the right information at the time of need, we developed a SmartSearch application that is able to extract a user's situational interests from usage data. It automatically constructs search queries based on situational interests. We extract the situational interests automatically without prior training and user involvement. In addition, SmartSearch is a client-side solution running completely on a mobile device to protect user's privacy.
Henry Song, Swaroop Kalasapur, Sangoh Jeong, Doreen Cheng
CCNC4
2009 Non-collaborative interest mining for personal devices
abstract
In our daily life we frequently use mobile devices to interact with the people and things on the Internet. However, finding the right things when needed is getting difficult and frustrating. In this paper, we introduce a relatively new problem of non-collaborative personal interest mining using contexts and ratings available for items of interest. We present multi-step algorithms to extract personal situational interests from mobile phone usage logs without depending on other people's data. The algorithms are based on clustering or a direct analogy from collaborative filtering. We provide extensive experimental results with our accuracy measure for synthetic data sets. The main advantages of our algorithms are: 1) no need for the user to train the phone actively, 2) no need for prior knowledge of the situations contained in a data set, 3) light-weight and running completely on a personal mobile phone and 4) good performance over low data densities. We also present a SmartSearch application. Upon user request, it automatically constructs search queries based on learned user interests and obtains information and advertisements for the user that suit the user's situation.
Sangoh Jeong, Doreen Cheng, Henry Song, Swaroop Kalasapur
CIDM2
2009 Clustering and Naive Bayesian Approaches for Situation-Aware Recommendation on Mobile Devices
abstract
In this paper, we target the problem of the situation-aware application (task) recommendation on mobile devices. To tackle this problem, we develop both supervised and unsupervised approaches. We use Naive Bayesian as a supervised approach, and co-clustering and vector quantization (VQ) as unsupervised approaches. We evaluate the performance of the proposed approaches with both synthetic and actual user log data that we have collected for six months. Our initial experiment shows that the co-clustering-based approach results in comparable purity performance with much less computation time than VQ. Therefore, the co-clustering approach can be practical for high dimensional data. Furthermore, we characterize the recommendation performance of the proposed approaches in terms of the receiver-operating-characteristics (ROC). One interesting observation is that the unsupervised approaches perform well with a single identical threshold over all applications, while the supervised approach does better with a different threshold for each application.
Sangoh Jeong, Swaroop Kalasapur, Doreen Cheng, Henry Song, Hyuk Cho
ICMLA3
2009 Extracting Co-locator context
abstract
Having reliable context sources is very important for context-aware applications and the devices around a user can be a useful context for many applications. While the importance of ‘devices around’ as a context has been highlighted many times, to the best of our knowledge, there is no systematic me
Swaroop Kalasapur, Henry Song, Doreen Cheng
MobiQuitous3
2008 Internet Search on TV
abstract
In today's content-rich homes, people use the television as a primary interface to enjoy various kinds of media, partly due to the screen size of today's TVs and partly due to its central placement in many homes. On the other hand, the Internet has become an immense source of information and entertainment content. A traditional TV viewer's experience can be greatly enhanced by enabling seamless search and access for relevant Internet content on the TV. In this demo, we illustrate a novel system and methodology that enables TV users to access relevant Internet content while watching TV.
Alan Messer, Anugeetha Kunjithapatham, Priyang Rathod, Mithun Sheshagiri, Doreen Cheng, Simon Gibbs
CCNC6
2008 SeeNSearch: A Context Directed Search Facilitator for Home Entertainment Devices
abstract
The Internet has become an extremely popular source of entertainment and information. But, despite the growing amount of media content, most Web sites today are designed for access via web browsers on the PC, making it difficult for home consumers to access Internet content on their TVs or other devices that lack keyboards. As a result, the Internet is generally restricted to access on the PC or via cumbersome interfaces on non-PC devices. In this paper, we present unobtrusive and assistive technologies enabling home users to easily find and access Internet content related to the TV program they are watching. Using these technologies, the user is now able to access relevant information and video content on the Internet while watching TV.
Alan Messer, Anugeetha Kunjithapatham, Priyang Rathod, Mithun Sheshagiri, Doreen Cheng, Simon Gibbs
PerCom6
2008 SeeNSearch: A context directed search facilitator for home entertainment devices
Alan Messer, Anugeetha Kunjithapatham, Priyang Rathod, Mithun Sheshagiri, Doreen Cheng, Simon Gibbs
Pervasive Mob. Comput.6
2007 Web Service Discovery Using General-Purpose Search Engines
abstract
WSDL provides the potential for Web services to enrich consumers' lives. However, it has had only limited success in enterprise environments and even less in the mass market. Apart from the difficulties and high costs involved in current approaches, another reason is the low precision for Web Services discovery using the most widely known tool: search engines. Seeking for effective ways to change the situation, we conducted an experiment to examine better approaches of using general-purpose search engines to discover Web Services. We used nine different approaches for publishing Web Services and two groups of total 18 queries for retrieving them using Yahoo and Google search engines. The queries were fired to each search engine daily over a week and the top 100 search results returned from every search are collected and analyzed. The results show that for both search engines, embedding a WSDL specification in a Web page that provides semantic description of the service yield the best results.
Henry Song, Doreen Cheng, Alan Messer, Swaroop Kalasapur
ICWS2
1994 NAS experiences with a prototype cluster of workstations
abstract
This paper discusses the year-long activity at NAS to implement a large, loose cluster of workstations from the existing Silicon Graphics, Inc. (SGI) pool of systems. Issues related to establishing a loosely coupled cluster of workstations are presented. Included are steps needed to resolve system management issues intended to provided reasonable cycle recovery from these systems without disrupting the primary system users. Performance evaluation tests were run based on the NAS Parallel Benchmarks (NPB) and other codes, including OVERFLOW-PVM, a full-fledged computational fluid dynamics (CFD) application. This paper summarizes the activities related to the prototype cluster and identifies areas that need improvement, development, and research in order to make workstation clusters a viable computing environment for solving aeroscience problems.>
Karen Castagnera, Doreen Cheng, Rod A. Fatoohi, Edward Hook, William T. Kramer, Craig Manning, John Musch, Charles Niggley, William Saphir, Douglas Sheppard, Merritt Smith, Ian Stockdale, Shaun Welch, Rita Williams, David Yip
SC2
1994 A portable debugger for parallel and distributed programs
abstract
We describe the design and implementation of a portable debugger for parallel and distributed programs. The design incorporates a client server model in order to isolate nonportable debugger code from the user interface. The precise definition of a protocol for client server interaction facilitates a high degree of client portability. Replication of server components permits the implementation of a debugger for distributed computations. Portability across message passing implementations is achieved with a protocol that specifies the interaction between a message passing library and the debugger. This permits the same debugger to be used both on PVM and MPI programs. The process abstractions used for debugging message passing programs can be adapted to debug HPF programs at the source level. This permits the meaningful display of information obscured in tool generated code.>
Doreen Cheng, Robert Hood
SC1
1993 Heterogeneous distributed computing (Mini symposium)
abstract
No abstract available.
Doreen Cheng
SC1
1991 An evaluation of automatic and interactive parallel programming tools
abstract
Article Free Access Share on An evaluation of automatic and interactive parallel programming tools Authors: Doreen Y. Cheng Computer Science Co., NASA Ames Research Center, MS 258-6, Moffett Field, CA Computer Science Co., NASA Ames Research Center, MS 258-6, Moffett Field, CAView Profile , Douglas M. Pase Formerly at NASA (CSC), Cray Research, Inc., 655F Lone Oak Dr., Eagan, MN Formerly at NASA (CSC), Cray Research, Inc., 655F Lone Oak Dr., Eagan, MNView Profile Authors Info & Claims Supercomputing '91: Proceedings of the 1991 ACM/IEEE conference on SupercomputingAugust 1991 Pages 412–423https://doi.org/10.1145/125826.126052Online:01 August 1991Publication History 15citation284DownloadsMetricsTotal Citations15Total Downloads284Last 12 Months11Last 6 weeks5 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteeReaderPDF
Doreen Cheng, Douglas M. Pase
SC1
1989 HDB-a high level debugging
abstract
This paper presents a new high level debugging tool, HDB, for debugging large scientific programs running on a moderate number of processors. The unique feature of HDB is that checksums are used to compress arrays and groups of variables without losing meaningful information for debugging. Using checksums makes it possible to use invariance assertions to detect misbehavior of a program at a place near the source of the error. Tracing the checksums allows the tracing of a large amount of data with a small amount of output. Comparing the traced checksums of a program and the traced checksums of its reference copy can rapidly reduce the potential error sources to a small number of subroutines. These subroutines can then be directly probed for further investigation. If desired, a debugger providing break points and single stepping source code can be used in conjunction with HDB. Examples show that using the HDB method can rapidly uncover bugs hidden in both parallel and sequential programs.
Doreen Cheng
SC1