EDBT 2026 Demo / reviewers in the wild / expert
Thomas Phan
dblp:74/5955
· DBLP profile ↗
22ranked-venue papers
11as first author
0since 2021 · last 2015
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 7 · 4 first-authorDatabases, data management, data science and information retrieval · 7 · 3 first-authorSystems, architecture and hardware · 5 · 2 first-authorArtificial intelligence and machine learning · 1Human-computer interaction and ubiquitous computing · 1Applied, interdisciplinary, general and emerging computing · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Databases, data mining, and information retrieval
3 papers |
Recommender systems · 23% Data mining · 23% Web and social media mining · 23% | |
| Computer architecture, parallel and distributed computing, and storage systems
6 papers |
Cloud and datacenter computing · 35% Distributed systems · 20% Parallel and multicore computing · 19% | |
| Software engineering, system software, and programming languages
3 papers |
Services computing and microservices · 85% Compilers and program optimization · 15% |
Topics — the 23 heaviest of 27, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Recommender systems
collaborative filtering |
0.2 | 1 | 2015 | Who, What, When, and Where: Multi-Dimensional Collaborative Recommendations Using Tensor Factorization on Sparse User-Generated Data · WWW 2015 |
Data mining › multidimensional data analysis › multiway data analysis › tensor analysis
tensor factorization |
0.2 | 1 | 2015 | Who, What, When, and Where: Multi-Dimensional Collaborative Recommendations Using Tensor Factorization on Sparse User-Generated Data · WWW 2015 |
Cloud and datacenter computing › cloud service models
software as a service |
0.1 | 1 | 2009 | Frontiers in Information and Software as Services · ICDE 2009 |
Indexing and storage engines
adaptive indexing |
0.1 | 1 | 2008 | Dynamic Materialization of Query Views for Data Warehouse Workloads · ICDE 2008 |
Indexing and storage engines
index management |
0.1 | 1 | 2008 | Dynamic Materialization of Query Views for Data Warehouse Workloads · ICDE 2008 |
Query processing and optimization › materialized view
materialized view selection |
0.1 | 1 | 2008 | Dynamic Materialization of Query Views for Data Warehouse Workloads · ICDE 2008 |
Services computing and microservices
service-oriented architecture |
0.1 | 1 | 2008 | A request-routing framework for SOA-based enterprise computing · Proc. VLDB Endow. 2008 |
Parallel and multicore computing
load balancing |
0.1 | 1 | 2008 | A request-routing framework for SOA-based enterprise computing · Proc. VLDB Endow. 2008 |
Services computing and microservices
middleware |
0.0 | 1 | 2004 | Middleware Support for Reconciling Client Updates and Data Transcoding · MobiSys 2004 |
Embedded and real-time systems
mobile computing |
0.0 | 1 | 2003 | iMASH: Interactive Mobile Application Session Handoff · MobiSys 2003 |
Distributed systems
grid computing |
0.0 | 1 | 2002 | Challenge: : integrating mobile wireless devices into the computational grid · MobiCom 2002 |
Distributed systems › distributed resource management
resource aggregation |
0.0 | 1 | 2002 | Challenge: : integrating mobile wireless devices into the computational grid · MobiCom 2002 |
Compilers and program optimization
compiler analysis |
0.0 | 1 | 1999 | Compiler-Supported Simulation of Highly Scalable Parallel Applications · SC 1999 |
Performance modeling and evaluation › simulation › architectural simulation
execution-driven simulation |
0.0 | 1 | 1999 | Performance Prediction of Large Parallel Applications using Parallel Simulations · PPoPP 1999 |
Parallel and multicore computing › parallel computing
parallel application simulation |
0.0 | 1 | 1999 | Compiler-Supported Simulation of Highly Scalable Parallel Applications · SC 1999 |
Performance modeling and evaluation › simulation › parallel and distributed simulation
parallel simulation |
0.0 | 1 | 1999 | Performance Prediction of Large Parallel Applications using Parallel Simulations · PPoPP 1999 |
Performance modeling and evaluation
performance prediction |
0.0 | 1 | 1999 | Performance Prediction of Large Parallel Applications using Parallel Simulations · PPoPP 1999 |
Performance modeling and evaluation
simulation |
0.0 | 1 | 1999 | Compiler-Supported Simulation of Highly Scalable Parallel Applications · SC 1999 |
Edge and fog computing › mobile edge computing › computation offloading › mobile computation offloading
cyber foraging |
0.0 | 1 | 2003 | iMASH: Interactive Mobile Application Session Handoff · MobiSys 2003 |
Wireless networking
mobile computing |
0.0 | 1 | 2002 | Challenge: : integrating mobile wireless devices into the computational grid · MobiCom 2002 |
Internet of things and sensor networks
resource-constrained devices |
0.0 | 1 | 2002 | Challenge: : integrating mobile wireless devices into the computational grid · MobiCom 2002 |
Distributed systems › operating system support › interprocess communication
message-passing systems |
0.0 | 1 | 1999 | Compiler-Supported Simulation of Highly Scalable Parallel Applications · SC 1999 |
Performance modeling and evaluation › simulation › discrete-event simulation
parallel discrete event simulation |
0.0 | 1 | 1999 | Performance Prediction of Large Parallel Applications using Parallel Simulations · PPoPP 1999 |
Methods — techniques the papers use, named apart from their topics
coupled tensor and matrix factorization · 0.2survey · 0.1reconciliation rules · 0.1proxy · 0.1middleware · 0.1genetic algorithm · 0.1content adaptation · 0.1LRU cache · 0.1economic model · 0.1static task graph analysis · 0.0compiler-supported simulation · 0.0parallel discrete-event simulation · 0.0direct execution · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2015 | Who, What, When, and Where: Multi-Dimensional Collaborative Recommendations Using Tensor Factorization on Sparse User-Generated DataabstractGiven the abundance of online information available to mobile users, particularly tourists and weekend travelers, recommender systems that effectively filter this information and suggest interesting participatory opportunities will become increasingly important. Previous work has explored recommending interesting locations; however, users would also benefit from recommendations for activities in which to participate at those locations along with suitable times and days. Thus, systems that provide collaborative recommendations involving multiple dimensions such as location, activities and time would enhance the overall experience of users.The relationship among these dimensions can be modeled by higher-order matrices called tensors which are then solved by tensor factorization. However, these tensors can be extremely sparse. In this paper, we present a system and an approach for performing multi-dimensional collaborative recommendations for Who (User), What (Activity), When (Time) and Where (Location), using tensor factorization on sparse user-generated data. We formulate an objective function which simultaneously factorizes coupled tensors and matrices constructed from heterogeneous data sources. We evaluate our system and approach on large-scale real world data sets consisting of 588,000 Flickr photos collected from three major metro regions in USA. We compare our approach with several state-of-the-art baselines and demonstrate that it outperforms all of them. Preeti Bhargava, Thomas Phan, Juhan Lee |
WWW | 2 |
| 2014 | Sensor fusion of physical and social data using Web SocialSense on smartphone mobile browsersabstractModern smartphones offer a rich selection of onboard sensors, where sensor access is typically performed through API calls provided by the phone's operating system. In this paper we evaluate the viability of implementing sensor processing entirely in the Web browser layer with Web SocialSense, a JavaScript framework for Tizen smartphones that uses a graph topology-based paradigm. This framework enables programmers to write personalized, context-aware applications that can dynamically fuse time-series signals from physical sensors (such as the accelerometer and geolocation services) and social software sensors (such as social network services and personal information management applications). To demonstrate the framework we implemented components for physical sensing and social software sensing to drive two context-aware applications, ActVertisements and Social Map. Thomas Phan, Swaroop Kalasapur, Anugeetha Kunjithapatham |
CCNC | 1 |
| 2012 | Editorial SI: Mobile Applications and Services
Thomas Phan, Rebecca Montanari, Petros Zerfos |
Mob. Networks Appl. | 1 |
| 2009 | Frontiers in Information and Software as ServicesabstractThe high cost of creating and maintaining software and hardware infrastructures for delivering services to businesses has led to a notable trend toward the use of third-party service providers, which rent out network presence, computation power, and data storage space to clients with infrastructural needs. These third party service providers can act as data stores as well as entire software suites for improved availability and system scalability, reducing small and medium businesses' burden of managing complex infrastructures. This is called information/application outsourcing or software as a service (SaaS). Emergence of enabling technologies, such as service oriented architectures (SOA), virtual machines, and cloud computing, contribute to this trend. Scientific Grid computing, on-line software services, and business service networks are typical examples leveraging database and software as service paradigm. In this paper, we survey the technologies used to enable SaaS paradigm as well as the current offerings on the market. We also outline research directions in the field. K. Selçuk Candan, Wen-Syan Li, Thomas Phan, Minqi Zhou |
ICDE | 3 |
| 2008 | Load distribution of analytical query workloads for database cluster architecturesabstractEnterprises may have multiple database systems spread across the organization for redundancy or for serving different applications. In such systems, query workloads can be distributed across different servers for better performance. A materialized view, or Materialized Query Table (MQT), is an auxiliary table with pre-computed data that can be used to significantly improve the performance of a database query. In this paper, we propose a framework for coordinating execution of OLAP query workloads across a database cluster with shared nothing architecture. Such coordination is complex since we need to consider (1) the time to build the MQTs, (2) the query execution impact of the MQTs, (3) whether the MQTs can fit in the disk space limitation, (4) server computation power, and (5) the effectiveness of the scheduling and placement algorithms in deriving a combination of configurations so that the workload can be completed in the shortest time period. We frame the problem as a combinatorial problem with a solution space that is exponential in the number of queries, MQTs, and servers. We provide a stochastic search heuristic that finds a near-optimal mapping of queries-to-servers and MQTs-to-servers within an arbitrarily bounded time and compare our solution with an exhaustive search and three standard greedy algorithms. Our search implementation produced schedules within 9% of the optimal found through an exhaustive search and produced better solutions than typical greedy algorithms for both TPC-H and synthetic benchmarks under a variety of experiments. For a key trial where disk space is limited, it produced 15% better results than the next best competitor, corresponding to an absolute wall clock advantage of over 10 hours. Thomas Phan, Wen-Syan Li |
EDBT | 1 |
| 2008 | Dynamic Materialization of Query Views for Data Warehouse WorkloadsabstractA materialized view, or Materialized Query Table (MQT), is an auxiliary table with precomputed data that can be used to significantly improve the performance of a database query. Previous research efforts have focused on finding the best candidate MQT set, with a common static heuristic being to greedily pre-materialize the MQTs prior to executing the workload. While this approach is sound when the size of the MQT set on disk is small, it will not be able to pre-materialize all MQTs and indexes when faced with real-world disk limits and view maintenance costs, and thus a static heuristic will fail to exploit the potentially large benefits of those MQTs not selected for materialization. In this paper we present an automated, dynamic MQT management scheme that materializes views and creates indexes in an on-demand fashion as a workload executes and manages them with an LRU cache. In order to maximize the benefit of executing queries with MQTs, the scheme makes an adaptive tradeoff between the MQT materializations, the base table accesses, and the benefit of MQT hits in the cache. To find the workload permutation that produces the overall highest net benefit, we use a genetic algorithm to search the N! solution space, and to avoid materializing seldom-used MQTs, we prune the set of MQT candidates. We ran our dynamic management on a TPC-H workload and found that our scheme produces higher benefit across a variety of scenarios; we demonstrate over 60% improvement in MQT benefit under harsh conditions when the size of available MQTs with indexes is larger than the cache. Thomas Phan, Wen-Syan Li |
ICDE | 1 |
| 2008 | A request-routing framework for SOA-based enterprise computingabstractEnterprises may use a service-oriented architecture (SOA) to provide a streamlined interface to their business processes. To scale up the system, each tier in a composite service usually deploys multiple servers for load distribution and fault tolerance. Such load distribution across multiple servers within the same tier can be viewed as horizontal load distribution. One limitation of this approach is that load cannot be further distributed when all servers in the same tier are fully loaded. In complex multi-tiered systems, a single business process may actually be implemented by multiple different computation pathways among the tiers, each with different components, in order to provide resiliency and scalability. Such SOA-based enterprise computing with multiple implementation options gives opportunities for vertical load distribution across tiers. In this paper, we propose a requestrouting framework for SOA-based enterprise computing that takes into consideration both horizontal and vertical load distribution. Through experimentation we show that our algorithm and methodology scale well up to a large system configuration comprising up to 1000 workflow requests to a complex composite service with multiple implementations. We also show that a combination of both horizontal and vertical load distributions gives the maximum flexibility to improve performance and fault tolerance. Thomas Phan, Wen-Syan Li |
Proc. VLDB Endow. | 1 |
| 2007 | Middleware and Performance Issues for Computational Finance Applications on Blue Gene/LabstractWe discuss real-world case studies involving the implementation of a Web services middleware tier for the IBM Blue Gene/L supercomputer to support financial business applications. These programs that are representative of a class of modern financial analytics that take part in distributed business workflows and are heavily database-centric with input and output data stored in external SQL data warehouses. We describe the design issues related to the development of our middleware tier that provides a number of core features, including an automated SQL data extraction and staging gateway, a standardized high-level job specification schema, a well-defined Web services (SOAP) API for interoperability with other applications, and a secure HTML/JSP Web-based interface suitable for general users. Further, we provide observations on performance optimizations to support the relevant data movement requirements. Thomas Phan, Ramesh Natarajan, Satoki Mitsumori |
IPDPS | 1 |
| 2007 | A grid-based approach for enterprise-scale data mining
Ramesh Natarajan, Radu Sion, Thomas Phan |
Future Gener. Comput. Syst. | 3 |
| 2006 | XG: A Grid-Enabled Query Processing Engine
Radu Sion, Ramesh Natarajan, Inderpal Narang, Thomas Phan |
EDBT | 4 |
| 2006 | TypeCast: Type-Based Routing in Wireless Ad-hoc NetworksabstractType-based communication is proposed as an effective paradigm to enable group communication in wireless ad-hoc networks (MANETs). In this paradigm, type is used as the fundamental construct for addressing and routing messages. Type hierarchies are used to dynamically control group size; and object-oriented principles such as subtyping and multiple inheritance are utilized to construct new groups from existing ones. We present the design of TypeCast, a routing protocol that directly supports type-based communication. TypeCast leverages efficiency and mobility management provided by MANET multicast protocols and extends them by adding a Bloom filter-based type encoding and routing mechanism. TypeCast is fully decentralized and supports subtyping and type-composition. We implement TypeCast on top of ODMRP and conduct a detailed performance and scalability study of TypeCast through simulation. The results show that TypeCast demonstrates good resiliency to mobility and group size. When the number of types in the network increases, TypeCast achieves good scalability thanks to type aggregation provided by Bloom filters Jinsong Lin, Thomas Phan, Rajive L. Bagrodia |
MobiQuitous | 2 |
| 2005 | XG: A Data-Driven Computation Grid for Enterprise-Scale Mining
Radu Sion, Ramesh Natarajan, Inderpal Narang, Wen-Syan Li, Thomas Phan |
DEXA | 5 |
| 2005 | Evolving Toward the Perfect Schedule: Co-scheduling Job Assignments and Data Replication in Wide-Area Systems Using a Genetic Algorithm
Thomas Phan, Kavitha Ranganathan, Radu Sion |
JSSPP | 1 |
| 2004 | Middleware Support for Reconciling Client Updates and Data TranscodingabstractIn mobile Internet applications, data can be transcoded, updated, and transferred across heterogenous clients. The problem then arises where updates made in the context of an initial transcoding results in content too stringently transcoded for subsequent clients, thereby causing loss of semantic value. We solve this problem by suggesting that the updates themselves can be transformed so that they can be applied directly to the original data instead of to the transcoded data; this approach allows the data to preserve as much semantic value as possible across all heterogeneous clients without unnecessary transcoding artifacts. We define reconciliation rules that can govern the interaction between client updates and transcoding, demonstrate a complete middleware architecture that supports our methodology, and provide two case studies using content-transferring applications. We show that our resulting middleware system executes our reconciliation approach with acceptable latency (under 5 seconds for 200 kbytes of layered content), good scalability, and well-organised modularity. Thomas Phan, George Zorpas, Rajive L. Bagrodia |
MobiSys | 1 |
| 2003 | iMASH: Interactive Mobile Application Session HandoffabstractMobile computing research has often focused on untethering an in-use computing device, rather than enabling the mobility of the computation task itself. This paper presents an architecture, implementation, and experimental evidence that together validate a new continuous computing concept, application session handoff. The iMASH architecture leverages previous work on proxies, content adaptation, and client awareness to provide a unique, middleware-enable capability for continuous computing. Implementation in both socket- and RPC-based environments shows that very fast, secure session handoff of non-trivial client/server applications across heterogeneous client devices and network is feasible: experiments on a number of applications yielded handoff latencies ranging from 0.5s to 2s. Rajive L. Bagrodia, S. Bhattacharyya, Fred Cheng, Steve Gerding, Glenn Glazer, Richard G. Guy, Zhengrong Ji, Jinsong Lin, Thomas Phan, Erik Skow, Maneesh Varshney, George Zorpas |
MobiSys | 9 |
| 2003 | A Scalable, Distributed Middleware Service Architecture to Support Mobile Internet Applications
Rajive L. Bagrodia, Thomas Phan, Richard G. Guy |
Wirel. Networks | 2 |
| 2002 | A security architecture for application session handoffabstractUbiquitous computing across a variety of wired and wireless connections still lacks an effective security architecture. In our research work, we address the specific issue of designing and building a security architecture for application session handoff, a functionality which we envision will be a key component enabling ubiquitous computing. Our architecture incorporates a number of proven approaches into the new context of ubiquitous computing. We employ the Bell-LaPadula (1976) and capability models to realise access control and adopt public key infrastructure (PKI)-based approaches to provide efficient and authenticated end-to-end security. To demonstrate the effectiveness of our design, we implemented an application enabled with this security architecture and showed that it incurred low latency. Erik Skow, Jiejun Kong, Thomas Phan, Fred Cheng, Richard G. Guy, Rajive L. Bagrodia, Mario Gerla, Songwu Lu |
ICC | 3 |
| 2002 | An Extensible and Scalable Content Adaptation Pipeline Architecture to Support Heterogeneous ClientsabstractThe importance of middleware and content adaptation has previously been demonstrated for pervasive use of Web-based applications. In this paper we propose a modular extensible, and scalable middleware component called the Content Adaptation Pipeline that performs content adaptation on arbitrarily complex data types not limited to text and graphic images. Furthermore, the architecture can be used as part of many client-server applications, not just Web browsers. In our work we leverage the XML language as a uniform means to describe all the elements in our architecture, including the client device and user profiles, the data characteristics, the transcoding operations performed on the data, and the resultant adapted data. We illustrate the flexibility of our architecture to support new data types and adaptation operations by first showing its use with data from a real-world medical application and then extending its capabilities to handle animated graphics and also real-time streaming RTP data. Finally, we demonstrate scalability in our architecture by executing the Content Adaptation Pipeline over a distributed set of servers running an efficient protocol. Thomas Phan, George Zorpas, Rajive L. Bagrodia |
ICDCS | 1 |
| 2002 | Challenge: : integrating mobile wireless devices into the computational gridabstractOne application domain the mobile computing community has not yet entered is that of grid computing -- the aggregation of network-connected computers to form a large-scale, distributed system used to tackle complex scientific or commercial problems. In this paper we present the challenge of harvesting the increasingly widespread availability of Internet connected wireless mobile devices such as PDAs and laptops to be beneficially used within the emerging national and global computational grid. The integration of mobile wireless consumer devices into the Grid initially seems unlikely due to the inherent limitations typical of mobile devices, such as reduced CPU performance, small secondary storage, heightened battery consumption sensitivity, and unreliable low-bandwidth communication. However, the millions of laptops and PDAs sold annually suggest that this untapped abundance should not be prematurely dismissed. Given that the benefits of combining the resources of mobile devices with the computational grid are potentially enormous, one must compensate for the inherent limitations of these devices in order to successfully utilise them in the Grid. In this paper we identify the research challenges arising from this problem and propose our vision of a potential architectural solution. We suggest a proxy based, clustered system architecture with favourable deployment, interoperability, scalability, adaptivity, and fault-tolerance characteristics as well as an economic model to stimulate future research in this emerging field. Thomas Phan, Lloyd Huang, Chris Dulan |
MobiCom | 1 |
| 2001 | Handoff of application sessions across time and spaceabstractPersonal computing on mobile platforms such as laptops and personal digital assistants, rather than in a traditional desktop environment, is becoming increasingly more common. We address the issue of application session transfer for uninterrupted data access across this diverse range of platforms. This work is part of the iMASH project, a multi-year, multi-discipline collaborative effort focused on enabling mobile client platforms and incorporating them into existing legacy networked systems for use by medical practitioners. We have developed a tiered architecture that includes a middleware server layer positioned between existing application servers and multiple clients to make session transfer transparent to the user. Any client application executing our middleware-aware remote code library can save and restore its session by interacting with a middleware server. As a proof of concept, we have implemented the transfer of bookmarks, history, Web cache, and user preferences with the Mozilla open source Web browser, from this effort we have established baseline performance metrics and have found that the overhead is within reasonable bounds of just a few seconds of latency. Thomas Phan, Kaixin Xu, Richard G. Guy, Rajive L. Bagrodia |
ICC | 1 |
| 1999 | Performance Prediction of Large Parallel Applications using Parallel SimulationsabstractAccurate simulation of large parallel applications can be facilitated with the use of direct execution and parallel discrete event simulation. This paper describes the use of COMPASS, a direct execution-driven, parallel simulator for performance prediction of programs that include both communication and I/O intensive applications. The simulator has been used to predict the performance of such applications on both distributed memory machines like the IBM SP and shared-memory machines like the SGI Origin 2000. The paper illustrates the usefulness of COMPASS as a versatile performance prediction tool. We use both real-world applications and synthetic benchmarks to study application scalability, sensitivity to communication latency, and the interplay between factors like communication pattern and parallel file system caching on application performance. We also show that the simulator is accurate in its predictions and that it is also efficient in its ability to use parallel simulation to reduce its own execution time which, in some cases, has yielded a nearlinear speedup. Rajive L. Bagrodia, Ewa Deelman, Steven Docy, Thomas Phan |
PPoPP | 4 |
| 1999 | Compiler-Supported Simulation of Highly Scalable Parallel ApplicationsabstractIn this paper, we propose and evaluate practical, automatic techniques that exploit compiler analysis to facilitate simulation of very large message-passing systems.We use a compilersynthesized static task graph model to identify the control-flow and the subset of the computations that determine the parallelism, communication and synchronization of the code, and to generate symbolic estimates of sequential task execution times.This information allows us to avoid executing or simulating large portions of the computational code during the simulation.We have used these techniques to integrate the MPI-Sim parallel simulator at UCLA with the Rice dHPF compiler infrastructure.The integrated system can simulate unmodified High Performance Fortran (HPF) programs compiled to the Message-Passing Interface standard (MPI) by the dHPF compiler, and we expect to simulate MPI programs as well.We evaluate the accuracy and benefits of these techniques for three standard benchmarks on a wide range of problem and system sizes.Our results show that the optimized simulator has errors of less than 17% compared with direct program measurement in all the cases we studied, and typically much smaller errors.Furthermore, it requires factors of 5 to 2000 less memory and up to a factor of 10 less time to execute than the original simulator.These dramatic savings allow us to simulate systems and problem sizes 10 to 100 times larger than is possible with the original simulator. Vikram S. Adve, Rajive L. Bagrodia, Ewa Deelman, Thomas Phan, Rizos Sakellariou |
SC | 4 |