Diego Puppin

dblp:90/2475 · DBLP profile ↗
← Back
10ranked-venue papers
5as first author
0since 2021 · last 2010
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 7 · 3 first-authorDatabases, data management, data science and information retrieval · 3 · 2 first-authorArtificial intelligence and machine learning · 2 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
3 papers
Cloud and datacenter computing · 52% Parallel and multicore computing · 26% Hardware accelerators and domain-specific architectures · 11%
Databases, data mining, and information retrieval
1 paper
Information retrieval · 100%
Software engineering, system software, and programming languages
1 paper
Compilers and program optimization · 100%

Topics — the 10 heaviest of 11, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Information retrieval
distributed information retrieval
0.112010
Tuning the capacity of search engines: Load-driven routing and incremental caching to reduce and balance the load · ACM Trans. Inf. Syst. 2010
Cloud and datacenter computing › job scheduling
batch scheduling
0.112007
A job scheduling framework for large computing farms · SC 2007
Cloud and datacenter computing
cluster resource management and scheduling
0.112007
A job scheduling framework for large computing farms · SC 2007
Parallel and multicore computing › task scheduling
learning-based scheduling
0.012003
Adapting convergent scheduling using machine learning · PPoPP 2003
Parallel and multicore computing
parallel scheduling
0.012003
Adapting convergent scheduling using machine learning · PPoPP 2003
Compilers and program optimization
instruction scheduling
0.012002
Convergent scheduling · MICRO 2002
Processor architecture and microarchitecture › clustered architecture
cluster assignment
0.012002
Convergent scheduling · MICRO 2002
Hardware accelerators and domain-specific architectures
spatial architecture
0.012002
Convergent scheduling · MICRO 2002
Information retrieval
search engines
0.012010
Tuning the capacity of search engines: Load-driven routing and incremental caching to reduce and balance the load · ACM Trans. Inf. Syst. 2010
Cloud and datacenter computing
job scheduling
0.012007
A job scheduling framework for large computing farms · SC 2007

Methods — techniques the papers use, named apart from their topics

query-vector document model · 0.1load balancing · 0.1earliest deadline first · 0.1backfilling · 0.1machine learning · 0.0
YearPublicationVenuePosition
2010 Tuning the capacity of search engines: Load-driven routing and incremental caching to reduce and balance the load
abstract
This article introduces an architecture for a document-partitioned search engine, based on a novel approach combining collection selection and load balancing, called load-driven routing . By exploiting the query-vector document model, and the incremental caching technique, our architecture can compute very high quality results for any query, with only a fraction of the computational load used in a typical document-partitioned architecture. By trading off a small fraction of the results, our technique allows us to strongly reduce the computing pressure to a search engine back-end; we are able to retrieve more than 2/3 of the top-5 results for a given query with only 10% the computing load needed by a configuration where the query is processed by each index partition. Alternatively, we can slightly increase the load up to 25% to improve precision and get more than 80% of the top-5 results. In fact, the flexibility of our system allows a wide range of different configurations, so as to easily respond to different needs in result quality or restrictions in computing power. More important, the system configuration can be adjusted dynamically in order to fit unexpected query peaks or unpredictable failures. This article wraps up some recent works by the authors, showing the results obtained by tests conducted on 6 million documents, 2,800,000 queries and real query cost timing as measured on an actual index.
Diego Puppin, Fabrizio Silvestri, Raffaele Perego 0001, Ricardo Baeza-Yates
ACM Trans. Inf. Syst.1
2007 A job scheduling framework for large computing farms
abstract
In this paper, we propose a new method, called Convergent Scheduling, for scheduling a continuous stream of batch jobs on the machines of large-scale computing farms. This method exploits a set of heuristics that guide the scheduler in making decisions. Each heuristics manages a specific problem constraint, and contributes to carry out a value that measures the degree of matching between a job and a machine. Scheduling choices are taken to meet the QoS requested by the submitted jobs, and optimizing the usage of hardware and software resources. We compared it with some of the most common job scheduling algorithms, i.e. Backfilling, and Earliest Deadline First. Convergent Scheduling is able to compute good assignments, while being a simple and modular algorithm.
Gabriele Capannini, Ranieri Baraglia, Diego Puppin, Laura Ricci, Marco Pasquali
SC3
2006 The query-vector document model
abstract
No abstract available.
Diego Puppin, Fabrizio Silvestri
CIKM1
2006 Toward a search architecture for software components
abstract
Abstract The Grid and its related technologies enable large‐scale sharing of resources of various types. We envision that in the near future applications will be completely built in a bottom‐up fashion using software components deployed on various locations and interconnected to form a workflow graph. In this paper, we make some proposals on the design of a component search service, enabling users to locate the components they need to deploy an application. Copyright © 2005 John Wiley & Sons, Ltd.
Fabrizio Silvestri, Diego Puppin, Domenico Laforenza, Salvatore Orlando 0001
Concurr. Comput. Pract. Exp.2
2005 A Grid Information Service Based on Peer-to-Peer
Diego Puppin, Stefano Moncelli, Ranieri Baraglia, Nicola Tonellotto, Fabrizio Silvestri
Euro-Par1
2004 An evaluation of component-based software design approaches
abstract
Component-oriented software design of Grid applications is commanding growing attention for business and scientific problems. The goal is create applications by assembling together independently developed software components. The components are independently developed, composable, reusable, substitutable software solutions, with clearly defined interface and behaviour.
Diego Puppin, Fabrizio Silvestri, Domenico Laforenza
CCGRID1
2004 Topic 6: Grid and Cluster Computing
Thierry Priol, Craig A. Lee, Uwe Schwiegelshohn, Diego Puppin
Euro-Par4
2004 A Search Architecture for Grid Software Components
abstract
Today, the development of Grid applications is considered a nightmare, due to lack of grid programming environments, standards, off-the-shelf software components, and so on.
Fabrizio Silvestri, Diego Puppin, Domenico Laforenza, Salvatore Orlando 0001
Web Intelligence2
2003 Adapting convergent scheduling using machine learning
abstract
No abstract available.
Diego Puppin
PPoPP1
2002 Convergent scheduling
abstract
Convergent scheduling is a general framework for cluster assignment and instruction scheduling on spatial architectures. A convergent scheduler is composed of independent passes, each implementing a heuristic that addresses a particular problem or constraint. The passes share a simple, common interface that provides spatial and temporal preference for each instruction. Preferences are not absolute; instead, the interface allows a pass to express the confidence of its preferences, as well as preferences for multiple space and time slots. A pass operates by modifying these preferences. By applying a series of passes that address all the relevant constraints, the convergent scheduler can produce a schedule that satisfies all the important constraints. Because all passes are independent and need to understand only one interface to interact with each other, convergent scheduling simplifies the problem of handling multiple constraints and co-developing different heuristics. We have applied convergent scheduling to two spatial architectures: the Raw processor and a clustered VLIW machine. It is able to successfully handle traditional constraints such as parallelism, load balancing, and communication minimization, as well as constraints due to preplaced instructions, which are instructions with predetermined cluster assignment. Convergent scheduling is able to obtain an average performance improvement of 21% over the existing space-time scheduler of the Raw processor, and an improvement of 14% over state-of-the-art assignment and scheduling techniques on a clustered VLIW architecture.
Walter Lee, Diego Puppin, Shane Swenson, Saman P. Amarasinghe
MICRO2