Marcel-Catalin Rosu

dblp:r/MarcelCatalinRosu · DBLP profile ↗
← Back
18ranked-venue papers
8as first author
0since 2021 · last 2015
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 9 · 5 first-authorSecurity and privacy · 4Software engineering, systems software and programming languages · 2 · 1 first-authorHuman-computer interaction and ubiquitous computing · 2 · 1 first-authorDatabases, data management, data science and information retrieval · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Network and information security
3 papers
Cryptographic primitives and cryptanalysis · 100%
Computer architecture, parallel and distributed computing, and storage systems
6 papers
Embedded and real-time systems · 22% Storage systems · 21% Interconnection networks and networks-on-chip · 16%
Human-computer interaction and pervasive computing
1 paper
Ubiquitous computing and smart environments · 44% Interaction techniques and input · 44% User interface design and tools · 13%
Computer networks
2 papers
Internet architecture and protocols · 57% Network performance modeling · 43%
Software engineering, system software, and programming languages
2 papers
Operating systems · 100%

Topics — the 19 heaviest of 26, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Cryptographic primitives and cryptanalysis
searchable encryption
0.422014
Dynamic Searchable Encryption in Very-Large Databases: Data Structures and Implementation · NDSS 2014
Highly-Scalable Searchable Symmetric Encryption with Support for Boolean Queries · CRYPTO (1) 2013
Cryptographic primitives and cryptanalysis › searchable encryption › searchable symmetric encryption
dynamic searchable symmetric encryption
0.212014
Dynamic Searchable Encryption in Very-Large Databases: Data Structures and Implementation · NDSS 2014
Cryptographic primitives and cryptanalysis › searchable encryption
searchable symmetric encryption
0.212013
Outsourced symmetric private information retrieval · CCS 2013
Interaction techniques and input
cross-device interaction
0.112006
Inverted Browser: A Novel Approach towards Display Symbiosis · PerCom 2006
Ubiquitous computing and smart environments › public displays
public display interaction
0.112006
Inverted Browser: A Novel Approach towards Display Symbiosis · PerCom 2006
Cryptographic primitives and cryptanalysis
symmetric cryptography
0.012013
Highly-Scalable Searchable Symmetric Encryption with Support for Boolean Queries · CRYPTO (1) 2013
Internet architecture and protocols › world wide web
web proxy
0.022003
An evaluation of TCP splice benefits in web proxy servers · WWW 2002
Kernel Support for Faster Web Proxies · USENIX ATC, General Track 2003
Operating systems › network stack
kernel networking
0.012003
Kernel Support for Faster Web Proxies · USENIX ATC, General Track 2003
High-performance computing › cluster computing
network of workstations
0.021998
Supporting Parallel Applications on Clusters of Workstations: The Intelligent Network Interface Approach · HPDC 1997
Sender Coordination in the Distributed Virtual Communication Machine · HPDC 1998
User interface design and tools › interactive systems
web browser
0.012006
Inverted Browser: A Novel Approach towards Display Symbiosis · PerCom 2006
Interconnection networks and networks-on-chip
network interface
0.011997
Supporting Parallel Applications on Clusters of Workstations: The Intelligent Network Interface Approach · HPDC 1997
Embedded and real-time systems
real-time scheduling
0.011997
CPU Reservations and Time Constraints: Efficient, Predictable Scheduling of Independent Activities · SOSP 1997
Distributed systems
distributed coordination and fault tolerance
0.011996
Early-Stopping Terminating Reliable Broadcast Protocol for General Omission Failures (Abstract) · PODC 1996
Electronic design automation › hardware verification and test › fault modeling
omission failures
0.011996
Early-Stopping Terminating Reliable Broadcast Protocol for General Omission Failures (Abstract) · PODC 1996
Distributed systems › fault tolerance › fault-tolerant protocols
reliable broadcast
0.011996
Early-Stopping Terminating Reliable Broadcast Protocol for General Omission Failures (Abstract) · PODC 1996
Processor architecture and microarchitecture
clustered architecture
0.012003
On Network CoProcessors for Scalable, Predictable Media Services · IEEE Trans. Parallel Distributed Syst. 2003
High-performance computing
collective communication
0.011998
Sender Coordination in the Distributed Virtual Communication Machine · HPDC 1998
Distributed systems › distributed system architecture
communication architecture
0.011998
Sender Coordination in the Distributed Virtual Communication Machine · HPDC 1998
Operating systems › resource management › process management
CPU scheduling
0.011997
CPU Reservations and Time Constraints: Efficient, Predictable Scheduling of Independent Activities · SOSP 1997

Methods — techniques the papers use, named apart from their topics

searchable symmetric encryption · 0.2oblivious transfer · 0.2web services · 0.1prototype implementation · 0.1real-time scheduling · 0.0experimental evaluation · 0.0socket-level implementation · 0.0precomputed schedule · 0.0network emulation · 0.0benchmarking · 0.0firmware coprocessor · 0.0active backplane · 0.0virtual communication machine · 0.0
YearPublicationVenuePosition
2015 Rich Queries on Encrypted Data: Beyond Exact Matches
Sky Faber, Stanislaw Jarecki, Hugo Krawczyk, Quan Nguyen 0006, Marcel-Catalin Rosu, Michael Steiner 0001
ESORICS (2)5
2014 Dynamic Searchable Encryption in Very-Large Databases: Data Structures and Implementation
David Cash, Joseph Jaeger, Stanislaw Jarecki, Charanjit S. Jutla, Hugo Krawczyk, Marcel-Catalin Rosu, Michael Steiner 0001
NDSS6
2013 Outsourced symmetric private information retrieval
abstract
In the setting of searchable symmetric encryption (SSE), a data owner D outsources a database (or document/file collection) to a remote server E in encrypted form such that D can later search the collection at E while hiding information about the database and queries from E. Leakage to E is to be confined to well-defined forms of data-access and query patterns while preventing disclosure of explicit data and query plaintext values. Recently, Cash et al. presented a protocol, OXT, which can run arbitrary boolean queries in the SSE setting and which is remarkably efficient even for very large databases.
Stanislaw Jarecki, Charanjit S. Jutla, Hugo Krawczyk, Marcel-Catalin Rosu, Michael Steiner 0001
CCS4
2013 Highly-Scalable Searchable Symmetric Encryption with Support for Boolean Queries
David Cash, Stanislaw Jarecki, Charanjit S. Jutla, Hugo Krawczyk, Marcel-Catalin Rosu, Michael Steiner 0001
CRYPTO (1)5
2008 Accessing speech documents on smartphones
abstract
This paper introduces BBSearch, which is an experimental system for exploring the challenges of ubiquitous access to recorded speech data. BBSearch applies information retrieval techniques to transcripts obtained by automatic speech recognition and it aims at providing a uniform user experience acro
Marcel-Catalin Rosu
MobiQuitous1
2007 A-SOAP: Adaptive SOAP Message Processing and Compression
abstract
Adaptive SOAP (A-SOAP) is a practical approach to SOAP message compression. A-SOAP separates mechanisms from policies and allows for incremental deployment. This paper focuses on its three underlying mechanisms: accelerating message composition, reducing parsing overheads, and compressing messages, which leverages the previous two mechanisms. In contrast to existing dictionary-based compression techniques, ASOAP does not require dictionaries to be exchanged between the two endpoints in advance; its dictionaries are built incrementally, as the communication progresses. ASOAP endpoints agree on the dictionary management policy using a mechanism similar to HTTP content negotiation, possibly using a dedicated HTTP header field. In experiments with short messages and a simple policy, an A-SOAP prototype reduces processing overheads by half and message sizes by an order of magnitude without increasing message latencies.
Marcel-Catalin Rosu
ICWS1
2006 Inverted Browser: A Novel Approach towards Display Symbiosis
abstract
In this paper we introduce the inverted browser, a novel approach to enable mobile users to view content from their personal devices on public displays. The inverted browser is a network service to start and control a browser that is then used to view the content. In contrast to a traditional Web browser, which runs on the client device and pulls content from a server, content is pushed to the inverted browser from a personal data source upon user input. This approach allows a wide variety of personal content to be viewed by facilitating symbiotic relationships between mobile devices and intelligent displays in the environment. Our initial inverted browser prototype is based on a Web services wrapper around a traditional Web browser. Our experiments show that the inverted browser approach is superior to other solutions in terms of user convenience, ease of use, energy consumption, and privacy, but interaction latencies need improvement.
Mandayam T. Raghunath, Nishkam Ravi, Marcel-Catalin Rosu, Chandrasekhar Narayanaswami 0001
PerCom3
2003 Kernel Support for Faster Web Proxies
Marcel-Catalin Rosu, Daniela Rosu 0001
USENIX ATC, General Track1
2003 On Network CoProcessors for Scalable, Predictable Media Services
abstract
This paper presents the embedded realization and experimental evaluation of a media stream scheduler on network interface (NI) CoProcessor boards. When using media frames as scheduling units, the scheduler is able to operate in real-time on streams traversing the CoProcessor, resulting in its ability to stream video to remote clients at real-time rates. This paper presents a detailed evaluation of the effects of placing application or kernel-level functionality, like packet scheduling on NIs, rather than the host machines to which they are attached. The main benefits of such placement are: 1) that traffic is eliminated from the host bus and memory subsystem, thereby allowing increased host CPU utilization for other tasks, and 2) that NI-based scheduling is immune to host-CPU loading, unlike host-based media schedulers that are easily affected even by transient load conditions. An outcome of this work is a proposed cluster architecture for building scalable media servers by distributing schedulers and media stream producers across the multiple NIs used by a single server and by clustering a number of such servers using commodity network hardware and software.
Raj Krishnamurthy, Karsten Schwan, Richard West, Marcel-Catalin Rosu
IEEE Trans. Parallel Distributed Syst.4
2002 An evaluation of TCP splice benefits in web proxy servers
abstract
This study is the first to evaluate the performance benefits of using the recently proposed TCP Splice kernel service in Web proxy servers. Previous studies show that splicing client and server TCP connections in the IP layer improves the throughput of proxy servers like firewalls and content routers by reducing the data transfer overheads. In a Web proxy server, data transfer overheads represent a relatively large fraction of the request processing overheads, in particular when content is not cacheable or the proxy cache is memory-based. The study is conducted with a socket-level implementation of TCP Splice. Compared to IP-level implementations, socket-level implementations make possible the splicing of connections with different TCP characteristics, and improve response times by reducing recovery delay after a packet loss. The experimental evaluation is focused on HTTP request types for which the proxy can fully exploit the TCP Splice service, which are the requests for non-cacheabl.content and SSL tunneling. The experimental testbed includes an emulated WAN environment and benchmark applications for HTTP/1.0 Web client, Web server, and Web proxy running on AIX RS/6000 machines. Our experiments demonstrate that TCP Splice enables reductions in CPU utilization of 10-43% of the CPU, depending on file sizes and request rates. Larger relative reductions are observed when tunneling SSL connections, in particular for small file transfers. Response times are also reduced by up to 1.8sec.
Marcel-Catalin Rosu, Daniela Rosu 0001
WWW1
2000 A Network Co-Processor-Based Approach to Scalable Media Streaming in Servers
abstract
This paper presents the embedded construction and experimental results for a media scheduler on i960 RD equipped I20 Network Interfaces (NI) used for streaming. We utilize the Distributed Virtual Communication Machine (DVCM) infrastructure developed by us which allows run-time extensions to provide scheduling for streams that may require it. The scheduling overhead of such a scheduler is /spl ap/65 /spl mu/s with the ability to stream MPEG video to remote clients at requested rates. Moreover, placement of scheduler action 'close' to the network on the Network Interface (NI) allows tighter coupling of computation and communication, eliminating traffic from the host bus and memory subsystem, allowing increased host CPU utilization for other tasks without being affected by host-CPU loading. Architectures to build scalable media scheduling servers are explored-by distributing media schedulers and media stream producers among NIs within a server and clustering a number of such servers using commodity hardware and software.
Raj Krishnamurthy, Karsten Schwan, Richard West, Marcel-Catalin Rosu
ICPP4
2000 Support for Recoverable Memory in the Distributed Virtual Communication Machine
abstract
Distributed Virtual Communication Machine (DVCM) is a software communication architecture for clusters of workstations equipped with programmable network interfaces (Nls) for high-speed networks. DVCM is an extensible architecture, which promotes the transfer of application modules to the NI. By executing "closer" to the network, on the NI CoProcessor, these modules can communicate with significantly higher message rates and lower latencies than achievable at the CPU-level. This paper describes how DVCM modules can be used to enhance the performance of the Cluster Recoverable Memory system (CRMem), a transaction-processing kernel for memory-resident databases. By using the NI CoProcessor for CRMem's remote operations, our implementation achieves more than 3,000 trans/sec on a simplified TpcB benchmark.
Marcel-Catalin Rosu, Karsten Schwan
IPDPS1
1998 Sender Coordination in the Distributed Virtual Communication Machine
abstract
The Distributed Virtual Communication Machine (DVCM) is an extensible communication architecture for tightly-coupled clusters of workstations (COWs) connected by high-speed networks. The DVCM is designed for off-the-shelf network interface cards equipped with communication coprocessors. Its main component is an active backplane implemented in firmware running on the coprocessors. This backplane can be extended with modules that implement application-specific-functionality and have access to some of the application's state. Consequently non-trivial collective computations can be implemented as DVCM extensions. We present a DVCM extension module that provides application-specific network flow control by coordinating the resource-competing components of a parallel application running on an ATM LAN. Our experiments show that this extension module helps eliminate message loss and achieve high link bandwidth utilization when there is significant link contention.
Marcel-Catalin Rosu, Karsten Schwan
HPDC1
1997 Supporting Parallel Applications on Clusters of Workstations: The Intelligent Network Interface Approach
abstract
This paper presents a novel networking architecture designed for communication intensive parallel applications running on clusters of workstations (COWs) connected by high speed network. This architecture permits: (1) the transfer of selected communication-related functionality the host machine to the network interface coprocessor and (2) the exposure of this functionality directly to applications as instructions of a Virtual Communication Machine (VCM) implemented by the coprocessor. The user-level code interacts directly with the network coprocessor as the host kernel only 'connects' the application to the VCM and does not participate in the data transfers. The distinctive feature of our design is its flexibility: the integration of the network with the application can be varied to maximize performance. The resulting communication architecture is characterized by a very low overhead on the host processor by latency and bandwidth close to the hardware limits, and by an application interface which enables zero-copy messaging and eases the port of some shared-memory parallel applications to COWs. The architecture admits low cost implementations based only on off-the-shelf hardware components. Additionally, its current ATM-based implementation can be used to communicate with any ATM-enabled host.
Marcel-Catalin Rosu, Karsten Schwan, Richard M. Fujimoto
HPDC1
1997 CPU Reservations and Time Constraints: Efficient, Predictable Scheduling of Independent Activities
abstract
Workstations and personal computers are increasingly being used for applications with real-time characteristics such as speech understanding and synthesis, media computations and I/O. and animation, often concurrently executed with traditional nonreal-time workloads.This paper presents a system that can schedule multiple independent activities so that: activities can obta& minimum guaranteed execution rates with application-specified reservation granularities via CPU Reservations, CPU Reservations, which are of the form "reserve X units of time out of every Y units", provide not just an average case execution rate of X/Y over long periods of time, but the stronger guarantee that from any instant of time, by Y time units later, the activity will have executedfor at least X time units, applications can use Time Constraints to schedule tasks by deadlines, with on-time completion guaranteed for tasks with accepted constraints, and both CPU Reservations and Time Constraints are implemented very efficiently.In particular, CPU scheduling overhead is bounded by a constant and is not a function of the number of schedulable tasks.Other key scheduler properties are: l activities cannot violate other activities' guarantees, l time constraints and CPU reservations may be used together, separately, or not at all (which gives a round-robin schedule), with well-defined interactions between all combinations, and l spare CPU time is fairly shared among all activities.The Rialto operating system, developed at Microsoft Research, achieves these goals by using a precomputed schedule, which is the fundamental basis of this work.
Michael B. Jones, Daniela Rosu 0001, Marcel-Catalin Rosu
SOSP3
1997 Efficient Message Passing Interface (MPI) for Parallel Computing on Clusters of Workstations
Jehoshua Bruck, Danny Dolev, C. T. Howard Ho, Marcel-Catalin Rosu, Ray Strong
J. Parallel Distributed Comput.4
1996 Early-Stopping Terminating Reliable Broadcast Protocol for General Omission Failures (Abstract)
abstract
No abstract available.
Marcel-Catalin Rosu
PODC1
1995 Efficient Message Passing Interface (MPI) for Parallel Computing on Clusters of Workstations
abstract
Parallel computing on clusters of workstations and personal computers has very high \npotential, since it leverages existing hardware and software. Parallel programming \nenvironments offer the user a convenient way to express parallel computation and communication. \nIn fact, recently, a Message Passing Interface (MPI) has been proposed as an industrial \nstandard for writing "portable" message-passing parallel programs. The communication \npart of MPI consists of the usual point-to-point communication as well as collective \ncommunication. However, existing implementations of programming environments for clusters \nare built on top of a point-to-point communication layer (send and receive) over local \narea networks (LANs) and, as a result, suffer from poor performance in the collective \ncommunication part. \nIn this paper, we present an efficient design and implementation of the collective \ncommunication part in MPI that is optimized for clusters of workstations. Our system consists \nof two main components: the MPI-CCL layer that includes the collective communication \nfunctionality of MPI and a User-level Reliable Transport Protocol (URTP) that interfaces \nwith the LAN Data-link layer and leverages the fact that the LAN is a broadcast medium. \nOur system is integrated with the operating system via an efficient kernel extension \nmechanism that we developed. The kernel extension significantly improves the performance of \nour implementation as it can handle part of the communication overhead without involving \nuser space. \nWe have implemented our system on a collection of IBM RS/6000 workstations con- \nnected via a lOMbit Ethernet LAN. Our performance measurements are taken from typical \nscientific programs that run in a parallel mode by means of the MPI. The hypothesis behind \nour design is that system's performance will be bounded by interactions between the kernel \nand user space rather than by the bandwidth delivered by the LAN Data-Link Layer. Our \nresults indicate that the performance of our MPI Broadcast (on top of Ethernet) is about \ntwice as fast as a recently published software implementation of broadcast on top of ATM.
Jehoshua Bruck, Danny Dolev, C. T. Howard Ho, Marcel-Catalin Rosu, Ray Strong
SPAA4