Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Jiachen Xue

dblp:66/7819 · DBLP profile ↗
← Back
7ranked-venue papers
4as first author
0since 2021 · last 2020
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 3 · 2 first-authorSoftware engineering, systems software and programming languages · 2Artificial intelligence and machine learning · 1Computer networks · 1 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
2 papers
Interconnection networks and networks-on-chip · 44% Cloud and datacenter computing · 41% Distributed systems · 7%
Computer networks
1 paper
Datacenter networks · 100%
Computer graphics and multimedia
1 paper
Audio and music processing · 50% Multimedia analysis and retrieval · 50%

Topics — the 13 heaviest of 15, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Datacenter networks
incast
0.412020
Dart: Divide and Specialize for Fast Response to Congestion in RDMA-Based Datacenter Networks · IEEE/ACM Trans. Netw. 2020
Datacenter networks
RDMA
0.412020
Dart: Divide and Specialize for Fast Response to Congestion in RDMA-Based Datacenter Networks · IEEE/ACM Trans. Netw. 2020
Datacenter networks › RDMA
RDMA congestion control
0.412020
Dart: Divide and Specialize for Fast Response to Congestion in RDMA-Based Datacenter Networks · IEEE/ACM Trans. Netw. 2020
Cloud and datacenter computing
datacenter network
0.412020
Network Interface Architecture for Remote Indirect Memory Access (RIMA) in Datacenters · ACM Trans. Archit. Code Optim. 2020
Interconnection networks and networks-on-chip
network interface
0.412020
Network Interface Architecture for Remote Indirect Memory Access (RIMA) in Datacenters · ACM Trans. Archit. Code Optim. 2020
Interconnection networks and networks-on-chip
remote direct memory access
0.412020
Network Interface Architecture for Remote Indirect Memory Access (RIMA) in Datacenters · ACM Trans. Archit. Code Optim. 2020
Datacenter networks
load balancing
0.112020
Dart: Divide and Specialize for Fast Response to Congestion in RDMA-Based Datacenter Networks · IEEE/ACM Trans. Netw. 2020
Memory systems
memory management
0.112020
Network Interface Architecture for Remote Indirect Memory Access (RIMA) in Datacenters · ACM Trans. Archit. Code Optim. 2020
Distributed systems › distributed communication
remote memory access
0.112020
Network Interface Architecture for Remote Indirect Memory Access (RIMA) in Datacenters · ACM Trans. Archit. Code Optim. 2020
Cloud and datacenter computing › resource provisioning
dynamic resource provisioning
0.112011
Dynamic server provisioning to minimize cost in an IaaS cloud · SIGMETRICS 2011
Cloud and datacenter computing › resource management
resource management and scheduling
0.112011
Dynamic server provisioning to minimize cost in an IaaS cloud · SIGMETRICS 2011
Multimedia analysis and retrieval › audio retrieval
sound event retrieval
0.112010
Segmentation, Indexing, and Retrieval for Environmental and Natural Sounds · IEEE Trans. Speech Audio Process. 2010
Performance modeling and evaluation
workload characterization
0.012011
Dynamic server provisioning to minimize cost in an IaaS cloud · SIGMETRICS 2011

Methods — techniques the papers use, named apart from their topics

in-order flow deflection · 0.4direct apportioning of sending rates · 0.4RDMA · 0.4NIC microarchitecture · 0.4DCQCN · 0.4statistical analysis · 0.1simulation · 0.1spectral clustering · 0.1hidden markov model · 0.1dynamic bayesian network · 0.1
YearPublicationVenuePosition
2020 Network Interface Architecture for Remote Indirect Memory Access (RIMA) in Datacenters
abstract
Remote Direct Memory Access (RDMA) fabrics such as InfiniBand and Converged Ethernet report latency shorter by a factor of 50 than TCP. As such, RDMA is a potential replacement for TCP in datacenters (DCs) running low-latency applications, such as Web search and memcached. InfiniBand’s Shared Receive Queues (SRQs), which use two-sided send/recv verbs (i.e., channel semantics ), reduce the amount of pre-allocated, pinned memory (despite optimizations such as InfiniBand’s on-demand paging (ODP)) for message buffers. However, SRQs are limited fundamentally to a single message size per queue, which incurs either memory wastage or significant programmer burden for typical DC traffic of an arbitrary number (level of burstiness) of messages of arbitrary size. We propose remote indirect memory access (RIMA) , which avoids these pitfalls by providing (1) network interface card (NIC) microarchitecture support for novel queue semantics and (2) a new “verb” called append . To append a sender’s message to a shared queue, the receiver NIC atomically increments the queue’s tail pointer by the incoming message’s size and places the message in the newly created space. As in traditional RDMA, the NIC is responsible for pointer lookup, address translation, and enforcing virtual memory protections. This indirection of specifying a queue (and not its tail pointer, which remains hidden from senders) handles the typical DC traffic of an arbitrary sender sending an arbitrary number of messages of arbitrary size. Because RIMA’s simple hardware adds only 1--2 ns to the multi-\mu s message latency, RIMA achieves the same message latency and throughput as InfiniBand SRQ with unlimited buffering. Running memcached traffic on a 30-node InfiniBand cluster, we show that at similar, low programmer effort, RIMA achieves significantly smaller memory footprint than SRQ. However, while SRQ can be crafted to minimize memory footprint by expending significant programming effort, RIMA provides those benefits with little programmer effort. For memcached traffic, a high-performance key-value cache ( FastKV ) using RIMA achieves either 3× lower 96 th-percentile latency or significantly better throughput or memory footprint than FastKV using RDMA.
Jiachen Xue, T. N. Vijaykumar, Mithuna Thottethodi
ACM Trans. Archit. Code Optim.1
2020 Dart: Divide and Specialize for Fast Response to Congestion in RDMA-Based Datacenter Networks
abstract
Though Remote Direct Memory Access (RDMA) promises to reduce datacenter network latencies significantly compared to TCP (e.g., 10x), end-to-end congestion control in the presence of incasts is a challenge. Targeting the full generality of the congestion problem, previous schemes rely on slow, iterative convergence to the appropriate sending rates (e.g., TIMELY takes 50 RTTs). Several papers have shown that even in oversubscribed datacenter networks most congestion occurs at the receiver. Accordingly, we propose a divide-and-specialize approach, called Dart, which isolates the common case of receiver congestion and further subdivides the remaining in-network congestion into the simpler spatially-localized and the harder spatially-dispersed cases. For receiver congestion, we propose direct apportioning of sending rates (DASR) in which a receiver for n senders directs each sender to cut its rate by a factor of n, converging in only one RTT. For the spatially-localized case, Dart provides fast (under one RTT) response by adding novel switch hardware for in-order flow deflection (IOFD) because RDMA disallows packet reordering on which previous load balancing schemes rely. For the uncommon spatially-dispersed case, Dart falls back to DCQCN. Small-scale testbed measurements and at-scale simulations, respectively, show that Dart achieves 60% (2.5x) and 79% (4.8x) lower 99t'-percentile latency, and similar and 58% higher throughput than InfiniBand, and TIMELY and DCQCN.
Jiachen Xue, Muhammad Usama Chaudhry, Balajee Vamanan, T. N. Vijaykumar, Mithuna Thottethodi
IEEE/ACM Trans. Netw.1
2013 PreTrans: Reducing TLB CAM-search via page number prediction and speculative pre-translation
abstract
The need for fast address translation within tight time constraints (before L1 tag check but after effective address computation) imposes many design constraints. The freedom from such constraints can potentially lead to lower TLB energy costs. In this paper, we observe that (1) data accesses commonly use base-displacement addressing modes in which the effective address is computed as the sum of a base and a displacement, and (2) the effective page numbers are predictable once the base address is known. Further, it is easy to cache address translations alongside the predicted page numbers thus enabling speculative address translation that can filter accesses to the TLB. The two observations enable our PreTrans design in which (a) a speculative translation is available based solely on the base address, and (b) the translation is available simultaneously with the effective (virtual) address. PreTrans replaces most of the energy-expensive CAM-lookups for TLB access with RAM lookups, which translates to significant power improvements in the TLB.
Jiachen Xue, Mithuna Thottethodi
ISLPED1
2012 Selective commitment and selective margin: Techniques to minimize cost in an IaaS cloud
abstract
Cloud computing holds the exciting potential of elastically scaling computation to match time-varying demand, thus eliminating the need to provision for peak demand. However, the uncertainty of variable loads necessitate the use of margins - servers that must be held active to absorb unpredictable potential load bursts - which can be a significant fraction of overall cost. Further, naively switching to an on-demand cloud model can actually degrade true costs (server costs that would be incurred even if margin costs disappeared) because of the fundamental economic rule wherein on-demand services/goods cost more compared to reserved services/goods where the user bears some commitment. On-demand customers pay a premium in exchange for not undertaking the fixed-cost risk that committed customers undertake. This paper addresses the twin challenges of minimizing margin costs and true costs in an Infrastructure-as-a-Service (IaaS) cloud. Our paper makes the following two contributions. First, rather than use a fixed margin, we observe that the margin may be selectively used depending on load levels. Based on the above observation, we develop ShrinkWrap-opt which is a dynamic programming algorithm that achieves optimal margin cost while satisfying the desired (statistical) response time guarantees. Second, we propose commitment straddling - the selective use of some reserved machines in conjunction with on-demand machines - to achieve optimal true-cost. Simulations with real Web server load traces using the Amazon EC2 cost model reveal that our techniques save between 13% and 29% (21% on average) in cost while satisfying response-time targets.
Yu-Ju Hong, Jiachen Xue, Mithuna Thottethodi
ISPASS2
2011 Dynamic server provisioning to minimize cost in an IaaS cloud
abstract
Cloud computing holds the exciting potential of elastically scaling computation to match time-varying demand, thus eliminating the need to provision for peak demand to satisfy response-time requirements. Moreover, cloud vendors often offer several commitment levels for their machine instances (e.g., users can choose to pay an upfront premium for the discounted hourly usage price). Because cost is a major concern that may limit the cloud adoption, two key challenges are to determine (a) the number of machines to provision and (b) the commitment level at which the machine instances should be acquired, to minimize cost while satisfying response-time targets. This paper address the above two challenges in an Infrastructure-as-a-Service (IaaS) cloud. Our simulations with real Web server load traces reveal that our techniques offer a cost reduction between 13% and 29% (21% on average) under Amazon EC2 pricing models.
Yu-Ju Hong, Jiachen Xue, Mithuna Thottethodi
SIGMETRICS2
2010 Segmentation, Indexing, and Retrieval for Environmental and Natural Sounds
abstract
We propose a method for characterizing sound activity in fixed spaces through segmentation, indexing, and retrieval of continuous audio recordings. Regardingsegmentation, we present a dynamic Bayesian network (DBN) that jointly infers onsets and end times of the most prominent sound events in the space, along with an extension of the algorithm for covering large spaces with distributed microphone arrays. Each segmented sound event isindexedwith a hidden Markov model (HMM) that models the distribution of example-based queries that a user would employ toretrievethe event (or similar events). In order to increase the efficiency of the retrieval search, we recursively apply a modified spectral clustering algorithm to group similar sound events based on the distance between their corresponding HMMs. We then conduct a formal user study to obtain the relevancy decisions necessary for evaluation of our retrieval algorithm on both automatically and manually segmented sound clips. Furthermore, our segmentation and retrieval algorithms are shown to be effective in both quiet indoor and noisy outdoor recording conditions.
Gordon Wichern, Jiachen Xue, Harvey D. Thornburg, Brandon Mechtley, Andreas Spanias
IEEE Trans. Speech Audio Process.2
2008 Fast query by example of environmental sounds via robust and efficient cluster-based indexing
abstract
There has been much recent progress in the technical infrastructure necessary to continuously characterize and archive all sounds, or more precisely auditory streams, that occur within a given space or human life. Efficient and intuitive access, however, remains a considerable challenge. In specifically musical domains, i.e., melody retrieval, query-by-example (QBE) has found considerable success in accessing music that matches a specific query. We propose an extension of the QBE paradigm to the broad class of natural and environmental sounds, which occur frequently in continuous recordings. We explore several cluster-based indexing approaches, namely non-negative matrix factorization (NMF) and spectral clustering to efficiently organize and quickly retrieve archived audio using the QBE paradigm. Experiments on a test database compare the performance of the different clustering algorithms in terms of recall, precision, and computational complexity. Initial results indicate significant improvements over both exhaustive search schemes and traditional K- means clustering, and excellent overall performance in the example-based retrieval of environmental sounds.
Jiachen Xue, Gordon Wichern, Harvey D. Thornburg, Andreas Spanias
ICASSP1