VLDB 2026 Research / reviewers in the wild / expert
Angela Wang
dblp:77/4883
· DBLP profile ↗
9ranked-venue papers
2as first author
1since 2021 · last 2024
—ORCID · unresolved
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 3 · 1 first-authorSystems, architecture and hardware · 2 · 1 since 2021Security and privacy · 2 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1Human-computer interaction and ubiquitous computing · 1Applied, interdisciplinary, general and emerging computing · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
2 papers |
Hardware accelerators and domain-specific architectures · 34% Memory systems · 34% Reconfigurable computing and FPGAs · 10% | |
| Computer networks
2 papers |
Internet architecture and protocols · 61% Network performance modeling · 21% Network management and operations · 18% |
Topics — the 7 heaviest of 10, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Hardware accelerators and domain-specific architectures
machine learning accelerator |
0.8 | 1 | 2024 | SambaNova SN40L: Scaling the AI Memory Wall with Dataflow and Composition of Experts · MICRO 2024 |
Memory systems
memory hierarchy |
0.8 | 1 | 2024 | SambaNova SN40L: Scaling the AI Memory Wall with Dataflow and Composition of Experts · MICRO 2024 |
Reconfigurable computing and FPGAs › reconfigurable computing
reconfigurable dataflow |
0.2 | 1 | 2024 | SambaNova SN40L: Scaling the AI Memory Wall with Dataflow and Composition of Experts · MICRO 2024 |
Cloud and datacenter computing › cloud service management
cloud service deployment |
0.1 | 1 | 2011 | Estimating the performance of hypothetical cloud service deployments: A measurement-based approach · INFOCOM 2011 |
Performance modeling and evaluation
performance prediction |
0.1 | 1 | 2011 | Estimating the performance of hypothetical cloud service deployments: A measurement-based approach · INFOCOM 2011 |
Internet architecture and protocols
domain name system |
0.1 | 1 | 2010 | A DNS Reflection Method for Global Traffic Management · USENIX ATC 2010 |
Network performance modeling › performance prediction
measurement-based prediction |
0.0 | 1 | 2011 | Estimating the performance of hypothetical cloud service deployments: A measurement-based approach · INFOCOM 2011 |
Methods — techniques the papers use, named apart from their topics
streaming dataflow · 0.8expert composition · 0.8speedtest · 0.2active web content · 0.2CDN infrastructure · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | SambaNova SN40L: Scaling the AI Memory Wall with Dataflow and Composition of ExpertsabstractMonolithic large language models (LLMs) like GPT-4 have paved the way for modern generative AI applications. Training, serving, and maintaining monolithic LLMs at scale, however, remains prohibitively expensive and challenging. The disproportionate increase in compute-to-memory ratio of modern AI accelerators have created a memory wall, necessitating new methods to deploy AI. Recent research has shown that a composition of many smaller expert models, each with several orders of magnitude fewer parameters, can match or exceed the capabilities of monolithic LLMs. Composition of Experts (CoE) is a modular approach that lowers the cost and complexity of training and serving. However, this approach presents two key challenges when using conventional hardware: (1) without fused operations, smaller models have lower operational intensity, which makes high utilization more challenging to achieve; and (2) hosting a large number of models can be either prohibitively expensive or slow when dynamically switching between them. In this paper, we describe how combining CoE, streaming dataflow, and a three-tier memory system scales the AI memory wall. We describe Samba-CoE, a CoE system with 150 experts and a trillion total parameters. We deploy Samba-CoE on the SambaNova SN40L Reconfigurable Dataflow Unit (RDU) -a commercial dataflow accelerator architecture that has been codesigned for enterprise inference and training applications. The chip introduces a new three-tier memory system with on-chip distributed SRAM, on-package HBM, and off-package DDR DRAM. A dedicated inter-RDU network enables scaling up and out over multiple sockets. We demonstrate speedups ranging from 2× to 13× on various benchmarks running on eight RDU sockets compared with an unfused baseline. We show that for CoE inference deployments, the 8-socket RDU Node reduces machine footprint by up to 19 ×, speeds up model switching time by 15× to 31×, and achieves an overall speedup of 3.7× over a DGX H100 and 6.6× over a DGX A100. Raghu Prabhakar, Ram Sivaramakrishnan, Darshan Gandhi, Mingran Wang, Kejie Zhang, Tianren Gao, Angela Wang, Yongning Sheng, Joshua Brot, Denis Sokolov, Apurv Vivek, Calvin Leung, Arjun Sabnis, Jiayu Bai, Tuowen Zhao, Mark Gottscho, Mark Luttrell, Manish K. Shah, Zhengyu Chen 0002, Kaizhao Liang, Swayambhoo Jain, Urmish Thakker, Dawei Huang, Sumti Jairath, Kevin J. Brown, Kunle Olukotun |
MICRO | 9 |
| 2019 | Post-Processing of Word Representations via Variance Normalization and Dynamic EmbeddingabstractLanguage processing becomes more and more important in multimedia processing. Although embedded vector representations of words offer impressive performance on many natural language processing (NLP) applications, the information of ordered input sequences is lost to some extent if only context-based samples are used in the training. For further performance improvement, two new post-processing techniques, called post-processing via variance normalization (PVN) and post-processing via dynamic embedding (PDE), are proposed in this work. The PVN method normalizes the variance of principal components of word vectors, while the PDE method learns orthogonal latent variables from ordered input sequences. The PVN and the PDE methods can be integrated to achieve better performance. We apply these post-processing techniques to several popular word embedding methods to yield their post-processed representations. Extensive experiments are conducted to demonstrate the effectiveness of the proposed post-processing techniques. Fenxiao Chen, Angela Wang, C.-C. Jay Kuo |
ICME | 2 |
| 2011 | Estimating the performance of hypothetical cloud service deployments: A measurement-based approachabstractTo optimize network performance, cloud service providers have a number of options available to them, including co-locating production servers in well-connected Internet eXchange (IX) points, deploying data centers in additional locations, or contracting with external Content Distribution Networks (CDNs). Some of these options can be very costly, and some may or may not improve performance significantly. Cloud service providers would clearly like to be able to estimate a priori performance gain of the various options before sinking significant capital expenditures into major infrastructure changes. In this paper we take a measurement-oriented approach and develop methodologies that accurately predict the performance improvement for making major infrastructure changes. Our methodologies leverage active web content, existing large-scale CDN infrastructures, and the SpeedTest network. We then apply our methodologies and a CloudBeacon tool to the problem of locating satellite data centers throughout the world. The results show that for North America, a deployment limited to 11 locations will be sufficient. However, in order to provide good latency and throughput performance on a global scale, somewhere between a total of 36 and 72 cloud-service locations with good peering connections is most likely needed. Angela Wang, Cheng Huang 0002, Jin Li 0001, Keith W. Ross |
INFOCOM | 1 |
| 2010 | Measuring and Evaluating TCP Splitting for Cloud Services
Abhinav Pathak, Angela Wang, Cheng Huang 0002, Albert G. Greenberg, Y. Charlie Hu, Randy Kern, Jin Li 0001, Keith W. Ross |
PAM | 2 |
| 2010 | A DNS Reflection Method for Global Traffic Management
Cheng Huang 0002, Nic Holt, Angela Wang, Albert G. Greenberg, Jin Li 0001, Keith W. Ross |
USENIX ATC | 3 |
| 2009 | Queen: Estimating Packet Loss Rate between Arbitrary Internet Hosts
Angela Wang, Cheng Huang 0002, Jin Li 0001, Keith W. Ross |
PAM | 1 |
| 2008 | Understanding hybrid CDN-P2P: why limelight needs its own Red SwooshabstractIn this paper, we quantify the potential gains of hybrid CDN-P2P for two of the leading CDN companies, Akamai and Limelight. We first develop novel measurement methodology for mapping the topologies of CDN networks. We then consider ISP-friendly P2P distribution schemes which work in conjunction with the CDNs to localize traffic within regions of ISPs. To evaluate these schemes, we use two recent, real-world traces: a video-on-demand trace and a large-scale software update trace. We find that hybrid CDN-P2P can significantly reduce the cost of content distribution, even when peer sharing is localized within ISPs and further localized within regions of ISPs. We conclude that hybrid CDN-P2P distribution can economically satisfy the exponential growth of Internet video content without placing an unacceptable burden on regional ISPs. Cheng Huang 0002, Angela Wang, Jin Li 0001, Keith W. Ross |
NOSSDAV | 2 |
| 2007 | Discovery of In-Band Streaming Services in Peer-to-Peer OverlaysabstractPeer-to-peer overlays can be used for service discovery over a global network fabric. We describe and evaluate a new service indexing mechanism for in-band streaming services such as application relays, mixers, and media transcoders. For this type of service, the location of the service in the network and service admission status are key attributes. We describe and analyze a service indexing mechanism which uses network position-based advertisement. We show that this mechanism gives good service selection, provides a close to uniform advertisement distribution in the overlay, reduces message overhead, and exhibits acceptable stability and setup delay. John F. Buford, Angela Wang, Xiaojun Hei, Yong Liu 0013, Keith W. Ross |
GLOBECOM | 2 |
| 1989 | The role of the bipolar cell in retinal signal processingabstractBipolar cells in the retina of the tiger salamander were studied by intracellular recording. Their receptive fields consisted of a central region and a surround which, when stimulated, antagonized the response of the center. The spatial distributions of the center and surround components were determined, using drugs that selectively eliminated the surround. The receptive field is shown to be a linear combination of the two components. Because of the spatial antagonism the receptive field acts as a bandpass spatial filter. The importance of this filtering in photon detection is discussed.> W. Geoffrey Owen, William A. Hare, Angela Wang |
SMC | 3 |