EDBT 2026 Demo / reviewers in the wild / expert
Weiyang Wang
dblp:145/9134
· DBLP profile ↗
14ranked-venue papers
4as first author
9since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 8 · 1 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 1 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Checkmate: Zero Performance Overhead Model Checkpointing via Network Gradient Replication
Ankit Bhardwaj 0002, Weiyang Wang, Jeremy Carin, Adam Belay, Manya Ghobadi |
NSDI | 2 |
| 2026 | MonkeyTree: Near-Minimal Congestion for Multi-tenant Training via MigrationabstractWe present MonkeyTree, the first system to mitigate network congestion in multi-tenant GPU clusters through job-migration based defragmentation rather than network-layer techniques. As cloud operators co-locate ML training jobs on shared, oversubscribed networks, congestion degrades training throughput for over a third of jobs. Prior approaches either rely on routing and flow scheduling—which we show have fundamental limits when traffic exceeds capacity, or require costly full-bisection bandwidth topologies with packet spraying. Anton A. Zabreyko, Weiyang Wang, Manya Ghobadi |
SIGCOMM | 2 |
| 2026 | Multimodal Dynamic Cost Matrix Adaptation under data imbalance for multimodal sentiment analysis
Weiyang Wang |
Pattern Recognit. Lett. | 1 |
| 2025 | Efficient Direct-Connect Topologies for Collective Communications
Liangyu Zhao, Siddharth Pal, Tapan Chugh, Weiyang Wang, Jason Fantl, Prithwish Basu, Joud Khoury, Arvind Krishnamurthy |
NSDI | 4 |
| 2024 | Rail-only: A Low-Cost High-Performance Network for Training LLMs with Trillion ParametersabstractThis paper presents a low-cost network architecture for training large language models (LLMs) at hyperscale. We study the optimal parallelization strategy of LLMs and propose a novel datacenter network design tailored to LLM's unique communication pattern. We show that LLM training generates sparse communication patterns in the network and, therefore, does not require any-to-any full-bisection network to complete efficiently. As a result, our design eliminates the spine layer in traditional GPU clusters. We name this design a Rail-only network and demonstrate that it achieves the same training performance while reducing the network cost by 38% to 77% and network power consumption by 37% to 75% compared to a conventional GPU datacenter. Our architecture also supports Mixture-of-Expert (MoE) models with all-to-all communication through forwarding, with only 4.1% to 5.6% completion time overhead for all-to-all traffic. We study the failure robustness of Rail-only networks and provide insights into the performance impact of different network and training parameters. Weiyang Wang, Manya Ghobadi, Kayvon Shakeri, Ying Zhang 0022, Naader Hasani |
HOTI | 1 |
| 2024 | Efficient ship detection in sar images with dynamic feature smoothing and visual module using omni-dimensional dynamic large-scale convolution
Weiyang Wang, Huachun Zhang, Anlin Xu |
Multim. Tools Appl. | 1 |
| 2023 | TopoOpt: Co-optimizing Network Topology and Parallelization Strategy for Distributed Training Jobs
Weiyang Wang, Moein Khazraee, Zhizhen Zhong, Manya Ghobadi, Dheevatsa Mudigere, Ying Zhang 0022, Anthony Kewitsch |
NSDI | 1 |
| 2023 | Category Alignment Adversarial Learning for Cross-Modal RetrievalabstractCross-modal retrieval aims to retrieve one semantically similar media from multiple media types based on queries entered by another type of media. An intuitive idea is to map different media data into a common space and then directly measure content similarity between different types of data. In this paper, we present a novel method, called Category Alignment Adversarial Learning (CAAL) for cross-modal retrieval. It aims to find a common representation space supervised by category information, in which the samples from different modalities can be compared directly. Specifically, CAAL firstly employs two parallel encoders to generate common representations for image and text features respectively. Furthermore, we employ two parallel GANs with category information to generate fake image and text features which next will be utilized with already generated embedding to reconstruct the common representation. At last, two joint discriminators are utilized to reduce the gap between the mapping of the first stage and the embedding of the second stage. Comprehensive experimental results on four widely-used benchmark datasets demonstrate the superior performance of our proposed method compared with the state-of-the-art approaches. Shiyuan He, Weiyang Wang, Zheng Wang 0044, Xing Xu 0001, Yang Yang 0002, Heng Tao Shen |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2022 | Time-division TCP for reconfigurable data center networksabstractRecent proposals for reconfigurable data center networks have shown that providing multiple time-varying paths can improve network capacity and lower physical latency. However, existing TCP variants are ill-suited to utilize available capacity because their congestion control cannot react quickly enough to drastic variations in bandwidth and latency. Shawn Shuoshuo Chen, Weiyang Wang, Christopher Canel, Srinivasan Seshan, Alex C. Snoeren, Peter Steenkiste |
SIGCOMM | 2 |
| 2020 | Adapting TCP for Reconfigurable Datacenter Networks
Matthew K. Mukerjee, Christopher Canel, Weiyang Wang, Daehyeok Kim, Srinivasan Seshan, Alex C. Snoeren |
NSDI | 3 |
| 2019 | Foxnet: A Multi-Face Alignment MethodabstractMulti-face alignment aims to identify geometry structures of multiple faces in an image, and its performance is essential for the many practical tasks, such as face recognition, face tracking, and face animation. In this work, we present a fast bottom-up multi-face alignment approach, which can simultaneously localize multi-person facial landmarks with high precision. In more detail, our bottom-up architecture maps the landmarks to the high-dimensional space with which landmarks of all faces are represented. By clustering the features belonging to the same face, our approach can align the multi-person facial landmarks synchronously. Extensive experiments show that our method can achieve high performance in the multi-face landmark alignment task while our model is extremely fast. Moreover, we propose a new multi-face dataset to compare the speed and precision of bottom-up face alignment method with top-down methods. Our dataset is publicly available at1. Yuxiang Wu, Zehua Cheng, Weiyang Wang |
ICIP | 6 |
| 2016 | Per-Line Power Controlled Lattice-Reduction Aided Zero-Forcing Precoding for G.fast DownstreamabstractLinear precoding for Far-End CrossTalk (FEXT) cancellation is supported by the first release of the next generation DSL standard G.fast utilizing bandwidth up to 106 MHz. The cancellation precoding for the second release (bandwidth up to 212 MHz with much stronger FEXT) has not been standardized yet. Non-linear Tomlinson-Harashima Precoding (THP) is a promising candidate as it outperforms the linear precoding. Nevertheless, sequential processing of THP introduces unfavorable additional delay proportional to the number of lines. In this paper, we focus on non-linear precoding based on Lattice Reduction aided Zero-Forcing (ZF-LR) which can be implemented in parallel without any sequential processing delay. ZF-LR precoding has been proposed in wireless scenario using scalar scaling to adjust average transmit power. We modify this scheme to comply with per-line per-carrier power constraint and enable power control of each line individually (e.g., to adapt power for different line lengths). Optimized Power Allocation (OPA) for the modified scheme maximizing weighted sum-rate leads to the signomial optimization problem. Similarly to THP, ZF-LR precoding with OPA outperforms the linear precoding as well. Miroslav Hekrdla, Andrea Matera, Umberto Spagnolini, Weiyang Wang |
GLOBECOM | 4 |
| 2015 | Ordered Tomlinson-Harashima Precoding in G.fast DownstreamabstractG.fast is an upcoming next generation DSL standard envisioned to use bandwidth up to 212 MHz. Far-end crosstalk (FEXT) at these frequencies greatly overcomes direct links. Its cancellation based on non-linear Tomlinson-Harashima Precoding (THP) proved to show significant advantage over standard linear precoding. This paper proposes a novel THP structure in which ordering of successive interference pre-cancellation can be optimized for downstream with non-cooperating receivers. The optimized scheme is compared to existing THP structure denoted as equal-rate THP which is widely adopted in wireless downlink. Structure and performance of both methods differ significantly favoring the proposed scheme. The ordering that maximizes the minimum rate (max-min fairness) for each tone of the discrete multi-tone modulation is the familiar V-BLAST ordering. However, V- BLAST does not lead to the global maximum when applied independently on each tone. The proposed novel Dynamic Ordering (DO) strategy takes into account asymmetric channel statistics to yield the highest minimum aggregated rate. Miroslav Hekrdla, Andrea Matera, Weiyang Wang, Dong Wei 0013, Umberto Spagnolini |
GLOBECOM | 3 |
| 2015 | VOBC Data Storage and Online Diagnosing System Based on Data CloudabstractDue to the inaccessibility to VOBC (Vehicle Onboard Controller) diagnosing data, urban railway transit operation suffers from transient failures and long time failures during operation time. This essay put forward the idea of a real time diagnosing data storage and cloud sharing system, to collect, store and sharing them, thus reproduce the transient failures for debugging, and proved a quick and effective solution to failures during operation time from a distance. Youqi Gu, Xiaoqing Zeng, Tuo Shen, Weiyang Wang |
ISADS | 4 |