VLDB 2026 Research / reviewers in the wild / expert
Fan Yang 0024
dblp:29/3081-24
· DBLP profile ↗
74ranked-venue papers
3as first author
50since 2021 · last 2026
0000-0002-0378-060XORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 25 · 20 since 2021Systems, architecture and hardware · 20 · 13 since 2021Artificial intelligence and machine learning · 18 · 17 since 2021Computer networks · 9 · 2 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 1 first-author · 5 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MetaAttention: A Unified and Performant Attention Framework across Hardware BackendsabstractComputing attention is the backbone of transformer-based models like large language models. However, the increasing diversity of attention algorithms presents significant challenges for unleashing hardware performance. State-of-the-art variants like FlashAttention target a specific attention algorithm or hardware platform, which fail to generalize to other algorithms and platforms. Yu Cheng 0030, Lei Wang 0222, Yuqing Xia, Ziming Miao, Lingxiao Ma, Fan Yang 0024, Jilong Xue, Zhi Yang 0001, Mao Yang 0004, Xingda Wei, Haibo Chen 0001 |
PPoPP | 7 |
| 2026 | RetroInfer: A Vector Storage Engine for Scalable Long-Context LLM Inference
Yaoqi Chen, Jinkai Zhang, Baotong Lu, Qianxi Zhang, Chengruidong Zhang, Jingjia Luo, Huiqiang Jiang, Qi Chen 0009, Bailu Ding, Xiao Yan 0002, Jiawei Jiang 0001, Chen Chen 0067, Cheng Li 0001, Yuqing Yang 0001, Fan Yang 0024, Mao Yang 0004 |
Proc. VLDB Endow. | 18 |
| 2025 | LUT-DLA: Lookup Table as Efficient Extreme Low-Bit Deep Learning AcceleratorabstractThe emergence of neural network capabilities invariably leads to a significant surge in computational demands due to expanding model sizes and increased computational complexity. To reduce model size and lower inference costs, recent research has focused on simplifying models and designing hardware accelerators using low-bit quantization. However, due to numerical representation limits, scalar quantization cannot reduce bit width lower than 1-bit, diminishing its benefits. To break through these limitations, we introduce LUT-DLA, a Look-Up Table (LUT) Deep Learning Accelerator Framework that utilizes vector quantization to convert neural network models into LUTs, achieving extreme low-bit quantization. The LUT-DLA framework facilitates efficient and cost-effective hardware accelerator designs and supports the LUTBoost algorithm, which helps to transform various DNN models into LUT-based models via multistage training, drastically cutting both computational and hardware overhead. Additionally, through co-design space exploration, LUT-DLA assesses the impact of various model and hardware parameters to fine-tune hardware configurations for different application scenarios, optimizing performance and efficiency. Our comprehensive experiments show that LUT-DLA achieves improvements in power efficiency and area efficiency with gains of 1.4~7.0× and 1.5~146.1×, respectively, while maintaining only a modest accuracy drop. For CNNs, accuracy decreases by 0.1%~3.1% using the L2distance similarity, 0.1%~3.4% with the L1distance similarity, and 0.1%~3.8% when employing the Chebyshev distance similarity. For transformer-based models, the accuracy drop ranges from 1.4% to 3.0%. Shengyu Ye, Chunyun Chen, Yang Wang 0053, Fan Yang 0024, Ting Cao 0003, Cheng Liu 0008, Mohamed M. Sabry, Mao Yang 0004 |
HPCA | 5 |
| 2025 | Automated Proof Generation for Rust Code via Self-EvolutionabstractEnsuring correctness is crucial for code generation. Formal verification offers a
definitive assurance of correctness, but demands substantial human effort in proof
construction and hence raises a pressing need for automation. The primary obsta-
cle lies in the severe lack of data—there is much fewer proofs than code snippets
for Large Language Models (LLMs) to train upon. In this paper, we introduce
SAFE, a framework that overcomes the lack of human-written proofs to enable
automated proof generation of Rust code. SAFE establishes a self-evolving cycle
where data synthesis and fine-tuning collaborate to enhance the model capability,
leveraging the definitive power of a symbolic verifier in telling correct proofs from
incorrect ones. SAFE also re-purposes the large number of synthesized incorrect
proofs to train the self-debugging capability of the fine-tuned models, empowering
them to fix incorrect proofs based on the verifier’s feedback. SAFE demonstrates
superior efficiency and precision compared to GPT-4o. Through tens of thousands
of synthesized proofs and the self-debugging mechanism, we improve the capa-
bility of open-source models, initially unacquainted with formal verification, to
automatically write proofs for Rust code. This advancement leads to a signifi-
cant improvement in performance, achieving a 52.52% accuracy rate in a bench-
mark crafted by human experts, a significant leap over GPT-4o’s performance of
14.39%. Tianyu Chen 0006, Shan Lu 0001, Yeyun Gong, Chenyuan Yang, Xuheng Li, Md Rakib Hossain Misu, Hao Yu 0016, Nan Duan 0001, Peng Cheng 0005, Fan Yang 0024, Shuvendu K. Lahiri, Tao Xie 0001, Lidong Zhou |
ICLR | 11 |
| 2025 | Mutual Reasoning Makes Smaller LLMs Stronger Problem-SolverabstractThis paper introduces rStar, a self-play mutual reasoning approach that significantly improves reasoning capabilities of small language models (SLMs) without fine-tuning or superior models. rStar decouples reasoning into a self-play mutual generation-discrimination process. First, a target SLM augments the Monte Carlo Tree Search (MCTS) with a rich set of human-like reasoning actions to construct higher quality reasoning trajectories. Next, another SLM, with capabilities similar to the target SLM, acts as a discriminator to verify each trajectory generated by the target SLM. The mutually agreed reasoning trajectories are considered mutual consistent, thus are more likely to be correct. Extensive experiments across five SLMs demonstrate rStar can effectively solve diverse reasoning problems, including GSM8K, GSM-Hard, MATH, SVAMP, and StrategyQA. Remarkably, rStar boosts GSM8K accuracy from 12.51\% to 63.91\% for LLaMA2-7B, from 36.46\% to 81.88\% for Mistral-7B, from 74.53\% to 91.13\% for LLaMA3-8B-Instruct. Code is available at https://github.com/zhentingqi/rStar. Zhenting Qi, Mingyuan Ma, Jiahang Xu, Li Lyna Zhang, Fan Yang 0024, Mao Yang 0004 |
ICLR | 5 |
| 2025 | rStar-Math: Small LLMs Can Master Math Reasoning with Self-Evolved Deep ThinkingabstractWe present rStar-Math to demonstrate that small language models (SLMs) can rival or even surpass the math reasoning capability of OpenAI o1, without distillation from superior models. rStar-Math achieves this by exercising “deep thinking” through Monte Carlo Tree Search (MCTS), where a math policy SLM performs test-time search guided by an SLM-based process reward model. rStar-Math introduces three innovations to tackle the challenges in training the two SLMs: (1) a novel code-augmented CoT data synthesis method, which performs extensive MCTS rollouts to generate step-by-step verified reasoning trajectories used to train the policy SLM; (2) a novel process reward model training method that avoids naïve step-level score annotation, yielding a more effective process preference model (PPM); (3) a self-evolution recipe in which the policy SLM and PPM are built from scratch and iteratively evolved to improve reasoning capabilities. Through 4 rounds of self-evolution with millions of synthesized solutions for 747k math problems, rStar-Math boosts SLMs’ math reasoning to state-of-the-art levels. On MATH benchmark, it improves Qwen2.5-Math-7B from 58.8% to 90.0%, surpassing o1-preview by +4.5%. On the USA Math Olympiad (AIME), rStar-Math solves an average of 53.3% (8/15) of problems, ranking among the top 20% of the brightest high school math students. Code and data are available at https://github.com/microsoft/rStar. Li Lyna Zhang, Youran Sun, Fan Yang 0024, Mao Yang 0004 |
ICML | 7 |
| 2025 | LongRoPE2: Near-Lossless LLM Context Window ScalingabstractLongRoPE2 is a novel approach that extends the effective context window of pre-trained large language models (LLMs) to the target length, while preserving the performance on the original shorter context window. This is achieved by three contributions: (1) a hypothesis that insufficient training in higher RoPE dimensions contributes to the persistent out-of-distribution (OOD) issues observed in existing methods; (2) an effective RoPE rescaling algorithm that adopts evolutionary search guided by "needle-driven" perplexity to address the insufficient training problem; (3) a mixed context window training approach that fine-tunes model weights to adopt rescaled RoPE for long-context sequences while preserving the short-context performance with the original RoPE. Extensive experiments on LLaMA3-8B and Phi3-mini-3.8B across various benchmarks validate the hypothesis and demonstrate the effectiveness of LongRoPE2. Remarkably, LongRoPE2 extends LLaMA3-8B to achieve a 128K effective context length while retaining over 98.5% of short-context performance, using only 10B tokens – 80x fewer than Meta’s approach, which fails to reach the target effective context length. Li Lyna Zhang, Siyuan Wang 0003, Gaokai Zhang, Gilsinia Lopez, Fan Yang 0024, Weizhu Chen, Mao Yang 0004 |
ICML | 6 |
| 2025 | LUT Tensor Core: A Software-Hardware Co-Design for LUT-Based Low-Bit LLM InferenceabstractLarge Language Model (LLM) inference becomes resource-intensive, prompting a shift toward low-bit model weights to reduce the memory footprint and improve efficiency.Such low-bit LLMs necessitate the mixed-precision matrix multiplication (mpGEMM), an important yet underexplored operation involving the multiplication of lower-precision weights with higher-precision activations.Off-theshelf hardware does not support this operation natively, leading to indirect, thus inefficient, dequantization-based implementations.In this paper, we study the lookup table (LUT)-based approach for mpGEMM and find that a conventional LUT implementation fails to achieve the promised gains.To unlock the full potential of LUT-based mpGEMM, we propose LUT Tensor Core, a softwarehardware co-design for low-bit LLM inference.LUT Tensor Core differentiates itself from conventional LUT designs through: 1) * Work is done during internship at Microsoft Research. Zhiwen Mo, Lei Wang 0222, Jianyu Wei, Zhichen Zeng 0002, Shijie Cao, Lingxiao Ma, Naifeng Jing, Ting Cao 0003, Jilong Xue, Fan Yang 0024, Mao Yang 0004 |
ISCA | 10 |
| 2025 | SeerAttention: Self-distilled Attention Gating for Efficient Long-context PrefillingabstractAttention is the cornerstone of modern Large Language Models (LLMs). Yet its quadratic complexity hinders efficiency and scalability, especially for long-context processing. A promising approach is to leverage sparsity in attention. However, existing sparsity-based solutions predominantly rely on predefined patterns or heuristics at the attention head level, struggling to adapt dynamically to different contexts efficiently. We propose SeerAttention, a simple yet effective attention mechanism that directly learns the block-level attention sparsity from the LLM itself. Inspired by the gating mechanism in Mixture of Experts (MoE), SeerAttention augments the conventional attention with a **learnable gate** that **selectively activates important blocks** within the attention map. Specifically, the gate first pools the query (Q) and key (K) tensors along the sequence dimension and processes them through learnable linear layers. The resulting matrices are then multiplied together to produce the gating scores, which are used to predict block-level attention sparsity. Combined with our block-sparse FlashAttention kernel, SeerAttention can achieve significant speedup on GPUs. When applied to pre-trained LLMs, SeerAttention only requires training the gate parameters in a lightweight self-distillation manner, allowing rapid convergence. Our evaluation results demonstrate that SeerAttention achieves better model accuracy and lower latency for long-context pre-filling compared to prior methods. Code is available at: https://github.com/microsoft/SeerAttention. Yizhao Gao 0002, Zhichen Zeng 0002, Dayou Du, Shijie Cao, Peiyuan Zhou, Jiaxing Qi, Junjie Lai, Hayden Kwok-Hay So, Ting Cao 0003, Fan Yang 0024, Mao Yang 0004 |
NeurIPS | 10 |
| 2025 | RetrievalAttention: Accelerating Long-Context LLM Inference via Vector RetrievalabstractTransformer-based Large Language Models (LLMs) have become increasingly important. However, scaling LLMs to longer contexts incurs slow inference speed and high GPU memory consumption for caching key-value (KV) vectors. This paper presents RetrievalAttention, a training-free approach to both accelerate the decoding phase and reduce GPU memory consumption by pre-building KV vector indexes for fixed contexts and maintaining them in CPU memory for efficient retrieval. Unlike conventional KV cache methods, RetrievalAttention integrate approximate nearest neighbor search (ANNS) indexes into attention computation. We observe that off-the-shelf ANNS techniques often fail due to the out-of-distribution (OOD) nature of query and key vectors in attention mechanisms. RetrievalAttention overcomes this with an attention-aware vector index. Our evaluation shows RetrievalAttention achieves near full attention accuracy while accessing only 1-3\% of the data, significantly reducing inference costs. Remarkably, RetrievalAttention enables LLMs with 8B parameters to handle 128K tokens on a single NVIDIA RTX4090 (24GB), achieving a decoding speed of 0.107 seconds per token. Baotong Lu, Huiqiang Jiang, Zhenhua Han, Qianxi Zhang, Qi Chen 0009, Chengruidong Zhang, Bailu Ding, Chen Chen 0067, Fan Yang 0024, Yuqing Yang 0001, Lili Qiu |
NeurIPS | 12 |
| 2025 | rStar-Coder: Scaling Competitive Code Reasoning with a Large-Scale Verified DatasetabstractAdvancing code reasoning in large language models (LLMs) is fundamentally limited by the scarcity of high-difficulty datasets, especially those with verifiable input-output test cases necessary for rigorous solution validation at scale. We introduce rStar-Coder, which significantly improves LLM code reasoning capabilities by constructing a large-scale, verified dataset of 418K competition-level code problems, 580K long-reasoning solutions along with rich test cases of varying difficulty. This is achieved through three core contributions: (1) we curate competitive programming code problems and solutions to synthesize new, solvable problems; (2) we introduce a reliable input-output test case synthesis pipeline that decouples the generation into a three-step input generation method and a mutual verification mechanism for effective output labeling; (3) we augment problems with high-quality, test-case-verified long-reasoning solutions. Extensive experiments on Qwen models (1.5B-14B) across various code reasoning benchmarks demonstrate the superiority of rStar-Coder dataset, achieving leading performance comparable to frontier reasoning LLMs with significantly smaller model sizes. On LiveCodeBench, rStar-Coder improves Qwen2.5-7B from 17.4% to an impressive 57.3%, and Qwen2.5-14B from 23.3% to 62.5%, surpassing o3-mini (low) by 3.1%. On the more challenging USA Computing Olympiad, our 7B model achieves an average pass@1 accuracy of 16.15%, outperforming the frontier-level QWQ-32B. rStar-Coder dataset is publicly available at https://huggingface.co/datasets/microsoft/rStar-Coder. Li Lyna Zhang, Bingcheng Dong, Fan Yang 0024, Cheng Li 0001, Mao Yang 0004 |
NeurIPS | 7 |
| 2025 | PipeThreader: Software-Defined Pipelining for Efficient DNN Execution
Yu Cheng 0030, Lei Wang 0222, Yining Shi 0001, Yuqing Xia, Lingxiao Ma, Jilong Xue, Yang Wang 0053, Zhiwen Mo, Fan Yang 0024, Mao Yang 0004, Zhi Yang 0001 |
OSDI | 10 |
| 2025 | WaferLLM: Large Language Model Inference at Wafer Scale
Congjie He, Yeqi Huang, Pei Mu 0003, Ziming Miao, Jilong Xue, Lingxiao Ma, Fan Yang 0024, Luo Mai |
OSDI | 7 |
| 2025 | TrainVerify: Equivalence-Based Verification for Distributed LLM TrainingabstractTraining large language models (LLMs) at scale requires parallel execution across thousands of devices, incurring enormous computational costs. Yet, these costly distributed trainings are prone to correctness bugs, causing silent errors and potentially wasting millions of GPU hours. These bugs are challenging to expose through testing. Yunchi Lu, Youshan Miao, Cheng Tan 0005, Peng Huang 0005, Xian Zhang 0001, Fan Yang 0024 |
SOSP | 7 |
| 2025 | AutoVerus: Automated Proof Generation for Rust CodeabstractGenerative AI has shown its value for many software engineering tasks. Still in its infancy, large language model (LLM)-based proof generation lags behind LLM-based code generation. In this paper, we present A uto V erus . A uto V erus uses LLMs to automatically generate correctness proof for Rust code. A uto V erus is designed to match the unique features of Verus, a verification tool that can prove the correctness of Rust code using proofs and specifications also written in Rust. A uto V erus consists of a network of agents that are crafted and orchestrated to mimic human experts’ three phases of proof construction: preliminary proof generation, proof refinement guided by generic tips, and proof debugging guided by verification errors. To thoroughly evaluate A uto V erus and help foster future research in this direction, we have built a benchmark suite of 150 non-trivial proof tasks, based on existing code-generation benchmarks and verification benchmarks. Our evaluation shows that A uto V erus can automatically generate correct proof for more than 90% of them, with more than half of them tackled in less than 30 seconds or 3 LLM calls. Chenyuan Yang, Xuheng Li, Md Rakib Hossain Misu, Jianan Yao, Weidong Cui, Yeyun Gong, Chris Hawblitzel, Shuvendu K. Lahiri, Jacob R. Lorch, Fan Yang 0024, Ziqiao Zhou, Shan Lu 0001 |
Proc. ACM Program. Lang. | 11 |
| 2024 | Amanda: Unified Instrumentation Framework for Deep Neural NetworksabstractThe success of deep neural networks (DNNs) has sparked efforts to analyze (e.g., tracing) and optimize (e.g., pruning) them. These tasks have specific requirements and ad-hoc implementations in current execution backends like TensorFlow/PyTorch, which require developers to manage fragmented interfaces and adapt their codes to diverse models. In this study, we propose a new framework called Amanda to streamline the development of these tasks. We formalize the implementation of these tasks as neural network instrumentation, which involves introducing instrumentation into the operator level of DNNs. This allows us to abstract DNN analysis and optimization tasks as instrumentation tools on various DNN models. We build Amanda with two levels of APIs to achieve a unified, extensible, and efficient instrumentation design. The user-level API provides a unified operator-grained instrumentation API for different backends. Meanwhile, internally, we design a set of callback-centric APIs for managing and optimizing the execution of original and instrumentation codes in different backends. Through these design principles, the Amanda framework can accommodate a broad spectrum of use cases, such as tracing, profiling, pruning, and quantization, across different backends (e.g., TensorFlow/PyTorch) and execution modes (graph/eager mode). Moreover, our efficient execution management ensures that the performance overhead is typically kept within 5%. Yue Guan 0003, Yuxian Qiu, Jingwen Leng, Fan Yang 0024, Shuo Yu 0006, Yunxin Liu 0001, Yu Feng 0007, Yuhao Zhu 0001, Lidong Zhou, Yun Liang 0001, Chen Zhang 0001, Chao Li 0009, Minyi Guo |
ASPLOS (1) | 4 |
| 2024 | IRGen: Generative Modeling for Image Retrieval
Ting Zhang 0002, Dong Chen 0003, Yujing Wang 0002, Qi Chen 0009, Xing Xie 0001, Hao Sun 0015, Qi Zhang 0066, Fan Yang 0024, Mao Yang 0004, Qingmin Liao, Jingdong Wang 0001, Baining Guo |
ECCV (15) | 10 |
| 2024 | Fewer is More: Boosting Math Reasoning with Reinforced Context PruningabstractLarge Language Models (LLMs) have shown impressive capabilities, yet they still struggle with math reasoning.In this work, we propose CoT-Influx, a novel approach that pushes the boundary of few-shot Chain-of-Thoughts (CoT) learning to improve LLM mathematical reasoning.Motivated by the observation that adding more concise CoT examples in the prompt can improve LLM reasoning performance, CoT-Influx employs a coarse-to-fine pruner to maximize the input of effective and concise CoT examples.The pruner first selects as many crucial CoT examples as possible and then prunes unimportant tokens to fit the context window.A math reasoning dataset with diverse difficulty levels and reasoning steps is used to train the pruner, along with a math-specialized reinforcement learning approach.As a result, by enabling more CoT examples with double the context window size in tokens, CoT-Influx significantly outperforms various prompting baselines across various LLMs (LLaMA2-7B, 13B, 70B) and 6 math datasets, achieving up to 4.40% absolute improvements.Remarkably, without any fine-tuning, LLaMA2-70B with CoT-Influx surpasses GPT-3.5 and a wide range of larger LLMs (PaLM, Minerva 540B, etc.) on GSM8K.CoT-Influx is a plug-and-play module for LLMs, adaptable in various scenarios.It's compatible with advanced reasoning prompting techniques, such as self-consistency, and supports different long-context LLMs, including Mistral-7B-v0.3-32K and Yi-6B-200K.Codes are available at https://github.com/HuangOwen/CoT-Influx Xijie Huang, Li Lyna Zhang, Kwang-Ting Cheng, Fan Yang 0024, Mao Yang 0004 |
EMNLP | 4 |
| 2024 | Aceso: Efficient Parallel DNN Training through Iterative Bottleneck AlleviationabstractMany parallel mechanisms, including data parallelism, tensor parallelism, and pipeline parallelism, have been proposed and combined together to support training increasingly large deep neural networks (DNN) on massive GPU devices. Given a DNN model and GPU cluster, finding the optimal configuration by combining these parallelism mechanisms is an NP-hard problem. Widely adopted mathematical programming approaches have been proposed to search in a configuration subspace, but they are still too costly when scaling to large models over numerous devices. Youshan Miao, Xiaoxiang Shi, Saeed Maleki, Fan Yang 0024, Yungang Bao, Sa Wang |
EuroSys | 6 |
| 2024 | Tessel: Boosting Distributed Execution of Large DNN Models via Flexible Schedule SearchabstractIncreasingly complex and diverse deep neural network (DNN) models necessitate distributing the execution across multiple devices for training and inference tasks, and also require carefully planned schedules for performance. However, existing practices often rely on predefined schedules that may not fully exploit the benefits of emerging diverse model-aware operator placement strategies. Handcrafting high-efficiency schedules can be challenging due to the large and varying schedule space. This paper presents Tessel, an automated system that searches for efficient schedules for distributed DNN training and inference for diverse operator placement strategies. To reduce search costs, Tessel leverages the insight that the most efficient schedules often exhibit repetitive pattern (repetend) across different data inputs. This leads to a two-phase approach: repetend construction and schedule completion. By exploring schedules for various operator placement strategies, Tessel significantly improves both training and inference performance. Experiments with representative DNN models demonstrate that Tessel achieves up to 5.5 x training performance speedup and up to 38 % inference latency reduction. Youshan Miao, Guanbin Xu, Cheng Li 0001, Olli Saarikivi, Saeed Maleki, Fan Yang 0024 |
HPCA | 7 |
| 2024 | Integer or Floating Point? New Outlooks for Low-Bit Quantization on Large Language ModelsabstractEfficient deployment of Large Language Models (LLMs) requires low-bit quantization to reduce model size and inference cost. Besides low-bit integer formats (e.g., INT8/INT4) used in previous quantization works, emerging low-bit floating-point formats (e.g., FP8/FP4) supported by advanced hardware like NVIDIA’s H100 GPU offer an alternative. Our study finds that introducing floating-point formats significantly improves LLMs quantization. We also discover that the optimal quantization format varies across layers. Therefore, we select the optimal format for each layer, which we call the Mixture of Formats Quantization (MoFQ) method. Our MoFQ method achieves better or comparable results over current methods in weight-only (W-only) and weight-activation (WA) post-training quantization scenarios across various tasks, with no additional hardware overhead. Lingran Zhao, Shijie Cao, Ting Cao 0003, Fan Yang 0024, Mao Yang 0004, Shanghang Zhang, Ningyi Xu |
ICME | 7 |
| 2024 | LongRoPE: Extending LLM Context Window Beyond 2 Million TokensabstractLarge context window is a desirable feature in large language models (LLMs). However, due to high fine-tuning costs, scarcity of long texts, and catastrophic values introduced by new token positions, current extended context windows are limited to around 128k tokens. This paper introduces LongRoPE that, for the first time, extends the context window of pre-trained LLMs to an impressive 2048k tokens, with up to only 1k fine-tuning steps at within 256k training lengths, while maintaining performance at the original short context window. This is achieved by three key innovations: (i) we identify and exploit two forms of non-uniformities in positional interpolation through an efficient search, providing a better initialization for fine-tuning and enabling an 8x extension in non-fine-tuning scenarios; (ii) we introduce a progressive extension strategy that first fine-tunes a 256k length LLM and then conducts a second positional interpolation on the fine-tuned extended LLM to achieve a 2048k context window; (iii) we readjust LongRoPE on 8k length to recover the short context window performance. Extensive experiments on LLaMA2 and Mistral across various tasks demonstrate the effectiveness of our method. Models extended via LongRoPE retain the original architecture with minor modifications to the positional embedding, and can reuse most pre-existing optimizations. Code is available at https://github.com/microsoft/LongRoPE Yiran Ding, Li Lyna Zhang, Chengruidong Zhang, Jiahang Xu, Fan Yang 0024, Mao Yang 0004 |
ICML | 7 |
| 2024 | Understanding the Weakness of Large Language Model Agents within a Complex Android EnvironmentabstractLarge language models (LLMs) have empowered intelligent agents to execute intricate tasks within domain-specific software such as browsers and games. However, when applied to general-purpose software systems like operating systems, LLM agents face three primary challenges. Firstly, the action space is vast and dynamic, posing difficulties for LLM agents to maintain an up-to-date understanding and deliver accurate responses. Secondly, real-world tasks often require inter-application cooperation, demanding farsighted planning from LLM agents. Thirdly, agents need to identify optimal solutions aligning with user constraints, such as security concerns and preferences. These challenges motivate AndroidArena, an environment and benchmark designed to evaluate LLM agents on a modern operating system. To address high-cost of manpower, we design a scalable and semi-automated method to construct the benchmark. In the task evaluation, AndroidArena incorporates accurate and adaptive metrics to address the issue of non-unique solutions. Our findings reveal that even state-of-the-art LLM agents struggle in cross-APP scenarios and adhering to specific constraints. Additionally, we identify a lack of four key capabilities, i.e. understanding, reasoning, exploration, and reflection, as primary reasons for the failure of LLM agents. Furthermore, we provide empirical analysis on the failure of reflection, and improve the success rate by 27% with our proposed exploration strategy. This work is the first to present valuable insights in understanding fine-grained weakness of LLM agents, and offers a path forward for future research in this area. Environment, benchmark, prompt, and evaluation code for AndroidArena are released at https://github.com/AndroidArenaAgent/AndroidArena. Mingzhe Xing, Rongkai Zhang 0005, Hui Xue 0004, Qi Chen 0009, Fan Yang 0024 |
KDD | 5 |
| 2024 | Parrot: Efficient Serving of LLM-based Applications with Semantic Variable
Chaofan Lin, Zhenhua Han, Chengruidong Zhang, Yuqing Yang 0001, Fan Yang 0024, Chen Chen 0067, Lili Qiu |
OSDI | 5 |
| 2024 | nnScaler: Constraint-Guided Parallelization Plan Generation for Deep Learning Training
Youshan Miao, Quanlu Zhang, Fan Yang 0024, Cheng Li 0001, Saeed Maleki, Yilei Yang, Weijiang Xu, Mao Yang 0004, Lidong Zhou |
OSDI | 4 |
| 2024 | Ladder: Enabling Efficient Low-Precision Deep Learning Computing through Hardware-aware Tensor Transformation
Lei Wang 0222, Lingxiao Ma, Shijie Cao, Quanlu Zhang, Jilong Xue, Yining Shi 0001, Ningxin Zheng, Ziming Miao, Fan Yang 0024, Ting Cao 0003, Yuqing Yang 0001, Mao Yang 0004 |
OSDI | 9 |
| 2024 | Uncovering Nested Data Parallelism and Data Reuse in DNN Computation with FractalTensorabstractTo speed up computation, deep neural networks (DNNs) usually rely on highly optimized tensor operators. Despite the effectiveness, tensor operators are often defined empirically with ad hoc semantics. This hinders the analysis and optimization across operator boundaries. FractalTensor is a programming framework that addresses this challenge. At the core, FractalTensor is a nested list-based abstract data type (ADT), where each element is a tensor with static shape or another FractalTensor (i.e., nested). DNNs are then de-fined by high-order array compute operators like map/reduce/scan and array access operators like window/stride on FractalTensor. This new way of DNN definition explicitly exposes nested data parallelism and fine-grained data access patterns, opening new opportunities for whole program analysis and optimization. To exploit these opportunities, from the FractalTensor-based code the compiler extracts a nested multi-dimensional dataflow graph called Extended Task Dependence Graph (ETDG), which provides a holistic view of data dependency across different granularity. The ETDG is then transformed into an efficient implementation through graph coarsening, data reordering, and access materialization. Evaluation on six representative DNNs like RNN and FlashAttention on NVIDIA A100 shows that Fractal-Tensor achieves speedup by up to 5.45x and 2.14x on average through a unified solution for diverse optimizations. Siran Liu, Chengxiang Qi, Chao Yang 0002, Weifang Hu, Xuanhua Shi, Fan Yang 0024, Mao Yang 0004 |
SOSP | 7 |
| 2024 | Efficient Schedule Construction for Distributed Execution of Large DNN ModelsabstractIncreasingly complex and diverse deep neural network (DNN) models necessitate distributing the execution across multiple devices for training and inference tasks, and also require carefully planned schedules for performance. However, existing practices often rely on predefined schedules that may not fully exploit the benefits of emerging diverse model-aware operator placement strategies. Handcrafting high-efficiency schedules can be challenging due to the large and varying schedule space. This paper presents Tessel, an automated system that searches for efficient schedules for distributed DNN training and inference for diverse operator placement strategies. To reduce search costs, Tessel leverages the insight that the most efficient schedules often exhibit repetitive pattern (repetend) across different data inputs. This leads to a two-phase approach: repetend construction and schedule completion. By exploring schedules for various operator placement strategies, Tessel significantly improves both training and inference performance. Experiments with representative DNN models demonstrate that Tessel achieves up to 5.5× training performance speedup and up to 38% inference latency reduction. Youshan Miao, Guanbin Xu, Cheng Li 0001, Olli Saarikivi, Saeed Maleki, Fan Yang 0024 |
IEEE Trans. Parallel Distributed Syst. | 7 |
| 2023 | NUWA-XL: Diffusion over Diffusion for eXtremely Long Video GenerationabstractShengming Yin, Chenfei Wu, Huan Yang, Jianfeng Wang, Xiaodong Wang, Minheng Ni, Zhengyuan Yang, Linjie Li, Shuguang Liu, Fan Yang, Jianlong Fu, Ming Gong, Lijuan Wang, Zicheng Liu, Houqiang Li, Nan Duan. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Shengming Yin, Chenfei Wu, Huan Yang 0005, Xiaodong Wang 0023, Minheng Ni, Zhengyuan Yang, Fan Yang 0024, Jianlong Fu, Ming Gong 0001, Zicheng Liu 0001, Houqiang Li, Nan Duan 0001 |
ACL (1) | 10 |
| 2023 | ElasticFlow: An Elastic Serverless Training Platform for Distributed Deep LearningabstractThis paper proposes ElasticFlow, an elastic serverless training platform for distributed deep learning. ElasticFlow provides a serverless interface with two distinct features: (i) users specify only the deep neural network (DNN) model and hyperparameters for a job, but not the number of GPUs; (ii) users specify the deadline for a job, but not the amount of time to occupy GPUs. In contrast to existing server-centric platforms, ElasticFlow provides performance guarantees in terms of meeting deadlines while alleviating tedious, low-level, and manual resource management for deep learning developers. The characteristics of distributed training introduce two challenges. First, the training throughput scales non-linearly with the number of GPUs. Second, the scaling efficiency is affected by worker placement. To address these challenges, we propose Minimum Satisfactory Share to capture the resource usage of training jobs to meet deadlines, and ElasticFlow performs admission control based on it. We develop a greedy algorithm that dynamically allocates resources to admitted jobs based on diminishing returns. We apply buddy allocation to worker placement to eliminate the effect of topology. Evaluation results on a cluster of 128 GPUs show that ElasticFlow increases the number of jobs that can meet their deadlines by 1.46–7.65× compared to existing solutions. Diandian Gu, Yinmin Zhong, Yifan Xiong 0001, Zhenhua Han, Peng Cheng 0005, Fan Yang 0024, Gang Huang 0001, Xin Jin 0008, Xuanzhe Liu |
ASPLOS (2) | 7 |
| 2023 | Adam Accumulation to Reduce Memory Footprints of Both Activations and Gradients for Large-Scale DNN TrainingabstractRunning out of GPU memory has become a main bottleneck for large-scale DNN training. How to reduce the memory footprint during training has received intensive research attention. We find that previous gradient accumulation reduces activation memory but fails to be compatible with gradient memory reduction due to a contradiction between preserving gradients and releasing gradients. To address this issue, we propose a novel optimizer accumulation method for Adam, named Adam Accumulation (AdamA), which enables reducing both activation and gradient memory. Specifically, AdamA directly integrates gradients into optimizer states and accumulates optimizer states over micro-batches, so that gradients can be released immediately after use. We mathematically and experimentally demonstrate AdamA yields the same convergence properties as Adam. Evaluated on transformer-based models, AdamA achieves up to 23% memory reduction compared to gradient accumulation with less than 2% degradation in training throughput. Notably, AdamA can work together with memory reduction methods for optimizer states to fit 1.26×~3.14× larger models over PyTorch and DeepSpeed baseline on GPUs with different memory capacities. Yibo Han, Shijie Cao, Guohao Dai 0001, Youshan Miao, Ting Cao 0003, Fan Yang 0024, Ningyi Xu |
ECAI | 7 |
| 2023 | SiloD: A Co-design of Caching and Scheduling for Deep Learning ClustersabstractDeep learning training on cloud platforms usually follows the tradition of the separation of storage and computing. The training executes on a compute cluster equipped with GPUs/TPUs while reading data from a separate cluster hosting the storage service. To alleviate the potential bottleneck, a training cluster usually leverages its local storage as a cache to reduce the remote IO from the storage cluster. However, existing deep learning schedulers do not manage storage resources thus fail to consider the diverse caching effects across different training jobs. This could degrade scheduling quality significantly. Zhenhua Han, Zhi Yang 0001, Quanlu Zhang, Mingxia Li, Fan Yang 0024, Qianxi Zhang, Binyang Li, Yuqing Yang 0001, Lili Qiu, Lidong Zhou |
EuroSys | 6 |
| 2023 | Learning 3D Photography Videos via Self-supervised Diffusion on Single Imagesabstract3D photography renders a static image into a video with appealing 3D visual effects. Existing approaches typically first conduct monocular depth estimation, then render the input frame to subsequent frames with various viewpoints, and finally use an inpainting model to fill those missing/occluded regions. The inpainting model plays a crucial role in rendering quality, but it is normally trained on out-of-domain data. To reduce the training and inference gap, we propose a novel self-supervised diffusion model as the inpainting module. Given a single input image, we automatically construct a training pair of the masked occluded image and the ground-truth image with random cycle rendering. The constructed training samples are closely aligned to the testing instances, without the need for data annotation. To make full use of the masked images, we designed a Masked Enhanced Block (MEB), which can be easily plugged into the UNet and enhance the semantic conditions. Towards real-world animation, we present a novel task: out-animation, which extends the space and time of input objects. Extensive experiments on real datasets show that our method achieves competitive results with existing SOTA methods. Xiaodong Wang 0023, Chenfei Wu, Shengming Yin, Minheng Ni, Zhengyuan Yang, Fan Yang 0024, Zicheng Liu 0001, Yuejian Fang, Nan Duan 0001 |
IJCAI | 8 |
| 2023 | OliVe: Accelerating Large Language Models via Hardware-friendly Outlier-Victim Pair QuantizationabstractTransformer-based large language models (LLMs) have achieved great success with the growing model size. LLMs' size grows by 240× every two years, which outpaces the hardware progress and makes model inference increasingly costly. Model quantization is a promising approach to mitigate the widening gap between LLM size and hardware capacity. However, the existence of outliers, values with significant magnitudes, in LLMs makes existing quantization methods less effective. Prior outlier-aware quantization schemes adopt sparsity encoding techniques to separate outliers from normal values where the process requires global coordination (e.g., a global sparsity coordination list). This incurs complex encoding/decoding hardware logics and an extra orchestration controller for the computation between outlier and normal values. As such, it is not hardware-efficient and hence only achieves sub-optimal quantization benefits. Cong Guo 0003, Weiming Hu 0005, Jingwen Leng, Chen Zhang 0001, Fan Yang 0024, Yunxin Liu 0001, Minyi Guo, Yuhao Zhu 0001 |
ISCA | 6 |
| 2023 | Model-enhanced Vector IndexabstractEmbedding-based retrieval methods construct vector indices to search for document representations that are most similar to the query representations. They are widely used in document retrieval due to low latency and decent recall performance. Recent research indicates that deep retrieval solutions offer better model quality, but are hindered by unacceptable serving latency and the inability to support document updates. In this paper, we aim to enhance the vector index with end-to-end deep generative models, leveraging the differentiable advantages of deep retrieval models while maintaining desirable serving efficiency. We propose Model-enhanced Vector Index (MEVI), a differentiable model-enhanced index empowered by a twin-tower representation model. MEVI leverages a Residual Quantization (RQ) codebook to bridge the sequence-to-sequence deep retrieval and embedding-based models. To substantially reduce the inference time, instead of decoding the unique document ids in long sequential steps, we first generate some semantic virtual cluster ids of candidate documents in a small number of steps, and then leverage the well-adapted embedding vectors to further perform a fine-grained search for the relevant documents in the candidate virtual clusters. We empirically show that our model achieves better performance on the commonly used academic benchmarks MSMARCO Passage and Natural Questions, with comparable serving latency to dense retrieval solutions. Hailin Zhang 0004, Yujing Wang 0002, Qi Chen 0009, Ruiheng Chang, Ting Zhang 0002, Ziming Miao, Yingyan Hou, Xupeng Miao, Bochen Pang, Yuefeng Zhan, Hao Sun 0015, Qi Zhang 0066, Fan Yang 0024, Xing Xie 0001, Mao Yang 0004, Bin Cui 0001 |
NeurIPS | 16 |
| 2023 | On Modular Learning of Distributed Systems for Predicting End-to-End Latency
Chieh-Jan Mike Liang, Zilin Fang, Yuqing Xie 0005, Fan Yang 0024, Zhao Lucis Li, Li Lyna Zhang, Mao Yang 0004, Lidong Zhou |
NSDI | 4 |
| 2023 | Welder: Scheduling Deep Learning Memory Access via Tile-graph
Yining Shi 0001, Zhi Yang 0001, Jilong Xue, Lingxiao Ma, Yuqing Xia, Ziming Miao, Yuxiao Guo 0001, Fan Yang 0024, Lidong Zhou |
OSDI | 8 |
| 2023 | Optimizing Dynamic Neural Networks with Brainstorm
Weihao Cui, Zhenhua Han, Lingji Ouyang, Yichuan Wang 0002, Ningxin Zheng, Lingxiao Ma, Yuqing Yang 0001, Fan Yang 0024, Jilong Xue, Lili Qiu, Lidong Zhou, Quan Chen 0002, Haisheng Tan, Minyi Guo |
OSDI | 8 |
| 2023 | Cocktailer: Analyzing and Optimizing Dynamic Control Flow in Deep Learning
Chen Zhang 0001, Lingxiao Ma, Jilong Xue, Yining Shi 0001, Ziming Miao, Fan Yang 0024, Jidong Zhai, Zhi Yang 0001, Mao Yang 0004 |
OSDI | 6 |
| 2023 | VBASE: Unifying Online Vector Similarity Search and Relational Queries via Relaxed Monotonicity
Qianxi Zhang, Shuotao Xu, Qi Chen 0009, Guoxin Sui, Jiadong Xie 0002, Zhizhen Cai, Yaoqi Chen, Yinxuan He, Yuqing Yang 0001, Fan Yang 0024, Mao Yang 0004, Lidong Zhou |
OSDI | 10 |
| 2023 | SPFresh: Incremental In-Place Update for Billion-Scale Vector SearchabstractApproximate Nearest Neighbor Search (ANNS) on high dimensional vector data is now widely used in various applications, including information retrieval, question answering, and recommendation. As the amount of vector data grows continuously, it becomes important to support updates to vector index, the enabling technique that allows for efficient and accurate ANNS on vectors. Yuming Xu, Hengyu Liang, Jin Li 0050, Shuotao Xu, Qi Chen 0009, Qianxi Zhang, Cheng Li 0001, Ziyue Yang 0002, Fan Yang 0024, Yuqing Yang 0001, Peng Cheng 0005, Mao Yang 0004 |
SOSP | 9 |
| 2023 | PIT: Optimization of Dynamic Sparse Deep Learning Models via Permutation Invariant TransformationabstractDynamic sparsity, where the sparsity patterns are unknown until runtime, poses a significant challenge to deep learning. The state-of-the-art sparsity-aware deep learning solutions are restricted to pre-defined, static sparsity patterns due to significant overheads associated with preprocessing. Efficient execution of dynamic sparse computation often faces the misalignment between the GPU-friendly tile configuration for efficient execution and the sparsity-aware tile shape that minimizes coverage wastes (non-zero values in tensor). Ningxin Zheng, Huiqiang Jiang, Quanlu Zhang, Zhenhua Han, Lingxiao Ma, Yuqing Yang 0001, Fan Yang 0024, Chengruidong Zhang, Lili Qiu, Mao Yang 0004, Lidong Zhou |
SOSP | 7 |
| 2022 | NÜWA: Visual Synthesis Pre-training for Neural visUal World creAtion
Chenfei Wu, Lei Ji 0001, Fan Yang 0024, Yuejian Fang, Daxin Jiang, Nan Duan 0001 |
ECCV (16) | 4 |
| 2022 | Nesting Forward Automatic Differentiation for Memory-Efficient Deep Neural Network TrainingabstractAn activation function is an element-wise mathematical function and plays a crucial role in deep neural networks (DNN). Many novel and sophisticated activation functions have been proposed to improve the DNN accuracy but also consume massive memory in the training process with back-propagation. In this study, we propose the nested forward automatic differentiation (Forward-AD), specifically for the element-wise activation function for memory-efficient DNN training. We deploy nested Forward-AD in two widely-used deep learning frameworks, TensorFlow and PyTorch, which support the static and dynamic computation graph, respectively. Our evaluation shows that nested Forward-AD reduces the memory footprint by up to 1.97× than the baseline model and outperforms the recomputation by 20% under the same memory reduction ratio. Cong Guo 0003, Yuxian Qiu, Jingwen Leng, Chen Zhang 0001, Quanlu Zhang, Yunxin Liu 0001, Fan Yang 0024, Minyi Guo |
ICCD | 8 |
| 2022 | SQuant: On-the-Fly Data-Free Quantization via Diagonal Hessian Approximation
Cong Guo 0003, Yuxian Qiu, Jingwen Leng, Xiaotian Gao, Chen Zhang 0001, Yunxin Liu 0001, Fan Yang 0024, Yuhao Zhu 0001, Minyi Guo |
ICLR | 7 |
| 2022 | ANT: Exploiting Adaptive Numerical Data Type for Low-bit Deep Neural Network QuantizationabstractQuantization is a technique to reduce the computation and memory cost of DNN models, which are getting increasingly large. Existing quantization solutions use fixed-point integer or floating-point types, which have limited benefits, as both require more bits to maintain the accuracy of original models. On the other hand, variable-length quantization uses low-bit quantization for normal values and high-precision for a fraction of outlier values. Even though this line of work brings algorithmic benefits, it also introduces significant hardware overheads due to variable-length encoding and decoding.In this work, we propose a fixed-length a daptive n umerical data t ype called ANT to achieve low-bit quantization with tiny hardware overheads. Our data type ANT leverages two key innovations to exploit the intra-tensor and inter-tensor adaptive opportunities in DNN models. First, we propose a particular data type, flint, that combines the advantages of float and int for adapting to the importance of different values within a tensor. Second, we propose an adaptive framework that selects the best type for each tensor according to its distribution characteristics. We design a unified processing element architecture for ANT and show its ease of integration with existing DNN accelerators. Our design results in $2.8\times $ speedup and $2.5\times $ energy efficiency improvement over the state-of-the-art quantization accelerators. Cong Guo 0003, Chen Zhang 0001, Jingwen Leng, Zihan Liu 0002, Fan Yang 0024, Yunxin Liu 0001, Minyi Guo, Yuhao Zhu 0001 |
MICRO | 5 |
| 2022 | SparTA: Deep-Learning Model Sparsity via Tensor-with-Sparsity-Attribute
Ningxin Zheng, Quanlu Zhang, Lingxiao Ma, Yuqing Yang 0001, Fan Yang 0024, Yang Wang 0053, Mao Yang 0004, Lidong Zhou |
OSDI | 6 |
| 2022 | ROLLER: Fast and Efficient Tensor Compilation for Deep Learning
Hongyu Zhu 0003, Yijia Diao, Shanbin Ke, Chen Zhang 0001, Jilong Xue, Lingxiao Ma, Yuqing Xia, Fan Yang 0024, Mao Yang 0004, Lidong Zhou, Asaf Cidon, Gennady Pekhimenko |
OSDI | 11 |
| 2022 | Distill-VQ: Learning Retrieval Oriented Vector Quantization By Distilling Knowledge from Dense EmbeddingsabstractVector quantization (VQ) based ANN indexes, such as Inverted File System (IVF) and Product Quantization (PQ), have been widely applied to embedding based document retrieval thanks to the competitive time and memory efficiency. Originally, VQ is learned to minimize the reconstruction loss, i.e., the distortions between the original dense embeddings and the reconstructed embeddings after quantization. Unfortunately, such an objective is inconsistent with the goal of selecting ground-truth documents for the input query, which may cause severe loss of retrieval quality. Recent works identify such a defect, and propose to minimize the retrieval loss through contrastive learning. However, these methods intensively rely on queries with ground-truth documents, whose performance is limited by the insufficiency of labeled data. In this paper, we propose Distill-VQ, which unifies the learning of IVF and PQ within a knowledge distillation framework. In Distill-VQ, the dense embeddings are leveraged as "teachers'', which predict the query's relevance to the sampled documents. The VQ modules are treated as the "students'', which are learned to reproduce the predicted relevance, such that the reconstructed embeddings may fully preserve the retrieval result of the dense embeddings. By doing so, Distill-VQ is able to derive substantial training signals from the massive unlabeled data, which significantly contributes to the retrieval quality. We perform comprehensive explorations for the optimal conduct of knowledge distillation, which may provide useful insights for the learning of VQ based ANN index. We also experimentally show that the labeled data is no longer a necessity for high-quality vector quantization, which indicates Distill-VQ's strong applicability in practice. The evaluations are performed on MS MARCO and Natural Questions benchmarks, where Distill-VQ notably outperforms the SOTA VQ methods in Recall and MRR. Our code is avaliable at https://github.com/staoxiao/LibVQ. Shitao Xiao, Zheng Liu 0011, Weihao Han, Jianjin Zhang, Defu Lian, Yeyun Gong, Qi Chen 0009, Fan Yang 0024, Hao Sun 0015, Yingxia Shao, Xing Xie 0001 |
SIGIR | 8 |
| 2022 | PilotFish: Harvesting Free Cycles of Cloud Gaming with Deep Learning Training
Wei Zhang 0149, Binghao Chen, Zhenhua Han, Quan Chen 0002, Peng Cheng 0005, Fan Yang 0024, Ran Shu 0001, Yuqing Yang 0001, Minyi Guo |
USENIX ATC | 6 |
| 2020 | Capuchin: Tensor-based GPU Memory Management for Deep LearningabstractIn recent years, deep learning has gained unprecedented success in various domains, the key of the success is the larger and deeper deep neural networks (DNNs) that achieved very high accuracy. On the other side, since GPU global memory is a scarce resource, large models also pose a significant challenge due to memory requirement in the training process. This restriction limits the DNN architecture exploration flexibility. Xuanhua Shi, Hulin Dai, Hai Jin 0001, Weiliang Ma, Qian Xiong, Fan Yang 0024, Xuehai Qian |
ASPLOS | 7 |
| 2020 | XGLUE: A New Benchmark Datasetfor Cross-lingual Pre-training, Understanding and GenerationabstractYaobo Liang, Nan Duan, Yeyun Gong, Ning Wu, Fenfei Guo, Weizhen Qi, Ming Gong, Linjun Shou, Daxin Jiang, Guihong Cao, Xiaodong Fan, Ruofei Zhang, Rahul Agrawal, Edward Cui, Sining Wei, Taroon Bharti, Ying Qiao, Jiun-Hung Chen, Winnie Wu, Shuguang Liu, Fan Yang, Daniel Campos, Rangan Majumder, Ming Zhou. Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP). 2020. Yaobo Liang, Nan Duan 0001, Yeyun Gong, Ning Wu 0013, Fenfei Guo, Weizhen Qi, Ming Gong 0001, Linjun Shou, Daxin Jiang, Guihong Cao, Xiaodong Fan, Ruofei Zhang, Rahul Agrawal, Edward Dong Bo Cui, Sining Wei, Taroon Bharti, Jiun-Hung Chen, Winnie Wu, Fan Yang 0024, Daniel Campos, Rangan Majumder, Ming Zhou 0001 |
EMNLP (1) | 21 |
| 2020 | Rammer: Enabling Holistic Deep Learning Compiler Optimizations with rTasks
Lingxiao Ma, Zhi Yang 0001, Jilong Xue, Youshan Miao, Wenxiang Hu, Fan Yang 0024, Lidong Zhou |
OSDI | 8 |
| 2020 | Retiarii: A Deep Learning Exploratory-Training Framework
Quanlu Zhang, Zhenhua Han, Fan Yang 0024, Yuge Zhang, Mao Yang 0004, Lidong Zhou |
OSDI | 3 |
| 2020 | HiveD: Sharing a GPU Cluster for Deep Learning with Guarantees
Zhenhua Han, Zhi Yang 0001, Quanlu Zhang, Fan Yang 0024, Lidong Zhou, Mao Yang 0004, Francis C. M. Lau 0001, Yifan Xiong 0001 |
OSDI | 5 |
| 2019 | Analysis of Large-Scale Multi-Tenant GPU Clusters for DNN Training Workloads
Myeongjae Jeon, Shivaram Venkataraman, Amar Phanishayee, Junjie Qian, Wencong Xiao, Fan Yang 0024 |
USENIX ATC | 6 |
| 2019 | Scaling out NUMA-Aware Applications with RDMA-Based Distributed Shared Memory
Yang Hong 0007, Fan Yang 0024, Binyu Zang, Haibing Guan, Haibo Chen 0001 |
J. Comput. Sci. Technol. | 3 |
| 2018 | Scheduling CPU for GPU-based Deep Learning JobsabstractDeep learning (DL) is popular in data-center as an important workload for artificial intelligence. With the recent breakthrough of using graphics accelerators and the popularity of DL framework, GPU server cluster dominates DL training in current practice. Cluster scheduler simply treats DL jobs as black-boxes and allocates GPUs as per job request specified by a user. However, other resources, e.g. CPU, are often allocated with workload-agnostic approaches. Kubeflow[1] performs heuristic static CPU resource assignment based on task types (e.g., worker, parameter-server), while [2] evenly divides CPUs of a server to each GPU. Despite the traditional impression that GPU is critical in DL, our observation suggests that the importance of CPU is undervalued. Identifying an appropriate CPU core number in a heterogeneous cluster is challenging yet performance critical to DL jobs. The diverse CPU usage characteristic is not well recognized in the following three aspects. Wencong Xiao, Zhenhua Han, Quanlu Zhang, Fan Yang 0024, Lidong Zhou |
SoCC | 6 |
| 2018 | Gandiva: Introspective Cluster Scheduling for Deep Learning
Wencong Xiao, Romil Bhardwaj, Ramachandran Ramjee, Muthian Sivathanu, Nipun Kwatra, Zhenhua Han, Pratyush Patel, Quanlu Zhang, Fan Yang 0024, Lidong Zhou |
OSDI | 11 |
| 2015 | GraM: scaling graph computation to the trillionsabstractGraM is an efficient and scalable graph engine for a large class of widely used graph algorithms. It is designed to scale up to multicores on a single server, as well as scale out to multiple servers in a cluster, offering significant, often over an order-of-magnitude, improvement over existing distributed graph engines on evaluated graph algorithms. GraM is also capable of processing graphs that are significantly larger than previously reported. In particular, using 64 servers (1,024 physical cores), it performs a PageRank iteration in 140 seconds on a synthetic graph with over one trillion edges, setting a new milestone for graph engines. Ming Wu 0007, Fan Yang 0024, Jilong Xue, Wencong Xiao, Youshan Miao, Haoxiang Lin, Yafei Dai, Lidong Zhou |
SoCC | 2 |
| 2015 | ImmortalGraph: A System for Storage and Analysis of Temporal GraphsabstractTemporal graphs that capture graph changes over time are attracting increasing interest from research communities, for functions such as understanding temporal characteristics of social interactions on a time-evolving social graph. ImmortalGraph is a storage and execution engine designed and optimized specifically for temporal graphs. Locality is at the center of ImmortalGraph’s design: temporal graphs are carefully laid out in both persistent storage and memory, taking into account data locality in both time and graph-structure dimensions. ImmortalGraph introduces the notion of locality-aware batch scheduling in computation, so that common “bulk” operations on temporal graphs are scheduled to maximize the benefit of in-memory data locality. The design of ImmortalGraph explores an interesting interplay among locality, parallelism, and incremental computation in supporting common mining tasks on temporal graphs. The result is a high-performance temporal-graph system that is up to 5 times more efficient than existing database solutions for graph queries. The locality optimizations in ImmortalGraph offer up to an order of magnitude speedup for temporal iterative graph mining compared to a straightforward application of existing graph engines on a series of snapshots. Youshan Miao, Ming Wu 0007, Fan Yang 0024, Lidong Zhou, Vijayan Prabhakaran, Enhong Chen |
ACM Trans. Storage | 5 |
| 2014 | Chronos: a graph engine for temporal graph analysisabstractTemporal graphs capture changes in graphs over time and are becoming a subject that attracts increasing interest from the research communities, for example, to understand temporal characteristics of social interactions on a time-evolving social graph. Chronos is a storage and execution engine designed and optimized specifically for running in-memory iterative graph computation on temporal graphs. Locality is at the center of the Chronos design, where the in-memory layout of temporal graphs and the scheduling of the iterative computation on temporal graphs are carefully designed, so that common "bulk" operations on temporal graphs are scheduled to maximize the benefit of in-memory data locality. The design of Chronos further explores the interesting interplay among locality, parallelism, and incremental computation in supporting common mining tasks on temporal graphs. The result is a high-performance temporal-graph system that offers up to an order of magnitude speedup for temporal iterative graph mining compared to a straightforward application of existing graph engines on a series of snapshots. Youshan Miao, Ming Wu 0007, Fan Yang 0024, Lidong Zhou, Vijayan Prabhakaran, Enhong Chen |
EuroSys | 5 |
| 2012 | Kineograph: taking the pulse of a fast-changing and connected worldabstractKineograph is a distributed system that takes a stream of incoming data to construct a continuously changing graph, which captures the relationships that exist in the data feed. As a computing platform, Kineograph further supports graph-mining algorithms to extract timely insights from the fast-changing graph structure. To accommodate graph-mining algorithms that assume a static underlying graph, Kineograph creates a series of consistent snapshots, using a novel and efficient epoch commit protocol. To keep up with continuous updates on the graph, Kineograph includes an incremental graph-computation engine. We have developed three applications on top of Kineograph to analyze Twitter data: user ranking, approximate shortest paths, and controversial topic detection. For these applications, Kineograph takes a live Twitter data feed and maintains a graph of edges between all users and hashtags. Our evaluation shows that with 40 machines processing 100K tweets per second, Kineograph is able to continuously compute global properties, such as user ranks, with less than 2.5-minute timeliness guarantees. This rate of traffic is more than 10 times the reported peak rate of Twitter as of October 2011. Raymond Cheng 0001, Aapo Kyrola, Youshan Miao, Xuetian Weng, Ming Wu 0007, Fan Yang 0024, Lidong Zhou, Feng Zhao 0001, Enhong Chen |
EuroSys | 7 |
| 2007 | Modeling path capacity in multi-hop IEEE 802.11 networks for QoS servicesabstractQoS provisioning in multi-hop IEEE 802.11 networks is very challenging due to the interference nature of wireless medium and the contention-based behavior among neighboring nodes. In such networks, one of the key questions for QoS support is: given a specific topology and traffic condition, how much bandwidth can be utilized along a path in the network without violating QoS demand of existing traffic? Considering that in general QoS-sensitive traffic has the well-controlled sending rate, one key observation is that the network unsaturated condition should be considered. Another observation is that, not only the interaction between the new traffic and the existing ones that can be sensed (by the new one), but also the interaction between the new traffic and the traffic that is hidden but can have influence upon the new one should be studied. Based upon the above observations, we propose an analytical model for multi-hop IEEE 802.11 networks to calculate how much bandwidth can be utilized along a path without violating the QoS requirements of existing traffic. A notion, "free channel time", which is the time allowed for a wireless link to transmit data, is introduced to analyze the path capacity. Simulation results demonstrate that our proposed analytical model can accurately predict the path capacity under various network conditions without breaking QoS demands of all existing traffic Kun Wang 0005, Fan Yang 0024, Qian Zhang 0001, Yinlong Xu 0001 |
IEEE Trans. Wirel. Commun. | 2 |
| 2006 | On Improving the Throughput of Media Delivery Applications in Heterogeneous Overlay NetworkabstractOverlay multicast is becoming increasingly popular for multimedia delivery over Internet. The available bandwidths between sender and different receivers are generally heterogeneous. In this paper, we propose a scheme to improve the throughput of a media delivery session with heterogeneous receivers by organizing the receivers into layered data distribution meshes and sending substreams to each mesh using layered coding. Our solution utilizes alternative paths and network coding in each mesh. We first formulate the problem into a mathematical programming, whose optimal solution requires global information. We therefore present a distributed heuristic algorithm. The heuristic progressively organizes the receivers into layered meshes. Each receiver can subscribe to a proper number of meshes to maximize its throughput by fully utilizing its available bandwidth. The benefits of organizing the topology into layered mesh and using network coding are demonstrated through extensive simulations. Numerical results indicate that the average throughput of a media delivery session is significantly improved. Fan Yang 0024, Qian Zhang 0001, Zhensheng Zhang |
GLOBECOM | 2 |
| 2006 | Impact of Power and Rate Selection on the Throughput of Ad Hoc NetworksabstractWith the advance of wireless technology, wireless devices are capable of adjusting transmit power and physical (PHY) layer data-rate. In this paper, we investigate the problem of how to adjust the power level and the PHY rate in order to maximize the network throughput in wireless ad hoc networks. Solving this problem can help compute the network capacity of ad hoc networks, which has drawn a lot of attention recently. In our study, we find that there exist intertwined relationships among maximum network throughput, power, and PHY rate. These intertwined relationships make computing the maximum network throughput a difficult problem. To get around the coupled relationships among power, PHY rate and network throughput, we take a simulation-based optimization approach, i.e., use a recursive randomized algorithm to find the solution. We study the convergence and complexity of our algorithm. Our results show that our algorithm always converges and the computation complexity is polynomial. The simulations also show that our algorithm can iteratively improve the network throughput, given an initial feasible power and rate setting. Cong Peng 0007, Fan Yang 0024, Qian Zhang 0001, Dapeng Oliver Wu, Ming Zhao 0001, Yan Yao 0002 |
ICC | 2 |
| 2006 | Modeling Path Capacity in Multi-hop IEEE 802.11 Networks for QoS Servicesabstractwe propose an analytical model for multi-hop IEEE 802.11 networks to calculate how much bandwidth can be utilized along a path without violating the QoS requirements of existing rate-controlled traffic flows. A notion, "free channel time", which is the time allowed for a wireless link to transmit data, is introduced to analyze the path capacity. To achieve the goal, the proposed model effectively characterizes the unsaturated traffic condition. It could also depict the interaction between the newly injected traffic and the hidden traffic that could have influence upon the new traffic. Simulation results demonstrate that our proposed analytical model can accurately predict the path capacity under various network conditions without breaking QoS demands of all existing traffic. Kun Wang 0005, Fan Yang 0024, Qian Zhang 0001, Yinlong Xu 0001 |
MASS | 2 |
| 2006 | Distributed cooperative rate adaptation for energy efficiency in IEEE 802.11-based multi-hop networksabstractIn this paper we study the problem of using the rate adaptation technique to achieve energy efficiency in an IEEE 802.11-based multi-hop network. Specifically, we formulate it as an optimization problem, i.e., minimizing the total transmission power over transmission data rates, subject to the traffic requirements of all the nodes in a multi-hop network. Interestingly, we can show that this problem is actually a well-known multiple-choice knapsack problem, which is proven to be an NP-hard problem. So, instead of finding an optimal solution, which is NP-hard, we seek a sub-optimal solution. Our key technique to attack this problem is distributed cooperative rate adaptation. Here, we promote node cooperation due to our observation that the inequality in non-cooperative channel contention among nodes caused by hidden terminal phenomenon in a multi-hop network tends to result in energy inefficiency. Under this design philosophy, we propose a distributed cooperative rate adaptation (CRA) scheme and prove that it converges. Simulation results show that our CRA scheme can reduce the power consumption up to 86% as compared to the existing (non-cooperative) algorithm. Kun Wang 0005, Fan Yang 0024, Qian Zhang 0001, Dapeng Oliver Wu, Yinlong Xu 0001 |
QSHINE | 2 |
| 2006 | Distributed Channel Assignment and Routing in Multiradio Multichannel Multihop Wireless NetworksabstractIn this paper, we first identify several challenges in designing a joint channel assignment and routing (JCAR) protocol in heterogeneous multiradio multichannel multihop wireless networks (M3WNs) using commercial hardware [e.g., IEEE 802.11 Network Interface Card (NIC)]. We then propose a novel software solution, called Layer 2.5 JCAR, which resides between the MAC layer and routing layer. JCAR jointly coordinates the channel selection on each wireless interface and the route selection among interfaces based on the traffic information measured and exchanged among the two-hop neighbors. Since interference is one of the major factors that constrain the performance in a M3WN, in this paper, we introduce an important channel cost metric (CCM) which actually reflects the interference cost and is defined as the sum of expected transmission time weighted by the channel utilization over all interfering channels (for each node). In CCM, both the interference and the diverse channel characteristics are taken into account. An expression for CCM is derived in terms of equivalent fraction of air time by explicitly taking the radio heterogeneity into consideration. Using CCM as one of the key performance measures, we propose a distributed algorithm (heuristic) that produces near-optimal JCAR solution. To evaluate the efficacy of our heuristics, we conduct extensive simulations using the network simulator NS2. To demonstrate implementation feasibility, we conducted various experiments for the proposed distributed JCAR algorithm on a multihop wireless network testbed with nine wireless nodes, each is equipped with single/multiple 802.11a/g cards. Both experimental and simulation results demonstrate the effectiveness and implementation easiness of our proposed software solution Fan Yang 0024, Kun Tan 0001, Jie Chen 0013, Qian Zhang 0001, Zhensheng Zhang |
IEEE J. Sel. Areas Commun. | 2 |
| 2006 | LION: Layered Overlay Multicast With Network CodingabstractRecent advances in information theory show that the throughput of a multicast session can be improved using network coding. In overlay networks, the available bandwidth between sender and different receivers are different. In this paper, we propose a solution to improve the throughput of an overlay multicast session with heterogeneous receivers by organizing the receivers into layered data distribution meshes and sending substreams to each mesh using layered coding. Our solutions utilize alternative paths and network coding in each mesh. We first formulate the problem into a mathematical programming, whose optimal solution requires global information. We therefore present a distributed heuristic algorithm. The heuristic progressively organizes the receivers into layered meshes. Each receiver can subscribe to a proper number of meshes to maximize its throughput by fully utilizing its available bandwidth. The benefits of organizing the topology into layered mesh and using network coding are demonstrated through extensive simulations. Numerical results indicate that the average throughput of a multicast session is significantly improved (up to 50% to 60%) with only slightly higher delay and network resource consumption Fan Yang 0024, Qian Zhang 0001, Zhensheng Zhang, Fuyan Zhang |
IEEE Trans. Multim. | 2 |
| 2005 | AMTP: a multipath multimedia streaming protocol for mobile ad hoc networksabstractMultipath streaming is a promising technique for multimedia transport over mobile ad hoc networks. In this paper, we propose AMTP, an ad hoc multipath streaming protocol for multimedia delivery. Coupled with a QoS multipath routing protocol, AMTP can accurately differentiate packet losses due to different network conditions, which is useful for real time multimedia streaming applications to perform correct error control and resource allocation. It also selects multiple maximally disjointed paths with best QoS to maximize aggregate end-to-end throughput. In case of path broken, AMTP can seamlessly switch to a proper path and therefore maintain high streaming quality. Simulation results demonstrate the effectiveness of the proposed AMTP. Kultida Rojviboonchai, Fan Yang 0024, Qian Zhang 0001, Hitoshi Aida, Wenwu Zhu 0001 |
ICC | 2 |
| 2004 | Streaming and Bit Allocation for Scalable Video over Mobile Wireless InternetabstractWith the convergence of wired line Internet and mobile wireless networks as well as tremendous demand on video application in mobile wireless Internet, it's essential to design an effective video streaming protocol and an efficient resource allocation scheme for video delivery over wireless Internet. In this paper, we employ WMSTFP, an end-to-end TCP-friendly multimedia streaming protocol, to detect the status of the wired and wireless part of the wireless Internet, where only the last hop is wireless link. By accurately distinguishing the packet losses due to transmission errors from the congestive losses and smoothing out the pathologic round-trip-time values caused by the highly dynamic wireless environment, in WMSTFP higher throughput in wireless Internet can be achieved and rate can be adjusted in a smooth and TCP-friendly manner. Based upon WMSTFP, we propose a novel loss pattern differentiated bit allocation scheme while applying unequal loss protection (ULP) for scalable video streaming over wireless Internet. Specifically, a rate-distortion (R-D) based bit allocation scheme which considers both wired and wireless network status is proposed to minimize the expected end-to-end distortion. The optimal solution of global optimization for the bit allocation scheme is obtained by a local search algorithm taking the characteristics of progressive fine granularity scalable (PFGS) video into account. Analytical and simulation results demonstrate the effectiveness of our proposed schemes Fan Yang 0024, Qian Zhang 0001, Wenwu Zhu 0001, Ya-Qin Zhang |
INFOCOM | 1 |
| 2004 | End-to-end TCP-friendly streaming protocol and bit allocation for scalable video over wireless InternetabstractWith the convergence of wired-line Internet and mobile wireless networks, as well as the tremendous demand on video applications in mobile wireless Internet, it is essential to an design effective video streaming protocol and resource allocation scheme for video delivery over wireless Internet. Taking both network conditions in the Internet and wireless networks into account, in this paper, we first propose an end-to-end transmission control protocol (TCP)-friendly multimedia streaming protocol for wireless Internet, namely WMSTFP, where only the last hop is wireless. WMSTFP can effectively differentiate erroneous packet losses from congestive losses and filter out the abnormal round-trip time values caused by the highly varying wireless environment. As a result, WMSTFP can achieve higher throughput in wireless Internet and can perform rate adjustment in a smooth and TCP-friendly manner. Based upon WMSTFP, we then propose a novel loss pattern differentiated bit allocation scheme, while applying unequal loss protection for scalable video streaming over wireless Internet. Specifically, a rate-distortion-based bit allocation scheme which considers both the wired and the wireless network status is proposed to minimize the expected end-to-end distortion. The global optimal solution for the bit allocation scheme is obtained by a local search algorithm taking the characteristics of the progressive fine granularity scalable video into account. Analytical and simulation results demonstrate the effectiveness of our proposed schemes. Fan Yang 0024, Qian Zhang 0001, Wenwu Zhu 0001, Ya-Qin Zhang |
IEEE J. Sel. Areas Commun. | 1 |
| 2003 | An end-to-end TCP-friendly streaming protocol for multimedia over wireless InternetabstractWith the convergence of wired line Internet and mobile wireless networks, it's important to study its impacts on continuous media delivery and media streaming protocols. In this paper, we propose an end-to-end (wireless) multimedia streaming TCP-friendly protocol for media delivery over wireless Internet (WMSTFP). WMSTFP can effectively differentiate erroneous packet losses from congestive losses and filter out the abnormal round-trip-time values due to the highly varying wireless environment. Analytical and simulation results show that WMSTFP can achieve higher throughput in wireless Internet and can perform rate adjustment in a smooth and TCP-friendly manner. Fan Yang 0024, Qian Zhang 0001, Wenwu Zhu 0001, Ya-Qin Zhang |
ICME | 1 |