EDBT 2026 Demo / reviewers in the wild / expert
Yang Li 0183
dblp:37/4190-183
· DBLP profile ↗
18ranked-venue papers
7as first author
9since 2021 · last 2026
0000-0002-9052-9308ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 9 · 4 first-author · 1 since 2021Artificial intelligence and machine learning · 4 · 1 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 2 first-author · 3 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
3 papers |
Efficient and distributed learning · 49% Language models and text generation · 39% Probabilistic and Bayesian machine learning · 12% | |
| Computer architecture, parallel and distributed computing, and storage systems
7 papers |
Energy-efficient computing · 32% Memory systems · 19% GPUs and heterogeneous computing · 13% | |
| Databases, data mining, and information retrieval
2 papers |
Data mining · 64% Information retrieval · 36% | |
| Network and information security
1 paper |
Privacy and data protection · 50% Cryptographic protocols and secure computation · 50% |
Topics — the 30 heaviest of 36, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Efficient and distributed learning › model compression
low-rank approximation |
1.0 | 1 | 2026 | IMPACT: Importance-Aware Activation Space Reconstruction · ACL (1) 2026 |
Machine learning › Efficient and distributed learning
model compression |
1.0 | 1 | 2026 | IMPACT: Importance-Aware Activation Space Reconstruction · ACL (1) 2026 |
Natural language and speech › Language models and text generation › large language model training
data mixing |
0.9 | 1 | 2025 | AutoMixer: Checkpoint Artifacts as Automatic Data Mixers · ACL (1) 2025 |
Machine learning › Efficient and distributed learning
data selection |
0.9 | 1 | 2025 | AutoMixer: Checkpoint Artifacts as Automatic Data Mixers · ACL (1) 2025 |
Natural language and speech › Language models and text generation
large language model fine-tuning |
0.9 | 1 | 2025 | AutoMixer: Checkpoint Artifacts as Automatic Data Mixers · ACL (1) 2025 |
Natural language and speech › Language models and text generation
large language model training |
0.8 | 1 | 2024 | Target-Aware Language Modeling via Granular Data Sampling · EMNLP 2024 |
Machine learning › Probabilistic and Bayesian machine learning
sampling |
0.8 | 1 | 2024 | Target-Aware Language Modeling via Granular Data Sampling · EMNLP 2024 |
Information retrieval › retrieval models › neural retrieval
embedding-based retrieval |
0.8 | 1 | 2024 | GPU-based Private Information Retrieval for On-Device Machine Learning Inference · ASPLOS (1) 2024 |
Privacy and data protection › privacy-preserving machine learning
privacy-preserving machine learning inference |
0.8 | 1 | 2024 | GPU-based Private Information Retrieval for On-Device Machine Learning Inference · ASPLOS (1) 2024 |
Cryptographic protocols and secure computation
private information retrieval |
0.8 | 1 | 2024 | GPU-based Private Information Retrieval for On-Device Machine Learning Inference · ASPLOS (1) 2024 |
GPUs and heterogeneous computing
GPU computing |
0.8 | 1 | 2024 | GPU-based Private Information Retrieval for On-Device Machine Learning Inference · ASPLOS (1) 2024 |
Data mining
clustering |
0.7 | 1 | 2023 | Forecasting COVID-19 Dynamics: Clustering, Generalized Spatiotemporal Attention, and Impacts of Mobility and Geographic Proximity · ICDE 2023 |
Data mining
spatiotemporal data mining |
0.7 | 1 | 2023 | Forecasting COVID-19 Dynamics: Clustering, Generalized Spatiotemporal Attention, and Impacts of Mobility and Geographic Proximity · ICDE 2023 |
Energy-efficient computing
datacenter power management |
0.6 | 2 | 2019 | A Scalable Priority-Aware Approach to Managing Data Center Server Power · HPCA 2019 SizeCap: Efficiently handling power surges in fuel cell powered data centers · HPCA 2016 |
Energy-efficient computing
power management |
0.4 | 1 | 2019 | A Scalable Priority-Aware Approach to Managing Data Center Server Power · HPCA 2019 |
Energy-efficient computing › datacenter power management
server power capping |
0.4 | 1 | 2019 | A Scalable Priority-Aware Approach to Managing Data Center Server Power · HPCA 2019 |
Memory systems › cache management
cache isolation |
0.3 | 1 | 2018 | dCat: dynamic cache management for efficient, performance-sensitive infrastructure-as-a-service · EuroSys 2018 |
Memory systems
cache management |
0.3 | 1 | 2018 | dCat: dynamic cache management for efficient, performance-sensitive infrastructure-as-a-service · EuroSys 2018 |
Memory systems › cache management
cache partitioning |
0.3 | 1 | 2018 | dCat: dynamic cache management for efficient, performance-sensitive infrastructure-as-a-service · EuroSys 2018 |
Cloud and datacenter computing
performance isolation |
0.3 | 1 | 2018 | dCat: dynamic cache management for efficient, performance-sensitive infrastructure-as-a-service · EuroSys 2018 |
Cloud and datacenter computing
virtualization |
0.3 | 1 | 2018 | dCat: dynamic cache management for efficient, performance-sensitive infrastructure-as-a-service · EuroSys 2018 |
Energy-efficient computing › power management
power capping |
0.2 | 1 | 2016 | SizeCap: Efficiently handling power surges in fuel cell powered data centers · HPCA 2016 |
Machine learning › Efficient and distributed learning
data-efficient learning |
0.2 | 1 | 2024 | Target-Aware Language Modeling via Granular Data Sampling · EMNLP 2024 |
Medical and health informatics › public health › public health informatics
infectious disease forecasting |
0.2 | 1 | 2023 | Forecasting COVID-19 Dynamics: Clustering, Generalized Spatiotemporal Attention, and Impacts of Mobility and Geographic Proximity · ICDE 2023 |
Integrated circuit design
3d integration |
0.2 | 1 | 2013 | An accurate semi-analytical framework for full-chip TSV-induced stress modeling · DAC 2013 |
Integrated circuit design
analog and mixed-signal circuits |
0.2 | 1 | 2013 | Time-domain segmentation based massively parallel simulation for ADCs · DAC 2013 |
Electronic design automation
circuit simulation |
0.2 | 1 | 2013 | Time-domain segmentation based massively parallel simulation for ADCs · DAC 2013 |
High-performance computing › large-scale simulation
massively parallel simulation |
0.2 | 1 | 2013 | Time-domain segmentation based massively parallel simulation for ADCs · DAC 2013 |
Electronic design automation › circuit simulation › device and circuit simulation
transistor-level simulation |
0.2 | 1 | 2013 | Time-domain segmentation based massively parallel simulation for ADCs · DAC 2013 |
Electronic design automation › circuit simulation › reduced-order modeling
macromodeling |
0.1 | 1 | 2011 | A novel framework for passive macro-modeling · DAC 2011 |
Methods — techniques the papers use, named apart from their topics
PIR-ML co-design · 2.3deep learning · 1.3clustering · 1.3importance-weighted covariance · 1.0gradient-based importance · 1.0data mixing · 0.9checkpoint analysis · 0.9target-aware sampling · 0.8spatiotemporal attention · 0.7spatio-temporal attention · 0.7priority-aware scheduling · 0.4power capping · 0.4dynamic cache management · 0.3Intel CAT · 0.3trace-driven simulation · 0.2semi-analytical modeling · 0.2parallel computing · 0.2linear superposition · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | IMPACT: Importance-Aware Activation Space ReconstructionabstractLarge language models (LLMs) achieve strong performance across diverse domains but remain difficult to deploy in resource-constrained environments due to their size.Low-rank compression is a common remedy, typically minimizing weight reconstruction error under the assumption that weights are low-rank.However, this assumption often does not hold in LLMs.In contrast, LLM activations exhibit a more pronounced low-rank structure, motivating approaches that minimize activation reconstruction error.This shift alone, however, is not sufficient: different activation dimensions contribute unequally to model performance, and treating them uniformly can lead to accuracy loss.We introduce IMPACT, an importance-aware activation reconstruction framework that links compression to its effect on model performance.IMPACT formulates compression as an optimization problem that integrates activation structure with gradient-based importance, deriving a closed-form solution where reconstruction bases arise from an importance-weighted activation covariance matrix.This yields lowrank compression explicitly optimized for accuracy preservation.Experiments across multiple models and tasks demonstrate that IMPACT achieves up to 55.4% greater model size reduction while maintaining accuracy comparable to or better than state-of-the-art baselines. Md Mokarram Chowdhury, Daniel Agyei Asante, Ernie Chang, Yang Li 0183 |
ACL (1) | 4 |
| 2025 | AutoMixer: Checkpoint Artifacts as Automatic Data MixersabstractErnie Chang, Yang Li, Patrick Huber, Vish Vogeti, David Kant, Yangyang Shi, Vikas Chandra. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Ernie Chang, Yang Li 0183, Patrick Huber, Vish Vogeti, David Kant, Yangyang Shi, Vikas Chandra |
ACL (1) | 2 |
| 2025 | Forecasting Graph-Based Time-Dependent Data with Graph Sequence AttentionabstractForecasting graph-based, time-dependent data has broad practical applications but presents challenges. Effective models must capture both spatial and temporal dependencies in the data, while also incorporating auxiliary information to enhance prediction accuracy. In this article, we identify limitations in current state-of-the-art models regarding temporal dependency handling. To overcome this, we introduce GSA-Forecaster, a new deep learning model designed for forecasting in graph-based, time-dependent contexts. GSA-Forecaster utilizes graph sequence attention, a new attention mechanism proposed in this article, to effectively manage temporal dependencies. GSA-Forecaster integrates the data’s graph structure directly into its architecture, addressing spatial dependencies. Additionally, it incorporates auxiliary information to refine its predictions further. We validate its performance using real-world graph-based, time-dependent datasets, where it demonstrates superior effectiveness compared to existing state-of-the-art models. Yang Li 0183, Di Wang 0003, José M. F. Moura |
ACM Trans. Knowl. Discov. Data | 1 |
| 2024 | GPU-based Private Information Retrieval for On-Device Machine Learning InferenceabstractOn-device machine learning (ML) inference can enable the use of private user data on user devices without revealing them to remote servers. However, a pure on-device solution to private ML inference is impractical for many applications that rely on embedding tables that are too large to be stored on-device. In particular, recommendation models typically use multiple embedding tables each on the order of 1--10 GBs of data, making them impractical to store on-device. To overcome this barrier, we propose the use of private information retrieval (PIR) to efficiently and privately retrieve embeddings from servers without sharing any private information. As off-the-shelf PIR algorithms are usually too computationally intensive to directly use for latency-sensitive inference tasks, we 1) propose novel GPU-based acceleration of PIR, and 2) co-design PIR with the downstream ML application to obtain further speedup. Our GPU acceleration strategy improves system throughput by more than 20× over an optimized CPU PIR implementation, and our PIR-ML co-design provides an over 5× additional throughput improvement at fixed model quality. Together, for various on-device ML applications such as recommendation and language modeling, our system on a single V100 GPU can serve up to 100,000 queries per second---a > 100× throughput improvement over a CPU-based baseline---while maintaining model accuracy. Maximilian Lam, Jeff Johnson 0004, Wenjie Xiong 0001, Kiwan Maeng, Udit Gupta 0001, Yang Li 0183, Liangzhen Lai, Ilias Leontiadis, Minsoo Rhu, Hsien-Hsin S. Lee, Vijay Janapa Reddi, Gu-Yeon Wei, David Brooks 0001, G. Edward Suh |
ASPLOS (1) | 6 |
| 2024 | Target-Aware Language Modeling via Granular Data SamplingabstractErnie Chang, Pin-Jie Lin, Yang Li, Changsheng Zhao, Daeil Kim, Rastislav Rabatin, Zechun Liu, Yangyang Shi, Vikas Chandra. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024. Ernie Chang, Pin-Jie Lin, Yang Li 0183, Changsheng Zhao 0002, Daeil Kim, Rastislav Rabatin, Zechun Liu, Yangyang Shi, Vikas Chandra |
EMNLP | 3 |
| 2024 | In-Context Prompt Editing for Conditional Audio GenerationabstractDistributional shift is a central challenge in the deployment of machine learning models as they can be ill-equipped for real-world data. This is particularly evident in text-to-audio generation where the encoded representations are easily undermined by unseen prompts, which leads to the degradation of generated audio — the limited set of the text-audio pairs remains inadequate for conditional audio generation in the wild as user prompts are under-specified. In particular, we observe a consistent audio quality degradation in generated audio samples with user prompts, as opposed to training set prompts. To this end, we present a retrieval-based in-context prompt editing framework that leverages the training captions as demonstrative exemplars to revisit the user prompts. We show that the framework enhanced the audio quality across the set of collected user prompts, which were edited with reference to the training captions as exemplars. Ernie Chang, Pin-Jie Lin, Yang Li 0183, Sidd Srinivasan, Gaël Le Lan, David Kant, Yangyang Shi, Forrest N. Iandola, Vikas Chandra |
ICASSP | 3 |
| 2024 | Folding Attention: Memory and Power Optimization for On-Device Transformer-Based Streaming Speech RecognitionabstractTransformer-based models excel in speech recognition. Existing efforts to optimize Transformer inference, typically for long-context applications, center on simplifying attention score calculations. However, streaming speech recognition models usually process a limited number of tokens each time, making attention score calculation less of a bottleneck. Instead, the bottleneck lies in the linear projection layers of multi-head attention and feedforward networks, constituting a substantial portion of the model size and contributing significantly to computation, memory, and power usage.To address this bottleneck, we propose folding attention, a technique targeting these linear layers, significantly reducing model size and improving memory and power efficiency. Experiments on on-device Transformer-based streaming speech recognition models show that folding attention reduces model size (and corresponding memory consumption) by up to 24% and power consumption by up to 23%, all without compromising model accuracy or computation overhead. Yang Li 0183, Liangzhen Lai, Yuan Shangguan, Forrest N. Iandola, Zhaoheng Ni, Ernie Chang, Yangyang Shi, Vikas Chandra |
ICASSP | 1 |
| 2023 | Factorized Blank Thresholding for Improved Runtime Efficiency of Neural TransducersabstractWe show how factoring the RNN-T’s output distribution can significantly reduce the computation cost and power consumption for on-device ASR inference with no loss in accuracy. With the rise in popularity of neural-transducer type models like the RNN-T for on-device ASR, optimizing RNN-T’s runtime efficiency is of great interest. While previous work has primarily focused on the optimization of RNN-T’s acoustic encoder and predictor, this paper focuses the attention on the joiner. We show that despite being only a small part of RNN-T, the joiner has a large impact on the overall model’s runtime efficiency. We propose to utilize HAT-style joiner factorization for the purpose of skipping the more expensive non-blank computation when the blank probability exceeds a certain threshold. Since the blank probability can be computed very efficiently and the RNN-T output is dominated by blanks, our proposed method leads to a 26-30% decoding speed-up and 43-53% reduction in on-device power consumption, all the while incurring no accuracy degradation and being relatively simple to implement. Frank Seide, Yang Li 0183, Kjell Schubert, Ozlem Kalinli, Michael L. Seltzer |
ICASSP | 4 |
| 2023 | Forecasting COVID-19 Dynamics: Clustering, Generalized Spatiotemporal Attention, and Impacts of Mobility and Geographic ProximityabstractForecasting the dynamics of COVID-19 enables government agencies and public health administrators to take proactive measures to combat the pandemic. This forecasting task faces several key challenges: First, the dynamics of COVID-19 exhibit complex spatial and temporal dependencies. The current growing trend at a location may be similar to that at another location in the past. Second, numerous factors, such as population mobility and geographic proximity between regions, mask usage, vaccine coverage, etc., significantly impact the dynamics. Third, we need to find the appropriate granularity for the forecasting task. The granularity should not be too coarse that we ignore the idiosyncrasies of individual regions. Still, the granularity should not be too fine that the prediction results are seriously vulnerable to noise.This paper addresses these challenges. We propose a simple but effective clustering algorithm that finds the appropriate granularity for the forecasting task. We invent generalized spatiotemporal attention, an attention mechanism that is generalized enough to capture the complex spatial and temporal dependencies and to flexibly account for intra- and inter-region characteristics such as geographic proximity and population mobility. Based on this generalized spatiotemporal attention, we designed COVID-Forecaster, a lightweight deep learning model for forecasting the dynamics of COVID-19. Experimental results demonstrate that COVID-Forecaster significantly outperforms state-of-the-art models. For example, COVID-Forecaster reduces the mean absolute percentage error (MAPE) by 6.8% and the weighted absolute percentage error (WAPE) by 13.5% in forecasting the COVID-19 dynamics at the 3141 counties of the United States. Yang Li 0183, José M. F. Moura |
ICDE | 2 |
| 2020 | Forecaster: A Graph Transformer for Forecasting Spatial and Time-Dependent DataabstractSpatial and time-dependent data is of interest in many applications. This task is difficult due to its complex spatial dependency, long-range temporal dependency, data non-stationarity, and data heterogeneity. To address these challenges, we propose Forecaster, a graph Transformer architecture. Specifically, we start by learning the structure of the graph that parsimoniously represents the spatial dependency between the data at different locations. Based on the topology of the graph, we sparsify the Transformer to account for the strength of spatial dependency, long-range temporal dependency, data non-stationarity, and data heterogeneity. We evaluate Forecaster in the problem of forecasting taxi ride-hailing demand and show that our proposed architecture significantly outperforms the state-of-the-art baselines. Yang Li 0183, José M. F. Moura |
ECAI | 1 |
| 2019 | A Scalable Priority-Aware Approach to Managing Data Center Server PowerabstractPower management is a key component of modern data center design. Power managers must (1) ensure the costand energy-efficient utilization of the data center infrastructure, (2) maintain availability of the services provided by the center, and (3) address environmental concerns associated with the center's power consumption. While several power management techniques have been proposed and deployed in production data centers, there are still many challenges to comprehensive data center power management. This is particularly true in public cloud environments, where different jobs have different priority levels, and where high availability is critical. One example of the challenges facing public cloud data centers involves power capping. As power delivery must be highly reliable and tolerate wide variation in the load drawn by the data center components, the power infrastructure (e.g., power supplies, circuit breakers, UPS) has high redundancy and overprovisioning. During normal operation (i.e., typical server power demands, and no failures in the center), the power infrastructure is significantly underutilized. Power capping is a common solution to reduce this underutilization, by allowing more servers to be added safely (i.e., without power shortfalls) to the existing power infrastructure, and throttling power consumption in the infrequent cases where the demanded power exceeds the provisioned power capacity to avoid shortfalls. However, state-of-the-art power capping solutions are (1) not directly applicable to the redundant power infrastructure used in highly-available data centers; and (2) oblivious to differing workload priorities across the entire center when power consumption needs to be throttled, which can unnecessarily slow down high-priority work. To address this need, we develop CapMaestro, a new power management architecture with three key features for public cloud data centers. First, CapMaestro is designed to work with multiple power feeds (i.e., sources), and exploits server-level power capping to independently cap the load on each feed of a server. Second, CapMaestro uses a scalable, global priority-aware power capping approach, which accounts for power capacity at each level of the power distribution hierarchy. It exploits the underutilization of commonly-employed redundant power infrastructure at each level of the hierarchy to safely accommodate a much greater number of servers. Third, CapMaestro exploits stranded power (i.e., power budgets that are not utilized) in redundant power infrastructure to boost the performance of workloads in the data center. We add CapMaestro to a real cloud data center control plane, and demonstrate the effectiveness of all three key features. Using a large-scale data center simulation, we demonstrate that CapMaestro significantly and safely increases the number of servers for existing infrastructure. We also call out other key technical challenges the industry faces in data center power management. Yang Li 0183, Charles Lefurgy, Karthick Rajamani, Malcolm Allen-Ware, Guillermo J. Silva, Daniel D. Heimsoth, Saugata Ghose, Onur Mutlu |
HPCA | 1 |
| 2018 | dCat: dynamic cache management for efficient, performance-sensitive infrastructure-as-a-serviceabstractIn the modern multi-tenant cloud, resource sharing increases utilization but causes performance interference between tenants. More generally, performance isolation is also relevant in any multi-workload scenario involving shared resources. Last level cache (LLC) on processors is shared by all CPU cores in x86, thus the cloud tenants inevitably suffer from the cache flush by their noisy neighbors running on the same socket. Intel Cache Allocation Technology (CAT) provides a mechanism to assign cache ways to cores to enable cache isolation, but its static configuration can result in underutilized cache when a workload cannot benefit from its allocated cache capacity, and/or lead to sub-optimal performance for workloads that do not have enough assigned capacity to fit their working set. Karthick Rajamani, Wes Felter, Juan C. Rubio, Yang Li 0183 |
EuroSys | 6 |
| 2017 | Utility-Based Hybrid Memory ManagementabstractWhile the memory footprints of cloud and HPC applications continue to increase, fundamental issues with DRAM scaling are likely to prevent traditional main memory systems, composed of monolithic DRAM, from greatly growing in capacity. Hybrid memory systems can mitigate the scaling limitations of monolithic DRAM by pairing together multiple memory technologies (e.g., different types of DRAM, or DRAM and non-volatile memory) at the same level of the memory hierarchy. The goal of a hybrid main memory is to combine the different advantages of the multiple memory types in a cost-effective manner while avoiding the disadvantages of each technology. Memory pages are placed in and migrated between the different memories within a hybrid memory system, based on the properties of each page. It is important to make intelligent page management (i.e., placement and migration) decisions, as they can significantly affect system performance.In this paper, we propose utility-based hybrid memory management (UH-MEM), a new page management mechanism for various hybrid memories, that systematically estimates the utility (i.e., the system performance benefit) of migrating a page between different memory types, and uses this information to guide data placement. UH-MEM operates in two steps. First, it estimates how much a single application would benefit from migrating one of its pages to a different type of memory, by comprehensively considering access frequency, row buffer locality, and memory-level parallelism. Second, it translates the estimated benefit of a single application to an estimate of the overall system performance benefit from such a migration.We evaluate the effectiveness of UH-MEM with various types of hybrid memories, and show that it significantly improves system performance on each of these hybrid memories. For a memory system with DRAM and non-volatile memory, UH-MEM improves performance by 14% on average (and up to 26%) compared to the best of three evaluated state-of-the-art mechanisms across a large number of data-intensive workloads. Yang Li 0183, Saugata Ghose, Jongmoo Choi, Onur Mutlu |
CLUSTER | 1 |
| 2016 | SizeCap: Efficiently handling power surges in fuel cell powered data centersabstractFuel cells are a promising power source for future data centers, offering high energy efficiency, low greenhouse gas emissions, and high reliability. However, due to mechanical limitations related to fuel delivery, fuel cells are slow to adjust to sudden increases in data center power demands, which can result in temporary power shortfalls. To mitigate the impact of power shortfalls, prior work has proposed to either perform power capping by throttling the servers, or to leverage energy storage devices (ESDs) that can temporarily provide enough power to make up for the shortfall while the fuel cells ramp up power generation. Both approaches have disadvantages: power capping conservatively limits server performance and can lead to service level agreement (SLA) violations, while ESD-only solutions must significantly overprovision the energy storage device capacity to tolerate the shortfalls caused by the worst-case (i.e., largest) power surges, which greatly increases the total cost of ownership (TCO). We propose SizeCap, the first ESD sizing framework for fuel cell powered data centers, which coordinates ESD sizing with power capping to enable a cost-effective solution to power shortfalls in data centers. SizeCap sizes the ESD just large enough to cover the majority of power surges, but not the worst-case surges that occur infrequently, to greatly reduce TCO. It then uses the smaller capacity ESD in conjunction with power capping to cover the power shortfalls caused by the worst-case power surges. As part of our new flexible framework, we propose multiple power capping policies with different degrees of awareness of fuel cell and workload behavior, and evaluate their impact on workload performance and ESD size. Using traces from Microsoft's production data center systems, we demonstrate that SizeCap significantly reduces the ESD size (by 85%ofor a workload with infrequent yet large power surges, and by 50% for a workload with frequent power surges) without violating any SLAs. Yang Li 0183, Di Wang 0003, Saugata Ghose, Jie Liu 0001, Sriram Govindan, Sean James, Eric Peterson, John Siegler, Rachata Ausavarungnirun, Onur Mutlu |
HPCA | 1 |
| 2015 | Domain-Alternated Optimization for Passive MacromodelingabstractPassivity enforcement is an important issue for macromodeling for passive systems from measured or simulated data. Existing passivity enforcement techniques based on iteratively fixing the passivity either suffer from convergence issue or lack optimality that will sometimes lead to unacceptable error. In addition to the traditional two-stage (fitting plus enforcement) schemes, we propose a postenforcement optimization, which takes a passive yet not necessarily accurate model as the starting point, and performs local search to find the local optimum. A new technique, called domain-alternated optimization is proposed to eliminate passivity constraints while still guarantees strict passivity during the optimization. Experiments show that taking the models generated from existing enforcement methods, the proposed method can provide significant improvement on accuracy. The proposed method is efficient and can deal with problems up to a few tens of thousands of variables. Zuochang Ye, Yang Li 0183 |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2013 | An accurate semi-analytical framework for full-chip TSV-induced stress modelingabstractTSV-induced stress is an important issue in 3D IC design since it leads to serious reliability problems and influences device performance. Existing finite element method can provide accurate analysis for the stress of simple TSV placement, but is not scalable to larger designs due to its expensive memory consumption and high run time. On the contrary, linear superposition method is efficient to analyze stress in full-chip scale, but sometimes it fails to provide an accurate estimation since it neglects the stress induced by interactions between TSVs. In this paper we propose an accurate two-stage semi-analytical framework for full-chip TSV-induced stress modeling. In addition to the linear superposition, we characterize the stress induced by interactions between TSVs to provide more accurate full-chip modeling. Experimental results demonstrate that the proposed framework can significantly improve the accuracy of linear superposition method with reasonable overhead in run time. Yang Li 0183, David Z. Pan |
DAC | 1 |
| 2013 | Time-domain segmentation based massively parallel simulation for ADCsabstractThe great availability of massively parallel computing platforms gives rise a question to the EDA industry--how can this be really helping the productivity of circuit designs. Scalability of traditional parallel methods have shown to be limited as the computational resources keep increasing. In this paper we propose a time-domain segmentation method for massively parallel transistor-level simulation for short-memory circuits. SNDR simulation for ADCs is selected as the application as ADCs are typical short-memory circuits and the SNDR simulation is very time consuming. Experiments with realistic Flash and SAR ADCs demonstrate 64x-78x speed-ups with 100 CPU cores. With minor, yet important modifications, the proposed method can even be applied to simulation of Σ-Δ modulator, which does not satisfy the short-memory condition due to the presence of integrator, and 52x speed-up is observed with 100 CPU cores. The implementation of the proposed method is extremely simple and no modification to simulator is needed. Zuochang Ye, Bichen Wu, Yang Li 0183 |
DAC | 4 |
| 2011 | A novel framework for passive macro-modelingabstractPassivity enforcement is an important issue for macro-modeling for passive systems from measured or simulated data. Existing convex programming based methods are too expensive and thus are ruled out for realistic application. Other methods based on iteratively fixing the passivity through perturbing the eigenvalues of the Hamiltonian matrix either suffer from convergence issue or lack optimality which will sometimes lead to unacceptable error. In this paper we propose a novel framework for macro-modeling. In addition to the traditional two-stage (fixing plus enforcement) schemes, we propose a post-enforcement optimization, which takes a passive, while potentially not-so-accurate model, as the starting point, and performs local search to find the local optimum with passivity constraint or build-in passivity guarantee. A simple yet stable passive modeling generator is proposed to produce the starting model for optimization. Two algorithms are proposed for performing constrained and unconstrained optimizations. Experiments show that the accuracy of passivity-fixed model can be significantly improved with the proposed methods. Zuochang Ye, Yang Li 0183, Mingzhi Gao, Zhiping Yu |
DAC | 2 |