EDBT 2026 Demo / reviewers in the wild / expert
Yinghai Lu
dblp:68/7450
· DBLP profile ↗
16ranked-venue papers
5as first author
2since 2021 · last 2024
—ORCID · unresolved
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 14 · 5 first-author · 1 since 2021Software engineering, systems software and programming languages · 2 · 1 first-author · 1 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Databases, data mining, and information retrieval
2 papers |
Recommender systems · 58% Transaction processing and concurrency control · 38% Database system architecture and tuning · 4% | |
| Software engineering, system software, and programming languages
1 paper |
Runtime systems and virtual machines · 50% Compilers and program optimization · 50% | |
| Computer architecture, parallel and distributed computing, and storage systems
4 papers |
Electronic design automation · 71% Parallel and multicore computing · 18% Hardware reliability and fault tolerance · 8% | |
| Artificial intelligence
2 papers |
Deep learning architectures and training · 100% |
Topics — the 17 heaviest of 21, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Deep learning architectures and training
transformer |
0.8 | 1 | 2024 | Actions Speak Louder than Words: Trillion-Parameter Sequential Transducers for Generative Recommendations · ICML 2024 |
Recommender systems
generative recommendation |
0.8 | 1 | 2024 | Actions Speak Louder than Words: Trillion-Parameter Sequential Transducers for Generative Recommendations · ICML 2024 |
Recommender systems
sequential recommendation |
0.8 | 1 | 2024 | Actions Speak Louder than Words: Trillion-Parameter Sequential Transducers for Generative Recommendations · ICML 2024 |
Runtime systems and virtual machines › dynamic compilation
just-in-time compilation |
0.8 | 1 | 2024 | PyTorch 2: Faster Machine Learning Through Dynamic Python Bytecode Transformation and Graph Compilation · ASPLOS (2) 2024 |
Transaction processing and concurrency control
concurrency control |
0.3 | 1 | 2018 | Rethinking Concurrency Control for In-Memory OLAP DBMSs · ICDE 2018 |
Transaction processing and concurrency control › synchronization
lock-free algorithms |
0.3 | 1 | 2018 | Rethinking Concurrency Control for In-Memory OLAP DBMSs · ICDE 2018 |
Transaction processing and concurrency control › isolation levels
snapshot isolation |
0.3 | 1 | 2018 | Rethinking Concurrency Control for In-Memory OLAP DBMSs · ICDE 2018 |
Machine learning › Deep learning architectures and training › deep learning systems
deep learning framework |
0.2 | 1 | 2024 | PyTorch 2: Faster Machine Learning Through Dynamic Python Bytecode Transformation and Graph Compilation · ASPLOS (2) 2024 |
Electronic design automation
design optimization |
0.2 | 2 | 2010 | Multicore Parallelization of Min-Cost Flow for CAD Applications · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2010 Multicore parallel min-cost flow algorithm for CAD applications · DAC 2009 |
Electronic design automation › design optimization
min-cost flow |
0.2 | 2 | 2010 | Multicore Parallelization of Min-Cost Flow for CAD Applications · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2010 Multicore parallel min-cost flow algorithm for CAD applications · DAC 2009 |
Electronic design automation
timing analysis |
0.2 | 2 | 2011 | Optimal multi-domain clock skew scheduling · DAC 2011 Statistical reliability analysis under process variation and aging effects · DAC 2009 |
Electronic design automation › physical design › clock network synthesis
clock skew optimization |
0.1 | 1 | 2011 | Optimal multi-domain clock skew scheduling · DAC 2011 |
Electronic design automation
physical design |
0.1 | 1 | 2011 | Optimal multi-domain clock skew scheduling · DAC 2011 |
Parallel and multicore computing › parallel computing
multicore parallelism |
0.1 | 1 | 2010 | Multicore Parallelization of Min-Cost Flow for CAD Applications · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2010 |
Parallel and multicore computing
parallel programming models |
0.0 | 1 | 2010 | Multicore Parallelization of Min-Cost Flow for CAD Applications · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2010 |
Parallel and multicore computing
transactional memory |
0.0 | 1 | 2010 | Multicore Parallelization of Min-Cost Flow for CAD Applications · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2010 |
Integrated circuit design › process-voltage-temperature variation
threshold voltage variation |
0.0 | 1 | 2009 | Statistical reliability analysis under process variation and aging effects · DAC 2009 |
Methods — techniques the papers use, named apart from their topics
scaling laws · 1.5just-in-time compilation · 1.5graph compilation · 1.5generative modeling · 1.5pruning · 0.1integer linear programming · 0.1runtime scheduling · 0.1nondeterministic transactional models · 0.1statistical analysis · 0.1nondeterministic transactional algorithm · 0.1negative bias temperature instability modeling · 0.1multicore parallelization · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | PyTorch 2: Faster Machine Learning Through Dynamic Python Bytecode Transformation and Graph CompilationabstractThis paper introduces two extensions to the popular PyTorch machine learning framework, TorchDynamo and TorchInductor, which implement the torch.compile feature released in PyTorch 2. TorchDynamo is a Python-level just-in-time (JIT) compiler that enables graph compilation in PyTorch programs without sacrificing the flexibility of Python. It achieves this by dynamically modifying Python bytecode before execution and extracting sequences of PyTorch operations into an FX graph, which is then JIT compiled using one of many extensible backends. TorchInductor is the default compiler backend for TorchDynamo, which translates PyTorch programs into OpenAI's Triton for GPUs and C++ for CPUs. Results show that TorchDynamo is able to capture graphs more robustly than prior approaches while adding minimal overhead, and TorchInductor is able to provide a 2.27× inference and 1.41× training geometric mean speedup on an NVIDIA A100 GPU across 180+ real-world models, which outperforms six other compilers. These extensions provide a new way to apply optimizations through compilers in eager mode frameworks like PyTorch. Jason Ansel, Edward Z. Yang, Horace He, Natalia Gimelshein, Animesh Jain, Michael Voznesensky, Bin Bao, Peter Bell 0008, David Berard, Evgeni Burovski, Geeta Chauhan, Anjali Chourdia, Will Constable, Alban Desmaison, Zach DeVito, Elias Ellison, Will Feng, Jiong Gong, Michael Gschwind, Brian Hirsh, Sherlock Huang, Kshiteej Kalambarkar, Laurent Kirsch, Michael Lazos, Mario Lezcano Casado, Yanbo Liang, Jason Liang, Yinghai Lu, C. K. Luk, Bert Maher, Yunjie Pan, Christian Puhrsch, Matthias Reso, Mark Saroufim, Marcos Yukio Siraichi, Helen Suk, Shunting Zhang, Michael Suo, Phil Tillet, Xu Zhao 0004, Eikan Wang, Keren Zhou 0001, Richard Zou, Ajit Mathews, Xiaoquan Wen, Gregory Chanan, Peng Wu 0001, Soumith Chintala |
ASPLOS (2) | 28 |
| 2024 | Actions Speak Louder than Words: Trillion-Parameter Sequential Transducers for Generative RecommendationsabstractLarge-scale recommendation systems are characterized by their reliance on high cardinality, heterogeneous features and the need to handle tens of billions of user actions on a daily basis. Despite being trained on huge volume of data with thousands of features, most Deep Learning Recommendation Models (DLRMs) in industry fail to scale with compute. Inspired by success achieved by Transformers in language and vision domains, we revisit fundamental design choices in recommendation systems. We reformulate recommendation problems as sequential transduction tasks within a generative modeling framework (``Generative Recommenders''), and propose a new architecture, HSTU, designed for high cardinality, non-stationary streaming recommendation data. HSTU outperforms baselines over synthetic and public datasets by up to 65.8% in NDCG, and is 5.3x to 15.2x faster than FlashAttention2-based Transformers on 8192 length sequences. HSTU-based Generative Recommenders, with 1.5 trillion parameters, improve metrics in online A/B tests by 12.4% and have been deployed on multiple surfaces of a large internet platform with billions of users. More importantly, the model quality of Generative Recommenders empirically scales as a power-law of training compute across three orders of magnitude, up to GPT-3/LLaMa-2 scale, which reduces carbon footprint needed for future model developments, and further paves the way for the first foundation models in recommendations. Jiaqi Zhai, Lucy Liao, Xuan Cao, Leon Gao, Zhaojie Gong, Fangda Gu, Jiayuan He 0008, Yinghai Lu |
ICML | 11 |
| 2018 | Rethinking Concurrency Control for In-Memory OLAP DBMSsabstractAlthough OLTP and OLAP database systems have fundamentally disparate architectures, most research work on concurrency control is geared towards transactional systems and simply adopted by OLAP DBMSs. In this paper we describe a new concurrency control protocol specifically designed for analytical DBMSs that can provide Snapshot Isolation for distributed in-memory OLAP database systems, called Append-Only Snapshot Isolation (AOSI). Unlike previous work, which are either based on multiversion concurrency control (MVCC) or Two Phase Locking (2PL), AOSI is completely lock-free and always maintains a single version of each data item. In addition, it removes the need for per-record timestamps of traditional MVCC implementations and thus considerably reduces the memory overhead incurred by concurrency control. In order to support these characteristics, the protocol sacrifices flexibility and removes support for a few operations, particularly record updates and single record deletions; however, we argue that even though these operations are essential in a pure transactional system, they are not strictly required in most analytic pipelines and OLAP systems. We also present an experimental evaluation of AOSI's current implementation within the Cubrick in-memory OLAP DBMS at Facebook, and show that lock-free single-version Snapshot Isolation can be achieved with low memory overhead and minor impact in query latency. Pedro Pedreira, Yinghai Lu, Sergey Pershin, Chris Croswhite |
ICDE | 2 |
| 2014 | Optimal and Efficient Algorithms for Multidomain Clock Skew SchedulingabstractClock skew scheduling is an effective technique to improve the performance of sequential circuits. However, with process variations, it becomes more difficult to implement a large number of clock delays in a precise manner. Multidomain clock skew scheduling (MDCSS) is one way to overcome this limitation. In this paper, we prove the NP-completeness of multidomain clock scheduling problem and design a practical optimal algorithm to solve it. Given the domain number, we bound the number of all possible skew assignments and develop an optimal algorithm with efficient pruning techniques as well as a very efficient heuristics based on the optimal framework. The experimental results on ISCAS89 sequential benchmarks show the optimality and efficiency of our method compared with the most recent approaches to MDCSS. Li Li 0021, Yinghai Lu, Hai Zhou 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2013 | Retiming for Soft Error Minimization Under Error-Latching Window ConstraintsabstractSoft error has become a critical reliability issue in nanoscale integrated circuits, especially in sequential circuits where a latched error will be propagated for many cycles and affect many outputs at different time. Retiming is a structural operation that relocates registers in a circuit without changing its functionality. In this paper, the effect of retiming on soft error rate (SER) of a sequential circuit is investigated considering both logic masking and timing masking. A minimum observability retiming problem under error-latching window constraints is formulated to reduce the SER of the circuit. And an efficient algorithm is proposed to solve the problem optimally. Experimental results show on average a 32.7% reduction on SER from the original circuits and a 15% improvement over the existing method. Yinghai Lu, Hai Zhou 0001 |
DATE | 1 |
| 2013 | Post-routing layer assignment for double patterning with timing critical paths consideration
Jian Sun 0005, Yinghai Lu, Hai Zhou 0001, Changhao Yan, Xuan Zeng 0001 |
Integr. | 2 |
| 2012 | Optimal prescribed-domain clock skew schedulingabstractClock skew scheduling is an efficient technique to minimize the cycle period by properly assigning clock delays to registers in a circuit. But its effectiveness is limited by the difficulty in implementing a large number of arbitrary clock skews. Multi-domain clock skew scheduling and prescribed-domain clock skew scheduling are two alternatives to overcome this shortage by restricting the number of clock domains. While multi-domain clock skew scheduling has been proved to be NP-hard, the hardness of prescribed-domain clock skew scheduling algorithm remains evasive. In this paper, we give a positive answer to the open question by presenting the first efficient and optimal algorithm for prescribed-domain clock skew scheduling. Besides the runtime improvement over the previous method, the experimental results on ISCAS89 benchmarks show comparable quality to those generated by optimal multi-domain clock skew scheduling. Li Li 0021, Yinghai Lu, Hai Zhou 0001 |
ASP-DAC | 2 |
| 2012 | An efficient algorithm for library-based cell-type selection in high-performance low-power designsabstractIn this paper, we present a complete framework for cell-type selection in modern high-performance low-power designs with library-based timing model. Our framework can be divided into three stages. First, the best design performance with all possible cell-types is achieved by a Minimum Clock Period Lagrangian Relaxation (Min-Clock LR) method, which extends the traditional LR approach to conquer the difficulties in discrete scenario. Min-Clock LR fully leverages the prevalent many-core systems as the main body of its workload is composed of independent tasks. Upon a timing-valid design, we solve the timing-constrained power optimization problem by min-cost network flow. Especially, we identify and address the core issues in applying network flow technique to library-based timing model. Finally, a power prune technique is developed to take advantage of the residual slacks due to the conservative network flow construction. Experiments on ISPD 2012 benchmarks show that on average we can save 13% more leakage power on designs with fast timing constraints compared to start-of-the-art techniques. Moreover, our algorithm shows a linear empirical runtime, finishing the largest benchmark with one million cells in 1.5 hours. Li Li 0021, Yinghai Lu, Hai Zhou 0001 |
ICCAD | 3 |
| 2012 | Efficient design space exploration for component-based system designabstractAs the technology scaling down continues to go beyond 22nm, the increasing transistor density on a single die is leading towards more and more complex systems-on-chip. Designers are faced with the challenge of how to efficiently design such a complicated system with tight time-to-market constraints. Component-based system design and design space exploration are two key techniques to overcoming the challenge. In this paper, we model the design space exploration of a system with difference constraints as a bi-criteria convex cost flow problem and develop an efficient solver for it based on parametric simplex method. Furthermore, considering the high cost of synthesizing the underlying soft IP cores, we propose an online algorithm to incrementally refine the system-level Pareto curves as more component-wise sampling points are added. The experimental results demonstrate the efficiency and effectiveness of the proposed algorithms. Yinghai Lu, Hai Zhou 0001 |
ICCAD | 1 |
| 2011 | Low power discrete voltage assignment under clock skew schedulingabstractMultiple Supply Voltage (MSV) assignment has emerged as an appealing technique in low power IC design, due to its flexibility in balancing power and performance. However, clock skew scheduling, which has great impact on criticality of combinational paths in sequential circuit, has not been explored in the merit of MSV assignment. In this paper, we propose a discrete voltage assignment algorithm for sequential circuit under clock scheduling. The sequential MSV assignment problem is first formulated as a convex cost dual network flow problem, which can be optimally solved in polynomial time assuming delay of each gate can be chosen in continuous domain. Then a mincut-based heuristic is designed to convert the unfeasible continuous solution into feasible discrete solution while largely preserving the global optimality. Besides, we revisit the hardness of the general discrete voltage assignment problem and point out some misunderstandings on the approximability of this problem in previous related work. Benchmark test for our algorithm shows 9.2% reduction in power consumption on average, in compared with combinational MSV assignment. Referring to the continuous solution obtained from network flow as the lower bound, the gap between our solution and the lower bound is only 1.77%. Li Li 0021, Jian Sun 0005, Yinghai Lu, Hai Zhou 0001, Xuan Zeng 0001 |
ASP-DAC | 3 |
| 2011 | Post-routing layer assignment for double patterningabstractDouble patterning lithography, where one-layer layout is decomposed into two masks, is believed to be inevitable for 32nm technology node of the ITRS roadmap. However, post-routing layer assignment, which decides the layout pattern on each layer, thus having great impact on double patterning related parameters, has not been explored in the merit of double patterning. In this paper, we propose a post-routing layer assignment algorithm for double patterning optimization. Our solution consists of three major phases: multi-layer assignment, single-layer double patterning, and via reduction. For phase one and three, multi-layer graph is constructed and dynamic programming is employed to solve optimization problem on this graph. In the second phase, single-layer double patterning is proved NP-hard and existing algorithm is implemented to optimize single layer double patterning problem. The proposed method is tested on CBL (Collaborative Benchmarking Laboratory) benchmarks and shows great performance. In comparison with single-layer double patterning, our method achieves 73% and 27% average reduction for unresolvable conflicts and stitches respectively, with only 9% increase of via number. When double patterning is constrained on only the bottom two metal layers as in current technology, these numbers become 62%, 8% and 0.42%. Jian Sun 0005, Yinghai Lu, Hai Zhou 0001, Xuan Zeng 0001 |
ASP-DAC | 2 |
| 2011 | Parallel cross-layer optimization of high-level synthesis and physical designabstractIntegrated circuit (IC) design automation has traditionally followed a hierarchical approach. Modern IC design flow is divided into sequentially-addressed design and optimization layers; each successively finer in design detail and data granularity while increasing in computational complexity. Eventual agreement across the design layers signals design closure. Obtaining design closure is a continual problem, as lack of awareness and interaction between layers often results in multiple design flow iterations. In this work, we propose parallel cross-layer optimization, in which the boundaries between design layers are broken, allowing for a more informed and efficient exploration of the design space. We leverage the heterogeneous parallel computational power in current and upcoming multi-core/many-core computation platforms to suite the heterogeneous characteristics of multiple design layers. Specifically, we unify the highlevel and physical synthesis design layers for parallel cross-layer IC design optimization. In addition, we introduce a massively-parallel GPU floorplanner with local and global convergence test as the proposed physical synthesis design layer. Our results show average performance gains of 11X speed-up over state-of-the-art. James Williamson, Yinghai Lu, Hai Zhou 0001, Xuan Zeng 0001 |
ASP-DAC | 2 |
| 2011 | Optimal multi-domain clock skew schedulingabstractClock skew scheduling is an effective technique to improve the performance of sequential circuits. However, with process variations, it becomes more difficult to implement a large number of clock delays in a precise manner. Multi-domain clock skew scheduling is one way to overcome this limitation. In this paper, we prove the NP-completeness of multi-domain clock scheduling problem, and design a practical optimal algorithm to solve it. Given the domain number, we bound the number of all possible skew assignments and develop an optimal algorithm with efficient pruning techniques. Experiment results on ISCAS89 sequential benchmarks show the optimality and efficiency of our method compared with existing approaches. Li Li 0021, Yinghai Lu, Hai Zhou 0001 |
DAC | 2 |
| 2010 | Multicore Parallelization of Min-Cost Flow for CAD ApplicationsabstractComputational complexity has been the primary challenge of many very large scale integration computer-aided design (CAD) applications. The emerging multicore and many-core microprocessors have the potential to offer scalable performance improvements. How to explore the multicore resources to speed up CAD applications is thus a natural question but also a huge challenge for CAD researchers. This paper proposes a methodology to explore concurrency via nondeterministic transactional models, and to program them on multicore processors for CAD applications. Various run-time scheduling implementations on multicore shared-memory machines are discussed and the most efficient one is identified. The proposed methodology is applied to the min-cost flow problem which has been identified as the key problem in many design optimizations, from wire-length optimization in detailed placement to timing-constrained voltage assignment. A concurrent algorithm for min-cost flow has been developed based on the methodology. Experiments on voltage island generation in floorplanning have demonstrated its efficiency and scalable speedup over different numbers of cores. Yinghai Lu, Hai Zhou 0001, Xuan Zeng 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2009 | Statistical reliability analysis under process variation and aging effectsabstractCircuit reliability is affected by various fabrication-time and run-time effects. Fabrication-induced process variation has significant impact on circuit performance and reliability. Various aging effects, such as negative bias temperature instability, cause continuous performance and reliability degradation during circuit run-time usage. In this work, we present a statistical analysis framework that characterizes the lifetime reliability of nanometer-scale integrated circuits by jointly considering the impact of fabrication-induced process variation and run-time aging effects. More specifically, our work focuses on characterizing circuit threshold voltage lifetime variation and its impact on circuit timing due to process variation and the negative bias temperature instability effect, a primary aging effect in nanometer-scale integrated circuits. The proposed work is capable of characterizing the overall circuit lifetime reliability, as well as efficiently quantifying the vulnerabilities of individual circuit elements. This analysis framework has been carefully validated and integrated into an iterative design flow for circuit lifetime reliability analysis and optimization. Yinghai Lu, Hai Zhou 0001, Hengliang Zhu, Fan Yang 0001, Xuan Zeng 0001 |
DAC | 1 |
| 2009 | Multicore parallel min-cost flow algorithm for CAD applicationsabstractComputational complexity has been the primary challenge of many VLSI CAD applications. The emerging multicore and many-core microprocessors have the potential to offer scalable performance improvement. How to explore the multicore resources to speed up CAD applications is thus a natural question but also a huge challenge for CAD researchers. Indeed, decades of work on general-purpose compilation approaches that automatically extracts parallelism from a sequential program has shown limited success. Past work has shown that programming model and algorithm design methods have a great influence on usable parallelism. In this paper, we propose a methodology to explore concurrency via nondeterministic transactional algorithm design, and to program them on multicore processors for CAD applications. We apply the proposed methodology to the min-cost flow problem which has been identified as the key problem in many design optimizations, from wire-length optimization in detailed placement to timing-constrained voltage assignment. A concurrent algorithm and its implementation on multicore processors for min-cost flow have been developed based on the methodology. Experiments on voltage island generation in floorplanning demonstrated its efficiency and scalable speedup over different number of cores. Yinghai Lu, Hai Zhou 0001, Xuan Zeng 0001 |
DAC | 1 |