VLDB 2026 Research / reviewers in the wild / expert
Lei Lan
dblp:49/3267
· DBLP profile ↗
20ranked-venue papers
9as first author
15since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 17 · 8 first-author · 14 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Systems, architecture and hardware · 1Computer networks · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | ElastoGen: 4D Generative ElastodynamicsabstractWe present ElastoGen, a knowledge-driven AI model that generates physically accurate 4D elastodynamics. Unlike deep models that learn from video- or image-based observations, ElastoGen leverages the principles of physics and learns from established mathematical and optimization procedures. The core idea of ElastoGen is converting the differential equation, corresponding to the nonlinear force equilibrium, into a series of iterative local convolution-like operations, which naturally fit deep architectures. We carefully build our network module following this overarching design philosophy. ElastoGen is much more lightweight in terms of both training requirements and network scale than deep generative models. Because of its alignment with actual physical procedures, ElastoGen efficiently generates accurate dynamics for a wide range of hyperelastic materials and can be easily integrated with upstream and downstream deep modules to enable end-to-end 4D generation. Yutao Feng, Yintong Shang, Xiang Feng 0004, Lei Lan, Shandian Zhe, Tianjia Shao, Hongzhi Wu, Kun Zhou 0001, Chenfanfu Jiang, Yin Yang 0002 |
AAAI | 4 |
| 2026 | JGS2-GQ: Training-free 2nd Jacobi with Gaussian QuadratureabstractJGS2 is a Jacobi-like GPU simulation algorithm. It avoids the overshooting issue by augmenting each subproblem with a perturbation subspace that predicts the global influence of the local solve. The efficiency of JGS2 is due to Cubature-based subspace integration at each subproblem. Being a data-driven method, Cubature requires a set of representative deformed poses that cover deformations likely to occur in the simulation. This requirement is unlikely for high-resolution deformation with rich local details. Therefore, simulation performance and convergence degenerate when Cubature extrapolates. This paper proposes a training-free subspace integration algorithm based on classic Gaussian quadrature (GQ). We leverage the fact that the subproblem's subspace bases can be well-approximated by a low-degree multivariable polynomial, which suggests GQ an excellent candidate for Cubature substitute. To this end, we introduce a novel algorithm that adaptively generates the integration region for each subproblem. As a result, GQ integration can be analytically retrieved without cumbersome data generation and training. We also show how to handle frictional contact by modifying the pre-computed perturbation subspace. The resulting JGS2-GQ framework is more versatile than the vanilla JGS2 method. It is more stable for large and novel deformations, and is free of data generation and expensive training, while maintaining a near second-order convergence that is comparable to Newton. Performance-wise, JGS2-GQ is as efficient as JGS2, which is three orders faster than classic CPU methods and up to two orders faster than classic GPU algorithms. When novel deformation occurs, JGS2-GQ outperforms JGS2 over 50%. Dewen Guo, Yuqi Meng, Lei Lan, Weiwei Xu 0003, Chenfanfu Jiang, Yin Yang 0002 |
ACM Trans. Graph. | 6 |
| 2025 | JGS2: Near Second-order Converging Jacobi/Gauss-Seidel for GPU ElastodynamicsabstractIn parallel simulation, convergence and parallelism are often seen as inherently conflicting objectives. Improved parallelism typically entails lighter local computation and weaker coupling, which unavoidably slow the global convergence. This paper presents a novel GPU algorithm that achieves convergence rates comparable to fullspace Newton's method while maintaining good parallelizability just like the Jacobi method. Our approach is built on a key insight into the phenomenon of overshoot. Overshoot occurs when a local solver aggressively minimizes its local energy without accounting for the global context, resulting in a local update that undermines global convergence. To address this, we derive a theoretically second-order optimal solution to mitigate overshoot. Furthermore, we adapt this solution into a pre-computable form. Leveraging Cubature sampling, our runtime cost is only marginally higher than the Jacobi method, yet our algorithm converges nearly quadratically as Newton's method. We also introduce a novel full-coordinate formulation for more efficient pre-computation. Our method integrates seamlessly with the incremental potential contact method and achieves second-order convergence for both stiff and soft materials. Experimental results demonstrate that our approach delivers high-quality simulations and outperforms state-of-the-art GPU methods with 50× to 100× better convergence. Lei Lan, Chun Yuan 0001, Weiwei Xu 0003, Hao Su 0001, Huamin Wang 0001, Chenfanfu Jiang, Yin Yang 0002 |
ACM Trans. Graph. | 1 |
| 2025 | High-performance CPU Cloth Simulation Using Domain-decomposed Projective DynamicsabstractWhenever the concept of high-performance cloth simulation is brought up, GPU acceleration is almost always the first that comes to mind. Leveraging immense parallelization, GPU algorithms have demonstrated significant success recently, whereas CPU methods are somewhat overlooked. Indeed, the need for an efficient CPU simulator is evident and pressing. In many scenarios, high-end GPUs may be unavailable or are already allocated to other tasks, such as rendering and shading. A high-performance CPU alternative can greatly boost the overall system capability and user experience. Inspired by this demand, this paper proposes a CPU algorithm for high-resolution cloth simulation. By partitioning the garment model into multiple (but not massive) sub-meshes or domains, we assign per-domain computations to individual CPU processors. Borrowing the idea of projective dynamics that breaks the computation into global and local steps, our key contribution is a new parallelization paradigm at domains for both global and local steps so that domain-level calculations are sequential and lightweight. The CPU has much fewer processing units than a GPU. Our algorithm mitigates this disadvantage by wisely balancing the scale of the parallelization and convergence. We validate our method in a wide range of simulation problems involving high-resolution garment models. Performance-wise, our method is at least one order faster than existing CPU methods, and it delivers a similar performance compared with the state-of-the-art GPU algorithms in many examples, but without using a GPU. Lei Lan, Huamin Wang 0001, Yuko Ishiwaka, Chenfanfu Jiang, Kui Wu 0003, Yin Yang 0002 |
ACM Trans. Graph. | 3 |
| 2024 | XPBI: Position-Based Dynamics with Smoothing Kernels Handles Continuum InelasticityabstractPBD and its extension, XPBD, have been predominantly applied to compliant constrained elastodynamics, with their potential in finite strain (visco-) elastoplasticity remaining underexplored. XPBD is often perceived to stand in contrast to other meshless methods, such as the MPM. MPM is based on discretizing the weak form of governing partial differential equations within a continuum domain, coupled with a hybrid Lagrangian-Eulerian method for tracking deformation gradients. In contrast, XPBD formulates specific constraints, whether hard or compliant, to positional degrees of freedom. We revisit this perception by investigating the potential of XPBD in handling inelastic materials that are described with classical continuum mechanics-based yield surfaces and elastoplastic flow rules. Our inspiration is that a robust estimation of the velocity gradient is a sufficiently useful key to effectively tracking deformation gradients in XPBD simulations. By further incorporating implicit inelastic constitutive relationships, we introduce a plasticity in-the-loop updated Lagrangian augmentation to XPBD. This enhancement enables the simulation of elastoplastic, viscoplastic, and granular substances following their standard constitutive laws. We demonstrate the effectiveness of our method through high-resolution and real-time simulations of diverse materials such as snow, sand, and plasticine, and its integration with standard XPBD simulations of cloth and water. Chang Yu 0005, Xuan Li 0015, Lei Lan, Yin Yang 0002, Chenfanfu Jiang |
SIGGRAPH Asia | 3 |
| 2024 | A path splitting and pruning strategy on list decoder for PAC codesabstractAbstract Recently, Arıkan proposed a polarization‐adjusted convolutional (PAC) codes, demonstrating their superior error correction performance over polar codes at short block lengths. It was confirmed that PAC codes approached the optimal performance achievable with limited code length. This paper proposes a novel low‐complexity list decoding algorithm for PAC codes, incorporating path splitting and pruning strategies based on a set of highly reliable information bits. Simulation results reveal that the proposed algorithm significantly reduces sorting complexity and average list size, all while incurring negligible performance loss. Unlike previous pruning algorithms designed for polar codes, the proposed strategy eliminates the need to individually assess the reliability of decoding paths in each decoding process. Instead, the algorithm minimizes redundant decoding paths through a high‐reliability information bit set, constructed using Monte Carlo experiments. Lei Lan, Zhongpeng Wang, Lijuan Zhang 0004 |
IET Commun. | 1 |
| 2024 | Efficient GPU Cloth Simulation with Non-distance Barriers and Subspace ReuseabstractThis paper pushes the performance of cloth simulation, making the simulation interactive even for high-resolution garment models while keeping every triangle untangled. The penetration-free guarantee is inspired by the interior point method, which converts the inequality constraints to barrier potentials. We propose a major overhaul of this modality within the projective dynamics framework by leveraging an adaptive weighting mechanism inspired by barrier formulation. This approach does not depend on the distance between mesh primitives, but on the virtual life span of a collision event and thus keeps all the vertices within feasible region. Such a non-distance barrier model allows a new way to integrate collision resolution into the simulation pipeline. Another contributor to the performance boost comes from the subspace reuse strategy. This is based on the observation that low-frequency strain propagation is near orthogonal to the deformation induced by collisions or self-collisions, often of high frequency. Subspace reuse then takes care of low-frequency residuals, while high-frequency residuals can also be effectively smoothed by GPU-based iterative solvers. We show that our method outperforms existing fast cloth simulators by at least one order while producing high-quality animations of high-resolution models. Lei Lan, Jingyi Long, Chun Yuan 0001, Xuan Li 0015, Xiaowei He 0004, Huamin Wang 0001, Chenfanfu Jiang, Yin Yang 0002 |
ACM Trans. Graph. | 1 |
| 2024 | Volumetric Homogenization for Knitwear SimulationabstractThis paper presents volumetric homogenization, a spatially varying homogenization scheme for knitwear simulation. We are motivated by the observation that macro-scale fabric dynamics is strongly correlated with its underlying knitting patterns. Therefore, homogenization towards a single material is less effective when the knitting is complex and non-repetitive. Our method tackles this challenge by homogenizing the yarn-level material locally at volumetric elements. Assigning a virtual volume of a knitting structure enables us to model bending and twisting effects via a simple volume-preserving penalty and thus effectively alleviates the material nonlinearity. We employ an adjoint Gauss-Newton formulation[Zehnder et al. 2021] to battle the dimensionality challenge of such per-element material optimization. This intuitive material model makes the forward simulation GPU-friendly. To this end, our pipeline also equips a novel domain-decomposed subspace solver crafted for GPU projective dynamics, which makes our simulator hundreds of times faster than the yarn-level simulator. Experiments validate the capability and effectiveness of volumetric homogenization. Our method produces realistic animations of knitwear matching the quality of full-scale yarn-level simulations. It is also orders of magnitude faster than existing homogenization techniques in both the training and simulation stages. Chun Yuan 0001, Haoyang Shi, Lei Lan, Yuxing Qiu, Cem Yuksel, Huamin Wang 0001, Chenfanfu Jiang, Kui Wu 0003, Yin Yang 0002 |
ACM Trans. Graph. | 3 |
| 2023 | Subspace-Preconditioned GPU Projective Dynamics with Contact for Cloth SimulationabstractWe propose an efficient cloth simulation method that combines the merits of two drastically different numerical procedures, namely the subspace integration and parallelizable iterative relaxation. We show those two methods can be organically coupled within the framework of projective dynamics (PD), where both low- and high-frequency cloth motions are effectively and efficiently computed. Our method works seamlessly with the state-of-the-art contact handling algorithm, the incremental potential contact (IPC), to offer the non-penetration guarantee of the resulting animation. Our core ingredient centers around the utilization of subspace for the expedited convergence of Jacobi-PD. This involves solving the reduced global system and smartly employing its precomputed factorization. In addition, we incorporate a time-splitting strategy to handle the frictional self-contacts. Xuan Li 0015, Yu Fang 0010, Lei Lan, Huamin Wang 0001, Yin Yang 0002, Minchen Li, Chenfanfu Jiang |
SIGGRAPH Asia | 3 |
| 2023 | Second-order Stencil Descent for Interior-point HyperelasticityabstractIn this paper, we present a GPU algorithm for finite element hyperelastic simulation. We show that the interior-point method, known to be effective for robust collision resolution, can be coupled with non-Newton procedures and be massively sped up on the GPU. Newton's method has been widely chosen for the interior-point family, which fully solves a linear system at each step. After that, the active set associated with collision/contact constraints is updated. Mimicking this routine using a non-Newton optimization (like gradient descent or ADMM) unfortunately does not deliver expected accelerations. This is because the barrier functions employed in an interior-point method need to be updated at every iteration to strictly confine the search to the feasible region. The associated cost (e.g., per-iteration CCD) quickly overweights the benefit brought by the GPU, and a new parallelism modality is needed. Our algorithm is inspired by the domain decomposition method and designed to move interior-point-related computations to local domains as much as possible. We minimize the size of each domain (i.e., a stencil) by restricting it to a single element, so as to fully exploit the capacity of modern GPUs. The stencil-level results are integrated into a global update using a novel hybrid sweep scheme. Our algorithm is locally second-order offering better convergence. It enables simulation acceleration of up to two orders over its CPU counterpart. We demonstrate the scalability, robustness, efficiency, and quality of our algorithm in a variety of simulation scenarios with complex and detailed collision geometries. Lei Lan, Minchen Li, Chenfanfu Jiang, Huamin Wang 0001, Yin Yang 0002 |
ACM Trans. Graph. | 1 |
| 2022 | A unified newton barrier method for multibody dynamicsabstractWe present a simulation framework for multibody dynamics via a universal variational integration. Our method naturally supports mixed rigid-deformables and mixed codimensional geometries, while providing guaranteed numerical convergence and accurate resolution of contact, friction, and a wide range of articulation constraints. We unify (1) the treatment of simulation degrees of freedom for rigid and soft bodies by formulating them both in terms of Lagrangian nodal displacements, (2) the handling of general linear equality joint constraints through an efficient change-of-variable strategy, (3) the enforcement of nonlinear articulation constraints based on novel distance potential energies, (4) the resolution of frictional contact between mixed dimensions and bodies with a variational Incremental Potential Contact formulation, and (5) the modeling of generalized restitution through semi-implicit Rayleigh damping. We conduct extensive unit tests and benchmark studies to demonstrate the efficacy of our method. Yunuo Chen 0001, Minchen Li, Lei Lan, Hao Su 0001, Yin Yang 0002, Chenfanfu Jiang |
ACM Trans. Graph. | 3 |
| 2022 | Affine body dynamics: fast, stable and intersection-free simulation of stiff materialsabstractSimulating stiff materials in applications where deformations are either not significant or else can safely be ignored is a fundamental task across fields. Rigid body modeling has thus long remained a critical tool and is, by far, the most popular simulation strategy currently employed for modeling stiff solids. At the same time, rigid body methods continue to pose a number of well known challenges and trade-offs including intersections, instabilities, inaccuracies, and/or slow performances that grow with contact-problem complexity. In this paper we revisit the stiff body problem and present ABD, a simple and highly effective affine body dynamics framework, which significantly improves state-of-the-art for simulating stiff-body dynamics. We trace the challenges in rigid-body methods to the necessity of linearizing piecewise-rigid trajectories and subsequent constraints. ABD instead relaxes the unnecessary (and unrealistic) constraint that each body's motion be exactly rigid with a stiff orthogonality potential, while preserving the rigid body model's key feature of a small coordinate representation. In doing so ABD replaces piecewise linearization with piecewise linear trajectories. This, in turn, combines the best of both worlds: compact coordinates ensure small, sparse system solves, while piecewise-linear trajectories enable efficient and accurate constraint (contact and joint) evaluations. Beginning with this simple foundation, ABD preserves all guarantees of the underlying IPC model we build it upon, e.g., solution convergence, guaranteed non-intersection, and accurate frictional contact. Over a wide range and scale of simulation problems we demonstrate that ABD brings orders of magnitude performance gains (two- to three-orders on the CPU and an order more when utilizing the GPU, obtaining 10, 000× speedups) over prior IPC-based methods, while maintaining simulation quality and nonintersection of trajectories. At the same time ABD has comparable or faster timings when compared to state-of-the-art rigid body libraries optimized for performance without guarantees, and successfully and efficiently solves challenging simulation problems where both classes of prior rigid body simulation methods fail altogether. Lei Lan, Danny M. Kaufman, Minchen Li, Chenfanfu Jiang, Yin Yang 0002 |
ACM Trans. Graph. | 1 |
| 2022 | Penetration-free projective dynamics on the GPUabstractWe present a GPU algorithm for deformable simulation. Our method offers good computational efficiency and penetration-free guarantee at the same time, which are not common with existing techniques. The main idea is an algorithmic integration of projective dynamics (PD) and incremental potential contact (IPC). PD is a position-based simulation framework, favored for its robust convergence and convenient implementation. We show that PD can be employed to handle the variational optimization with the interior point method e.g., IPC. While conceptually straightforward, this requires a dedicated rework over the collision resolution and the iteration modality to avoid incorrect collision projection with improved numerical convergence. IPC exploits a barrier-based formulation, which yields an infinitely large penalty when the constraint is on the verge of being violated. This mechanism guarantees intersection-free trajectories of deformable bodies during the simulation, as long as they are apart at the rest configuration. On the downside, IPC brings a large amount of nonlinearity to the system, making PD slower to converge. To mitigate this issue, we propose a novel GPU algorithm named A-Jacobi for faster linear solve at the global step of PD. A-Jacobi is based on Jacobi iteration, but it better harvests the computation capacity on modern GPUs by lumping several Jacobi steps into a single iteration. In addition, we also re-design the CCD root finding procedure by using a new minimum-gradient Newton algorithm. Those saved time budgets allow more iterations to accommodate stiff IPC barriers so that the result is both realistic and collision-free. Putting together, our algorithm simulates complicated models of both solids and shells on the GPU at an interactive rate or even in real time. Lei Lan, Guanqun Ma, Yin Yang 0002, Changxi Zheng, Minchen Li, Chenfanfu Jiang |
ACM Trans. Graph. | 1 |
| 2021 | Medial IPC: accelerated incremental potential contact with medial elasticsabstractWe propose a framework of efficient nonlinear deformable simulation with both fast continuous collision detection and robust collision resolution. We name this new framework Medial IPC as it integrates the merits from medial elastics, for an efficient and versatile reduced simulation, as well as incremental potential contact, for a robust collision and contact resolution. We leverage medial axis transform to construct a kinematic subspace. Instead of resorting to projective dynamics, we use classic hyperelastics to embrace real-world nonlinear materials. A novel reduced continuous collision detection algorithm is presented based on the medial mesh. Thanks to unique geometric properties of medial axis and medial primitives, we derive closed-form formulations for identifying between-primitive collision within the reduced medial space. In the meantime, the implicit barrier energy that generates necessary repulsion forces for collision resolution is also formulated with the medial coordinate. In other words, Medial IPC exploits a universal reduced coordinate for simulation, continuous self-/collision detection, and IPC-based collision resolution. Continuous collision detection also allows more aggressive time stepping. In addition, we carefully implement our system with a heterogeneous CPU-GPU deployment such that massively parallelizable computations are carried out on the GPU while few sequential computations are on the CPU. Such implementation also frees us from generating training poses for selecting Cubature points and pre-computing their weights. We have tested our method on complicated deformable models and collision-rich simulation scenarios. Due to the reduced nature of our system, the computation is faster than fullspace IPC or other fullspace methods using continuous collision detection by at least one order. The simulation remains high-quality as the medial subspace captures intriguing and local deformations with sufficient realism. Lei Lan, Yin Yang 0002, Danny M. Kaufman, Junfeng Yao, Minchen Li, Chenfanfu Jiang |
ACM Trans. Graph. | 1 |
| 2021 | High-order differentiable autoencoder for nonlinear model reductionabstractThis paper provides a new avenue for exploiting deep neural networks to improve physics-based simulation. Specifically, we integrate the classic Lagrangian mechanics with a deep autoencoder to accelerate elastic simulation of deformable solids. Due to the inertia effect, the dynamic equilibrium cannot be established without evaluating the second-order derivatives of the deep autoencoder network. This is beyond the capability of off-the-shelf automatic differentiation packages and algorithms, which mainly focus on the gradient evaluation. Solving the nonlinear force equilibrium is even more challenging if the standard Newton's method is to be used. This is because we need to compute a third-order derivative of the network to obtain the variational Hessian. We attack those difficulties by exploiting complex-step finite difference, coupled with reverse automatic differentiation. This strategy allows us to enjoy the convenience and accuracy of complex-step finite difference and in the meantime, to deploy complex-value perturbations as collectively as possible to save excessive network passes. With a GPU-based implementation, we are able to wield deep autoencoders (e.g., 10+ layers) with a relatively high-dimension latent space in real-time. Along this pipeline, we also design a sampling network and a weighting network to enable weight-varying Cubature integration in order to incorporate nonlinearity in the model reduction. We believe this work will inspire and benefit future research efforts in nonlinearly reduced physical simulation problems. Yin Yang 0002, Tianjia Shao, He Wang 0002, Chenfanfu Jiang, Lei Lan, Kun Zhou 0001 |
ACM Trans. Graph. | 6 |
| 2020 | Medial Elastics: Efficient and Collision-Ready Deformation via Medial Axis TransformabstractWe propose a framework for the interactive simulation of nonlinear deformable objects. The primary feature of our system is the seamless integration of deformable simulation and collision culling, which are often independently handled in existing animation systems. The bridge connecting them is the medial axis transform (MAT), a high-fidelity volumetric approximation of complex 3D shapes. From the physics simulation perspective, MAT leads to an expressive and compact reduced nonlinear model. We employ a semireduced projective dynamics formulation, which well captures high-frequency local deformations of high-resolution models while retaining a low computation cost. Our key observation is that the most compelling (nonlinear) deformable effects are enabled by the local constraints projection, which should not be aggressively reduced, and only apply model reduction at the global stage. From the collision detection (CD)/collision culling (CC) perspective, MAT is geometrically versatile using linear-interpolated spheres (i.e., the so-called medial primitives (MPs)) to approximate the boundary of the input model. The intersection test between two MPs is formulated as a quadratically constrained quadratic program problem. We give an algorithm to solve this problem exactly, which returns the deepest penetration between a pair of intersecting MPs. When coupled with spatial hashing, collision (including self-collision) can be efficiently identified on the GPU within a few milliseconds even for massive simulations. We have tested our system on a variety of geometrically complex and high-resolution deformable objects, and our system produces convincing animations with all of the collisions/self-collisions well handled at an interactive rate. Lei Lan, Ran Luo 0001, Marco Fratarcangeli, Weiwei Xu 0003, Huamin Wang 0001, Xiaohu Guo, Junfeng Yao, Yin Yang 0002 |
ACM Trans. Graph. | 1 |
| 2017 | Medial-axis-driven shape deformation with volume preservation
Lei Lan, Junfeng Yao, Xiaohu Guo |
Vis. Comput. | 1 |
| 2015 | Realistic and stable animation of clothabstractAbstract Researchers have proposed various techniques for cloth simulation in the last decades. The crucial problem in interactive animation for cloth is how to speed up simulation with a stable system. In this paper, we describe a realistic and stable scheme for cloth simulation based on mass‐spring system. This scheme modifies semi‐implicit integration used in cloth simulation system with an efficient damping method. Semi‐implicit integration methods have been widely used in dynamic simulation because of acceptable speed and stability. However, internal damping forces are generated with respect to rotational rigid motions, which are only depending on relative velocities. Undesirable change in global movement of dynamic cloth could result in damping artifact. We replace the internal damping forces with an optimal damping method which is based on iterative spring damping but the limit can be computed directly. The method provides simple damping parameter estimation and guarantees conservation of global movement. As a result, complex clothes can be realistically and stably simulated in real time. Copyright © 2015 John Wiley & Sons, Ltd. Junfeng Yao, Lei Lan, Wenlin Lin, Yingying She |
Comput. Animat. Virtual Worlds | 2 |
| 2006 | LOMARC: Lookahead Matchmaking for Multiresource Coscheduling on Hyperthreaded CPUsabstractJob scheduling typically focuses on the CPU with little work existing to include I/O or memory. Time-shared execution provides the chance to hide I/O and long-communication latencies though potentially creating a memory conflict. Hyperthreaded CPUs support coscheduling without any context switches and provide additional options for CPU-internal resource sharing. We present an approach that includes all possible resources into the schedule optimization and improves utilization by coscheduling two jobs if feasible. Our LOMARC approach partially reorders the queue by lookahead to increase the potential to find good matches. In simulations based on the workload model of Lublin and Feitelson, we have obtained improvements between 30 percent and 50 percent in both response times and relative bounded response times on hyperthreaded CPUs (i.e., cut times to two third or to half) Angela C. Sodan, Lei Lan |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2004 | LOMARC - Lookahead Matchmaking for Multi-resource CoschedulingabstractJob scheduling typically focuses on the CPU with little work existing to include I/O or memory. Time-shared execution provides the chance to hide I/O and long-communication latencies though potentially creating a memory conflict. We consider two different cases: standard local CPU scheduling and coscheduling on hyperthreaded CPUs. The latter supports coscheduling without any context switches and provides additional options for CPU-internal resource sharing. We present an approach that includes all possible resources into the schedule optimization and improves utilization by coscheduling two jobs if feasible. Our LOMARC approach partially reorders the queue by lookahead to increase the potential to find good matches. In simulations based on the workload model of [12], we have obtained improvements of about 50% in both response times and relative bounded response times on hyperthreaded CPUs (i.e. cut times by half) and of about 25% on standard CPUs for our LOMARC scheduling approach. These keywords were added by machine and not by the authors. This process is experimental and the keywords may be updated as the learning algorithm improves. Angela C. Sodan, Lei Lan |
JSSPP | 2 |