VLDB 2026 Research / reviewers in the wild / expert
Liang Deng
dblp:75/300
· DBLP profile ↗
38ranked-venue papers
11as first author
19since 2021 · last 2027
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 23 · 6 first-author · 11 since 2021Artificial intelligence and machine learning · 7 · 1 first-author · 5 since 2021Security and privacy · 3 · 3 first-authorSoftware engineering, systems software and programming languages · 2 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 2 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2027 | Retrofitting temporal GNN training with decoder-only Transformers on large-scale graphs
Qiang Huang 0009, Ke Liu 0014, Liang Deng, Xiao Yan 0002, Chuang Hu, Quanqing Xu, Wentao Zhang 0001, Jiawei Jiang 0001 |
Expert Syst. Appl. | 3 |
| 2026 | DM-RAG: Enhancing User Support in Dameng Databases with Retrieval-Augmented Generation
Qiang Huang 0009, Ke Liu 0014, Liang Deng, Sijing Zhang, Chuang Hu, Tieyun Qian, Xiao Yan 0002, Jiawei Jiang 0001 |
ICDE | 3 |
| 2026 | Dual-channel machine learning proxy for pseudo-two-dimensional model with enhanced extrapolation correction
Yaxuan Wang, Shilong Guo, Liang Deng, Junfu Li |
Eng. Appl. Artif. Intell. | 4 |
| 2026 | GNNRL-smoothing: A prior-free reinforcement learning model for mesh optimization
Xinhai Chen 0001, Chunye Gong, Bo Yang 0023, Liang Deng, Yufei Pang, Xiang Zhang 0008, Jie Liu 0002 |
Neural Networks | 5 |
| 2026 | Towards Efficient Symmetric Sparse Matrix-Vector Multiplication on Multi-CoresabstractExploiting matrix symmetry to halve memory footprint offers a substantial opportunity for accelerating memory-bound computations like Sparse Matrix-Vector Multiplication (SpMV). However, symmetric SpMV incurs data conflicts when concurrently writing the output vector. Previous approaches fail to address this issue efficiently, i.e., either are non-scalable or yield poor performance for large high-bandwidth irregular matrices. This article extends DCS-SpMV , a D ivide-and- C onquer (DC) based shared-memory implementation of S ymmetric SpMV. The key idea of DCS-SpMV is to recursively divide and reorder the matrix-induced conflict graph into independent subgraphs for parallel execution, and construct separate subgraphs to avoid data conflicts. The DC approach naturally transforms the input matrix into a low-conflict part and a high-conflict part, which motivates us to design a conflict-aware hybrid solution DCH-SpMV that executes these two parts using DCS-SpMV and the standard SpMV, respectively. We also develop a machine learning model for DCH-SpMV to predict the optimal number of DC recursions on a given matrix and architecture. In this work, we further optimize the hybrid DC implementation by reducing data conflicts before the DC preprocessing. First, we present a conflict-pruning strategy to decouple certain highly dense columns or rows from the conflict graph of a symmetric matrix. Second, we implement a heuristic to adaptively select the lower or upper triangular part of a symmetric matrix, leading to fewer data conflicts. Our optimizations not only facilitate the DC preprocessing, but also improve the performance of DCH-SpMV. We evaluate our work on both x86 and ARM multi-core CPUs using 298 symmetric sparse matrices from the SuiteSparse Matrix Collection. Our new optimizations improve the performance of previous version [ 42 ] by up to 4.89×, demonstrating significant speedup over the state-of-the-art approaches including the vendor-tuned Intel oneMKL library. Haozhong Qiu, Chuanfu Xu, Jianbin Fang, Jian Zhang 0115, Liang Deng, Yue Ding 0001, Zhimeng Han, Yonggang Che |
ACM Trans. Archit. Code Optim. | 6 |
| 2025 | Me-MPK: Accelerating Krylov Subspace Solvers via Memory-efficient Matrix-Power KernelabstractThis paper focuses on optimizing the Matrix-Power Kernel (MPK), which relies on a series of Sparse Matrix-Vector multiplications (SpMVs) using the same sparse matrix. MPK is a crucial component of Krylov subspace methods for solving large sparse linear systems in various fields, including circuit simulations. MPK offers a potential for matrix reuse in cache, which can accelerate memory-bound sparse solvers. Additionally, many sparse matrices encountered in applications are symmetric, allowing us to reduce the memory footprint for SpMVs by half. However, reusing the matrix introduces data dependencies between subsequent SpMVs, and symmetric SpMVs can result in data conflicts during shared-memory parallelization. Previous research has often focused on either matrix reuse or symmetry, failing to leverage both aspects effectively. This paper proposes a unified, memory-efficient approach called Me-MPK that takes advantage of both cache reuse and matrix symmetry for MPK on shared-memory multi-core systems. We first introduce a unified dependency graph for a sparse matrix, which represents all potential data dependencies and conflicts. Next, we perform architecture-aware recursive partitioning on this graph to create subgraphs and formulate a separating subgraph that decouples all dependencies and conflicts among the subgraphs. These independent subgraphs are then scheduled for parallel execution of SpMV or symmetric SpMV in a specified order to optimize cache reuse. We apply Me-MPK in two s-Step Krylov subspace solvers, and our evaluations show that Me-MPK significantly outperforms the current state-of-the-art solutions, delivering an average speedup of up to 2.00X and 1.86X on X86 and ARM CPUs, respectively. As a result, we achieve overall speedup in the sparse solvers of up to $\mathbf{1. 6 5 X}$ and $\mathbf{1. 5 8 X}$. Haozhong Qiu, Chuanfu Xu, Jianbin Fang, Shengguo Li, Liang Deng, Jian Zhang 0115, Yue Ding 0001, Zhimeng Han, Yonggang Che, Jie Liu 0002 |
DAC | 5 |
| 2025 | PEINR: A Physics-enhanced Implicit Neural Representation for High-Fidelity Flow Field ReconstructionabstractImplicit neural representation (INR) has now been thrust into the limelight with its flexibility in high-fidelity flow field reconstruction tasks. However, the lack of standard benchmarking datasets and the grid independence assumption for INR-based methods hinder progress and adoption in real-world simulation scenarios. Moreover, naive adoptions of existing INR frameworks suffer from limited accuracy in capturing fine-scale structures and spatiotemporal dynamics. Tacking these issues, we first introduce HFR-Beach, a 5.4 TB public large-scale CFD dataset with 33,600 unsteady 2D and 3D vector fields for reconstructing high-fidelity flow fields. We further present PEINR, a physics-enhanced INR framework, to enrich the flow fields by concurrently enhancing numerical-precision and grid-resolution. Specifically, PEINR is mainly composed of physical encoding and transformer-based spatiotemporal fuser (TransSTF). Physical encoding decouples temporal and spatial components, employing Gaussian coordinate encoding and localized encoding techniques to capture the nonlinear characteristics of spatiotemporal dynamics and the stencil discretization of spatial dimensions, respectively. TransSTF fuses both spatial and temporal information via transformer for capturing long-range temporal dependencies. Qualitative and quantitative experiments and demonstrate that PEINR outperforms state-of-the-art INR-based methods in reconstruction quality. Liming Shen, Liang Deng, Chongke Bi, Xinhai Chen 0001, Yueqing Wang, Jie Liu 0002 |
ICML | 2 |
| 2025 | BiFTVis: A Bidirectional Feature-Tracking Method for Visual Analytics of Flow FieldsabstractAccurate analysis of time-dependent flow fields generated by numerical simulations requires effective interpretation of temporal information. Visualizations offer an exceptional ability to convey complex data and have been widely used for such analysis. However, current research mainly focuses on flow features at individual time steps, lacking exploration of how these features evolve over time, leading to insufficient investigation and understanding of fluid motion laws. To overcome this limitation, this study proposes BiFTVis, an interactive visualization framework that enhances the complete flow field analysis pipeline. The framework includes a volume rendering module that represents flow features extracted using the Q criterion method. An additional forward tracking module is provided to track feature events by utilizing a graph optimization-based feature tracking algorithm. A timeline-based visual encoding method has been designed to convey multiple feature events along the time axis, facilitating the examination of events in specific time periods. Furthermore, a novel radial layout glyph along with a Directed Acyclic Graph has been designed for encoding multi-facet feature attributes, enabling fast identification of various feature types and reverse tracking analysis of feature events in a backward tracking module. BiFTVis has been evaluated through a case study, demonstrating its effectiveness in promoting new insights into potential causes of feature events and avoiding certain feature event occurrences, ultimately enhancing flow field analysis. Chongke Bi, Peipu Pan, Liang Deng |
SMC | 5 |
| 2025 | DCSolver: Accelerating Sparse Iterative Solvers via Divide-and-Conquer on GPUsabstractSparse iterative solvers are commonly used in various fields. However, certain essential kernels of these solvers, such as sparse triangular solves (SpTRSV), present significant challenges for efficient parallelization due to data dependencies . Previous methods, like level-scheduling or multi-coloring, typically involve creating a Task Dependency Graph (TDG) to represent data dependencies and identify independent sets from the TDG for parallel execution. However, these approaches often result in limited parallelism with substantial synchronization overheads or negatively impact the solver convergence rate. This article introduces DCSolver , a Divide-and-Conquer (DC) framework designed to efficiently parallelize sparse solvers with data dependencies on GPUs. To achieve this, we break down the solver TDG into independent subgraphs, allowing us to exploit both coarse-grained and fine-grained parallelism. To efficiently allocate GPU threads for subgraphs with varying degrees of parallelism, we have developed an adaptive in-warp scheduling strategy. Additionally, we propose a hybrid parallelization scheme in DCSolver, which involves employing different parallel approaches for different DC recursions to achieve a more optimal balance between parallelism and convergence for solvers. To evaluate the effectiveness of DCSolver, we apply it to two preconditioned Krylov subspace solvers and an unstructured mesh Computational Fluid Dynamics (CFD) solver. Our results show that when compared with the state-of-the-art methods, DCSolver accelerates the time-to-solution of solvers by an average speedup of up to 26.19X. Haozhong Qiu, Chuanfu Xu, Jianbin Fang, Jian Zhang 0115, Liang Deng, Yue Ding 0001, Zhimeng Han, Yonggang Che, Jie Liu 0002 |
ACM Trans. Archit. Code Optim. | 5 |
| 2024 | DaCP: Accelerating Synchronization-Free SpTRSV via GPU-Friendly Data Communication and Parallelism Strategies
Mingfeng Guo, Liang Deng, Ruitian Li, Gaofeng Lin, Jie Liu 0002 |
NPC (1) | 2 |
| 2024 | Towards Scalable Unstructured Mesh Computations on Shared Memory Many-CoresabstractDue to data conflicts or data dependences, exploiting shared memory parallelism on unstructured mesh applications is highly challenging. The prior approaches are neither general nor scalable on emerging many-core processors. This paper presents a general and scalable shared memory approach for unstructured mesh computations. We recursively divide and reorder an unstructured mesh to construct a task dependency tree (TDT), where massive parallelism is exposed and data conflicts as well as data dependences are respected. We propose two recursion strategies to support popular programming models on both CPUs and GPUs for TDT. We evaluate our approach by applying it to an industrial unstructured Computational Fluid Dynamics (CFD) software. Experimental results show that our approach significantly outperforms the prior shared memory approaches, delivering up to 8.1× performance improvement over the engineer-tuned implementations. Haozhong Qiu, Chuanfu Xu, Jianbin Fang, Liang Deng, Jian Zhang 0115, Yue Ding 0001, Yonggang Che, Shizhao Chen, Jie Liu 0002 |
PPoPP | 4 |
| 2024 | A Conflict-aware Divide-and-Conquer Algorithm for Symmetric Sparse Matrix-Vector MultiplicationabstractExploiting matrix symmetry to halve memory footprint offers an opportunity for accelerating memory-bound computations like Sparse Matrix-Vector Multiplication (SpMV). However, symmetric SpMV incurs data conflicts when concurrently writing the output vector. Previous approaches fail to address this issue efficiently. This paper proposes DCS-SpMV, a Divide-and-Conquer (DC) algorithm for efficient Symmetric SpMV. The key idea is to recursively divide the matrix-induced conflict graph into independent subgraphs for parallel execution, and construct separate subgraphs to avoid data conflicts. Our DC algorithm transforms the input matrix into a low-conflict part and a high-conflict part, which motivates us to design a conflict-aware hybrid solution that executes these two parts using DCS-SpMV and traditional SpMV respectively. We develop a machine learning model to predict an optimal hybrid implementation for a given matrix and architecture. We evaluate our work on both X86 and ARM CPUs, demonstrating significant performance improvement over the state-of-the-art. Haozhong Qiu, Chuanfu Xu, Jianbin Fang, Jian Zhang 0115, Liang Deng, Yue Ding 0001, Shizhao Chen, Yonggang Che, Jie Liu 0002 |
SC | 5 |
| 2024 | Evaluating performance portability of five shared-memory programming models using a high-order unstructured CFD solver
Liang Deng, Yonggang Che, Yueqing Wang |
J. Parallel Distributed Comput. | 2 |
| 2024 | Massively parallel simulations of multi-stage compressors on Sunway TaihuLight
Liang Deng, Fengshun Lu, Zhaolin Fan, Xiong Jiang |
J. Supercomput. | 3 |
| 2023 | Developing a proxy application for an industrial unstructured CFD software: preliminary resultsabstractAs programming models and architectures evolve in the exa-scale era, porting large-scale HPC applications are becoming increasingly difficult and expensive. In HPC community, mini-apps are often developed to mimic real-world applications, and it offers an easy way to benchmark new HPC platforms. In this paper, we design and implement a mini-app MiniFS as a proxy for an industry-level unstructured Computational Fluid Dynamics (CFD) software FlowStar. The main purpose of MiniFS is to evaluate different shared memory approaches on emerging multi/many-core architectures, because data conflicts and data dependencies in unstructured CFD pose tough challenges for shared memory parallelization. Results show that our mini-app can represent the performance characteristic of the original application. However, the existing approaches are unscalable on modern multi/many-cores. It is imperative to develop novel scalable shared memory approaches for unstructured CFD. Chuanfu Xu, Jian Zhang 0115, Liang Deng, Haozhong Qiu, Weixi Dai, Yongzhen Lin, Yue Ding 0001, Yonggang Che |
ICPADS | 4 |
| 2023 | Achieving high performance and portable parallel GMRES algorithm for compressible flow simulations on unstructured grids
Jian Zhang 0115, Liang Deng, Ruitian Li, Jie Liu 0002 |
J. Supercomput. | 2 |
| 2023 | An Intelligent Method for Predicting the Pressure Coefficient Curve of Airfoil-Based Conditional Generative Adversarial NetworksabstractRecently, extensive studies have focused on analyzing aerodynamic performance due to its important impact on aircraft design. Most of these works compute the aerodynamic coefficient of the airfoil through computational fluid dynamics (CFD) simulation, which is too time-consuming. To reduce the computational time required, some intelligence-based methods have been presented. However, these methods also suffer from certain issues. First, most of them directly implement existing machine learning methods used to predict the aerodynamic coefficient without adding any improvements. Second, some methods convert the airfoil shape and aerodynamic curves into images, which may lead to curve distortion and the introduction of noise. Third, some methods learn the relationship between the airfoil shape and aerodynamic coefficients but ignore the influence of initial inflow conditions. Accordingly, to address these issues, we propose an intelligent method for predicting the pressure coefficients (Cp) of airfoil based on a conditional generative adversarial network (cGAN). More specifically, we first present a two-step data augmentation strategy designed to expand the original airfoil dataset. Subsequently, we design a novel cGAN-based neural network to predict the Cp curve. To the best of our knowledge, this is the first work to apply generative adversarial network (GAN) to aerodynamic coefficient prediction. Moreover, we design a new loss function to train our network. Extensive experimental results demonstrate that the Cp curve predicted by our method is very close to that generated via CFD simulation. More importantly, our method achieves a speedup close to 1000x compared with CFD simulation. Yueqing Wang, Liang Deng, Yunbo Wan, Zhigong Yang, Wenxiang Yang, Cheng Chen 0005, Dan Zhao 0002, Fang Wang 0004 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2022 | A novel neural network approach for airfoil mesh quality evaluation
Xinhai Chen 0001, Chunye Gong, Jie Liu 0002, Yufei Pang, Liang Deng, Lihua Chi, Kenli Li 0001 |
J. Parallel Distributed Comput. | 5 |
| 2021 | A rapid vortex identification method using fully convolutional segmentation network
Yueqing Wang, Liang Deng, Zhigong Yang, Dan Zhao 0002, Fang Wang 0004 |
Vis. Comput. | 2 |
| 2017 | Dancing with Wolves: Towards Practical Event-driven VMM MonitoringabstractThis paper presents a novel framework that enables practical event-driven monitoring for untrusted virtual machine monitors (VMMs) in cloud computing. Unlike previous approaches for VMM monitoring, our framework neither relies on a higher privilege level nor requires any special hardware support. Instead, we place the trusted monitor at the same privilege level and in the same address space with the untrusted VMM to achieve superior efficiency, while proposing a unique mutual-protection mechanism to ensure the integrity of the monitor. Our security analysis demonstrates that our framework can provide high-assurance for event-driven VMM monitoring, even if the highest-privilege VMM is fully compromised. The experimental results show that our framework only incurs trivial performance overhead for enforcing event-driven monitoring policies, exhibiting tremendous performance improvement on previous approaches. Liang Deng, Peng Liu 0005, Jun Xu 0024, Ping Chen 0003, Qingkai Zeng 0002 |
VEE | 1 |
| 2016 | Modified Elman neural network based neural adaptive inverse control of rate-dependent hysteresisabstractA modified Elman neural network (MENN) based neural adaptive inverse control scheme is proposed for trajectory tracking of rate-dependent hysteresis. To attenuate the influence of the rate-dependent hysteresis, a modified inverse backlash operator (MIBO) is developed to act as the hidden layer neuron of the MENN to describe the dynamic behavior of the inverse rate-dependent hysteresis. The diagonal recurrent weights of context layer and the recurrent weights from hidden layer to context layer are designed to enhance the dynamic learning capability of the MENN. To determine an appropriate structure and the parameters of the MENN for adaptive inverse control, a MENN is trained first based on the data of inverse rate-dependent hysteresis. In view of the nonsmooth characteristics of the MIBOs, a restricted step proximal bundle (RSPB) method is employed to search the appropriate subgradients at the nonsmooth vertexes of the MIBOs. The relevant Levenberg-Marquardt (L-M) algorithm is developed to acquire an appropriate MENN that is used as the initial controller to implement adaptive inverse control for rate-dependent hysteresis via gradient descent learning algorithm. Numerical control results on a Duhem model of piezoelectric actuators have validated the effectiveness of the proposed method. Liang Deng, Rudolf J. Seethaler, YangQuan Chen, Qiming Cheng |
IJCNN | 1 |
| 2016 | Evaluating Multi-core and Many-Core Architectures through Accelerating an Alternating Direction Implicit CFD SolverabstractIn this paper, we accelerate a double-precision alternating direction implicit (ADI) solver for three-dimensional compressible Navier-Stokes equations from our in-house computational fluid dynamics (CFD) software on the latest multi-core and many-core architectures (Intel Ivy Bridge CPU, Intel Xeon Phi 7110P coprocessor and NVIDIA Kepler K20c GPU). For the GPU platform, both the OpenACC-based and the CUDA-based versions of the ADI solver are developed. To achieve high performance, we use a series of optimization techniques. For the Ivy Bridge CPU and Xeon Phi, we focus on three categories of optimization techniques: thread parallelism for multi-/many-core scaling, data parallelism to exploit the SIMD mechanism and improving on-chip data reuse, to maximize the performance. Also, we provide an in-depth analysis on the performance differences between Ivy Bridge and Xeon Phi. Our numerical experiments show that the proposed CUDA-based ADI solver can achieve a speedup of 9.7 on a Kepler GPU in contrast to a single naive serial version and our optimization techniques can improve the performance of the ADI solver by 2.5x on two Ivy Bridge CPUs and 1.7x on the Intel Xeon Phi coprocessor. We also notice that the OpenACC-based version runs around 29% slower than the CUDA-based one with careful manual optimizations. Besides, we systematically evaluate the programmability of the three platforms. Our insights facilitate the programmers to select a right platform with a suitable programming model according to their target applications. Liang Deng, Jianbin Fang, Fang Wang 0004, Hanli Bai |
ISPDC | 1 |
| 2016 | An efficient Uniform Integrated Advection algorithm for Finite Time Lyapunov Exponent field computation on GPU and MICabstractFinite Time Lyapunov Exponent (FTLE) is widely used in Lagrangian Coherent Structure extraction and research on unsteady flow field. FTLE is computationally expensive due to the flow map calculation on dense samples in the field. In order to improve the computation efficiency, we investigate the most time-consuming flow map and propose UIA, a Uniform Integrated Advection algorithm based on piecewise linear hypothesis. FTLE coverts multiple spatiotemporal interpolations and multi-step time advection into single matrix multiplication so as to dramatically reduce computation and be applied to any type of grids. We also design an UIA based FTLE algorithm on GUP and MIC to further reduce computation time, and present corresponding performance optimization strategies. The correctness and accuracy of algorithm are also experimentally verified by applying the algorithm to an analytical 3-dimensional field. The speedup ratio of the proposed algorithm ranges from 6 to 143 compared with FTLE execution time based on traditional RK4 flow map, which demonstrates that UIA has significantly improve the efficiency of FTLE computation and visualization. Fang Wang 0004, Liang Deng, Dan Zhao 0002, Sikun Li |
SNPD | 2 |
| 2016 | Exception-oriented programming: retrofitting code-reuse attacks to construct kernel malwareabstractCommodity operating system kernels are vulnerable to a wide range of attacks due to the large code base and broad attack surface. Mitigation mechanisms such as code signing, W⊕X, and code integrity protection have raised the bar for kernel security. In turn, attack mechanisms have also become increasingly advanced. They have evolved from simple injection of malicious code into more sophisticated code‐reuse attacks [e.g. return‐oriented programming (ROP)]. In this study, the authors describe exception‐oriented programming (EOP), a novel code‐reuse method to construct kernel malware. Unlike previous ROP that can only reuse a limited part of existing code (gadgets), EOP is able to reuse any instruction in existing code and chain the instructions in any order to generate malicious programmes. As a result, EOP can provide the attackers with more powerful capabilities and less complexity for building kernel malware. Liang Deng, Qingkai Zeng 0002 |
IET Inf. Secur. | 1 |
| 2015 | ISboxing: An Instruction Substitution Based Data Sandboxing for x86 Untrusted Libraries
Liang Deng, Qingkai Zeng 0002, Yao Liu 0011 |
SEC | 1 |
| 2014 | EqualVisor: Providing Memory Protection in an Untrusted Commodity HypervisorabstractIn cloud computing, hypervisor is the all-powerful software running in the highest privilege layer, thus attackers who compromise a hypervisor may jeopardize the whole cloud, especially cause memory corruption of any sensitive workloads within the cloud. In this paper, we propose a novel architecture and approach to provide memory protection from an untrusted hypervisor on current x86 platforms. Unlike previous approaches such as nested virtualization, we do not place another higher privilege TCB below the hypervisor. Instead, our approach introduces a properly isolated tiny TCB running in the same privilege level and the same address space with the hypervisor, and uses this TCB to intercept and validate hypervisor's privilege actions for memory protection. In this way, we can enforce further memory security policies only relying on the TCB even if the hypervisor is fully compromised. Liang Deng, Qingkai Zeng 0002, Yao Liu 0011 |
TrustCom | 1 |
| 2013 | Integrating Orientation Cue With EOH-OLBP-Based Multilevel Features for Human DetectionabstractDetecting pedestrians efficiently and accurately is a fundamental step for many computer vision applications, such as smart cars and robotics. In this paper, we introduce a pedestrian detection system to extract human objectives using an on-board monocular camera. First of all, we use an experiment to demonstrate that the orientation information is critical in human detection. Secondly, the local binary patterns-based feature, oriented LBP (OLBP), is discussed. The OLBP feature integrates pixel intensity difference with texture orientation information to capture salient object features. Thirdly, a set of edge orientation histogram (EOH) and OLBP-based intrablock and interblock features is presented to describe cell-level and block-level structure information. These multilevel features capture larger-scale structure information which is more informative for pedestrian localization. Experiments on the Institut national de recherche en informatique et en automatique (INRIA) dataset and the Caltech pedestrian detection benchmark demonstrate that the new pedestrian detection system is not only comparable to the existing pedestrian detectors, but also performs at a faster speed. Liang Deng, Xiankai Chen |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2009 | Wire shaping is practicalabstractWire shaping for delay/power minimization has been extensively studied. Due to the perceived high design and manufacturing costs for using non-uniform wire shapes, wire shaping is generally considered to be impractical. In this paper, we present a practical wire shaping methodology. Non-uniform wire shapes are directly implemented on silicon wafer instead of in GDSII during design. We present novel enhancements to existing OPC technology to accurately print non-uniform wire shapes. Experimental results show that the post-OPC mask complexities of uniform wire and non-uniform wire are comparable. With minimal impact on the design and manufacturing flows and minimal additional design and manufacturing costs, we demonstrate that wire shaping can help to obtain substantial reduction of interconnect dynamic power without affecting timing closure. Our wire shaping methodology is an excellent example of Manufacturing for Design. Hongbo Zhang 0001, Martin D. F. Wong, Kai-Yuan Chao, Liang Deng |
ISPD | 4 |
| 2008 | Fast Dummy-Fill Density Analysis With Coupling ConstraintsabstractIn modern very large scale integration manufacturing processes, dummy fills are widely used to adjust local metal density in order to improve layout uniformity and yield optimization. However, the introduction of a large amount of dummy features also affects wire electrical properties. In this paper, we propose the first coupling-constrained dummy-fill analysis algorithm which identifies feasible locations for dummy fills such that the fill-induced coupling capacitance can be bounded within the given coupling threshold of each wire segment. A speedup approach is presented based on the cache concept. The algorithm also makes efforts to maximize ground dummy fills, which are more robust and predictable. The output of the algorithm can be treated as the upper bound for dummy-fill insertion, and it can be easily adopted in density models to guide dummy-fill insertion without disturbing the existing design. Hua Xiang 0001, Liang Deng, Ruchir Puri, Kai-Yuan Chao, Martin D. F. Wong |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2007 | Coupling-aware Dummy Metal Insertion for LithographyabstractAs integrated circuits manufacturing technology is advancing into 65nm and 45nm nodes, extensive resolution enhancement techniques (RETs) are needed to correctly manufacture a chip design. The widely used RET called off-axis illumination (OAI) introduces forbidden pitches which lead to very complex design rules. It has been observed that imposing uniformity on layout designs can substantially improve printability under OAI. For metal layers, uniformity can be achieved simply by inserting dummy metal wire segments at all free spaces. Simulation results indeed show significant improvement in printability with such a dummy metal insertion approach. To minimize mask cost, it is advantageous to use dummy metal segments that are of the same size as regular metal wires due to their simple geometry. But these dummy wires are printable and hence increase coupling capacitances and potentially affect yield. The alternative is to use a set of parallel sub-resolution thin wires (which is not printed) to replace a printable dummy wire segment. These invisible dummy metal segments do not increase coupling capacitances but bring a higher lithography cost, which includes mask cost and RET/process expense. This paper presents a strategy for dummy metal insertion that can optimally trade off lithography cost and coupling capacitance. In particular, we present an optimal algorithm that can minimize lithography cost subject to any given coupling capacitance bound. Moreover, this dummy metal insertion achieves a highly uniform density because of the locality of coupling capacitance, which automatically ameliorates chemical mechanical polish (CMP) problem. Liang Deng, Martin D. F. Wong, Kai-Yuan Chao, Hua Xiang 0001 |
ASP-DAC | 1 |
| 2007 | Fast and Accurate OPC for Standard-Cell LayoutsabstractModel based optical proximity correction (OPC) has become necessary at 90nm technology node and beyond. Cell-wise OPC is an attractive technique to reduce the mask data size as well as the prohibitive runtime of full-chip OPC. As feature dimensions have gotten smaller, the radius of influence for edge features has extended further into neighboring cells such that it is no longer sufficient to perform cellwise OPC independent of neighboring cells, especially for the critical layers. The methodology described in this work accounts for features in neighboring cells and allows a cellwise approach to be applied to cells with a printed gate length of 45nm with the projection that it can also be applied to future technology nodes. OPC-ready cells are generated at library creation (independent of placement) using a boundary-based technique. Each cell has a tractable number of OPC-ready versions due to an intelligent characterization of standard cell layout features. Total number of cells with boundaries in the OPC-ready library only increases linearly with the number of cells in the original library. Results are very promising: the average edge placement error (EPE) for all metal1 features in 100 layouts is 0.731nm which is less than 1 % of metal1 width, creating similar levels of lithographic accuracy while obviating any of the drawbacks inherent in layout specific full-chip model-based OPC. For even small circuits, there were runtime reductions of up to 100times and a potential 35times decrease in mask data size. David M. Pawlowski, Liang Deng, Martin D. F. Wong |
ASP-DAC | 2 |
| 2007 | Dummy fill density analysis with coupling constraintsabstractIn modern VLSI manufacturing processes, dummy fills are widely used to adjust local metal density in order to improve layout uniformity and yield optimization. However, the introduction of a large amount of dummy features also affects wire electrical properties. In this paper, we propose the first Coupling constrained Dummy Fill (CDF) analysis algorithm which identifies feasible locations for dummy fills such that the fill induced coupling capacitance can be bounded within the given coupling threshold of each wire segment. The algorithm also makes efforts to maximize ground dummy fills, which are more robust and predictable. The output of the algorithm can be treated as the upper bound for dummy fill insertion, and it can be easily adopted in density models to guide dummy fill insertion without disturbing the existing design. Hua Xiang 0001, Liang Deng, Ruchir Puri, Kai-Yuan Chao, Martin D. F. Wong |
ISPD | 2 |
| 2006 | An exact algorithm for the statistical shortest path problemabstractGraph algorithms are widely used in VLSI CAD. Traditional graph algorithms can handle graphs with deterministic edge weights. As VLSI technology continues to scale into nanometer designs, we need to use probability distributions for edge weights in order to model uncertainty due to parameter variations. In this paper, we consider the statistical shortest path (SSP) problem. Given a graph G, the edge weights of G are random variables. For each path P in G, let L/sub P/ be its length, which is the sum of all edge weights on P. Clearly L/sub P/ is a random variable and we let /spl mu//sub P/, and /spl omega//sub P//sup 3/ be its mean and variance, respectively. In the SSP problem, our goal is to find a path P connecting two given vertices to minimize the cost function /spl mu//sub p/, + /spl Phi/ (/spl omega//sub P//sup 2/) where /spl Phi/ is an arbitrary function. (For example, if /spl Phi/ (/spl times/) /spl equiv/ the cost function is /spl mu//sub P/, + 3/spl omega//sub P/.) To minimize uncertainty in the final result, it is meaningful to look for paths with bounded variance, i.e., /spl omega//sub P//sup 2/ /spl les/ B for a given fixed bound B. In this paper, we present an exact algorithm to solve the SSP problem in O(B(V + E)) time where V and E are the numbers of vertices and edges, respectively, in G. Our algorithm is superior to previous algorithms for SSP problem because we can handle: 1) general graphs (unlike previous works applicable only to directed acyclic graphs), 2) arbitrary edge-weight distributions (unlike previous algorithms designed only for specific distributions such as Gaussian), and 3) general cost function (none of the previous algorithms can even handle the cost function /spl mu//sub P/, + 3/spl omega//sub P/. Finally, we discuss applications of the SSP problem to maze routing, buffer insertions, and timing analysis under parameter variations. Liang Deng, Martin D. F. Wong |
ASP-DAC | 1 |
| 2006 | A fast simultaneous input vector generation and gate replacement algorithm for leakage power reductionabstractInput vector control (IVC) technique is based on the observation that the leakage current in a CMOS logic gate depends on the gate input state, and a good input vector is able to minimize the leakage when the circuit is in the sleep mode. The gate replacement technique is a very effective method to further reduce the leakage current. In this paper, we propose a fast algorithm to find a low leakage input vector with simultaneous gate replacement. Results on MCNC91 benchmark circuits show that our algorithm produces $14 %$ better leakage current reduction with several orders of magnitude speedup in runtime for large circuits compared to the previous state-of-the-art algorithm. In particular, the average runtime for the ten largest combinational circuits has been dramatically reduced from 1879 seconds to 0.34 seconds. Lei Cheng 0001, Liang Deng, Deming Chen, Martin D. F. Wong |
DAC | 2 |
| 2005 | Floorplanning for 3-D VLSI designabstractIn this paper we present a floorplanning algorithm for 3-D ICs. The problem can be formulated as that of packing a given set of 3-D rectangular blocks while minimizing a suitable cost function. Our algorithm is based on a generalization of the classical 2-D slicing floorplans to 3-D slicing floorplans. A new encoding scheme of slicing floorplans (2-D/3-D) and its associated set of moves form the basis of the new simulated annealing based algorithm. The bestknown algorithm for packing 3-D rectangular blocks is based on simulated annealing using sequence-triple floorplan representation. Experimental results show that our algorithm produces packing results on average 3% better than the sequence-triple-based algorithm under the same annealing parameters, and our algorithm runs much faster (17 times for problems containing 100 blocks) than the sequence-triple. Moreover, our algorithm can be extended to consider various types of placement constraints and thermal distribution while the existing sequence-triple-based algorithm does not have such capabilities. Finally, when specializing to 2-D problems, our algorithm is a new 2-D slicing floorplanning algorithm. We are excited to report the surprising results that our new 2-D floorplanner has produced slicing floorplans for the two largest MCNC benchmarks ami33 and ami49 which have the smallest areas (among all slicing/nonslicing floorplanning algorithms) ever reported in the literature. Lei Cheng 0001, Liang Deng, Martin D. F. Wong |
ASP-DAC | 2 |
| 2005 | Energy optimization in memory address bus structure for application-specific systemsabstractEnergy optimization for high-capacitance on-chip buses has become a critical problem in VLSI design, especially for embedded or SoC systems. Coupling effects between bus wires make this issue even more urgent. Coding schemes have been proposed to reduce the energy dissipation. However, the circuits overhead increases significantly when the coding schemes consider the inter-wire capacitances. In this paper, we present a novel method for energy optimization in memory address bus (MAB). The data on application specified MAB has different characters to the data bus, which has high repetition vectors and unevenly distributed switch activity. Thus a combined method is proposed to optimize the energy consumption by both self capacitance and inter-wire capacitance. First, we lower the switch activity by an efficient coding scheme. Based on the statistical data, a modified bus-invert coding scheme can intelligently divide bus lines into groups and apply bus-invert coding. It brings ultra-low area or timing penalty because of the simple circuit structure. Then the energy consumption of coupling capacitances is optimized by net reordering technique. Implemented with table-look-up technique, a fast simulated annealing algorithm is proposed to solve the net reordering problem. The experimental results show that our combined method is very efficient to reduce the energy consumption in memory address bus for varieties of applications. Liang Deng, Martin D. F. Wong |
ACM Great Lakes Symposium on VLSI | 1 |
| 2005 | Buffer insertion under process variations for delay minimizationabstractThis paper considers the buffer insertion problem under process variations. With continued technology scaling, it is necessary to model the physical parameters to be random variables. One approach to the buffer insertion problem under variations is to use the mean values of these parameters and solve the problem using traditional buffer insertion techniques for delay minimization. Another approach is to find a buffer insertion solution using a new method that can handle the probability distributions. Thus, the performance can be optimized with some yield constraint. In this paper, we present both analytical and experimental results to show that the two approaches give almost identical solutions. In other words, the more expensive statistical methods are not needed for the buffer insertion in delay minimization problem. Liang Deng, Martin D. F. Wong |
ICCAD | 1 |
| 2004 | Optimal Algorithm for Minimizing the Number of Twists in an On-Chip BusabstractComplementary bus architecture is used to achieve higher speed and lower power in VLSI chips. However, in deep submicron circuit design, the effects of crosstalk become more and more serious, especially in the bus structure where wires are placed close to each other. Complementary bus architecture with twisted wires can reduce the coupling noise. But in current chip design flow, engineering change order (ECO) happens commonly to meet improvement requirement. Layout changes due to ECO introduce obstacles to the twists, which could reduce the number of twists and increase the coupling noise. In this paper, an ECO algorithm for generating twisted complementary architecture is proposed based on the shortest path algorithm. Our algorithm guarantees to give the minimum number of twists along the bus wires under noise constraints. Experimental results show that the twist patterns generated by our algorithm can effectively reduce the capacitive coupling noises. Liang Deng, Martin D. F. Wong |
DATE | 1 |