Jiang Lin

dblp:20/1473 · DBLP profile ↗
← Back
26ranked-venue papers
11as first author
10since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 16 · 7 first-author · 4 since 2021Software engineering, systems software and programming languages · 5 · 3 first-authorGraphics, computer vision, multimedia, augmented reality and games · 5 · 2 first-author · 3 since 2021Artificial intelligence and machine learning · 4 · 3 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2026 Computational identification of lineage-committed precursors in mammalian organogenesis reveals a novel hematopoietic enhancer regulating Bhlhe41 expression
abstract
Lineage-committed precursors are essential yet rarely identified in mammalian organogenesis, as they lack definitive molecular signatures required for conventional marker-based approaches. Herein, we developed iCommitted, an integrated multi-omics computational pipeline for precise identification of these precursors. iCommitted first reconstructs in vivo organogenesis by modeling the in vitro differentiation trajectory spanning naïve to terminally differentiated cells. It then integrates epigenomic (ATAC-seq/DNase-seq) and transcriptomic (RNA-seq) data to achieve standardized developmental staging and precursor identification. Applied to mammalian hematopoiesis, iCommitted robustly identified hematopoietic progenitors as the hematopoietic lineage-committed precursors. Subsequent cis-regulatory annotation generated a high-confidence atlas of 16 774 hematopoietic cis-regulatory elements. Functional analysis of the atlas further pinpointed a 218-bp hematopoietic enhancer (chr6:145 855 899-145 856 116) that regulates Bhlhe41 expression during lineage commitment. This study establishes a valuable approach for identifying lineage-committed precursors and elucidating regulatory mechanisms in mammalian organogenesis, offering broad utility in developmental biology.
Lihui Jin, Zhenyuan Han, Rebecca Hannah, Junxin Huang, Jiang Lin
Briefings Bioinform.8
2025 Exploring the Relationship Between Samples and Masks for Robust Defect Localization
abstract
Defect detection aims to detect and localize regions out of the normal distribution. The previous approaches often explicitly incorporate the defect detection concept, such as by utilizing self-supervised ground truth or manually defined feature comparison. The aforementioned processes involve modeling the distribution of normal samples, and they rely on the modeled normality for accurate inference. This reliance may hinder their ability to generalize to unseen test scenarios or the test set that deviates from the training distribution. In this paper, we propose a one-stage framework that detects defective patterns directly without the modeling process. This ability is adopted through the joint efforts of three parties: a generative adversarial network (GAN), a newly proposed scaled pattern loss, and a dynamic correction mechanism that allows the network to self-correct. In training, explicit information that could indicate the position of defects is intentionally excluded to avoid learning any direct mapping. Experimental results show that the proposed method performs superior in comparison with the previous SOTA methods in various test scenarios.
Jiang Lin, Fanxiu Sun, Yaping Yan
AAAI1
2025 OmniStyle: Filtering High Quality Style Transfer Data at Scale
abstract
In this paper, we introduce OmniStyle-1M, a large-scale paired style transfer dataset comprising over one million content-style-stylized image triplets across 1,000 diverse style categories, each enhanced with textual descriptions and instruction prompts. We show that OmniStyle-1M can not only enable efficient and scalable of style transfer models through supervised training but also facilitate precise control over target stylization. Especially, to ensure the quality of the dataset, we introduce OmniFilter, a comprehensive style transfer quality assessment framework, which filters high-quality triplets based on content preservation, style consistency, and aesthetic appeal. Building upon this foundation, we propose OmniStyle, a framework based on the Diffusion Transformer (DiT) architecture designed for high-quality and efficient style transfer. This framework supports both instruction-guided and imageguided style transfer, generating high resolution outputs with exceptional detail. Extensive qualitative and quantitative evaluations demonstrate OmniStyle’s superior performance compared to existing approaches, highlighting its efficiency and versatility. OmniStyle-1M and its accompanying methodologies provide a significant contribution to advancing high-quality style transfer, offering a valuable resource for the research community.
Jiang Lin, Zili Yi, Rui Ma 0011
CVPR3
2025 FreeControl: Efficient, Training-Free Structural Control via One-Step Attention Extraction
abstract
Controlling the spatial and semantic structure of diffusion-generated images remains a challenge. Existing methods like ControlNet rely on handcrafted condition maps and retraining, limiting flexibility and generalization. Inversion-based approaches offer stronger alignment but incur high inference cost due to dual-path denoising. We present \textbf{FreeControl}, a training-free framework for semantic structural control in diffusion models. Unlike prior methods that extract attention across multiple timesteps, FreeControl performs \textit{one-step attention extraction} from a single, optimally chosen timestep and reuses it throughout denoising. This enables efficient structural guidance without inversion or retraining. To further improve quality and stability, we introduce \textit{Latent-Condition Decoupling (LCD)}: a principled separation of the timestep condition and the noised latent used in attention extraction. LCD provides finer control over attention quality and eliminates structural artifacts. FreeControl also supports compositional control via reference images assembled from multiple sources, enabling intuitive scene layout design and stronger prompt alignment. FreeControl introduces a new paradigm for test-time control—enabling structurally and semantically aligned, visually coherent generation directly from raw images, with the flexibility for intuitive compositional design and compatibility with modern diffusion models at ~5\% additional cost.
Jiang Lin, Zhiqiu Zhang, Jizhi Zhang, Qiang Tang 0002, Zili Yi
NeurIPS1
2024 A Comprehensive Augmentation Framework for Anomaly Detection
abstract
Data augmentation methods are commonly integrated into the training of anomaly detection models. Previous approaches have primarily focused on replicating real-world anomalies or enhancing diversity, without considering that the standard of anomaly varies across different classes, potentially leading to a biased training distribution. This paper analyzes crucial traits of simulated anomalies that contribute to the training of reconstructive networks and condenses them into several methods, thus creating a comprehensive framework by selectively utilizing appropriate combinations. Furthermore, we integrate this framework with a reconstruction-based approach and concurrently propose a split training strategy that alleviates the overfitting issue while avoiding introducing interference to the reconstruction process. The evaluations conducted on the MVTec anomaly detection dataset demonstrate that our method outperforms the previous state-of-the-art approach, particularly in terms of object classes. We also generate a simulated dataset comprising anomalies with diverse characteristics, and experimental results demonstrate that our approach exhibits promising potential for generalizing effectively to various unseen anomalies encountered in real-world scenarios.
Jiang Lin, Yaping Yan
AAAI1
2024 GC-UNet: Enhance UNet with GCN for Periodontitis Segmentation
abstract
The intelligent segmentation of periodontitis regions plays a crucial role in the field of dental medicine, significantly aiding in the diagnosis and treatment of the disease. In the domain of medical image segmentation, the UNet architecture, composed of convolutional networks (CNNs), has shown outstanding performance, yielding favorable results in many segmentation tasks. However, when dealing with periodontitis segmentation images that feature both complex overall structures and intricate details with numerous focal points, a UNet based on CNNs suffers from insufficient receptive fields and inadequate capability to capture global feature relationships. Therefore, in this paper, we propose GC-UNet, which designs a GCN-based global feature receptor. This receptor embeds the feature map information obtained by the UNet encoder into graph data and employs a method to calculate edge weights based on spatial distance. By cleverly converting the original spatial positional relationships into an adjacency matrix, our network can capture the associations of global information. To this end, we design a series of experiments to compare the performance of our method with existing methods on the periodontitis segmentation task and to investigate the impact of various key parameter variables on network performance. The experiments demonstrate that GC-UNet exhibits superiority over existing methods in the periodontitis segmentation task.
Wenchong Xu, Wei Li 0100, Jiang Lin
BIBM4
2021 Speech emotion recognition using emotion perception spectral feature
abstract
Summary Speech emotion recognition is an important technique for human‐computer interface applications. Due to contain rich information of emotion, the spectral feature is widely used for emotion recognition. However, the recognition performance is limited because of imprecise extracted rule and uncertain size of resolution of spectral feature. To address this issue, motivated by speech coding, we introduced psychoacoustics model, provided a perception spectral subband partition method for obtaining more precise frequency resolution. Moreover, we also provided a new spectral feature on the divided subband frequency signals. The proposed feature includes emotional perception entropy, spectral inclination, and spectral flatness. Then, a Support Vector Machine classifier is used to recognize emotion categories. The experiment results show that the proposed spectral feature is superior to the traditional MFCC feature, and also better than the state‐of‐the‐art Fourier feature and multi‐resolution amplitude feature.
Jiang Lin, Xingbao Liu
Concurr. Comput. Pract. Exp.1
2021 A novel infant cry recognition system using auditory model-based robust feature and GMM-UBM
abstract
Summary Recognizing infant cry is a meaning work, which can help new parent to understand infant's needs. Mostly, the motivation of existed recognition features is based on the psychoacoustic model. However, weak features are not enough to represent the details of infant cry. To address this issue, we propose a novel infant cry recognition system. In our system, the feature extraction method derived from auditory model, this model can address the auditory neural active representation. Additionally, we designed the intelligent system by a classifier named Gaussian Mixture Model‐Universal Background Model, which has a robust recognition performance to channel imbalance and corroded signal. The experiment results have shown a superior performance. Compared with the High‐order Spectral feature, a state‐of‐the‐art feature, when using Gaussian Mixture Model and General Background Model as classifier, the recognition accuracy improved by 14.6% and 13.7% under clear and corroded signal, respectively.
Jiang Lin, Yumei Yi, Defeng Chen, Xingbao Liu
Concurr. Comput. Pract. Exp.1
2021 Foreword to the special issue of the intelligent systems for the Internet of Things (ISIT2018)
abstract
The purpose of this special issue is to assemble a selection of best research articles that were presented at the third annual workshop on Intelligent Systems for the Internet of Things held in Wuhan. A number of practitioners, scholars, researchers, and vendors attended this workshop. In this workshop, some academic ideas, engineering technologies, future collaborations, and new research directions on IoT were discussed. With the advent of the fourth Industrial Revolution, many technologies, including big data analytics, blockchain, artificial intelligence, and data mining, have also been used for exploring and analyzing data generated from the IoT field. In addition to these, intelligent technologies such as colony optimization, particle swarm optimization, and simulated annealing also provide many solutions for the IoT applications. These provided solutions not only can enhance the performance of an IoT system and its devices, but also can make a system aware of events occurred. This special issue presents many excellent works on how scholars, researchers, vendors, and engineers are collaborating to address high-performance edge complicated research challenges. The scope of this special issue is broad and is representative of the multidisciplinary nature of the IoT. In addition to submissions that deal with intelligent technologies for the IoT and their applications, this issue also includes articles that address practical challenges with IoT-related technologies, such as edge computing, cloud computing, blockchain, computer vision, and deep learning technologies. In the article Intelligent cloud computing platform for 3D sound reproduction,1 3D sound reproduction is a hot topic in virtual reality. Both Dolby and DTS pay great attention to 3D sound reproduction research. Although there exist several 3D sound reproduction methods, few techniques are practical. The requirements in state-of-the-art techniques are critical, such as that all the loudspeakers are on a sphere, the calculation is complicated. The authors develop a novel 3D sound reproduction method and all the parameters in the proposed system are generated on a cloud platform. It is convenient to configure loudspeaker array for users and it is also helpful for the popularization of 3D sound reproduction. In the article Speech emotion recognition using emotion perception spectral feature,2 a new speech spectral feature is provided for speech emotion recognition. Speech emotion recognition is an important technique for human–computer interface applications. Due to its rich information of emotion, the spectral feature is widely used for emotion recognition. However, the recognition performance is limited because of the imprecise extracted rule and uncertain size of resolution of the spectral feature. To address this issue, the authors were motivated by speech coding, they introduced psychoacoustics model, and provided a perception spectral subband partition method for obtaining more precise frequency resolution. Moreover, they also provided a new spectral feature on the divided subband frequency signals. The proposed feature includes emotional perception entropy, spectral inclination, and spectral flatness. Then, a support vector machine classifier is used to recognize emotion categories. The experimental results show that the proposed spectral feature is superior to the traditional MFCC feature, and also better than the state-of-the-art Fourier feature and the multiresolution amplitude feature. This article proposed a new ideal for emotion feature extracting and broaden solution of speech emotion recognition. In the article Real-time action feature extraction via fast PCA-Flow,3 a novel real-time action feature extraction method for human action recognition is provided. Video-based human action recognition can be widely used in network video retrieval, video surveillance analysis, medical video monitoring, and so forth. So it has attracted the attention of scholars who majored in image analysis. The difficulty of action recognition researchers is how to improve the accuracy of action features while reducing the complexity of feature calculation. The authors developed a novel real-time video feature extraction technique by exploiting the fast PCA-Flow algorithm. Experimental results indicate that the proposed method properly balances the accuracy of action features and the time of feature computation. In the article A novel infant cry recognition system using auditory model-based robust feature and GMM-UBM,4 the authors provided a novel infant cry recognition system. Recognizing infant cry is a meaningful work, which can help a new parent understand an infant's needs. Mostly, the motivation of existed recognition features is based on the psychoacoustic model. However, weak features are not enough to represent the details of infant cry. To address this issue, the authors propose a novel infant cry recognition system. In their system, the feature extraction method was derived from the auditory model. This model can address the auditory neural active representation. In addition, they designed the intelligent system by a classifier named Gaussian mixture model-universal background model, which has a robust recognition performance to channel imbalance and corroded signal. The experiment results have shown a superior performance. Compared with the high-order spectral feature, a state-of-the-art feature, when using Gaussian mixture model and general background model as the classifier, the recognition accuracy improved by 14.6% and 13.7% under clear and corroded signals, respectively. This article on infant cry recognition is an interesting work.
Zhibo Wang 0002, Jiang Lin, Bilial Suman
Concurr. Comput. Pract. Exp.2
2021 Intelligent cloud computing platform for three-dimensional sound reproduction
abstract
Summary Three‐dimensional (3D) audio reproduction techniques reproduce realistic sound sources and spatial perception for listeners. However, it is a rather difficult work to configure all the parameters in the previous 3D reproduction systems since ordinary users hardly possess professional acoustic knowledge. In this article, we developed a sound source distance reproduction model and design an online parameter computing framework for users to generate reproduction parameters based on a cloud computing platform. The proposed model reproduces sound direction and distance cues under specific conditions. Subjective experiments are executed to assess the spatial reproduction performance in a real environment and objective experiments are performed in simulated scenarios. Both the subjective and objective experiments show that the proposed method improves the spatial perception of sound events. Since the reproduction model depends on environment and loudspeaker array configuration, a cloud parameter generation framework is developed for users to generate acoustic parameters for the reproduction system.
Maosheng Zhang, Ruimin Hu, Jiang Lin, Xiaochen Wang 0001
Concurr. Comput. Pract. Exp.3
2014 Mini-Rank: A Power-EfficientDDRx DRAM Memory Architecture
abstract
Memory power consumption has become a severe concern in multi-core computer platforms. As memory data rate, capacity and bandwidth are being pushed higher and higher, the power consumption of memory systems becomes a significant part in the overall system power profile. Conventional memory systems do not provide an efficient mechanism for managing its power and performance tradeoff. We propose a novel mini-rank architecture for DDRx memories to reduce memory power consumption by breaking each DRAM rank into multiple narrow mini-ranks and activating fewer devices for each request. We also propose a heterogeneous mini-rank design to further improve the performance-power tradeoff for each workload based on its memory access behavior and bandwidth requirement. The evaluation results show that homogeneous mini-rank significantly reduces memory power with small performance loss. For instance, using four-core multiprogramming workloads, a x32 mini-rank configuration reduces memory power by 19.5 percent with 1.3 percent performance loss on average for memory-intensive workloads. Heterogeneous mini-rank further improves the balance between the performance and power saving. For instance, it reduces the memory power by up to 38.0 percent with an average performance loss of 2.4 percent, compared with a conventional memory system. In comparison, the x32 homogeneous mini-rank reduces memory power by up to 25.4 percent; while the x8 homogeneous mini-rank incurs performance loss by up to 19.3 percent. Furthermore, heterogeneous mini-rank achieves consistently good performance-power tradeoff for workloads made by programs of diverse memory access behavior and bandwidth requirement.
Kun Fang 0005, Hongzhong Zheng, Jiang Lin, Zhao Zhang 0010, Zhichun Zhu
IEEE Trans. Computers3
2013 Thermal Modeling and Management of DRAM Systems
abstract
With increasing data rate and power density, high-performance memories have started to require dynamic thermal management (DTM), following the trend of processor and hard drive. There are also lack of a memory thermal model and simulation tools to facilitate the research of memory DTM. This study investigates the approach of coordinating processor, which is the source of memory access requests, and memory to improve system performance and/or power efficiency during memory thermal emergency. Two such schemes, namely adaptive core gating (DTM-ACG) and coordinated DVFS (DTM-CDVFS), are proposed and evaluated on a real server platform. DTM-ACG gates processor cores and DTM-CDVFS scales down the frequency and voltage level of processor cores according to memory thermal emergency level. Their combination, namely DTM-COMB, is also evaluated. The experimental results show that the two schemes, while successfully controlling memory activities and handling thermal emergencies, improve performance significantly under the given thermal envelope. The measurement results from an Intel SR1500AL server testbed show that on average, DTM-ACG and DTM-CDVFS improve performance by 6.7 and 15.3 percent, respectively, over a prior memory bandwidth throttling scheme. DTM-CDVFS also reduces the processor power rate by 15.5 percent and system (including processor and memory) energy by 22.7 percent. Additionally, we propose a DRAM thermal model and validate it with measurement on the instrumented server platform. We find that our proposed model faithfully catches the dynamic DRAM temperature changes; the average difference between the modeled and measured temperature is less than $(1^{\circ}{\rm C})$.
Jiang Lin, Hongzhong Zheng, Zhichun Zhu, Zhao Zhang 0010
IEEE Trans. Computers1
2011 Collision-streams: fast GPU-based collision detection for deformable models
abstract
We present a fast GPU-based streaming algorithm to perform collision queries between deformable models. Our approach is based on hierarchical culling and reduces the computation to generating different streams. We present a novel stream registration method to compact the streams and efficiently compute the potentially colliding pairs of primitives. We also use a deferred front tracking method to lower the memory overhead. The overall algorithm has been implemented on different GPUs and we have evaluated its performance on non-rigid and deformable simulations. We highlight our speedups over prior CPU-based and GPU-based algorithms. In practice, our algorithm can perform inter-object and intra-object computations on models composed of hundreds of thousands of triangles in tens of milliseconds.
Min Tang 0001, Dinesh Manocha, Jiang Lin, Ruofeng Tong 0001
SI3D3
2010 Enigma: architectural and operating system support for reducing the impact of address translation
abstract
Most modern microprocessors provide hardware support for rapidly translating a program logical address to a system physical address (PA). Translation typically sits on the critical path of every memory access, since an access cannot usually be performed until after it has been translated. Enigma is a novel approach to address translation that defers the bulk of the work associated with address translation until data must be retrieved from physical memory. Enigma replaces the address translation unit that exists in each conventional core with a simpler unit to translate from the logical address space to a new intermediate address (IA) space. Intermediate addresses are unique across the entire system except where sharing is required or desired, and their use sidesteps the "synonym" problem present in logically tagged caches. All cache addressing, as well as I/O and coherence traffic, is carried out using IA. Enigma translates an IA to a PA only when no cache in the entire CMP can satisfy the request and memory or I/O must be accessed. A central translation unit attached to the system bus performs translations on IA that must be resolved to a PA. Deferring the bulk of address translation work and removing it from each individual processor core in this manner affords many benefits.
Lixin Zhang 0002, William Evan Speight, Ramakrishnan Rajamony, Jiang Lin
ICS4
2009 Soft-OLP: Improving Hardware Cache Performance through Software-Controlled Object-Level Partitioning
abstract
Performance degradation of memory-intensive programs caused by the LRU policy's inability to handle weak-locality data accesses in the last level cache is increasingly serious for two reasons. First, the last-level cache remains in the CPU's critical path, where only simple management mechanisms, such as LRU, can be used, precluding some sophisticated hardware mechanisms to address the problem. Second, the commonly used shared cache structure of multi-core processors has made this critical path even more performance-sensitive due to intensive inter-thread contention for shared cache resources. Researchers have recently made efforts to address the problem with the LRU policy by partitioning the cache using hardware or OS facilities guided by run-time locality information. Such approaches often rely on special hardware support or lack enough accuracy. In contrast, for a large class of programs, the locality information can be accurately predicted if access patterns are recognized through small training runs at the data object level. To achieve this goal, we present a system-software framework referred to as Soft-OLP (Software-based Object-Level cache Partitioning). We first collect per-object reuse distance histograms and inter-object interference histograms via memory-trace sampling. With several low-cost training runs, we are able to determine the locality patterns of data objects. For the actual runs, we categorize data objects into different locality types and partition the cache space among data objects with a heuristic algorithm, in order to reduce cache misses through segregation of contending objects. The object-level cache partitioning framework has been implemented with a modified Linux kernel, and tested on a commodity multi-core processor. Experimental results show that in comparison with a standard L2 cache managed by LRU, Soft-OLP significantly reduces the execution time by reducing L2 cache misses across inputs for a set of single- and multi-threaded programs from the SPEC CPU2000 benchmark suite, NAS benchmarks and a computational kernel set.
Qingda Lu, Jiang Lin, Xiaoning Ding, Zhao Zhang 0010, Xiaodong Zhang 0001, P. Sadayappan
PACT2
2009 Decoupled DIMM: building high-bandwidth memory system using low-speed DRAM devices
abstract
The widespread use of multicore processors has dramatically increased the demands on high bandwidth and large capacity from memory systems. In a conventional DDR2/DDR3 DRAM memory system, the memory bus and DRAM devices run at the same data rate. To improve memory bandwidth, we propose a new memory system design called decoupled DIMM that allows the memory bus to operate at a data rate much higher than that of the DRAM devices. In the design, a synchronization buffer is added to relay data between the slow DRAM devices and the fast memory bus; and memory access scheduling is revised to avoid access conflicts on memory ranks. The design not only improves memory bandwidth beyond what can be supported by current memory devices, but also improves reliability, power efficiency, and cost effectiveness by using relatively slow memory devices. The idea of decoupling, precisely the decoupling of bandwidth match between memory bus and a single rank of devices, can also be applied to other types of memory systems including FB-DIMM.
Hongzhong Zheng, Jiang Lin, Zhao Zhang 0010, Zhichun Zhu
ISCA2
2009 Enabling software management for multicore caches with a lightweight hardware support
abstract
The management of shared caches in multicore processors is a critical and challenging task. Many hardware and OS-based methods have been proposed. However, they may be hardly adopted in practice due to their non-trivial overheads, high complexities, and/or limited abilities to handle increasingly complicated scenarios of cache contention caused by many-cores.
Jiang Lin, Qingda Lu, Xiaoning Ding, Zhao Zhang 0010, Xiaodong Zhang 0001, P. Sadayappan
SC1
2008 Gaining insights into multicore cache partitioning: Bridging the gap between simulation and real systems
abstract
Cache partitioning and sharing is critical to the effective utilization of multicore processors. However, almost all existing studies have been evaluated by simulation that often has several limitations, such as excessive simulation time, absence of OS activities and proneness to simulation inaccuracy. To address these issues, we have taken an efficient software approach to supporting both static and dynamic cache partitioning in OS through memory address mapping. We have comprehensively evaluated several representative cache partitioning schemes with different optimization objectives, including performance, fairness, and quality of service (QoS). Our software approach makes it possible to run the SPEC CPU2006 benchmark suite to completion. Besides confirming important conclusions from previous work, we are able to gain several insights from whole-program executions, which are infeasible from simulation. For example, giving up some cache space in one program to help another one may improve the performance of both programs for certain workloads due to reduced contention for memory bandwidth. Our evaluation of previously proposed fairness metrics is also significantly different from a simulation-based study. The contributions of this study are threefold. (1) To the best of our knowledge, this is a highly comprehensive execution- and measurement-based study on multicore cache partitioning. This paper not only confirms important conclusions from simulation-based studies, but also provides new insights into dynamic behaviors and interaction effects. (2) Our approach provides a unique and efficient option for evaluating multicore cache partitioning. The implemented software layer can be used as a tool in multicore performance evaluation and hardware design. (3) The proposed schemes can be further refined for OS kernels to improve performance.
Jiang Lin, Qingda Lu, Xiaoning Ding, Zhao Zhang 0010, Xiaodong Zhang 0001, P. Sadayappan
HPCA1
2008 Memory Access Scheduling Schemes for Systems with Multi-Core Processors
abstract
On systems with multi-core processors, the memory access scheduling scheme plays an important role not only in utilizing the limited memory bandwidth but also in balancing the program execution on all cores. In this study, we propose a scheme, called ME-LREQ, which considers the utilization of both processor cores and memory subsystem. It takes into consideration both the long-term and short-term gains of serving a memory request by prioritizing requests hitting on the row buffers and from the cores that can utilize memory more efficiently and have fewer pending requests. We have also thoroughly evaluated a set of memory scheduling schemes that differentiate and prioritize requests from different cores. Our simulation results show that for memory-intensive, multiprogramming workloads, the new policy improves the overall performance by 10.7% on average and up to 17.7% on a four-core processor, when compared with scheme that serves row buffers hit memory requests first and allows memory reads bypassing writes; and by up to 9.2% (6.4% on average) when compared with the scheme that serves requests from the core with the fewest pending requests first.
Hongzhong Zheng, Jiang Lin, Zhao Zhang 0010, Zhichun Zhu
ICPP2
2008 Mini-rank: Adaptive DRAM architecture for improving memory power efficiency
abstract
The widespread use of multicore processors has dramatically increased the demand on high memory bandwidth and large memory capacity. As DRAM subsystem designs stretch to meet the demand, memory power consumption is now approaching that of processors. However, the conventional DRAM architecture prevents any meaningful power and performance trade-offs for memory-intensive workloads. We propose a novel idea called mini-rank for DDRx (DDR/DDR2/DDR3) DRAMs, which uses a small bridge chip on each DRAM DIMM to break a conventional DRAM rank into multiple smaller mini-ranks so as to reduce the number of devices involved in a single memory access. The design dramatically reduces the memory power consumption with only a slight increase on the memory idle latency. It does not change the DDRx bus protocol and its configuration can be adapted for the best performance-power trade-offs. Our experimental results using four-core multiprogramming workloads show that using x32 mini-ranks reduces memory power by 27.0% with 2.8% performance penalty and using x16 mini-ranks reduces memory power by 44.1% with 7.4% performance penalty on average for memory-intensive workloads, respectively.
Hongzhong Zheng, Jiang Lin, Zhao Zhang 0010, Eugene Gorbatov, Howard David, Zhichun Zhu
MICRO2
2008 Software thermal management of dram memory for multicore systems
abstract
Thermal management of DRAM memory has become a critical issue for server systems. We have done, to our best knowledge, the first study of software thermal management for memory subsystem on real machines. Two recently proposed DTM (Dynamic Thermal Management) policies have been improved and implemented in Linux OS and evaluated on two multicore servers, a Dell PowerEdge 1950 server and a customized Intel SR1500AL server testbed. The experimental results first confirm that a system-level memory DTM policy may significantly improve system performance and power efficiency, compared with existing memory bandwidth throttling scheme. A policy called DTM-ACG (Adaptive Core Gating) shows performance improvement comparable to that reported previously. The average performance improvements are 13.3% and 7.2% on the PowerEdge 1950 and the SR1500AL (vs. 16.3% from the previous simulation-based study), respectively. We also have surprising findings that reveal the weakness of the previous study: the CPU heat dissipation and its impact on DRAM memories, which were ignored, are significant factors. We have observed that the second policy, called DTM-CDVFS (Coordinated Dynamic Voltage and Frequency Scaling), has much better performance than previously reported for this reason. The average improvements are 10.8% and 15.3% on the two machines (vs. 3.4% from the previous study), respectively. It also significantly reduces the processor power by 15.5% and energy by 22.7% on average.
Jiang Lin, Hongzhong Zheng, Zhichun Zhu, Eugene Gorbatov, Howard David, Zhao Zhang 0010
SIGMETRICS1
2007 A Scheme for System Multiplexing and Program Component Identification
abstract
A novel system multiplexing and program component identification scheme is proposed, which has been adopted by AVS standard working group. Multiplexer periodically inserts program element information table (PEIT) describing ordering information of transport packets of program elements contained within transport stream. Demultiplexer extracts PEIT transport packets by a unique PEIT indicator and parses PEIT. With the position and count information recovered from PEIT, transport packets containing program specific information and elementary streams are identified, de-multiplexed and sent to matched decoders for further processing. To improve system error resilience, Packet Link Table (PLT) is introduced.
Yaqiang Ding, Jiang Lin, Fuhuei Lin, Yi Kang
ICME2
2007 Thermal modeling and management of DRAM memory systems
abstract
With increasing speed and power density, high-performance memories, including FB-DIMM (Fully Buffered DIMM) and DDR2 DRAM, now begin to require dynamic thermal management(DTM) as processors and hard drives did. The DTM of memories, nevertheless, is different in that it should take the processor performance and power consumption into consideration. Existing schemes have ignored that. In this study, we investigate a new approach that controls the memory thermal issues from the source generating memory activities - the processor. It will smooth the program execution when compared with shutting down memory abruptly, and therefore improve the overall system performance and power efficiency. For multicore systems, we propose two schemes called adaptive core gating and coordinated DVFS. The first scheme activates clock gating on selected processor cores and the second one scales down the frequency and voltage levels of processor cores when the memory is to be over-heated. They can successfully control the memory activities and handle thermal emergency. More importantly, they improve performance significantly under the given thermal envelope. Our simulation results show that adaptive coregating improves performance by up to 23.3% (16.3% on average) on a four-core system with FB-DIMM when compared with DRAM thermal shutdown; and coordinated DVFS with control-theoretic methods improves the performance by up to 18.5% (8.3% on average).
Jiang Lin, Hongzhong Zheng, Zhichun Zhu, Howard David, Zhao Zhang 0010
ISCA1
2007 DRAM-Level Prefetching for Fully-Buffered DIMM: Design, Performance and Power Saving
abstract
We have studied DRAM-level prefetching for the fully buffered DIMM (FB-DIMM) designed for multi-core processors. FB-DIMM has a unique two-level interconnect structure, with FB-DIMM channels at the first-level connecting the memory controller and Advanced Memory Buffers (AMBs); and DDR2 buses at the second-level connecting the AMBs with DRAM chips. We propose an AMB prefetching method that prefetches memory blocks from DRAM chips to AMBs. It utilizes the redundant bandwidth between the DRAM chips and AMBs but does not consume the crucial channel bandwidth. The proposed method fetches K memory blocks of L2 cache block sizes around the demanded block, where K is a small value ranging from two to eight. The method may also reduce the DRAM power consumption by merging some DRAM precharges and activations. Our cycle-accurate simulation shows that the average performance improvement is 16% for single-core and multi-core workloads constructed from memory-intensive SPEC2000 programs with software cache prefetching enabled; and no workload has negative speedup. We have found that the performance gain comes from the reduction of idle memory latency and the improvement of channel bandwidth utilization. We have also found that there is only a small overlap between the performance gains from the AMB prefetching and the software cache prefetching. The average of estimated power saving is 15%.
Jiang Lin, Hongzhong Zheng, Zhichun Zhu, Zhao Zhang 0010, Howard David
ISPASS1
2005 Performance Characterization of Java Applications on SMT Processors
abstract
As Java is emerging as one of the major programming languages in software development, studying how Java applications behave on recent SMT processors is of great interest. This paper characterizes the performance of Java applications on an Intel Pentium 4 hyper-threading processor. Using the performance counters provided by Pentium 4, we quantitatively evaluate micro-architecture metrics while running various types of Java applications. The experimental results reveal that: (1) Hyper-threading can indeed improve the performance of multithreaded Java programs; (2) The resource contentions within Pentium 4 are the major reason of pipeline inefficiency, which prevents better performance promised by SMT; (3) The static partition design of hyper-threading causes considerable performance loss for many single-thread Java programs; (4) Most multiprogrammed Java benchmarks can achieve decent combined speedups on hyper-threading processors
Wei Huang 0032, Jiang Lin, Zhao Zhang 0010, J. Morris Chang
ISPASS2
2005 Towards Pairing Java Applications on SMT Processors
abstract
This paper investigates various issues of pairing Java applications for multithreaded execution on Intel's hyper-threading Pentium 4 processor. We first quantify the overall performance of multiprogrammed Java applications using a metric called combined speedup. Using the performance counters provided by the Pentium 4, we then quantitatively evaluate the performance of underneath micro-architecture components and their implications to the combined speedup. A statistical model is proposed to analyze the collected data. This novel approach reveals that trace cache is the major factor determining the pairing performance. In particular, we find that the trace cache miss rates of Java applications can be utilized to predict the combined speedups. Three new scheduling strategies are proposed based on these observations and then evaluated. The experimental results show that the proposed strategies have better performance than the conventional round-robin scheduling scheme. Overall, our best strategy enables a reduction in execution time of 10.5% over the serial execution, comparing with a reduction of 5.92% achieved by the round-robin scheduling. The improvement will be increasingly significant on future SMT processors.
Wei Huang 0032, Jiang Lin, Zhao Zhang 0010, J. Morris Chang
MASCOTS2