VLDB 2026 Research / reviewers in the wild / expert
Jieyi Long
dblp:95/2892
· DBLP profile ↗
25ranked-venue papers
12as first author
12since 2021 · last 2026
0009-0007-4646-7131ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 15 · 11 first-author · 2 since 2021Databases, data management, data science and information retrieval · 5 · 5 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 4 since 2021Artificial intelligence and machine learning · 3 · 3 since 2021Security and privacy · 3 · 1 first-author · 3 since 2021Software engineering, systems software and programming languages · 3 · 3 first-authorGraphics, computer vision, multimedia, augmented reality and games · 3 · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | ShuttleCross: An Efficient Cross-Chain Smart Contract Invocation Framework
Rongkai Zhang 0005, Qiuyu Ding, Qianyi Liu, Shengjie Guan, Jieyi Long |
DSN | 6 |
| 2026 | Interstellar: Efficient GKR-Based IVC Scheme with Privacy-Preserving Collaborative Folding
Jieyi Long |
PKC (3) | 1 |
| 2026 | Alzo: Auto-Tuning with Reinforcement Learning for DAG-based BlockchainsabstractAs critical infrastructure for Web 3.0, DAG-based blockchains promise high throughput for DeFi, IoT, and DApps. However, realizing this potential is challenging, as system performance is dictated by a multitude of interdependent parameters across network, node, and consensus layers. Manual configuration fails to adapt to dynamic workloads, leading to suboptimal performance. We introduce Alzo, a novel auto-tuner that employs hierarchical reinforcement learning (HRL) to navigate this complex configuration space. By decomposing the DAG blockchain's workflow into distinct stages, Alzo's HRL policy learns from stage-level performance metrics to control critical parameters governing consensus, execution, and graph topology in real-time. Furthermore, we employ a shadow-control loop to ensure the safety of all parameter adjustments. Our experiments show that Alzo significantly outperforms other configurations, achieving higher throughput and lower latency under variable workloads with minimal overhead. Qiuyu Ding, Rongkai Zhang 0005, Qinnan Zhang, Jieyi Long, Mingchao Wan, Jin Dong 0004 |
WWW | 5 |
| 2026 | DARA: Few-shot Budget Allocation in Online Advertising via In-Context Decision Making with RL-Finetuned LLMsabstractOptimizing the advertiser's cumulative value of winning impressions under budget constraints poses a complex challenge in online advertising, under the paradigm of AI-Generated Bidding (AIGB). Advertisers often have personalized objectives but limited historical interaction data, resulting in few-shot scenarios where traditional reinforcement learning (RL) methods struggle to perform effectively. Large Language Models (LLMs) offer a promising alternative for AIGB by leveraging their in-context learning capabilities to generalize from limited data. However, they lack the numerical precision required for fine-grained optimization. To address this limitation, we introduce GRPO-Adaptive, an efficient LLM post-training strategy that enhances both reasoning and numerical precision by dynamically updating the reference policy during training. Built upon this foundation, we further propose DARA, a novel dual-phase framework that decomposes the decision-making process into two stages: a few-shot reasoner that generates initial plans via in-context prompting, and a fine-grained optimizer that refines these plans using feedback-driven reasoning. This separation allows DARA to combine LLMs' in-context learning strengths with precise adaptability required by AIGB tasks. Extensive experiments on both real-world and synthetic data environments demonstrate that our approach consistently outperforms existing baselines in terms of cumulative advertiser value under budget constraints. Mingxuan Song, Yusen Huo, Shenglin Yin, Jieyi Long, Zhilin Zhang 0003, Chuan Yu 0002 |
WWW | 6 |
| 2026 | Nexus: A Novel Transaction Processing Framework for Permissioned BlockchainabstractThe transaction execution layer is a key determinant of throughput in permissioned blockchains. While recent Shared Memory Pools (SMP)-based approaches improve throughput by enabling all consensus nodes to participate in transaction packaging, they face two fundamental limitations. First, the performance bottleneck shifts from the consensus layer to the transaction execution layer as transaction number confirmed in a round increases. Second, these approaches are vulnerable to “transaction duplication” attacks where malicious clients can simultaneously send the same transaction to multiple consensus nodes, thereby decreasing the number of valid transactions in block proposals. To address these limitations, this paper introducesNexus, a novel blockchain transaction processing framework with high scalability.Nexusleverages the idle computational resources of full nodes to enable transaction execution in parallel with the consensus. Moreover,Nexusallows each node to handle only a fraction of the total transactions and share execution results with others. This approach reduces overall transaction execution time, increases throughput, and decreases latency. Lastly,Nexusintroduces a transaction partitioning mechanism that effectively addresses the “transaction duplication” attack and achieves load balancing between clients and consensus nodes. Our implementation ofNexusdemonstrates significant improvements: throughput increases by 4x to 15x, and latency is reduced by 50% to 70%. Shengjie Guan, Rongkai Zhang 0005, Qiuyu Ding, Mingxuan Song, Jieyi Long, Mingchao Wan, Taifu Yuan, Jin Dong 0004 |
IEEE Trans. Parallel Distributed Syst. | 6 |
| 2025 | MMGDreamer: Mixed-Modality Graph for Geometry-Controllable 3D Indoor Scene GenerationabstractControllable 3D scene generation has extensive applications in virtual reality and interior design, where the generated scenes should exhibit high levels of realism and controllability in terms of geometry. Scene graphs provide a suitable data representation that facilitates these applications. However, current graph-based methods for scene generation are constrained to text-based inputs and exhibit insufficient adaptability to flexible user inputs, hindering the ability to precisely control object geometry. To address this issue, we propose MMGDreamer, a dual-branch diffusion model for scene generation that incorporates a novel Mixed-Modality Graph, visual enhancement module, and relation predictor. The mixed-modality graph allows object nodes to integrate textual and visual modalities, with optional relationships between nodes. It enhances adaptability to flexible user inputs and enables meticulous control over the geometry of objects in the generated scenes. The visual enhancement module enriches the visual fidelity of text-only nodes by constructing visual representations using text embeddings. Furthermore, our relation predictor leverages node representations to infer absent relationships between nodes, resulting in more coherent scene layouts. Extensive experimental results demonstrate that MMGDreamer exhibits superior control of object geometry, achieving state-of-the-art scene generation performance. Zhifei Yang 0004, Keyang Lu, Jiaxing Qi, Hanqi Jiang, Ruifei Ma, Shenglin Yin, Yifan Xu 0028, Mingzhe Xing, Jieyi Long, Xiangde Liu, Guangyao Zhai |
AAAI | 11 |
| 2025 | AERO: Enhancing Sharding Blockchain via Deep Reinforcement Learning for Account MigrationabstractSharding blockchain networks face significant scalability challenges due to high frequencies of cross-shard transactions and uneven workload distributions among shards. To address these scalability issues, account migration offers a promising solution. However, existing migration solutions struggle with the high computational overhead and insufficient capture of complex transaction patterns. We propose AERO, a deep reinforcement learning framework to facilitate efficient account migration in sharding blockchains. AERO employs a prefix-based grouping strategy to enable group-level migration decisions and capture complex transaction patterns and relationships between accounts. We also implement a sharding blockchain system called AEROChain, which integrates AERO and aligns with the blockchain decentralization principle. Extensive evaluation with real Ethereum transaction data demonstrates that AERO improves the system throughput by 31.77% compared to existing solutions, effectively reducing cross-shard transactions and balancing shard workloads. Mingxuan Song, Pengze Li, Shenglin Yin, Jieyi Long |
WWW | 6 |
| 2024 | Adversarial Distillation Based on Slack Matching and Attribution Region AlignmentabstractAdversarial distillation (AD) is a highly effective method for enhancing the robustness of small models. Contrary to expectations, a high-performing teacher model does not always result in a more robust student model. This is due to two main reasons. First, when there are significant differences in predictions between the teacher model and the student model, exact matching of predicted values using KL divergence interferes with training, leading to poor performance of existing methods. Second, matching solely based on the output prevents the student model from fully understanding the behavior of the teacher model. To address these challenges, this paper proposes a novel AD method named SmaraAD. During the training process, we facilitate the student model in better understanding the teacher model's behavior by aligning the attribution region that the student model focuses on with that of the teacher model. Concurrently, we relax the condition of exact matching in KL divergence and replace it with a more flexible matching criterion, thereby enhancing the model's robustness. Extensive experiments substantiate the effectiveness of our method in improving the robustness of small models, out-performing previous SOTA methods. Shenglin Yin, Mingxuan Song, Jieyi Long |
CVPR | 4 |
| 2024 | Embracing Adaptation: An Effective Dynamic Defense Strategy Against Adversarial ExamplesabstractExisting adversarial example defense methods are static, meaning they remain unchanged once training is completed, regardless of how attack methods change. Consequently, static defense methods are highly vulnerable to adaptive attacks. We argue that to counter more formidable attacks, models should continually adapt to various attack methods. We propose a novel dynamic defense approach. Initially, we use Gaussian Mixture Models (GMM) to obtain structural information of the data, which is combined with model prediction information to generate pseudo-labels for optimizing inputs. Subsequently, we employ information maximization and enhanced mean predictions as optimization objectives, utilizing a hierarchical optimization approach to refine the model. Meanwhile, we propose a sample-efficient optimization strategy that reduces the total number of samples in the test data stream for reverse updating and improves the efficiency. Notably, our method can be directly applied to pre-trained models without the need for accessing training data or retraining the model. Therefore, our approach is training-data-agnostic and model-agnostic, easily applicable to existing adversarially trained models, significantly enhancing the resilience of various models against white-box, black-box, and adaptive attacks across diverse datasets. We have conducted extensive experiments to validate the state-of-the-art of our proposed method. The pseudo-code can be found in the appendix. Shenglin Yin, Kelu Yao, Jieyi Long |
ACM Multimedia | 4 |
| 2024 | Presto: Optimizing Cross-Shard Transactions in Sharded Blockchain ArchitectureabstractBlockchain sharding technology has been used to enhance the scalability of blockchain systems. As the number of shards increases, the high latency inherent in cross-shard transactions gradually becomes a bottleneck, hindering improvements in overall system efficiency. Therefore, reducing the latency of cross-shard transactions is significantly important. However, existing mechanisms for handling cross-shard transactions fail to minimize the latency of cross-shard transactions and have not fully used the bandwidth available within shards. In this paper, we introduce Presto, a protocol designed for the account-state-based blockchain, which reduces the latency of handling cross-shard transactions. Presto leverages the concept of optimistic pre-execution along with pending tree to optimize cross-shard transaction processing. Presto also employs predistribution of cross-shard transactions with Erasure Coding to efficiently utilize bandwidth resources. We have developed an prototype and conducted extensive experiments on a cloud platform. The evaluation results indicate that Presto surpasses existing solutions in terms of system throughput, transaction confirmation latency, and mempool queue size, demonstrating Presto's potential to significantly improve blockchain scalability and user experience. Qiuyu Ding, Rongkai Zhang 0005, Shenglin Yin, Pengze Li, Shengjie Guan, Jieyi Long |
SRDS | 7 |
| 2024 | SPRING: Improving the Throughput of Sharding Blockchain via Deep Reinforcement Learning Based State PlacementabstractSharding provides an opportunity to overcome the inherent scalability challenges of the blockchain, which is the infrastructure for the next generation of the Web. In a sharding blockchain, the state is partitioned into smaller groups known as "shards." Since the states are placed on different shards, cross-shard transactions are inevitable, which is detrimental to the performance of the sharding blockchain. Existing solutions place states based on heuristic algorithms or redistribute states via graph-partitioning-based methods, which are either less effective or costly. In this paper, we present SPRING, the first deep-reinforcement-learning(DRL)-based sharding framework for state placement. SPRING formulates the state placement as a Markov Decision Process, which considers the cross-shard transaction ratio and workload balancing and employs DRL to learn the effective state placement policy. Experimental results based on real Ethereum transaction data demonstrate the superiority of SPRING compared to other state placement solutions. In particular, it decreases the cross-shard transaction ratio by up to 26.63% and boosts throughput by up to 36.03%, all without unduly sacrificing the workload balance among shards. Moreover, updating the training model and making decisions takes only 0.1s and 0.002s, respectively, which shows the overhead is acceptable. Pengze Li, Mingxuan Song, Mingzhe Xing, Qiuyu Ding, Shengjie Guan, Jieyi Long |
WWW | 7 |
| 2023 | A Dual-Agent Scheduler for Distributed Deep Learning Jobs on Public Cloud via Reinforcement LearningabstractPublic cloud GPU clusters are becoming emerging platforms for training distributed deep learning jobs. Under this training paradigm, the job scheduler is a crucial component to improve user experiences, i.e., reducing training fees and job completion time, which can also save power costs for service providers. However, the scheduling problem is known to be NP-hard. Most existing work divides it into two easier sub-tasks, i.e., ordering task and placement task, which are responsible for deciding the scheduling orders of jobs and placement orders of GPU machines, respectively. Due to the superior adaptation ability, learning-based policies can generally perform better than traditional heuristic-based methods. Nevertheless, there are still two main challenges that have not been well-solved. First, most learning-based methods only focus on ordering or placement policy independently, while ignoring their cooperation. Second, the unbalanced machine performances and resource contention impose huge overhead and uncertainty on job duration, but rarely be considered in existing work. To tackle these issues, this paper presents a dual-agent scheduler framework abstracted from the two sub-tasks to jointly learn the ordering and placement policies and make better-informed scheduling decisions. Specifically, we design an ordering agent with a scalable squeeze-and-communicate strategy for better cooperation; for the placement agent, we propose a novel Random Walk Gaussian Process to learn the performance similarities of GPU machines while being aware of the uncertain performance fluctuation. Finally, the dual-agent is jointly optimized with multi-agent reinforcement learning. Extensive experiments conducted on the real-world production cluster trace demonstrate the superiority of our model. Mingzhe Xing, Hangyu Mao, Shenglin Yin, Lichen Pan, Zhengchao Zhang, Jieyi Long |
KDD | 7 |
| 2013 | Theory and Analysis for Optimization of On-Chip Thermoelectric Cooling SystemsabstractWe established a novel theoretical analysis framework for optimizing the cooling system configuration of chips employing thermoelectric cooling (TEC) elements by extending the theory of inverse-positive matrices and the eigenvalue/eigenvector theory in linear algebra. In this brief, we present a new theorem and its formal proof, which is the key enabler to achieving a provably optimal solution for configuring bias current levels of TEC devices. Jieyi Long, Seda Ogrenci Memik, Semail Ülgen |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2010 | A framework for optimizing thermoelectric active cooling systemsabstractThin-film thermoelectric cooling is a promising technology for mitigating heat dissipation in high performance chips. In this paper, we present an optimization framework for an active cooling system that is comprised of an array of thin-film thermoelectric coolers. We observe a set of constraints of the cooling system design. Firstly, integrating an excessive amount of coolers increases the chip package cost. Moreover, thermoelectric coolers are active devices, which dissipate heat in the chip package when they are in operation. Hence, setting the supply current level to operate the cooler improperly can actually lead to overheating of the chip package. Besides, the supply current needs to be delivered to the integrated cooler devices via dedicated pins. However, extra pins available on high-performance chip packages are limited. Observing these constraints, we propose an optimization framework for configuring the active cooling system, which minimizes the maximum silicon temperature. This includes determining the amount of coolers to deploy and their locations, the mapping of supply pins to the coolers, and determining the current levels of each pin. We propose algorithms to tackle the optimal configuration problem. We found that only a small portion of the silicon die needs to be covered by TEC devices (18% on average). Our experiments show that our algorithms are able to reduce the temperatures of the hot spots by as much as 10.6 °C (compared to the cases without integrated thermoelectric coolers). The average temperature reduction is 8.6 °C when 4 dedicated pins are available on the package. The total power consumption of the resulting active cooling system is reasonably small (~ 2 W). Our experiments also reveal that our framework maximizes the efficiency of the cooling devices. In the ideal case where hundreds of pins are available to tune the supply level of each individual cooler, the additional average reduction of the hot spot temperature is only 0.3 °C. Jieyi Long, Seda Ogrenci Memik |
DAC | 1 |
| 2010 | Optimization of the bias current network for accurate on-chip thermal monitoringabstractMicroprocessor chips employ increasingly larger number of thermal sensing devices. These devices are networked by an underlying infrastructure, which provides bias currents to sensing devices and collects measurements. In this work, we address the optimization of the bias current distribution network utilized by the sensing devices. We show that the choice between two fundamental topologies (the 2-wire and the 4-wire measurement) for this network has a non-negligible impact on the precision of the monitoring system. We also show that the 4-wire measurement principle supports the remote sensing technique better. However, it requires more routing resources. We thus propose a novel routing algorithm to minimize its routing cost. We also present a detailed evaluation of the quality of the resulting system in presence of process and thermal variations. Our Monte Carlo simulations using the IBM 10SF 65nm SPICE models show that the monitoring accuracies can be as high as 0.6°C under considerable amount of process and temperature variation. Moreover, by adopting a customized routing approach for the current mirror network, the total wire length of the bias current network can be reduced by as much as 42.74% and by 27.65% on average. Jieyi Long, Seda Ogrenci Memik |
DATE | 1 |
| 2010 | Inversed Temperature Dependence aware clock skew scheduling for sequential circuitsabstractWe present an Inversed Temperature Dependence (ITD) aware clock skew scheduling framework. Specifically, we demonstrate how our framework can assist dual-Vthassignment in preventing timing violations arising due to ITD effect. We formulate the ITD aware synthesis problem and prove that it is NP-Hard. Then, we propose an algorithm for synergistic temperature aware clock skew scheduling and dual-Vthassignment. Experiments on ISCAS89 benchmarks reveal that several circuits synthesized by the traditional high-temperature corner based flow with a commercial tool exhibit timing violations in the low temperature range while all circuits generated using our methodology for the same timing constraints have guaranteed timing. Jieyi Long, Seda Ogrenci Memik |
DATE | 1 |
| 2010 | Optimization of an on-chip active cooling system based on thin-film thermoelectric coolersabstractIn this paper, we explore the design and optimization of an on-chip active cooling system based on thin-film thermoelectric coolers (TEC). We start our investigation by establishing the compact thermal model for the chip package with integrated thin-film TEC devices. We observe that deploying an excessive number of TEC devices and/or providing the TEC devices with an improper supply current might adversely result in the overheating of the chip, rendering the cooling system ineffective. A large amount of supply current could even cause the thermal runaway of the system. Motivated by this observation, we formulate the deployment of the integrated TEC devices and their supply current setting as a system-level design problem. We propose a greedy algorithm to determine the deployment of TEC devices and a convex programming based scheme for setting the supply current levels. Leveraging the theory of inverse-positive matrix, we provide an optimality condition for the current setting algorithm. We have tested our algorithms on various benchmarks. We observe that our algorithms are able to determine the proper deployment and supply current level of the TEC devices which reduces the temperatures of the hot spots by as much as 7.5 °C compared to the cases without integrated TEC devices. Jieyi Long, Seda Ogrenci Memik, Matthew Grayson |
DATE | 1 |
| 2010 | SACTA: A Self-Adjusting Clock Tree Architecture for Adapting to Thermal-Induced Delay VariationabstractAggressive technology scaling down and low-power design techniques lead to uneven distributed power density, which translates into heat flow in the chips, causing significant temperature variations in both spatial and temporal terms. In order to mitigate the negative impacts of temperature variations on circuit timing, we propose SACTA, a self-adjusting clock tree architecture, which performs temperature-dependent dynamic clock skew scheduling to prevent timing violations in a pipelined circuit. The dynamic and adaptive features of SACTA are enabled by our proposed automatic temperature-adjustable skew buffers and temperature-insensitive skew buffers. These special delay elements are carefully tuned to ensure resilience of the entire circuit against temperature variation. To determine their configurations, we proposed an efficient and general clock tree design and optimization framework. Furthermore, we show that SACTA is applicable across a wide spectrum of circuits, including multi-${V}_{\rm dd}/{V}_{\rm th}$designs. Experimental results show that a pipeline supported by SACTA is able to prevent thermal-induced timing violations within a significantly larger range of operating temperatures (on average, the violation-free range can be enhanced by over 15$^{\circ}\hbox {C}$). Jieyi Long, Ja Chun Ku, Seda Ogrenci Memik, Yehea I. Ismail |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2008 | Automated design of self-adjusting pipelinesabstractWe propose a self-adjusting pipeline structure to enhance chip performance and robustness considering the effects of process variations. We achieve this by introducing delay sensors to monitor internal timing violations within a pipeline stage and variable clock skew buffers to adjust the timing of the pipeline stage based on the feedback from the delay sensors. Furthermore, we formulate the delay sensor insertion and variable clock skew configuration problem as a stochastic mixed-integer programming problem and propose a simulated-annealing based algorithm to solve it. A comparison between the designs with and without the self-adjusting enhancement reveals that, we are able to improve the average performance of a batch of chips by 9.5%. Jieyi Long, Seda Ogrenci Memik |
DAC | 1 |
| 2008 | An O(nlogn) edge-based algorithm for obstacle-avoiding rectilinear steiner tree constructionabstractObstacle-avoiding Steiner tree construction is a fundamental problem in VLSI physical design. In this paper, we provide a new approach for rectilinear Steiner tree construction in the presence of obstacles. We propose a novel algorithm, which generates sparse obstacle-avoiding spanning graphs efficiently. We design a fast algorithm for the minimum terminal spanning tree construction, which is the bottleneck step of several existing approaches in terms of running time. We adopt an edge-based heuristic, which enables us to perform both local and global refinement, leading to Steiner trees with small lengths. The time complexity of our algorithm is O(nlogn). Hence, our technique is the most efficient one to the best of our knowledge. Experimental results on various benchmarks show that our algorithm achieves 25.8 times speedup on average, while the average length of the resulting obstacle-avoiding rectilinear Steiner trees is only 1.58% larger than the best existing solution Jieyi Long, Hai Zhou 0001, Seda Ogrenci Memik |
ISPD | 1 |
| 2008 | Thermal monitoring mechanisms for chip multiprocessorsabstractWith large-scale integration and increasing power densities, thermal management has become an important tool to maintain performance and reliability in modern process technologies. In the core of dynamic thermal management schemes lies accurate reading of on-die temperatures. Therefore, careful planning and embedding of thermal monitoring mechanisms into high-performance systems becomes crucial. In this paper, we propose three techniques to create sensor infrastructures for monitoring the maximum temperature on a multicore system. Initially, we extend a nonuniform sensor placement methodology proposed in the literature to handle chip multiprocessors (CMPs) and show its limitations. We then analyze a grid-based approach where the sensors are placed on a static grid covering each core and show that the sensor readings can differ from the actual maximum core temperature by as much as 12.6°C when using 16 sensors per core. Also, as large as 10.6% of the thermal emergencies are not captured using the same number of sensors. Based on this observation, we first develop an interpolation scheme, which estimates the maximum core temperature through interpolation of the readings collected at the static grid points. We show that the interpolation scheme improves the measurement accuracy and emergency coverage compared to grid-based placement when using the same number of sensors. Second, we present a dynamic scheme where only a subset of the sensor readings is collected to predict the maximum temperature of each core. Our results indicate that, we can reduce the number of active sensors by as much as 50%, while maintaining similar measurement accuracy and emergency coverage compared to the case where the entire sensor set on the grid is sampled at all times. Jieyi Long, Seda Ogrenci Memik, Gokhan Memik, Rajarshi Mukherjee |
ACM Trans. Archit. Code Optim. | 1 |
| 2008 | EBOARST: An Efficient Edge-Based Obstacle-Avoiding Rectilinear Steiner Tree Construction AlgorithmabstractObstacle-avoiding Steiner routing has arisen as a fundamental problem in the physical design of modern VLSI chips. In this paper, we present EBOARST, an efficient four-step algorithm to construct a rectilinear obstacle-avoiding Steiner tree for a given set of pins and a given set of rectilinear obstacles. Our contributions are fourfold. First, we propose a novel algorithm, which generates sparse obstacle-avoiding spanning graphs efficiently. Second, we present a fast algorithm for the minimum terminal spanning tree construction step, which dominates the running time of several existing approaches. Third, we present an edge-based heuristic, which enables us to perform both local and global refinements, leading to Steiner trees with small lengths. Finally, we discuss a refinement technique called segment translation to further enhance the quality of the trees. The time complexity of our algorithm isO(nlogn). Experimental results on various benchmarks show that our algorithm achieves 16.56 times speedup on average, while the average length of the resulting obstacle-avoiding rectilinear Steiner trees is only 0.46% larger than the best existing solution. Jieyi Long, Hai Zhou 0001, Seda Ogrenci Memik |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2008 | Optimizing Thermal Sensor Allocation for MicroprocessorsabstractHigh-performance microprocessor families employ dynamic-thermal-management techniques to cope with the increasing thermal stress resulting from peaking power densities. These techniques operate on feedback generated from on-die thermal sensors. The allocation and the placement of thermal-sensing elements directly impact the effectiveness of the dynamic management mechanisms. In this paper, we propose systematic techniques for determining the optimal locations for thermal sensors to provide high-fidelity thermal monitoring of a complex microprocessor system. Our strategies can be divided into two main categories: uniform sensor allocation and nonuniform sensor allocation. In the uniform approach, the sensors are placed on a regular grid. The nonuniform allocation identifies an optimal physical location for each sensor such that the sensor's attraction toward steep thermal gradients is maximized, which can result in uneven concentrations of sensors on different locations of the chip. We also present a hybrid algorithm that shows the tradeoffs associated with number of sensors and expected accuracy. Our experimental results show that our uniform approach using interpolation can detect the chip temperature with a maximum error of 5.47degC and an average maximum error of 1.05degC . On the other hand, our nonuniform strategy is able to create a sensor distribution for a given microprocessor architecture, providing thermal measurements with a maximum error of 3.18degC and an average maximum error of 1.63degC across a wide set of applications. Seda Ogrenci Memik, Rajarshi Mukherjee, Min Ni, Jieyi Long |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2007 | A self-adjusting clock tree architecture to cope with temperature variationsabstractEnsuring resilience against environmental variations is becoming one of the great challenges of chip design. In this paper, we propose a self adjusting clock tree architecture, SACTA, to improve chip performance and reliability in the presence of on-chip temperature variations. SACTA performs temperature dependent dynamic clock skew scheduling to prevent timing violations in a pipelined circuit. We present an automatic temperature adjustable skew buffer design, which enables the adaptive feature of SACTA. Furthermore, we propose an efficient and general optimization framework to determine the configuration of these special delay elements. Experimental results show that a pipeline supported by SACTA is able to prevent thermal induced timing violations within a significantly larger range of operating temperatures (enhancing the violation-free range by as much as 45°C). Jieyi Long, Ja Chun Ku, Seda Ogrenci Memik, Yehea I. Ismail |
ICCAD | 1 |
| 2005 | An improved test access mechanism structure and optimization technique in system-on-chipabstractThis paper presents a new test access mechanism (TAM) architecture and optimization method based on an improved flexible-width test bus. The method is first to set up the test time lower bound that is not depends on TAM architecture, then to construct a bus assignment that makes test time up to the lower bound. We present experimental results on our improved flexible-width test buses for four benchmark SOCs. Experiment results in a significant reduction of the test time, and is better than the proposed traditional methods in test time. Jianhua Feng, Jieyi Long, Wenhua Xu, Hongfei Ye |
ASP-DAC | 2 |