Xin Zhan

dblp:41/3368 · DBLP profile ↗
← Back
29ranked-venue papers
9as first author
16since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 12 · 12 since 2021Systems, architecture and hardware · 11 · 7 first-authorGraphics, computer vision, multimedia, augmented reality and games · 10 · 1 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 1 since 2021Software engineering, systems software and programming languages · 1
YearPublicationVenuePosition
2026 Cycle-CFM: An unsupervised framework for robust multimodal anomaly detection in industrial settings
Yikang Shi, Xin Zhan, Zhongqiang Wu, Wenming Zhang
Expert Syst. Appl.2
2025 HAFUNet: A Hierarchical Attention Fusion Network for Monocular Depth Estimation Integrating Event and Frame Data
abstract
In robotics and autonomous driving, accurate depth estimation is vital yet challenging under dynamic scenes and extreme lighting. Conventional frame-based cameras offer rich context but suffer from motion blur and limited dynamic range, while event cameras provide high temporal resolution and dynamic range but lack global scene structure. Therefore, recent studies explore frame-event fusion depth estimation methods to leverage these two complementary modalities to achieve robust performance. However, due to the mismatch in temporal and spatial resolution, there is an inherent contradiction between high spatial resolution frames captured at sparse temporal intervals and event streams characterized by spatial sparsity but high temporal resolution, rendering cross-modal feature fusion ineffective. Moreover, the limited availability of frame-event depth datasets further undermines the model's generalization capability across different scenes. To address the above challenges, we propose HAFUNet, a Hierarchical Attention Fusion Network for depth estimation via frame-event fusion. Our method contains: (1) a pre-trained Dual-Stream Encoder (DSEer) to extract complementary features from frame and event inputs; (2) a Cross-modal Feature Interaction Module (CFIM) that aligns and fuses spatial-channel features across modalities; and (3) a Hierarchical Attention Decoder (HADer) that progressively refines depth predictions via attention-guided convolution. Experiments on synthetic and real-world datasets show that HAFUNet surpasses existing methods in depth accuracy and robustness. These results demonstrate the strength of our fusion strategy in diverse environments. Code is available at https://github.com/SiYZhangwh/HAFUNet.
Xiaoping Wang 0001, Jiang Li 0004, Weibin Feng, Xin Zhan, Hongzhi Huang
ACM Multimedia5
2025 Event denoising for dynamic vision sensor using residual graph neural network with density-based spatial clustering
Weibin Feng, Xiaoping Wang 0001, Xin Zhan, Hongzhi Huang
Neurocomputing3
2025 Low-Resolution Self-Attention for Semantic Segmentation
abstract
Semantic segmentation tasks naturally require high-resolution information for pixel-wise segmentation and global context information for class prediction. While existing vision transformers demonstrate promising performance, they often utilize high-resolution context modeling, resulting in a computational bottleneck. In this work, we challenge conventional wisdom and introduce the Low-Resolution Self-Attention (LRSA) mechanism to capture global context at a significantly reduced computational cost, i.e., FLOPs. Our approach involves computing self-attention in a fixed low-resolution space, regardless of the input image's resolution, with additional $\text{3}\times \text{3}$3×3 depth-wise convolutions to capture fine details in the high-resolution space. We demonstrate the effectiveness of our LRSA approach by building the LRFormer, a vision transformer with an encoder-decoder structure. Extensive experiments on the ADE20 K, COCO-Stuff, and CityScapes datasets demonstrate that LRFormer outperforms state-of-the-art models.
Yu-Huan Wu, Shi-Chen Zhang, Yun Liu 0011, Le Zhang 0001, Xin Zhan, Daquan Zhou, Jiashi Feng, Ming-Ming Cheng, Liangli Zhen
IEEE Trans. Pattern Anal. Mach. Intell.5
2025 Producing Considerate Responses: Progressive Staged Training for Emotional Support Conversation
abstract
Emotional support conversation (ESC) aims to alleviate the negative emotions of help-seekers by providing psychological assistance. Existing approaches typically overlook the abundant annotations contained in the ESC dataset, such as the situation descriptions and feedback scores of seekers, which limits their performance. In an effort to utilize the annotation information to enhance the emotional support ability of the backbone, we propose a three-stage training method called BlenderBot-ThTra for ESC systems. The proposed BlenderBot-ThTra involves the following three training processes: fine-tuning with supplemental feedback utterance, fine-tuning with auxiliary situation restoration, and calibration with the helpfulness estimation. The first stage aims to intensify the backbone's perception of conversational context, the second stage propels the backbone into excavating the causes of the emotional distress faced by the seeker. In the third stage, we leverage a Bayesian method based on the seeker's feedback scores to train a helpfulness evaluation model, then exploit a contrastive learning method to calibrate the ESC backbone. We conduct experiments on the standard multiturn ESC dataset, and the results demonstrate that BlenderBot-ThTra has a significant advantage in generating more supportive and adaptive responses.
Guoqing Lv, Jiang Li 0004, Xiaoping Wang 0001, Xin Zhan, Zhigang Zeng
IEEE Trans. Comput. Soc. Syst.4
2025 Selection and guidance: high-dimensional identity consistency preservation for face inpainting
Xin Zhan, Wenming Zhang
Vis. Comput.2
2024 Construction of Paired Knowledge Graph - Text Datasets Informed by Cyclic Evaluation
abstract
Datasets that pair Knowledge Graphs (KG) and text together (KG-T) can be used to train forward and reverse neural models that generate text from KG and vice versa. However models trained on datasets where KG and text pairs are not equivalent can suffer from more hallucination and poorer recall. In this paper, we verify this empirically by generating datasets with different levels of noise and find that noisier datasets do indeed lead to more hallucination. We argue that the ability of forward and reverse models trained on a dataset to cyclically regenerate source KG or text is a proxy for the equivalence between the KG and the text in the dataset. Using cyclic evaluation we find that manually created WebNLG is much better than automatically created TeKGen and T-REx. Informed by these observations, we construct a new, improved dataset called LAGRANGE using heuristics meant to improve equivalence between KG and text and show the impact of each of the heuristics on cyclic evaluation. We also construct two synthetic datasets using large language models (LLMs), and observe that these are conducive to models that perform significantly well on cyclic generation of text, but less so on cyclic generation of KGs, probably because of a lack of a consistent underlying ontology.
Ali Mousavi 0003, Xin Zhan, He Bai 0002, Peng Shi 0010, Theodoros Rekatsinas, Benjamin Han, Yunyao Li 0001, Jeffrey Pound, Joshua M. Susskind, Natalie Schluter, Ihab F. Ilyas, Navdeep Jaitly
LREC/COLING2
2024 Quality assessment of identity inpainting based on multidimensional discrimination
Xin Zhan, Wenming Zhang
Multim. Syst.2
2023 PUPS: Point Cloud Unified Panoptic Segmentation
abstract
Point cloud panoptic segmentation is a challenging task that seeks a holistic solution for both semantic and instance segmentation to predict groupings of coherent points. Previous approaches treat semantic and instance segmentation as surrogate tasks, and they either use clustering methods or bounding boxes to gather instance groupings with costly computation and hand-craft designs in the instance segmentation task. In this paper, we propose a simple but effective point cloud unified panoptic segmentation (PUPS) framework, which use a set of point-level classifiers to directly predict semantic and instance groupings in an end-to-end manner. To realize PUPS, we introduce bipartite matching to our training pipeline so that our classifiers are able to exclusively predict groupings of instances, getting rid of hand-crafted designs, e.g. anchors and Non-Maximum Suppression (NMS). In order to achieve better grouping results, we utilize a transformer decoder to iteratively refine the point classifiers and develop a context-aware CutMix augmentation to overcome the class imbalance problem. As a result, PUPS achieves 1st place on the leader board of SemanticKITTI panoptic segmentation task and state-of-the-art results on nuScenes.
Shihao Su, Jianyun Xu, Zhenwei Miao, Xin Zhan, Dayang Hao
AAAI5
2023 Research on Defect Detection Method of Nonwoven Fabric Mask Based on Machine Vision
abstract
During the production, transportation and storage of nonwoven fabric mask, there are many damages caused by human or nonhuman factors. Therefore, checking the defects of nonwoven fabric mask in a timely manner to ensure the reliability and integrity, which plays a positive role in the safe use of nonwoven fabric mask. At present, the wide application of machine vision technology provides a technical mean for the defect detection of nonwoven fabric mask. On the basis of the pre-treatment of the defect images, it can effectively simulate the contour fluctuation grading and gray value change of the defect images, which is helpful to realize the segmentation, classification and recognition of nonwoven fabric mask defect features. First, in order to accurately obtain the image information of the nonwoven fabric mask, the binocular vision calibration method of the defect detection system is discussed. On this basis, the defect detection mechanism of the nonwoven fabric mask is analyzed, and the model of image processing based on spatial domain and Hough transform is established, respectively. The original image of the nonwoven fabric mask is processed by region processing and edge extraction. Second, the defect detection algorithm of nonwoven fabric mask is established and the detection process is designed. Finally, a fast defect detection system for nonwoven fabric mask is designed, and the effectiveness of the detection method for nonwoven fabric mask is analyzed with an example. The results show that this detection method has positive engineering significance for improving the detection efficiency of defects in nonwoven fabric mask.
Jingde Huang, Zhangyu Huang, Xin Zhan
Int. J. Pattern Recognit. Artif. Intell.3
2023 P2T: Pyramid Pooling Transformer for Scene Understanding
abstract
Recently, the vision transformer has achieved great success by pushing the state-of-the-art of various vision tasks. One of the most challenging problems in the vision transformer is that the large sequence length of image tokens leads to high computational cost (quadratic complexity). A popular solution to this problem is to use a single pooling operation to reduce the sequence length. This paper considers how to improve existing vision transformers, where the pooled feature extracted by a single pooling operation seems less powerful. To this end, we note that pyramid pooling has been demonstrated to be effective in various vision tasks owing to its powerful ability in context abstraction. However, pyramid pooling has not been explored in backbone network design. To bridge this gap, we propose to adapt pyramid pooling to Multi-Head Self-Attention (MHSA) in the vision transformer, simultaneously reducing the sequence length and capturing powerful contextual features. Plugged with our pooling-based MHSA, we build a universal vision transformer backbone, dubbed Pyramid Pooling Transformer (P2T). Extensive experiments demonstrate that, when applied P2T as the backbone network, it shows substantial superiority in various vision tasks such as image classification, semantic segmentation, object detection, and instance segmentation, compared to previous CNN- and transformer-based networks. The code will be released at https://github.com/yuhuan-wu/P2T.
Yu-Huan Wu, Yun Liu 0011, Xin Zhan, Ming-Ming Cheng
IEEE Trans. Pattern Anal. Mach. Intell.3
2022 BE-STI: Spatial-Temporal Integrated Network for Class-agnostic Motion Prediction with Bidirectional Enhancement
abstract
Determining the motion behavior of inexhaustible categories of traffic participants is critical for autonomous driving. In recent years, there has been a rising concern in performing class-agnostic motion prediction directly from the captured sensor data, like LiDAR point clouds or the combination of point clouds and images. Current motion prediction frameworks tend to perform joint semantic segmentation and motion prediction and face the trade-off between the performance of these two tasks. In this paper, we propose a novel Spatial-Temporal Integrated network with Bidirectional Enhancement, BE-STI, to improve the temporal motion prediction performance by spatial semantic features, which points out an efficient way to combine semantic segmentation and motion prediction. Specifically, we propose to enhance the spatial features of each individual point cloud with the similarity among temporal neighboring frames and enhance the global temporal features with the spatial difference among non-adjacent frames in a coarse-to-fine fashion. Extensive experiments on nuScenes and Waymo Open Dataset show that our proposed framework outperforms all state-of-the-art LiDAR-based and RGB+LiDAR-based methods with remarkable margins by using only point clouds as input.11The code will be released at https://github.com/be-sti/be-sti.
Yunlong Wang 0009, Hongyu Pan, Yu-Huan Wu, Xin Zhan, Kun Jiang 0002, Diange Yang
CVPR5
2022 LIFT: Learning 4D LiDAR Image Fusion Transformer for 3D Object Detection
abstract
LiDAR and camera are two common sensors to collect data in time for 3D object detection under the autonomous driving context. Though the complementary information across sensors and time has great potential of benefiting 3D perception, taking full advantage of sequential cross-sensor data still remains challenging. In this paper, we propose a novel LiDAR Image Fusion Transformer (LIFT) to model the mutual interaction relationship of cross-sensor data over time. LIFT learns to align the input 4D sequential cross-sensor data to achieve multi-frame multi-modal information aggregation. To alleviate computational load, we project both point clouds and images into the bird-eye-view maps to compute sparse grid-wise self-attention. LIFT also benefits from a cross-sensor and cross-time data augmentation scheme. We evaluate the proposed approach on the challenging nuScenes and Waymo datasets, where our LIFT performs well over the state-of-the-art and strong baselines.
Yihan Zeng, Chunwei Wang, Zhenwei Miao, Xin Zhan, Dayang Hao, Chao Ma 0004
CVPR6
2022 SP-Net: Slowly Progressing Dynamic Inference Networks
Wenhu Zhang, Shihao Su, Hui Wang 0107, Zhenwei Miao, Xin Zhan, Xi Li 0001
ECCV (11)6
2022 INT: Towards Infinite-Frames 3D Detection with an Efficient Framework
Jianyun Xu, Zhenwei Miao, Hongyu Pan, Peihan Hao, Zhengyang Sun, Xin Zhan
ECCV (9)10
2021 PVGNet: A Bottom-Up One-Stage 3D Object Detector With Integrated Multi-Level Features
abstract
Quantization-based methods are widely used in LiDAR points 3D object detection for its efficiency in extracting context information. Unlike image where the context information is distributed evenly over the object, most LiDAR points are distributed along the object boundary, which means the boundary features are more critical in LiDAR points 3D detection. However, quantization inevitably introduces ambiguity during both the training and inference stages. To alleviate this problem, we propose a one-stage and voting-based 3D detector, named Point-Voxel-Grid Network (PVGNet). In particular, PVGNet extracts point, voxel and grid-level features in a unified backbone architecture and produces point-wise fusion features. It segments Li-DAR points into foreground and background, predicts a 3D bounding box for each foreground point, and performs group voting to get the final detection results. Moreover, we observe that instance-level point imbalance due to occlusion and observation distance also degrades the detection performance. A novel instance-aware focal loss is proposed to alleviate this problem and further improve the detection ability. We conduct experiments on the KITTI and Waymo datasets. Our proposed PVGNet outperforms previous state-of-the-art methods and ranks at the top of KITTI 3D/BEV detection leaderboards.
Zhenwei Miao, Jikai Chen, Hongyu Pan, Ruiwen Zhang, Peihan Hao, Xin Zhan
CVPR9
2019 Taming the Stability-Constrained Performance Optimization Challenge of Distributed On-Chip Voltage Regulation
abstract
Distributed on-chip voltage regulation is promising for addressing many IC power delivery challenges. However, complex interactions between active regulators and the surrounding parasitic passive RLC network cause stability concern. The recently developed hybrid stability technique provides a unique opportunity for coping with stability of distributed on-chip regulation and enabling efficient localized system design. However, the inherent conservativeness of the hybrid stability theorem (HST) leads to large pessimism in stability evaluation and hence causes overdesign. In this paper, the above challenge is addressed by extending the HST with an optimal frequency-dependent system partitioning technique which can significantly reduce the amount of pessimism in stability analysis. To put the proposed approach on a firm theoretical footing, we prove that the partitioning technique removes the conservativeness without altering the physical system and key theoretical properties of the partitioned blocks are maintained under certain constraints. Upon this, an efficient stability-ensuring power delivery design methodology using an automated design flow is developed to significantly improve power delivery performance. Within a large design space, the proposed approach ensures stability and improves system performance by up to 53%, measured by a figure of merit (FOM), when compared to the classical phase margin design approach, which provides no guarantee of stability. Furthermore, on average our approach boosts the FOM by 113% while consuming 11% less power compared to a reference hybrid stability approach.
Xin Zhan, Peng Li 0001, Edgar Sánchez-Sinencio
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2019 Power Management for Multicore Processors via Heterogeneous Voltage Regulation and Machine Learning Enabled Adaptation
abstract
This work is based on the vision that the ultimate power integrity and efficiency may be best achieved via a heterogeneous chain of voltage processing starting from onboard switching voltage regulators (VRs), to on-chip switching VRs, and finally to networks of distributed on-chip linear VRs. As such, we propose a heterogeneous voltage regulation (HVR) architecture encompassing regulators with complimentary characteristics in response time, size, and efficiency. By exploring the rich heterogeneity and tunability in HVR, we develop systematic workload-aware power management policies to adapt heterogeneous VRs with respect to workload change at multiple temporal scales to significantly improve system power efficiency while providing a guarantee for power integrity. The proposed techniques are further supported by hardware-accelerated machine learning (ML) prediction of nonuniform spatial workload distributions for more accurate HVR adaptation at fine time granularity. Our evaluations based on the PARSEC benchmark suite show that the proposed adaptive three-stage HVR reduces the total system energy dissipation by up to 23.9% and 15.7% on average compared with the conventional static two-stage voltage regulation using off-chip and on-chip switching VRs. Compared with the three-stage static HVR, our runtime control reduces system energy by up to 17.9% and 12.2% on average. Furthermore, the proposed ML prediction offers up to 4.1% reduction of system energy.
Xin Zhan, Edgar Sánchez-Sinencio, Peng Li 0001
IEEE Trans. Very Large Scale Integr. Syst.1
2018 QoR-aware power capping for approximate big data processing
abstract
To limit the peak power consumption of a cluster, a centralized power capping system typically assigns power caps to the individual servers, which are then enforced using local capping controllers. Consequently, the performance and throughput of the servers are affected, and the runtime of jobs is extended as a result. We observe that servers in big data processing clusters often execute big data applications that have different tolerance for approximate results. To mitigate the impact of power capping, we propose a new power-Capping aware resource manager for Approximate Big data processing (CAB) that takes into consideration the minimum Quality-of-Result (QoR) of the jobs. We use industry-standard feedback power capping controllers to enforce a power cap quickly, while, simultaneously modifying the resource allocations to various jobs based on their progress rate, target minimum QoR, and the power cap such that the impact of capping on runtime is minimized. Based on the applied cap and the progress rates of jobs, CAB dynamically allocates the computing resources (i.e., number of cores and memory) to the jobs to mitigate the impact of capping on the finish time. We implement CAB in Hadoop-2.7.3 and evaluate its improvement over other methods on a state-of-the-art 28-core Xeon server. We demonstrate that CAB minimizes the impact of power capping on runtime by up to 39.4% while meeting the minimum QoR constraints.
Seyed Morteza Nabavinejad, Xin Zhan, Maziar Goudarzi, Sherief Reda
DATE2
2018 Design Space Exploration of Distributed On-Chip Voltage Regulation Under Stability Constraint
Xin Zhan, Joseph Riad, Peng Li 0001, Edgar Sánchez-Sinencio
IEEE Trans. Very Large Scale Integr. Syst.1
2017 Fast Decentralized Power Capping for Server Clusters
abstract
Power capping is a mechanism to ensure that the power consumption of clusters does not exceed the provisioned resources. A fast power capping method allows for a safe over-subscription of the rated power distribution devices, provides equipment protection, and enables large clusters to participate in demand-response programs. However, current methods have a slow response time with a large actuation latency when applied across a large number of servers as they rely on hierarchical management systems. We propose a fast decentralized power capping (DPC) technique that reduces the actuation latency by localizing power management at each server. The DPC method is based on a maximum throughput optimization formulation that takes into account the workloads priorities as well as the capacity of circuit breakers. Therefore, DPC significantly improves the cluster performance compared to alternative heuristics. We implement the proposed decentralized power management scheme on a real computing cluster. Compared to state-of-the-art hierarchical methods, DPC reduces the actuation latency by 72% up to 86% depending on the cluster size. In addition, DPC improves the system throughput performance by 16%, while using only 0.02% of the available network bandwidth. We describe how to minimize the overhead of each local DPC agent to a negligible amount. We also quantify the traffic and fault resilience of our decentralized power capping approach.
Masoud Badiei, Xin Zhan, Na Li 0002, Sherief Reda
HPCA3
2016 DiBA: Distributed Power Budget Allocation for Large-Scale Computing Clusters
abstract
Power management has become a central issue inlarge-scale computing clusters where a considerable amount ofenergy is consumed and a large operational cost is incurredannually. Traditional power management techniques have a centralizeddesign that creates challenges for scalability of computingclusters. In this work, we develop a framework for distributedpower budget allocation that maximizes the utility of computingnodes subject to a total power budget constraint. To eliminate the role of central coordinator in the primaldualtechnique, we propose a distributed power budget allocationalgorithm (DiBA) which maximizes the combined performanceof a cluster subject to a power budget constraint in a distributedfashion. Specifically, DiBA is a consensus-based algorithm inwhich each server determines its optimal power consumptionlocally by communicating its state with neighbors (connectednodes) in a cluster. We characterize a synchronous primal-dualtechnique to obtain a benchmark for comparison with thedistributed algorithm that we propose. We demonstrate numericallythat DiBA is a scalable algorithm that outperforms theconventional primal-dual method on large scale clusters in termsof convergence time. Further, DiBA eliminates the communicationbottleneck in the primal-dual method. We thoroughly evaluatethe characteristics of DiBA through simulations of large-scaleclusters. Furthermore, we provide results from a proof-of-conceptimplementation on a real experimental cluster.
Masoud Badiei, Xin Zhan, Sherief Reda, Na Li 0002
CCGrid2
2016 Creating Soft Heterogeneity in Clusters Through Firmware Re-configuration
abstract
Customizing server hardware to adapt to its workload has the potential to improve both runtime and energy efficiency. In a cluster that caters to diverse workloads, employing servers with customized hardware components leads to heterogeneity, which is not scalable. In this paper, we seek to create soft heterogeneity from existing servers with homogenous hardware components through customizing the firmware configuration. We demonstrate that firmware configurations have a large impact on runtime, power, and energy efficiency of workloads. Since finding the firmware configuration that minimizes runtime and/or energy efficiency grows exponentially as a function of the number of firmware settings, we propose a methodology called FXplore that helps complete the exploration with a quadratic time complexity. Furthermore, FXplore enables system administrators to manage the degree of the heterogeneity by deriving firmware configurations for sub-clusters that can cater to multiple workloads with similar characteristics. Thus, during online operation, incoming workloads to the cluster can be mapped to appropriate sub-clusters with pre-configured firmware settings. FXplore also finds the best firmware settings in case of co-runners on the same server. We validate our methodology on a fully-instrumented cluster under a large range of parallel workloads that are representative of both high-performance compute clusters and datacenters. Compared to enabling all firmware options, our method improves average runtime and energy consumption by 11% and 15%, respectively.
Xin Zhan, Mohammed Shoaib, Sherief Reda
CCGrid1
2016 Distributed on-chip regulation: theoretical stability foundation, over-design reduction and performance optimization
abstract
While distributed on-chip voltage regulation offers an appealing solution to power delivery, designing power delivery networks (PDNs) with distributed on-chip voltage regulators with guaranteed stability is challenging because of the complex interactions between active regulators and the bulky passive network. The recently developed hybrid stability theory provides an efficient stability checking and design approach, giving rise to highly desirable localized design of PDNs. However, the inherent conservativeness of the hybrid stability criteria can lead to pessimism in stability evaluation and hence large over-design. We address this challenge by proposing an optimal frequency-dependent system partitioning technique to significantly reduce the amount of pessimism in stability analysis. With theoretical rigor, we show how to partition a PDN system by employing optimal frequency-dependent impedance splitting between the passive network and voltage regulators while maintaining the desired theoretical properties of the partitioned system blocks upon which the hybrid stability principle is anchored. We demonstrate a new stability-ensuring PDN design approach with the proposed over-design reduction technique using an automated optimization flow which significantly boosts regulation performance and power efficiency.
Xin Zhan, Peng Li 0001, Edgar Sánchez-Sinencio
DAC1
2015 Power Budgeting Techniques for Data Centers
abstract
The development of cloud computing and data science result in rapid increases of number and scale of data centers. Because of cost and sustainability concerns, energy efficiency has been a major goal for data center architects. Focusing on reducing the cooling power and making full use of available computing power, power budgeting is an increasingly important requirement for data center operations. In this paper, we present a framework of power budgeting, considering both computing power and cooling power, in data centers to maximize the system normalized performance (SNP) of the entire center under a total power budget. Maximizing the SNP for a given power budget is equivalent to maximizing the energy efficiency. We propose a method to partition the total power budget among the cooling and computing infrastructure in a self-consistent way, where the cooling power is sufficient to extract the heat of the computing power. Intertwinedly, we devise an optimal computing power budgeting technique based on dynamic programming algorithm to determine the optimal power caps for the individual servers such that the available power could be efficiently translated to performance improvements. The optimal computing budgeting technique leverages a proposed online throughput predictor based on performance counter measurements to estimate the change in throughput of heterogeneous workloads as a function of allocated server power caps. We demonstrate that our proposed power budgeting method outperforms previous methods by 3-4 percent in terms of SNP using our data center simulation environment. While maintaining the improvement of SNP, our method improve fairness at best by 57 percent. We also evaluate the performance of our method in power saving scenario and dynamic power budgeting case.
Xin Zhan, Sherief Reda
IEEE Trans. Computers1
2014 Thermal-aware layout planning for heterogeneous datacenters
abstract
Cooling power represents a significant portion of total power consumption in datacenters. Heterogeneous datacenters deploy clusters of servers with different hardware configurations, each offering its own performance and power characteristics. We observe that heterogeneous datacenters offer a unique opportunity to reduce cooling power through appropriate planning. In this paper we formulate the problem of rack layout for planning of heterogeneous datacenters, where the goal is to identify the best locations of the server racks with different hardware capabilities to improve the supply temperatures of the CRAC units and the total cooling power. We provide optimal solutions that take into account the impact of varying utilizations of datacenters and job scheduling methods. Using state-of-the-art thermal modeling tools, we prove that our methods lead to datacenter layouts with significant improvements in cooling power reduction, between 15.5%-38.5% based on the datacenter utilizations and an average of 28.3% without any negative side effects.
Xin Zhan, Sherief Reda
ISLPED2
2013 Techniques for energy-efficient power budgeting in data centers
abstract
We propose techniques for power budgeting in data centers, where a large power budget is allocated among the servers and the cooling units such that the aggregate performance of the entire center is maximized. Maximizing the performance for a given power budget automatically maximizes the energy efficiency. We first propose a method to partition the total power budget among the cooling and computing units in a self-consistent way, where the cooling power is sufficient to extract the heat of the computing power. Given the computing power budget, we devise an optimal computing budgeting technique based on knapsack-solving algorithms to determine the power caps for the individual servers. The optimal computing budgeting technique leverages a proposed on-line throughput predictor based on performance counter measurements to estimate the change in throughput of heterogeneous workloads as a function of allocated server power caps. We set up a simulation environment for a data center, where we simulate the air flow and heat transfer within the center using computational fluid dynamic simulations to derive accurate cooling estimates. The power estimates for the servers are derived from measurements on a real server executing heterogeneous workload sets. Our budgeting method delivers good improvements over previous power budgeting techniques.
Xin Zhan, Sherief Reda
DAC1
2013 Remote sensing image compression based on double-sparsity dictionary learning and universal trellis coded quantization
abstract
In this paper, we propose a novel remote sensing image compression method based on double-sparsity dictionary learning and universal trellis coded quantization (UTCQ). Recent years have seen a growing interest in the study of natural image compression based on sparse representation and dictionary learning. We show that using the double-sparsity model to learn a dictionary gives much better compression results for remote sensing images, the texture of which is much richer than that of natural images. We also show that the compression performance is improved significantly when advanced quantization and entropy coding strategies are used for encoding the sparse representation coefficients. The proposed method outperforms the existing dictionary-based image coding algorithms. Additionally, our method results in better ratedistortion performance and structural similarity results than CCSDS and JPEG2000 standard.
Xin Zhan, Rong Zhang 0004, Anzhou Hu, Wenlong Hu
ICIP1
2013 SAR Image Compression Using Multiscale Dictionary Learning and Sparse Representation
abstract
In this letter, we focus on a new compression scheme for synthetic aperture radar (SAR) amplitude images. The last decade has seen a growing interest in the study of dictionary learning and sparse representation, which have been proved to perform well on natural image compression. Because of the special techniques of radar imaging, SAR images have some distinct properties when compared with natural images that can affect the design of a compression method. First, we introduce SAR properties, sparse representation, and dictionary learning theories. Second, we propose a novel SAR image compression scheme by using multiscale dictionaries. The experimental results carried out on amplitude SAR images reveal that, when compared with JPEG, JPEG2000, and a single-scale dictionary-based compression scheme, the proposed method is better for preserving the important features of SAR images with a competitive compression performance.
Xin Zhan, Chengfu Huo
IEEE Geosci. Remote. Sens. Lett.1