Haihang You

dblp:89/4494 · DBLP profile ↗
← Back
29ranked-venue papers
1as first author
20since 2021 · last 2026
0000-0003-1432-827XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 12 · 9 since 2021Artificial intelligence and machine learning · 6 · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 4 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Unary Positional System: Flexible Balance of Hardware Area and Performance
abstract
Modern computer architectures face challenges in balancing hardware overhead and performance. Binary computing, known for its compactness, requires hardware area that scales quadratically with precision, while unary computing, despite its simplicity, suffers from exponentially increasing computation time. This paper introduces Unary Positional System (UPS), a paradigm that combines spatial and temporal characteristics to address this trade-off. We develop UPS-based architectures to perform fundamental arithmetic operations, and apply it to GEMM and superconductor FFT processor. Experimental results show that UPS bridges the gap between binary and unary computing, offering a balanced solution with flexibility for further optimization.
Zeshi Liu, Zheng Weng, Ruijie Tan, Guangming Tang, Haihang You
DATE5
2025 See Further When Clear: Curriculum Consistency Model
abstract
Significant advances have been made in the sampling efficiency of diffusion and flow matching models, driven by Consistency Distillation (CD), which trains a student model to mimic the output of a teacher model at a later timestep. However, we found that the knowledge discrepancy between student and teacher varies significantly across different timesteps, leading to suboptimal performance in CD. To address this issue, we propose the Curriculum Consistency Model (CCM), which stabilizes and balances the knowledge discrepancy across timesteps. Specifically, we regard the distillation process at each timestep as a curriculum and introduce a metric based on the Peak Signal-to-Noise Ratio (PSNR) to quantify the knowledge discrepancy of this curriculum, then ensure that the curriculum maintains consistent knowledge discrepancy across different timesteps by having the teacher model iterate more steps when the noise intensity is low. Our method achieves competitive single-step sampling Fréchet Inception Distance (FID) scores of 1.64 on CIFAR-10 and 2.18 on ImageNet 64x64. Moreover, we have extended our method to large-scale text-to-image models and confirmed that it generalizes well to both diffusion models (Stable Diffusion XL) and flow matching models (Stable Diffusion 3). The generated samples demonstrate improved image-text alignment and semantic structure since CCM enlarges the distillation step at large timesteps and reduces the accumulated error.
Boxiao Liu, Yi Zhang 0108, Xingzhong Hou, Guanglu Song, Yu Liu 0015, Haihang You
CVPR7
2025 Maximizing Energy Efficiency in Spiking Neural Networks: A Dynamic Joint Pruning Framework
abstract
Spiking Neural Networks (SNNs) face increasing challenges for efficient deployment as architectures grow in complexity, necessitating network pruning to improve energy and computational efficiency. Existing pruning methods primarily focus on a single form of sparsity, overlooking the importance of joint pruning, which is critical for minimizing synaptic operations (SOPs) and enhancing energy efficiency. This paper presents a novel dynamic joint pruning framework that leverages both spatiotemporal spike sparsity and weight sparsity to minimize SOPs in SNN inference. Based on a comprehensive analysis of the SOPs model, we introduce an integrated solution that combines a multi-stage masking mechanism for fine-grained neuron firing threshold control, a temporal attention batch normalization (TABN) module with learnable time scaling factors, and a dynamic sparse strategy that adjusts importance coefficients based on real-time computational impact. Experimental results on CIFAR-10, CIFAR-100, and ImageNet validate the effectiveness of proposed framework. Our method achieves up to $126.38 \times$ compression ratio of SOPs on CIFAR-10 with minimal accuracy loss, establishing a new state-of-the-art in energy-efficient SNN pruning.
Zeshi Liu, Haihang You
DAC3
2025 Parallel Dynamic Partitioning for Datapath Combinational Equivalence Checking
abstract
Combinational Equivalence Checking (CEC) is a crucial technique in electronic design automation for verifying the functional equivalence of combinational circuits. Recently, combinational circuit design increasingly incorporates more complex arithmetic structures, commonly known as datapath circuits. However, existing state-of-the-art tools often exhibit subpar performance in solving datapath CEC problems. To further advance the exploration on datapath CEC process, this study introduces PDP-CEC (Parallel Dynamic Partitioning Combinational Equivalence Checking), a novel parallel CEC approach integrating circuit partitioning and dynamic task scheduling into the CEC process, enhancing the efficiency of CEC for datapath circuits. PDP-CEC introduces an innovative method for selecting critical nodes to split the search space of the CEC problem, facilitating the efficient generation of numerous independent subproblems. Meanwhile, a dynamic task scheduling strategy is implemented in PDP-CEC to ensure load balancing and prevent hard-to-solve subproblems from stalling the entire process. Compared to the most advanced tools such as ABC and HybridCEC, PDP-CEC significantly accelerates CEC process, achieving speedups ranging from $5.11 x$ to $125.27 x$, while effectively solving approximately three times more datapath CEC problems. With excellent scalability, PDP-CEC shows substantial improvements in combinational equivalence checking for datapath circuits, offering an efficient parallel approach to meet the demands of large-scale datapath CEC tasks.
Xindi Zhang 0001, Zite Jiang, Haihang You, Shaowei Cai 0001
DAC5
2025 Brain-Inspired Efficient Pruning: Exploiting Criticality in Spiking Neural Networks
abstract
ABSTRACT Spiking neural networks (SNNs) have gained significant attention due to their energy‐efficient and multiplication‐free characteristics. Despite these advantages, deploying large‐scale SNNs on edge hardware is challenging due to limited resource availability. Network pruning offers a viable approach to compress the network scale and reduce hardware resource requirements for model deployment. However, existing SNN pruning methods cause high pruning costs and performance loss because they lack efficiency in processing the sparse spike representation of SNNs. In this paper, inspired by the critical brain hypothesis in neuroscience and the high biological plausibility of SNNs, we explore and leverage criticality to facilitate efficient pruning in deep SNNs. We first explain criticality in SNNs from the perspective of maximizing feature information entropy. Second, we propose a low‐cost metric to assess neuron criticality in feature transmission and design a pruning‐regeneration method that incorporates this criticality into the pruning process. Experimental results demonstrate that our method achieves higher performance than the current state‐of‐the‐art (SOTA) method with up to 95.26% reduction in pruning cost. The criticality‐based regeneration process efficiently selects potential structures and facilitates consistent feature representation. Our code is available at https://github.com/MmPple/pruning‐criticality .
Zeshi Liu, Haihang You
Concurr. Comput. Pract. Exp.3
2025 RHINO: An Efficient Serverless Container System for Small-Scale HPC Applications
abstract
Serverless computing, characterized by its pay-as-you-go and auto-scaling features, offers a promising alternative for High Performance Computing (HPC) applications, as traditional HPC clusters often face long waiting times and resources over/under-provisioning. However, current serverless platforms struggle to support HPC applications due to restricted inter-function communication and high coupling runtime. To address these issues, we introduce RHINO, which offers end-to-end support for the development and deployment of serverless HPC. Using the Two-Step Adaptive Build strategy, the HPC code is packaged into lightweight, scalable functions. The Rhino Function Execution Model decouples HPC applications from the underlying infrastructures. The Auto-scaling Engine dynamically scales cloud resources and schedules tasks based on performance and cost requirements. We deploy RHINO on AWS Fargate and evaluate it on both benchmarks and real-world workloads. Experimental results show that, when compared to the traditional VM clusters, RHINO can achieve a performance improvement of 10%-30% for small-scale applications and more than 40% cost reduction.
Haihang You
IEEE Trans. Parallel Distributed Syst.3
2024 EasyDrag: Efficient Point-Based Manipulation on Diffusion Models
abstract
Generative models are gaining increasing popularity, and the demand for precisely generating images is on the rise. However, generating an image that perfectly aligns with users' expectations is extremely challenging. The shapes of objects, the poses of animals, the structures of landscapes, and more may not match the user's desires, and this applies to real images as well. This is where point-based image editing becomes essential. An excellent image editing method needs to meet the following criteria: user-friendly interaction, high performance, and good generalization capability. Due to the limitations of StyleGAN, DragGAN exhibits limited robustness across diverse scenarios, while DragDiffusion lacks user-friendliness due to the necessity of LoRA fine-tuning and masks. In this paper, we introduce a novel interactive point-based image editing framework, called EasyDrag, that leverages pretrained diffusion models to achieve high-quality editing outcomes and user-friendship. Extensive experimentation demonstrates that our approach surpasses DragDiffusion in terms of both image quality and editing precision for point-based image manipulation tasks. The code will be available on https://github.com/Ace-Pegasus/EasyDrag.
Xingzhong Hou, Boxiao Liu, Yi Zhang 0108, Jihao Liu, Yu Liu 0015, Haihang You
CVPR6
2024 IVE: Accelerating Enumeration-Based Subgraph Matching via Exploring Isolated Vertices
abstract
The performance of the enumeration-based sub-graph matching, which searches all isomorphic subgraphs in the data graph, is crucial to various applications. The upper bound of the complexity for the enumeration-based method is exponential to the number of query graph vertices, denoted as$n$. We propose a novel subgraph matching algorithm called the Isolated Vertices Exploration (IVE). The IVE leverages isolated vertices during the reordering and enumeration phases, thereby significantly accelerating the subgraph matching process. During the enumeration, the isolated vertices can be matched by using a quick bipartite graph matching algorithm. Consequently, the complexity of matching the remaining non-isolated vertices is exponential to the number of non-isolated vertices, denoted as$n^{\prime}$. For the reordering, we designed the Maximum Deleted Edges (MDE) to minimize$n^{\prime}$. MDE iteratively selects the query vertex with the maximum edges. According to the experimental results,$n^{\prime}$is less than$0.8n$for 99.8% of arbitrary graphs. Moreover, IVE outperforms the state-of-the-art algorithms in various scenarios with different sizes, sparsities and fields, achieving a performance speedup of up to 80.3x.
Zite Jiang, Shuai Zhang 0040, Xingzhong Hou, Mengting Yuan 0001, Haihang You
ICDE5
2024 A Novel Cloud-Native and Multi-Platform Parallelized SBAS-INSAR Algorithm
abstract
Facing massive Synthetic Aperture Radar (SAR) data, the challenge of achieving rapid and efficient data processing has garnered attention. Currently, most time-series InSAR processing solutions are designed to operate on a single computing platform, leading to drawbacks such as inflexible deployment and slow data transmission. Leveraging the open-source Kubernetes platform, we have employed cloud-native technology to deploy a cross-platform, multi-level parallelization algorithm across multiple nodes. The algorithm is designed based on modular concept, and its modules can successfully achieve cross-platform deployment by relying on both the Network File System (NFS) protocol and the Common Internet File System (CFIS) protocol. This overcomes the limitations of the previous data file storage system, which solely depended on the NFS protocol, leading to deployment on a single operating system. Deploying the algorithm on different operating systems, our results show that the speeds of Linux platform's algorithm parallel modules were improved 16% and 30%, respectively. The successful cross-platform operation of the algorithm enables the data processing workflow to assimilate the strengths of different operating platforms, enhancing data visualization capabilities and gaining support from diverse platform resources. This introduces a new approach for large-scale time-series InSAR processing.
Peichen Yu, Chao Wang 0004, Yixian Tang, Haihang You, Shaoyang Guan, Lichuan Zou, Hong Zhang 0001
IGARSS5
2024 Fast Subgraph Matching by Dynamic Graph Editing
abstract
Subgraph matching is a challenging NP-complete problem that involves finding identical subgraphs of a query graph$q$in a larger data graph$G$. It has numerous applications in diverse fields, including social and biological networks. However, existing subgraph matching algorithms assume that the graph structure is fixed, which limits their performance in solving more difficult matching cases. To address this issue, we propose a novel approach called Dynamic Graph Editing (DGE), which dynamically edits the query graph to optimize the subgraph matching algorithm. Based on this approach, we introduce an efficient enumeration method called Dynamic Graph Editing Enumeration, which significantly improves the performance of the algorithm. Our experimental results show that DGE outperforms current state-of-the-art algorithms in terms of computational efficiency and ability to solve more complex subgraph matching cases.
Zite Jiang, Shuai Zhang 0040, Boxiao Liu, Xingzhong Hou, Mengting Yuan 0001, Haihang You
IEEE Trans. Serv. Comput.6
2023 SUSHI: Ultra-High-Speed and Ultra-Low-Power Neuromorphic Chip Using Superconducting Single-Flux-Quantum Circuits
abstract
The rapid single-flux-quantum (RSFQ) superconducting technology is highly promising due to its ultra-high-speed computation with ultra-low-power consumption, making it an ideal solution for the post-Moore era. In superconducting technology, information is encoded and processed based on pulses that resemble the neuronal pulses present in biological neural systems. This has led to a growing research focus on implementing neuromorphic processing using superconducting technology. However, current research on superconducting neuromorphic processing does not fully leverage the advantages of superconducting circuits due to incomplete neuromorphic design and approach. Although they have demonstrated the benefits of using superconducting technology for neuromorphic hardware, their designs are mostly incomplete, with only a few components validated, or based solely on simulation. This paper presents SUSHI (Superconducting neUromorphic proceSsing cHIp) to fully leverage the potential of superconducting neuromorphic processing. Based on three guiding principles and our architectural and methodological designs, we address existing challenges and enables the design of verifiable and fabricable superconducting neuromorphic chips. We fabricate and verify a chip of SUSHI using superconducting circuit technology. Successfully obtaining the correct inference results of a complete neural network on the chip, this is the first instance of neural networks being completely executed on a superconducting chip to the best of our knowledge. Our evaluation shows that using approximately 105 Josephson junctions, SUSHI achieves a peak neuromorphic processing performance of 1,355 giga-synaptic operations per second (GSOPS) and a power efficiency of 32,366 GSOPS per Watt (GSOPS/W). This power efficiency outperforms the state-of-the-art neuromorphic chips TrueNorth and Tianjic by 81 and 50 times, respectively.
Zeshi Liu, Peiyao Qu, Huanli Liu, Minghui Niu, Liliang Ying, Guangming Tang, Haihang You
MICRO9
2023 Improving the performance of stochastic local search for maximum vertex weight clique problem using programming by optimization
Yi Chu, Chuan Luo 0002, Holger H. Hoos, Haihang You
Expert Syst. Appl.4
2023 A hybrid-order local search algorithm for set k-cover problem in wireless sensor networks
Boxiao Liu, Mengting Yuan 0001, Haihang You
Frontiers Comput. Sci.3
2023 A heterogeneous processing-in-memory approach to accelerate quantum chemistry simulation
Zeshi Liu, Wenqian Dong, Mengting Yuan 0001, Haihang You, Dong Li 0001
Parallel Comput.5
2023 DRONE: An Efficient Distributed Subgraph-Centric Framework for Processing Large-Scale Power-law Graphs
abstract
Nowadays, the ever-increasing volume of graph-structured data such as social networks, graph databases and knowledge graphs requires to be processed efficiently and scalably. These natural graphs commonly found in the real world have highly skewed power-law degree distribution and are called power-law graphs. The subgraph-centric programming model is a promising approach applied in many state-of-the-art distributed graph computing frameworks. However, the performance of subgraph-centric frameworks is limited when processing large-scale power-law graphs. When deployed to the subgraph-centric framework, existing graph partitioning algorithms are not suitable for power-law graphs. In this paper, we present a novel distributed graph computing framework, DRONE (Distributed gRaph cOmputiNg Engine), which leverages the subgraph-centric model and the vertex-cut graph partitioning strategy. DRONE also supports the fault tolerance mechanism to accommodate the increasing scale of machines with negligible overhead (6.48% on average). We further study the execution workflow of DRONE and propose an efficient and balanced graph partition algorithm (EBV) for DRONE. Experiments show that DRONE reduces the running time on real-world graphs by 25.6%, on average, compared to the state-of-the-art distributed graph computing frameworks. In addition, the EBV graph partition algorithm reduces the replication factor by at least 21.8% than other self-based partition algorithms. Our results indicate that DRONE has excellent potential in processing large-scale power-law graphs.
Shuai Zhang 0040, Zite Jiang, Xingzhong Hou, Mengting Yuan 0001, Haihang You
IEEE Trans. Parallel Distributed Syst.6
2022 Dynamic Weighted Semantic Correspondence for Few-Shot Image Generative Adaptation
abstract
Few-shot image generative adaptation, which finetunes well-trained generative models on limited examples, is of practical importance. The main challenge is that the few-shot model easily becomes overfitting. It can be attributed to two aspects: the lack of sample diversity for the generator and the failure of fidelity discrimination for the discriminator. In this paper, we introduce two novel methods to solve the diversity and fidelity respectively. Concretely, we propose dynamic weighted semantic correspondence to keep the diversity for the generator, which benefits from the richness of samples generated by source models. To prevent discriminator overfitting, we propose coupled training paradigm across the source and target domains to keep the feature extraction capability of the discriminator backbone. Extensive experiments show that our method outperforms previous methods both on image quality and diversity significantly.
Xingzhong Hou, Boxiao Liu, Shuai Zhang 0040, Lulin Shi, Zite Jiang, Haihang You
ACM Multimedia6
2022 Fast and efficient parallel breadth-first search with power-law graph transformation
Zite Jiang, Shuai Zhang 0040, Mengting Yuan 0001, Haihang You
Frontiers Comput. Sci.5
2021 Switchable K-class Hyperplanes for Noise-Robust Representation Learning
abstract
Optimizing the K-class hyperplanes in the latent space has become the standard paradigm for efficient representation learning. However, it’s almost impossible to find an optimal K-class hyperplane to accurately describe the latent space of massive noisy data. For this potential problem, we constructively propose a new method, named Switchable K-class Hyperplanes (SKH), to sufficiently describe the latent space by the mixture of K-class hyperplanes. It can directly replace the conventional single K-class hyperplane optimization as the new paradigm for noise-robust representation learning. When collaborated with the popular ArcFace on million-level data representation learning, we found that the switchable manner in SKH can effectively eliminate the gradient conflict generated by real-world label noise on a single K-class hyperplane. Moreover, combined with the margin-based loss functions (e.g. ArcFace), we propose a simple Posterior Data Clean strategy to reduce the model optimization deviation on clean dataset caused by the reduction of valid categories in each K-class hyperplane. Extensive experiments demonstrate that the proposed SKH easily achieves new state-of-the-art on IJB-B and IJB-C by encouraging noise-robust representation learning. Our code will be available at https://github.com/liubx07/SKH.git.
Boxiao Liu, Guanglu Song, Manyuan Zhang, Haihang You, Yu Liu 0015
ICCV4
2021 An Efficient and Balanced Graph Partition Algorithm for the Subgraph-Centric Programming Model on Large-scale Power-law Graphs
abstract
Nowadays, the parallel processing of power-law graphs is one of the biggest challenges in the field of graph computation. The subgraph-centric programming model is a promising approach and has been applied in many state-of-the-art distributed graph computing frameworks. The graph partition algorithm plays an important role in the overall performance of subgraph-centric frameworks. However, traditional graph partition algorithms have significant difficulties in processing large-scale power-law graphs. The major problem is the communication bottleneck found in many subgraph-centric frameworks. Detailed analysis indicates that the communication bottleneck is caused by the huge communication volume or the extreme message imbalance among partitioned subgraphs. The traditional partition algorithms do not consider both factors at the same time, especially on power-law graphs. In this paper, we propose a novel efficient and balanced vertex-cut graph partition algorithm (EBV) which grants appropriate weights to the overall communication cost and communication balance. We observe that the number of replicated vertices and the balance of edge and vertex assignment have a great influence on communication patterns of distributed subgraph-centric frameworks, which further affect the overall performance. Based on this insight, We design an evaluation function that quantifies the proportion of replicated vertices and the balance of edges and vertices assignments as important parameters. Besides, we sort the order of edge processing by the sum of end-vertices' degrees from small to large. Experiments show that EBV reduces replication factor and communication by at least 21.8% and 23.7% respectively than other self-based partition algorithms. When deployed in the subgraph-centric framework, it reduces the running time on power-law graphs by an average of 16.8% compared with the state-of-the-art partition algorithm. Our results indicate that EBV has a great potential in improving the performance of subgraph-centric frameworks for the parallel large-scale power-law graph processing.
Shuai Zhang 0040, Zite Jiang, Xingzhong Hou, Zhen Guan, Mengting Yuan 0001, Haihang You
ICDCS6
2021 Parallel CS-InSAR for Mapping Nationwide Deformation in China
abstract
Synthetic aperture radar (SAR) interferometer (InSAR) is now a key geodetic tool for monitoring the surface displacement. Thanks to ESA's Sentinel-1 sensors with IW mode as its default acquisition mode for land observations and its free access data policy, which have global coverage at moderate resolution with about 20m, national scale InSAR-based deformation is being studied in recent years by using big data techniques such as high performance computing and cloud computing. In this paper, we proposed the time series InSAR technique called Coherent-Scatterers InSAR (CS-InSAR) and its parallel solution for processing the whole CS-InSAR chain of Sentinel-1 data automatically and efficiently, considering the characteristics of CS-InSAR algorithm, such as frequent I/O data flow and heavy computation. By developing the parallelized CS-InSAR algorithm on the Big Earth Data Platform, 11922 satellite SAR data from September 2018 to December 2019 over China were processed, and the preliminary national InSAR-based surface deformation mapping for 2018–2019 was produced, with the deformation accuracy better than 0.6 cm in urban area.
Yixian Tang, Chao Wang 0004, Hong Zhang 0001, Haihang You, Wei Duan 0005, Jing Wang 0057, Longkai Dong
IGARSS4
2019 C-MIDN: Coupled Multiple Instance Detection Network With Segmentation Guidance for Weakly Supervised Object Detection
abstract
Weakly supervised object detection (WSOD) that only needs image-level annotations has obtained much attention recently. By combining convolutional neural network with multiple instance learning method, Multiple Instance Detection Network (MIDN) has become the most popular method to address the WSOD problem and been adopted as the initial model in many works. We argue that MIDN inclines to converge to the most discriminative object parts, which limits the performance of methods based on it. In this paper, we propose a novel Coupled Multiple Instance Detection Network (C-MIDN) to address this problem. Specifically, we use a pair of MIDNs, which work in a complementary manner with proposal removal. The localization information of the MIDNs is further coupled to obtain tighter bounding boxes and localize multiple objects. We also introduce a Segmentation Guided Proposal Removal (SGPR) algorithm to guarantee the MIL constraint after the removal and ensure the robustness of C-MIDN. Through a simple implementation of the C-MIDN with online detector refinement, we obtain 53.6% and 50.3% mAP on the challenging PASCAL VOC 2007 and 2012 benchmarks respectively, which significantly outperform the previous state-of-the-arts.
Gao Yan, Boxiao Liu, Nan Guo 0003, Xiaochun Ye, Fang Wan 0001, Haihang You, Dongrui Fan
ICCV6
2019 Empirical investigation of stochastic local search for maximum satisfiability
Yi Chu, Chuan Luo 0002, Shaowei Cai 0001, Haihang You
Frontiers Comput. Sci.4
2017 A new two-dimensional Fourier transform algorithm based on image sparsity
abstract
With the coming age of big data, the image signals play more and more important role in our life due to the extraordinary advance of network communication technology, and the corresponding high efficiency image processing techniques are demanded urgently. The Fourier transform is an important image processing tool which is used in a wide range of applications. Traditional Fourier transform algorithm computes on the value of each point of image, regardless of their properties in frequency domain. However, most image signals possess sparsity in frequency domain. In this paper, we present a new fast two-dimensional Fourier transform based on image sparsity. With hash function including a series of procedures such as random spectrum permutation, filtering and subsampling in frequency domain, the algorithm could identify and estimate the k largest coefficients quickly. In most sparse cases, the resulting algorithm performs faster than state-of-the-art fast Fourier transform algorithm, FFTW.
Sheng Shi, Runkai Yang, Haihang You
ICASSP3
2017 Hard Neighboring Variables Based Configuration Checking in Stochastic Local Search for Weighted Partial Maximum Satisfiability
abstract
Weighted partial maximum satisfiability (WPMS) is a significant generalization of maximum satisfiability (MAXSAT), weighted maximum satisfiability (weighted MAX-SAT) and unweighted partial maximum satisfiability (PMS), and WPMS can be widely used in real-world application domains. Recently, great breakthroughs have been made on stochastic local search (SLS) for solving MAX-SAT, weighted MAX-SAT, PMS and WPMS, resulting in several state-of-the-art SLS algorithms, such as CCLS, Dist and CCEHC. Indeed, boosting the practical performance on solving WPMS is of great interest, and the performance of SLS algorithms on solving WPMS could be further improved. In this paper, we follow this research direction, and propose a new SLS algorithm named CCHNV for solving WPMS. CCHNV adopts the framework of CCEHC, and employs a new forbidding strategy of configuration checking, named hard neighboring variables based configuration checking (HNVCC). Extensive experiments on a broad range of WPMS instances present that CCHNV pushes forward the state-of-the-art performance of SLS algorithms on solving WPMS, and is complementary to a state-of-the-art complete algorithm for solving WPMS.
Yi Chu, Chuan Luo 0002, Haihang You, Dongrui Fan
ICTAI4
2012 Comprehensive Workload Analysis and Modeling of a Petascale Supercomputer
Haihang You
JSSPP1
2010 Principles and construction of MSD adder in ternary optical computer
Yi Jin 0009, Yunfu Shen, Shiyi Xu, Guangtai Ding, Dongjian Yue, Haihang You
Sci. China Inf. Sci.7
2008 A comparison of search heuristics for empirical code optimization
abstract
This paper describes the application of various search techniques to the problem of automatic empirical code optimization. The search process is a critical aspect of auto-tuning systems because the large size of the search space and the cost of evaluating the candidate implementations makes it infeasible to find the true optimum point by brute force. We evaluate the effectiveness of Nelder-Mead Simplex, Genetic Algorithms, Simulated Annealing, Particle Swarm Optimization, Orthogonal search, and Random search in terms of the performance of the best candidate found under varying time limits.
Keith Seymour, Haihang You, Jack J. Dongarra
CLUSTER2
2008 The impact of paravirtualized memory hierarchy on linear algebra computational kernels and software
abstract
Previous studies have revealed that paravirtualization imposes minimal performance overhead on High Performance Computing (HPC) workloads, while exposing numerous benefits for this field. In this study, we are investigating the memory hierarchy characteristics of paravirtualized systems and their impact on automatically-tuned software systems. We are presenting an accurate characterization of memory attributes using hardware counters and user-process accounting. For that, we examine the proficiency of ATLAS, a quintessential example of an autotuning software system, in tuning the BLAS library routines for paravirtualized systems. In addition, we examine the effects of paravirtualization on the performance boundary. Our results show that the combination of ATLAS and Xen paravirtualization delivers native execution performance and nearly identical memory hierarchy performance profiles. Our research thus exposes new benefits to memory-intensive applications arising from the ability to slim down the guest OS without influencing the system performance. In addition, our findings support a novel and very attractive deployment scenario for computational science and engineering codes on virtual clusters and computational clouds.
Lamia Youseff, Keith Seymour, Haihang You, Jack J. Dongarra, Richard Wolski
HPDC3
2007 POET: Parameterized Optimizations for Empirical Tuning
abstract
The excessive complexity of both machine architectures and applications have made it difficult for compilers to statically model and predict application behavior. This observation motivates the recent interest in performance tuning using empirical techniques. We present a new embedded scripting language, POET (parameterized optimization for empirical tuning), for parameterizing complex code transformations so that they can be empirically tuned. The POET language aims to significantly improve the generality, flexibility, and efficiency of existing empirical tuning systems. We have used the language to parameterize and to empirically tune three loop optimizations - interchange, blocking, and unrolling - for two linear algebra kernels. We show experimentally that the time required to tune these optimizations using POET, which does not require any program analysis, is significantly shorter than that when using a full compiler-based source-code optimizer which performs sophisticated program analysis and optimizations.
Qing Yi, Keith Seymour, Haihang You, Richard W. Vuduc, Daniel J. Quinlan
IPDPS3