Qian Zhang 0020

dblp:04/2024-20 · DBLP profile ↗
← Back
21ranked-venue papers
9as first author
6since 2021 · last 2024
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 13 · 7 first-author · 1 since 2021Software engineering, systems software and programming languages · 11 · 4 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2024 MetaSim: A Search Engine for Finding Similar GitHub Repositories
abstract
How can we find other repositories on GitHub that are functionally similar to a specific repository? While GitHub offers keyword-based search functionality, there is a lack of a tool that can perform query by example to search and compare functionally similar repositories. To address this challenge, we present MetaSim: a search engine that finds similar GitHub repositories based on repository metadata features. MetaSim employs a customized technique to represent repository metadata in the embedding space for efficient indexing and searching. We construct a curated dataset of 267.6K public GitHub repositories to support our search engine. We evaluate our tool through a manual assessment on a set of 202 query by example repository and their corresponding matching pairs. Experiment results demonstrate that Readme alone can achieve high similarity precision (90.1%), which we define later. In contrast, the combined usage of Description, Topics, and Readme yields the best overall performance with similarity precision of 97.8%. To foster both research and practical applications, we open source our research artifacts through the MetaSim platform at https://metasim-app.github.io. The demonstration video of MetaSim is available at https://youtu.be/HnFnN3JclQw.
Md Rayhanul Masud, Md Omar Faruk Rokon, Qian Zhang 0020, Michalis Faloutsos
ICSME3
2024 Calico: Automated Knowledge Calibration and Diagnosis for Elevating AI Mastery in Code Tasks
abstract
Recent advancements in large language models (LLMs) have exhibited promising capabilities in addressing various tasks such as defect detection and program repair. Despite their prevalence, LLMs still face limitations in effectively handling these tasks. Common strategies to adapt them and improve their performance for specific tasks involve fine-tuning models based on user data or employing in-context learning with examples of desired inputs and outputs. However, they pose challenges for practical adoption due to the need for extensive computational resources, high-quality data, and continuous maintenance. Furthermore, neither strategy can explain or reason about the deficiencies of LLMs in the given tasks. We propose Calico to address the high cost of fine-tuning, eliminate the necessity for task-specific examples, and provide explanations of LLM deficiency. At the heart of Calico is an evolutionary approach that interleaves knowledge calibration and AI deficiency diagnosis. The key essence of Calico is as follows. First, it focuses on identifying knowledge gaps in LLMs’ program comprehension. Second, it conducts automated code refactoring to integrate the overlooked knowledge into the source code for mitigating those gaps. Third, it employs what-if analysis and counterfactual reasoning to determine a minimum set of overlooked knowledge necessary to improve the performance of LLMs in code tasks. We have extensively evaluated Calico over 8,938 programs on three most commonly seen code tasks. Our experimental results show that vanilla ChatGPT cannot fully understand code structures. With knowledge calibration, Calico improves it by 20% and exhibits comparable proficiency compared to fine-tuned LLMs. Deficiency diagnosis contributes to 8% reduction in program sizes while ensuring performance. These impressive results demonstrate the feasibility of utilizing a vanilla LLM for automated software engineering (SE) tasks, thereby avoiding the high computational costs associated with a fine-tuned model.
Yuxin Qiu, Jie Hu 0031, Qian Zhang 0020, Heng Yin 0001
ISSTA3
2023 Leveraging Hardware Probes and Optimizations for Accelerating Fuzz Testing of Heterogeneous Applications
abstract
There is a growing interest in the computer architecture community to incorporate heterogeneity and specialization to improve performance. Developers can create heterogeneous applications that consist of both host code and kernel code, where compute-intensive kernels can be offloaded from CPU to hardware accelerators. Testing such applications on real heterogeneous architectures is extremely challenging as kernels are black boxes, providing no information about the kernels’ internal execution to diagnose issues such as silent hangs or unexpected results. Additionally, inputs for heterogeneous applications are often large matrices, leading to a vast search space for identifying bug-revealing inputs.
Qian Zhang 0020, Hongbo Rong, Guoqing Harry Xu, Miryung Kim
ESEC/SIGSOFT FSE2
2022 HeteroGen: transpiling C to heterogeneous HLS code with automated test generation and program repair
abstract
Despite the trend of incorporating heterogeneity and specialization in hardware, the development of heterogeneous applications is limited to a handful of engineers with deep hardware expertise. We propose HeteroGen that takes C/C++ code as input and automatically generates an HLS version with test behavior preservation and better performance. Key to the success of HeteroGen is adapting the idea of search-based program repair to the heterogeneous computing domain, while addressing two technical challenges. First, the turn-around time of HLS compilation and simulation is much longer than the usual C/C++ compilation and execution time; therefore, HeteroGen applies pattern-oriented program edits guided by common fix patterns and their dependences. Second, behavior and performance checking requires testing, but test cases are often unavailable. Thus, HeteroGen auto-generates test inputs suitable for checking C to HLS-C conversion errors, while providing high branch coverage for the original C code.
Qian Zhang 0020, Guoqing Harry Xu, Miryung Kim
ASPLOS1
2021 QDiff: Differential Testing of Quantum Software Stacks
abstract
Over the past few years, several quantum software stacks (QSS) have been developed in response to rapid hardware advances in quantum computing. A QSS includes a quantum programming language, an optimizing compiler that translates a quantum algorithm written in a high-level language into quantum gate instructions, a quantum simulator that emulates these instructions on a classical device, and a software controller that sends analog signals to a very expensive quantum hardware based on quantum circuits. In comparison to traditional compilers and architecture simulators, QSSes are difficult to tests due to the probabilistic nature of results, the lack of clear hardware specifications, and quantum programming complexity.This work devises a novel differential testing approach for QSSes, named QDiff with three major innovations: (1) We generate input programs to be tested via semantics-preserving, source to source transformation to explore program variants. (2) We speed up differential testing by filtering out quantum circuits that are not worthwhile to execute on quantum hardware by analyzing static characteristics such as a circuit depth, 2-gate operations, gate error rates, and T1 relaxation time. (3) We design an extensible equivalence checking mechanism via distribution comparison functions such as Kolmogorov–Smirnov test and cross entropy.We evaluate QDiff with three widely-used open source QSSes: Qiskit from IBM, Cirq from Google, and Pyquil from Rigetti. By running QDiff on both real hardware and quantum simulators, we found several critical bugs revealing potential instabilities in these platforms. QDiff’s source transformation is effective in producing semantically equivalent yet not-identical circuits (i.e., 34% of trials), and its filtering mechanism can speed up differential testing by 66%.
Qian Zhang 0020, Guoqing Harry Xu, Miryung Kim
ASE2
2021 HeteroFuzz: fuzz testing to detect platform dependent divergence for heterogeneous applications
abstract
As specialized hardware accelerators like FPGAs become a prominent part of the current computing landscape, software applications are increasingly constructed to leverage heterogeneous architectures. Such a trend is already happening in the domain of machine learning and Internet-of-Things (IoT) systems built on edge devices. Yet, debugging and testing methods for heterogeneous applications are currently lacking. These applications may look similar to regular C/C++ code but include hardware synthesis details in terms of preprocessor directives. Therefore, their behavior under heterogeneous architectures may diverge significantly from CPU due to hardware synthesis details. Further, the compilation and hardware simulation cycle takes an enormous amount of time, prohibiting frequent invocations required for fuzz testing.
Qian Zhang 0020, Miryung Kim
ESEC/SIGSOFT FSE1
2020 HeteroRefactor: refactoring for heterogeneous computing with FPGA
abstract
Heterogeneous computing with field-programmable gate-arrays (FPGAs) has demonstrated orders of magnitude improvement in computing efficiency for many applications. However, the use of such platforms so far is limited to a small subset of programmers with specialized hardware knowledge. High-level synthesis (HLS) tools made significant progress in raising the level of programming abstraction from hardware programming languages to C/C++, but they usually cannot compile and generate accelerators for kernel programs with pointers, memory management, and recursion, and require manual refactoring to make them HLS-compatible. Besides, experts also need to provide heavily handcrafted optimizations to improve resource efficiency, which affects the maximum operating frequency, parallelization, and power efficiency.
Jason Lau, Aishwarya Sivaraman, Qian Zhang 0020, Muhammad Ali Gulzar, Jason Cong, Miryung Kim
ICSE3
2020 BigFuzz: Efficient Fuzz Testing for Data Analytics Using Framework Abstraction
abstract
As big data analytics become increasingly popular, data-intensive scalable computing (DISC) systems help address the scalability issue of handling large data. However, automated testing for such data-centric applications is challenging, because data is often incomplete, continuously evolving, and hard to know a priori. Fuzz testing has been proven to be highly effective in other domains such as security; however, it is nontrivial to apply such traditional fuzzing to big data analytics directly for three reasons: (1) the long latency of DISC systems prohibits the applicability of fuzzing: naïve fuzzing would spend 98% of the time in setting up a test environment; (2) conventional branch coverage is unlikely to scale to DISC applications because most binary code comes from the framework implementation such as Apache Spark; and (3) random bit or byte level mutations can hardly generate meaningful data, which fails to reveal real-world application bugs.
Qian Zhang 0020, Muhammad Ali Gulzar, Rohan Padhye, Miryung Kim
ASE1
2020 ARSketch: Sketch-Based User Interface for Augmented Reality Glasses
abstract
Hand gesture interaction is a key component in Augmented Reality (AR) / Mixed Reality (MR). Users usually interact with AR/MR devices, e.g., Microsoft HoloLens, etc., via hand gestures to express their intentions and the devices will recognize the gestures and respond accordingly to users. However, the use of such technique so far is limited to only a few less-expressive hand gestures, which, unfortunately, are insufficient or inadequate to input complex information.
Qian Zhang 0020
ACM Multimedia3
2020 ApproxIt: A Quality Management Framework of Approximate Computing for Iterative Methods
abstract
Approximate computing, being able to tradeoff computation quality (e.g., accuracy) and computational effort (e.g., energy) for error-tolerant applications such as media processing and the emerging recognition, mining, and synthesis (RMS) applications, has gained significant traction in recent years. With approximate computing, we expect to obtain acceptable results, but how do we make sure the quality of the final results are good enough? This challenging problems remains largely unexplored. As many of the RMS applications employ iterative methods (IMs) for solution-finding, wherein a sequence of improving approximate solutions are generated before reaching the final converged solution, in this paper, we propose ApproxIt, a novel quality management framework of approximate computing dedicated for IMs with quality guarantees. ApproxIt is comprised of two stages: 1) offline stage and 2) online stage. To be specific, at offline stage, we first analyze the manifold of parameter space to identify the given problem as convex case or nonconvex case at the offline stage. And for each case, we propose the corresponding runtime dynamic quality calibration scheme and reconfiguration control policy. Then during runtime, our proposed lightweight quality estimator will evaluate the intermediate quality at specific calibration iteration, which is determined by the novel Markov model-based calibration scheme. If quality violation occurs, the configuration control policy will select the most appropriate approximate computing mode for the following iterations. With the proposed dynamic effort scaling technique, ApproxIt is able to dramatically improve application energy efficiency under quality guarantees, as demonstrated in our experimental results.
Qian Zhang 0020, Qiang Xu 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2018 Lookup table allocation for approximate computing with memory under quality constraints
abstract
Computation kernels in emerging recognition, mining, and synthesis (RMS) applications are inherently error-resilient, where approximate computing can be applied to improve their energy efficiency by trading off computational effort and output quality. One promising approximate computing technique is to perform approximate computing with memory, which stores a subset of function responses in a lookup table (LUT), and avoids redundant computation when encountering similar input patterns. Limited by the memory space, most existing solutions simply store values for those frequently-appeared input patterns, without considering output quality and/or intrinsic characteristic of the target kernel. In this paper, we propose a novel LUT allocation technique for approximate computing with memory, which is able to dramatically improve the hit rate of LUT and hence achieves significant energy savings under given quality constraints. We also present how to apply the proposed LUT allocation solution for multiple computation kernels. Experimental results show the efficacy of our proposed methodology.
Ye Tian 0010, Qian Zhang 0020, Ting Wang 0008, Qiang Xu 0001
DATE2
2017 On resilient task allocation and scheduling with uncertain quality checkers
abstract
Many emerging applications are inherently error-resilient and hence do not require exact computation. Previous work on resilience-aware task allocation and scheduling problem first generates an initial energy-efficient task schedule on voltage-scalable multiprocessor system at design-time, and then conducts voltage adjustment according to runtime quality checking result. While energy efficiency improvements are quite encouraging, the final quality requirement might be violated because quality checkers are usually designed based on partial information and they are not always correct. In this paper, we propose to address the uncertainty issue of quality checkers in resilient task allocation and scheduling. To be specific, given the initial task schedule and quality checkers, we propose (i) a solution that ensures the final output quality with maximized probability, which can be applied to any approximate computing quality management systems; and (ii) a greedy runtime algorithm to achieve optimized energy efficiency gains. Experimental results on various task graphs demonstrate the efficacy of our proposed technique.
Qian Zhang 0020, Ting Wang 0008, Qiang Xu 0001
ASP-DAC1
2017 ApproxQA: A unified quality assurance framework for approximate computing
abstract
Approximate computing, being able to trade off computation quality and computational effort (e.g., energy) by exploiting the inherent error-resilience of emerging applications (e.g., recognition and mining), has garnered significant attention recently. No doubt to say, quality assurance is indispensable for satisfactory user experience with approximate computing, but this issue has remained largely unexplored in the literature. In this work, we propose a novel framework namely ApproxQA to tackle this problem, in which approximation mode tuning and rollback recovery are considered in a unified manner. To be specific, ApproxQA resorts to a two-level controller, in which the high-level approximation controller tunes approximation modes at a coarse-grained scale based on Q-learning while the low-level rollback controller judiciously determines whether to perform rollback recovery at a fine-grained scale based on the target quality requirement. Experimental results on various benchmark applications demonstrate that it significantly outperforms existing solutions in terms of energy efficiency with quality assurance.
Ting Wang 0008, Qian Zhang 0020, Qiang Xu 0001
DATE2
2017 ApproxLUT: A novel approximate lookup table-based accelerator
abstract
Computing with memory, which stores function responses of some input patterns into lookup tables offline and retrieves their values when encountering similar patterns (instead of performing online calculation), is a promising energy-efficient computing technique. No doubt to say, with a given lookup table size, the efficiency of this technique depends on which function responses are stored and how they are organized. In this paper, we propose a novel adaptive approximate lookup table based accelerator, wherein we store function responses in a hierarchical manner with increasing fine-grained granularity and accuracy. In addition, the proposed accelerator provides lightweight compensation on output results at different precision levels according to input patterns and output quality requirements. Moreover, our accelerator conducts adaptive lookup table search by exploiting input locality. Experimental results on various computation kernels show significant energy savings of the proposed accelerator over prior solutions.
Ye Tian 0010, Ting Wang 0008, Qian Zhang 0020, Qiang Xu 0001
ICCAD3
2016 ApproxMap: On task allocation and scheduling for resilient applications
abstract
Many emerging applications are inherently error-resilient and hence do not require exact computation. In this paper, we consider the task allocation and scheduling problem for mapping such applications to voltage-scalable multiprocessor systems. The proposed solution, namely ApproxMap, judiciously determines the mapping and execution sequence of resilient tasks to minimize the energy consumption of the application while meeting their target quality requirements and timing constraints. To be specific, ApproxMap generates energy-efficient yet flexible task schedule at design-time, and conducts lightweight online adjustment according to runtime dynamics for further energy-efficiency improvement. Experimental results on various task graphs demonstrate the efficacy of ApproxMap.
Juan Yi, Qian Zhang 0020, Ye Tian 0010, Ting Wang 0008, Weichen Liu 0001, Edwin H.-M. Sha, Qiang Xu 0001
ASP-DAC2
2016 On Effective and Efficient Quality Management for Approximate Computing
abstract
Approximate computing, where computation quality is traded off for better performance and/or energy savings, has gained significant tractions from both academia and industry. With approximate computing, we expect to obtain acceptable results, but how do we make sure the quality of the final results are acceptable? This challenging problem remains largely unexplored. In this paper, we propose an effective and efficient quality management framework to achieve controlled quality-efficiency tradeoffs. To be specific, at the offline stage, our solution automatically selects an appropriate approximator configuration considering rollback recovery for large occasional errors with minimum cost under the target quality requirement. Then during the online execution, our framework judiciously determines when and how to rollback, which is achieved with cost-effective yet accurate quality predictors that synergistically combine the outputs of several basic light-weight predictors. Experimental results demonstrate that our proposed solution can achieve 11% to 23% energy savings compared to existing solutions under the target quality requirement.
Ting Wang 0008, Qian Zhang 0020, Nam Sung Kim, Qiang Xu 0001
ISLPED2
2015 ApproxANN: an approximate computing framework for artificial neural network
Qian Zhang 0020, Ting Wang 0008, Ye Tian 0010, Qiang Xu 0001
DATE1
2015 ApproxMA: Approximate Memory Access for Dynamic Precision Scaling
abstract
Motivated by the inherent error-resilience of emerging recognition, mining, and synthesis (RMS) applications, approximate computing techniques such as precision scaling has been advocated for achieving energy-efficiency gains at the cost of small accuracy loss. Most existing solutions, however, focus on the approximation of on-chip computations without considering that of off-chip data accesses, whose energy consumption may contribute to a significant portion of the total energy. In this work, we propose a novel approximate memory access technique for dynamic precision scaling, namely ApproxMA. To be specific, by taking both runtime data precision constraints and error-resilient capabilities of the application into consideration, ApproxMA determines the precision of data accesses and loads scaled data from off-chip memory for computation. Experimental results with mixture model-based clustering algorithms demonstrate the efficacy of the proposed methodology.
Ye Tian 0010, Qian Zhang 0020, Ting Wang 0008, Qiang Xu 0001
ACM Great Lakes Symposium on VLSI2
2015 ApproxEigen: An Approximate Computing Technique for Large-Scale Eigen-Decomposition
abstract
Recognition, Mining, and Synthesis (RMS) applications are expected to make up much of the computing workloads of the future. Many of these applications (e.g., recommender systems and search engine) are formulated as finding eigenvalues/vectors of large-scale matrices. These applications are inherently error-tolerant, and it is often unnecessary, sometimes even impossible, to calculate all the eigenpairs. Motivated by the above, in this work, we propose a novel approximate computing technique for large-scale eigen-decomposition, namely ApproxEigen, wherein we focus on the practically-used Krylov subspace methods to find finite number of eigenpairs. With ApproxEigen, we provide a set of computation kernels with different levels of approximation for data pre-processing and solution finding, and conduct accuracy tuning under given quality constraints. Experimental results demonstrate that ApproxEigen is able to achieve significant energy-efficiency improvement while keeping high accuracy.
Qian Zhang 0020, Ye Tian 0010, Ting Wang 0008, Qiang Xu 0001
ICCAD1
2014 ApproxIt: An Approximate Computing Framework for Iterative Methods
abstract
Approximate computing, being able to tradeoff computation quality (e.g., accuracy) and computational effort (e.g., energy) for error-tolerant applications such as media processing and the emerging Recognition, Mining, and Synthesis (RMS) applications, has gained significant traction in recent years. Many of these applications employ iterative methods for solution-finding, wherein a sequence of improving approximate solutions are generated before reaching the final converged solution. In this work, we propose ApproxIt, a novel approximate computing framework for iterative methods with quality guarantees. To be specific, we present a lightweight quality estimator that is able to capture the solution quality of each iteration and use it to guide the selection of approximate computing mode in the next iteration. With the proposed dynamic effort scaling technique, ApproxIt is able to dramatically improve application energy efficiency under quality guarantees, as demonstrated in our experimental results.
Qian Zhang 0020, Rong Ye, Qiang Xu 0001
DAC1
2014 On hybrid memory allocation for FPGA behavioral synthesis (abstract only)
abstract
FPGA behavioral synthesis has gained significant momentum recently with the growing interests in accelerating high-performance computing applications. While the latest generation of high-level synthesis (HLS) tools has made significant progress, they still lack the support for certain high-level language features such as dynamic memory allocation, despite the fact that efficiently utilization of the on-chip memory resources in FPGAs is critical to achieve the performance and power consumption target for many designs.
Qian Zhang 0020, Chenfei Ma, Qiang Xu 0001
FPGA1