Yu Zhang 0086

dblp:50/671-86 · DBLP profile ↗
← Back
66ranked-venue papers
10as first author
40since 2021 · last 2026
0000-0001-6638-6442ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 18 · 2 first-author · 14 since 2021Software engineering, systems software and programming languages · 17 · 1 first-author · 11 since 2021Artificial intelligence and machine learning · 12 · 12 since 2021Applied, interdisciplinary, general and emerging computing · 9 · 2 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 7 since 2021Theory of computation · 4 · 1 since 2021Computer networks · 2 · 2 since 2021Security and privacy · 2Databases, data management, data science and information retrieval · 2 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 first-author · 1 since 2021
YearPublicationVenuePosition
2026 An empirical study of CGO usage in Go projects - Distribution, purposes, patterns and critical issues
Jinbao Chen, Boyao Ding, Yu Zhang 0086, Qingwei Li, Fugen Tang
J. Syst. Softw.3
2025 GoFree: Reducing Garbage Collection via Compiler-Inserted Freeing
abstract
In a memory-managed programming language, programmers allocate memory by creating new objects, but programmers never free memory. A garbage collector (GC) periodically reclaims memory used by unreachable objects. As an optimization based on escape analysis, some memory can be freed explicitly by instructions inserted by the compiler. This optimization reduces the cost of garbage collection, without changing the programming model. We designed and implemented this explicit freeing optimization for the Go language. We devised a new escape analysis that is both powerful and fast (𝑂(𝑁^2) time). Our escape analysis identifies short-lived heap objects that can be safely explicitly deallocated. We also implemented a freeing primitive that is safe for use in concurrent environments. We evaluated our system, GoFree, on 6 open-source Go programs. GoFree did not observably slow down compilation. At run time, GoFree deallocated on average 14% of allocated heap memory. It reduced GC frequency by 7%, GC time by 13%, wall-clock time by 2%, and heap size by 4%. We open-source GoFree.
Yu Zhang 0086, Michael D. Ernst, Jinbao Chen, Boyao Ding
CGO2
2025 Self-Attention to Operator Learning-based 3D-IC Thermal Simulation
abstract
Thermal management in 3D ICs is increasingly challenging due to higher power densities. Traditional PDESolving based methods, while accurate, are too slow for iterative design. Machine learning approaches like FNO provide faster alternatives but suffer from high-frequency information loss and high-fidelity data dependency. We introduce Self-Attention UNet Fourier Neural Operator (SAU-FNO), a novel framework combining self-attention and U-Net with FNO to capture longrange dependencies and model local high-frequency features effectively. Transfer learning is employed to fine-tune low-fidelity data, minimizing the need for extensive high-fidelity datasets and speeding up training. Experiments demonstrate that SAUFNO achieves state-of-the-art thermal prediction accuracy and provides an $842 \times$ speedup over traditional FEM methods, making it an efficient tool for advanced 3D IC thermal simulations.
Zhen Huang 0007, Wenkai Yang, Muxi Tang, Depeng Xie, Ting-Jung Lin, Yu Zhang 0086, Wei W. Xing, Lei He 0001
DAC7
2025 SpatialSplat: Efficient Semantic 3D from Sparse Unposed Images
abstract
A major breakthrough in 3D reconstruction is the feedforward paradigm to generate pixel-wise 3D points or Gaussian primitives from sparse, unposed images. To further incorporate semantics while avoiding the significant memory and storage costs of high-dimensional semantic features, existing methods extend this paradigm by associating each primitive with a compressed semantic feature vector. However, these methods have two major limitations: (a) the naively compressed feature compromises expressiveness, affecting the model's ability to capture fine-grained semantics, and (b) the pixel-wise primitive prediction introduces redundancy in overlapping areas, causing unnecessary memory overhead. To this end, we introduce \textbf{SpatialSplat}, a feedforward framework that produces redundancy-aware Gaussians and capitalizes on a dual-field semantic representation. Particularly, with the insight that primitives within the same instance exhibit high semantic consistency, we decompose the semantic representation into a coarse feature field that encodes uncompressed semantics with minimal primitives, and a fine-grained yet low-dimensional feature field that captures detailed inter-instance relationships. Moreover, we propose a selective Gaussian mechanism, which retains only essential Gaussians in the scene, effectively eliminating redundant primitives. Our proposed Spatialsplat learns accurate semantic information and detailed instances prior with more compact 3D Gaussians, making semantic 3D reconstruction more applicable. We conduct extensive experiments to evaluate our method, demonstrating a remarkable 60\% reduction in scene representation parameters while achieving superior performance over state-of-the-art methods. The code is available at https://github.com/shengyuuu/SpatialSplat.git
Yu Sheng, Jiajun Deng, Yu Zhang 0086, Bei Hua, Yanyong Zhang, Jianmin Ji
ICCV4
2025 CAFE-AD: Cross-Scenario Adaptive Feature Enhancement for Trajectory Planning in Autonomous Driving
abstract
Imitation learning based planning tasks on the nuPlan dataset have gained great interest due to their potential to generate human-like driving behaviors. However, open-loop training on the nuPlan dataset tends to cause causal confusion during closed-loop testing, and the dataset also presents a longtail distribution of scenarios. These issues introduce challenges for imitation learning. To tackle these problems, we introduce CAFE-AD, a Cross-Scenario Adaptive Feature Enhancement for Trajectory Planning in Autonomous Driving method, designed to enhance feature representation across various scenario types. We develop an adaptive feature pruning module that ranks feature importance to capture the most relevant information while reducing the interference of noisy information during training. Moreover, we propose a cross-scenario feature interpolation module that enhances scenario information to introduce diversity, enabling the network to alleviate overfitting in dominant scenarios. We evaluate our method CAFEAD, on the challenging public nuPlan Test14-Hard closed-loop simulation benchmark. The results demonstrate that CAFEAD outperforms state-of-the-art methods including rule-based and hybrid planners, and exhibits the potential in mitigating the impact of long-tail distribution within the dataset. Additionally, we further validate its effectiveness in real-world environments. The code and models will be made available at https://github.com/AlniyatRui/CAFE-AD.
Junrui Zhang 0012, Chenjie Wang, Jie Peng 0002, Jianmin Ji, Yu Zhang 0086, Yanyong Zhang
ICRA6
2025 Improving Efficiency of Answer Set Planning with Rough Solutions from Large Language Models for Robotic Task Planning
abstract
Answer Set Programming (ASP) planning can be used to refine the rough solutions generated by Large Language Models (LLMs) to handle specific restrictions of actions, i.e., reconstruct the rough solutions to be executable, for robotic task planning. However, it is still challenging to efficiently solve ASP programs that have multiple variables with large domains, which prevents the above application of ASP planning from real-world task planning problems. In this paper, we consider how to reduce the domains of variables without losing possible solutions for ASP planning, while given these rough solutions from LLMs. Based on the above reduction, we introduce CLMASP, an approach that couples LLMs with ASP for robotic task planning. We evaluate CLMASP on the VirtualHome platform for common indoor tasks, demonstrating a significant improvement in the executable rate from under 10% to nearly 90% and reducing average ASP planning time from over 2 hours to under 5 seconds. Code is available at https://github.com/CLMASP/CLMASP.
Xinrui Lin, Yangfan Wu, Yu Zhang 0086, Jianmin Ji, Yanyong Zhang
IJCAI5
2025 Enhancing LLM to Decompile Optimized PTX to Readable CUDA for Tensor Programs
abstract
The growing demand for high-performance tensor programs on GPUs, especially for large language models (LLMs), necessitates advanced compilation and optimization techniques. However, the critical task of analyzing optimized, low-level PTX code for performance tuning or understanding poses significant challenges. While LLMs hold promise for PTX-to-CUDA de-compilation to improve code intelligibility, their effectiveness is severely limited by the scarcity of aligned training data and the inherent complexity of highly optimized, unrolled PTX code.In this work, we explore methodologies to significantly enhance LLM capabilities for accurate and readable PTX-to-CUDA decompilation and present PtxDec, a decompilation prototype implementing our approach. To overcome the critical barrier of data scarcity, we develop a compiler-based data augmentation framework coupled with rigorous post-processing, enabling the creation of a large-scale, high-quality dataset of 400K aligned CUDA-PTX kernel pairs for effective LLM training. Furthermore, to empower LLMs to handle the complexity of optimized PTX, we introduce Rolled-PTX—an intermediate representation generated through heuristic loop rerolling during preprocessing. Rolled-PTX condenses unrolled patterns, drastically simplifying the input structure presented to the LLM and aligning it better with higher-level loop constructs.Comprehensive evaluation demonstrates that PtxDec achieves substantial performance gains: our approach yields a 2.3×–3.1× improvement in functional accuracy over baseline methods, alongside significant enhancements in generated code readability and scheduling consistency with the original optimized kernels. Ablation studies further validate the contribution of each proposed component to the overall performance.To the best of our knowledge, this is the first work tackling PTX-to-CUDA decompilation, specifically focusing on and demonstrating effective strategies for augmenting LLMs to overcome the key challenges in this domain.
Fugen Tang, Yu Zhang 0086, Chengru Song
ASE3
2025 UrgenGo: Urgency-Aware Transparent GPU Kernel Launching for Autonomous Driving
abstract
The rapid advancements in autonomous driving have introduced increasingly complex, real-time GPU-bound tasks critical for reliable vehicle operation. However, the proprietary nature of these autonomous systems and closed-source GPU drivers hinder fine-grained control over GPU executions, often resulting in missed deadlines that compromise vehicle performance. To address this, we present UrgenGo, a non-intrusive, urgency-aware GPU scheduling system that operates without access to application source code. UrgenGo implicitly prioritizes GPU executions through transparent kernel launch manipulation, employing task-level stream binding, delayed kernel launching, and batched kernel launch synchronization. We conducted extensive real-world evaluations in collaboration with a self-driving startup, developing 11 GPU-bound task chains for a realistic autonomous navigation application and implementing our system on a self-driving bus. Our results show a significant 61% reduction in the overall deadline miss ratio, compared to the state-of-the-art GPU scheduler that requires source code modifications.
Hanqi Zhu, Wuyang Zhang, Ziyang Tao, Xinrui Lin, Yu Zhang 0086, Jianmin Ji, Yanyong Zhang
MobiCom6
2025 OA-DET3D: Embedding Object Awareness As A General Plug-in for Multi-Camera 3D Object Detection
Xiaomeng Chu, Jiajun Deng, Jianmin Ji, Yu Zhang 0086, Houqiang Li, Yanyong Zhang
Int. J. Comput. Vis.4
2025 PauliForest: Connectivity-Aware Synthesis and Pauli-Oriented Qubit Mapping for Near-Term Quantum Simulation
abstract
Quantum simulation is the foundation for the design of many algorithms which share subroutines known as quantum simulation kernels. Optimizing the compilation of these kernels is crucial, involving two key components: 1) circuit synthesis and 2) qubit mapping. However, existing circuit synthesis methods either overlook qubit connectivity constraints (QCCs) or prioritize minimizing gate count over optimizing circuit depth. Similarly, current qubit mapping techniques do not work well with circuit synthesis methods. To address these limitations, we propose PauliForest, which comprises a connectivity-aware circuit synthesis algorithm and a Pauli-oriented qubit mapping algorithm. The synthesis algorithm employs heuristic strategies to generate shallower circuits, while the qubit mapping algorithm seamlessly collaborates with the circuit synthesis process. Compared to the state-of-the-art Paulihedral compiler, our approach significantly reduces both CNOT gate counts (by 13%) and circuit depths (by 25%). Experiments on a noisy simulator and a real superconducting quantum computer show that our algorithm can improve the fidelity of quantum circuit execution compared to Paulihedral.
Yongshang Li, Yu Zhang 0086, Haoning Deng, Mingyu Chen 0009
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2025 A Survey of Multilanguage Interoperability and Its Program Analysis
abstract
Since multilanguage programming has the strength of interoperating languages with different features and paradigms, and it also enables the reuse of existing libraries, developers often use multilanguage interoperability in software systems of different application domains. Program analysis is an effective way to maintain the reliability and safety of software. However, numerous challenges appear when using program analysis to analyze multilanguage interoperability. However, there still lacks a systematic overview of the multilanguage interoperability and the corresponding program analysis (e.g., what are the research trends, how existing works solve them). To bridge this gap, we conducted a comprehensive investigation of the 195 research works related to multilanguage interoperability and its program analysis in the past 38 years (1987–2024). In this article, we classify these works into three categories with 10 research perspectives, including foreign interface design, interface definition and generation, intermediate representation (IR), semantics, static analysis of memory management, type system, exception handling and concurrency, dynamic analysis, and others. Then, we evaluate research trends that affect multilanguage interoperability and key static/dynamic language futures of multilanguage program analysis. Finally, we discuss the open challenges and future research directions of multilanguage interoperability.
Mingzhe Hu, Le Yu 0002, Yu Zhang 0086, Liping Han
IEEE Trans. Reliab.3
2025 Productively Generating a High-Performance Linear Algebra Library on FPGAs
abstract
Linear algebra computations can be greatly accelerated using spatial accelerators on FPGAs. As a standard building block of linear algebra applications, BLAS covers a wide range of compute patterns that vary vastly in data reuse, bottleneck resources, matrix storage layouts, and data types. However, existing implementations of BLAS routines on FPGAs are stuck in the dilemma of productivity and performance. They either require extensive human effort or fail to leverage the properties of routines for acceleration. We introduce Lasa, a framework composed of a programming model and a compiler, designed to address the dilemma by abstracting (for productivity) and specializing (for performance) the architecture of a spatial accelerator. The programming model realizes systolic arrays using uniform recurrence equations and space-time transforms. Streaming tensors, an intuitive dataflow-style abstraction, is proposed to uniformly describe the movement, storage, and transpose of input and output data across the spatial components. According to streaming tensors, a customized memory hierarchy is automatically built on an FPGA by our compiler. The compiler further specializes the architecture with transparent optimizations on FPGAs. Using this framework, we develop a complete BLAS library, demonstrating performance in parity with expert-written HLS code for BLAS level 3 routines, 76%–94% machine peak for level 1 and 2 routines, and 1.6X–13X speedup by leveraging the matrix properties such as symmetry, triangularity, and bandness.
Xiaochen Hao, Mingzhe Zhang 0002, Ce Sun 0001, Zhuofu Tao, Hongbo Rong, Yu Zhang 0086, Lei He 0001, Eric Petit 0002, Yun Liang 0001
ACM Trans. Reconfigurable Technol. Syst.6
2024 SDAC: A Multimodal Synthetic Dataset for Anomaly and Corner Case Detection in Autonomous Driving
abstract
Nowadays, closed-set perception methods for autonomous driving perform well on datasets containing normal scenes. However, they still struggle to handle anomalies in the real world, such as unknown objects that have never been seen while training. The lack of public datasets to evaluate the model performance on anomaly and corner cases has hindered the development of reliable autonomous driving systems. Therefore, we propose a multimodal Synthetic Dataset for Anomaly and Corner case detection, called SDAC, which encompasses anomalies captured from multi-view cameras and the LiDAR sensor, providing a rich set of annotations for multiple mainstream perception tasks. SDAC is the first public dataset for autonomous driving that categorizes anomalies into object, scene, and scenario levels, allowing the evaluation under different anomalous conditions. Experiments show that closed-set models suffer significant performance drops on anomaly subsets in SDAC. Existing anomaly detection methods fail to achieve satisfactory performance, suggesting that anomaly detection remains a challenging problem. We anticipate that our SDAC dataset could foster the development of safe and reliable systems for autonomous driving.
Yu Zhang 0086, Yingqing Xia, Yanyong Zhang, Jianmin Ji
AAAI2
2024 Crop: An Analytical Cost Model for Cross-Platform Performance Prediction of Tensor Programs
abstract
Learn-based cost models used for tensor compiler auto-tuning often suffer from poor performance when trained on one hardware platform and applied to another. This issue necessitates collecting performance data for each potential platform during model deployment, incurring significant overhead.
Yu Zhang 0086, Shuo Liu 0019, Yi Zhai 0005
DAC2
2024 Graph-Specific Schema-Guided Query Optimization
Chaijun Xu, Yunlong Liang, Yu Zhang 0086, Hairong Hu, Yanyong Zhang
DASFAA (1)3
2024 Map++: Towards User-Participatory Visual SLAM Systems with Efficient Map Expansion and Sharing
abstract
Constructing precise 3D maps is crucial for the development of future map-based systems such as self-driving and navigation. However, generating these maps in complex environments, such as multi-level parking garages or shopping malls, remains a formidable challenge. In this paper, we introduce a participatory sensing approach that delegates map-building tasks to map users, thereby enabling cost-effective and continuous data collection. The proposed method harnesses the collective efforts of users, facilitating the expansion and ongoing update of the maps as the environment evolves.
Hanqi Zhu, Yifan Duan, Wuyang Zhang, Longfei Shangguan, Yu Zhang 0086, Jianmin Ji, Yanyong Zhang
MobiCom6
2024 Enabling Tensor Language Model to Assist in Generating High-Performance Tensor Programs for Deep Learning
Yi Zhai 0005, Keyu Pan, Renwei Zhang, Shuo Liu 0019, Zichun Ye, Jianmin Ji, Jie Zhao 0002, Yu Zhang 0086, Yanyong Zhang
OSDI10
2024 MEA2: A Lightweight Field-Sensitive Escape Analysis with Points-to Calculation for Golang
abstract
Escape analysis plays a crucial role in garbage-collected languages as it enables the allocation of non-escaping variables on the stack by identifying the dynamic lifetimes of objects and pointers. This helps in reducing heap allocations and alleviating garbage collection pressure. However, Go, as a garbage-collected language, employs a fast yet conservative escape analysis, which is field-insensitive and omits point-to-set calculation to expedite compilation. This results in more variables being allocated on the heap. Empirical statistics reveal that field access and indirect memory access are prevalent in real-world Go programs, suggesting potential opportunities for escape analysis to enhance program performance. In this paper, we propose MEA 2 , an escape analysis framework atop GoLLVM (an LLVM-based Go compiler), which combines field sensitivity and points-to analysis. Moreover, a novel generic function summary representation is designed to facilitate fast inter-procedural analysis. We evaluated it by using MEA 2 to perform stack allocation in 12 wildly-use open-source projects. The results show that, compared to Go’s escape analysis, MEA 2 can reduce heap allocation sites by 7.9 % on average (up to 25.7 % ) while reducing the dynamic memory allocation size by 11.6 % on average (up to 35.5 % ). All this is achieved while keeping the time overhead of escape analysis within 1 % of the compilation process.
Boyao Ding, Qingwei Li, Yu Zhang 0086, Fugen Tang, Jinbao Chen
Proc. ACM Program. Lang.3
2023 TLP: A Deep Learning-Based Cost Model for Tensor Program Tuning
abstract
Tensor program tuning is a non-convex objective optimization problem, to which search-based approaches have proven to be effective. At the core of the search-based approaches lies the design of the cost model. Though deep learning-based cost models perform significantly better than other methods, they still fall short and suffer from the following problems. First, their feature extraction heavily relies on expert-level domain knowledge in hardware architectures. Even so, the extracted features are often unsatisfactory and require separate considerations for CPUs and GPUs. Second, a cost model trained on one hardware platform usually performs poorly on another, a problem we call cross-hardware unavailability.
Yi Zhai 0005, Yu Zhang 0086, Shuo Liu 0019, Xiaomeng Chu, Jie Peng 0002, Jianmin Ji, Yanyong Zhang
ASPLOS (2)2
2023 Bi-LRFusion: Bi-Directional LiDAR-Radar Fusion for 3D Dynamic Object Detection
abstract
LiDAR and Radar are two complementary sensing approaches in that LiDAR specializes in capturing an object's 3D shape while Radar provides longer detection ranges as well as velocity hints. Though seemingly natural, how to efficiently combine them for improved feature representation is still unclear. The main challenge arises from that Radar data are extremely sparse and lack height information. Therefore, directly integrating Radar features into LiDAR-centric detection networks is not optimal. In this work, we introduce a bi-directional LiDAR-Radar fusion framework, termed Bi-LRFusion, to tackle the challenges and improve 3D detection for dynamic objects. Technically, Bi-LRFusion involves two steps: first, it enriches Radar's local features by learning important details from the LiDAR branch to alleviate the problems caused by the absence of height information and extreme sparsity; second, it combines LiDAR features with the enhanced Radar features in a unified bird's-eye-view representation. We conduct extensive experiments on nuScenes and ORR datasets, and show that our Bi-LRFusion achieves state-of-the-art performance for detecting dynamic objects. Notably, Radar data in these two datasets have different formats, which demonstrates the generalizability of our method. Codes will be published.
Yingjie Wang 0005, Jiajun Deng, Yao Li 0016, Jinshui Hu, Cong Liu 0006, Yu Zhang 0086, Jianmin Ji, Wanli Ouyang, Yanyong Zhang
CVPR6
2023 Lasa: Abstraction and Specialization for Productive and Performant Linear Algebra on FPGAs
abstract
Linear algebra can often be significantly expedited by spatial accelerators on FPGAs. As a broadly-adopted linear algebra library, BLAS requires extensive optimizations for routines that vary vastly in data reuse, bottleneck resources, matrix storage layouts, and data types. Existing solutions are stuck in the dilemma of productivity and performance. We introduce Lasa, a framework composed of a programming model and a compiler, that addresses the dilemma by abstracting (for productivity) and specializing (for performance) the architecture of a spatial accelerator. Lasa abstracts a compute and its I/O as two dataflow graphs. A compiler maps the graphs onto systolic arrays and a customized memory heirarchy. The compiler further specializes the architecture transparently. In this framework, we develop 14 key BLAS routines, and demonstrate performance in parity with expert-written HLS code for BLAS level 3 routines, >=80% machine peak performance for level 2 and 1 routines, and 1.6X-7X speed up by taking advantage of matrix properties of symmetry, triangularity and bandness.
Xiaochen Hao, Mingzhe Zhang 0002, Ce Sun 0001, Zhuofu Tao, Hongbo Rong, Yu Zhang 0086, Lei He 0001, Eric Petit 0002, Yun Liang 0001
FCCM6
2023 P3O: Transferring Visual Representations for Reinforcement Learning via Prompting
abstract
It is important for deep reinforcement learning (DRL) algorithms to transfer their learned policies to new environments that have different visual inputs. In this paper, we introduce Prompt based Proximal Policy Optimization (P3O), a three-stage DRL algorithm that transfers visual representations from a target to a source environment by applying prompting. The process of P3O consists of three stages: pre-training, prompting, and predicting. In particular, we specify a prompt-transformer for representation conversion and propose a two-step training process to train the prompt-transformer for the target environment, while the rest of the DRL pipeline remains unchanged. We implement P3O and evaluate it on the OpenAI CarRacing video game. The experimental results show that P3O outperforms the state-of-the-art visual transferring schemes. In particular, P3O allows the learned policies to perform well in environments with different visual inputs, which is much more effective than retraining the policies in these environments.
Guoliang You, Xiaomeng Chu, Yifan Duan, Jie Peng 0002, Jianmin Ji, Yu Zhang 0086, Yanyong Zhang
ICME6
2023 Reinforcement Learning for Robot Navigation with Adaptive Forward Simulation Time (AFST) in a Semi-Markov Model
abstract
Deep reinforcement learning (DRL) algorithms have proven effective in robot navigation, especially in unknown environments, by directly mapping perception inputs into robot control commands. However, most existing methods ignore the local minimum problem in navigation and thereby cannot handle complex unknown environments. In this paper, we propose the first DRL-based navigation method modeled by a semi-Markov decision process (SMDP) with continuous action space, named Adaptive Forward Simulation Time (AFST), to overcome this problem. Specifically, we reduce the dimensions of the action space and improve the distributed proximal policy optimization (DPPO) algorithm for the specified SMDP problem by modifying its GAE to better estimate the policy gradient in SMDPs. Experiments in various unknown environments demonstrate the effectiveness of AFST.
Yu'an Chen, Ruosong Ye, Ziyang Tao, Hongjian Liu, Guangda Chen, Jie Peng 0002, Jun Ma 0034, Yu Zhang 0086, Jianmin Ji, Yanyong Zhang
IROS8
2023 CluB: Cluster Meets BEV for LiDAR-Based 3D Object Detection
abstract
Currently, LiDAR-based 3D detectors are broadly categorized into two groups, namely, BEV-based detectors and cluster-based detectors. BEV-based detectors capture the contextual information from the Bird's Eye View (BEV) and fill their center voxels via feature diffusion with a stack of convolution layers, which, however, weakens the capability of presenting an object with the center point. On the other hand, cluster-based detectors exploit the voting mechanism and aggregate the foreground points into object-centric clusters for further prediction. In this paper, we explore how to effectively combine these two complementary representations into a unified framework. Specifically, we propose a new 3D object detection framework, referred to as CluB, which incorporates an auxiliary cluster-based branch into the BEV-based detector by enriching the object representation at both feature and query levels. Technically, CluB is comprised of two steps. First, we construct a cluster feature diffusion module to establish the association between cluster features and BEV features in a subtle and adaptive fashion. Based on that, an imitation loss is introduced to distill object-centric knowledge from the cluster features to the BEV features. Second, we design a cluster query generation module to leverage the voting centers directly from the cluster branch, thus enriching the diversity of object queries. Meanwhile, a direction loss is employed to encourage a more accurate voting center for each cluster. Extensive experiments are conducted on Waymo and nuScenes datasets, and our CluB achieves state-of-the-art performance on both benchmarks.
Yingjie Wang 0005, Jiajun Deng, Yuenan Hou, Yao Li 0016, Yu Zhang 0086, Jianmin Ji, Wanli Ouyang, Yanyong Zhang
NeurIPS5
2023 CGORewritter: A better way to use C library in G
abstract
CGO is a foreign function interface mechanism that enables the creation of Go packages that call C code. It provides a way to reuse legacy C code or high-performance C libraries. However, it is tedious and error-prone to manually write the bindings of C libraries in Go. To make better use of CGO to realize the use of C libraries in Go, we propose CGORewritter in this paper. For a given Go library and corresponding C library, the internal code of the Go API function can be rewritten while keeping the Go API unchanged, and the C API function can be called through CGO in a semi-automatic way so that the application layer program can upgrade without any change. We use this method to rewrite the go/crypto with the OpenSSL library. Experimental results show that the rewritten code maintains full functionality and can be used by application code without modification. Besides, the rewritten code gains a 0.97-3.07× speedup.
Boyao Ding, Yu Zhang 0086, Jinbao Chen, Mingzhe Hu, Qingwei Li
SANER2
2023 Cross-Language Call Graph Construction Supporting Different Host Languages
abstract
Modern software systems are increasingly multi-lingual, which consist of components developed in different programming languages to reuse existing libraries and com-bine language features. Foreign function interface (FFI) is a mechanism that enables interoperation between a host language and a guest language. CFFI interoperating with external C is part of the language standard for almost all languages. For example, Python/C API is the CFFI between Python and C/C++. Python host with C guest can achieve both productivity and performance, and is widely used by many mainstream software systems in different application domains. The popularity of these software systems makes high demands on program analysis of multilingual codebases. A fundamental challenge is to construct call graphs that capture the connectivity between host and guest languages. In this work, we present a novel approach to call graph construction for calls from different host languages to C/C++ foreign functions. The semantics of the foreign function declaration interfaces are modeled to establish cross-language call relationships, taking into account the semantic abstraction to support different host languages. Graph transformation and function node fusions are defined to build the complete call graphs. We demonstrate experimentally that static call graphs for Python and JavaScript calling C/C++ can be constructed effectively and automatically. The call graph construction can further efficiently work as an IDE language server and integrate with other tools.
Mingzhe Hu, Yu Zhang 0086, Yan Xiong 0001
SANER3
2023 Multi-Modal 3D Object Detection in Autonomous Driving: A Survey
Yingjie Wang 0005, Qiuyu Mao, Hanqi Zhu, Jiajun Deng, Yu Zhang 0086, Jianmin Ji, Houqiang Li, Yanyong Zhang
Int. J. Comput. Vis.5
2023 An empirical study of the Python/C API on evolution and bug patterns
abstract
Abstract Python is a popular programming language, and a large part of its appeal comes from diverse libraries and extension modules. In the bloom of data science and machine learning, Python frontend with C/C++ native implementation achieves both productivity and performance and has almost become the standard structure for many mainstream software systems. However, feature discrepancies between two languages such as exception handling, memory management, and type system can pose many safety hazards in the interface layer using the Python/C API. In this paper, we carry out an empirical study of the Python/C API on evolution and bug patterns. The evolution analysis includes Python/C API design in CPython compilers and its usage in mainstream software. By designing and applying a static analysis toolset, we reveal the evolution and usage statistics of the Python/C API and provide a summary and classification of 9 common bug patterns. In Pillow, a widely used Python imaging library, we find 48 bugs, 19 of which are undiscovered before. Our toolset can be easily extended to access different types of syntactic bug‐finding checkers, and our systematical taxonomy to classify bugs can guide the construction of more highly automated and high‐precision bug‐finding tools.
Mingzhe Hu, Yu Zhang 0086
J. Softw. Evol. Process.2
2023 Timing-Aware Qubit Mapping and Gate Scheduling Adapted to Neutral Atom Quantum Computing
abstract
As a less developed but potential quantum technology, neutral atoms (NAs) can provide advantages, including higher qubit connectivity, longer-range interactions, and much more native multicontrol gates than superconductivity. Long-range interactions, however, prevent parallelism of interacting qubit pairs with their surrounding restriction zones. Therefore, the quantum program cannot be run directly on an NA quantum computer (NAQC) unless compiled. The recent compiling study of NAQC applies simple layer-by-layer scheduling and does not consider the difference in gate duration. To address the above issues, we focus on the qubit mapping and quantum gate scheduling (M&S) problem of quantum circuits to meet hardware constraints of superconductivity and even NAs. Our goal is to shorten the execution time to mitigate decoherence noise. We propose a block-game-like abstraction mechanism TETRIS which is an abstract model of the M&S problem and aware of the duration difference of gates and several other NA characteristics. Based on the abstraction, we propose a heuristic greedy algorithm (HGA) to solve the M&S problem efficiently. To further speed up the execution of the circuit, we embed HGA into a Monte Carlo tree search (MCTS) framework to solve the M&S problem, which consumes more compiling time but achieves a better result. Comparing our MCTS algorithm with the only recent M&S algorithm for NAQC, the average speedup ratio is$1.75\times $for several quantum circuits collected from RevLib, and$1.18\times $for circuits from Qiskit Lib.
Yongshang Li, Yu Zhang 0086, Mingyu Chen 0009, Xiang-Yang Li 0001, Peng Xu 0029
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2023 TrajMatch: Toward Automatic Spatio-Temporal Calibration for Roadside LiDARs Through Trajectory Matching
abstract
Recently, deploying sensors such as LiDARs on the roadside to monitor the passing traffic and assist autonomous vehicle perception has become popular. However, unlike autonomous vehicle systems, roadside sensor systems involve sensors from different subsystems, resulting in a lack of synchronization in both time and space between the sensors. Calibration is a critical technology that enables the central server to fuse data generated by different location infrastructures, which vastly improves sensing range and detection robustness. Regrettably, existing calibration algorithms frequently assume that LiDARs have significant overlap or that temporal calibration has already been achieved. However, since these assumptions do not always hold in real-world scenarios, the calibration results obtained from existing algorithms are frequently unsatisfactory. In this paper, we propose TrajMatch - the first system that can automatically calibrate roadside LiDARs in both time and space. The main idea is to automatically calibrate the sensors based on the result of the detection/tracking task, rather than relying on extracting special features. Furthermore, we propose a novel mechanism for evaluating calibration parameters that align with our algorithm, and we demonstrate its effectiveness through experiments. This mechanism can also guide parameter iterations for multiple calibrations, further enhancing the accuracy and efficiency of our calibration method. Finally, to evaluate the performance of TrajMatch, we collected two datasets, one simulated dataset LiDARnet-sim 1.0 and one real-world dataset. The experimental results show that TrajMatch can achieve a spatial calibration error of less than$10cm$and a temporal calibration error of less than$1.5ms$.
Haojie Ren, Sha Zhang 0002, Sugang Li, Yao Li 0016, Xinchen Li, Jianmin Ji, Yu Zhang 0086, Yanyong Zhang
IEEE Trans. Intell. Transp. Syst.7
2023 VPFNet: Improving 3D Object Detection With Virtual Point Based LiDAR and Stereo Data Fusion
abstract
It has been well recognized that fusing the complementary information from depth-aware LiDAR point clouds and semantic-rich stereo images would benefit 3D object detection. Nevertheless, it is non-trivial to explore the inherently unnatural interaction between sparse 3D points and dense 2D pixels. To ease this difficulty, the recent approaches generally project the 3D points onto the 2D image plane to sample the image data and then aggregate the data at the points. However, these approaches often suffer from the mismatch between the resolution of point clouds and RGB images, leading to sub-optimal performance. Specifically, taking the sparse points as the multi-modal data aggregation locations causes severe information loss for high-resolution images, which in turn undermines the effectiveness of multi-sensor fusion. In this paper, we presentVPFNet—a new architecture that cleverly aligns and aggregates the point cloud and image data at the “virtual” points. Particularly, with their density lying between that of the 3D points and 2D pixels, the virtual points can nicely bridge the resolution gap between the two sensors, and thus preserve more information for processing. Moreover, we also investigate the data augmentation techniques that can be applied to both point clouds and RGB images, as the data augmentation has made non-negligible contribution towards 3D object detectors to date. We have conducted extensive experiments on KITTI dataset, and have observed good performance compared to the state-of-the-art methods. Remarkably, ourVPFNetachieves 83.21% moderate$AP_{3D}$and 91.86% moderate$AP_{BEV}$on the KITTI test set. The network design also takes computation efficiency into consideration – we can achieve a FPS of 15 on a single NVIDIA RTX 2080Ti GPU.
Hanqi Zhu, Jiajun Deng, Yu Zhang 0086, Jianmin Ji, Qiuyu Mao, Houqiang Li, Yanyong Zhang
IEEE Trans. Multim.3
2022 QCIR: Pattern Matching Based Universal Quantum Circuit Rewriting Framework
abstract
Due to multiple limitations of quantum computers in the NISQ era, quantum compilation efforts are required to efficiently execute quantum algorithms on NISQ devices Program rewriting based on pattern matching can improve the generalization ability of compiler optimization. However, it has rarely been explored for quantum circuit optimization, further considering physical features of target devices.
Mingyu Chen 0009, Yu Zhang 0086, Yongshang Li, Zhen Wang 0075, Jun Li 0001, Xiang-Yang Li 0001
ICCAD2
2022 PFilter: Building Persistent Maps through Feature Filtering for Fast and Accurate LiDAR-based SLAM
abstract
Simultaneous localization and mapping (SLAM) based on laser sensors has been widely adopted by mobile robots and autonomous vehicles. These SLAM systems are required to support accurate localization with limited computational resources. In particular, point cloud registration, i.e., the process of matching and aligning multiple LiDAR scans collected at multiple locations in a global coordinate framework, has been deemed as the bottleneck step in SLAM. In this paper, we propose a feature filtering algorithm, PFilter, that can filter out invalid features and can thus greatly alleviate this bottleneck. Meanwhile, the overall registration accuracy is also improved due to the carefully curated feature points. We integrate PFilter into the well-established scan-to-map LiDAR odometry framework, F-LOAM, and evaluate its performance on the KITTI dataset. The experimental results show that PFilter can remove about 48.4% of the points in the local feature map and reduce feature points in scan by 19.3% on average, which save 20.9% processing time per frame. In the mean time, we improve the accuracy by 9.4%.
Yifan Duan, Jie Peng 0002, Yu Zhang 0086, Jianmin Ji, Yanyong Zhang
IROS3
2021 Towards an Online RRT-based Path Planning Algorithm for Ackermann-steering Vehicles
abstract
It is challenging to develop an online path planning algorithm for Ackermann-steering vehicles to find collision-free and kinematically-feasible paths, that is efficient for dense environments, adaptable to various environments, and suitable for environments with narrow passages. In this paper, we propose a kinematically constrained RRT-based path planning algorithm integrating with a trajectory parameter space (TP-space) with three novel improvements to meet the above requirements. In specific, we introduce a new way to choose candidate nodes to expand the tree for an RRT-based algorithm, which can significantly increase the success rate of the expansion and improve the efficiency of the algorithm. We also introduce a procedure to incrementally adjust the step size for the expansion, which enables the algorithm to automatically adapt to various environments. At last, we integrate rapidly-exploring random vines (RRV) with a TP-space to handle kinematic constraints and improve the performance of the algorithm to expand the tree through a narrow passage. We also prove that the algorithm is probabilistic complete and asymptotically near-optimal. An ablation study shows that all three improvements can notably improve the performance of the RRT-based path planning algorithm. We also evaluate the algorithm in various environments. The experimental results show that our algorithm achieves competitive performance compared with the state-of-the-art. The source code is available at https://github.com/PengJieb/fastbkrrt.
Jie Peng 0002, Yu'an Chen, Yifan Duan, Yu Zhang 0086, Jianmin Ji, Yanyong Zhang
ICRA4
2021 DRQN-based 3D Obstacle Avoidance with a Limited Field of View
abstract
In this paper, we propose a map-based end-to-end DRL approach for three-dimensional (3D) obstacle avoidance in a partially observed environment, which is applied to achieve autonomous navigation for an indoor mobile robot using a depth camera with a narrow field of view. We first train a neural network with LSTM units in a 3D simulator of mobile robots to approximate the Q-value function in double DRQN. We also use a curriculum learning strategy to accelerate and stabilize the training process. Then we deploy the trained model to a real robot to perform 3D obstacle avoidance in its navigation. We evaluate the proposed approach both in the simulated environment and on a robot in the real world. The experimental results show that the approach is efficient and easy to be deployed, and it performs well for 3D obstacle avoidance with a narrow observation angle, which outperforms other existing DRL-based models by 15.5% on success rate.
Yu'an Chen, Guangda Chen, Lifan Pan, Jun Ma 0034, Yu Zhang 0086, Yanyong Zhang, Jianmin Ji
IROS5
2021 Static Type Inference for Foreign Functions of Python
abstract
Static type inference is an effective way to maintain the safety of programs written in a dynamically typed language. However, foreign functions implemented in another programming language are often outside the inference range. Python, a popular dynamically typed language, has a lot of widely used packages which follow the multilingual structure with C/C++ extension modules. Existing deterministic Python static type inference tools which are not based on type annotations can do nothing about these foreign functions. In this paper, we propose a novel method to infer the type signature of foreign functions by analyzing implicit information in the layer of foreign function interface. We design a static type inference system, its evaluation on CPython, NumPy and Pillow shows that our method soundly infers the number and type of arguments for most foreign functions. Our results can further work as a complement to the state-of-the-art Python static type inference tool and enable it to analyze programs with foreign function calls. We catch 48 bugs of mismatch between foreign function declaration and its implementation, which make a parameter-free foreign function take argument of any type. 8 of the bugs we reported have been confirmed and fixed by communities.
Mingzhe Hu, Yu Zhang 0086, Wenchao Huang 0001, Yan Xiong 0001
ISSRE2
2021 Encouraging Compiler Optimization Practice for Undergraduate Students through Competition
abstract
AI and other emerging applications demand domain-specific architectures which require compiler techniques such as back-end generation for different architectures and optimizations. However, traditional undergraduate compiler courses emphasize the front-end, while code generation and optimization are rarely involved. To motivate universities to have more industry-friendly compiler courses, we have designed a national compiler design competition for undergraduates to include compiler techniques beyond parsing. Moreover, we provided reliable, continuous cloud storage and an online evaluation platform for distributed competitors. In the 9-week Competition in 2020, each team (up to 4 students) was required to implement a compiler for a given SysY language and given target hardware (Raspberry Pi 4B). The performance was evaluated by executing code generated by the compiler on real hardware. Finally,21 of 72 teams successfully passed all functional test cases; 12 of 21 teams implemented optimizations showing significant speedup over gcc -O0; furthermore, compilers of the top 3 teams performed better than gcc -O2 on the given 10 performance test cases. Some advanced optimization techniques, such as multithreading and SIMD, were used by some teams. This paper summarizes the competition and further thoughts on compiler courses.
Yu Zhang 0086, Chunming Hu, Mingliang Zeng, Yitong Huang, Yuanwei Wang
ITiCSE (1)1
2021 Neighbor-Vote: Improving Monocular 3D Object Detection through Neighbor Distance Voting
abstract
As cameras are increasingly deployed in new application domains such as autonomous driving, performing 3D object detection on monocular images becomes an important task for visual scene understanding. Recent advances on monocular 3D object detection mainly rely on the "pseudo-LiDAR'' generation, which performs monocular depth estimation and lifts the 2D pixels to pseudo 3D points. However, depth estimation from monocular images, due to its poor accuracy, leads to inevitable position shift of pseudo-LiDAR points within the object. Therefore, the predicted bounding boxes may suffer from inaccurate location and deformed shape. In this paper, we present a novel neighbor-voting method that incorporates neighbor predictions to ameliorate object detection from severely deformed pseudo-LiDAR point clouds. Specifically, each feature point around the object forms their own predictions, and then the "consensus'' is achieved through voting. In this way, we can effectively combine the neighbors' predictions with local prediction and achieve more accurate 3D detection. To further enlarge the difference between the foreground region of interest (ROI) pseudo-LiDAR points and the background points, we also encode the ROI prediction scores of 2D foreground pixels into the corresponding pseudo-LiDAR points. We conduct extensive experiments on the KITTI benchmark to validate the merits of our proposed method. Our results on the bird's eye view detection outperform the state-of-the-art performance, especially for the "hard" level detection. The code is available at https://github.com/cxmomo/Neighbor-Vote.
Xiaomeng Chu, Jiajun Deng, Yao Li 0016, Zhenxun Yuan, Yanyong Zhang, Jianmin Ji, Yu Zhang 0086
ACM Multimedia7
2021 An Empirical Study for Common Language Features Used in Python Projects
abstract
As a dynamic programming language, Python is widely used in many fields. For developers, various language features affect programming experience. For researchers, they affect the difficulty of developing tasks such as bug finding and compilation optimization. Former research has shown that programs with Python dynamic features are more change-prone. However, we know little about the use and impact of Python language features in real-world Python projects. To resolve these issues, we systematically analyze Python language features and propose a tool named PYSCAN to automatically identify the use of 22 kinds of common Python language features in 6 categories in Python source code. We conduct an empirical study on 35 popular Python projects from eight application domains, covering over 4.3 million lines of code, to investigate the the usage of these language features in the project. We find that single inheritance, decorator, keyword argument, for loops and nested classes are top 5 used language features. Meanwhile different domains of projects may prefer some certain language features. For example, projects in DevOps use exception handling frequently. We also conduct in-depth manual analysis to dig extensive using patterns of frequently but differently used language features: exceptions, decorators and nested classes/functions. We find that developers care most about ImportError when handling exceptions. With the empirical results and in-depth analysis, we conclude with some suggestions and a discussion of implications for three groups of persons in Python community: Python designers, Python compiler designers and Python developers.
Yu Zhang 0086, Mingzhe Hu
SANER2
2021 Quingo: A Programming Framework for Heterogeneous Quantum-Classical Computing with NISQ Features
abstract
The increasing control complexity of Noisy Intermediate-Scale Quantum (NISQ) systems underlines the necessity of integrating quantum hardware with quantum software. While mapping heterogeneous quantum-classical computing (HQCC) algorithms to NISQ hardware for execution, we observed a few dissatisfactions in quantum programming languages (QPLs), including difficult mapping to hardware, limited expressiveness, and counter-intuitive code. In addition, noisy qubits require repeatedly performed quantum experiments, which explicitly operate low-level configurations, such as pulses and timing of operations. This requirement is beyond the scope or capability of most existing QPLs. We summarize three execution models to depict the quantum-classical interaction of existing QPLs. Based on the refined HQCC model, we propose the Quingo framework to integrate and manage quantum-classical software and hardware to provide the programmability over HQCC applications and map them to NISQ hardware. We propose a six-phase quantum program life-cycle model matching the refined HQCC model, which is implemented by a runtime system. We also propose the Quingo programming language, an external domain-specific language highlighting timer-based timing control and opaque operation definition, which can be used to describe quantum experiments. We believe the Quingo framework could contribute to the clarification of key techniques in the design of future HQCC systems.
Xiang Fu 0003, Hanru Jiang, Fucheng Cheng, Yihang Yang, Chunchao Hu, Anqi Huang 0003, Guangyao Huang 0001, Xiaogang Qiang, Mingtang Deng, Ping Xu 0004, Weixia Xu 0001, Wanwei Liu, Yu Zhang 0086, Yuxin Deng 0001, Junjie Wu 0003, Yuan Feng 0001
ACM Trans. Quantum Comput.20
2020 Codar: A Contextual Duration-Aware Qubit Mapping for Various NISQ Devices
abstract
Quantum computing devices in the NISQ era share common features and challenges like limited connectivity between qubits. Since two-qubit gates are allowed on limited qubit pairs, quantum compilers must transform original quantum programs to fit the hardware constraints. Previous works on qubit mapping assume different gates have the same execution duration, which limits them to explore the parallelism from the program. To address this drawback, we propose a Multi-architecture Adaptive Quantum Abstract Machine (maQAM) and a COntext-sensitive and Duration-Aware Remapping algorithm (Codar). The Codar remapper is aware of gate duration difference and program context, enabling it to extract more parallelism from programs and speed up the quantum programs by 1.23 in simulation on average in different architectures and maintain the fidelity of circuits when running on OriginQ quantum noisy simulator.
Haowei Deng, Yu Zhang 0086, Quanxi Li
DAC2
2020 Lightweight Map-Enhanced 3D Object Detection and Tracking for Autonomous Driving
abstract
3D object detection and tracking are crucial to the real-time and accurate perception of the surrounding environment for autonomous driving. Recent approaches on 3D object detection and tracking have made great progress, thanks to the rapid development of deep learning models. Even though these models have achieved superior performance on specific datasets, the actual self-driving systems still cannot deal with real-world driving situations properly, especially in complicated scenarios like road intersections. With the development of vehicle-infrastructure cooperation technology, scene information such as map is considered to have great potential in alleviating these problems. In this paper, we explore the potential of solving corner cases in real driving scenarios through the cooperation between autonomous vehicles and map information. We propose a holistic approach that integrates and utilizes the map information in system following the tracking-by-detection paradigm. In order to ensure that the use of map information does not bring much overhead to detection and tracking, we propose a representation method for concise information extracted from rich map. We show that our framework can improve the detection and tracking accuracy with mild or no increase of latency. Specifically, in some cases, our results demonstrate a MOTA improvement of nearly 2% .
Shunhong Wang, Yu Zhang 0086, Yanyong Zhang, Jianmin Ji
Internetware3
2020 A Close Look at Multi-tenant Parallel CNN Inference for Autonomous Driving
Yitong Huang, Yu Zhang 0086, Boyuan Feng, Yanyong Zhang, Yufei Ding 0001
NPC2
2020 The Python/C API: Evolution, Usage Statistics, and Bug Patterns
abstract
Python has become one of the most popular programming languages in the era of data science and machine learning, especially for its diverse libraries and extension modules. Python front-end with C/C++ native implementation achieves both productivity and performance, almost becoming the standard structure for many mainstream software systems. However, feature discrepancies between two languages can pose many security hazards in the interface layer using the Python/C API. In this paper, we applied static analysis to reveal the evolution and usage statistics of the Python/C API, and provided a summary and classification of its 10 bug patterns with empirical bug instances from Pillow, a widely used Python imaging library. Our toolchain can be easily extended to access different types of syntactic bug-finding checkers. And our systematical taxonomy to classify bugs can guide the construction of more highly automated and high-precision bug-finding tools.
Mingzhe Hu, Yu Zhang 0086
SANER2
2019 Optimizing Quantum Programs Against Decoherence: Delaying Qubits into Quantum Superposition
abstract
Quantum computing technology has reached a second renaissance in the last decade. However, in the NISQ era pointed out by John Preskill in 2018, quantum noise and decoherence, which affect the accuracy and execution effect of quantum programs, cannot be ignored and corrected by the near future NISQ computers. In order to let users more easily write quantum programs, the compiler and runtime system should consider underlying quantum hardware features such as decoherence. To address the challenges posed by decoherence, in this paper, we propose and prototype QLifeReducer to minimize the qubit lifetime in the input OpenQASM program by delaying qubits into quantum superposition. QLifeReducer includes three core modules, i.e., the parser, parallelism analyzer and transformer. It introduces the layered bundle format to express the quantum program, where a set of parallelizable quantum operations is packaged into a bundle. We evaluate quantum programs before and after transformed by QLifeReducer on both real IBM Q 5 Tenerife and the self-developed simulator. The experimental results show that QLifeReducer reduces the error rate of a quantum program when executed on IBMQ 5 Tenerife by 11%; and can reduce the longest qubit lifetime as well as average qubit lifetime by more than 20% on most quantum workloads.
Yu Zhang 0086, Haowei Deng, Quanxi Li, Haoze Song, Leihai Nie
TASE1
2018 A Scalable Pthreads-Compatible Thread Model for VM-Intensive Programs
Yu Zhang 0086, Jiankang Chen
ICA3PP (4)1
2018 Compiler Practice System Integrated with Real Open Source Compiler: (Abstract Only)
abstract
The rapidly growing scale of modern computer systems has been increasing the skill gaps between graduating students and industry expectations. The Compiler Course, as one of the core CS courses, is not only a course introducing the theory and practice of programming language translation, but also a comprehensive course cultivating students/ multidimensional competencies, such as programming, language design, software engineering, communication and collaboration, etc. Writing a compiler for a toy language is a common assignment in many compiler courses. Yet our compiler practice system differs from most of its peers in several aspects: integrated with real open source LLVM compiler, practice of some modern compilation mechanisms, process control and version management (using git), team work, etc. We designed two kinds of projects to integrate the LLVM compiler, one is coding class, e.g. developing an LLVM IR generator and a JIT-compiler on LLVM IR; the other is source code understanding class, e.g. providing guidance and issues for understanding the mechanisms of Clang parser or static analyzer. We also designed some team projects to let students investigate some modern language features and their implementation mechanisms, to discuss within and among teams, and finally to give team presentation. The poster describes the components of our compiler practice system, related practice support package and guidance. Some tradeoffs among difficulty, complexity, time and knowledge points are discussed. So far we have been practiced and improved the practice system for 4 years, and lessons as well as experiences are shared on the poster as well.
Yu Zhang 0086
SIGCSE1
2018 RepassDroid: Automatic Detection of Android Malware Based on Essential Permissions and Semantic Features of Sensitive APIs
abstract
Most current literature on Android malware pays particular attention to the features of applications. Much of them focus on permissions or APIs, neglecting the behavioral semantics of applications, and the literature considering behavioral semantics is often expensive and weak in extendibility. In this paper, we introduce RepassDroid - a relatively coarse-grained but faster tool for automatic Android malware detection. We define Generalized-sensitive API and emphasize on considering if the trigger points of generalized-sensitive APIs are UI-related or not. It analyzes the application by abstracting the generalized sensitive API with its trigger point as the semantic feature, with the addition of Really-essential Permission as the syntax feature. Then it utilizes machine learning to automatically determine whether an application is benign or malicious. We evaluate RepassDroid on 24288 samples in total, 20000 for training and 4288 for test. With the comparative experiments, we find that Random Forest is the optimal classification technique for our feature set, achieving 97.7% accuracy and 0.99 AUC, along with a malware classification precision as high as 99.3%. Our evaluation results confirm that our approach and the feature set are logical and effective for Android malware detection.
Niannian Xie, Fanping Zeng, Xiaoxia Qin, Yu Zhang 0086, Mingsong Zhou, Chengcheng Lv
TASE4
2017 CHAUS: Scalable VM-Based Channels for Unbounded Streaming
Yu Zhang 0086, Yu-Fen Yu, Hui-Fang Cao, Jian-Kang Chen, Qi-Liang Zhang
J. Comput. Sci. Technol.1
2016 Making User-Level VMM for Deterministic Parallelism Nonblocking and Efficient
abstract
Many parallel programs are intended to yield deterministic results, but unpredictable thread or process interleavings can lead to subtle bugs and nondeterminism. We proposed a producer-consumer virtual memory-SPMC-for efficient system-enforced deterministic parallelism, and prototyped the SPMC model and its software stack entirely in Linux user space, called DLinux. This paper summarizes the implementation policies and limitations in our previous DLinux. To reduce SPMC page fault overhead and suspend/resume overhead which severely degrade the performance of DLinux, we enhance the SPMC model with nonblocking test and direct read and write primitives. Based on the extended SPMC model, we improve the implementation of upper programming abstractions. Experimental results show that relative to the previous version, the new DLinux can improve the performance of NPB workloads up to 2.33X and 1.76X on 8 and 16 processes, respectively. For CG on 8 processes, its runtime relative to MPICH2 decreases from 4.12X to 1.77X.
Yu Zhang 0086, Jiange Zhang, Qiliang Zhang
PDCAT1
2015 Lightweight Function Pointer Analysis
Yu Zhang 0086
ISPEC2
2015 System-Enforced Deterministic Streaming for Efficient Pipeline Parallelism
Yu Zhang 0086, Zhaopeng Li, Hui-Fang Cao
J. Comput. Sci. Technol.1
2015 Program equivalence in linear contexts
Yuxin Deng 0001, Yu Zhang 0086
Theor. Comput. Sci.2
2014 A temporal programming model with atomic blocks based on projection temporal logic
Xiaoxiao Yang, Yu Zhang 0086, Ming Fu, Xinyu Feng 0001
Frontiers Comput. Sci.2
2013 The Buffered π-Calculus: A Model for Concurrent Languages
Xiaojie Deng, Yu Zhang 0086, Yuxin Deng 0001, Farong Zhong
LATA2
2013 A Shape Graph Logic and A Shape System
Zhaopeng Li, Yu Zhang 0086
J. Comput. Sci. Technol.2
2012 A Concurrent Temporal Programming Model with Atomic Blocks
Xiaoxiao Yang, Yu Zhang 0086, Ming Fu, Xinyu Feng 0001
ICFEM2
2012 Exploring Deterministic Shared Memory Programming Model
abstract
Deterministic parallelism promises many benifits for parallel programming. Exist deterministic runtimes, however, either require modifying programs to use a restricted set of synchronization primitives, or limit scalability by relying on a centralized, deterministic thread scheduler. We addressed these challenges with single producer multi-consumer (SPMC) virtual memory as a possible foundation for scalable deterministic parallelism. The SPMC foundation avoids the introduction of read/write and write/write data races, and permits peer-topeer “dialog” short-circuiting the process/thread hierarchy. In this paper, we explore high level runtime-DetSM-atop the SPMC foundation to support deterministic shared memory parallel programming. We introduce a generalized shared memory unit or SMU abstraction to represent memory deterministically shared among threads, and a specific DetArray abstraction to allow concurrent threads updating disjoint continuous elements in an array, with certain programmability and efficiency tradeoffs. Preliminary results suggest that DetSM atop the SPMC maybe realistic and useful, achieving better run-time than Dthreads for dedup, and better performance than the original Determinator for quick sort and merge sort.
Yu Zhang 0086
PDCAT1
2010 Reasoning about Optimistic Concurrency Using a Program Logic for History
Ming Fu, Xinyu Feng 0001, Zhong Shao 0001, Yu Zhang 0086
CONCUR5
2010 Just-in-Time Compiler Assisted Object Reclamation and Space Reuse
Yu Zhang 0086, Lina Yuan, Tingpeng Wu, Wen Peng, Quanlong Li
NPC1
2010 Formal verification of concurrent programs with read-write locks
Ming Fu, Yu Zhang 0086
Frontiers Comput. Sci. China2
2010 Formal Reasoning About Lazy-STM Programs
Yu Zhang 0086, Ming Fu
J. Comput. Sci. Technol.2
2009 Verifying Anonymous Credential Systems in Applied Pi Calculus
Xiangxi Li, Yu Zhang 0086, Yuxin Deng 0001
CANS2
2009 Formal Reasoning about Concurrent Assembly Code with Reentrant Locks
abstract
This paper focuses on the problem of reasoning about concurrent assembly code with reentrant locks. Our verification technique is based on concurrent separation logic (CSL). In CSL, locks are treated as non-reentrant locks and each lock is associated with a resource invariant, the lock-protected resources are obtained and released through acquiring and releasing the lock respectively. In order to accommodate for reentrancy, we introduce some additional notions into our specification language to describe reentrant level for each acquiring and releasing lock operation. Keeping track of the reentrant level for each lock in the pre- and post- conditions enables the program logic to ensure that resources are not reacquired upon reentrancy, thus resources owned by a thread are prevented from reintroducing in the postcondition. Our framework is fully mechanized. Its soundness has been verified using the Coq proof assistant. We demonstrate the usage of our framework through giving a safety proof of a simple program.
Ming Fu, Yu Zhang 0086
TASE2
2009 Certifying Concurrent Programs Using Transactional Memory
Yu Zhang 0086
J. Comput. Sci. Technol.2
2008 A Decision Procedure for XPath Satisfiability in the Presence of DTD Containing Choice
Yu Zhang 0086, Yihua Cao, Xunhao Li
APWeb1