Weijie Fang

dblp:261/4609 · DBLP profile ↗
← Back
8ranked-venue papers
2as first author
8since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 5 · 2 first-author · 5 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Computer networks · 1 · 1 since 2021Security and privacy · 1 · 1 since 2021
YearPublicationVenuePosition
2026 HoPart: Hop-Constrained Partitioning with Routing Support for Multi-FPGA Systems
abstract
Multi-FPGA platforms are indispensable for VLSI emulation and prototyping, but remain fundamentally constrained by limited inter-FPGA I/O bandwidth. Techniques such as time-division multiplexing and FPGA hopping partially alleviate this bottleneck but substantially increase partitioning and routing complexity and exacerbate timing closure. As modern FPGA-based applications impose stringent timing budgets, design flows must be explicitly delay-aware. In this paper, we present HoPart, a Hop-constrained partitioning approach that enforces per-path hop limits. A core ingredient of our approach is the joint optimization of path delay and congestion during partitioning. In addition, we propose a routing algorithm that adaptively adjusts the number of edges (hops) along each path based on real-time criticality metrics. This strategy reduces interconnect resource usage on non-critical paths while minimizing delay on timing-critical ones. Extensive experiments on public benchmark suites demonstrate that HoPart reduces maximum path delay by up to 30% compared with the state-of-the-art MaPart, while maintaining efficient utilization of inter-FPGA interconnect.
Longkun Guo, Weijie Fang
DATE3
2026 Near-Optimal TDM Ratio Assignment for Die-Level Routing in Multi-FPGA Systems
abstract
Modern multi-FPGA systems often integrate multiple dies to expand logic capacity and address the increasing complexity of integrated circuit designs. To overcome the limitations of physical I/O pins, these systems typically employ time-division multiplexing (TDM) technology. However, higher TDM ratios introduce considerable signal delays, resulting in higher critical connection delays. This paper focuses on optimizing the TDM ratio to tackle this challenge. We formulate the TDM ratio assignment problem as a block-angular convex program and solve it using Lagrangian decomposition, obtaining a (1 + ϵ)approximate solution for any given ϵ > 0. We further introduce a delay-aware TDM wire assignment scheme to achieve efficient signal assignment. Experimental results demonstrate that our method enables efficient, high-quality die-level routing in modern multi-FPGA systems, achieving up to 10.8% reduction in critical connection delay compared to the state-of-the-art approaches.
Longkun Guo, Weijie Fang
DATE3
2026 Exploiting Partial JPEG Decoding to Mitigate On-Device Image Processing
abstract
Device-cloud collaborative inference is often necessary for resource-constrained IoT devices that cannot support full on-device models. To minimize bandwidth and support concurrency, existing methods typically compress images before transmission. However, these approaches often ignore the significant overhead of decoding native JPEG camera output, especially for high-resolution frames. Our measurements show that the on-device (Raspberry Pi 4B) decoding overhead for 700KB JPEG format is$\sim$14.4x the latency of on-cloud (GeForce RTX 3090) meter recognition inference. To reduce on-device decoding overhead, we design DC Camera, which is built upon a JPEG camera and leverages partial JPEG decoding to efficiently extract DC features from high-resolution images, significantly mitigating on-device image processing overhead. These DC features can preserve structural information better than conventional downsampled images. We utilize DC Camera to implement fast meter recognition system and deploy the system in material science laboratory to monitor multiple meters. Our evaluation demonstrates that compared to state-of-the-art (SOTA) methods, DC Camera can reduce on-device computation overhead by$\sim$5.8x and decrease transmission volume by$\sim$90.9x, without inference accuracy degradation.
Kaijie Gong, Hao Wang 0238, Yi Gao 0001, Weijie Fang, Wei Dong 0001
IEEE Trans. Mob. Comput.5
2026 LP-Based Area Assignment for Length-Matching Routing of Complex Multilayer PCBs with Any-Direction Wires
abstract
Any-direction wires and complex obstacle environments in emerging Printed Circuit Board (PCB) routing applications impose new challenges for length matching. To address these challenges, we develop a suite of area assignment approaches that leverage network flow theory and Linear Programming (LP) techniques. First, we partition the initial PCB region into a set of subregions and propose an LP formulation to model the area assignment of subregions to the wires in sparse PCB layouts. We then refine the LP formulation to handle dense PCB scenarios effectively. To enhance the algorithm’s performance further, we introduce utility constraints that consider complex obstacles and propose an Integer Linear Programming (ILP) model to optimize the subregion area assignment. Given the high computational complexity of solving the ILP model, we develop a combinatorial algorithm based on minimum-weight hierarchical flow to efficiently tackle the area assignment problem. By leveraging LP primal-dual techniques, we demonstrate that the proposed algorithm achieves near-optimal solutions even under upper bound constraints, and we further extend the approach to accommodate lower bound constraints. Notably, our algorithm allows the modification of the wire topology to generate better area assignment solutions. In addition, the proposed methodology is extensible to length-matching tasks in multilayer PCB designs. Lastly, we conduct extensive experiments to evaluate the effectiveness of our approach, demonstrating significant performance improvements over existing state-of-the-art methods, particularly in complex PCB routing scenarios.
Longkun Guo, Weijie Fang
ACM Trans. Design Autom. Electr. Syst.3
2025 Acceleration of Timing-Aware Gate-Level Logic Simulation Through One-Pass GPU Parallelism
abstract
Witnessing the advancements in the scale and complexity of chip design, along with the benefits from high-performance computing technologies, the simulation of Very Large Scale Integration (VLSI) circuits increasingly demands acceleration through parallel computing with GPU devices. However, conventional parallel strategies fail to fully leverage modern GPU capabilities, introducing new challenges in GPU-based parallelism for VLSI simulations despite previous demonstrations of significant acceleration. In this paper, we propose a novel approach for accelerating the simulation of 4-value logic timing-aware gate-level circuits through waveform-based GPU parallelism. Our approach introduces an innovative strategy that effectively manages task dependencies during the parallelism of combinational circuits, significantly reducing the synchronization requirement between CPU and GPU. The proposed approach achieves one-pass parallelism by requiring only a single round of data transfer. Moreover, to address the implementation challenges associated with our strategy on GPU devices, we have developed and optimized a series of data structures that dynamically allocate and store newly generated outputs of uncertain scale. Finally, we conduct experiments on industrial-scale open-source benchmarks to demonstrate our approach’s performance gains over several state-of-the-art baselines.
Weijie Fang, Yanggeng Fu, Jiaquan Gao, Longkun Guo, Gregory Z. Gutin, Xiaoyan Zhang 0001
IEEE Trans. Computers1
2024 Obstacle-Aware Length-Matching Routing for Any-Direction Traces in Printed Circuit Board
abstract
Emerging applications in Printed Circuit Board (PCB) routing impose new challenges on automatic length matching, including adaptability for any-direction traces with their original routing preserved for interactiveness. The challenges can be addressed through two orthogonal stages: assign non-overlapping routing regions to each trace and meander the traces within their regions to reach the target length. In this paper, mainly focusing on the meandering stage, we propose an obstacle-aware detailed routing approach to optimize the utilization of available space and achieve length matching while maintaining the original routing of traces. Furthermore, our approach incorporating the proposed Multi-Scale Dynamic Time Warping (MSDTW) method can also handle differential pairs against common decoupled problems. Experimental results demonstrate that our approach has effective length-matching routing ability and compares favorably to previous approaches under more complicated constraints.
Weijie Fang, Longkun Guo, Silu Xiong, Jianli Chen
DAC1
2024 FaceObfuscator: Defending Deep Learning-based Privacy Attacks with Gradient Descent-resistant Features in Face Recognition
Shuaifan Jin, He Wang 0005, Zhibo Wang 0001, Jiahui Hu 0001, Zhongjie Ba, Weijie Fang, Shuhong Yuan, Kui Ren 0001
USENIX Security Symposium9
2021 EBRB cascade classifier for imbalanced data via rule weight updating
Yanggeng Fu, Hong-Yun Huang, Ying-Ming Wang 0001, Wenxi Liu, Weijie Fang
Knowl. Based Syst.6