EDBT 2026 Demo / reviewers in the wild / expert
Junqi Yuan
dblp:161/4748
· DBLP profile ↗
6ranked-venue papers
2as first author
1since 2021 · last 2024
0000-0002-7996-6574ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 5 · 2 first-authorArtificial intelligence and machine learning · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
1 paper |
Electronic design automation · 80% Reconfigurable computing and FPGAs · 20% |
Topics — the 5 heaviest of 5, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Reconfigurable computing and FPGAs
FPGA compilation |
0.4 | 1 | 2019 | ARBSA: Adaptive Range-Based Simulated Annealing for FPGA Placement · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2019 |
Electronic design automation › physical design › placement › circuit placement
FPGA placement |
0.4 | 1 | 2019 | ARBSA: Adaptive Range-Based Simulated Annealing for FPGA Placement · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2019 |
Electronic design automation
physical design |
0.4 | 1 | 2019 | ARBSA: Adaptive Range-Based Simulated Annealing for FPGA Placement · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2019 |
Electronic design automation › physical design
placement |
0.4 | 1 | 2019 | ARBSA: Adaptive Range-Based Simulated Annealing for FPGA Placement · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2019 |
Electronic design automation
simulated annealing |
0.4 | 1 | 2019 | ARBSA: Adaptive Range-Based Simulated Annealing for FPGA Placement · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2019 |
Methods — techniques the papers use, named apart from their topics
range-limiting strategy · 0.4adaptive range-based simulated annealing · 0.4
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Unveiling the potential of progressive training diffusion model for defect image generation and recognition in industrial processes
Yalin Wang 0003, Zexiong Zhou, Xujie Tan, Yuqing Pan, Junqi Yuan, Zhifeng Qiu, Chenliang Liu |
Neurocomputing | 5 |
| 2019 | ARBSA: Adaptive Range-Based Simulated Annealing for FPGA PlacementabstractPlacement has always been the most time-consuming part of the field programmable gate array (FPGA) compilation flow. Conventional simulated annealing has been unable to keep pace with ever increasing sizes of designs and FPGA chip resources. Without utilizing information of the circuit topology, it relies on large amounts of random swap operations, which are time-costly. This paper proposes an adaptive range-based algorithm to improve the behavior of swap operations and limit the swap distances by introducing the concept of range-limiting strategy for nets. It avoids unnecessary design space exploration, and thus can converge to near-optimal solutions much more quickly. The experimental results are based on the Titan benchmarks, which contain 4K to 30K blocks, including logic array blocks, inputs and outputs, digital signal processors, and random access memories. This approach achieves$2.82\boldsymbol \times $speed up, 4.8% reduction on wire length, 4.1% improvement on critical path compared with the SA from VTR with wire length-driven optimization, and$1.78\boldsymbol \times $speed up, 10% reduction on wire length, 2% reduction on critical path with path timing-driven optimization. It also manifests better scalability on larger benchmarks. Junqi Yuan, Jialing Chen, Lingli Wang, Xuegong Zhou, Yinshui Xia |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2017 | RBSA: Range-based simulated annealing for FPGA placementabstractPlacement has always been the most time-consuming part in the FPGA compilation flow. Traditional simulated annealing has been unable to keep pace with ever increasing sizes of designs and FPGA chip resources. Without utilizing information of the circuit topology, it relies on large amounts of random swap operations, which are time-costly. This paper proposes a range-based algorithm to improve the behavior of swap operations and limit the swap distances by introducing the concept of range limiting for nets. It avoids unnecessary design space exploration, and thus can converge to near-optimal solutions much more quickly. The Titan benchmarks we have tested on contains 4K to 30K blocks, which include LABs, IOs, DSPs and RAMs. This approach achieves 2.05X speed up on average compared with the SA from VTR while preserving the placement quality of both the wire length and critical path. It also manifests better scalability towards larger benchmarks. Junqi Yuan, Lingli Wang, Xuegong Zhou, Yinshui Xia |
FPT | 1 |
| 2017 | Lossless Compression Decoders for Bitstreams and Software Binaries Based on High-Level SynthesisabstractAs the density of field-programmable gate arrays continues to increase, the size of configuration bitstreams grows accordingly. Compression techniques can reduce memory size and save external memory bandwidth. To accelerate the configuration process and reduce the software startup time, four open-source lossless compression decoders developed using high-level synthesis techniques are presented. Moreover, in order to balance the objectives of compression ratio, decompression throughput, and hardware resource overhead, various improvements and optimizations are proposed. Full bitstreams and software binaries have been collected as a benchmark, and 33 partial bitstreams have also been developed and integrated into the benchmark. Evaluations of the synthesizable compression decoders are demonstrated on a Xilinx ZC706 board, showing higher decompression throughput than those of the existing lossless compression decoders using our benchmark. The proposed decoders can reduce software startup time by up to 31.23% in embedded systems and 69.83% reduction of reconfiguration time for partial reconfigurable systems. Jian Yan 0002, Junqi Yuan, Philip H. W. Leong, Wayne Luk, Lingli Wang |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2016 | Connect on the fly: Enhancing and prototyping of cycle-reconfigurable modulesabstractThis paper introduces cycle-reconfigurable modules that enhance FPGA architectures with efficient support for dynamic data accesses: data accesses with accessed data size and location known only at runtime. The proposed module adopts new reconfiguration strategies based on dynamic FIFOs, dynamic caches, and dynamic shared memories to significantly reduce configuration generation and routing complexity. We develop a prototype FPGA chip with the proposed cycle-reconfigurable module in the SMIC 130-nm technology. The integrated module takes less than the chip area of 39 CLBs, and reconfigures thousands of runtime connections in 1.2 ns. Applications for large-scale sorting, sparse matrix-vector multiplication, and Memcached are developed. The proposed modules enable 1.4 and 11 times reduction in area-delay product compared with those applications mapped to previous architectures and conventional FPGAs. Xinyu Niu, Junqi Yuan, Lingli Wang, Wayne Luk |
FPL | 3 |
| 2014 | Design space exploration for FPGA-based hybrid multicore architectureabstractThis paper presents a parameterized system-level design framework, which enables rapid and powerful research for hybrid multicore architecture exploration and hardware/software co-design. The framework comprises the component-based hardware design and application compiler, which make it easy for a designer to build stream-oriented applications with FPGA-based hybrid multicore architectures. The high modularity and parameterization of the framework supports fast multicore architecture exploration of different topologies, routing schemes, processor types, customized hardware processing units and memory system organizations. The compiler tool chain is used to map C/C++ based applications onto the soft processing units. Experimental results targeting the JPEG encoding application demonstrate the feasibility and performance improvement of this framework. Jian Yan 0002, Junqi Yuan, Ying Wang 0032, Philip H. W. Leong, Lingli Wang |
FPT | 2 |