EDBT 2026 Demo / reviewers in the wild / expert
Yuan Dai
dblp:125/6076
· DBLP profile ↗
27ranked-venue papers
12as first author
24since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 21 · 10 first-author · 18 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 first-author · 3 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Computer networks · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Lora: Towards Improved Applicability of Reconfigurable Architecture for Versatile Nonlinear Functions
Yuan Dai, Guibin Zou, Yuanda Yang, Jiahang Lou, Yiwen Luo, Xinyu Cai, Wenbo Yin, Wai-Shing Luk, Lingli Wang |
ISCA | 1 |
| 2026 | Live Demonstration: An Agile FPGA-Overlayed CGRA SoC for High-Efficiency Computing
Jiahang Lou, Jianrong Zhang, Yuan Dai, Zewei Zhong, Wenbo Yin, Lingli Wang |
ISCAS | 3 |
| 2026 | MOE: An Efficient Multicasting and One-hot Encoding Hybrid Configuration Compression Technique for CGRAs
Yuan Dai, Wenbo Yin, Lingli Wang |
ISCAS | 2 |
| 2026 | Dependency-Aware Data Parallelism on Spatial CGRA via Constraint Satisfaction and Graph ColoringabstractCoarse-grained Reconfigurable Architecture (CGRA) is a competitive accelerator architecture for computation-intensive loop kernels. Spatial CGRA is a typical CGRA that performs all the operations spatially to reduce reconfiguration costs within a single iteration, demanding high data parallelism. To achieve this goal, one of the main challenges is the loop-carried dependency between memory accesses. Many existing CGRA compilers struggle to precisely analyze the dependency distance, especially when accesses involve complex address patterns. Consequently, these compilers often default to setting the distance to one, based on a worst-case assumption, leading to degraded performance. However, we observe that a precise distance can improve performance significantly, raising the requirement for an efficient distance calculation approach. Another challenge is the performance constraints of single-bank memory, which necessitate the designer partitioning the original data into a multi-bank memory. However, we observe that the mapping result can cause the inter-iteration conflict, thereby invalidating the memory partition scheme. Therefore, an efficient post-mapping conflict detection is required. In this paper, we develop a constraint satisfaction problem (CSP)-based approach for calculating dependency distance and detecting conflicts, which determines the maximum available dependency distance and identifies conflicts within both intra- and inter-iterations. Besides, we formulate access scheduling as a graph coloring problem, which can minimize conflicts and improve performance. Overall, we develop a comprehensive end-to-end framework with architectural and compiler support for efficient data parallelism on spatial CGRA. We conduct extensive experiments to systematically evaluate the impact of different approaches on performance and compilation. Evaluation results show that our architecture can achieve 13.16× and 1.19× (up to 1.68×) average performance improvements compared to a RISC-V CPU and a state-of-the-art CGRA SoC, respectively. Besides, our architecture has 7.38× and 1.18× (up to 1.65×) average energy efficiency gains compared to these two architectures. Yuan Dai, Xuchen Gao, Wenbo Yin, Wai-Shing Luk, Lingli Wang |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2026 | Toward Efficient Edge AI With Heterogeneous Computing and Multilevel OptimizationabstractThe rapid progress of artificial intelligence (AI) has brought increasing demands on hardware accelerators, particularly as modern models combine dense linear operations with a growing number of irregular, nonlinear, and control-intensive operators. While tensor cores and systolic arrays offer high throughput for regular computations, they often struggle to efficiently support the diverse operations emerging in recent model structures. Coarse-grained reconfigurable arrays (CGRAs), with their spatial parallelism and reconfigurability, may serve as a natural complement to dense accelerators in such heterogeneous workloads. In this work, we propose EUREKA, a heterogeneous acceleration framework that integrates tensor cores with CGRAs through a unified instruction set, cross-architecture data scheduling, tailored hardware support for nonlinear operators, and optimizations at the instruction, task, and operator levels to exploit parallelism. At the software level, we introduce a hierarchical compilation strategy that combines graph-level optimizations with tensor-level scheduling techniques. To address the large design space of hardware–software co-optimization, we further develop a Bayesian optimization-based exploration scheme enhanced with kernel compression methods, which provides an efficient means of identifying promising hardware configurations and scheduling strategies. Experiment results on representative AI benchmarks show that EUREKA improves execution efficiency, achieving an average$12.6\times $normalized performance gain over state-of-the-art frameworks. Jingyuan Li 0003, Xinyu Cai, Yuan Dai, Wenbo Yin, Lingli Wang |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2025 | Towards Efficient Data Parallelism on Spatial CGRA via Constraint Satisfaction and Graph ColoringabstractCoarse-Grained Reconfigurable Architecture (CGRA) is a competitive accelerator architecture for computation-intensive loop kernels. Spatial CGRA is a typical CGRA that performs all the operations spatially, demanding high data parallelism. Given the performance limitations of single-bank memory, partitioning original data into multi-bank memory within the spatial CGRA is favored. However, we observe that the mapping result can cause the inter-iteration conflict, thereby invalidating the memory partition scheme. Yuan Dai, Xuchen Gao, Bingbing Peng, Wenbo Yin, Wai-Shing Luk, Lingli Wang |
ASP-DAC | 1 |
| 2025 | Adora Compiler: End-to-End Optimization for High-Efficiency Dataflow Acceleration and Task Pipelining on CGRAsabstractTo fully harness emerging computing architectures, compilers must provide intuitive input handling alongside powerful code optimization to unlock maximum performance. Coarse-Grained Reconfigurable Arrays (CGRAs) — highly energy-efficient for nested-loop applications — have lacked a compiler capable of meeting these objectives. This paper introduces the Adora compiler [1], which effectively bridges user-friendly, lightweight coding inputs with high-performance acceleration on the CGRA SoC. Adora utilizes CGRA-target loop transformations to achieve efficient data-flow level execution while optimizing data communication and task pipelining at the task-flow level. Additionally, it incorporates a comprehensive automated algorithm with a thoughtfully designed optimization sequence. A series of comprehensive experiments highlights the exceptional efficiency and scalability of the Adora compiler, demonstrating its transformative impact in leveraging CGRA capabilities for acceleration in edge computing. Jiahang Lou, Qilong Zhu, Yuan Dai, Zewei Zhong, Wenbo Yin, Lingli Wang |
DAC | 3 |
| 2025 | COFFA: A Co-Design Framework for Fused-Grained Reconfigurable Architecture Towards Efficient Irregular Loop HandlingabstractCoarse-Grained Reconfigurable Architecture (CGRA) emerges as a competitive accelerator due to its high flexibility and energy efficiency. However, most CGRAs are effective for computation-intensive applications with regular loops but struggle with irregular loops containing control flows. These loops introduce fine-grained logic operations and are costly to execute by coarse-grained arithmetic units in CGRA. Efficiently handling such logic operations necessitates incorporating Boolean algebra optimization, which can improve logic density and reduce logic depth. Unfortunately, no previous research has incorporated it into the compilation flow to support irregular loops efficiently.We proposeCOFFA, an open-source framework for heterogeneous architecture with a RISC-V CPU and a fused-grained reconfigurable accelerator, which integrates coarse-grained arithmetic and fine-grained logic units, along with flexible IO units and distributed interconnects. As a software/hardware co-design framework,COFFAhas a powerful compiler that extracts and optimizes fine-grained logic operations from irregular loops, performs coarse-grained arithmetic and memory optimizations, and offloads the loops to the accelerator.Across various challenging benchmarks with irregular loops,COFFAachieves significant performance and energy efficiency improvements over an in-order, an out-of-order RISC-V CPUs, and a recent FPGA, respectively. Moreover, compared with the state-of-the-art CGRAUE-CGRAandHycube,COFFAcan achieve 2.5× and 3.5× performance gains, respectively. Yuan Dai, Xuchen Gao, Yunhui Qiu, Jingyuan Li 0003, Yuhang Cao, Yiqing Mao, Sichao Chen, Wenbo Yin, Wai-Shing Luk, Lingli Wang |
IEEE Trans. Computers | 1 |
| 2025 | Soil Moisture Affects Multitemporal InSAR Deformation Monitoring via Dielectric Property ChangesabstractSynthetic Aperture Radar Interferometry (InSAR) is utilized to evaluate slope stability, revealing pronounced periodic oscillations in the deformation time series. Although such periodic patterns have conventionally been ascribed to stratified tropospheric delays or seasonal precipitation in prior studies, periodic signals persist in the deformation results even after atmospheric phase removal by a linear iterative model. This observation underscores the limitations of conventional deformation interpretations and prompts further investigation into the underlying physical mechanisms. To address this, an interferometric phase correction model accounting for variations in surface dielectric property has been proposed. This model effectively removes phase delays induced by dielectric property changes from the raw interferometric phase, enabling the extraction of linear trend signals that reflect actual surface deformation. To verify the reliability of the correction model, Sentinel-1A (C-band) and TerraSAR-X (X-band) radar data is employed to quantify deformation patterns in the slope of non-sliding section around the Huangnibazi landslide. The analysis consistently identifies periodic deformation signals in both SAR datasets after mitigating atmospheric influences. Our findings indicate that dielectric property changes constitute a critical factor in InSAR deformation monitoring that cannot be overlooked. The observed periodic fluctuations in deformation time series are attributed to dielectric-induced phase modulation effects. Furthermore, systematic correlation analyses confirm a strong coherence between soil moisture variations and deformation fluctuations, with minimal temporal hysteresis. In contrast, seasonal precipitation exhibits a weaker correlation with deformation and longer hysteresis time. These results robustly support the theoretical framework linking soil moisture, dielectric property, penetration depth, and phase delay. This insight holds significant implications for enhancing the effectiveness of InSAR technology in slope disaster monitoring and early warning systems. Meng Ao, Xiangben Zhang, Yuan Dai, Lianhuan Wei, Xiaosong Feng, Shanjun Liu, Mingsheng Liao, Lu Zhang 0034, Cristiano Tolomei |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2025 | MoDAF: A Multi-objective Divide-and-Conquer Parameter Tuning Framework for CGRAsabstractCoarse-grained reconfigurable architectures (CGRAs) are gaining increasing attention as domain-specific accelerators due to their high flexibility and energy efficiency. These architectures offer a compelling solution for applications that require custom hardware performance while retaining a degree of programmability. However, the design space of CGRAs is inherently vast and complex, presenting significant challenges for architects to explore design choices efficiently and systematically. Existing design space exploration (DSE) methodologies for CGRAs are often time-demanding and struggle to deliver optimal solutions when confronted with high-dimensional and multi-objective design space. Therefore, we consider constructing a CGRA parameter tuning framework called MoDAF. MoDAF initializes the design space using the most representative and diverse samples. It adopts a divide-and-conquer approach, utilizing Monte Carlo Tree Search (MCTS) and space partitioning techniques to dynamically break down the complex design space into more manageable subspaces. A hybrid model handles local fluctuations within each subspace, while a dual sampling algorithm is designed to increase sampling efficiency. MoDAF also incorporates a fast evaluation model to estimate CGRA throughput and area, significantly speeding up the exploration process. Compared with previous approaches, experiments show that our proposed framework reduces the average distance from the reference set by 53.0% and the hypervolume deviation by 64.2%, while also cutting wall time by 57.5%. Jingyuan Li 0003, Yuan Dai, Wenbo Yin, Lingli Wang |
ACM Trans. Design Autom. Electr. Syst. | 2 |
| 2024 | MDCRA: A Reconfigurable Accelerator Framework for Multiple Dataflow LanesabstractCoarse-grained reconfigurable architecture (CGRA) is a type of reconfigurable computing architecture suitable for emerging applications that require dynamic compilation hardware. However, the resource utilization of existing CGRA is low due to the lack of flexibility across varied application granularity. In this paper, we propose a CGRA framework for multiple dataflow lanes (MDCRA). It supports post-silicon computational granularity adjustments. Evaluated with Polybench, Machsuite and Express, the speedup of MDCRA is$24.83\times$higher than CPU CVA6, and$2.08\times$higher than vector processor Ara. Compared with TRAM and DSAGEN, MDCRA achieves an area reduction of 27% and 47% respectively with the same speedup. Besides, compared with OpenCGRA, the average utilization of function units is improved by 20.05%. Shaoyang Sun, Boyin Jin, Jiahang Lou, Yuhang Cao, Jingyuan Li 0003, Yuan Dai, Wenbo Yin, Wai-Shing Luk, Lingli Wang |
ASAP | 8 |
| 2024 | A CGRA Front-end Compiler Enabling Extraction of General Control and Dedicated OperatorsabstractCoarse-grained reconfigurable architecture (CGRA) gradually becomes an extraordinarily promising accelerator due to its flexibility and power efficiency. However, most CGRA front-end compilers focus on the innermost body of regular loops with a pure data flow. Therefore, we propose CO-Compiler, an LLVM-based CGRA front-end compiler to generate an optimized control-data flow graph (CDFG), which can handle versatile loops in C/C++, including general control flow, arbitrary nested levels, and imperfect statements. Then we extract multi-dimension memory access patterns and various dedicated operators adapting to concrete hardware functions. In addition, we analyze variable loop bounds which are settled at runtime, and realize the SoC runtime configuration of CGRA. The feasibility of our methodology is verified by a RISC-V based SoC simulation. The experimental results demonstrate that our dedicated operator extraction can reduce 43% PE resources and decrease 84% initiation interval (II) on a TRAM architecture. Furthermore, compared with state-of-the-art (SOTA) CGRA front-end compilers, CO-Compiler has the highest 88.1% success rate in CDFG generation for a wide range of benchmarks. Moreover, by using the same back-end mappers, our work can reach 78% reduction for II and $2.06\times$ PE spatio-temporal utilization in contrast with their own front-end compilers. Xuchen Gao, Yunhui Qiu, Yuan Dai, Wenbo Yin, Lingli Wang |
ASPDAC | 3 |
| 2024 | MemFormer: A memory based unified model for anomaly detection on metro railway tracks
Ruikang Liu, Weiming Liu 0003, Mengfei Duan, Wei Xie 0021, Yuan Dai, Xianzhe Liao |
Expert Syst. Appl. | 5 |
| 2024 | Background subtraction for video sequence using deep neural network
Yuan Dai, Long Yang 0001 |
Multim. Tools Appl. | 1 |
| 2024 | HierCGRA: A Novel Framework for Large-scale CGRA with Hierarchical Modeling and Automated Design Space ExplorationabstractCoarse-grained reconfigurable arrays (CGRAs) are promising design choices in computation-intensive domains, since they can strike a balance between energy efficiency and flexibility. A typical CGRA comprises processing elements (PEs) that can execute operations in applications and interconnections between them. Nevertheless, most CGRAs suffer from the ineffectiveness of supporting flexible architecture design and solving large-scale mapping problems. To address these challenges, we introduce HierCGRA, a novel framework that integrates hierarchical CGRA modeling, Chisel-based Verilog generation, LLVM-based data flow graph (DFG) generation, DFG mapping, and design space exploration (DSE). With the graph homomorphism (GH) mapping algorithm, HierCGRA achieves a faster mapping speed and higher PE utilization rate compared with the existing state-of-the-art CGRA frameworks. The proposed hierarchical mapping strategy achieves 41× speedup on average compared with the ILP mapping algorithm in CGRA-ME. Furthermore, the automated DSE based on Bayesian optimization achieves a significant performance improvement by the heterogeneity of PEs and interconnections. With these features, HierCGRA enables the agile development for large-scale CGRA and accelerates the process of finding a better CGRA architecture. Sichao Chen, Su Zheng, Guowei Zhu, Jingyuan Li 0003, Yazhou Yan, Yuan Dai, Wenbo Yin, Lingli Wang |
ACM Trans. Reconfigurable Technol. Syst. | 8 |
| 2024 | HETA: A Heterogeneous Temporal CGRA Modeling and Design Space Exploration via Bayesian OptimizationabstractDue to its high energy efficiency and flexibility, coarse-grained reconfigurable architecture (CGRA) has gained increasing attention. Temporal CGRA is a typical category of CGRA that supports single-cycle context switching and time-multiplexing hardware resources to perform spatial and temporal computations. Although multiple temporal CGRAs have been proposed, an architecture with rich design parameters and heterogeneous modeling is still lacking. To this end, we propose a highly parameterized heterogeneous temporal CGRA, called HETA. However, the highly parameterized and heterogeneous design introduces a challenging design space for manual exploration. To address this challenge, we introduce a Bayesian-optimization (BO)-based design space exploration (DSE) of homogeneous and heterogeneous architectures. Different from other DSE processes that require defining the heterogeneous exploration strategy, our approach adopts a searching-pruning-based method without manual intervention. To improve the efficiency of DSE, we develop a fast statistic model for area evaluation, whose error is below 1%. In addition, a pipeline mapping (PiPMap) algorithm is developed to alleviate the restrictions caused by data synchronization and unleash the potential of the proposed architecture. Experimental results show that HETA can achieve 89%, 52%, and 47% improvement in throughput, area efficiency, and energy efficiency over the neighbor-to-neighbor (N2N)-based interconnect CGRA, respectively. Compared with the Switch-based interconnect CGRA, HETA’s area efficiency is increased by 61%. Furthermore, compared with the homogeneous architecture of HETA, the optimized heterogeneous architecture improves area efficiency and energy efficiency by 14.7% and 4.8%, respectively. Yuan Dai, Jingyuan Li 0003, Qilong Zhu, Yunhui Qiu, Yihan Hu 0003, Wenbo Yin, Lingli Wang |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2023 | UPTRA: An Ultra-Parameterized Temporal CGRA Modeling and OptimizationabstractTemporal Coarse-Grained Reconfigurable Architecture (CGRA) is a typical category of CGRA that supports single-cycle context switching and time-multiplexing hardware resources to perform both spatial and temporal computations. Compared with the spatial CGRA, it can be used in area and power budget-constrained scenarios, with the sacrifice of the throughput. Therefore, achieving minimum Initialization Interval (II) for higher throughput is the main objective in many works for temporal CGRA mapping. Yuan Dai, Yunhui Qiu, Qilong Zhu, Jingyuan Li 0003, Wenbo Yin, Lingli Wang |
FCCM | 1 |
| 2023 | PRAD: A Bayesian Optimization-based DSE Framework for Parameterized Reconfigurable Architecture DesignabstractCoarse-Grained Reconfigurable Architecture (CGRA) is a domain-specific reconfigurable architecture. Generally, the CGRA architecture consists of IO, memory, coarse-grained processing element (PE), and interconnect. Usually, ALU in PE contains a relatively complete set of operations and most of the interconnects adopt neighbor-to-neighbor (N2N) [1], switch-based [2], and combination of the connection box and switch box (CB-SB) patterns [3]. However, the complex operation sets and switch-based/CB-SB fully-connected interconnects provide sufficient reconfigurability at the cost of resource overhead. Thus, it is important to build a parameterized architecture of CGRA to achieve a balance among hardware overhead, flexibility and performance through automatic design space exploration (DSE). Bingbing Peng, Shaoyang Sun, Yuan Dai, Jingyuan Li 0003, Yunhui Qiu, Kaihang Wang, Wenbo Yin, Lingli Wang |
FCCM | 3 |
| 2023 | STEV: Stretchable Triboelectric E-skin enabled Proprioceptive Vibration Sensing for Soft RobotabstractVibration perception is essential for robotic sensing and dynamic control. Nevertheless, due to the rigorous demand for sensor conformability and stretchability, enabling soft robots with proprioceptive vibration sensing remains challenging. This paper proposes a novel liquid metal-based stretchable e-skin via a kirigami-inspired design to enable soft robot proprioceptive vibration sensing. The e-skin is fabricated into 0.1mm ultrathin thickness, ensuring its negligible influence on the overall stiffness of the soft robot. Moreover, the working mechanism of the e-skin is based on the ubiquitous triboelectrification effect, which transduces mechanical stimuli without external power supply. To demonstrate the practicability of the e-skin, we built a soft gripper consisting of three soft robotic fingers with proprioceptive vibration sensing. Our experiment shows that the gripper can accurately distinguish the grain category (six grains with the same mass, 99.9% accuracy) and the packaging quality (100% accuracy) by simply shaking the gripped bottle. In summary, a soft robotic proprioceptive vibration sensing solution is proposed; it helps soft robots to have a more comprehensive awareness of their self-state and may inspire further research on soft robotics. Kai-Chong Lei, Huaze Tang, Shoujie Li, Yuan Dai, Wenbo Ding 0001, Xiao-Ping Zhang 0002 |
ICRA | 5 |
| 2023 | Skin-Integrated Haptic Interfaces Enabled by Scalable Mechanical Actuators for Virtual RealityabstractThe very recent concept of metaverse highlights the importance of virtual reality (VR) and augmented reality (AR), which associates with a wide variety of applications in entertainment, medical treatment, and human–machine interfaces. The current VR/AR technologies mainly rely on visual interaction, while immersive experience in VR and AR highly demands sensational feedback, such as haptic and temperature with noticeable quality in wearable or even skin-integrated formarts. In this article, we report a wearable and flexible haptic interface based on electromagnetic vibrotactile actuators with high wearability and stability. By adopting double layers of copper (Cu) coils at the top and bottom of the magnetic disc, an enhanced electromagnetic field can be generated. Additionally, the intensity of the haptic feedback can be modulated according to sensed pressure in the virtual world by adjusting the value of power input and frequency. The actuator exhibits high stability and tolerance upon environmental, cyclic, and impact resistance tests. Finally, the actuators are developed into the soft VR interfaces for mounting on forearms, fingers, and hands to verify their superiority over conventional haptic actuators in the aspects of performance and applications. Chun Ki Yiu, Xingcan Huang, Wooyoung Park, Jingyou Su, Jingkun Zhou, Tsz Hung Wong, Kuanming Yao, Pu Fan, Yuan Dai, Zhengbao Yang, Xinge Yu |
IEEE Internet Things J. | 16 |
| 2022 | TRAM: An Open-Source Template-based Reconfigurable Architecture Modeling FrameworkabstractCoarse-grained reconfigurable architecture (CGRA) is a promising accelerator design choice due to its high performance and power efficiency in the computation or data-intensive application domains, such as security, multimedia, digital signal processing, machine learning, and high-performance computing. CGRA consists of coarse-grained processing elements (PEs) and interconnects that determine the architecture flexibility to support different applications and also affect the performance and power efficiency significantly. Although multiple types of interconnects have been proposed, a parameterized unified model is still lacking. In this paper, we propose a flexible and scalable CGRA template with a novel interconnect model that can unify the typical neighbor-to-neighbor, switch-based, and FPGA-like interconnects. Furthermore, we present TRAM, an open-source template-based reconfigurable architecture modeling framework that integrates the Chisel-based CGRA modeling, architecture intermediate representation (IR) and Verilog generation, dataflow graph (DFG) mapping, simulation, and evaluation. The mapping flow contains graph-based placement and routing, critical-path-driven data synchronization, and simulated-annealing-based optimization. We evaluate the impacts of the rich design parameters, which demonstrate the significance of such a flexible template to facilitate architecture optimization. Compared with the related work, TRAM can achieve a 4.1× smaller DFG latency and a faster mapping speed for both the 8×8 and 16×16 CGRAs. Moreover, TRAM is able to attain an extremely high PE utilization of 94.4 % on average by architecture tuning. Yunhui Qiu, Yuhang Cao, Yuan Dai, Wenbo Yin, Lingli Wang |
FPL | 3 |
| 2022 | Detecting moving object from dynamic background video sequences via simulating heat conduction
Yuan Dai, Long Yang 0001 |
J. Vis. Commun. Image Represent. | 1 |
| 2022 | CSR-Net: Learning Adaptive Context Structure Representation for Robust Feature CorrespondenceabstractFeature matching, which refers to identifying and then corresponding the same or similar visual pattern from two or more images, is a key technique in any image processing task that requires establishing good correspondences between images. Given potential correspondences (matches) in two scenes, a novel whole-part deep learning framework, termed as Context Structure Representation Network (CSR-Net), is designed to infer the probabilities of arbitrary correspondences being inliers. Traditional approaches commonly build the local relation between correspondences by manually engineered criteria. Different from existing attempts, the main idea of our work is to learn explicitly neighborhood structure of each correspondence, allowing us to formulate the matching problem into a dynamic local structure consensus evaluation in an end-to-end fashion. For this purpose, we propose a permutation-invariant STructure Representation (STR) learning module, which can easily merge different types of networks into a unified architecture to deal with sparse matches directly. By the collaborative use of STR, we introduce a Context-Aware Attention (CAA) mechanism to adaptively re-calibrate structure features via a rotation-invariant context aware encoding and simple feature gating, thus arising the ability of fine-grained patterns recognition. Moreover, to further weaken the cost of establishing reliable correspondences, the CSR-Net is formulated as whole-part consensus learning, where the aim of whole level is compensating rigid transformations. In order to demonstrate our CSR-Net can effectively boost the baselines, we intensively experiment on image matching and other visual tasks. The results of the experiment confirm that the matching performances of CSR-Net have significantly improved over nine state-of-the-art competitors. Jiaxuan Chen 0002, Shuang Chen 0008, Xiaoxian Chen, Yuan Dai, Yang Yang 0032 |
IEEE Trans. Image Process. | 4 |
| 2021 | APIR-DSP: An approximate PIR-DSP architecture for error-tolerant applicationsabstractIn error-tolerant applications such as low-precision DNNs and digital filters, approximate arithmetic circuits can significantly reduce hardware resource utilization. In this work we propose an embedded block for field-programmable gate arrays, called APIR-DSP, which incorporates an approximate 9×9 hard multiplier based on the PIR-DSP architecture to improve speed and reduce area. In addition, a DSP unit evaluation platform based on Yosys and VPR which packs multiply accumulate operations into DSP blocks is developed. Using this tool we synthesis designs from Verilog implementations of matrix multiplication in DeepBench and the DoReFaNet low-precision neural network and show that APIR-DSP significantly reduces DSP resources and improves hardware utilization and performance compared with the Xilinx DSP48E2 embedded block. Compared with exact multiplication, it is shown that accuracy loss is optimized with the SNR of an FIR filter being reduced by 1.03 dB. For DNNs, accuracy loss for AlexNet is 0.31% on CIFAR10 dataset and no accuracy loss for LeNet on MNIST dataset is observed. Synthesis results show that the APIR-DSP enjoys an area reduction of 21.60%, critical path reduction of 4.85% and power consumption is reduced by 2.80%, compared with PIR-DSP. Yuan Dai, Hao Zhou 0008, Seyedramin Rasoulinezhad, Philip H. W. Leong, Lingli Wang |
FPT | 1 |
| 2016 | Graphics processing unit-accelerated joint-bitplane belief propagation algorithm in DSC
Yuan Dai, Yong Fang 0001, Long Yang 0001, Gwanggil Jeon |
J. Supercomput. | 1 |
| 2014 | Accelerating 2D orthogonal matching pursuit algorithm on GPU
Yuan Dai, Dongjian He, Yong Fang 0001, Long Yang 0001 |
J. Supercomput. | 1 |
| 2013 | Parallel design for error-resilient entropy coding algorithm on GPU
Yuan Dai, Yong Fang 0001, Dongjian He, Bormin Huang |
J. Parallel Distributed Comput. | 1 |