Di Mou

dblp:303/6463 · DBLP profile ↗
← Back
4ranked-venue papers
1as first author
4since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 1 first-author · 4 since 2021
YearPublicationVenuePosition
2025 Decentralized Control of the Decoupled Modular Multi-Active-Bridge Converters for Modular Scalability
abstract
Modular multi-active-bridge (MMAB) converters have emerged as promising solutions for efficient power conversion in integrating various distributed energy resources, energy storage systems, and loads. However, existing studies predominantly rely on centralized controllers, which limit modular scalability due to the lack of software modularity. This also poses significant computational burdens for the centralized controllers. To address these challenges, this paper proposes a decentralized control method based on a decoupled MMAB converter. The decoupling is achieved by eliminating the inductance from one port, thereby simplifying the control complexity. Based on the decoupled MMAB converter, a decentralized control scheme is proposed, which is composed of a PI controller for frequency synchronization and a proportional controller for direct phase-shift regulation. This combination ensures accurate module synchronization and rapid transient responses. A detailed parameter design is obtained using small-signal analysis. Additionally, a thorough inductance design methodology is provided to maintain the phase shift within a stable operating range. Simulation results validate the effectiveness and superior dynamic performance of the proposed decentralized control strategy.
Kai Sun 0004, Di Mou, Adrià Junyent-Ferré
IECON3
2024 DISC: Exploiting Data Parallelism of Non-Stencil Computations on CGRAs via Dynamic Iteration Scheduling
abstract
Memory partitioning is commonly used to enhance data parallelism of Coarse-grain reconfigurable arrays (CGRAs), typically targeting stencil computation with regular memory access patterns. However, many important workloads, such as linear algebra and signal processing, include non-stencil computation with irregular memory access patterns where memory partitioning is not feasible, leading to memory access conflicts and poor data parallelism. In this paper, we propose a Dynamic Iteration Scheduling CGRA (DISC) that can dynamically exploit data parallelism from non-stencil computations. Via dynamic scheduling on loop iterations, DISC can select conflict-free iterations from an iteration buffer for parallel data access, while holding data dependence. To further enhance the ability to find conflict-free iterations, dynamic data reuse is also introduced to reduce the number of memory references. The experimental results show that DISC can achieve 1.41× performance and 2.75× energy efficiency while consuming much less area and power overhead, as compared to dynamic-scheduling CGRA which supports dynamic operator scheduling.
Yue Liang 0004, Di Mou, Dajiang Liu
ICCAD2
2024 SC-CGRA: An Energy-Efficient CGRA Using Stochastic Computing
abstract
Stochastic Computing (SC) offers a promising computing paradigm for low-power and cost-effective applications, with the added advantage of high error tolerance. In parallel, Coarse-Grained Reconfigurable Arrays (CGRA) prove to be a highly promising platform for domain-specific applications due to their combination of energy efficiency and flexibility. Intuitively, introducing SC to CGRA would significantly reinforce the strengths of both paradigms. However, existing SC-based architectures often encounter inherent computation errors, while the stochastic number generators employed in SC result in exponentially growing latency, which is deemed unacceptable in CGRA. In this work, we propose an SC-based CGRA by replacing the exact multiplication in traditional CGRA with an SC-based multiplication. To improve the accuracy of SC and shorten the latency of Stochastic Number Generators (SNG), we introduce the leading zero shifting and comparator truncation, while keeping the length of bitstream fixed. In addition, due to the flexible interconnections among PEs, we propose a quality scaling strategy that combines neighbor PEs to achieve high-accuracy operations without switching costs like power-gating. Compared to the state-of-the-art approximate computing design of CGRA, our proposed CGRA can averagely achieve a 65.3% reduction in output error while having a 21.2% reduction in energy consumption and a noteworthy 28.37% area savings.
Di Mou, Dajiang Liu
IEEE Trans. Parallel Distributed Syst.1
2023 DARIC: A Data Reuse-Friendly CGRA for Parallel Data Access via Elastic FIFOs
abstract
Coarse-Grained Reconfigurable Arrays (CGRAs) are a promising architecture for data-intensive applications. For parallel data accesses, uniform memory partitioning is usually introduced to CGRA for better pipelining performance. However, uniform memory partitioning not only suffers from a local minimum, but also introduces non-negligible overhead for banking function, which may greatly degrade the performance of CGRA. To this end, this paper introduces non-uniform memory partitioning and proposes a data-reuse-friendly CGRA (DARIC). With well elaborated configurable bank groups cooperated with register chains, elastic FIFOs can be achieved for non-uniform memory partitioning. Based on the resource graph of DARIC, a mapping algorithm supporting path sharing is proposed. Finally, the experimental results show that DARIC can achieve 2.35 × throughput and 2.59 × energy efficiency while having even less area and power overhead, as compared to the state-of-the-art.
Dajiang Liu, Di Mou, Yan Zhuang 0003, Jiaxing Shang, Shouyi Yin
DAC2