Sen Ma

dblp:84/7485 · DBLP profile ↗
← Back
12ranked-venue papers
5as first author
1since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 8 · 4 first-author · 1 since 2021Software engineering, systems software and programming languages · 2 · 1 first-authorSecurity and privacy · 1Applied, interdisciplinary, general and emerging computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
3 papers
Distributed systems · 63% Reconfigurable computing and FPGAs · 14% Energy-efficient computing · 12%
Network and information security
1 paper
Blockchain and cryptocurrency security · 100%
Databases, data mining, and information retrieval
1 paper
Transaction processing and concurrency control · 100%

Topics — the 10 heaviest of 12, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Blockchain and cryptocurrency security › blockchain scalability
sharding
1.012026
Secure and Efficient Read-Write Synchronization in Re-Sharding Via Lightweight Global State Tree · IEEE Trans. Computers 2026
Distributed systems › fault tolerance
byzantine fault tolerance
1.012026
Secure and Efficient Read-Write Synchronization in Re-Sharding Via Lightweight Global State Tree · IEEE Trans. Computers 2026
Distributed systems
consensus
1.012026
Secure and Efficient Read-Write Synchronization in Re-Sharding Via Lightweight Global State Tree · IEEE Trans. Computers 2026
Parallel and multicore computing
domain-specific language
0.212016
Just In Time Assembly of Accelerators · FPGA 2016
Reconfigurable computing and FPGAs
FPGA accelerator
0.212016
Just In Time Assembly of Accelerators · FPGA 2016
Energy-efficient computing
clock gating
0.212014
On energy efficiency and amdahl's law in FPGA based chip heterogeneous multiprocessor systems (abstract only) · FPGA 2014
Energy-efficient computing
power management
0.212014
On energy efficiency and amdahl's law in FPGA based chip heterogeneous multiprocessor systems (abstract only) · FPGA 2014
Compilers and program optimization › hardware compilation
high-level synthesis
0.112016
Just In Time Assembly of Accelerators · FPGA 2016
Performance modeling and evaluation › parallel system performance › speedup modeling
amdahl's law
0.112014
On energy efficiency and amdahl's law in FPGA based chip heterogeneous multiprocessor systems (abstract only) · FPGA 2014
Performance modeling and evaluation › parallel system performance
speedup modeling
0.112014
On energy efficiency and amdahl's law in FPGA based chip heterogeneous multiprocessor systems (abstract only) · FPGA 2014

Methods — techniques the papers use, named apart from their topics

non-blocking coordination · 3.0global state tree · 3.0runtime interpretation · 0.5partial reconfiguration · 0.5runtime measurement · 0.2
YearPublicationVenuePosition
2026 Secure and Efficient Read-Write Synchronization in Re-Sharding Via Lightweight Global State Tree
abstract
State re-sharding can reduce cross-shard transaction ratios, which improves the scalability of blockchain systems. However, unavoidable cross-shard transactions and account-locking mechanisms can lead to security risks (read-write conflicts) and performance bottlenecks (low synchronization efficiency). Therefore, this paper proposes a secure and efficient read-write synchronization model for cross-shard transactions in blockchain state re-sharding via a lightweight Global State Tree ($\mathcal{GT}$). The model consists of intra-shard and inter-shard state consistency modules. The intra-shard module includes two methods: account state read-write and account record update. The former allows local shard committees to track account state changes and prevent the use of expired account states, while the latter incorporates account records within maximum latency into an account state data structure, thereby enhancing the traceability and verification efficiency of update history. In the inter-shard module, a transaction processing method with a global takeover mechanism is proposed during the re-sharding window. By using validated data in the$\mathcal{GT}$, the method achieves non-blocking global coordination and account reallocation. Experimental results demonstrate that the proposed model increases transaction throughput and reduces transaction latency compared to bases under Byzantine conditions.
Peiyun Zhang, Sen Ma, Qinglin Zhao, Haibin Zhu 0001
IEEE Trans. Computers3
2020 Sliding-Window Based Batch Forwarding using Intra-Flow Random Linear Network Coding
abstract
Batch forwarding using intra-flow random linear network coding (RLNC) has been used to improve the performance of a wireless network constituent of lossy links. However, existing batch-based forwarding mechanisms in this aspect can lead to a lot of bandwidth waste and thus reduced transmission efficiency. In this paper, we design a Sliding WIndow based Multiple batch forwarding mechanism (SWIM) using RLNC. In SWIM, multiple batches are allowed to be sent out simultaneously in a way that the forwarding process is managed by a sliding window. In SWIM, adaptive rate assignment is used to assign bandwidth resources to different batches based on their decoding states at the destination, in order to make full use of the bandwidth resources. Simulation results show that SWIM can achieve improved throughput performance as compared with existing work.
Sen Ma, Xiulian Liu, Yan Yan 0009, Baoxian Zhang, Jun Zheng 0002
IWCMC1
2019 Learning-Assisted Optimization in Mobile Crowd Sensing: A Survey
abstract
Mobile crowd sensing (MCS) is a relatively new paradigm for collecting real-time and location-dependent urban sensing data. Given its applications, it is crucial to optimize the MCS process with the objective of maximizing the sensing quality and minimizing the sensing cost. While earlier studies mainly tackle this issue by designing different combinatorial optimization algorithms, there is a new trend to further optimize MCS by integrating learning techniques to extract knowledge, such as participants' behavioral patterns or sensing data correlation. In this paper, we perform an extensive literature review of learning-assisted optimization approaches in MCS. Specifically, from the perspective of the participant and the task, we organize the existing work into a conceptual framework, present different learning and optimization methods, and describe their evaluation. Furthermore, we discuss how different techniques can be combined to form a complete solution. In the end, we point out existing limitations, which can inform and guide future research directions.
Jiangtao Wang 0001, Yasha Wang, Daqing Zhang 0001, Jorge Gonçalves 0001, Denzil Ferreira, Aku Visuri, Sen Ma
IEEE Trans. Ind. Informatics7
2018 CoBOT: static C/C++ bug detection in the presence of incomplete code
abstract
To obtain precise and sound results, most of existing static analyzers require whole program analysis with complete source code. However, in reality, the source code of an application always interacts with many third-party libraries, which are often not easily accessible to static analyzers. Worse still, more than 30% of legacy projects [1] cannot be compiled easily due to complicated configuration environments (e.g., third-party libraries, compiler options and macros), making ideal "whole-program analysis" unavailable in practice. This paper presents CoBOT [2], a static analysis tool that can detect bugs in the presence of incomplete code. It analyzes function APIs unavailable in application code by either using function summarization or automatically downloading and analyzing the corresponding library code as inferred from the application code and its configuration files. The experiments show that CoBOT is not only easy to use, but also effective in detecting bugs in real-world programs with incomplete code. Our demonstration video is at: https://youtu.be/bhjJp3e7LPM.
Sen Ma, Sihao Shao, Yulei Sui, Fuyao Duan, Shikun Zhang
ICPC2
2016 Run time interpretation for creating custom accelerators
Sen Ma, Zeyad Aklah, David Andrews 0001
DATE1
2016 Just In Time Assembly of Accelerators
abstract
Despite the significant advancements that have been made in High Level Synthesis, the reconfigurable computing community has failed at getting programmers to use Field Programmable Gate Arrays (FPGAs). Existing barriers that prevent programmers from using FPGAs include the need to work within vendor specific CAD tools, knowledge of hardware programming models, and the requirement to pass each design through synthesis, place and route. In this paper we present a new approach that takes these barriers out of the design flows for programmers. Synthesis is eliminated from the application programmers path by becoming part of the initial coding process when creating the programming patterns that define a Domain Specific Language. Programmers see no difference between creating software or hardware functionality when using the DSL. A run time interpreter is introduced that assembles hardware accelerators within a configurable tile array of partially reconfigurable slots at run time. Initial results show the approach allows hardware accelerators to be compiled 100x faster compared to the time required to synthesize the same functionality. Initial performance results further show a compilation/interpretation approach can achieve approximately equivalent performance for matrix operations and filtering compared to synthesizing a custom accelerator.
Sen Ma, Zeyad Aklah, David Andrews 0001
FPGA1
2015 A run time interpretation approach for creating custom accelerators
abstract
The world of software development has the notion of just-in-time compilation, run time binary translation, and language interpretation. These dynamic run time techniques support increased code portability and designer productivity. There are no such equivalences to increase the productivity or portability of creating new hardware components within Field Programmable Gate Arrays (FPGAs). Instead, creating a new hardware component requires hardware design skills and the overhead of running through synthesis, place and route. If a change is made to even a single line of code, the synthesis, place and route steps must be repeated. In this paper we present a new approach that allows hardware accelerators to be built and run using compilation and run time interpretation. Our results show the approach can enable software programmers without any hardware skills to create hardware accelerators at productivity levels consistent with software development and compilation. The same accelerator can be compiled 100× faster than synthesis. Even though the approach is focused on productivity, our observed performance results are promising. Our initial application test cases show the same accelerator written by a software programmer and synthesized through Vivado HLS or written using our DSL and compiled within our approach achieves equivalent performance.
Sen Ma, Zeyad Aklah, David Andrews 0001
FPL1
2014 A Hierarchical Memory Architecture with NoC Support for MPSoC on FPGAs
abstract
This work presents a memory hierarchy with the support of network-on-chip (NoC) for MPSoC systems. The memory hierarchy consists of a shared global memory and private local memories as shown in Figure 1. Each core in the system is equipped with two local memories, one for instructions and one for data. The MicroBlaze soft core used in this work connects the main bus through the PLB interface and connects the local memory modules through the LMB interface. Further it connects to a 4x4 mesh NoC through the FSL interface, as shown in Figure 2(a). We built the generic NoC (NoC-g) using the open-source router designed by the Concurrent VLSI Architecture group at the Stanford University [2]. Each router has 5 input ports and 5 output ports. Each input physical channel and each output physical channel is connected to 4 input virtual channels and 4 output virtual channels, respectively. The 40 virtual channels are connected to an internal crossbar switch for routing. We designed the adapter to connect the MicroBlaze processor to the router.
Miaoqing Huang, Hongyuan Ding, Sen Ma
FCCM4
2014 On energy efficiency and amdahl's law in FPGA based chip heterogeneous multiprocessor systems (abstract only)
abstract
This poster presents our preliminary findings on the relationship between speedup and energy efficiency on FPGA based Chip Heterogeneous Multiprocessor Systems (CHMPs). While researchers have investigated how to tailor combinations of heterogeneous compute engines within a CHMP system to best meet the performance needs of specific applications, exploring how these optimized architectures also effect energy efficiency is not as well studied. We show that a simple relationship exists between the speedup these systems gain and their associated energy efficiency. We show that the simple relationship between Amdahl's law and energy efficiency. All the experiments result achieved through actual run time measurements on homogeneous and heterogeneous multiprocessor systems implemented within a Xilinx Virtex6 FPGA. We further show how a systems with 6 MicroBlaze soft processors' dynamic power and hence the overall energy efficiency of the system can be effected through transparent operating system control of the compute resources. We also present how to use clock gating to control the dynamic power consumption for each processor and with this careful power-aware management unit, the system's dynamic power consumption can follow the requirements of each application.
Sen Ma, David Andrews 0001
FPGA1
2014 Achieving portability and efficiency over chip heterogeneous multiprocessor systems
abstract
Emerging programming models for chip heterogeneous multiprocessor (CHMP) systems elevate architecture details up into the source code. This eliminates portability and requires designers to navigate a multidimensional search space when trying to optimize designs. In this paper, we present an approach that reinstates portability through a combination of polymorphic functions and an adaptive runtime system. Together they enable runtime profiling and dynamic scheduling of unaltered source code across systems with different combinations of heterogeneous resources. Our results verify the ability of our programming model and runtime system to re-enable the notion of writing code once and run anywhere. Runtime results show how runtime tuning can increase resource utilization and provide performance increases as the number and heterogeneity of computing resources increases.
Eugene Cartwright, Alborz Sadeghian, Sen Ma, David Andrews 0001
FPL3
2012 Automating the design of mLUT MPSoPC FPGAs in the cloud
abstract
Modern platform FPGAs are over the million-LUT level, large enough to support complete heterogeneous Multiprocessor System-On-Chips (MPSoCs). Constructing systems with 10's of processors is currently feasible using existing manual methods within vendor-specific CAD tools. However these manual, by-hand, approaches will not be feasible for constructing future systems with 100's to 1,000's of processors. Instead, new automated system assembly approaches will be required to handle these levels of system complexity and diversity. In this paper we present a new automated design flow for creating such next generation heterogeneous MPSoCs. An integral part of the MPSoPC system created is the inclusion of a general purpose PThreads-compliant HW/SW co-designed operating system and heterogeneous compiler. Our design flow has been placed in the cloud and is freely accessible across the Internet.
Eugene Cartwright, Azad Fahkari, Sen Ma, Christina Smith, Miaoqing Huang, David Andrews 0001, Jason Agron
FPL3
2009 Deadlock Detection Based on Resource Allocation Graph
abstract
Deadlock occurs randomly and is difficult to detect, it always has a negative impact on the effective execution of operating system. This paper uses the principle of adjacency matrix, path matrix and strongly-connected component of simple directed graph in graph theory, gives a model of detecting deadlock by exploring strongly-connected component from resource allocation graph. The experiment shows that it can detect resources and processes involved in deadlock effectively by this detection method. The paper provides a new idea for the research of operating system algorithms, and a new way for auxiliary teaching and practical engineering.
Qinqin Ni, Weizhen Sun, Sen Ma
IAS3