Pao-Ann Hsiung

dblp:94/6830 · DBLP profile ↗
← Back
86ranked-venue papers
34as first author
7since 2021 · last 2025
0000-0002-3639-1467ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 39 · 13 first-author · 2 since 2021Software engineering, systems software and programming languages · 27 · 15 first-authorApplied, interdisciplinary, general and emerging computing · 6 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 4 · 4 since 2021Artificial intelligence and machine learning · 3 · 1 first-author · 2 since 2021Computer networks · 3 · 3 first-authorSecurity and privacy · 1Human-computer interaction and ubiquitous computing · 1Theory of computation · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Theoretical computer science
6 papers
Automated reasoning and model checking · 88% Mathematical optimization · 12%
Computer architecture, parallel and distributed computing, and storage systems
4 papers
Embedded and real-time systems · 53% Electronic design automation · 32% Parallel and multicore computing · 15%
Software engineering, system software, and programming languages
5 papers
Program verification · 57% Requirements engineering and software design · 24% Concurrent programming · 19%

Topics — the 16 heaviest of 21, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Automated reasoning and model checking
model checking
0.542014
Accelerating Coverage Estimation Through Partial Model Checking · IEEE Trans. Computers 2014
Model Checking Prioritized Timed Systems · IEEE Trans. Computers 2012
Counterexample-Guided Assume-Guarantee Synthesis through Learning · IEEE Trans. Computers 2011
Mathematical optimization › submodular optimization
coverage problem
0.212014
Accelerating Coverage Estimation Through Partial Model Checking · IEEE Trans. Computers 2014
Automated reasoning and model checking › model checking › real-time model checking
timed automata model checking
0.112012
Model Checking Prioritized Timed Systems · IEEE Trans. Computers 2012
Automated reasoning and model checking › compositional verification
assume-guarantee reasoning
0.112011
Counterexample-Guided Assume-Guarantee Synthesis through Learning · IEEE Trans. Computers 2011
Automated reasoning and model checking
compositional verification
0.112011
Counterexample-Guided Assume-Guarantee Synthesis through Learning · IEEE Trans. Computers 2011
Requirements engineering and software design
model-driven engineering
0.112007
Model Checking Safety-Critical Systems Using Safecharts · IEEE Trans. Computers 2007
Electronic design automation › hardware verification and test
hardware verification
0.112014
Accelerating Coverage Estimation Through Partial Model Checking · IEEE Trans. Computers 2014
Program verification › system verification
embedded software verification
0.012004
VERTAF: An Application Framework for the Design and Verification of Embedded Real-Time Software · IEEE Trans. Software Eng. 2004
Program verification
model checking
0.012004
VERTAF: An Application Framework for the Design and Verification of Embedded Real-Time Software · IEEE Trans. Software Eng. 2004
Embedded and real-time systems
component-based design
0.012004
VERTAF: An Application Framework for the Design and Verification of Embedded Real-Time Software · IEEE Trans. Software Eng. 2004
Embedded and real-time systems › embedded software
embedded software design
0.012004
VERTAF: An Application Framework for the Design and Verification of Embedded Real-Time Software · IEEE Trans. Software Eng. 2004
Parallel and multicore computing › dataflow computing › dataflow scheduling
quasi-static scheduling
0.012004
VERTAF: An Application Framework for the Design and Verification of Embedded Real-Time Software · IEEE Trans. Software Eng. 2004
Embedded and real-time systems
real-time scheduling
0.012004
VERTAF: An Application Framework for the Design and Verification of Embedded Real-Time Software · IEEE Trans. Software Eng. 2004
Electronic design automation › hardware verification and test › formal verification
timed automata
0.012012
Model Checking Prioritized Timed Systems · IEEE Trans. Computers 2012
Program verification
modular verification
0.012002
Efficient and User-Friendly Verification · IEEE Trans. Computers 2002
Embedded and real-time systems › critical systems
safety-critical systems
0.012007
Model Checking Safety-Critical Systems Using Safecharts · IEEE Trans. Computers 2007

Methods — techniques the papers use, named apart from their topics

partial model checking · 0.4mutation analysis · 0.4learning · 0.4DBM subtraction · 0.3DBM merging · 0.3l* learning · 0.2counterexample elimination · 0.2timed automata · 0.2difference bound matrix · 0.1difference bound matrices · 0.1binary decision diagrams · 0.1binary decision diagram · 0.1model checking · 0.1UML · 0.1state-graph merging and reduction · 0.1group theory · 0.1
YearPublicationVenuePosition
2025 SegMAgNet: A Comparative Study of Segmentation Models for Defect Detection in Additively Manufactured Mg Alloys
abstract
Magnesium (Mg) alloys are increasingly used in biomedical applications due to their biodegradability, biocompatibility, and mechanical compatibility with human bone. However, fabricating these alloys using additive manufacturing techniques like Selective Laser Melting (SLM) is challenging due to their high reactivity and the potential for internal defects such as gas porosity and lack of fusion. This study presents a comprehensive evaluation of deep learning-based image segmentation methods for detecting such defects in Mg-based alloys fabricated via SLM at 80 W laser power and 500 mm/s scanning speed. X-ray Computed Tomography (XCT) was used to scan the printed samples, and 83 non-identically distributed image slices were selected to ensure variation in defect morphology. Two types of annotations—manual (using LabelImg) and threshold-based (using ImageJ)—were used to prepare datasets. These datasets were analyzed using two models: a custom U-Net with five-fold cross-validation and a U-Net with ResNet-50 backbone enhanced by patchify augmentation. The threshold-based dataset combined with U-Net + ResNet50 achieved the highest segmentation accuracy with an IoU of 86%, outperforming manual annotation-based training. This study emphasizes the importance of annotation strategy and model selection in defect detection workflows for reactive metal additive manufacturing.
Ayush Pratap, P. Karthikeyan 0004, Neha Sardana, Pao-Ann Hsiung
AVSS6
2025 LightMOT: Lightweight and anchor-free solution for tracking multiple objects in dense populations
P. Karthikeyan 0004, Yong-Hong Liu, Pao-Ann Hsiung
Future Gener. Comput. Syst.3
2024 ACNNE: An Adaptive Convolution Engine for CNNs Acceleration Exploiting Partial Reconfiguration on FPGAs
abstract
In this work, we propose a dynamical function exchange Convolutional Neural Networks (CNN) accelerator architecture named Adaptive CNN engine (ACNNE) that can reconfigure specific convolution layer hardware blocks according to model parameters at runtime. We mainly focus on exploiting reconfigurability for inferencing large-scale CNN on resource- constrained FPGAs. The proposed ACNNE can accelerate the process of convolution layers based on a nested-loop algorithm while a data buffering scheme is presented to reduce the iterations of memory accesses. As a study case, the VGG16 model was implemented on Xilinx ZU3EG MPSoC that can achieve 21.09 giga operations per second in the frequency of 100 MHz with a device resource utilization of 25% to 35%. Experiments show a single-image inference can be completed in 2.5 seconds with an average power consumption of 2.23 W, corresponding to a power efficiency of 9.46 GOPS/W.
Chun-Hsian Huang, Shao-Wei Tang, Pao-Ann Hsiung
ISCAS3
2023 Take Expert Advice Judiciously: Combining Groupwise Calibrated Model Probabilities with Expert Predictions
abstract
Training the machine learning (ML) models require a large amount of data, still the capacity of these models is limited. To enhance model performance, recent literature focuses on combining ML models’ predictions with that of human experts, a setting popularly known as the human-in-the-loop or human-AI teams. Human experts can complement the ML models as they are well-equipped with vast real-world experience and sometimes have access to private information that may not be accessible while training the ML model. Existing approaches for combining an expert and ML model either require end-to-end training of the combined model or require expert annotations for every task. End-to-end training further needs a custom loss function and human annotations, which is cumbersome, results in slower convergence, and may adversely impact the ML model’s accuracy. On the other hand, using expert annotations for every task is also cost-ineffective. We propose a novel technique that optimizes the cost of seeking the expert’s advice while utilizing the ML model’s predictions to improve accuracy. Our model considers two intrinsic parameters: the expert’s cost for each prediction and the misclassification cost of the combined human-AI model. Further, we present the impact of group-wise calibration on the combined model that improves the overall model’s performance. Experimental results on our combined model with group-wise calibration show a significant increase in accuracy with limited expert advice against different established ML models for the image classification task. In addition, the combined model’s accuracy is always greater than that of the ML model, irrespective of the expert’s accuracy, the expert’s cost, and the misclassification cost.
Shweta Jain 0002, Shashi Shekhar Jha, Pao-Ann Hsiung, Ming-Hung Wang
ECAI4
2023 depthUNet: A Dehazing Model with Adaptive Depth Attention for Natural Images
abstract
Most image dehazing deep learning models target synthetic datasets of hazy images, resulting in not considering features in natural hazy images. Leveraging on depth attention with adaptation, we propose a novel dehazing network called depthUNet, that is focused on natural images. Utilizing the correlation between depth information and haze distribution, our network enhances its generalization performance on natural images. Furthermore, our method improves the PSNR for the non-homogeneous realistic haze dataset NH-HAZE from 20.66 (DeHamer's result) to 20.74, using only 1.6% of the parameters. Similarly, for the outdoor scenes realistic haze dataset O-HAZE, our method enhances the PSNR from 24.36 (MSBDN's result) to 25.77, with just 6% of the parameters. On natural road images with haze from dataset RTTS, our method improved the vehicle detection rate by 10% in terms of the R-squared value. In summary, our method outperforms all state-of-the-art methods while utilizing the least number of parameters.
Kai-Yu Wang, Yu-Hsiang Lin, Pao-Ann Hsiung
VCIP3
2023 Prediction-based peer-to-peer energy transaction market design for smart grids
I. Chien, P. Karthikeyan 0004, Pao-Ann Hsiung
Eng. Appl. Artif. Intell.3
2023 Labor exploitation investigation using statistical and multiple object tracking assessment methods
P. Karthikeyan 0004, Chih-Chun Chang, Pao-Ann Hsiung
Multim. Tools Appl.3
2020 Deep Neural Network-Based Data Reconstruction for Landslide Detection
abstract
Landslides could cause huge threats to lives and cause property damages. In the landslide prediction system, environmental information can be collected through sensors to detect the possibility of landslide occurrences. However, the data collected may be lost due to sensor failures, external interferences or other environmental factors, which may affect the accuracy of landslide predictions. In order to solve the problem of missing data, we propose a data reconstruction method based on rainfall intensity and soil moisture, which reconstructs missing data based on temporal relationships. It is based on the data trend in the past period of time. A Long Short-Term Memory (LSTM) deep neural network is trained to predict the data value in missing time slots. We use the predicted data to compensate for the missing data so as the elevate the accuracy not only of data, but also landslide predictions. Our method is compared with other reconstruction methods. The proposed LSTM model exhibit a smaller RMSE than the Linear Extrapolation (LE) method. Even if 90% of random data is lost, the RMSE results for the data reconstruction by LE and LSTM are, respectively, 0.033 and 0.036 for rainfall data and 0.029 and 0.032 for soil moisture data.
Darmawan Utomo, Liang-Cheng Hu, Pao-Ann Hsiung
IGARSS3
2019 Data Reconstruction for Cyber-Physical Landslide Detection System
abstract
Wireless Sensor Network (WSN) systems are often used to collect data from the environment for predicting landslide occurrence. However, due to the instability of wireless communications in WSN systems, data loss could occur often. To reconstruct the missing soil moisture data, heterogeneous data with spatio-temporal relations are used, such as soil moisture and rainfall intensity. In the spatial reconstruction method, we take into consideration not only the distance but also the difference of rainfall intensity between two locations. In the temporal reconstruction method, soil moisture will decrease because of evaporation if there is no rainfall. We use historical evaporation rate to reconstruct missing soil moisture data when there is no rainfall. If there is rain, we take a period of sensed soil moisture data to reconstruct any missing data. As a result, we can improve the accuracy of data reconstruction by selecting either spatial or temporal reconstruction results that have a smaller estimation error. The Root Mean Squared Error of spatiotemporal reconstruction is below 2%, even when there is 80% random missing soil moisture data. This is due to both spatiotemporal considerations, as well as, the use of heterogeneous data (rainfall intensity).
Pao-Ann Hsiung, Chih-Chen Lin
CloudCom1
2016 Feedback Control Optimization for Performance and Energy Efficiency on CPU-GPU Heterogeneous Systems
Feng-Sheng Lin, Po-Ting Liu, Ming-Hua Li, Pao-Ann Hsiung
ICA3PP4
2016 Model Predictive Optimization for distribution management in smart grids
abstract
The traditional centralized power system is gradually being replaced by smart grids. However, an important design issue is how to perform accurate demand-response such that the power distribution management is effective. This includes two sub-problems, namely the accurate prediction of future electricity demand-response situations and the optimization of power distribution. In this work, we propose a novel Model Predictive Optimization (MPO) method for the advanced distribution management system in smart grids. Future electricity situations (surplus/deficit) are predicted using a customized Autoregressive Integrated Moving Average (ARIMA) model. Pairing between buyers and sellers of electricity are performed based on not only the current situation, but also considering future situations. As a result, trading pairs with overall near-optimal cost are found through concurrent and multiple instances of Particle Swarm Optimization (PSO), along with conflict resolution. Experimental results on 30 micro-grids show the error rate of the ARIMA prediction model to be less than 10%. The proposed MPO method saves totally 19.38% overall trading cost, if predictions are made for 4 future time slots.
Hung-Lin Chao, Pei-Chi Hsieh, Tsai-Chen Yang, Pao-Ann Hsiung
IECON4
2016 SysML-Based Requirement Management to Improve Software Development
abstract
Among the various steps in the life cycle of software development, system requirement management is an essential but often neglected step. Comprehensive requirement management can not only help developers to work on a system to meet the requirements of a project, but can also play a vital role in the communications among stakeholders. In general, natural languages are often used to describe and record user requirements; however, this results in ambiguity, inconsistency, imprecision and incompleteness. To increase the accuracy of requirement modeling and analysis, it is important to have appropriate management methods and tools such that the requirement engineering process can be supported within the project. In this work, we propose a System Modeling Language (SysML)-based requirement management methodology to assist in the collection and the modeling of user requirements. We also provide a convenient procedure and a prototype tool to model, analyze, validate and verify the recorded system requirements, and consequently to ensure that the system can satisfy users’ requirements.
Chih-Hung Chang, Chih-Wei Lu, William C. Chu, Pao-Ann Hsiung, Dong-Meau Chang
Int. J. Softw. Eng. Knowl. Eng.4
2016 Introduction to the special issue on reconfigurable cyber-physical and embedded system design
Pao-Ann Hsiung, Tei-Wei Kuo, Yuan-Hao Chang 0001, Chun-Hsian Huang
J. Syst. Archit.1
2016 Auto-tuning for GPGPU applications using performance and energy model
Chih-Sheng Lin, Shih-Meng Teng, Pao-Ann Hsiung
J. Syst. Archit.3
2016 Dynamic Task Mapping with Congestion Speculation for Reconfigurable Network-on-Chip
abstract
Network-on-Chip (NoC) has been proposed as a promising communication architecture to replace the dedicated interconnections and shared buses for future embedded system platforms. In such a parallel platform, mapping application tasks to the NoC is a key issue because it affects throughput significantly due to the problem of communication congestion. Increased communication latency, low system performance, and low resource utilization are some side-effects of a bad mapping. Current mapping algorithms either do not consider link utilizations or consider only the current utilizations. Besides, to design an efficient NoC platform, mapping task to computation nodes and scheduling communication should be taken into consideration. In this work, we propose an efficient algorithm for dynamic task mapping with congestion speculation (DTMCS) that not only includes the conventional application mapping, but also further considers future traffic patterns based on the link utilization. The proposed algorithm can reduce overall congestion, instead of only improving the current packet blocking situation. Our experiment results have demonstrated that compared to the state-of-the-art congestion-aware Path Load algorithm, the proposed DTMCS algorithm can reduce up to 40.5% of average communication latency, while the maximal communication latency can be reduced by up to 67.7%.
Hung-Lin Chao, Sheng-Ya Tung, Pao-Ann Hsiung
ACM Trans. Reconfigurable Technol. Syst.3
2014 Compositional Synthesis of Concurrent Systems through Causal Model Checking and Learning
Shangwei Lin 0001, Pao-Ann Hsiung
FM2
2014 Accelerating Coverage Estimation Through Partial Model Checking
abstract
In model checking a system design against a set of properties, coverage estimation is frequently used to measure the amount of system behavior being checked by the properties. A popular coverage estimation method is to mutate the system model and check if the mutation can be detected by the given properties. For each mutation and each property, a full model check is required by some state-of-the-art coverage estimation methods. With such repeated model checking, mutation-based coverage estimation becomes significantly time-consuming. To alleviate this problem, a partial model checking (PMC) technique is proposed to recheck only those system states that were affected by a mutation, thus unnecessary rechecking of a large portion of the system states is avoided and time is saved. The PMC method has been integrated into the State Graph Manipulators model checker. Applying the proposed method to several examples showed that PMC has a saving of 50% to 70% in the coverage estimation time, and a reduction of 90% in mode visits.
Yean-Ru Chen, Jia-Jen Yeh, Pao-Ann Hsiung, Sao-Jie Chen
IEEE Trans. Computers3
2014 Reasoning and Learning-Based Dynamic Codec Reconfiguration for Varying Processing Requirements in Network-on-Chip
abstract
Crosstalk interferences and high dynamic power consumption in a network-on-chip (NoC) are two increasingly problematic design issues. Using data codecs can reduce the switching activities on wires that cause crosstalk interferences and high dynamic power. However, data codecs have different overheads in terms of area and performance, and varying capabilities in reducing crosstalk and dynamic power. To adapt to the wide range of processing requirements incurred by applications and operating environments, a reasoning and learning (REAL) framework is proposed for a reconfigurable NoC. REAL dynamically investigates the tradeoffs among reliability, dynamic power reduction, performance, and hardware resource usages to configure the reconfigurable NoC with an appropriate data codec at runtime. As a proof of concept, a 3 × 3 reconfigurable NoC was implemented on Xilinx Virtex-4 field-programmable gate array, which required 8.2% lesser number of slices compared with a conventional NoC. Experiments show that at the same overheads of performance and hardware resources the reconfigurable NoC induces a higher probability toward the reduction of crosstalk interferences and dynamic power consumption.
Jih-Sheng Shen, Pao-Ann Hsiung
IEEE Trans. Very Large Scale Integr. Syst.2
2013 The Architecture of Parallelized Cloud-Based Automatic Testing System
abstract
Software testing is the key of software quality control. However, software testing requires plenty of time, manpower and resources in hardware and software. Unfortunately, plenty of human errors may cause software testing become more difficult. To increase efficiency and reduce costs, automatic software testing is relatively important. It is possible to dynamically adjust the resources of hardware and software from the actual needs by introducing cloud technology. In this research, we proposed the paralleled cloud-based automatic testing system (PCATS) which have the advantages of real-time software testing and automatically computation scaling. PCATS can parallel the tests at the same time with distinct servers. The main contributions of this paper are automatically: (1) parse the source code to perform statistics analysis (2) generate test drivers and test cases (3) testing in virtual environments (4) paralleled testing (5) profile the consumed resources.
Chorng-Shiuh Koong, Chihhsiong Shih, Chang-Chung Wu, Pao-Ann Hsiung
CISIS4
2013 Real-Time Object Detection for Multi-Camera on Heterogeneous Parallel Processing Systems
abstract
In recent years, the need for object detection has significantly increased for multi-camera systems. However, the detection methods in such systems incur high computational cost, which leads to a major challenge in real-time applications. In this work, we propose a Scissor Algorithm for object detection using a multi-core CPU and a graphic processing unit (GPU). Leveraging the features of both the CPU and the GPU, the object detection method was enhanced in two stages: (a) pixel-to-pixel color filtering and (b) grouping. The proposed algorithm can effectively shrink the search area for detection and further improve the process of detection, thus effectively increasing the frame rate for real-time applications. Experimental results demonstrate the real-time performance of the proposed algorithm.
Chih-Sheng Lin, Shih-Meng Teng, Yen-Ting Chen, Pao-Ann Hsiung
CISIS4
2013 Spatio-Temporally-Shared Reconfigurable Fast Fourier Transform architecture design
abstract
The Fast Fourier Transform (FFT) has been one of the most popular and widely-used transform functions in communication hardware designs. With growing digital convergence, a single device needs to support multiple communication protocols, all of which need FFT computations. Currently, most FFT designs are either not shareable across applications or only among a fixed set of applications. This work proposes a novel reconfigurable FFT design called Spatio-Temporally-shAred Reconfigurable Fast Fourier Transform (STARFFT), which leverages on the partial dynamic reconfiguration technology such that it can be shared across arbitrary set of applications. STARFFT has a software driver that checks feasibility, schedules applications, and reconfigures the hardware. STARFFT hardware has several radix-2 pipelines that are time-multiplexed among applications such that significant reductions in hardware resource requirements and in power consumption are achieved. Experimental results show that STARFFT can reduce the total hardware resource usage by nearly 88% and the power consumption requirements by about 90%.
Hung-Lin Chao, Chun-Yang Peng, Cheng-Chien Wu, Ken-Shin Huang, Chun-Hsien Lu, Jih-Sheng Shen, Pao-Ann Hsiung
FPT7
2013 Backward probing deadlock detection for networks-on-chip
abstract
To accurately detect deadlocks in Network-on-Chip (NoC) as early as possible, a novel deadlock detection mechanism called Backward-probing Deadlock Detection (BDD) is proposed in this work, which can detect and resolve all existing deadlocks. It was realized using probe systems that generate probes for deadlock detection. A probe system includes a probe System Manager (SM) for turning on probe system, a probe Generator (GEN) for generating probes, a Link Selection (LS) connected to a Switch Allocation (SA), which is used for copying the generated probes, transmitting probes backward, and discarding probes when the probes find that the traversal path is just a congestion not a deadlock or when probe congestion occurs. There is also a TB Calculation (TBC) in LS for TB settings. Finally, a probe comparator (PB Comparator) is used for claiming deadlocks. Note that each port except the local one in a router has its own probe system.
Yean-Ru Chen, Zi-Rong Wangt, Pao-Ann Hsiung, Sao-Jie Chen, Meng-Hsun Tsai
NOCS3
2013 Multi-objective exploitation of pipeline parallelism using clustering, replication and duplication in embedded multi-core systems
Chih-Sheng Lin, Chao-Sheng Lin, Yu-Shin Lin, Pao-Ann Hsiung, Chihhsiong Shih
J. Syst. Archit.4
2013 Virtualizable hardware/software design infrastructure for dynamically partially reconfigurable systems
abstract
In most existing works, reconfigurable hardware modules are still managed as conventional hardware devices. Further, the software reconfiguration overhead incurred by loading corresponding device drivers into the kernel of an operating system has been overlooked until now. As a result, the enhancement of system performance and the utilization of reconfigurable hardware modules are still quite limited. This work proposes a virtualizable hardware/software design infrastructure (VDI) for dynamically partially reconfigurable systems. Besides the gate-level hardware virtualization provided by the partial reconfiguration technology, VDI supports the device-level hardware virtualization. In VDI, a reconfigurable hardware module can be virtualized such that it can be accessed efficiently by multiple applications in an interleaving way. A Hot-Plugin Connector (HPC) replaces the conventional device driver, such that it not only assists the device-level hardware virtualization but can also be reused across different hardware modules. To facilitate hardware/software communication and to enhance system scalability, the proposed VDI is realized as a hierarchical design framework. User-designed reconfigurable hardware modules can be easily integrated into VDI, and are then executed as hardware tasks in an operating system for reconfigurable systems (OS4RS). A dynamically partially reconfigurable network security system was designed using VDI, which demonstrated a higher utilization of reconfigurable hardware modules and a reduction by up to 12.83% of the processing time required by using the conventional method in a dynamically partially reconfigurable system.
Chun-Hsian Huang, Pao-Ann Hsiung
ACM Trans. Reconfigurable Technol. Syst.2
2012 Congestion-aware scheduling for NoC-based reconfigurable systems
abstract
Network-on-Chip (NoC) is becoming a promising communication architecture in place of dedicated interconnections and shared buses for embedded systems. Nevertheless, it has also created new design issue such as communication congestion and power consumption. A major factor leading to communication congestion is mapping of application tasks to NoC. Latency, throughput, and overall execution time are all affected by task mapping. As a solution, an efficient run-time Congestion-Aware Scheduling (CWS) is proposed for NoC-based reconfigurable systems, which predicts traffic pattern based on the link utilization. The proposed algorithm alleviates the overall congestion, instead of only improving the current packet blocking situation. Our experiment results have demonstrated that compared to other existing congestion-aware algorithm, the proposed CWS algorithm can reduce the average communication latency by 66%, increase the average throughput by 32%, reduce the energy consumption by 23%, and decrease the overall execution by 32%.
Hung-Lin Chao, Yean-Ru Chen, Sheng-Ya Tong, Pao-Ann Hsiung, Sao-Jie Chen
DATE4
2012 Automatic Generation of Provably Correct Embedded Systems
Shangwei Lin 0001, Yang Liu 0003, Pao-Ann Hsiung, Jun Sun 0001, Jin Song Dong 0001
ICFEM3
2012 Automatic testing environment for multi-core embedded software - ATEMES
Chorng-Shiuh Koong, Chihhsiong Shih, Pao-Ann Hsiung, Hung-Jui Lai, Chih-Hung Chang, William C. Chu, Nien-Lin Hsueh, Chao-Tung Yang
J. Syst. Softw.3
2012 Model Checking Prioritized Timed Systems
abstract
Real-time systems modeled by timed automata are often symbolically verified using Difference Bound Matrix (DBM) and Binary Decision Diagram (BDD) operations. When designing concurrent real-time systems with two or more processes sharing a resource, priorities are often used to schedule processes and to resolve conflicting resource requests. Concurrent real-time systems can thus be modeled by timed automata with priorities. However, model checking timed automata with priorities needs the DBM subtraction operation, whose result may not be convex, i.e., DBMs are not closed under subtraction. Thus, a partition of the resulting DBM is required. In this work, we propose Prioritized Timed Automata (PTA) and resolve all the issues related to the model checking of PTA. Two algorithms are proposed including an optimal DBM subtraction algorithm that produces the minimal number of DBM partitions, and a DBM merging algorithm that reduces the DBM partitions after a series of DBM subtractions. Application examples show the advantages of the proposed method in terms of support for the efficient verification of prioritized timed systems.
Shangwei Lin 0001, Pao-Ann Hsiung
IEEE Trans. Computers2
2011 Network-on-Chip router design with Buffer-Stealing
abstract
Communication in a Network-on-Chip (NoC) can be made more efficient by designing faster routers, using larger buffers, larger number of ports and channels, and adaptive routing, all of which incur significant overheads in hardware costs. As a more economic solution, we try to improve communication efficiency without increasing the buffer size. A Buffer-Stealing (BS) mechanism is proposed, which enables the input channels that have insufficient buffer space to utilize at runtime the unused input buffers from other input channels. Implementation results of the proposed BS design for a 64-bit 5-input-buffer router show a reduction of the average packet transmission latency by up to 10.17% and an increase of the average throughput by up to 23.47%, at an overhead of 22% more hardware resources.
Wan-Ting Su, Jih-Sheng Shen, Pao-Ann Hsiung
ASP-DAC3
2011 Pattern-based framework for modularized software development and evolution robustness
Chih-Hung Chang, Chih-Wei Lu, Pao-Ann Hsiung
Inf. Softw. Technol.3
2011 VERTAF/Multi-Core: A SysML-Based Application Framework for Multi-Core Embedded Software Development
Chao-Sheng Lin, Chun-Hsien Lu, Shangwei Lin 0001, Yean-Ru Chen, Pao-Ann Hsiung
J. Comput. Sci. Technol.5
2011 Counterexample-Guided Assume-Guarantee Synthesis through Learning
abstract
Assume-guarantee reasoning (AGR) is a promising compositional verification technique that can address the state space explosion problem associated with model checking. Since the construction of assumptions usually requires nontrivial human efforts, a framework was already proposed for generating assumptions automatically using the L* algorithm. However, if the framework shows that a system model does not satisfy a given specification, the designer has to manually refine the system model. To automate this refinement process, we propose a framework that can automatically eliminate all counterexamples from a system model such that the synthesized model satisfies a given safety specification. Further, the framework for synthesis is not only automatic, but is also an iterative L*-based compositional process, i.e., the global state space of the system is never generated in the synthesis process. When a model checker shows that a system model does not satisfy a specification by giving a counterexample, the proposed framework eliminates a class of equivalent counterexamples, that is, the set of counterexamples that transit to the error state through the same final transition. Then, AGR is applied again to check if there is another counterexample. The action of eliminating counterexamples continues until all classes of counterexamples are eliminated from the system model. We prove that the synthesized model satisfies the specification and the synthesis flow terminates after a finite number of iterations. Due to compositional synthesis, our target model for synthesis, namely the component models, is much smaller than the global system state graph.
Shangwei Lin 0001, Pao-Ann Hsiung
IEEE Trans. Computers2
2011 Model-Based Verification and Estimation Framework for Dynamically Partially Reconfigurable Systems
abstract
Unified Modeling Language (UML), an industry de-facto standard, has been used to analyze dynamically partially reconfigurable systems (DPRS) that can reconfigure their hardware functionalities on-demand at runtime. To make model-driven architecture (MDA) more realistic and applicable to the DPRS design in an industrial setting, a model-based verification and estimation (MOVE) framework is proposed in this work. By taking advantage of the inherent features of DPRS and considering real-time system requirements, a semiautomatic model translator converts the UML models of DPRS into timed automata models with transition urgency semantics for model checking. Furthermore, a UML-based hardware/software co-design platform (UCoP) is proposed to support the direct interaction between the UML models and the real hardware architecture. The two-phase verification process, including exhaustive functional verification and physical-aware performance estimation, is completely model-based, thus reducing system verification efforts. We used a dynamically partially reconfigurable network security system (DPRNSS) as a case study. The related experiments have demonstrated that the model checker in MOVE can alleviate the impact of the state-space-explosion problem. Compared to the synthesis-based estimation method having inaccuracies ranging from -43.4% to 18.4%, UCoP can provide accurate and efficient platform-specific verification and estimation through actual time measurements.
Chun-Hsian Huang, Pao-Ann Hsiung
IEEE Trans. Ind. Informatics2
2010 Supporting Design Enhancement by Pattern-Based Transformation
abstract
In general, a design pattern is usually documented in the form of an essay, with descriptions and rough design such as intent, motivation, structure, behavior, applicability and consequence, etc. Even though there are tools supporting pattern application, developers still may misuse patterns since misunderstanding. It may result failures of systems because of inconsistencies or design errors. In fact, the refinement process by applying a design pattern is merely the addition or removal of model elements in structure view. The refinement process for each design pattern is almost constant whenever the same pattern is applied. In this paper, we propose an approach for design pattern application and assisting the design enhancement by model transformation. Furthermore, we demonstrate our approach by a case study on a real-world multi-core embedded system PVE (Parallel Video Encoder), where a design pattern Command Pipeline is designed for the design enhancement.
Nien-Lin Hsueh, Peng-Hua Chu, Pao-Ann Hsiung, Min-Ju Chuang, William C. Chu, Chih-Hung Chang, Chorng-Shiuh Koong, Chihhsiong Shih
COMPSAC3
2010 Learning-based adaptation to applications and environments in a reconfigurable Network-on-Chip
abstract
The set of applications communicating via a Network-on-Chip (NoC) and the NoC itself both have varying run-time requirements on reliability and power-efficiency. To meet these requirements, we propose a novel Power-aware and Reliable Encoding Schemes Supported reconfigurable Network-on-Chip (PRESSNoC) architecture which allows processing elements, routers, and data encoding methods to be reconfigured at runtime. Further, an intelligent selection of encoding methods is achieved through a REasoning And Learning (REAL) framework at run-time. An instance of PRESSNoC was implemented on a Xilinx Virtex 4 FPGA device, which required 25.5% lesser number of slices compared to a conventional NoC with a full-fledged encoding method. The average benefit to overhead ratio of the proposed architecture is greater than that of a conventional NoC by 71%, 32%, and 277% when we consider the individual effects of interference rate per instruction, application domains, and system characteristics, respectively. Experiments have thus shown that PRESSNoC induces a higher probability toward the reduction of crosstalk interferences and dynamic power consumption, at the same amount of overheads in performance and hardware usage.
Jih-Sheng Shen, Chun-Hsian Huang, Pao-Ann Hsiung
DATE3
2010 Innovative Application of RFID Systems to Special Education Schools
abstract
Innovation is a new way of doing something. It may be incremental, radical, or revolutionary changes in thinking, products, processes, or organizations. Different from invention, which is an idea made manifest, innovation is ideas applied successfully. In this work, we strive to apply innovative Radio Frequency Identification (RFID) systems to special education school campus because in this modern age of science and technology, there still exists a wide digital gap in special education schools such that they have not yet benefited from technology advancements such as RFID. Supported by the ministry of education in Taiwan, we successfully designed and deployed RFID technology to the campus of a special education school at Chiayi in Taiwan. Though the technology was applied to eight different use case scenarios, we will focus on five of the more innovative ones in this work, including student temperature monitoring (STM), body weight monitoring (BWM), garbage disposal monitoring (GDM), mopping course recording (MCR), and campus visitor monitoring (CVM). Both active and passive tags and readers were employed to implement these five systems within the same campus. The benefits obtained from these systems by the students, teachers, and administrators were three-folds. First, student health monitoring through STM and BWM systems allowed the teachers and administration real-time control over changing health conditions that significantly affects such students. Second, course monitoring and recording through GDM and MCR allowed teachers to easily grasp and tune the learning curve of each student and also to implement a more guided training based on past learning efforts. Last but not least, campus safety monitoring through CVM allowed the administration to monitor the location of visitors in the campus and thus safeguard the students and teachers from dangerous or troublesome visitors. Novel techniques and creative methods were employed in the five systems, including temperature correction algorithm in STM, BMI-based weight tuning strategy in BWM, multiple route-tracking in GDM, learning improvement through history analysis in MCR, and face detection in CVM. The project was successfully deployed and is currently in use by the Chiayi School of Special Education which has more than 300 students and 150 administration staff and faculty.
Shu-Hui Yang, Pao-Ann Hsiung
NAS2
2010 A Self-Adaptive Hardware/Software System Architecture for Ubiquitous Computing Applications
Chun-Hsian Huang, Jih-Sheng Shen, Pao-Ann Hsiung
UIC3
2010 UML-based hardware/software co-design platform for dynamically partially reconfigurable network security systems
Chun-Hsian Huang, Pao-Ann Hsiung, Jih-Sheng Shen
J. Syst. Archit.2
2010 Model-based platform-specific co-design methodology for dynamically partially reconfigurable systems with hardware virtualization and preemption
Chun-Hsian Huang, Pao-Ann Hsiung, Jih-Sheng Shen
J. Syst. Archit.2
2010 Scheduling and Placement of Hardware/Software Real-Time Relocatable Tasks in Dynamically Partially Reconfigurable Systems
abstract
With the gradually fading distinction between hardware and software, it is now possible to relocate tasks from a microprocessor to reconfigurable logic and vice versa. However, existing hardware-software scheduling can rarely cope with such runtime task relocation. In this work, we propose a new Relocatable Hardware-Software Scheduling (RHSS) method that not only can be applied to dynamically relocatable hardware-software tasks, but also increases the reconfigurable hardware resource utilization, reduces the reconfigurable hardware resource fragmentation with realistic placement methods, and makes best efforts at meeting the real-time constraints of tasks. The feasibility of the proposed relocatable hardware-software scheduling algorithm was proved by applying it to some randomly generated examples and a real dynamically reconfigurable network security system example. Compared to the quadratic time complexity of the state-of-the-art Adaptive Hardware-Software Allocation (AHSA) method, RHSS is linear in time complexity, and improves the reconfigurable hardware utilization by as much as 117.8%. The scheduling and placement time and the memory usage are also drastically reduced by as much as 89.5% and 96.4%, respectively.
Pao-Ann Hsiung, Chun-Hsian Huang, Jih-Sheng Shen, Cheng-Chi Chiang
ACM Trans. Reconfigurable Technol. Syst.1
2009 A Model-Driven Multicore Software Development Environment for Embedded System
abstract
Multi-core programming is no more a luxury; it is now a necessity, because even embedded processors are becoming multi-core. However, the state-of-the-art techniques such as OpenMP and the Intel Threading Building Block (TBB) library are far from user-friendly due to the tedious work needed in explicitly designing multi-core programs and debugging. At the present days, a solution for above problems will be that to enhance the abstract level of multicore embedded software design. By leveraging on the expertise gained from Verifiable Embedded Real-Time Application Framework (VERTAF), we propose a Multi-Core version of VERTAF, called VERTAF/ Multi-core (VMC in short). VMC is an integrated development environment for multi-core embedded software architecture. Developers would be able to 1. describe their system requirements with SysML by using this environment, 2. model their design with SysML standard notation, 3. automatically apply a pattern structure into their design for a high quality multicore embedded system, 4. generate source code through a well-designed model; 5. map to different hardware architecture as assigned by the model, and 6. finally we can test the code.Using the model driven architecture (MDA) design flow in SysML, we saw a significantly improvement on productivity and quality of a multicore embedded programming over traditional approach.
Chihhsiong Shih, Chien-Ting Wu, Cheng-Yao Lin, Pao-Ann Hsiung, Nien-Lin Hsueh, Chih-Hung Chang, Chorng-Shiuh Koong, William C. Chu
COMPSAC (2)4
2009 VERTAF/Multi-Core: A SysML-Based Application Framework for Multi-Core Embedded Software Development
Pao-Ann Hsiung, Chao-Sheng Lin, Shangwei Lin 0001, Yean-Ru Chen, Chun-Hsien Lu, Sheng-Ya Tong, Wan-Ting Su, Chihhsiong Shih, Chorng-Shiuh Koong, Nien-Lin Hsueh, Chih-Hung Chang, William C. Chu
ICA3PP1
2009 On the Use of a UML-Based HW/SW Co-Design Platform for Reconfigurable Cryptographic Systems
abstract
In this work, we use our proposed UML-based HW/SW co-design platform (UCoP) to implement a reconfigurable cryptographic system for network multimedia applications. The UCoP is categorized into three reusable models, including software application, hardware configuration and system management models. Through the use of the reusable models, the proposed dynamically partially reconfigurable system can be easily implemented in UCoP. Furthermore, the direct interaction with the real system architecture in UCoP is very helpful to designers in validating and analyzing system correctness and performance at a high-level, which can significantly reduce system development efforts. We compare the use of UCoP with a lower-bound estimation method by implementing a network multimedia application with the data encryption/decryption. According to our experiment results, we can clearly see how UCoP plays a key role in helping designers to develop dynamically partially reconfigurable systems with hard real-time constraints.
Chun-Hsian Huang, Pao-Ann Hsiung
ISCAS2
2009 A 900 MHz to 5.2 GHz Dual-loop Feedback Multi-band LNA
abstract
This paper demonstrates a multi-band low noise amplifier (LNA) in 0.13µm CMOS process, which is configurable with switching capacitor in 900MHz, 1800MHz, 2.4GHz and 5.2GHz bands. A dual-loop feedback technique is used to enhance the performance in noise figure, power gain and power consumption. The noise figures are 2.6dB at 900MHz, 2dB at 1800MHz, 2.1dB at 2.4GHz and 3.5dB at 5.2GHz; and the power gains are 17dB at 900MHz, 21dB at 1800MHz, 26dB at 2.4GHz and 19dB at 5.2GHz in post-simulation. The S11in all of the bands is below −10dB by wideband matching. The LNA consumes 3.92mW from 1 V supply.
Jia-Wei Lin, Da-Tong Yen, Wei-Yi Hu, Chu Yu, Mao-Hsu Yen, Pao-Ann Hsiung, Sao-Jie Chen
ISCAS6
2009 Modeling and verification of real-time embedded systems with urgency
Pao-Ann Hsiung, Shangwei Lin 0001, Yean-Ru Chen, Chun-Hsian Huang, Chihhsiong Shih, William C. Chu
J. Syst. Softw.1
2008 Automatic synthesis and verification of real-time embedded software for mobile and ubiquitous systems
Pao-Ann Hsiung, Shangwei Lin 0001
Comput. Lang. Syst. Struct.1
2008 Perfecto: A systemc-based design-space exploration framework for dynamically reconfigurable architectures
abstract
To cope with increasing demands for higher computational power and greater system flexibility, dynamically and partially reconfigurable logic has started to play an important role in embedded systems and systems-on-chip (SoC). However, when using traditional design methods and tools, it is difficult to estimate or analyze the performance impact of including such reconfigurable logic devices into a system design. In this work, we present a system-level framework, called Perfecto, which is able to perform rapid exploration of different reconfigurable design alternatives and to detect system performance bottlenecks. This framework is based on the popular IEEE standard system-level design language SystemC, which is supported by most EDA and ESL tools. Given an architecture model and an application model, Perfecto uses SystemC transaction-level models (TLMs) to simulate the system design alternatives automatically. Different hardware-software copartitioning, coscheduling, and placement algorithms can be embedded into the framework for analysis; thus, Perfecto can also be used to design the algorithms to be used in an operating system for reconfigurable systems. Applications to a simple illustration example and a network security system have shown how Perfecto helps a designer make intelligent partition decisions, optimize system performance, and evaluate task placements.
Pao-Ann Hsiung, Chao-Sheng Lin, Chih-Feng Liao
ACM Trans. Reconfigurable Technol. Syst.1
2007 Real-Time Embedded Software Design for Mobile and Ubiquitous Systems
Pao-Ann Hsiung, Shangwei Lin 0001, Chin-Chieh Hung, Jih-Ming Fu, Chao-Sheng Lin, Cheng-Chi Chiang, Kuo-Cheng Chiang, Chun-Hsien Lu, Pin-Hsien Lu
EUC1
2007 Exploiting Hardware and Software Low Power Techniques for Energy Efficient Co-scheduling in Dynamically Reconfigurable Systems
abstract
Currently, the hardware and the software tasks in reconfigurable systems are either scheduled separately at run time or co-scheduled statically, which results in high power consumption and low performance. This work proposes runtime co-scheduling of hardware and software tasks by using the slack time, which is introduced due to reusing hardware task configurations, for dynamically scaling the processor voltage such that preceding software tasks consume lesser power. At the same time, the reuse of hardware task configurations also result in lower power consumption and higher performance due to fewer number of reconfigurations. The combined effects of hardware configuration reuse and software dynamic voltage scaling result in schedules with a lower power consumption and higher performance than that obtained through individual techniques applied to hardware and software separately. The proposed method was implemented in the SystemC-based Perfecto simulation environment for dynamically reconfigurable hardware software systems and TGFF was used for generating random task sets as input for Perfecto. We performed extensive experiments whose results show that irrespective of different slack ratios or hardware partitions, the schedules generated by our proposed method are more energy efficient than methods that either do not apply any runtime techniques or only apply hardware configuration prefetch and reuse.
Pao-Ann Hsiung, Chih-Wen Liu
FPL1
2007 Reconfigurable Hardware Module Sequencer - A Tradeoff Between Networked and Data Flow Architectures
abstract
Dynamically reconfigurable systems either adopt a processor-controlled networked architecture or a sequencer-controlled data flow architecture. In the networked architecture, the processor is overloaded with data transfer requests, whereas in the data flow architecture, the burden is completely shifted from the processor to the data sequencer. As a tradeoff between these two extremes, this work proposes a novel module sequencer architecture, which not only allows the processor and the sequencer to share the heavy data communication load, but is also more coherent with the conventional processor-FPGA architecture. Further, the architecture is highly flexible because it can be tuned to fit a particular application. Application examples show how the proposed architecture is superior to the networked architecture in terms of lower communication load and to the data flow architecture in terms of reduced system complexity.
Kai-Jung Shih, Chin-Chieh Hung, Pao-Ann Hsiung
FPT3
2007 From ISA to application design via RTOS - a course design framework for embedded software
abstract
Embedded systems have pervaded every aspect of our daily lives, however their design and verification are often accomplished using ad hoc and trial-and-error methods. Courses introducing systematic and more formal methods are required. However, currently there is little consensus on what a standard syllabus for an undergraduate course on embedded software design should cover. This paper proposes a course design that have undergone thorough experimentations and evaluations through the last four years in actual classes. The course starts from the ARM instruction set architecture and concludes with an introduction of Java-based wireless application design. The design of standalone, as well as, RTOS-based embedded software are all introduced. The course has culminated in the generation of embedded software engineers that significantly contribute to the technical industry in Taiwan, spanning from handheld devices to home appliances and from networked systems to personal computer accessories. We hope the proposed curriculum becomes a standard effort at training embedded software engineers in both theory and practice.
Pao-Ann Hsiung, Shangwei Lin 0001
ICPADS1
2007 Dynamically Swappable Hardware Design in Partially Reconfigurable Systems
abstract
In this work, we propose two wrapper designs for arbitrary digital hardware circuit designs such that they can be enhanced with the capability for dynamic swapping controlled by software. A hardware design with either of the proposed wrappers can thus be swapped out of the partially reconfigurable logic at runtime in some intermediate state of computation and then swapped in when required to continue from that state. The context data is saved to a buffer in the wrapper at interruptible states, and then the wrapper takes care of saving the hardware context to communication memory through a peripheral bus, and later restoring the hardware context after the design is swapped in. The overheads of the hardware standardization and the wrapper in terms of additional reconfigurable logic resources and the time for context switching are small and generally acceptable. With the capability for dynamic swapping, high priority hardware tasks can interrupt low priority tasks in real-time embedded systems so that the utilization of hardware space per unit time is increased.
Chun-Hsian Huang, Kai-Jung Shih, Chao-Sheng Lin, Shih-Shiue Chang, Pao-Ann Hsiung
ISCAS5
2007 Modeling and Automatic Failure Analysis of Safety-Critical Systems Using Extended Safecharts
Yean-Ru Chen, Pao-Ann Hsiung, Sao-Jie Chen
SAFECOMP2
2007 Automatic Failure Analysis Using Safecharts
abstract
With rapid developments in science and technology, we now see the ubiquitous use of different types of safety-critical systems in our daily lives such as in avionics, consumer electronics, and medical systems. In such systems, unintentional design faults might result in injury or even death to human beings. To avoid such mishaps, we need to verify safety-critical systems thoroughly and formal verification techniques such as model checking are a very promising approach. However, modeling the systems formally is a challenging task, which is further aggravated by the necessity to model faults and automatic repairs in safety-critical systems. Currently, there is no automatic technique in formal verification that can aid system designers in formally modeling the faults and repairs. This work contributes by proposing an extension to the Safecharts model so that faults and repairs are easily modeled and then the Safecharts are transformed into semantically equivalent Extended Timed Automata models that can be directly model checked. In this way, automatic failure analysis techniques are integrated into the SGM model checker. Application examples show the feasibility and benefits of the proposed model-driven verification of safety-critical systems.
Yean-Ru Chen, Pao-Ann Hsiung
Int. J. Softw. Eng. Knowl. Eng.2
2007 Model Checking Safety-Critical Systems Using Safecharts
abstract
With rapid developments in science and technology, we now see the ubiquitous use of different types of safety-critical systems in our daily lives such as in avionics, consumer electronics, and medical systems. In such systems, unintentional design faults might result in injury or even death to human beings. To make sure that safety-critical systems are really safe, there is a need to verify them formally. However, the verification of such systems is getting more and more difficult because designs are becoming very complex. To cope with high design complexity, currently, model-driven architecture design is becoming a well-accepted trend. However, existing methods of testing and standards conformance are restricted to implementation code, so they do not fit very well with model-based approaches. To bridge this gap, we propose a model-based formal verification technique for safety-critical systems. In this work, the model-checking paradigm is applied to the Safecharts model, which was used for modeling but not yet used for verification. Our contributions listed are as follows: first, the safety constraints in Safecharts are mapped to semantic equivalents in timed automata for verification. Second, the theory for safety constraint verification is proven and implemented in a compositional model checker (that is, the state-graph manipulator (SGM)). Third, prioritized and urgent transitions are implemented in SGM to model the risk semantics in Safecharts. Finally, it is shown that the priority-based approach to mutual exclusion of resource usage in the original Safecharts is unsafe and corresponding solutions are proposed. Application examples show the feasibility and benefits of the proposed model-driven verification of safety-critical systems
Pao-Ann Hsiung, Yean-Ru Chen, Yen-Hung Lin
IEEE Trans. Computers1
2006 Model Checking Timed Systems with Urgencies
Pao-Ann Hsiung, Shangwei Lin 0001, Yean-Ru Chen, Chun-Hsian Huang, Jia-Jen Yeh, Chao-Sheng Lin, Hsiao-Win Liao
ATVA1
2006 Perfecto: A Systemc-Based Performance Evaluation Framework for Dynamically Partially Reconfigurable Systems
abstract
To cope with increasing demands for higher computational power and flexibility, dynamically and partially reconfigurable logic has started to play an important role in embedded systems and systems-on-chip. However, when using traditional design methods and tools, it is difficult to estimate or analyze the performance impact of including such reconfigurable logic devices into a system design. In this work, we present an easy-to-use system-level framework, called Perfecto, which is able to perform rapid explorations of different reconfiguration alternatives and to detect system performance bottlenecks. This framework is based on the popular system-level design language SystemC, which is supported by most EDA and ESL tools. Different hardware-software co-partitioning, co-scheduling, and placement algorithms can all be embedded into the framework for analysis. Perfecto can also be used to design the algorithms to be used in an operating system for reconfigurable systems. Applications to some examples have shown advantages of having an evaluation framework such as Perfecto
Pao-Ann Hsiung, Chun-Hsian Huang, Chih-Feng Liao
FPL1
2005 Model Checking Prioritized Timed Automata
Shangwei Lin 0001, Pao-Ann Hsiung, Chun-Hsian Huang, Yean-Ru Chen
ATVA2
2005 Hardware Task Scheduling and Placement in Operating Systems for Dynamically Reconfigurable SoC
Yuan-Hsiu Chen, Pao-Ann Hsiung
EUC2
2005 UML-Based Design Flow and Partitioning Methodology for Dynamically Reconfigurable Computing Systems
Chih-Hao Tseng, Pao-Ann Hsiung
EUC2
2005 Modeling and Verification of Safety-Critical Systems Using Safecharts
Pao-Ann Hsiung, Yen-Hung Lin
FORTE1
2005 Model Checking Timed Systems with Priorities
abstract
Priorities are used to resolve conflicts such as in re-source sharing and in safety designs. The use of priorities has become indispensable in real-time system design such as in scheduling, synchronization, arbitration, and fairness guaranteeing. There are several modeling frameworks that show how timed systems with priorities are to be designed and how priority schedulers can be automatically synthesized. However, the verification of timed systems with priorities using model checking is still a relatively untouched area. We show what the issues are in model checking timed systems with priorities and how the issues are solved in this work. In the process, we propose an optimal zone subtraction algorithm. The method has been implemented into the SGM model checker and successfully applied to real-time embedded systems and safety-critical systems, which illustrate the feasibility and advantages of the proposed verification method.
Pao-Ann Hsiung, Shangwei Lin 0001
RTCSA1
2005 Model-based Verification of Safety-Critical Systems
Pao-Ann Hsiung, Yen-Hung Lin
SEKE1
2005 Device-Centric Low-Power Scheduling for Real-Time Embedded Systems
abstract
Existing low power schedules mainly try to decrease overall system energy usage by shutting down unused devices after process scheduling. This often leads to suboptimal energy usage due to the lack of process slack time to shut down unused devices in a fixed system schedule such as that generated by a rate-monotonic or an earliest deadline first scheduler in a real-time embedded system. In this work, we try to integrate process scheduling with low power scheduling such that a low-power real-time feasible schedule is obtained. The proposed method called Low-Power Quasi-Dynamic Scheduling (LQS) was implemented and applied to some examples, including a sensor network node and Bluetooth devices, to prove its benefits.
Pao-Ann Hsiung, Hsin-Chieh Kao
Int. J. Softw. Eng. Knowl. Eng.1
2005 SESAG: an object-oriented application framework for real-time systems
abstract
Abstract Advancements in hardware and software technologies have made possible the design of real‐time systems and applications where stringent timing constraints are imposed on critical tasks. The design of such systems is more complex than that of temporally unrestricted systems because system correctness depends on the satisfaction of functional as well as temporal requirements. To aid users in correctly and efficiently designing systems, object‐oriented frameworks provide a useful environment for significant reuse and reduction in design effort. In contrast to other application domains, there has been relatively little work on an application framework for the design of real‐time systems. Facing the growing need for real‐time applications, we propose a novel application framework called SESAG, which consists of five components, namely Specifier, Extractor, Scheduler, Allocator, and Generator. Within SESAG, several design patterns are proposed and used for the development of real‐time applications. A new evaluation metric called relative design effort is proposed for evaluating SESAG. Experiences in using SESAG show a significant increase in design productivity through design reuse and a significant decrease in design time and effort. Two complex application examples have been developed using SESAG and evaluated using the new evaluation metric. The examples demonstrate relative design efforts of at most 18% of the design efforts required by conventional methods. Copyright © 2005 John Wiley & Sons, Ltd.
Pao-Ann Hsiung, Trong-Yen Lee, Jih-Ming Fu, Win-Bin See
Softw. Pract. Exp.1
2004 Formal Design and Verification of Real-Time Embedded Software
Pao-Ann Hsiung, Shangwei Lin 0001
APLAS1
2004 Mutation Coverage Estimation for Model Checking
Te-Chang Lee, Pao-Ann Hsiung
ATVA2
2004 Automatic Synthesis and Verification of Real-Time Embedded Software
Pao-Ann Hsiung, Shangwei Lin 0001
EUC1
2004 VERTAF: An Application Framework for the Design and Verification of Embedded Real-Time Software
abstract
The growing complexity of embedded real-time software requirements calls for the design of reusable software components, the synthesis and generation of software code, and the automatic guarantee of nonfunctional properties such as performance, time constraints, reliability, and security. Available application frameworks targeted at the automatic design of embedded real-time software are poor in integrating functional and nonfunctional requirements. To bridge this gap, we reveal the design flow and the internal architecture of a newly proposed framework called verifiable embedded real-time application framework (VERTAF), which integrates software component-based reuse, formal synthesis, and formal verification. A formal UML-based embedded real-time object model is proposed for component reuse. Formal synthesis employs quasistatic and quasidynamic scheduling with automatic generation of multilayer portable efficient code. Formal verification integrates a model checker kernel from SGM, by adapting it for embedded software. The proposed architecture for VERTAF is component-based and allows plug-and-play for the scheduler and the verifier. Using VERTAF to develop application examples significantly reduced design effort and illustrated how high-level reuse of software components combined with automatic synthesis and verification can increase design productivity.
Pao-Ann Hsiung, Shangwei Lin 0001, Chih-Hao Tseng, Trong-Yen Lee, Jih-Ming Fu, Win-Bin See
IEEE Trans. Software Eng.1
2003 Quasi-Dynamic Scheduling for the Synthesis of Real-Time Embedded Software with Local and Global Deadlines
Pao-Ann Hsiung, Cheng-Yi Lin, Trong-Yen Lee
RTCSA1
2003 RESS: Real-Time Embedded Software Synthesis and Prototyping Methodology
Trong-Yen Lee, Pao-Ann Hsiung, I-Mu Wu, Feng-Shi Su
RTCSA2
2003 Software Platform for Embedded Software Development
Win-Bin See, Pao-Ann Hsiung, Trong-Yen Lee, Sao-Jie Chen
RTCSA2
2002 Formal Synthesis and Code Generation of Real-Time Embedded Software using Time-Extended Quasi-Static Scheduling
abstract
The rapid escalation in complexity of real-time embedded systems design has made embedded software an integral system part such that formal software synthesis has become an indispensable design automation technique. The current work takes one more step forward in this research direction by proposing a formal synthesis method for complex real-time embedded software. Compared to previous work, our method not only synthesizes embedded software with complex interrelated branching choices for execution within a user-given memory bound, but also tries to guarantee the satisfaction of all user-given local and global time constraints. Our proposed method called time-extended quasi-static scheduling (TEQSS) synthesizes real-time embedded software code from a set of time complex-choice Petri nets. The two most important issues in real-time embedded software, namely memory and time constraints are both elegantly and efficiently handled by TEQSS. We show the feasibility of our method through a master-slave role switch application which is a part of the Bluetooth wireless communication protocol.
Pao-Ann Hsiung, Trong-Yen Lee, Feng-Shi Su
APSEC1
2002 TCN: Scalable Hierarchical Hypercubes
abstract
Hierarchical hypercubes, such as extended hypercube (EH), hyperweave (HW), and extended hypercube with cross connections (EHC), have been proposed to overcome the scalability limitation of conventional hypercubes through the use of fixed dimension hypercubes of processing elements (PEs) as basic modules interconnected by network controllers (NC) which are themselves interconnected into hypercubes. The scalability of all these three hierarchical hypercube networks is still limited, because the average network communication load in each NC increases as the number of interconnected PEs become very large. In this work a generalization scheme is proposed for improving network scalability, namely transformer cube network (TCN). For illustration purposes, generalized TCN is presented only for EH, though the same scheme can be applied to HW and EHC as well. Several characteristics of TCN, such as topological properties, message routing complexity, fault tolerance, and scalability are analyzed We present a communication algorithm for one-to-one message passing in a fault-free case. Further, the application of TCN to a class of divide-and-conquer problems is shown to have a time complexity of O(log/sub 2/ N), where N is the total number of PEs.
Trong-Yen Lee, Pao-Ann Hsiung, Sao-Jie Chen
ICPADS2
2002 Efficient and User-Friendly Verification
abstract
A compositional verification method from a high-level resource-management standpoint is presented for dense-time concurrent systems and implemented in the tool of SGM (State-Graph Manipulators) with graphical user interface. SGM packages sophisticated verification technology into state-graph manipulators and provides a user interface which views state-graphs as basic data-objects. Hence, users do not have to be verification theory experts and do not have to trace inside state-graphs to analyze state and path properties to make the best use of verification theory. Instead, users can construct their own verification strategies based on observation on the state-graph complexity changes after experimenting with some combinations of manipulators. Moreover, SGM allows users to control the complexity of state-graphs through iterative state-graphs merging and reductions before they become out of control, Reduction techniques specially designed for the context of state-graph iteration composition and shared variable manipulations are developed and used in SGM. Experiments on different benchmarks to show SGM performance are reported. An algorithm based on group theory to pick a manipulator combination is presented.
Farn Wang, Pao-Ann Hsiung
IEEE Trans. Computers2
2001 Formal Verification of Embedded Real-Time Software in Component-Based Application Frameworks
abstract
Producing correct software is a major goal for application frameworks that are targeted at embedded real-time systems because incorrect software is of no use and may also cause severe system damage. It is shown how formal verification can be elegantly, seamlessly, and scalably integrated into a component-based object-oriented application framework for embedded real-time systems. Two issues in such technology integration are addressed: (1) the choice of a common system model, and (2) the integration of formal synthesis and model checking. Solutions are provided, respectively, in the form of (1) proposing a new formal object-oriented model (FOOM), and (2) the execution of model checkers within synthesis algorithms. Technically, we propose a compositional software verification framework, in which model checking is employed, with state-space reduction techniques adapted for embedded real-time software. A separate verifier component is proposed for modular integration as illustrated by its implementation in the VERTAF application framework. An example illustrates the success of our approach and the benefits gained through integrating formal verification.
Pao-Ann Hsiung, Win-Bin See, Trong-Yen Lee, Jih-Ming Fu, Sao-Jie Chen
APSEC1
2001 Formal Synthesis and Control of Soft Embedded Real-Time Systems
Pao-Ann Hsiung
FORTE1
2001 POSE: a parallel object-oriented synthesis environment
abstract
Design automation tools and methodologies always encounter a problem of how systems may be designed efficiently, including issues such as static modeling and dynamic manipulation of system parts. With the rapid progress of design technology, the continuously increasing number of different choices per system part and the growing complexity of today's systems, the efficiency of the design environment is not only a major concern now, but will also be a demanding problem in the near future. In contrast to heuristic methods, a novel environment called POSE is proposed that increases efficiency during design without losing optimality in the final design results. System parts are modeled using the popular object-oriented modeling technique and are dynamically manipulated using the parallel design technique. A complete integration of object-oriented and parallel techniques is one of the major feature of POSE. Common problems related to parallel design such as emptiness and deadlock are also elegantly solved within POSE. Experimental results and formal analysis based on POSE all show its practical and theoretical usefulness. POSE can be used at any level of synthesis as long as off-the-shelf building-blocks manipulation is required. POSE can be applied especially to system-level synthesis, whose targets can be parallel computer architectures, systems-on-chip, or embedded systems. We will show how POSE has been applied to ICOS, a recently proposed synthesis methodology. Furthermore, POSE can be easily integrated with other heuristic design methodologies to allow increased design efficiency.
Pao-Ann Hsiung
ACM Trans. Design Autom. Electr. Syst.1
2000 Concurrent Embedded Real-Time Software Verification
abstract
The verification of software is more complex than hardware due to inherent flexibilities (dynamic behavior) that incur a multitude of possible system states. The verification of Concurrent Embedded Real-Time Software (CERTS) is all the more difficult due to its concurrency and embeddedness. The work presented shows how the complexity of CERTS verification can be reduced significantly through answering common engineering questions such as when, where, and how one must verify embedded software. Application examples illustrate the usefulness of our technique in increasing verification scalability.
Pao-Ann Hsiung
COMPSAC1
2000 Embedded software verification in hardware-software codesign
Pao-Ann Hsiung
J. Syst. Archit.1
2000 CMAPS: a cosynthesis methodology for application-oriented parallel systems
abstract
Currently, a lot of research is devoted to system design , and little work is done on requirements analysis . Besides going from specification to design, one of our main objectives is to show how an application problem can be transformed into specifications. Working from the hardware-software codesign perspective, a system is designed starting from an application problem itself, rather than the detailed behavioral specifications. Given an application problem specified as a directed acyclic graph of elementary problems, a hardware-software solution is derived such that the synthesized software, a parallel pseudoprogram, can be scheduled and executed on the synthesized software, a parallel pseudoprogram, can be scheduled and executed on the synthesized hardware, a set of system-level parallel computer specifications, with heuristically optimal performance. This is known as system-level cosynthesis of application-oriented general-purpose parallel systems for which a novel methodology called Cosynthesis Methodology for Applicaton-Oriented Parallel Systems (CMAPS), is presented. Since parallel programs and multiprocessor architectures are largely interdependent, CMAPS explores the relationship between hardware designs and software algorithms by interleaving the modeling phases and the synthesis phases of both hardware and software. High scalability in terms of problem complexity and easy upgradability to new technologies are achieved through modularization of the input problem specification, of the software algorithms, and of the hardware subsystem models. The work presented in this paper will be beneficial to designers of general-purpose parallel computer systems which must be oriented toward solving some user-specified problem such as the global controller of an industry automation process or a multiprocessor video server. Some application examples are given to illustrate various codesign phases of CMAPS and its feasibility.
Pao-Ann Hsiung
ACM Trans. Design Autom. Electr. Syst.1
1999 Hardware-software coverification of concurrent embedded real-time systems
abstract
The results of hardware-software codesign of concurrent embedded real-time systems are often not verified or not easily verifiable. This has serious consequences when high-assurance systems are codesigned. The main difficulty lies in the different time-scales of the embedded hardware, of the embedded software, and of the environment. This difference makes hardware-software coverification not only a difficult task for most systems, but has also restricted coverification to the initial system specifications. Currently, most codesign tools or methodologies only support validation in the form of cosimulation and testing of design alternatives. Here, we propose a new formal coverification method based on linear hybrid automata. The basic problems found in most coverification tasks are presented and solved For complex systems, a simplification strategy is proposed to attack the state-space explosion occurring in formal coverification. Experimental results show the feasibility of our approach and the increase in verification scalability through the application of the proposed method.
Pao-Ann Hsiung
ECRTS1
1999 User-Friendly Verification
Pao-Ann Hsiung, Farn Wang
FORTE1
1999 Scheduling System Verification
Pao-Ann Hsiung, Farn Wang, Yue-Sun Kuo
TACAS1
1998 ICOS: an intelligent concurrent object-oriented synthesis methodology for multiprocessor systems
abstract
The design of multiprocessor architectures differs from uniprocessor systems in that the number of processors and their interconnection must be considered. This leads to an enormous increase in the design-space exploration time, which is exponential in the total number of system components. The methodology proposed here, called Intelligent Concurrent Object-Oriented Synthesis (ICOS) methodology, makes feasible the synthesis of complex multiprocessor systems through the application of several techiques that speed up the design process. ICOS is based on Performance Synthesis Methodology (PSM), a recently proposed object-oriented system-level design methodology. Four major techniques: object-oriented design, fuzzy design-space exploration, concurrent design, and intelligent reuse of complete subsystems are integrated in ICOS. First, object-oriented modeling and design, through the use of object-oriented relationships and operators, make the whole design process manageable and maintainable in ICOS. Second, fuzzy comparison applied to the specializations or instances of components reduces the exponential growth of design-space exploration in ICOS. Third, independent components from different design alternatives are synthesized in parallel; this design concurrency shortens the overall design time. Lastly, the resynthesis of complete subsystems can be avoided through the application of learning, thus making the methodology intelligent enough to reuse previous design configurations. Experiments show that all these applied techniques contribute to the synthesis efficiency and the degree of automation in ICOS.
Pao-Ann Hsiung, Chung-Hwang Chen, Trong-Yen Lee, Sao-Jie Chen
ACM Trans. Design Autom. Electr. Syst.1
1996 PSM: an object-oriented synthesis approach to multiprocessor system design
abstract
Although multiprocessor systems are becoming a trend today, few synthesis tools currently available can actually automate the design of multiprocessor systems. Performance synthesis methodology (PSM) is an object-oriented system-level synthesis approach to multiprocessor system design. Since PSM was designed specifically for the synthesis of multiprocessor systems, it is not only much more efficient when synthesizing parallel systems, but also produces better parallel systems than currently available uniprocessor system-level synthesis tools. Colored Petri nets used in modeling system components and object modeling technique used in the design process have both contributed to the shortening of system development time and to the reduction of design cost. First, user specification consisting of functional models and performance constraints is translated into architecture models. Then, the system is configured by selecting the method of control, the memory organization, the type of processor, and the type of system interconnection. Finally, a heuristic design space exploration algorithm is used to generate several near-optimal design alternatives. The best architecture is chosen by evaluating the design alternatives using a flexible performance estimation formula that mainly considers system level design features, such as system throughput, utilization, reliability, scalability, fault-tolerance, and cost. Several systems were successfully synthesized using this top-down object-oriented PSM, thus showing its feasibility as a design automation tool for parallel systems.
Pao-Ann Hsiung, Sao-Jie Chen, Tsung-Chien Hu, Shih-Chiang Wang
IEEE Trans. Very Large Scale Integr. Syst.1