EDBT 2026 Demo / reviewers in the wild / expert
Yixiao Yang
dblp:179/3243
· DBLP profile ↗
30ranked-venue papers
8as first author
22since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 11 · 10 since 2021Software engineering, systems software and programming languages · 7 · 5 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 1 first-author · 5 since 2021Artificial intelligence and machine learning · 5 · 3 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 4 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Theory of computation · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | The design and formal verification of a merging protocol for autonomous vehicles
Rui Wang 0024, Fang Qi, Yixiao Yang, Yi Wang 0003 |
Future Gener. Comput. Syst. | 4 |
| 2026 | Dual-Domain Fractional Fourier Transformer for Underwater Image Degradation RemovalabstractIn this letter, we propose DEFriT, a dual-domain framework for underwater image restoration targeting degradation caused by scattering and absorption. Supervised training is enabled by simulating underwater degradation from land images using a physics-based formation model. To address the spectrally non-stationary nature of underwater attenuation, we introduce the fractional Fourier transform (FrFT) to bridge spatial and spectral representations. A Fractional Transform Block (FrTB) with real-imaginary decomposition and Hybrid Time-Frequency Self-Attention (HTFSA) supports joint spatial-spectral modeling. Empirical analysis shows that degradation mainly affect the amplitude spectrum, with learned fractional orders concentrated in a narrow range ($p \in [0.48, 0.52]$). To the best of our knowledge, this is the first work to introduce FrFT into underwater image restoration and to systematically establish a fractional order prior. DEFriT achieves state-of-the-art performance on five benchmark datasets. Chuangxi Chen, Xudong Zhao 0003, Yixiao Yang, Ran Tao 0003, Binghua Su |
IEEE Signal Process. Lett. | 4 |
| 2026 | MFrodo: Efficient and Memory-Sensitive Simulink Code Generation via Redundancy EliminationabstractSimulink has emerged as the fundamental infrastructure that supports modeling, simulation, verification, and code generation for embedded software development. To improve the performance of the code generated from Simulink models, state-of-the-art code generators employ various optimization techniques, such as expression folding, variable reuse, and parallelism. However, they overlook the presence of redundant calculations within data-intensive models widely used to perform substantial data processing in embedded scenarios, which can significantly degrade the performance and introduce additional memory usage. This paper proposes MFRODO, an efficient and memory-sensitive code generator for data-intensive Simulink models through redundancy elimination. MFRODO begins by conducting model analysis to construct the dataflow graph and derive the I/O mapping of each block. Then, for each block within the dataflow graph, MFRODO recursively determines its calculation range by leveraging the I/O mapping of its subsequent blocks and marks optimization blocks whose calculation range is eliminated. For optimizable blocks, MFRODO eliminates the redundant calculations and reduces the memory space associated with these calculations. Finally, MFRODO rebuilds the I/O mappings of these optimizable blocks to ensure code correctness and synthesizes the embedded code for deployment. We implemented and evaluated MFRODO on benchmark Simulink models, in terms of execution duration, memory usage, and code generation overhead across different compilers and architectures. The results show that, compared with the Simulink Embedded Coder, DFSynth, and HCG, MFRODO achieves performance improvements ranging from 1.17× - 8.55×, while reducing BSS segment usage by 12.00% - 52.71%. Besides, MFRODO reduces compile time by 91.8% - 98.7% and code synthesis time by 94.3% - 99.6% compared with Simulink Embedded Coder, while incurring comparable overhead to DFSynth and HCG. Zehong Yu, Yixiao Yang, Zhuo Su 0005, Haowei Qiu, Rui Wang 0024, Aiguo Cui, Yu Jiang 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2025 | MonoInstance: Enhancing Monocular Priors via Multi-view Instance Alignment for Neural Rendering and ReconstructionabstractMonocular depth priors have been widely adopted by neural rendering in multi-view based tasks such as 3D reconstruction and novel view synthesis. However, due to the inconsistent prediction on each view, how to more effectively leverage monocular cues in a multi-view context remains a challenge. Current methods treat the entire estimated depth map indiscriminately, and use it as ground truth supervision, while ignoring the inherent inaccuracy and cross-view inconsistency in monocular priors. To resolve these issues, we propose MonoInstance, a general approach that explores the uncertainty of monocular depths to provide enhanced geometric priors for neural rendering and reconstruction. Our key insight lies in aligning each segmented instance depths from multiple views within a common 3D space, thereby casting the uncertainty estimation of monocular depths into a density measure within noisy point clouds. For high-uncertainty areas where depth priors are unreliable, we further introduce a constraint term that encourages the projected instances to align with corresponding instance masks on nearby views. MonoInstance is a versatile strategy which can be seamlessly integrated into various multi-view neural rendering frameworks. Our experimental results demonstrate that MonoInstance significantly improves the performance in both reconstruction and novel view synthesis under various benchmarks. Project page: https://wen-yuan-zhang.github.io/MonoInstance/. Yixiao Yang, Kanle Shi, Yu-Shen Liu, Zhizhong Han |
CVPR | 2 |
| 2025 | Self-supervised Underwater Color Restoration via Wavelet-Diffusion Model with Filtered Multi-Scale Feature DistillationabstractExisting underwater image processing methods often struggle due to the limited availability of real paired training data. Models trained on public datasets frequently fail to generalize across diverse underwater conditions and produce suboptimal color restoration. To address these challenges, we propose a self-supervised underwater color restoration framework based on a Wavelet-Diffusion Model with Filtered Multi-Scale Feature Distillation. Specifically, we introduce a wavelet-diffusion training paradigm on terrestrial images, guided by a stochastic underwater imaging model prior. This randomized control enables the model to learn diverse underwater imaging processes, facilitating effective generalization to real-world underwater images and achieving precise color restoration. Furthermore, to tackle feature entanglement in zero-shot domain generalization and mitigate the slow sampling and partial corruption issues of diffusion models, We integrate a Mamba-based U-shaped student network for multi-scale feature distillation. Additionally, we introduce a filtering mechanism to refine the diffusion sampled features, allowing the student model to outperform the teacher in both performance and image quality. Extensive experiments across multiple underwater datasets demonstrate that our approach effectively restores natural colors, eliminating water-induced distortions while achieving state-of-the-art performance in both qualitative and quantitative evaluations. Code and data for this paper are at https://github.com/zx826/FMFD Yixiao Yang, Haijun Xie, Haowen Yan, Hexiang Zhai, Binghua Su |
SIGGRAPH Asia | 3 |
| 2025 | Knight: Optimizing Code Generation for Simulink Models With Loop ReshapingabstractSimulink has become a pivotal infrastructure in embedded scenarios, including automotive systems and aerospace designs. To improve the performance of the code generated from Simulink models, state-of-the-art code generators employ various optimization techniques, such as expression folding, variable reuse, and parallelism. However, they struggle to generate efficient code for loop-semantic models which are crucial in substantial data processing tasks. This inefficiency manifests in numerous redundant calculations, such as array calculations and conditional statements. As a result, the performance of the generated code is limited. This article proposes Knight, an efficient code generator for loop-semantic Simulink models with loop reshaping. Knight first parses the Simulink model to extract essential content, such as block functionalities and connections. Knight then identifies blocks with internal states and implements the specific interaction rules to discern those that are state-dependent. For state-dependent blocks, Knight conducts forward inference to obtain their preceding blocks, which influence the internal state calculations. Subsequently, Knight isolates blocks that are optimizable and irrelevant to the internal states. These blocks are strategically relocated outside the loop semantics while preserving critical semantics related to code generation. We implemented and evaluated Knight on benchmark Simulink models across different compilers and architectures. Compared with the state-of-the-art code generators Simulink Embedded Coder, DFSynth, and HCG, the code generated by Knight is$16.58 \times $faster,$16.89 \times $faster, and$15.38 \times $faster in terms of execution duration on average, without incurring additional overhead of memory usage. Zehong Yu, Yixiao Yang, Zhuo Su 0005, Rui Wang 0024, Yu Jiang 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2025 | Chirplet Fourier Analysis Network for Cross-Scene Classification of Multisource Remote Sensing DataabstractThe joint application of multisource remote sensing (MSRS) data, such as hyperspectral image (HSI) and light detection and ranging (LiDAR), offers significant potential for accurate land cover classification. However, the existing applications often struggle with domain shifts across scenes caused by sensor, illumination, and phase variations. Focusing on this domain adaptation problem, a Chirplet Fourier analysis network (ChirpFAN) is proposed for cross-scene classification of MSRS data in this paper. Firstly, a fractional spatial-frequency-phase feature extraction module including the fractional Fourier transform and a learnable phase-aware weighting block is proposed to capture multi-domain features. Secondly, a Chirplet swin transformer (ChirpST) block integrates a Chirplet Fourier analysis (ChirpFA) layer within a Swin transformer is designed to analyze multi-scale textural and oscillatory patterns. Finally, a modality-shared network including ChirpST blocks is designed for inter-modal fusion and alignment. Extensive experiments demonstrate that the ChirpFAN framework achieves state-of-the-art performance with 3% average improvements on three challenging cross-scene MSRS datasets. Code will be released on GitHub. Xudong Zhao 0003, Qi Ming, Yixiao Yang, Wen-Shuai Hu, Wei Li 0032, Ran Tao 0003 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2024 | SQLPass: A Semantic Effective Fuzzing Method for DBMSabstractFuzzing, as an effective method for software system defect testing and bug mining, utilizes random data generated by mutation to execute programs to trigger potential bugs in the tested program. However, due to the strict syntax and semantic checks in DBMSs, existing mutation-based testing methods are difficult to generate test cases with correct syntax and semantics. To ensure the syntax and semantic correctness of the test case generated by mutation, this paper proposes the concepts of weak semantic correlation nodes and semantic relationship tables. At the same time, a set of mutation operators for weak semantic correlation nodes in the syntax tree is designed, and each time the “optimal” mutation operator is selected to replace, delete, and insert the target node subtree in the test case syntax tree. In response to semantic errors during the mutation, this paper checks and corrects them through a pre-extracted semantic relationship table, further ensuring the correctness of the semantics of the test cases generated by the mutation. This paper utilizes this semantic effective fuzzing method for DBMSs to implement a new fuzzing framework for DBMSs, SQLPass. We evaluated SQLPass on four popular DBMSs: SQLite3, MySQL, MariaDB and PostgreSQL. In our experiment, SQLPass achieves 5.7%-94.2% higher semantic correctness than state-of-the-art tools, and explores 1.3%-52% more code coverage than other tools. Notably, it discovered four unknown bugs on SQLite3 and MariaDB and has submitted them to the vendor for confirmation. Yixiao Yang, Zhi-Ping Shi 0002, Rui Wang 0024 |
COMPSAC | 2 |
| 2024 | Spectrum analysis for nonuniform sampling of bandlimited and multiband signals in the fractional Fourier domain
Yixiao Yang, Ran Tao 0003, Gang Li 0008, Chang Gao 0005 |
Signal Process. | 2 |
| 2024 | HSTCG: State-Aware Simulink Model Test Case Generation With Heuristic StrategyabstractSimulink has gained widespread recognition as a valuable tool for system design. As systems grow increasingly complex, particularly in terms of their internal states, this complexity poses new challenges for existing model testing methodologies. Traditional techniques such as constraint solving and random search encounter difficulties when attempting to explore the intricate logic embedded within these models. In this paper, we introduceHSTCG, a state-aware test case generation method for Simulink models with heuristic strategy.HSTCGsolves only one iteration of the model each time to get the test input that can cover a target branch, then executes the model once to obtain and update the new model state based on the solved input dynamically. Then, it solves the remaining branches based on the new model state iteratively until all the coverage requirements are satisfied. To improve the efficiency of test case generation, we also designed a heuristic strategy containing heuristic branch searching, repeated state filter and unreached branch filter to minimize the times of constraint solving. We implementedHSTCGand evaluated it on several benchmark Simulink models. Compared to the built-in Simulink Design Verifier and state-of-the-art academic work SimCoTest,HSTCGachieves an average improvement of 55% and 103% on Decision Coverage, 53% and 62% on Condition Coverage and 192% and 201% on Modified Condition Decision Coverage, respectively. We also validated the significant improvement of the heuristic strategy, which can improve the efficiency of test case generation by 62.2% on average. Zhuo Su 0005, Zehong Yu, Dongyan Wang, Yixiao Yang, Rui Wang 0024, Wanli Chang 0001, Aiguo Cui, Yu Jiang 0001 |
IEEE Trans. Software Eng. | 4 |
| 2023 | STCG: State-Aware Test Case Generation for Simulink ModelsabstractSimulink has been widely used in system design, which supports the efficient modeling and synthesis of embedded controllers, with automatic test case generation to simulate and validate the correctness of the constructed Simulink model. However, the increasing complexity of the model, especially the internal states, brings extra challenges to existing model testing techniques such as constraint solving and random search, which results in difficulties when trying to reach the deeper logic of the model effectively.In this paper, we propose STCG, a state-aware test case generation method for Simulink models. STCG solves only one iteration of the model each time to get the test input that can cover a target branch, then executes the model once to obtain and update the novel model state based on the solved input dynamically. Then, it solves the remaining branches based on the new model state iteratively until all the coverage requirements are satisfied. We implemented STCG and evaluated it on several benchmark Simulink models. Compared to the built-in Simulink Design Verifier and state-of-the-art academic work SimCoTest, STCG achieves an average improvement of 58% and 132% on Decision Coverage, 52% and 70% on Condition Coverage and 239% and 237% on Modified Condition Decision Coverage, respectively. Zhuo Su 0005, Zehong Yu, Dongyan Wang, Yixiao Yang, Rui Wang 0024, Wanli Chang 0001, Aiguo Cui, Yu Jiang 0001 |
DAC | 4 |
| 2023 | Single-Shot Fractional Fourier Phase RetrievalabstractTraditional phase retrieval is generally concerned with re-covering a signal from its Fourier magnitude measurements whose inherent ambiguities make this problem especially difficult. In this work, we present an efficient phase retrieval technique from the single fractional Fourier transform (FrFT) magnitude measurement. Specifically, the FrFT measurement can be well-combined with signal priors via a generalized alternating projection framework, which can effectively alleviate the ambiguities of phase retrieval and the stagnation problem of numerical iterative processes. Through numerical simulations, we demonstrate that reconstructing an image from the single FrFT measurement leads to a significant performance improvement over that from the Fourier transform magnitude by using the proposed method. The source code is available at https://github.com/Yixiao-Yang/SFrFPR. Yixiao Yang, Ran Tao 0003 |
ICASSP | 1 |
| 2023 | Binary Level Concolic Execution on Windows with Rich Instrumentation Based Taint Analysis
Yixiao Yang, Rui Wang 0024 |
SETTA | 1 |
| 2023 | PHCG: Optimizing Simulink Code Generation for Embedded System With SIMD InstructionsabstractSimulink is widely used for the model-driven design of embedded systems. It is able to generate optimized embedded control software code through expression folding, variable reuse, etc. However, for some commonly used computing-sensitive models, such as the models for signal processing applications, the efficiency of the generated code is still limited. In this article, we propose PHCG, an optimized code generator for the Simulink model with single-instruction–multiple-data (SIMD) instruction synthesis. It will select the optimal implementations for intensive computing actors based on adaptively precalculation of the input scales, and synthesize the appropriate SIMD instructions for batch computing actors based on the iterative dataflow graph mapping. In addition, actors of the same type that can be executed in parallel can be combined into batch computing actors as much as possible by merging isomorphic subgraphs. We implemented and evaluated its performance on benchmark Simulink models. Compared to the built-in Simulink Coder and the most recent DFSynth, the code generated by PHCG achieves an improvement of 38.9%–92.9% and 41.2%–76.8% in terms of execution time across different architectures and compilers, respectively. Zhuo Su 0005, Dongyan Wang, Zehong Yu, Yixiao Yang, Yu Jiang 0001, Rui Wang 0024, Wanli Chang 0001, Aiguo Cui, Jia-Guang Sun 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2022 | HCG: optimizing embedded code generation of simulink with SIMD instruction synthesisabstractSimulink is widely used for the model-driven design of embedded systems. It is able to generate optimized embedded control software code through expression folding, variable reuse, etc. However, for some commonly used computing-sensitive models, such as the models for signal processing applications, the efficiency of the generated code is still limited. Zhuo Su 0005, Zehong Yu, Dongyan Wang, Yixiao Yang, Yu Jiang 0001, Rui Wang 0024, Wanli Chang 0001, Jia-Guang Sun 0001 |
DAC | 4 |
| 2022 | Retinex-Based Low-Light Hyperspectral Restoration Using Camera Response ModelabstractSpectral quality is one of the most critical issues that has to be considered in real hyperspectral image (HSI) application. Denoising, destriping, inpainting, deblurring and super-resolution are common techniques to improve the quality of HSIs from different aspects. These techniques have attracted much attention that a diversity of methods, algorithms, tools have been well developed to facilitate the development of HSI restoration. Although effectively improving the quality of HSIs, these technologies mainly focus on recovering an HSI captured in the normal sunlight. It is acknowledged that HSIs are captured via passive imaging mechanisms covering the spectral bands from visible& near-infrared to shortwave infrared spectral range (i.e., around 400nm to 2500nm). The imaging condition limits HSI spectrometers to capture HSIs without sunlight (e.g., in dark environments or night time). In this work, a low-light HSI restoration method is proposed, where we borrow idea of intrinsic decomposition based on Retinex theory in natural image low-light enhancement. Additionally, camera response function that describe the spectral degradation of RGB image and relationship between irradiance and pixel values are employed, respectively. The experimental results validate the effectiveness of the proposed method. Na Liu 0014, Yinjian Wang, Yixiao Yang, Wei Li 0032, Ran Tao 0003 |
IGARSS | 3 |
| 2022 | Dynamic proximal unrolling network for compressive imaging
Yixiao Yang, Ran Tao 0003, Kaixuan Wei, Ying Fu 0001 |
Neurocomputing | 1 |
| 2022 | Code Synthesis for Dataflow-Based Embedded Software DesignabstractModel-driven methodology has been widely adopted in embedded software design, and Dataflow is a widely used computation model, with strong modeling and simulation ability supported in tools such as Ptolemy. However, its code synthesis support is quite limited, which restricts its applications in real industrial practice. In this article, we focus on the automatic code synthesis of Dataflow, and implementDFSynth, a code generator that could support most of the widely used modeling features, such as the expression type and Boolean switch, more efficiently. First, we disassemble the Dataflow model into actors embedded in if-else or switch-case statements based on the schedule analysis, which bridges the semantic gap between the code and the original Dataflow model. Then, we design well-designed templates for each actor, and synthesize well-structured executable C and Java codes with sequential code assembly. Compared to the existing C and Java code generators of Dataflow model in Ptolemy-II, and the C code generator in Simulink, the lines of code synthesized byDFSynthare decreased by an average of 99.7%, 81.4%, and 61.9%, and the execution time of the synthesized code byDFSynthis also decreased by an average of 76.2%, 56.8%, and 22.7%, respectively. Zhuo Su 0005, Dongyan Wang, Yixiao Yang, Yu Jiang 0001, Wanli Chang 0001, Liming Fang 0001, Jia-Guang Sun 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2022 | MDD: A Unified Model-Driven Design Framework for Embedded Control SoftwareabstractModel-driven methods are widely used in embedded control software development. Current design tools, such as Ptolemy-II and Simulink, have strong modeling capability but their simulation and code generation functionalities are challenged by the increasing complexity of control requirements. For simulation, emulating the triggering of the actor leads to additional time overhead and speed degradation. For code generation, generating redundant content degrades the code quality. Besides, current tools do not have a unified interface, which makes it difficult to cooperation. In this article, we propose a unified model-driven design framework MDD to facilitate embedded control software development. MDD can support the unification of models built by different modeling tools for high-efficiency simulation and high-quality code generation. The MDD framework supports the expansion of more modeling tools, and also supports the expansion of more uses, such as unified testing and verification. First, it offers a model intermediate representation (MIR) and several corresponding parsers, which facilitate a unified representation and cooperation for different design tools. Then, based on data flow schedule analysis of the original MIR, intermediate code representation will be generated for optimized code synthesis. Finally, a variety of code translators will synthesize the intermediate code representation into the code of actual use, such as code for simulation and code for deployment. For evaluation, we enhance two widely used design tools in industry, Ptolemy-II and Simulink, and apply them on the implementation of several benchmark models and a real-world self-driving control software of our industrial collaborator. Using MDD can help reduce their simulation time by 98.9% and 92.6%, the generated code by 99.7% and 69.9% in the number of lines, and 94.3% and 34.3% in code execution time, respectively. Zhuo Su 0005, Dongyan Wang, Yixiao Yang, Zehong Yu, Wanli Chang 0001, Aiguo Cui, Yu Jiang 0001, Jia-Guang Sun 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2022 | Mercury: Instruction Pipeline Aware Code Generation for Simulink ModelsabstractSimulink is a widely used model-driven design environment for supporting the simulation and code generation of embedded applications. To improve the quality of the code generated from Simulink models, state-of-the-art code generators employ various high-level optimizations, like eliminating local variables. However, they overlook the compatibility between code and the low-level processor architecture, especially the instruction pipeline. Consequently, instruction pipeline stalls occur frequently, leading to additional delays in instruction execution, as well as limited efficiency for deployed the embedded software. In this article, we propose Mercury, an instruction pipeline aware code generator for Simulink models which utilizes data dependencies between actors to decrease the instruction pipeline stalls of the generated code. First, Mercury collects data dependencies through model dataflow traversal and records the property of each actor. Then, Mercury approximately estimates the execution latency of required instructions fetched from corresponding actors and uses a topology-based method to obtain candidate actors for code synthesis. Finally, Mercury adopts the least penalty priority to iteratively select the most suitable actor for code synthesis and releases data dependencies with its subsequent actors. We implemented and evaluated Mercury on benchmark Simulink models (Su et al., 2021) as well as a real industrial model. Compared to the official tool Simulink Embedded Coder and the state-of-the-art academic tool DFSynth, Mercury outperformed them by 9.7%–33.4% and 9.2%–59.4% in terms of the execution time of the generated code across different architectures, respectively. The statistics also demonstrate that the generated code of Mercury increases utilization of pipeline slots by 11.0%–37.1% and 10.6%–50.0%, respectively. Zehong Yu, Zhuo Su 0005, Yixiao Yang, Jie Liang 0006, Yu Jiang 0001, Aiguo Cui, Wanli Chang 0001, Rui Wang 0024 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2021 | Rtkaller: State-aware Task Generation for RTOS FuzzingabstractA real-time operating system (RTOS) is an operating system designed to meet certain real-time requirements. It is widely used in embedded applications, and its correctness is safety-critical. However, the validation of RTOS is challenging due to its complex real-time features and large code base. In this paper, we propose Rtkaller , a state-aware kernel fuzzer for the vulnerability detection in RTOS. First, Rtkaller implements an automatic task initialization to transform the syscall sequences into initial tasks with more real-time information. Then, a coverage-guided task mutation is designed to generate those tasks that explore more in-depth real-time related code for parallel execution. Moreover, Rtkaller realizes a task modification to correct those tasks that may hang during fuzzing. We evaluated it on recent versions of rt-Linux, which is one of the most widely used RTOS. Compared to the state-of-the-art kernel fuzzers Syzkaller and Moonshine, Rtkaller achieves the same code coverage at the speed of 1.7X and 1.6X, gains an increase of 26.1% and 22.0% branch coverage within 24 hours respectively. More importantly, Rtkaller has confirmed 28 previously unknown vulnerabilities that are missed by other fuzzers. Yuheng Shen, Hao Sun 0021, Yu Jiang 0001, Heyuan Shi, Yixiao Yang, Wanli Chang 0001 |
ACM Trans. Embed. Comput. Syst. | 5 |
| 2021 | Formal Design of Multi-Function Vehicle Bus ControllerabstractData of the train communication network(TCN) is becoming more complicated, which results in higher requirements of the data processing unit-the multifunction vehicle bus controller (MVBC) connected within the TCN. Developing an MVBC is challenging because of the integrated hardware-software solutions to support reactions in real time and dynamic environment. Hence, there is an urgent need for a rigorous design framework to facilitate the development of MVBC. In this paper, we propose a design framework TooMVBC to generate executable MVBC code from formal verified computation model. TooMVBC uses formal computation model MVBChart to capture the specification of the MVBC at high level. First, primitive syntax of MVBChart is designed to model MVBC features (e.g. hierarchy structure, data flow of the encoder, the control logic of communication protocol), and semantics of MVBChart is formalized for simulation and verification. Then, semantics-preserving code generation algorithms are designed to generate VHDL code for partitioned hardware implementations and C code for partitioned software implementations from verified MVBChart model. The generated code can be loaded into the proposed flexible MVBC hardware architecture directly. Finally, supporting graphical model editor, simulator, verification translator, partitioning and code generator are implemented and seamlessly integrated into TooMVBC. When we apply TooMVBC to design MVBC with the highest class 5 according to the description of the standard IEC 61375, several critical ambiguousness or bugs in the standard are detected during formal verification of the constructed system model. Yu Jiang 0001, Zhuo Su 0005, Yixiao Yang, Huihui Wang 0001 |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2020 | Single Image Reflection Removal Through Cascaded RefinementabstractWe address the problem of removing undesirable reflections from a single image captured through a glass surface, which is an ill-posed, challenging but practically important problem for photo enhancement. Inspired by iterative structure reduction for hidden community detection in social networks, we propose an Iterative Boost Convolutional LSTM Network (IBCLN) that enables cascaded prediction for reflection removal. IBCLN is a cascaded network that iteratively refines the estimates of transmission and reflection layers in a manner that they can boost the prediction quality to each other, and information across steps of the cascade is transferred using an LSTM. The intuition is that the transmission is the strong, dominant structure while the reflection is the weak, hidden structure. They are complementary to each other in a single image and thus a better estimate and reduction on one side from the original image leads to a more accurate estimate on the other side. To facilitate training over multiple cascade steps, we employ LSTM to address the vanishing gradient problem, and propose residual reconstruction loss as further training guidance. Besides, we create a dataset of real-world images with reflection and ground-truth transmission layers to mitigate the problem of insufficient data. Comprehensive experiments demonstrate that the proposed method can effectively remove reflections in real and synthetic images compared with state-of-the-art reflection removal methods. Chao Li 0068, Yixiao Yang, Kun He 0001, Stephen Lin 0001, John E. Hopcroft |
CVPR | 2 |
| 2019 | Improve Language Modelling for Code Completion through Learning General Token Repetition of Source CodeabstractIn last few years, to solve the problem of code completion, using a language model such as LSTM to learn code token sequences is the state-of-art method.However, tokens in source code are more repetitive than words in natural languages.For example, once a variable is declared in a program, it may be used many times.Other elements such as generic types in templates also occur repeatedly.It is important to capture token repetition of code.For example, if usage patterns of variables are not captured, there is little chance for a model trained on one project to predict the name of an unseen variable in another project correctly.Capturing token repetition of source code is challenging because not only the repeated token but also the place at where the repetition should happen must be both decided at the same time.Hence, we propose a novel deep neural model named REP to capture the general token repetition of source code.The repetitions of code tokens are modeled as edges connecting between repeated tokens on a graph.The REP model is essentially a deep neural graph generation model.The experiments indicate that the proposed model outperforms stateof-arts in code completion. Yixiao Yang, Chen Xiang |
SEKE | 1 |
| 2019 | Improve Language Modelling for Code Completion by Tree Language Model with Tree Encoding of Context (S)abstractIn last few years, using a language model such as LSTM to train code token sequences is the state-of-art to get a code generation model.However, source code can be viewed not only as a token sequence but also as a syntax tree.Treating all source code tokens equally will lose valuable structural information.Recently, in code synthesis tasks, tree models such as Seq2Tree and Tree2Tree have been proposed to generate code and those models perform better than LSTMbased seq2seq methods.In those models, encoding model encodes user-provided information such as the description of the code, and decoding model decodes code based on the encoding results of user-provided information.When applying decoding model to decode code, current models pay little attention to the context of the already decoded code.According to experiments, using tree models to encode the already decoded code and predicting next code based on tree representations of the already decoded code can improve the decoding performance.Thus, in this paper, we propose a novel tree language model (TLM) which predicts code based on a novel tree encoding of the already decoded code (context).The experiments indicate that the proposed method outperforms state-of-arts in code completion. Yixiao Yang, Chen Xiang |
SEKE | 1 |
| 2019 | Improve Language Modeling for Code Completion Through Learning General Token Repetition of Source Code with Optimized MemoryabstractIn last few years, applying language model to source code is the state-of-the-art method for solving the problem of code completion. However, compared with natural language, code has more obvious repetition characteristics. For example, a variable can be used many times in the following code. Variables in source code have a high chance to be repetitive. Cloned code and templates, also have the property of token repetition. Capturing the token repetition of source code is important. In different projects, variables or types are usually named differently. This means that a model trained in a finite data set will encounter a lot of unseen variables or types in another data set. How to model the semantics of the unseen data and how to predict the unseen data based on the patterns of token repetition are two challenges in code completion. Hence, in this paper, token repetition is modelled as a graph, we propose a novel REP model which is based on deep neural graph network to learn the code toke repetition. The REP model is to identify the edge connections of a graph to recognize the token repetition. For predicting the token repetition of token [Formula: see text], the information of all the previous tokens needs to be considered. We use memory neural network (MNN) to model the semantics of each distinct token to make the framework of REP model more targeted. The experiments indicate that the REP model performs better than LSTM model. Compared with Attention-Pointer network, we also discover that the attention mechanism does not work in all situations. The proposed REP model could achieve similar or slightly better prediction accuracy compared to Attention-Pointer network and consume less training time. We also find other attention mechanism which could further improve the prediction accuracy. Yixiao Yang, Jia-Guang Sun 0001 |
Int. J. Softw. Eng. Knowl. Eng. | 1 |
| 2019 | Dependable Model-driven Development of CPS: From Stateflow Simulation to Verified ImplementationabstractSimulink is widely used for model-driven development (MDD) of cyber-physical systems. Typically, the Simulink-based development starts with Stateflow modeling, followed by simulation, validation, and code generation mapped to physical execution platforms. However, recent trends have raised the demands of rigorous verification on safety-critical applications to prevent intrinsic development faults and improve the system dependability, which is unfortunately challenging. Even though the constructed Stateflow model and the generated code pass the validation of Simulink Design Verifier and Simulink Polyspace, respectively, the system may still fail due to some implicit defects contained in the design model (design defect) and the generated code (implementation defects). In this article, we bridge the Stateflow-based MDD and a well-defined rigorous verification to reduce development faults. First, we develop a self-contained toolkit to translate a Stateflow model into timed automata, where major advanced modeling features in Stateflow are supported. Taking advantage of the strong verification capability of Uppaal, we can not only find bugs in Stateflow models that are missed by Simulink Design Verifier but also check more important temporal properties. Next, we customize a runtime verifier for the generated non-intrusive VHDL and C code of a Stateflow model for monitoring. The major strength of the customization is the flexibility to collect and analyze runtime properties with a pure software monitor, which offers more opportunities for engineers to achieve high reliability of the target system compared with the traditional act that only relies on Simulink Polyspace. In this way, safety-critical properties are both verified at the model level and at the consistent system implementation level with physical execution environment in consideration. We apply our approach to the development of a typical cyber-physical system-train communication controller based on the IEC standard 61375. Experiments show that more ambiguousness in the standard are detected and confirmed and more development faults and those corresponding errors that would lead to system failure have been removed. Furthermore, the verified implementation has been deployed on real trains. Yu Jiang 0001, Houbing Song, Yixiao Yang, Han Liu 0010, Ming Gu 0001, Jia-Guang Sun 0001, Lui Sha |
ACM Trans. Cyber Phys. Syst. | 3 |
| 2017 | A language model for statements of software codeabstractBuilding language models for source code enables a large set of improvements on traditional software engineering tasks. One promising application is automatic code completion. State-of-the-art techniques capture code regularities at token level with lexical information. Such language models are more suitable for predicting short token sequences, but become less effective with respect to long statement level predictions. In this paper, we have proposed PCC to optimize the token-level based language modeling. Specifically, PCC introduced an intermediate representation (IR) for source code, which puts tokens into groups using lexeme and variable relative order. In this way, PCC is able to handle long token sequences, i.e., group sequences, to suggest a complete statement with the precise synthesizer. Further more, PCC employed a fuzzy matching technique which combined genetic and longest common subsequence algorithms to make the prediction more accurate. We have implemented a code completion plugin for Eclipse and evaluated it on open-source Java projects. The results have demonstrated the potential of PCC in generating precise long statement level predictions. In 30%-60% of the cases, it can correctly suggest the complete statement with only six candidates, and 40%-90% of the cases with ten candidates. Yixiao Yang, Yu Jiang 0001, Ming Gu 0001, Jia-Guang Sun 0001, Jian Gao 0008, Han Liu 0010 |
ASE | 1 |
| 2016 | Verifying simulink stateflow model: timed automata approachabstractSimulink Stateflow is widely used for the model-driven development of software. However, the increasing demand of rigorous verification for safety critical applications brings new challenge to the Simulink Stateflow because of the lack of formal semantics. In this paper, we present STU, a self-contained toolkit to bridge the Simulink Stateflow and a well-defined rigorous verification. The tool translates the Simulink Stateflow into the Uppaal timed automata for verification. Compared to existing work, more advanced and complex modeling features in Stateflow such as the event stack, conditional action and timer are supported. Then, with the strong verification power of Uppaal, we can not only find design defects that are missed by the Simulink Design Verifier, but also check more important temporal properties. The evaluation on artificial examples and real industrial applications demonstrates the effectiveness. Yixiao Yang, Yu Jiang 0001, Ming Gu 0001, Jia-Guang Sun 0001 |
ASE | 1 |
| 2016 | From Stateflow Simulation to Verified Implementation: A Verification Approach and A Real-Time Train Controller DesignabstractSimulink is widely used for model driven development (MDD) of industrial software systems. Typically, the Simulink based development is initiated from Stateflow modeling, followed by simulation, validation and code generation mapped to physical execution platforms. However, recent industrial trends have raised the demands of rigorous verification on safety-critical applications, which is unfortunately challenging for Simulink. In this paper, we present an approach to bridge the Stateflow based model driven development and a well- defined rigorous verification. First, we develop a self- contained toolkit to translate Stateflow model into timed automata, where major advanced modeling features in Stateflow are supported. Taking advantage of the strong verification capability of Uppaal, we can not only find bugs in Stateflow models which are missed by Simulink Design Verifier, but also check more important temporal properties. Next, we customize a runtime verifier for the generated nonintrusive VHDL and C code of Stateflow model for monitoring. The major strength of the customization is the flexibility to collect and analyze runtime properties with a pure software monitor, which opens more opportunities for engineers to achieve high reliability of the target system compared with the traditional act that only relies on Simulink Polyspace. We incorporate these two parts into original Stateflow based MDD seamlessly. In this way, safety-critical properties are both verified at the model level, and at the consistent system implementation level with physical execution environment in consideration. We apply our approach on a train controller design, and the verified implementation is tested and deployed on a real hardware platform. Yu Jiang 0001, Yixiao Yang, Han Liu 0010, Hui Kong 0004, Ming Gu 0001, Jia-Guang Sun 0001, Lui Sha |
RTAS | 2 |