EDBT 2026 Demo / reviewers in the wild / expert
Boyu Chang
dblp:264/5312
· DBLP profile ↗
8ranked-venue papers
2as first author
8since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Security and privacy · 3 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | AbductiveMLLM: Boosting Visual Abductive Reasoning Within MLLMsabstractVisual abductive reasoning (VAR) is a challenging task that requires AI systems to infer the most likely explanation for incomplete visual observations. While recent MLLMs develop strong general-purpose multimodal reasoning capabilities, they remain fall short in abductive inference, as compared to human beings. To bridge this gap, we draw inspiration from the interplay between verbal and pictorial abduction in human cognition, and propose to strengthen abduction of MLLMs by mimicking such dual-mode behavior. Concretely, we introduce AbductiveMLLM comprising of two synergistic components: REASONER and IMAGINER. The REASONER operates in the verbal domain. It first explores a broad space of possible explanations using a blind LLM and then prunes visually incongruent hypotheses based on cross-modal causal alignment. The remaining hypotheses are introduced into the MLLM as targeted priors, steering its reasoning toward causally coherent explanations. The IMAGINER, on the other hand, further guides MLLMs by emulating human-like pictorial thinking. It conditions a text-to-image diffusion model on both the input video and the REASONER’s output embeddings to “imagine” plausible visual scenes that correspond to verbal explanation, thereby enriching MLLMs' contextual grounding. The two components are trained jointly in an end-to-end manner. Experiments on standard VAR benchmarks show that AbductiveMLLM achieves state-of-the-art performance, consistently outperforming traditional solutions and advanced MLLMs. Boyu Chang, Qi Wang 0009, Zhixiong Nan, Yazhou Yao, Tianfei Zhou |
AAAI | 1 |
| 2025 | SyzOrch: An Orchestration Framework for Resource-Aware and Composable Kernel FuzzingabstractKernel fuzzing plays a critical role in uncovering vulnerabilities, reproducing bugs, and testing patches in operating systems. While integrating external resources such as symbolic execution engines, static analyzers, and language models has proven effective in areas such as enhancing path exploration, optimizing seed generation, and improving seed mutation, existing approaches remain tightly coupled and task-specific, hindering the reuse, migration, scheduling, and composition of these external resources. This limitation further restricts the ability of researchers to explore flexible hybrid fuzzing strategies and hinders industry efforts to build stronger and more adaptable kernel fuzzers. We present SyzOrch to address this limitation. SyzOrch (1) decouples the kernel fuzzing workflow; (2) provides event-driven coordination between external resources and the fuzzer; (3) abstracts heterogeneous external resources through a generalized behavior model; and (4) supports user-defined dynamic control via a programmable DSL runner. We evaluate SyzOrch across diverse kernel fuzzing scenarios and show that it achieves a 30% speedup of directed kernel fuzzing by migrating existing techniques, improves coverage by 8.6% through hybrid composition with multiple external resources, and discovers previously unknown kernel bugs, including one assigned a CNNVD identifier. These results demonstrate SyzOrch’s effectiveness in orchestrating external resources to enhance kernel fuzzing. Lukai Xu, Bo Yu 0008, Boyu Chang, Shouling Ji, Danjun Liu, Lei Zhou 0023, Yaojia Yang |
ISSRE | 4 |
| 2025 | Firmrca: Towards Post-Fuzzing Analysis on ARM Embedded Firmware with Efficient Event-Based Fault LocalizationabstractWhile fuzzing has demonstrated its effectiveness in exposing vulnerabilities within embedded firmware, the discovery of crashing test cases is only the first step in improving the security of these critical systems. The subsequent fault localization process, which aims to precisely identify the root causes of observed crashes, is a crucial yet time-consuming post-fuzzing work. Unfortunately, the automated root cause analysis on embedded firmware crashes remains an underexplored area, which is challenging from several perspectives: (1) the fuzzing campaign towards the embedded firmware lacks adequate debugging mechanisms, making it hard to automatically extract essential runtime information for analysis; (2) the inherent raw binary nature of embedded firmware often leads to over-tainted and noisy suspicious instructions, which provides limited guidance for analysts in manually investigating the root cause and remediating the underlying vulnerability. To address these challenges, we design and implement FirmRCA, a practical fault localization framework tailored specifically for embedded firmware. FirmRCA introduces an event-based footprint collection approach that leverages concrete memory accesses in the crash reproducing process to aid and significantly expedite reverse execution. Next, to solve the complicated memory alias problem, FirmRCA proposes a history-driven method by tracking data propagation through the execution trace, enabling precise identification of deep crash origins. Finally, FirmRCA proposes a novel strategy to highlight key instructions related to the root cause, providing practical guidance in the final investigation. To demonstrate the efficacy of FirmRCA, we evaluate it with both synthetic and real-world targets, including 41 crashing test cases across 17 firmware images. The results show that FIRMRCA can effectively (92.7% success rate) identify the root cause of crashing test cases within the top 10 instructions. Compared to state-of-the-art works, FIRMRCA demonstrates its superiority in 27.8% improvement in full execution trace analysis capability, polynomial-level acceleration in overall efficiency and 73.2% higher success rate within the top 10 instructions in effectiveness. Boyu Chang, Peiyu Liu 0003, Yuan Tian 0001, Raheem A. Beyah, Shouling Ji |
SP | 1 |
| 2024 | SyzTrust: State-aware Fuzzing on Trusted OS Designed for IoT DevicesabstractTrusted Execution Environments (TEEs) embedded in IoT devices provide a deployable solution to secure IoT applications at the hardware level. By design, in TEEs, the Trusted Operating System (Trusted OS) is the primary component. It enables the TEE to use security-based design techniques, such as data encryption and identity authentication. Once a Trusted OS has been exploited, the TEE can no longer ensure security. However, Trusted OSes for IoT devices have received little security analysis, which is challenging from several perspectives: (1) Trusted OSes are closed-source and have an unfavorable environment for sending test cases and collecting feedback. (2) Trusted OSes have complex data structures and require a stateful workflow, which limits existing vulnerability detection tools.To address the challenges, we present SyzTrust, the first state-aware fuzzing framework for vetting the security of resource-limited Trusted OSes. SyzTrust adopts a hardware-assisted framework to enable fuzzing Trusted OSes directly on IoT devices as well as tracking state and code coverage non-invasively. SyzTrust utilizes composite feedback to guide the fuzzer to effectively explore more states as well as to increase the code coverage. We evaluate SyzTrust on Trusted OSes from three major vendors: Samsung, Tsinglink Cloud, and Ali Cloud. These systems run on Cortex M23/33 MCUs, which provide the necessary abstraction for embedded TEEs. We discovered 70 previously unknown vulnerabilities in their Trusted OSes, receiving 10 new CVEs so far. Furthermore, compared to the baseline, SyzTrust has demonstrated significant improvements, including 66% higher code coverage, 651% higher state coverage, and 31% improved vulnerability-finding capability. We report all discovered new vulnerabilities to vendors and open source SyzTrust. Qinying Wang, Boyu Chang, Shouling Ji, Yuan Tian 0001, Xuhong Zhang 0002, Chenyang Lyu, Mathias Payer, Wenhai Wang, Raheem A. Beyah |
SP | 2 |
| 2024 | TTFL: Towards Trustworthy Federated Learning with Arm Confidential ComputingabstractFederated learning (FL), as a distributed training paradigm, has drawn great attention from both academia and industry. Recently, privacy and security concerns have been raised for FL. Despite many efforts to protect privacy and security, an FL framework that can systematically provide privacy and security guarantees is lacking. In this work, we present TTFL, a trustworthy FL framework in practice to defend the security and privacy issues based on Arm Confidential Compute Architecture (CCA). TTFL has two core designs. (1) It achieves a high-availability privacy protection based on flexible Trusted Execution Environments (TEEs). It leverages the resource-rich and conveniently accessed features of the latest TEE on Arm CCA, combined with our TEE secure interconnection design, to enable the whole FL process performed in distributed TEEs, which efficiently protects parameter confidentiality and protocol integrity. (2) It achieves effective security protection by proposing an effective poisoning-resisted secure aggregation scheme and protecting it within TEE. The new proposed secure aggregation combines the advantages of existing defenses and is placed in the flexible TEE to ensure a secure, effective, and non-bypassable aggregation procedure. We implement a prototype of TTFL and evaluate it regarding security, privacy, and system performance. Evaluation results show that TTFL can comprehensively and efficiently address the main privacy and security threats in FL. For instance, compared with previous work, it improves the model accuracy by 1.9% and reduces the attack success rate by 79.7% on the CIFAR-10 dataset with only about 19.8% training time overhead. Lizhi Sun, Jingzhou Zhu, Boyu Chang, Yixin Xu 0003, Hao Wu 0067, Fengyuan Xu, Sheng Zhong 0002 |
TrustCom | 3 |
| 2024 | Multi-Label and Evolvable Dataset Preparation for Web-Based Object DetectionabstractIn this article, we focus on the emerging field of web-based object detection, which has gained considerable attention due to its ability to utilize large amounts of web data for training, thus eliminating the need for labor-intensive manual annotations. However, the noisy and ever-evolving nature of web data poses challenges in preparing high-quality datasets for web-based object detection. To address these challenges, we propose a fully automatic dataset preparation method in this article. Our proposed method incorporates a hierarchical clustering module that assigns multiple precise labels to each image. This module is based on our observation that web image data exhibits different distributions at varying granularities. Furthermore, an evolutionary relabeling module ensures the adaptability of both the prepared dataset and trained detection models to the ever-evolving web data. Extensive experiments demonstrate that our method outperforms other web-based methods, and achieves a comparable performance to those manually labeled benchmark datasets. Shucheng Li, Jingzhou Zhu, Boyu Chang, Hao Wu 0067, Fengyuan Xu, Sheng Zhong 0002 |
ACM Trans. Knowl. Discov. Data | 3 |
| 2023 | Dataset Preparation for Arbitrary Object Detection: An Automatic Approach based on Web Information in EnglishabstractAutomatic dataset preparation can help users avoid labor-intensive and costly manual data annotations. The difficulty in preparing a high-quality dataset for object detection involves three key aspects: relevance, naturality, and balance, which are not addressed by existing works. In this paper, we leverage information from the web, and propose a fully-automatic dataset preparation mechanism without any human annotation, which can automatically prepare a high-quality training dataset for the detection task with English text terms describing target objects. It contains three key designs, i.e., keyword expansion, data de-noising, and data balancing. Our experiments demonstrate that the object detectors trained with auto-prepared data are comparable to those trained with benchmark datasets and outperform other baselines. We also demonstrate the effectiveness of our approach in several more challenging real-world object categories that are not included in the benchmark datasets. Shucheng Li, Boyu Chang, Hao Wu 0067, Sheng Zhong 0002, Fengyuan Xu |
SIGIR | 2 |
| 2022 | Towards Practical and Efficient Long Video SummaryabstractRecently, video summarization (VS) techniques are widely used to alleviate huge processing pressure brought by numerous long videos. However, it is hard to summarize long videos efficiently since processing hundreds of frames is still time-consuming. In this paper, we find that the Kernel Temporal Segmentation (KTS) method designed for detecting the shot boundaries in SOTA VS methods is time-consuming while handling long videos. To address this issue, we propose the Distribution-based KTS (D-KTS) by fully considering the characteristic of shot length distribution. Furthermore, we propose the Hash-based Adaptive Frame Selection (HAFS) to improve the system performance by fully taking advantage of the temporal locality of long videos. Our experiments present that the proposed D-KTS is 92.70% faster and takes up 90.08% less memory than the baseline KTS method on average. Xiaopeng Ke, Boyu Chang, Hao Wu 0067, Fengyuan Xu, Sheng Zhong 0002 |
ICASSP | 2 |