VLDB 2026 Research / reviewers in the wild / expert
Miaomiao Yuan
dblp:312/1684
· DBLP profile ↗
5ranked-venue papers
0as first author
5since 2021 · last 2026
0009-0008-5978-8046ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Security and privacy · 4 · 4 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MCTS-SQL: Light-Weight LLMs Can Master the Text-to-SQL Through Monte Carlo Tree SearchabstractText-to-SQL is a fundamental yet challenging task in the NLP area, aiming at translating natural language questions into SQL queries. While recent advances in large language models have greatly improved performance, most existing approaches depend on models with tens of billions of parameters or costly APIs, limiting their applicability in resource-constrained environments. For real world, especially on edge devices, it is crucial for Text-to-SQL to ensure cost-effectiveness. Therefore, enabling the light-weight models for Text-to-SQL is of great practical significance. However, smaller LLMs often struggle with complicated user instruction, redundant schema linking or syntax correctness. To address these challenges, we propose MCTS-SQL, a novel framework that uses Monte Carlo Tree Search to guide SQL generation through multi-step refinement. Since the light-weight models' weak performance of single-shot prediction, we generate better results through several trials with feedback. However, directly applying MCTS-based methods inevitably leads to significant time and computational overhead. Driven by this issue, we propose a token-level prefix-cache mechanism that stores prior information during iterations, effectively improved the execution speed. Experiments results on the SPIDER and BIRD benchmarks demonstrate the effectiveness of our approach. Using a small open-source Qwen2.5-Coder-1.5B, our method outperforms ChatGPT-3.5. When leveraging a more powerful model Gemini 2.5 to explore the performance upper bound, we achieved results competitive with the SOTA. Our findings demonstrate that even small models can be effectively deployed in practical Text-to-SQL systems with the right strategy. Shuozhi Yuan, Miaomiao Yuan |
AAAI | 3 |
| 2026 | FCovFuzz: Enhancing Processor Fuzzing via Functional-Behavioral Coverage Guidance
Ruomin Fang, Yanqi Yang, Miaomiao Yuan, Dan Meng 0002 |
IEEE Trans. Inf. Forensics Secur. | 6 |
| 2026 | ModFuzz: Adaptive Module-Level Fuzzing of ProcessorsabstractHardware fuzzing has become a compelling automated verification method for efficiently identifying hardware bugs. However, current fuzzers predominantly focus on maximizing overall coverage, often overlooking the coverage of individual modules. This oversight leads to insufficient testing of low-coverage yet functionally critical modules and leaves essential inter-module dependencies unexplored. Consequently, effectively and efficiently verifying processor modules remains an unresolved challenge. In order to achieve focused exploration of low-coverage modules and effectively capture inter-module dependencies, we propose ModFuzz, a novel adaptive module-level processor fuzzer. We divide the processor into modules and dynamically adjust their priorities based on the Nondominated Sorting Genetic Algorithm II (NSGA-II). By selecting the highest-priority module and applying Inter-Module Dependency Matrix (IMDM)-driven seed selection, ModFuzz concentrates fuzz testing on low-coverage modules and high-dependency seeds. We evaluated ModFuzz on five popular open source RISC-V processors and discovered 16 new bugs with varying degrees of complexity, each of which received a CVE assignment. Compared to the representative CPU fuzzers DifuzzRTL and ProcessorFuzz, ModFuzz improves module coverage by an average of 4.35× and 4.44×, respectively, and increases overall coverage by an average of 4.16× and 3.97×. Our experimental results demonstrate that ModFuzz effectively detects processor bugs while significantly enhancing both the module and the overall coverage. Ruomin Fang, Miaomiao Yuan, Dan Meng 0002 |
IEEE Trans. Inf. Forensics Secur. | 5 |
| 2025 | Into the Unknown: Fuzzing CPU Non-standard Instructions with MystFuzzabstractModern CPU designs have become increasingly complex, making their comprehensive verification a significant challenge. Non-standard instructions, such as illegal, reserved, and hint instructions, are often overlooked during the verification process, potentially leading to critical bugs remaining undetected. Despite recent advancements in fuzzing techniques offering hope for CPU verification, the verification of nonstandard instructions remains a significant challenge. To fill the gap in the current field of CPU verification regarding non-standard instructions, we present MystFuzz, a fuzzing method specifically tailored for non-standard instructions. MystFuzz introduces an efficient instruction space constraint mechanism, supported by a lightweight instruction simulator, to generate large-scale non-standard instructions. This design enables CPU fuzzing without relying on an external golden reference model, and the constrained instruction space dynamically adjusts throughout the fuzzing process. Combined with an efficient exception handling and recovery mechanism, it supports large-scale fuzzing of non-standard instructions in CPUs. Experimental results show significant improvements in fuzzing non-standard instructions for CPUs. Compared to widely-used tools such as riscv-torture and riscv-dv, MystFuzz achieves 215.4x and 61.1x performance improvements in fuzzing, respectively. Even with the same number of non-standard instructions generated, MystFuzz achieves a more diverse range of instruction scenarios. We evaluate five RISC-V CPUs, including XiangShan, CVA6, Rocket, NutShell, Kronos, and discover 19 new bugs (with 10 CVEs assigned) caused by non-standard instructions, highlighting the security impact of non-standard instructions. Zihui Guo, Wenhao Cui, Miaomiao Yuan, Dan Meng 0002 |
ACSAC | 4 |
| 2025 | DiveFuzz: Enhancing CPU Fuzzing via Diverse Instruction ConstructionabstractComprehensive exploration of the CPU architectural states in fuzzing is akin to generating diverse test cases, which include a reasonable distribution of opcode and diversity in instruction execution results (typically measured through write-back data). However, our analysis of state-of-the-art CPU fuzzers reveals that they exhibit high repetition in write-back data and an imbalanced distribution of opcodes during fuzzing. This paper presents DiveFuzz, which diversifies write-back data by finely controlling the operands of instructions at runtime, coupled with correlated contextual semantics, to generate instruction streams with diverse write-back data and semantic associations. Furthermore, DiveFuzz introduces a novel mutator that monitors the fuzzing process to dynamically adjust opcode distribution and accurately eliminate false positives. Our evaluations show that DiveFuzz significantly increases the diversity of instruction write-back data and achieves a more balanced opcode distribution compared to state-of-the-art fuzzers. Across five common coverage metrics, DiveFuzz achieves coverage 204× faster than DifuzzRTL and 114× faster than Cascade. We evaluated DiveFuzz on four well-known open-source RISC-V CPUs—XiangShan, CVA6, Rocket, and NutShell—uncovering 26 new bugs, 15 of which have CVE identifiers. Zihui Guo, Miaomiao Yuan, Yanqi Yang, Dan Meng 0002 |
CCS | 2 |