VLDB 2026 Research / reviewers in the wild / expert
Siwei Wu
dblp:240/8368
· DBLP profile ↗
11ranked-venue papers
3as first author
11since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Security and privacy · 4 · 1 first-author · 4 since 2021Artificial intelligence and machine learning · 3 · 3 since 2021Software engineering, systems software and programming languages · 2 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | The Dark Side of Upgrades: Uncovering Insecurity in Smart Contract Upgrades
Dingding Wang 0003, Jianting He, Siwei Wu, Yajin Zhou, Lei Wu 0012, Cong Wang 0001 |
ACISP (1) | 3 |
| 2025 | OmniBench: Towards The Future of Universal Omni-Language ModelsabstractRecent advancements in multimodal large language models (MLLMs) have focused on integrating multiple modalities, yet their ability to simultaneously process and reason across different inputs remains underexplored. We introduce OmniBench, a novel benchmark designed to evaluate models’ ability to recognize, interpret, and reason across visual, acoustic, and textual inputs simultaneously. We define language models capable of such tri-modal processing as omni-language models (OLMs). OmniBench features high-quality human annotations that require integrated understanding across all modalities. Our evaluation reveals that: i) open-source OLMs show significant limitations in instruction-following and reasoning in tri-modal contexts; and ii) most baseline models perform poorly (below 50% accuracy) even with textual alternatives to image/audio inputs. To address these limitations, we develop OmniInstruct, an 96K-sample instruction tuning dataset for training OLMs. We advocate for developing more robust tri-modal integration techniques and training strategies to enhance OLM performance. Codes and data could be found at https://m-a-p.ai/OmniBench/. Ge Zhang 0009, Yinghao Ma, Ruibin Yuan, Kang Zhu, Hangyu Guo, Yiming Liang, Noah Wang, Jian Yang 0003, Siwei Wu, Xingwei Qu, Jinjie Shi, Xinyue Zhang 0005, Zhenzhu Yang, Yidan Wen, Yanghai Wang, Zhaoxiang Zhang 0001, Ruibo Liu, Emmanouil Benetos, Wenhao Huang 0001, Chenghua Lin 0002 |
NeurIPS | 11 |
| 2025 | Identifying cancer prognosis genes through causal learningabstractAccurate identification of causal genes for cancer prognosis is critical for estimating disease progression and guiding treatment interventions. In this study, we propose CPCG (Cancer Prognosis's Causal Gene), a two-stage framework identifying gene sets causally associated with patient prognosis across diverse cancer types using transcriptomic data. Initially, an ensemble approach models gene expression's impact on survival with parametric and semiparametric hazard models. Subsequently, an iterative conditional independence test combined with graph pruning is utilized to infer the causal skeleton, thereby pinpointing prognosis-related genes. Experiments on transcriptomic data from 18 cancer types sourced from The Cancer Genome Atlas Project demonstrate CPCG's effectiveness in predicting prognosis under four evaluation metrics. Validations on 24 additional datasets covering 12 cancer types from the Gene Expression Omnibus and the Chinese Glioma Genome Atlas Project further demonstrate CPCG's robustness and generalizability. CPCG identifies a concise but reliable set of genes, obviating the need for gene combination enumeration for survival time estimation. These genes are also proved closely linked to crucial biological processes in cancer. Moreover, CPCG constructs a stable causal skeleton and exhibits insensitivity to the order of data shuffling. Overall, CPCG is a powerful tool for extracting cancer prognostic biomarkers, offering interpretability, generalizability, and robustness. CPCG holds promise for facilitating targeted interventions in clinical treatment strategies. Siwei Wu, Chaoyi Yin, Yuezhu Wang, Huiyan Sun |
Briefings Bioinform. | 1 |
| 2025 | Expediting the discovery of promising photothermal cyanine molecules through a transfer learning approachabstractCyanine-based molecules have gained significant attention in photothermal therapy due to their unique fluorescence brightness and tunable spectral properties. However, the development of new photothermal agents is often constrained by the complexity of the chemical landscape and the need for biocompatibility. To address these challenges, we present an innovative transfer learning approach for rapidly identifying promising photothermal agent candidates with excellent photothermal properties, high synthetic feasibility, and superior biocompatibility. Using natural language processing, our pretrained model generated a molecular library based on cyanine scaffolds. The most promising candidates were screened rigorously through a weighted analysis of chemical indicators, such as photothermal performance and synthetic accessibility and biological indicators, including bio-toxicity. From these, three molecules were selected for retrosynthetic analysis. This artificial intelligence-driven approach provides a robust solution to the traditional challenges in photothermal agent design, significantly enhancing their potential applications in cancer bioimaging, mitochondrial phototherapy, and image-guided surgery. Siwei Wu, Liqiang He, Guining Cao, Jiacheng Tang, Zhenxing Pan, Zihui Huang, Andi Li, Shuting Cai, Xujie Liu |
Briefings Bioinform. | 2 |
| 2024 | Overview of the NLPCC 2024 Shared Task on Chinese Metaphor Generation
Xingwei Qu, Ge Zhang 0009, Siwei Wu, Chenghua Lin 0002 |
NLPCC (5) | 3 |
| 2024 | DeFiRanger: Detecting DeFi Price Manipulation AttacksabstractThe rapid growth of Decentralized Finance (DeFi) boosts the blockchain ecosystem. At the same time, attacks on DeFi applications (apps) are increasing. However, to the best of our knowledge, existing smart contract vulnerability detection tools cannot directly detect DeFi attacks. That's because they lack the capability to recover and understand high-level DeFi semantics, e.g., a user trades a token pairXandYin a Decentralized EXchange (DEX). In this work, we focus on the detection of two new types of price manipulation attacks. To this end, we propose a platform-independent method to identify high-level DeFi semantics. Specifically, we first construct the Cash Flow Tree (CFT) from a raw transaction and then lifting the low-level semantics to high-level ones, including five advanced DeFi actions. Finally, we use patterns expressed with the recovered DeFi semantics to detect price manipulation attacks. We implemented a prototype namedDeFiRangerthat detected 14zero-daysecurity incidents. These findings were reported to affected parties or/and the community for the first time. Furthermore, the backtest experiment discovered 15 unknown historical security incidents. We further performed an attack analysis to shed light on the root causes of vulnerabilities incurring price manipulation attacks. Siwei Wu, Zhou Yu 0002, Dabao Wang, Yajin Zhou, Lei Wu 0012, Haoyu Wang 0001, Xingliang Yuan |
IEEE Trans. Dependable Secur. Comput. | 1 |
| 2023 | Dense-ATOMIC: Towards Densely-connected ATOMIC with High Knowledge Coverage and Massive Multi-hop PathsabstractATOMIC is a large-scale commonsense knowledge graph (CSKG) containing everyday ifthen knowledge triplets, i.e., {head event, relation, tail event}.The one-hop annotation manner made ATOMIC a set of independent bipartite graphs, which ignored the numerous links between events in different bipartite graphs and consequently caused shortages in knowledge coverage and multi-hop paths.In this work, we aim to construct Dense-ATOMIC with high knowledge coverage and massive multi-hop paths.The events in ATOMIC are normalized to a consistent pattern at first.We then propose a CSKG completion method called Rel-CSKGC to predict the relation given the head event and the tail event of a triplet, and train a CSKG completion model based on existing triplets in ATOMIC.We finally utilize the model to complete the missing links in ATOMIC and accordingly construct Dense-ATOMIC.Both automatic and human evaluation on an annotated subgraph of ATOMIC demonstrate the advantage of Rel-CSKGC over strong baselines.We further conduct extensive evaluations on Dense-ATOMIC in terms of statistics, human evaluation, and simple downstream tasks, all proving Dense-ATOMIC's advantages in Knowledge Coverage and Multi-hop Paths.Both the source code of Rel-CSKGC and Dense-ATOMIC are publicly available on https://github.com/ NUSTM/Dense-ATOMIC. Xiangqing Shen, Siwei Wu |
ACL (1) | 2 |
| 2023 | Demystifying Random Number in Ethereum Smart Contract: Taxonomy, Vulnerability Identification, and Attack DetectionabstractRecent years have witnessed explosive growth in blockchain smart contract applications. As smart contracts become increasingly popular and carry trillion dollars worth of digital assets, they become more of an appealing target for attackers, who have exploited vulnerabilities in smart contracts to cause catastrophic economic losses. Notwithstanding a proliferation of work that has been developed to detect an impressive list of vulnerabilities, the bad randomness vulnerability is overlooked by many existing tools. In this article, we make the first attempt to provide a systematic analysis of random numbers in Ethereum smart contracts, by investigating the principles behind pseudo-random number generation and organizing them into a taxonomy. We also lucubrate various attacks against bad random numbers and group them into four categories. Furthermore, we presentRNVulDet– a tool that incorporates taint analysis techniques to automatically identify bad randomness vulnerabilities and detect corresponding attack transactions. To extensively verify the effectiveness ofRNVulDet, we construct three new datasets: i) 34 well-known contracts that are reported to possess bad randomness vulnerabilities, ii) 214 popular contracts that have been rigorously audited before launch and are regarded as free of bad randomness vulnerabilities, and iii) a dataset consisting of 47,668 smart contracts and 49,951 suspicious transactions. We compareRNVulDetwith three state-of-the-art smart contract vulnerability detectors, and our tool significantly outperforms them. Meanwhile,RNVulDetspends 2.98 s per contract on average, in most cases orders-of-magnitude faster than other tools.RNVulDetsuccessfully reveals 44,264 attack transactions. Our implementation and datasets are released, hoping to inspire others. Jianting He, Lingling Lu, Siwei Wu, Zhipeng Lu 0001, Lei Wu 0012, Yajin Zhou, Qinming He |
IEEE Trans. Software Eng. | 4 |
| 2022 | BSB: Bringing Safe Browsing to Blockchain Platform
Rongwei Yu, Siwei Wu, Shengwu Xiong 0001 |
NSS | 5 |
| 2022 | Penny Wise and Pound Foolish: Quantifying the Risk of Unlimited Approval of ERC20 Tokens on EthereumabstractThe prosperity of decentralized finance motivates many investors to profit via trading their crypto assets on decentralized applications (DApps for short) of the Ethereum ecosystem. Apart from Ether (the native cryptocurrency of Ethereum), many ERC20 (a widely used token standard on Ethereum) tokens obtain vast market value in the ecosystem. Specifically, the approval mechanism is used to delegate the privilege of spending users’ tokens to DApps. By doing so, the DApps can transfer these tokens to arbitrary receivers on behalf of the users. To increase the usability, unlimited approval is commonly adopted by DApps to reduce the required interaction between them and their users. However, as shown in existing security incidents, this mechanism can be abused to steal users’ tokens. Dabao Wang, Hang Feng, Siwei Wu, Yajin Zhou, Lei Wu 0012, Xingliang Yuan |
RAID | 3 |
| 2022 | Time-travel Investigation: Toward Building a Scalable Attack Detection Framework on EthereumabstractEthereum has been attracting lots of attacks, hence there is a pressing need to perform timely investigation and detect more attack instances. However, existing systems suffer from the scalability issue due to the following reasons. First, the tight coupling between malicious contract detection and blockchain data importing makes them infeasible to repeatedly detect different attacks. Second, the coarse-grained archive data makes them inefficient to replay transactions. Third, the separation between malicious contract detection and runtime state recovery consumes lots of storage. In this article, we propose a scalable attack detection framework named EthScope , which overcomes the scalability issue by neatly re-organizing the Ethereum state and efficiently locating suspicious transactions. It leverages the fine-grained state to support the replay of arbitrary transactions and proposes a well-designed schema to optimize the storage consumption. The performance evaluation shows that EthScope can solve the scalability issue, i.e., efficiently performing a large-scale analysis on billions of transactions, and a speedup of around \( \text{2,300}\times \) when replaying transactions. It also has lower storage consumption compared with existing systems. Further analysis shows that EthScope can help analysts understand attack behaviors and detect more attack instances. Siwei Wu, Lei Wu 0012, Yajin Zhou, Runhuai Li, Zhi Wang 0004, Xiapu Luo, Cong Wang 0001, Kui Ren 0001 |
ACM Trans. Softw. Eng. Methodol. | 1 |