VLDB 2026 Research / reviewers in the wild / expert
Yanghao Wang
dblp:146/8246
· DBLP profile ↗
11ranked-venue papers
4as first author
6since 2021 · last 2025
0009-0003-6038-8828ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 5 · 1 first-author · 1 since 2021Artificial intelligence and machine learning · 3 · 3 first-author · 3 since 2021Security and privacy · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Inversion Circle Interpolation: Diffusion-based Image Augmentation for Data-scarce ClassificationabstractData Augmentation (DA), i.e., synthesizing faithful and diverse samples to expand the original training set, is an effective strategy to improve the performance of various data-scarce tasks. With the powerful image generation ability, diffusion-based DA has shown strong performance gains on different image classification benchmarks. In this paper, we analyze today’s diffusion-based DA methods, and argue that they cannot take account of both faithfulness and diversity, which are two critical keys for generating high-quality samples and boosting classification performance. To this end, we propose a novel Diffusion-based DA method: Diff-II. Specifically, it consists of three steps: 1) Category concepts learning: Learning concept embeddings for each category. 2) Inversion interpolation: Calculating the inversion for each image, and conducting circle interpolation for two randomly sampled inversions from the same category. 3) Two-stage denoising: Using different prompts to generate synthesized images in a coarse-to-fine manner. Extensive experiments on various data-scarce image classification tasks (e.g., few-shot, long-tailed, and out-of-distribution classification) have demonstrated their effectiveness over state-of-the-art diffusion-based DA methods. Yanghao Wang, Long Chen 0016 |
CVPR | 1 |
| 2025 | Noise Matters: Optimizing Matching Noise for Diffusion ClassifiersabstractAlthough today's pretrained discriminative vision-language models (e.g., CLIP) have demonstrated strong perception abilities, such as zero-shot image classification, they also suffer from the bag-of-words problem and spurious bias. To mitigate these problems, some pioneering studies leverage powerful generative models (e.g., pretrained diffusion models) to realize generalizable image classification, dubbed Diffusion Classifier (DC). Specifically, by randomly sampling a Gaussian noise, DC utilizes the differences of denoising effects with different category conditions to classify categories. Unfortunately, an inherent and notorious weakness of existing DCs is noise instability: different random sampled noises lead to significant performance changes. To achieve stable classification performance, existing DCs always ensemble the results of hundreds of sampled noises, which significantly reduces the classification speed. To this end, we firstly explore the role of noise in DC, and conclude that: there are some ``good noises'' that can relieve the instability. Meanwhile, we argue that these good noises should meet two principles: 1) Frequency Matching: noise should destroy the specific frequency signals; 2) Spatial Matching: noise should destroy the specific spatial areas. Regarding both principles, we propose a novel Noise Optimization method to learn matching (i.e., good) noise for DCs: NoOp. For frequency matching, NoOp first optimizes a dataset-specific noise: Given a dataset and a timestep $t$, optimize one randomly initialized parameterized noise. For Spatial Matching, NoOp trains a Meta-Network that adopts an image as input and outputs image-specific noise offset. The sum of optimized noise and noise offset will be used in DC to replace random noise. Extensive ablations on various datasets demonstrated the effectiveness of NoOp. It is worth noting that our noise optimization is orthogonal to existing optimization methods (e.g., prompt tuning), our NoOP can even benefit from these methods to further boost performance. Yanghao Wang, Long Chen 0016 |
NeurIPS | 1 |
| 2023 | Random Boxes Are Open-world Object DetectorsabstractWe show that classifiers trained with random region proposals achieve state-of-the-art Open-world Object Detection (OWOD): they can not only maintain the accuracy of the known objects (w/training labels), but also considerably improve the recall of unknown ones (w/o training labels). Specifically, we propose RandBox, a Fast R-CNN based architecture trained on random proposals at each training iteration, surpassing existing Faster R-CNN and Transformer based OWOD. Its effectiveness stems from the following two benefits introduced by randomness. First, as the randomization is independent of the distribution of the limited known objects, the random proposals become the instrumental variable that prevents the training from being confounded by the known objects. Second, the unbiased training encourages more proposal explorations by using our proposed matching score that does not penalize the random proposals whose prediction scores do not match the known objects. On two benchmarks: Pascal-VOC/MS-COCO and LVIS, Rand-Box significantly outperforms the previous state-of-the-art in all metrics. We also detail the ablations on randomization and loss designs. Codes are available at https://github.com/scuwyh2000/RandBox. Yanghao Wang, Zhongqi Yue, Xian-Sheng Hua 0001, Hanwang Zhang |
ICCV | 1 |
| 2023 | APTBert: Abstract Generation and Event Extraction from APT Reports
Chenxin Zhou, Cheng Huang 0003, Yanghao Wang, Zheng Zuo |
ICDF2C (2) | 3 |
| 2021 | CyberRel: Joint Entity and Relation Extraction for Cybersecurity Concepts
Yongyan Guo, Cheng Huang 0003, Wangyuan Jing, Ziwang Wang, Yanghao Wang |
ICICS (1) | 7 |
| 2021 | Parallel Discrepancy Detection and Incremental DetectionabstractThis paper studies how to catch duplicates, mismatches and conflicts in the same process. We adopt a class of entity enhancing rules that embed machine learning predicates, unify entity resolution and conflict resolution, and are collectively defined across multiple relations. We detect discrepancies as violations of such rules. We establish the complexity of discrepancy detection and incremental detection problems with the rules; they are both NP-complete and W[1]-hard. To cope with the intractability and scale with large datasets, we develop parallel algorithms and parallel incremental algorithms for discrepancy detection. We show that both algorithms are parallelly scalable, i.e. , they guarantee to reduce runtime when more processors are used. Moreover, the parallel incremental algorithm is relatively bounded. The complexity bounds and algorithms carry over to denial constraints, a special case of the entity enhancing rules. Using real-life and synthetic datasets, we experimentally verify the effectiveness, scalability and efficiency of the algorithms. Wenfei Fan, Chao Tian 0001, Yanghao Wang, Qiang Yin 0002 |
Proc. VLDB Endow. | 3 |
| 2020 | Querying Shared Data with Security HeterogeneityabstractThere has been increasing need for secure data sharing. In practice a group of data owners often adopt a heterogeneous security scheme under which each pair of parties decide their own protocol to share data with diverse levels of trust. The scheme also keeps track of how the data is used. This paper studies distributed SQL query answering in the heterogeneous security setting. We define query plans by incorporating toll functions determined by data sharing agreements and reflected in the use of various security facilities. We formalize query answering as a bi-criteria optimization problem, to minimize both data sharing toll and parallel query evaluation cost. We show that this problem is PSPACE-hard for SQL and Σ_3^p-hard for SPC, and it is in NEXPTIME. Despite the hardness, we develop a set of approximate algorithms to generate distributed query plans that minimize data sharing toll and reduce parallel evaluation cost. Using real-life and synthetic data, we empirically verify the effectiveness, scalability and efficiency of our algorithms. Yang Cao 0012, Wenfei Fan, Yanghao Wang, Ke Yi 0001 |
SIGMOD Conference | 3 |
| 2017 | BEAS: Bounded Evaluation of SQL QueriesabstractWe demonstrate BEAS, a prototype system for querying relations with bounded resources. BEAS advocates an unconventional query evaluation paradigm under an access schema A, which is a combination of cardinality constraints and associated indices. Given an SQL query Q and a dataset D, BEAS computes Q(D) by accessing a bounded fraction DQ of D, such that Q(DQ) = Q(D) and DQ is determined by A and Q only, no matter how big D grows. It identifies DQ by reasoning about the cardinality constraints of A, and fetches DQ using the indices of A. We demonstrate the feasibility of bounded evaluation by walking through each functional component of BEAS. As a proof of concept, we demonstrate how BEAS conducts CDR analyses in telecommunication industry, compared with commercial database systems. Yang Cao 0012, Wenfei Fan, Yanghao Wang, Tengfei Yuan, Laura Yu Chen |
SIGMOD Conference | 3 |
| 2016 | Recommender systems based on ranking performance optimization
Richong Zhang, Han Bao 0005, Hailong Sun 0001, Yanghao Wang, Xudong Liu 0001 |
Frontiers Comput. Sci. | 4 |
| 2014 | AdaMF: Adaptive Boosting Matrix Factorization for Recommender System
Yanghao Wang, Hailong Sun 0001, Richong Zhang |
WAIM | 1 |
| 2014 | Discovering Semantic Mobility Pattern from Check-in Data
Ji Yuan, Xudong Liu 0001, Richong Zhang, Hailong Sun 0001, Xiaohui Guo, Yanghao Wang |
WISE (1) | 6 |