Shaomeng Wang

dblp:45/3632 · DBLP profile ↗
← Back
11ranked-venue papers
2as first author
11since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 5 · 2 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 5 since 2021Artificial intelligence and machine learning · 3 · 1 first-author · 3 since 2021
YearPublicationVenuePosition
2026 PC-Flow: Preference Alignment in Flow Matching via Classifier
abstract
Flow Matching (FM) is an efficient generative modeling framework, but aligning it with human preferences remains underexplored.~Although applying Direct Preference Optimization (DPO) to diffusion models has yielded improvements, directly extending DPO-like methods to FM poses three challenges: 1) Incompatibility with ODE-based models, 2) Heavy computational cost from full model fine-tuning, and 3) Reliance on reference model quality. To address these limitations, we propose Preference Classifier for Flow Matching (PC-Flow), a novel reference-free preference alignment framework. Specifically, we reinterpret FM’s deterministic ODE as an equivalent SDE to enable DPO-style learning. Then, we introduce a lightweight classifier to model relative preferences exclusively. This approach decouples alignment from the generative model, eliminating the need for costly fine-tuning or a reference model. Theoretically, PC-Flow guarantees consistent preference-guided distribution evolution, achieves a DPO-equivalent objective without a reference model, and progressively steers generation toward preferred outputs. Experiments show that PC-Flow achieves DPO-level alignment with significantly lower training costs.
Shaomeng Wang, He Wang 0054, Longquan Dai, Jinhui Tang 0001
AAAI1
2025 AccCtr: Accelerating Training-Free Conditional Control For Diffusion Models
abstract
In current training-free Conditional Diffusion Models (CDM), the sampling process is steered by the gradient, which measures the discrepancy between the guidance and the condition extracted by a pre-trained condition extraction network. These methods necessitate small guidance steps, resulting in longer sampling times. To address the issue of slow sampling, we introduce AccCtr, a method that simplifies the conditional sampling algorithm by maximizing the sum of two objectives. The local maximum set of one objective is contained within the local maximum set of the other. Leveraging this relationship, we decompose the joint optimization into two parts, alternately maximizing each objective. By analyzing the steps involved in optimizing these objectives, we identify the most time-consuming steps and recommend retraining condition extraction network—a relatively simple task—to reduce its computational cost. Integrating AccCtr into current CDMs is a seamless task that does not impose a significant computational burden. Extensive testing has demonstrated that AccCtr offers superior sample quality and faster generation times.
Longquan Dai, He Wang 0054, Shaomeng Wang, Jinhui Tang 0001
IJCAI4
2025 Conducting Conditional Diffusion by Estimating the Mean Vector of von Mises-Fisher Distribution
abstract
Recent diffusion model advancements aim to handle conditional generative tasks without extra training. Existing training-free methods add a correction term at each denoising step, but they often face computational instability and lack controllability, especially with limited samples and large noise. We propose a new approach using the von Mises-Fisher (vMF) distribution to model the denoised result, turning the conditional generation task into an estimation problem for vMF parameters. We formulate the conditional diffusion model as a mean vector estimation problem for the Gaussian distribution, noting that this can be seen as an estimation problem from noisy observations. When the sampling number is small, the estimation is unstable. To address this, we optimize the mean vector of the vMF distribution by minimizing the KL divergence between the prior and posterior distributions. This approach not only addresses the computational instability but also improves the controllability and quality of the generated results. Once these parameters are determined, the denoised result can be sampled directly from the vMF distribution. Estimating the parameters requires minimal additional code and incurs negligible computational overhead while significantly improving performance. Extensive experiments across various conditional generation tasks, including depth maps, edge detection, segmentation, and style guidance, demonstrate the superiority and versatility of our method. Our approach consistently outperforms existing training-free methods and even surpasses some training-required methods in terms of visual quality and controllability.
Longquan Dai, He Wang 0054, Xiaolu Wei, Shaomeng Wang, Jinhui Tang 0001
ACM Multimedia4
2025 Generative Semantic Probing for Vision-Language Models via Hierarchical Feature Optimization
abstract
Vision-language models (VLMs) has demonstrated impressive cross-modal alignment. However, their internal mechanisms of associating text concepts with visual patterns remain opaque. This opacity raises a critical question: What visual patterns do VLMs inherently associate with text concepts? Current methods for decoding representations of VLMs often produce suboptimal outputs, hindering to probe the clear visual patterns. To address this, we introduce Generative Semantic Probing (GSP), a novel training-free framework that synthesizes images to probe the implicit semantic preferences of VLMs. Our method generates visual patterns that maximize the similarity to the target text embeddings, through three core components: (1) Hierarchical Feature Decomposition, which decomposes the image generation across multi-scale feature levels; (2) Feature Space Constraint, which constrains the optimization within semantically meaningful feature subspace; (3) Quality Assessment Module, which ensures the generation of visually plausible outputs. Experiments validate our method's strengths in high-fidelity image generation and interpretable model analysis. Beyond text-to-image generation, style transfer and image editing applications, our framework enables unprecedented visualization of VLMs' decision boundaries. By exposing implicit preferences and systematic biases in the cross-modal association, our work provides a valuable insight for both understanding and improvement of the vision-language alignment.
He Wang 0054, Longquan Dai, Shihao Pu, Shaomeng Wang, Jinhui Tang 0001
ACM Multimedia4
2025 Aligning Text-to-Image Diffusion Models to Human Preference by Classification
abstract
Text-to-image diffusion models are typically trained on large-scale web data, often resulting in outputs that misalign with human preferences. Inspired by preference learning in large language models, we propose ABC (Alignment by Classification), a simple yet effective framework for aligning diffusion models with human preferences. In contrast to prior DPO-based methods that depend on suboptimal supervised fine-tuned (SFT) reference models, ABC assumes access to an ideal reference model perfectly aligned with human intent and reformulates alignment as a classification problem. Under this view, we recognize that preference data naturally forms a semi-supervised classification setting. To address this, we propose a data augmentation strategy that transforms preference comparisons into fully supervised training signals. We then introduce a classification-based ABC loss to guide alignment. Our alignment by classification approach could effectively steer the diffusion model toward the behavior of the ideal reference. Experiments on various diffusion models show that our ABC consistently outperforms existing baselines, offering a scalable and robust solution for preference-based text-to-image fine-tuning.
Longquan Dai, Xiaolu Wei, He Wang 0054, Shaomeng Wang, Jinhui Tang 0001
NeurIPS4
2024 Research on Ridge-Enhanced Flatted Grating Folded Waveguide SWS for G-Band TWT
abstract
Our previous research introduced a novel Flat-ted Grating Folded Waveguide (FGFW) slow wave structure (SWS) for G-band traveling wave tubes (TWT). Compared to conventional folded waveguide (FW), FGFW demonstrates better EM characteristics. To further enhance performance, this paper proposes the ridge-enhanced flat-ted grating folded waveguide (RE-FGFW), which significantly increases interaction impedance and saturated output power. The RE-FGFW demonstrates superior interaction impedance and amplification performance in TWT applications compared to both the FGFW and FW.
Jingrui Duan, Zhigang Lu 0006, CaiDong Xiong, Zhanliang Wang, Shaomeng Wang, Huarong Gong, Yubin Gong
TENCON9
2024 A Pencil Beam Electron-Optical System for Teahertz Vacuum Electronic Devices
Shaomeng Wang, Qingying Yi, Yubin Gong
TENCON3
2024 W-Band Backward Wave Oscillator Utilizing a Diverging Radial Sheet Electron Beam
Atif Jameel, Zhanliang Wang, Jibran Latif, M. Khawar Nadeem, Khalil Ud Din, Bilawal Ali, Shaomeng Wang, Yubin Gong
TENCON7
2024 Emission Gating of Sheet Electron Beam for High-Power Gridless Inductive Output Tube
abstract
The gridless inductive output tube (IOT) is a versatile variant of the conventional IOT. It primarily offers high-power operation in the MW-range, and a compact, flexible design. The electron beam is density-modulated at the cathode to achieve bunching with a non-intercepting anode. This paper presents an analysis on the emission gating of a sheet electron beam designed for a high-power gridless IOT. Here, the focusing electrode is utilized for this modulation by applying a potential difference, relative to the cathode. The field intensity required for emission is calculated using a modified Fowler-Norheim equation, which accounts for the transition into the space-charge limited regime, and compensates for the geometric effects of the emission surface. The beam voltage is 80 kV, beam current is 36.5A, and perveance is$1.6 \times 10^{-6}\mathrm{A}/\mathrm{V}^{3/2}$. The beam is confined with a magnetic field of 0.7 T. The compressed beam dimensions are$2.54\ \text{mm} \times 37.6\ \text{mm}$. The simulated current density at the cathode is 20.27 A/ cm2with a peak electric field of the order of$10^{6}\mathrm{V}/\mathrm{m}$. The simulation results show agreement with the Fowler-Nordheim model for the electric field intensity required to operate this beam.
Muhammad Khawar Nadeem, Shaomeng Wang, Atif Jameel, Jibran Latif, Bilawal Ali, Longfei Dang, Yubin Gong
TENCON2
2023 20736-node weighted max-cut problem solving by quadrature photonic spatial Ising machine
Xin Ye 0008, Shaomeng Wang, Xiaoxuan Yang 0002, Zuyuan He
Sci. China Inf. Sci.3
2023 Com-STAL: Compositional Spatio-Temporal Action Localization
abstract
Spatio-temporal action localization aims to locate the spatial and temporal positions of actors and classify their actions. However, prior research overlooks the fact that human actions often interact with novel objects in real-world scenarios, which neglects the various combinations of action-object, and considerably limits the generalization of the developed models. In this paper, we study the action-object combinations by researching multi-modal vision information of them. To this end, we propose a novel compositional spatio-temporal action localization (Com-STAL) task, which features non-overlapping action-object combinations in their training and test sets. Based on this, we construct a compositional action localization dataset (Com-AD). Beyond that, we propose a simple yet effective framework, Instance-Centric Interaction Network (ICIN), to reduce invalid induction biases within the visual modality and alleviate the combined distribution bias issue by leveraging additional modal information. The extensive experiment results on Com-AD demonstrate superior action localization performance of ICIN.
Shaomeng Wang, Rui Yan 0010, Guangzhao Dai, Yan Song 0005, Xiangbo Shu
IEEE Trans. Circuits Syst. Video Technol.1