VLDB 2026 Research / reviewers in the wild / expert
Siying Wu
dblp:167/0390
· DBLP profile ↗
13ranked-venue papers
4as first author
12since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 7 · 3 first-author · 6 since 2021Artificial intelligence and machine learning · 5 · 1 first-author · 4 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Enhancing zero-shot brain tumor subtype classification via fine-grained patch-text alignment
Lubin Gan, Jing Zhang 0165, Linhao Qu, Siying Wu, Xiaoyan Sun 0001 |
Expert Syst. Appl. | 5 |
| 2025 | Enhancing Visual Question Answering Via Clustered In-Context Sequence ConfigurationabstractRecent advances in Multimodal In-Context Learning (M-ICL) for Multimodal Large Language Models (MLLMs) have attracted considerable attention. These developments primarily focus on configuring an in-context sequence for a given test case based on instance-level semantic similarity. However, high similarity among demonstrations in the sequence introduces inductive biases, which may mislead MLLMs and ultimately degrade their overall performance. To address this, we propose a novel cluster-based in-context configuration method that adaptively groups candidate data and selects demonstrations from each cluster. This method enhances the diversity within the sequence while preserving semantic consistency, enabling MLLMs to focus on the main intent of the demonstrations. The experimental results on four Visual Question Answering (VQA) benchmarks, including OK-VQA, VQAv2, VizWiz, and TextVQA, demonstrate the effectiveness of our proposed method. Yijun Pan, Hebei Li, Feipeng Ma, Yansong Peng, Siying Wu, Xiaoyan Sun 0001 |
ICIP | 6 |
| 2025 | Create Anything Anywhere: Layout-Controllable Personalized Diffusion Model for Multiple SubjectsabstractDiffusion models have significantly advanced text-to-image generation, laying the foundation for the development of personalized generative frameworks. However, existing methods lack precise layout controllability and overlook the potential of dynamic features of reference subjects in improving fidelity. In this work, we propose Layout-Controllable Personalized Diffusion (LCP-Diffusion) model, a novel framework that integrates subject identity preservation with flexible layout guidance in a tuning-free approach. Our model employs a Dynamic-Static Complementary Visual Refining module to comprehensively capture the intricate details of reference subjects, and introduces a Dual Layout Control mechanism to enforce robust spatial control across both training and inference stages. Extensive experiments validate that LCP-Diffusion excels in both identity preservation and layout controllability. To the best of our knowledge, this is a pioneering work enabling users to "create anything anywhere". Hebei Li, Yansong Peng, Siying Wu, Yueyi Zhang 0001, Xiaoyan Sun 0001 |
ICME | 4 |
| 2025 | MeDKCoOp: Dual Knowledge-guided Graph Prompt Learning for Biomedical Vision-Language ModelsabstractThe rapid evolution of vision-language models (VLMs), such as CLIP, has demonstrated remarkable zero-shot capabilities in downstream tasks. Prompt learning paradigms like Context Optimization (CoOp) refine learnable prompts for efficient adaptation. However, their application in the biomedical domain remains limited due to the insufficient utilization of specialized biomedical knowledge and cross-modality structural relationships. To address these, we introduce MeDKCoOp, a Medical Dual Knowledge-guided graph adaptation method that leverages systematic integration of knowledge through three aspects: exploit domain-specific knowledge from both textual and visual branches, formalize it into graph-structured representations, and leverage knowledge-guided relation transfer for learning cross-modality fusion. By dynamically optimizing learnable prompts through relation learning process, our method achieves disentangled visual representation and enhances transferability to downstream tasks. Evaluations across 8 biomedical datasets spanning 7 imaging modalities demonstrate state-of-the-art cross-domain generalization, with an average 15.12% accuracy improvement over baselines. Our work establishes a new paradigm through graph prompt learning in medical vision-language models, advancing robust diagnostic AI in data-scarce clinical scenarios. Our code is available at: https://github.com/WangYijun-OUC/MeDKCoOp. Siying Wu, Lubin Gan, Zheyu Zhang 0002, Jing Zhang 0165, Zhangchi Hu, Huyue Zhu, Peixi Wu, Xiaoyan Sun 0001 |
ACM Multimedia | 2 |
| 2025 | Damage Analysis via Bidirectional Multi-Task Cascaded Multimodal FusionabstractDamage analysis in social media platforms such as Twitter is a comprehensive problem which involves different subtasks for mining damage-related information from tweets ( e.g., informativeness, humanitarian categories and severity assessment). The comprehensive information obtained by damage analysis enables to identify breaking events around the world in real-time and hence provides aids in emergency responses. Recently, with the rapid development of web technologies, multimodal damage analysis has received increasing attentions due to users' preference of posting multimodal information in social media. Multimodal damage analysis leverages the associated image modality to improve the identification of damage-related information in social media. However, existing works on multimodal damage analysis address each damage-related subtask individually and do not consider their joint training mechanism. In this work, we propose the Bidirectional Multi-task Cascaded multimodal Fusion (BiMCF) approach towards joint multimodal damage analysis. To this end, we introduce the cascaded multimodal fusion framework to separately integrate effective visual and text information for each task, considering that different tasks attend to different information. To exploit the interactions across tasks, bidirectional propagation of the attended image-text interactive information is implemented between tasks, which can lead to enhanced multimodal fusion. Comprehensive experiments are conducted to validate the effectiveness of the proposed approach. Code is available at https://github.com/tiggers23/BiMCF. Siying Wu, Junfeng Fang, Guowu Yang, Wenya Wang 0001, Fengmao Lv |
WWW | 2 |
| 2025 | MMSupcon: An image fusion-based multi-modal supervised contrastive method for brain tumor diagnosis
Jing Zhang 0165, Siying Wu, Xun Chen 0001, Yunwei Ou, Xiaoyan Sun 0001 |
Artif. Intell. Medicine | 3 |
| 2025 | Hierarchical Task-aware Temporal Modeling and Matching for few-shot action recognition
Yucheng Zhan, Yijun Pan, Siying Wu, Yueyi Zhang 0001, Xiaoyan Sun 0001 |
Neurocomputing | 3 |
| 2025 | Dynamic Challenge Cross-Selection Physical Unclonable Function Based on MRAMabstractThe rapid development of Internet of Things (IoT) devices has triggered massive data transmission. Meanwhile, advances in artificial intelligence (AI) introduce new security vulnerabilities in device interactions. These challenges demand lightweight yet robust security solutions. In this context, physical unclonable functions (PUFs) serve as critical hardware security primitives, enabling reliable authentication for edge devices. Nevertheless, PUF is increasingly susceptible to novel threats, notably machine learning attacks. To address this security vulnerability to attacks, we propose a novel double-layer dynamic challenge cross-selection magnetoresistive random access memory PUF (MPUF). This design leverages the inherent process variation in spin-transfer torque magnetoresistive random access memory (STT-MRAM) as an entropy source. The proposed structure incorporates an obfuscation decode circuit (ODC) that combinesxorgates and shift registers. It dynamically obfuscates interlayer relationships between two PUF arrays to enhance circuit nonlinearity. The simulation results demonstrate uniformity of 50.16%, uniqueness of 49.94%, a worst bit error rate (BER) of 2.34% for$- 25~^{\circ } $C to$125~^{\circ } $C and 1.56% for$0.5\sim 1.1$V. In addition, four common machine learning models are used to attack this PUF, achieving accuracies of 50.49%, 50.49%, 50.48%, and 58.41%, which are close to a random guess. Compared with traditional PUF implementations, this work exhibits higher reliability and enhanced security while maintaining low power consumption of approximately 9.975 fJ/bit. Siying Wu, Yu Gong 0002, Jiaao Dai, Shouzhong Peng, Yue Zhang 0010, You Wang 0002, Weiqiang Liu 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2024 | Semantic-Enhanced Point-Box Joint Prompting for Video Object SegmentationabstractThe Segment Anything Model (SAM) has demonstrated outstanding zero-shot performance in image segmentation through efficient point and box prompts. In this paper, we propose a SAM-based Semantic-enhanced Point-Box joint prompting (SAM-SPB) framework for Video Object Segmentation (VOS). SAM-SPB leverages the local structure information and the global semantic cues of interest objects, leading to strong and robust segmentation. To be specific, the local structure information of the objects is maintained by a point tracking branch, and the semantic consistency of the objects across frames are propagated through our proposed semantic-aware memory-based box tracking branch. Compared with previous SAM-based point-centric video segmentation method, we highlight the importance of point-box joint prompting for video object segmentation. The state-of-the-art experimental results on popular VOS benchmarks in the zero-shot setting demonstrate the strong zero-shot ability of the proposed method. Siying Wu, Yueyi Zhang 0001, Xiaoyan Sun 0001 |
ICIP | 2 |
| 2024 | Visual Perception by Large Language Model's WeightsabstractExisting Multimodal Large Language Models (MLLMs) follow the paradigm that perceives visual information by aligning visual features with the input space of Large Language Models (LLMs) and concatenating visual tokens with text tokens to form a unified sequence input for LLMs. These methods demonstrate promising results on various vision-language tasks but are limited by the high computational effort due to the extended input sequence resulting from the involvement of visual tokens. In this paper, instead of input space alignment, we propose a novel parameter space alignment paradigm that represents visual information as model weights. For each input image, we use a vision encoder to extract visual features, convert features into perceptual weights, and merge the perceptual weights with LLM's weights. In this way, the input of LLM does not require visual tokens, which reduces the length of the input sequence and greatly improves efficiency. Following this paradigm, we propose VLoRA with the perceptual weights generator. The perceptual weights generator is designed to convert visual features to perceptual weights with low-rank property, exhibiting a form similar to LoRA. The experimental results show that our VLoRA achieves comparable performance on various benchmarks for MLLMs, while significantly reducing the computational costs for both training and inference. Code and models are released at \url{https://github.com/FeipengMa6/VLoRA}. Feipeng Ma, Hongwei Xue, Yizhou Zhou, Guangting Wang, Fengyun Rao, Shilin Yan, Yueyi Zhang 0001, Siying Wu, Zheng Shou 0001, Xiaoyan Sun 0001 |
NeurIPS | 8 |
| 2024 | Vision-and-Language Navigation via Latent Semantic Alignment LearningabstractVision-and-Language Navigation (VLN) requires that an agent can comprehensively understand the given instructions and the immediate visual information obtained from the environment, so as to make correct actions to achieve the navigation goal. Therefore, semantic alignment across modalities is crucial for the agent understanding its own state during the navigation process. However, the potential of semantic alignment has not been systematically explored in current studies, which limits the further improvement of navigation performance. To address this issue, we propose a new Latent Semantic Alignment Learning method to develop the semantically aligned relationships contained in the environment. Specifically, we introduce three novel pre-training tasks: Trajectory-conditioned Masked Fragment Modeling, Action Prediction of Masked Observation, and Hierarchical Triple Contrastive Learning. The first two tasks are used to reason about cross-modal dependencies, while the third one is able to learn semantically consistent representations across modalities. In this way, the Latent Semantic Alignment Learning method establishes a consistent perception of the environment and makes the agent's actions easier to explain. Experiments on common benchmarks verify the effectiveness of our proposed methods. For example, we improve the Success Rate by 1.6% on the R2R validation unseen set and 4.3% on the R4R validation unseen set over the baseline model. Siying Wu, Xueyang Fu, Feng Wu 0005, Zhengjun Zha |
IEEE Trans. Multim. | 1 |
| 2022 | Cross-modal Semantic Alignment Pre-training for Vision-and-Language NavigationabstractVision-and-Language Navigation needs an agent to navigate to a target location by progressively grounding and following the relevant instruction conditioning on its memory and current observation. Existing works utilize the cross-modal transformer to pass the message between visual modality and textual modality. However, they are still limited to mining the fine-grained matching between the underlying components of trajectories and instructions. Inspired by the significant progress achieved by large-scale pre-training methods, in this paper, we propose CSAP, a new method of Cross-modal Semantic Alignment Pre-training for Vision-and-Language Navigation. It is designed to learn the alignment from trajectory-instruction pairs through two novel tasks, including trajectory-conditioned masked fragment modeling and contrastive semantic-alignment modeling. Specifically, the trajectory-conditioned masked fragment modeling encourages the agent to extract useful visual information to reconstruct the masked fragment. The contrastive semantic-alignment modeling is designed to align the visual representation with corresponding phrase embeddings. By showing experimental results on the benchmark dataset, we demonstrate that transformer architecture-based navigation agent pre-trained with our proposed CSAP outperforms existing methods on both SR and SPL scores. Siying Wu, Xueyang Fu, Feng Wu 0001, Zhengjun Zha |
ACM Multimedia | 1 |
| 2019 | Densely Supervised Hierarchical Policy-Value Network for Image Paragraph GenerationabstractImage paragraph generation aims to describe an image with a paragraph in natural language. Compared to image captioning with a single sentence, paragraph generation provides more expressive and fine-grained description for storytelling. Existing approaches mainly optimize paragraph generator towards minimizing word-wise cross entropy loss, which neglects linguistic hierarchy of paragraph and results in ``sparse" supervision for generator learning. In this paper, we propose a novel Densely Supervised Hierarchical Policy-Value (DHPV) network for effective paragraph generation. We design new hierarchical supervisions consisting of hierarchical rewards and values at both sentence and word levels. The joint exploration of hierarchical rewards and values provides dense supervision cues for learning effective paragraph generator. We propose a new hierarchical policy-value architecture which exploits compositionality at token-to-token and sentence-to-sentence levels simultaneously and can preserve the semantic and syntactic constituent integrity. Extensive experiments on the Stanford image-paragraph benchmark have demonstrated the effectiveness of the proposed DHPV approach with performance improvements over multiple state-of-the-art methods. Siying Wu, Zhengjun Zha, Zilei Wang, Houqiang Li, Feng Wu 0001 |
IJCAI | 1 |