VLDB 2026 Research / reviewers in the wild / expert
Yulu Wang
dblp:143/3574
· DBLP profile ↗
17ranked-venue papers
8as first author
11since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 8 · 2 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 4 first-author · 5 since 2021Databases, data management, data science and information retrieval · 4 · 3 first-author · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Neural Architecture for Fast and Reliable Coagulation Assessment in Clinical Settings: Leveraging ThromboelastographyabstractIn an ideal medical environment, real-time coagulation monitoring can enable early detection and prompt remediation of risks. However, traditional Thromboelastography (TEG), a widely employed diagnostic modality, can only provide such outputs after nearly 1 hour of measurement. The delay might lead to elevated mortality rates. These issues clearly point out one of the key challenges for medical AI development: Making reasonable predictions based on very small data sets and accounting for variation between different patient populations, a task where conventional deep learning methods typically perform poorly. We present Physiological State Reconstruction (PSR), a new algorithm specifically designed to take advantage of dynamic changes between individuals and to maximize useful information produced by small amounts of clinical data through mapping to reliable predictions and diagnosis. We develop MDFE to facilitate integration of varied temporal signals using multi-domain learning, and jointly learn high-level temporal interactions together with attentions via HLA; furthermore, the parameterized DAM we designed maintains the stability of the computed vital signs. PSR evaluates with 4 TEG-specialized data sets and establishes remarkable performance -- predictions of R^2 > 0.98 for coagulation traits and error reduction around half compared to the state-of-the-art methods, and halving the inferencing time too. Drift-aware learning suggests a new future, with potential uses well beyond thrombophilia discovery towards medical AI applications with data scarcity. Yulu Wang, Ziqian Zeng, Zhifeng Tang |
AAAI | 1 |
| 2026 | BLM-Guard: Explainable Multimodal Ad Moderation with Chain-of-Thought and Policy-Aligned RewardsabstractShort-video platforms now host vast multimodal ads whose deceptive visuals, speech and subtitles demand finer-grained, policy-driven moderation than community safety filters. We present BLM-Guard, a content-audit framework for commercial ads that fuses Chain-of-Thought reasoning with rule-based policy principles and a critic-guided reward. A rule-driven ICoT data-synthesis pipeline jump-starts training by generating structured scene descriptions, reasoning chains and labels, cutting annotation costs. Reinforcement learning then refines the model using a composite reward balancing causal coherence with policy adherence. A multitask architecture models intra-modal manipulations (e.g., exaggerated imagery) and cross-modal mismatches (e.g., subtitle–speech drift), boosting robustness. Experiments on real short-video ads show BLM-Guard surpasses strong baselines in accuracy, consistency and generalization. Xiong Ma, Yinghao Song, Xiangji Zeng, Yulu Wang, Shuai Cui, Zhaohan Gong, Jiefei Zhang |
AAAI | 9 |
| 2026 | Multi-scale pyramid-former network with multiple consistency constraints for semi-supervised video action detection
Zhengping Hu, Yulu Wang, Hehao Zhang, Jirui Di |
Inf. Process. Manag. | 3 |
| 2026 | Video and noise collaboratively guided semi-supervised diffusion model for video action detection
Zhengping Hu, Yulu Wang, Hehao Zhang, Jirui Di |
Pattern Recognit. | 3 |
| 2025 | Satisfying Complex Data Security Requirements in Digital Business EcosystemsabstractDigital Business Ecosystems (DBEs) involve collaboration and sharing of data across various independent parties. Data sharing comes with security requirements, e.g. who may see which data elements. Often, these security requirements can be satisfied by well-known techniques, such as access controls, but sometimes the traditional solutions are not sufficient. For example, in our use case there is a requirement to sum up the revenue of companies by the government to calculate the average revenue for an industry, without disclosing the revenue of each company. To satisfy these kinds of requirements without a trusted third party, advanced Privacy-Preserving Computation (PPC) techniques are needed. However, the field of PPC is technically difficult to understand for most people and is highly specialized. We are not aware of a unified, comprehensive framework that can guide the systematic selection and integration of PPC methods, given the security requirements of a DBE use case. Therefore, our research goal is to establish such a framework. In this paper, two motivating examples are given, taken from the music digital business ecosystem we participate in. Yulu Wang, Charlotte van de Velde, Sabine Oechsner, Jaap Gordijn |
RE | 1 |
| 2025 | Human body contour extraction method based on human skeleton key point guidanceabstractBy extracting human contours from 2D images captured by the camera and then obtaining human size data, the cost of garment custom measurement can be effectively reduced and the efficiency of custom measurement can be improved. The extraction of human contours plays an important role in the collection of online human size data. We propose a method to extract human contours by fusing the prior information of human skeleton key points into the salience target detection network. Specifically, the skeleton key point information extracted based on OpenPose is fused into the encoder-decoder network for rough detection of the human body target, and the residual refinement network is used to fine-adjust the human body matting, so as to achieve accurate human contour extraction. In this paper, the accuracy and superiority of the algorithm are verified in the public data set P3M-10K of human body matting and applied to the 2D body measurement WeChat applet on mobile phone and computer website. Zhongwei Hua, Yulu Wang, Zhuriyao Jin |
Connect. Sci. | 3 |
| 2025 | Multi-granularity hierarchical contrastive learning between foreground and background for semi-supervised video action detection
Zhengping Hu, Yulu Wang, Hehao Zhang, Jirui Di |
Knowl. Based Syst. | 3 |
| 2024 | Temporal refinement network: Combining dynamic convolution and multi-scale information for fine-grained action recognition
Jirui Di, Zhengping Hu, Shuai Bi, Hehao Zhang, Yulu Wang |
Image Vis. Comput. | 5 |
| 2023 | Hi-SIGIR: Hierachical Semantic-Guided Image-to-image Retrieval via Scene GraphabstractImage-to-image retrieval, a fundamental task, aims at matching similar images based on a query image. Existing methods with convolutional neural networks are usually sensitive to low-level visual features, and ignore high-level semantic relationship information. This makes retrieving complicated images with multiple objects and various relationships a significant challenge. Although some works introduce the scene graph to capture the global semantic features of the objects and their relations, they ignore the local visual representations. In addition, due to the fragility of individual modal representations, poisoning attacks in adversarial scenarios are easily achieved, hurting the robustness of the visual-guided foundation image retrieval model. To overcome these issues, we propose a novel hierarchical semantic-guided image-to-image retrieval method via scene graph, called Hi-SIGIR. Specifically, to begin with, our proposed method generates the scene graph of an image. Then, our model extracts and learns both the visual and semantic features of the nodes and relations within the scene graphs. Next, these features are fused to obtain local information and sent to the graph neural network to obtain global information. Using these information, the similarity between the scene graphs of several images is calculated at both the local and global levels to perform image retrieval. Finally, we introduce a surrogate that calculates relevance in a cross-modal manner to understand image content better. Experimental evaluations on several wildly-used benchmarks demonstrate the superiority of the proposed method. Yulu Wang, Pengwen Dai, Xiaojun Jia, Zhitao Zeng, Rui Li 0109, Xiaochun Cao |
ACM Multimedia | 1 |
| 2022 | Leveraging Sparse Coding for EEG Based Emotion Recognition in ShootingabstractEmotion recognition in shooting is of great importance for improving athletes’ training methods. However, there is no open and high confident electroencephalography (EEG) dataset about shooting due to the difficulty of data acquisition, which made it a challenge for related studies. In this paper, we collected EEG of novice shooters and high-level shooters in different emotion states, and established two shooting datasets. Furthermore, instead of adopting the common convolutional neural network, we are the first to leverage sparse coding for EEG based emotion recognition in shooting process. Our proposed method can effectively solve the problem of low accuracy caused by data with low signal-noise ratio and small training set. The experimental results demonstrate that our method outperforms other representative deep learning based methods. Yulu Wang, Changshui Zhang |
ICASSP | 1 |
| 2021 | FMA-ETA: Estimating Travel Time Entirely Based on FFN with AttentionabstractEstimated time of arrival (ETA) is one of the most important services in intelligent transportation systems (ITS) and becomes a challenging spatial-temporal (ST) data mining task in recent years. Nowadays, deep learning based methods, specifically recurrent neural networks (RNN) based ones are adapted to model the ST patterns from massive data for ETA and become the state-of-the-art. However, RNN is suffering from slow training and inference speed, as its structure is unfriendly to parallel computing. To solve this problem, we propose a novel, brief and effective framework mainly based on feed-forward network (FFN) for ETA, FFN with Multifactor Attention (FMA-ETA). The novel Multi-factor Attention mechanism is proposed to deal with different category features and aggregate the information purposefully. Extensive experimental results on the real-world vehicle travel dataset show FMA-ETA is competitive with state-of-the-art methods in terms of the prediction accuracy with significantly better inference speed. Yulu Wang, Kun Fu 0002, Zheng Wang 0010, Ziang Yan, Changshui Zhang, Jieping Ye |
ICASSP | 2 |
| 2020 | Adversarial Attacks on Deep Unfolded Networks for Sparse CodingabstractPrevious works have discovered that general DNNs are vulnerable under some subtle and specific perturbations on classification tasks. In recent years, deep neural networks (DNNs) unfolded through sparse coding algorithms have achieved great success in sparse coding problem. Some applications of sparse coding have important strategic significance in many cases, and the security of learned models is vital. However, it has not achieved enough attentions. Our paper is the first work to study the adversarial performance on unfolded DNNs for sparse coding. We first verify the effectiveness of the existing attack or defense strategies, and surprisingly discover the defense strategies are useless. In addition, we propose a special attack strategy to eliminate the element on certain dimension of the sparse output code of DNNs. Furthermore, we propose a succinct black-box attack strategy, which could generate adversarial perturbations without knowing the parameters of DNNs and data. Yulu Wang, Kailun Wu, Changshui Zhang |
ICASSP | 1 |
| 2020 | Constructing Geographic and Long-term Temporal Graph for Traffic ForecastingabstractTraffic forecasting influences various intelligent transportation system (ITS) services and is of great significance for user experience as well as urban traffic control. It is challenging due to the fact that the road network contains complex and time-varying spatial-temporal dependencies. Recently, deep learning based methods have achieved promising results by adopting graph convolutional network (GCN) to extract the spatial correlations and recurrent neural network (RNN) to capture the temporal dependencies. However, the existing methods often construct the graph only based on road network connectivity, which limits the interaction between roads. In this work, we propose Geographic and Long-term Temporal Graph Convolutional Recurrent Neural Network (GLT-GCRNN), a novel framework for traffic forecasting that learns the rich interactions between roads sharing similar geographic or longterm temporal patterns. Extensive experiments on a real-world traffic state dataset validate the effectiveness of our method by showing that GLT-GCRNN outperforms the state-of-the-art methods in terms of different metrics. Yulu Wang, Kun Fu 0002, Zheng Wang 0010, Changshui Zhang, Jieping Ye |
ICPR | 2 |
| 2017 | Partitioning and Segment Organization Strategies for Real-Time Selective Search on Document StreamsabstractThe basic idea behind selective search is to partition a collection into topical clusters, and for each query, consider only a subset of the clusters that are likely to contain relevant documents. Previous work on web collections has shown that it is possible to retain high-quality results while considering only a small fraction of the collection. These studies, however, assume static collections where it is feasible to run batch clustering algorithms for partitioning. In this work, we consider the novel formulation of selective search on document streams (specifically, tweets), where partitioning must be performed incrementally. In our approach, documents are partitioned into temporal segments and selective search is performed within each segment: these segments can either be clustered using batch or online algorithms, and at different temporal granularities. For efficiency, we take advantage of word embeddings to reduce the dimensionality of the document vectors. Experiments with test collections from the TREC Microblog Tracks show that we are able to achieve precision indistinguishable from exhaustive search while considering only around 5% of the collection. Interestingly, we observe no significant effectiveness differences between batch vs. online clustering and between hourly vs. daily temporal segments, despite them being very different index organizations. This suggests that architectural choices should be primarily guided by efficiency considerations. Yulu Wang, Jimmy Lin |
WSDM | 1 |
| 2015 | Assessor Differences and User Preferences in Tweet Timeline GenerationabstractIn information retrieval evaluation, when presented with an effectiveness difference between two systems, there are three relevant questions one might ask. First, are the differences statistically significant? Second, is the comparison stable with respect to assessor differences? Finally, is the difference actually meaningful to a user? This paper tackles the last two questions about assessor differences and user preferences in the context of the newly-introduced tweet timeline generation task in the TREC 2014 Microblog track, where the system's goal is to construct an informative summary of non-redundant tweets that addresses the user's information need. Central to the evaluation methodology is human-generated semantic clusters of tweets that contain substantively similar information. We show that the evaluation is stable with respect to assessor differences in clustering and that user preferences generally correlate with effectiveness metrics even though users are not explicitly aware of the semantic clustering being performed by the systems. Although our analyses are limited to this particular task, we believe that lessons learned could generalize to other evaluations based on establishing semantic equivalence between information units, such as nugget-based evaluations in question answering and temporal summarization. Yulu Wang, Garrick Sherman, Jimmy Lin, Miles Efron |
SIGIR | 1 |
| 2014 | The Impact of Future Term Statistics in Real-Time Tweet Search
Yulu Wang, Jimmy Lin |
ECIR | 1 |
| 2014 | Interactive video segmentation using occlusion boundaries and temporally coherent superpixelsabstractWe propose an interactive video segmentation system built on the basis of occlusion and long term spatio-temporal structure cues. User supervision is incorporated in a superpixel graph clustering framework that differs crucially from prior art in that it modifies the graph according to the output of an occlusion boundary detector. Working with long temporal intervals (up to 100 frames) enables our system to significantly reduce annotation effort with respect to state of the art systems. Even though the segmentation results are less than perfect, they are obtained efficiently and can be used in weakly supervised learning from video or for video content description. We do not rely on a discriminative object appearance model and allow extracting multiple foreground objects together, saving user time if more than one object is present. Additional experiments with unsupervised clustering based on occlusion boundaries demonstrate the importance of this cue for video segmentation and thus validate our system design. Radu Dondera, Vlad I. Morariu, Yulu Wang, Larry Davis 0001 |
WACV | 3 |