EDBT 2026 Demo / reviewers in the wild / expert
Wanyun Li
dblp:267/0070
· DBLP profile ↗
13ranked-venue papers
5as first author
13since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 10 · 3 first-author · 10 since 2021Artificial intelligence and machine learning · 3 · 2 first-author · 3 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | ClickVOS: Click Video Object SegmentationabstractVideo Object Segmentation (VOS) task aims to segment objects in videos. However, previous settings either require time-consuming manual masks of target objects at the first frame during inference or lack the flexibility to specify arbitrary objects of interest. To address these limitations, we propose the setting named Click Video Object Segmentation (ClickVOS) which segments objects of interest across the whole video according to a single click per object in the first frame. And we provide the extended datasets DAVIS-P and YouTubeVOS-P that with point annotations to support this task. ClickVOS is of significant practical applications and research implications due to its only 1-2 seconds interaction time for indicating an object, comparing annotating the mask of an object needs several minutes. However, ClickVOS also presents increased challenges. To address this task, we propose an end-to-end baseline approach named called Attention Before Segmentation (ABS), motivated by the attention process of humans. ABS utilizes the given point in the first frame to perceive the target object through a concise yet effective segmentation attention. Although the initial object mask is possibly inaccurate, in our ABS, as the video goes on, the initially imprecise object mask can self-heal instead of deteriorating due to error accumulation, which is attributed to our designed improvement memory that continuously records stable global object memory and updates detailed dense memory. In addition, we conduct various baseline explorations utilizing off-the-shelf algorithms from related fields, which could provide insights for the further exploration of ClickVOS. The experimental results demonstrate the superiority of the proposed ABS approach. Extended datasets and codes will be available at https://github.com/PinxueGuo/ClickVOS. Pinxue Guo, Lingyi Hong, Xinyu Zhou 0006, Shuyong Gao, Wanyun Li, Zhaoyu Chen 0001, Xiaoqiang Li 0002, Wei Zhang 0016 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2025 | Uncertainty Reactivation: Dynamic Contrastive Correction for Semi-Supervised Medical Image SegmentationabstractSemi-supervised medical image segmentation has advanced significantly by utilizing pseudo-labeled annotations. However, ensuring pseudo-label accuracy remains challenging, often causing misclassification and confirmation bias. Existing methods mainly use prediction uncertainty to exclude or downweight uncertain regions, but these areas frequently coincide with diagnostically important zones, such as lesion cores or tissue boundaries. Neglecting them can thus degrade segmentation performance. To address this, we propose a Dynamic Contrastive Correction Network (DCCN) that corrects uncertain regions instead of ignoring them. DCCN aligns features from high-uncertainty areas with dynamically assigned classes via contrastive learning, reconstructing the uncertain feature space. Additionally, a Multi-layer Sampling (MLS) module leverages boundary-aware sampling to focus contrastive learning on uncertain tissue boundaries. Experiments on two public datasets show that DCCN surpasses previous SOTA methods and effectively mitigates the challenges of high-uncertainty regions. Kexin Xie, Baoyao Yang, Wanyun Li, Fei Lyu 0004 |
BIBM | 3 |
| 2025 | A Review of Neural Radiation Field Based 3D Reconstruction Methods for Spatial Targets
Wanyun Li, Yuqiang Fang, Gege Sun |
ICIG (3) | 1 |
| 2025 | Simple but Effective: Sub-Volume Contrastive Learning for Class-Imbalanced Semi-Supervised 3D Medical Image SegmentationabstractMedical image segmentation is essential for precise anatomical delineation and clinical decision-making. However, fully supervised methods are limited by the substantial cost of acquiring pixel-level annotations, particularly for 3D volumetric data. Semi-supervised learning (SSL) alleviates this challenge by leveraging unlabeled data, yet it remains hindered by severe class imbalance, where dominant structures disproportionately occupy the voxel space, leading to feature degradation and unreliable pseudo-labels. To address this issue, we propose a simple but effective SSL framework, namely Sub-Volume Contrastive Learning (SuVCL), to enhance feature discriminability in imbalanced 3D medical image segmentation. Our approach incorporates localized contrastive learning through sub-volume sampling, which captures small but semantically informative regions to retain fine-grained structural details while mitigating computational overhead. Furthermore, we introduce a balanced memory bank mechanism, which dynamically maintains class-specific feature representations with adaptive updates guided by class-predictive confidence. Extensive experimental evaluations demonstrate that our method substantially enhances segmentation performance for minority classes, demonstrating substantial performance gains over existing SOTAs. Xianrun Xu, Baoyao Yang, Wanyun Li, Jingsong Lin, Yufei Xu |
ACM Multimedia | 3 |
| 2025 | Shallow Features Matter: Hierarchical Memory with Heterogeneous Interaction for Unsupervised Video Object SegmentationabstractUnsupervised Video Object Segmentation (UVOS) aims to predict pixel-level masks for the most salient objects in videos without any prior annotations. While memory mechanisms have been proven critical in various video segmentation paradigms, their application in UVOS yield only marginal performance gains despite sophisticated design. Our analysis reveals a simple but fundamental flaw in existing methods: over-reliance on memorizing high-level semantic features. UVOS inherently suffers from the deficiency of lacking fine-grained information due to the absence of pixel-level prior knowledge. Consequently, memory design relying solely on high-level features, which predominantly capture abstract semantic cues, is insufficient to generate precise predictions. To resolve this fundamental issue, we propose a novel hierarchical memory architecture to incorporate both shallow- and high-level features for memory, which leverages the complementary benefits of pixel and semantic information. Furthermore, to balance the simultaneous utilization of the pixel and semantic memory features, we propose a heterogeneous interaction mechanism to perform pixel-semantic mutual interactions, which explicitly considers their inherent feature discrepancies. Through the design of Pixel-guided Local Alignment Module (PLAM) and Semantic-guided Global Integration Module (SGIM), we achieve delicate integration of the fine-grained details in shallow-level memory and the semantic representations in high-level memory. Our Hierarchical Memory with Heterogeneous Interaction Network (HMHI-Net) consistently achieves state-of-the-art performance across all UVOS and video saliency detection benchmarks. Moreover, HMHI-Net consistently exhibits high performance across different backbones, further demonstrating its superiority and robustness. Project page: https://github.com/ZhengxyFlow/HMHI-Net . Songcheng He, Wanyun Li, Xiaoqiang Li 0002, Wei Zhang 0016 |
ACM Multimedia | 3 |
| 2025 | Predicting stock price movement using social network analytics: Posts are sometimes less usefulabstractContemporary research has leveraged social network data as a predictive tool for decision-making process in the capital market. Yet, its effectiveness may be compromised by social contagion. This study addresses this problem by introducing conversation-level measures that capture how interactions among investors affect market predictions. Drawing on social contagion theory, we identified three conversation conditions—argument similarity, sentiment similarity, and conversation size—and examined their association with the likelihood of abrupt stock price changes, which indicate a loss of collective wisdom. Our analysis of 18 million StockTwits posts for 859 Initial Public Offerings (2008–2017) reveals that conversations with highly similar arguments, highly similar sentiments, and larger size are significantly associated with an increased likelihood of abrupt stock price changes in the subsequent week. Moreover, out-of-sample tests confirm that monitoring conversational dynamics enhances the predictive power of social network analytics, offering valuable guidance for investors and practitioners. Our study extends the theoretical framework of social contagion by highlighting the importance of the conversation level and provides practical recommendations for refining trading strategies based on social media data. • Applying social contagion theory to study stock returns prediction by social media posts. • Examining cognitive and emotional contagion and extensity of social contagion. • Cautioning against over-reliance on social media data for stock returns prediction. • Emphasizing the significance of conversation-level metrics in social media analysis. Wanyun Li, Alvin Chung Man Leung, Ka Wai (Stanley) Choi, Shuk Ying Ho |
Decis. Support Syst. | 1 |
| 2024 | OneTracker: Unifying Visual Object Tracking with Foundation Models and Efficient TuningabstractVisual object tracking aims to localize the target object of each frame based on its initial appearance in the first frame. Depending on the input modility, tracking tasks can be divided into RGB tracking and RGB+X (e.g. RGB+N, and RGB+D) tracking. Despite the different input modalities, the core aspect of tracking is the temporal matching. Based on this common ground, we present a general framework to unify various tracking tasks, termed as One Tracker. One- Tracker first performs a large-scale pre-training on a RGB tracker called Foundation Tracker. This pretraining phase equips the Foundation Tracker with a stable ability to estimate the location of the target object. Then we regard other modality information as prompt and build Prompt Tracker upon Foundation Tracker. Through freezing the Foundation Tracker and only adjusting some additional trainable parameters, Prompt Tracker inhibits the strong localization ability from Foundation Tracker and achieves parameter- efficient finetuning on downstream RGB+X tracking tasks. To evaluate the effectiveness of our general framework OneTracker, which is consisted of Foundation Tracker and Prompt Tracker, we conduct extensive experiments on 6 popular tracking tasks across 11 benchmarks and our One- Tracker outperforms other models and achieves state-of-the-art performance. Lingyi Hong, Shilin Yan, Renrui Zhang, Wanyun Li, Xinyu Zhou 0006, Pinxue Guo, Kaixun Jiang, Zhaoyu Chen 0001 |
CVPR | 4 |
| 2024 | OneVOS: Unifying Video Object Segmentation with All-in-One Transformer Framework
Wanyun Li, Pinxue Guo, Xinyu Zhou 0006, Lingyi Hong, Yangji He, Wei Zhang 0016 |
ECCV (58) | 1 |
| 2024 | X-Prompt: Multi-modal Visual Prompt for Video Object Segmentation
Pinxue Guo, Wanyun Li, Lingyi Hong, Xinyu Zhou 0006, Zhaoyu Chen 0001, Kaixun Jiang, Wei Zhang 0016 |
ACM Multimedia | 2 |
| 2024 | HFVOS: History-Future Integrated Dynamic Memory for Video Object SegmentationabstractMemory-based methods have substantially enhanced the precision of video object segmentation (VOS) by storing features in an expanding memory bank. However, this comes at the cost of increased computational demands and storage overhead. While recent methods have sought to alleviate this issue via compression or selection strategies, their reliance solely on history cues and simple memory structures result in precision degradation and intrinsic limitations, such as error accumulation and poor robustness. In this paper, we introduce HFVOS, an efficient yet effective framework to bolster VOS performance in both speed and precision by meticulously considering the memory design with low redundancy, high accuracy, and adaptability. First, we construct a novel hierarchical memory update pipeline with the proposed Buffered Memory Mechanism, which incorporates both future and history cues to reduce redundancy and improve the utility of memory. Second, we propose an Adaptive Dual-stream Selection Network (ADSN) to carry out the adaptive selection and drop operations of the memory update, and integrate an ADSN based long-term memory to enhance the robustness, especially for long videos. Furthermore, to further boost HFVOS, a progressive selection loss is designed to facilitate ADSN gradually adapt to fewer features while preserving high precision. Experiments show that HFVOS achieves the state-of-the-art segmentation precision and speed on both short-term datasets (DAVIS-17 val: 86.8%J&Fand 33.0 FPS, DAVIS-16 val: 92.0%J&Fand 42.0 FPS) and long-term datasets (LVOS val: 58.0%J&Fand 37.4 FPS). Code will be available at https://github.com/L599wy/HFVOS. Wanyun Li, Jack Fan, Pinxue Guo, Lingyi Hong, Wei Zhang 0016 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2023 | Wavelet Transform-Assisted Adaptive Generative Modeling for ColorizationabstractUnsupervised deep learning has recently demonstrated the promise of producing high-quality samples. While it has tremendous potential to promote the image colorization task, the performance is limited owing to the high-dimension of data manifold and model capability. This study presents a novel scheme that exploits the score-based generative model in wavelet domain to address the issues. By taking advantage of the multi-scale and multi-channel representation via wavelet transform, the proposed model learns the richer priors from stacked coarse and detailed wavelet coefficient components jointly and effectively. This strategy also reduces the dimension of the original manifold and alleviates the curse of dimensionality, which is beneficial for estimation and sampling. Moreover, dual consistency terms in the wavelet domain, namely data-consistency and structure-consistency are devised to leverage colorization task better. Specifically, in the training phase, a set of multi-channel tensors consisting of wavelet coefficients is used as the input to train the network with denoising score matching. In the inference phase, samples are iteratively generated via annealed Langevin dynamics with data and structure consistencies. Experiments demonstrated remarkable improvements of the proposed method on both generation and colorization quality, particularly in colorization robustness and diversity. Jin Li 0031, Wanyun Li, Zichen Xu 0001, Yuhao Wang 0001, Qiegen Liu |
IEEE Trans. Multim. | 2 |
| 2023 | Joint intensity-gradient guided generative modeling for colorization
Kuan Xiong, Kai Hong, Jin Li 0031, Wanyun Li, Weidong Liao, Qiegen Liu |
Vis. Comput. | 4 |
| 2021 | The strategic role of CIOs in IT controls: IT control weaknesses and CIO turnover
Wanyun Li, Soon-Yeow Phang, Ka Wai (Stanley) Choi, Shuk Ying Ho |
Inf. Manag. | 1 |