EDBT 2026 Demo / reviewers in the wild / expert
Meiqi Wu
dblp:335/6876
· DBLP profile ↗
15ranked-venue papers
5as first author
14since 2021 · last 2026
0009-0007-3155-4013ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 10 · 1 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 4 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Omni-Effects: Unified and Spatially-Controllable Visual Effects GenerationabstractVisual effects (VFX) are essential visual enhancements fundamental to modern cinematic production. Although video generation models offer cost-efficient solutions for VFX production, current methods are constrained by per-effect LoRA training, which limits generation to single effects. This fundamental limitation impedes applications that require spatially controllable composite effects, i.e., the concurrent generation of multiple effects at designated locations. However, integrating diverse effects into a unified framework faces major challenges: interference from effect variations and spatial uncontrollability during multi-VFX joint training. To tackle these challenges, we propose Omni-Effects, a first unified framework capable of generating prompt-guided effects and spatially controllable composite effects. The core of our framework comprises two key innovations: (1) LoRA-based Mixture of Experts (LoRA-MoE), which employs a group of expert LoRAs, integrating diverse effects within a unified model while effectively mitigating cross-task interference. (2) Spatial-Aware Prompt (SAP) incorporates spatial mask information into the text token, enabling precise spatial control. Furthermore, we introduce an Independent-Information Flow (IIF) module integrated within the SAP, isolating the control signals corresponding to individual effects to prevent any unwanted blending. To facilitate this research, we construct a comprehensive VFX dataset Omni-VFX via a novel data collection pipeline combining image editing and First-Last Frame-to-Video (FLF2V) synthesis, and introduce a dedicated VFX evaluation framework for validating model performance. Extensive experiments demonstrate that Omni-Effects achieves precise spatial control and diverse effect generation, enabling users to specify both the category and location of desired effects. Fangyuan Mao, Aiming Hao, Dongxia Liu, Xiaokun Feng, Jiashu Zhu, Meiqi Wu, Chubin Chen, Jiahong Wu 0005, Xiangxiang Chu |
AAAI | 7 |
| 2026 | ImagerySearch: Adaptive Test-Time Search for Video Generation Beyond Semantic Dependency ConstraintsabstractVideo generation models have achieved remarkable progress, particularly excelling in realistic scenarios; however, their performance degrades notably in imaginative scenarios. These prompts often involve rarely co-occurring concepts with long-distance semantic relationships, falling outside training distributions. Existing methods typically apply test-time scaling for improving video quality, but their fixed search spaces and static reward designs limit adaptability to imaginative scenarios. To fill this gap, we propose ImagerySearch, a dynamic test-time scaling law strategy inspired by imagery that adaptively adjusts the inference search space and reward guided by prompts, effectively enhancing generation quality in imaginative scenarios. Furthermore, we introduce LDT-Bench, the first benchmark targeting long-distance semantic prompts, designed to evaluate the creativity of video generation models. It comprises 2,839 challenging concept pairs from diverse recognition datasets and incorporates an automatic evaluation protocol to assess creative capacity. Extensive experiments on LDT-Bench demonstrate that our approach consistently outperforms general generation models and test-time scaling approaches. Additionally, ImagerySearch achieves strong performance on VBench, confirming its effectiveness in improving video generation quality under diverse conditions. Meiqi Wu, Jiashu Zhu, Xiaokun Feng, Chubin Chen, Bingze Song, Fangyuan Mao, Jiahong Wu 0005, Xiangxiang Chu, Kaiqi Huang |
AAAI | 1 |
| 2026 | Calibration-Aware Policy Optimization for Reasoning LLMsabstractGroup Relative Policy Optimization (GRPO) enhances LLM reasoning but often induces overconfidence, where incorrect responses yield lower perplexity than correct ones, degrading relative calibration as described by the Area Under the Curve (AUC).Existing approaches either yield limited improvements in calibration or sacrifice gains in reasoning accuracy.We first prove that this degradation in GRPO-style algorithms stems from their uncertainty-agnostic advantage estimation, which inevitably misaligns optimization gradients with calibration.This leads to improved accuracy at the expense of degraded calibration.We then propose Calibration-Aware Policy Optimization (CAPO).It adopts a logistic AUC surrogate loss that is theoretically consistent and admits regret bound, enabling uncertainty-aware advantage estimation.By further incorporating a noise masking mechanism, CAPO achieves stable learning dynamics that jointly optimize calibration and accuracy.Experiments on multiple mathematical reasoning benchmarks show that CAPO-1.5Bsignificantly improves calibration by up to 15% while achieving accuracy comparable to or better than GRPO, and further boosts accuracy on downstream inference-time scaling tasks by up to 5%.Moreover, when allowed to abstain under low-confidence conditions, CAPO achieves a Pareto-optimal precision-coverage trade-off, highlighting its practical value for hallucination mitigation. Xingzhou Lou, Meiqi Wu, Zhengqi Wen, Junge Zhang |
ACL (1) | 3 |
| 2026 | DiffChar: A fast conditional diffusion model for air-writing Chinese character generation
Weixi Zhao, Meiqi Wu |
Pattern Recognit. | 2 |
| 2026 | Talk With Your Fingers: A Depth-Aware Benchmark for Air-Writing RecognitionabstractAir-writing has emerged as a promising communication modality for AR/VR and metaverse environments, enabling quiet, non-contact text input by translating finger movements into natural language. However, existing approaches typically project in-air writing onto a virtual 2D plane and assume characters are formed with a single continuous stroke–an oversimplification that neglects the rich 3D structure inherent in natural handwriting. In this work, we challenge the “single-stroke 2D” paradigm and explore the role of depth cues in enhancing air-writing recognition. To this end, we present DAAWBench, the first large-scale Depth-Aware Air-Writing dataset, featuring 8.8 million RGB-D frames annotated with 3,755 Chinese characters from the GB2312-80 Level-1 set. Our analysis reveals consistent depth variations at stroke boundaries, indicating that stroke segmentation and character recognition can benefit from depth modeling. Based on these insights, we propose DARec, a novel 3D trajectory-based recognition model that effectively leverages depth-aware priors. Extensive experiments across in-domain and out-of-domain settings, including evaluations with vision-language models (e.g., GPT-4o, Qwen-VL) and human baselines, show that DARec significantly outperforms 2D-only counterparts, achieving 87.73% accuracy versus 9.05%. Our findings demonstrate the critical importance of depth modeling in human-computer co-creative interfaces, and we will publicly release our dataset and code at https://github.com/wmeiqi/DAAWBench. Meiqi Wu, Yuzhong Zhao, Xuchen Li 0001, Yuanqiang Cai, Jiahong Wu 0005, Weiqiang Wang 0001, Kaiqi Huang |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2025 | Enhancing Vision-Language Tracking by Effectively Converting Textual Cues into Visual CuesabstractVision-Language Tracking (VLT) aims to localize a target in video sequences using a visual template and language description. While textual cues enhance tracking potential, current datasets typically contain much more image data than text, limiting the ability of VLT methods to align the two modalities effectively. To address this imbalance, we propose a novel plug-and-play method named CTVLT that leverages the strong text-image alignment capabilities of foundation grounding models. CTVLT converts textual cues into interpretable visual heatmaps, which are easier for trackers to process. Specifically, we design a textual cue mapping module that transforms textual cues into target distribution heatmaps, visually representing the location described by the text. Additionally, the heatmap guidance module fuses these heatmaps with the search image to guide tracking more effectively. Extensive experiments on mainstream benchmarks demonstrate the effectiveness of our approach, achieving state-of-the-art performance and validating the utility of our method for enhanced VLT. Xiaokun Feng, Dailing Zhang, Xuchen Li 0001, Meiqi Wu, Jing Zhang 0110, Xiaotang Chen, Kaiqi Huang |
ICASSP | 5 |
| 2025 | ATCTrack: Aligning Target-Context Cues with Dynamic Target States for Robust Vision-Language TrackingabstractVision-language tracking aims to locate the target object in the video sequence using a template patch and a language description provided in the initial frame. To achieve robust tracking, especially in complex long-term scenarios that reflect real-world conditions as recently highlighted by MGIT, it is essential not only to characterize the target features but also to utilize the context features related to the target. However, the visual and textual target-context cues derived from the initial prompts generally align only with the initial target state. Due to their dynamic nature, target states are constantly changing, particularly in complex long-term sequences. It is intractable for these cues to continuously guide Vision-Language Trackers (VLTs). Furthermore, for the text prompts with diverse expressions, our experiments reveal that existing VLTs struggle to discern which words pertain to the target or the context, complicating the utilization of textual cues. In this work, we present a novel tracker named ATCTrack, which can obtain multimodal cues Aligned with the dynamic target states through comprehensive Target-Context feature modeling, thereby achieving robust tracking. Specifically, (1) for the visual modality, we propose an effective temporal visual target-context modeling approach that provides the tracker with timely visual cues. (2) For the textual modality, we achieve precise target words identification solely based on textual content, and design an innovative context words calibration method to adaptively utilize auxiliary context words. (3) We conduct extensive experiments on mainstream benchmarks and ATCTrack achieves a new SOTA performance. The code and models will be released at: https://github.com/XiaokunFeng/ATCTrack. Xiaokun Feng, Xuchen Li 0001, Dailing Zhang, Meiqi Wu, Jing Zhang 0110, Xiaotang Chen, Kaiqi Huang |
ICCV | 5 |
| 2025 | VMBench: A Benchmark for Perception-Aligned Video Motion Generation
Xinran Ling, Meiqi Wu, Xiaokun Feng, Cundian Yang, Aiming Hao, Jiashu Zhu, Jiahong Wu 0005, Xiangxiang Chu |
ICCV | 3 |
| 2025 | CSTrack: Enhancing RGB-X Tracking via Compact Spatiotemporal FeaturesabstractEffectively modeling and utilizing spatiotemporal features from RGB and other modalities (e.g., depth, thermal, and event data, denoted as X) is the core of RGB-X tracker design. Existing methods often employ two parallel branches to separately process the RGB and X input streams, requiring the model to simultaneously handle two dispersed feature spaces, which complicates both the model structure and computation process. More critically, intra-modality spatial modeling within each dispersed space incurs substantial computational overhead, limiting resources for inter-modality spatial modeling and temporal modeling. To address this, we propose a novel tracker, CSTrack, which focuses on modeling Compact Spatiotemporal features to achieve simple yet effective tracking. Specifically, we first introduce an innovative Spatial Compact Module that integrates the RGB-X dual input streams into a compact spatial feature, enabling thorough intra- and inter-modality spatial modeling. Additionally, we design an efficient Temporal Compact Module that compactly represents temporal features by constructing the refined target distribution heatmap. Extensive experiments validate the effectiveness of our compact spatiotemporal modeling method, with CSTrack achieving new SOTA results on mainstream RGB-X benchmarks. The code and models will be released at: https://github.com/XiaokunFeng/CSTrack. Xiaokun Feng, Dailing Zhang, Xuchen Li 0001, Meiqi Wu, Jing Zhang 0110, Xiaotang Chen, Kaiqi Huang |
ICML | 5 |
| 2024 | MemVLT: Vision-Language Tracking with Adaptive Memory-based PromptsabstractVision-language tracking (VLT) enhances traditional visual object tracking by integrating language descriptions, requiring the tracker to flexibly understand complex and diverse text in addition to visual information. However, most existing vision-language trackers still overly rely on initial fixed multimodal prompts, which struggle to provide effective guidance for dynamically changing targets. Fortunately, the Complementary Learning Systems (CLS) theory suggests that the human memory system can dynamically store and utilize multimodal perceptual information, thereby adapting to new scenarios. Inspired by this, (i) we propose a Memory-based Vision-Language Tracker (MemVLT). By incorporating memory modeling to adjust static prompts, our approach can provide adaptive prompts for tracking guidance.
(ii) Specifically, the memory storage and memory interaction modules are designed in accordance with CLS theory. These modules facilitate the storage and flexible interaction between short-term and long-term memories, generating prompts that adapt to target variations.
(iii) Finally, we conduct extensive experiments on mainstream VLT datasets (e.g., MGIT, TNL2K, LaSOT and LaSOT$_{ext}$). Experimental results show that MemVLT achieves new state-of-the-art performance. Impressively, it achieves 69.4% AUC on the MGIT and 63.3% AUC on the TNL2K, improving the existing best result by 8.4% and 4.7%, respectively. Xiaokun Feng, Xuchen Li 0001, Dailing Zhang, Meiqi Wu, Jing Zhang 0110, Xiaotang Chen, Kaiqi Huang |
NeurIPS | 5 |
| 2024 | Beyond Accuracy: Tracking more like Human via Visual SearchabstractHuman visual search ability enables efficient and accurate tracking of an arbitrary moving target, which is a significant research interest in cognitive neuroscience. The recently proposed Central-Peripheral Dichotomy (CPD) theory sheds light on how humans effectively process visual information and track moving targets in complex environments. However, existing visual object tracking algorithms still fall short of matching human performance in maintaining tracking over time, particularly in complex scenarios requiring robust visual search skills. These scenarios often involve Spatio-Temporal Discontinuities (i.e., STDChallenge), prevalent in long-term tracking and global instance tracking. To address this issue, we conduct research from a human-like modeling perspective: (1) Inspired by the CPD, we pro- pose a new tracker named CPDTrack to achieve human-like visual search ability. The central vision of CPDTrack leverages the spatio-temporal continuity of videos to introduce priors and enhance localization precision, while the peripheral vision improves global awareness and detects object movements. (2) To further evaluate and analyze STDChallenge, we create the STDChallenge Benchmark. Besides, by incorporating human subjects, we establish a human baseline, creating a high- quality environment specifically designed to assess trackers’ visual search abilities in videos across STDChallenge. (3) Our extensive experiments demonstrate that the proposed CPDTrack not only achieves state-of-the-art (SOTA) performance in this challenge but also narrows the behavioral differences with humans. Additionally, CPDTrack exhibits strong generalizability across various challenging benchmarks. In summary, our research underscores the importance of human-like modeling and offers strategic insights for advancing intelligent visual target tracking. Code and models are available at https://github.com/ZhangDailing8/CPDTrack. Dailing Zhang, Xiaokun Feng, Xuchen Li 0001, Meiqi Wu, Jing Zhang 0110, Kaiqi Huang |
NeurIPS | 5 |
| 2024 | VS-LLM: Visual-Semantic Depression Assessment Based on LLM for Drawing Projection Test
Meiqi Wu, Yaxuan Kang, Xuchen Li 0001, Xiaotang Chen, Yunfeng Kang, Weiqiang Wang 0001, Kaiqi Huang |
PRCV (9) | 1 |
| 2024 | Finger in Camera Speaks Everything: Unconstrained Air-Writing for Real-WorldabstractAir-writing is a challenging task that combines the fields of computer vision and natural language processing, offering an intuitive and natural approach for human-computer interaction. However, current air-writing solutions face two primary challenges: (1) their dependency on complex sensors (e.g., Radar, EEGs and others) for capturing precise handwritten trajectories, and (2) the absence of a video-based air-writing dataset that covers a comprehensive vocabulary range. These limitations impede their practicality in various real-world scenarios, including the use on devices like iPhones and laptops. To tackle these challenges, we present the groundbreaking air-writing Chinese character video dataset (AWCV-100K), serving as a pioneering benchmark for video-based air-writing. This dataset captures handwritten trajectories in various real-world scenarios using commonly accessible RGB cameras, eliminating the need for complex sensors. AWCV-100K includes 8.8 million video frames, encompassing the complete set of 3,755 characters from the GB2312-80 level-1 set (GB1). Furthermore, we introduce our baseline approach, the video-based character recognizer (VCRec). VCRec adeptly extracts fingertip features from sparse visual cues and employs a spatio-temporal sequence module for analysis. Experimental results showcase the superior performance of VCRec compared to existing models in recognizing air-written characters, both quantitatively and qualitatively. This breakthrough paves the way for enhanced human-computer interaction in real-world contexts. Moreover, our approach leverages affordable RGB cameras, enabling its applicability in a diverse range of scenarios. The code and data examples will be made public at https://github.com/wmeiqi/AWCV. Meiqi Wu, Kaiqi Huang, Yuanqiang Cai, Yuzhong Zhao, Weiqiang Wang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2023 | A Multi-modal Global Instance Tracking Benchmark (MGIT): Better Locating Target in Complex Spatio-temporal and Causal RelationshipabstractTracking an arbitrary moving target in a video sequence is the foundation for high-level tasks like video understanding. Although existing visual-based trackers have demonstrated good tracking capabilities in short video sequences, they always perform poorly in complex environments, as represented by the recently proposed global instance tracking task, which consists of longer videos with more complicated narrative content. Recently, several works have introduced natural language into object tracking, desiring to address the limitations of relying only on a single visual modality. However, these selected videos are still short sequences with uncomplicated spatio-temporal and causal relationships, and the provided semantic descriptions are too simple to characterize video content.To address these issues, we (1) first propose a new multi-modal global instance tracking benchmark named MGIT. It consists of 150 long video sequences with a total of 2.03 million frames, aiming to fully represent the complex spatio-temporal and causal relationships coupled in longer narrative content. (2) Each video sequence is annotated with three semantic grains (i.e., action, activity, and story) to model the progressive process of human cognition. We expect this multi-granular annotation strategy can provide a favorable environment for multi-modal object tracking research and long video understanding. (3) Besides, we execute comparative experiments on existing multi-modal object tracking benchmarks, which not only explore the impact of different annotation methods, but also validate that our annotation method is a feasible solution for coupling human understanding into semantic labels. (4) Additionally, we conduct detailed experimental analyses on MGIT, and hope the explored performance bottlenecks of existing algorithms can support further research in multi-modal object tracking. The proposed benchmark, experimental results, and toolkit will be released gradually on http://videocube.aitestunion.com/. Dailing Zhang, Meiqi Wu, Xiaokun Feng, Xuchen Li 0001, Xin Zhao 0012, Kaiqi Huang |
NeurIPS | 3 |
| 2019 | A deep learning method to more accurately recall known lysine acetylation sitesabstractBACKGROUND: Lysine acetylation in protein is one of the most important post-translational modifications (PTMs). It plays an important role in essential biological processes and is related to various diseases. To obtain a comprehensive understanding of regulatory mechanism of lysine acetylation, the key is to identify lysine acetylation sites. Previously, several shallow machine learning algorithms had been applied to predict lysine modification sites in proteins. However, shallow machine learning has some disadvantages. For instance, it is not as effective as deep learning for processing big data. RESULTS: In this work, a novel predictor named DeepAcet was developed to predict acetylation sites. Six encoding schemes were adopted, including a one-hot, BLOSUM62 matrix, a composition of K-space amino acid pairs, information gain, physicochemical properties, and a position specific scoring matrix to represent the modified residues. A multilayer perceptron (MLP) was utilized to construct a model to predict lysine acetylation sites in proteins with many different features. We also integrated all features and implemented the feature selection method to select a feature set that contained 2199 features. As a result, the best prediction achieved 84.95% accuracy, 83.45% specificity, 86.44% sensitivity, 0.8540 AUC, and 0.6993 MCC in a 10-fold cross-validation. For an independent test set, the prediction achieved 84.87% accuracy, 83.46% specificity, 86.28% sensitivity, 0.8407 AUC, and 0.6977 MCC. CONCLUSION: The predictive performance of our DeepAcet is better than that of other existing methods. DeepAcet can be freely downloaded from https://github.com/Sunmile/DeepAcet . Meiqi Wu, Yingxi Yang |
BMC Bioinform. | 1 |