EDBT 2026 Demo / reviewers in the wild / expert
Jiaxin Wu 0001
dblp:06/6984-1
· DBLP profile ↗
23ranked-venue papers
9as first author
14since 2021 · last 2026
0000-0003-4074-3442ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 17 · 7 first-author · 11 since 2021Artificial intelligence and machine learning · 4 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 4 · 3 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Adaptive Multi-Agent Reasoning for Text-to-Video RetrievalabstractThe rise of short-form video platforms and the emergence of multimodal large language models (MLLMs) have amplified the need for scalable, effective, zero-shot text-to-video retrieval systems. While recent advances in large-scale pretraining have improved zero-shot cross-modal alignment, existing methods still struggle with query-dependent temporal reasoning, limiting their effectiveness on complex queries involving temporal, logical, or causal relationships. To address these limitations, we propose an adaptive multi-agent retrieval framework that dynamically orchestrates specialized agents over multiple reasoning iterations based on the demands of each query. The framework includes: (1) a retrieval agent for scalable retrieval over large video corpora, (2) a reasoning agent for zero-shot contextual temporal reasoning, and (3) a query reformulation agent for refining ambiguous queries and recovering performance for those that degrade over iterations. These agents are dynamically coordinated by an orchestration agent, which leverages intermediate feedback and reasoning outcomes to guide execution. We also introduce a novel communication mechanism that incorporates retrieval-performance memory and historical reasoning traces to improve coordination and decision-making. Experiments on three TRECVid benchmarks spanning eight years show that our framework achieves a twofold improvement over CLIP4Clip and significantly outperforms state-of-the-art methods by a large margin. The code is available at https://github.com/nikkiwoo-gh/multi-agent-retrieval. Jiaxin Wu 0001, Xiaoyong Wei, Qing Li 0001 |
ICMR | 1 |
| 2026 | RelAgent: a multi-agent solution for molecular relationship groundingabstractMOTIVATION: Molecular captions, patents, and medicinal-chemistry notes describe substructures and their relations in natural language, whereas computational models operate on formal representations such as SMILES. Bridging this semantic-structural gap is important for patent interpretation, structural relationship analysis, and controllable molecular editing, yet current large language models struggle to ground textual references to precise molecular components. RESULTS: We propose RelAgent, a cooperative multi-agent framework for molecular relationship grounding. RelAgent decomposes the task into three interpretable stages: entity extraction, substructure localization, and ontology-guided relationship reasoning, and then uses verifier agents to rank structurally plausible candidates. This design supports fine-grained reasoning over molecular substructure and substantially improves performance on the MolGround benchmark. RelAgent achieves 81.4% entity-extraction F1, 56.0% exact-match localization F1, and 54.6% relationship F1 on an open-source LLaMA3.1-8B model, improving the REL F1 from 0.1% to 54.6% and exceeding the vanilla Gemini-3.1-Pro baseline in our experiments. These results indicate that agentic, structure-aware reasoning is a practical direction for interpretable molecular understanding in bioinformatics. AVAILABILITY AND IMPLEMENTATION: The source code for RelAgent is available at https://github.com/Anya-RB-Chen/RelAgent. Rubing Chen, Jiaxin Wu 0001, Chen Zhang 0013, Xiaoyong Wei |
Bioinform. | 2 |
| 2026 | Self-Paced Learning for Images of Antinuclear AntibodiesabstractAntinuclear antibody (ANA) testing is a critical method for diagnosing autoimmune disorders such as Lupus, Sjögren's syndrome, and scleroderma. Despite its importance, manual ANA detection is slow, labor-intensive, and demands years of training. ANA detection is complicated by over 100 coexisting antibody types, resulting in vast fluorescent pattern combinations. Although machine learning and deep learning have enabled automation, ANA detection in real-world clinical settings presents unique challenges as it involves multi-instance, multi-label (MIML) learning. In this paper, a novel framework for ANA detection is proposed that handles the complexities of MIML tasks using unaltered microscope images without manual preprocessing. Inspired by human labeling logic, it identifies consistent ANA sub-regions and assigns aggregated labels accordingly. These steps are implemented using three task-specific components: an instance sampler, a probabilistic pseudo-label dispatcher, and self-paced weight learning rate coefficients. The instance sampler suppresses low-confidence instances by modeling pattern confidence, while the dispatcher adaptively assigns labels based on instance distinguishability. Self-paced learning adjusts training according to empirical label observations. Our framework overcomes limitations of traditional MIML methods and supports end-to-end optimization. Extensive experiments on one ANA dataset and three public medical MIML benchmarks demonstrate the superiority of our framework. On the ANA dataset, our model achieves up to +7.0% F1-Macro and +12.6% mAP gains over the best prior method, setting new state-of-the-art results. It also ranks top-2 across all key metrics on public datasets, reducing Hamming loss and one-error by up to 18.2% and 26.9%, respectively. The source code can be accessed at https://github.com/fletcherjiang/ANA-SelfPacedLearning. Guangwu Qian, Jiaxin Wu 0001, Qing Li 0001, Yongkang Wu, Xiaoyong Wei |
IEEE Trans. Medical Imaging | 3 |
| 2025 | Sound Bridge: Associating Egocentric and Exocentric Videos via Audio CuesabstractUnderstanding human behavior and environmental information in egocentric videos is very challenging due to the invisibility of some actions (e.g., laughing and sneezing) and the local nature of the first-person view. Leveraging the corresponding exocentric video to provide global context has shown promising results. However, existing visual-to-visual and visual-to-textual Ego-Exo video alignment methods struggle with the issue that some activities may have non-visual overlap. To address this, we propose using sound as a bridge, as audio is often consistent across Ego-Exo videos. However, direct audio-to-audio alignment lacks context. Thus, we introduce two context-aware sound modules: one aligns audio with vision via a visual-audio cross-attention module, and another aligns text with sound closed caption generated by LLM. Experimental results on two Ego-Exo video association benchmarks show that each of the proposed modules enhances the state-of-the-art methods. Moreover, the proposed sound-aware egocentric or exocentric representation boosts the performance of downstream tasks, such as action recognition of exocentric videos and scene recognition of egocentric videos. The code and models can be accessed at https://github.com/shhuangcoder/SoundBridge. Sihong Huang, Jiaxin Wu 0001, Xiaoyong Wei, Yi Cai 0001, Dongmei Jiang, Yaowei Wang 0001 |
CVPR | 2 |
| 2025 | Interactive Video Search with Multi-modal LLM Video Captioning
Yu-Tong Cheng, Jiaxin Wu 0001, Zhixin Ma 0001, Jiangshan He, Xiaoyong Wei, Chong-Wah Ngo |
MMM (5) | 2 |
| 2024 | Improving Interpretable Embeddings for Ad-hoc Video Search with Generative Captions and Multi-word Concept BankabstractAligning a user query and video clips in cross-modal latent space and that with semantic concepts are two mainstream approaches for ad-hoc video search (AVS). However, the effectiveness of existing approaches is bottlenecked by the small sizes of available video-text datasets and the low quality of concept banks, which results in the failures of unseen queries and the out-of-vocabulary problem. This paper addresses these two problems by constructing a new dataset and developing a multi-word concept bank. Specifically, capitalizing on a generative model, we construct a new dataset consisting of 7 million generated text and video pairs for pre-training. To tackle the out-of-vocabulary problem, we develop a multi-word concept bank based on syntax analysis to enhance the capability of a state-of-the- art interpretable AVS method in modelling relationships between query words. We also study the impact of current advanced features on the method. Experimental results show that the integration of the above-proposed elements doubles the R@1 performance of the AVS method on the MSRVTT dataset and improves the xinfAP on the TRECVid AVS query sets for 2016-2023 (eight years) by a margin from 2% to 77%, with an average about 20%. The code and model are available at https://github.com/nikkiwoo-gh/Improved-ITV. Jiaxin Wu 0001, Chong-Wah Ngo, Wing Kwong Chan |
ICMR | 1 |
| 2024 | Leveraging LLMs and Generative Models for Interactive Known-Item Video Search
Zhixin Ma 0001, Jiaxin Wu 0001, Chong-Wah Ngo |
MMM (4) | 2 |
| 2024 | (Un)likelihood Training for Interpretable EmbeddingabstractCross-modal representation learning has become a new normal for bridging the semantic gap between text and visual data. Learning modality agnostic representations in a continuous latent space, however, is often treated as a black-box data-driven training process. It is well known that the effectiveness of representation learning depends heavily on the quality and scale of training data. For video representation learning, having a complete set of labels that annotate the full spectrum of video content for training is highly difficult, if not impossible. These issues, black-box training and dataset bias, make representation learning practically challenging to be deployed for video understanding due to unexplainable and unpredictable results. In this article, we propose two novel training objectives, likelihood and unlikelihood functions, to unroll the semantics behind embeddings while addressing the label sparsity problem in training. The likelihood training aims to interpret semantics of embeddings beyond training labels, while the unlikelihood training leverages prior knowledge for regularization to ensure semantically coherent interpretation. With both training objectives, a new encoder-decoder network, which learns interpretable cross-modal representation, is proposed for ad-hoc video search. Extensive experiments on TRECVid and MSR-VTT datasets show that the proposed network outperforms several state-of-the-art retrieval models with a statistically significant performance margin. Jiaxin Wu 0001, Chong-Wah Ngo, Wing Kwong Chan, Zhijian Hou |
ACM Trans. Inf. Syst. | 1 |
| 2023 | Improving Query and Assessment Quality in Text-Based Interactive Video Retrieval EvaluationabstractDifferent task interpretations are a highly undesired element in interactive video retrieval evaluations. When a participating team focuses partially on a wrong goal, the evaluation results might become partially misleading. In this paper, we propose a process for refining known-item and open-set type queries, and preparing the assessors that judge the correctness of submissions to open-set queries. Our findings from recent years reveal that a proper methodology can lead to objective query quality improvements and subjective participant satisfaction with query clarity. Werner Bailer, Rahel Arnold, Vera Benz, Davide Coccomini, Anastasios Gkagkas, Gylfi Þór Guðmundsson, Silvan Heller, Björn Þór Jónsson 0001, Jakub Lokoc, Nicola Messina, Nick Pantelidis, Jiaxin Wu 0001 |
ICMR | 12 |
| 2023 | Reinforcement Learning Enhanced PicHunter for Interactive Search
Zhixin Ma 0001, Jiaxin Wu 0001, Weixiong Loo, Chong-Wah Ngo |
MMM (1) | 2 |
| 2022 | A Task Category Space for User-Centric Comparative Multimedia Search Evaluations
Jakub Lokoc, Werner Bailer, Kai Uwe Barthel, Cathal Gurrin, Silvan Heller, Björn Þór Jónsson 0001, Ladislav Peska, Luca Rossetto, Klaus Schöffmann, Lucia Vadicamo, Stefanos Vrochidis, Jiaxin Wu 0001 |
MMM (1) | 12 |
| 2022 | Reinforcement Learning-Based Interactive Video Search
Zhixin Ma 0001, Jiaxin Wu 0001, Zhijian Hou, Chong-Wah Ngo |
MMM (2) | 2 |
| 2021 | SQL-Like Interpretable Interactive Video Search
Jiaxin Wu 0001, Phuong Anh Nguyen 0002, Zhixin Ma 0001, Chong-Wah Ngo |
MMM (2) | 1 |
| 2021 | Is the Reign of Interactive Search Eternal? Findings from the Video Browser Showdown 2020abstractComprehensive and fair performance evaluation of information retrieval systems represents an essential task for the current information age. Whereas Cranfield-based evaluations with benchmark datasets support development of retrieval models, significant evaluation efforts are required also for user-oriented systems that try to boost performance with an interactive search approach. This article presents findings from the 9th Video Browser Showdown, a competition that focuses on a legitimate comparison of interactive search systems designed for challenging known-item search tasks over a large video collection. During previous installments of the competition, the interactive nature of participating systems was a key feature to satisfy known-item search needs, and this article continues to support this hypothesis. Despite the fact that top-performing systems integrate the most recent deep learning models into their retrieval process, interactive searching remains a necessary component of successful strategies for known-item search tasks. Alongside the description of competition settings, evaluated tasks, participating teams, and overall results, this article presents a detailed analysis of query logs collected by the top three performing systems, SOMHunter, VIRET, and vitrivr. The analysis provides a quantitative insight to the observed performance of the systems and constitutes a new baseline methodology for future events. The results reveal that the top two systems mostly relied on temporal queries before a correct frame was identified. An interaction log analysis complements the result log findings and points to the importance of result set and video browsing approaches. Finally, various outlooks are discussed in order to improve the Video Browser Showdown challenge in the future. Jakub Lokoc, Patrik Veselý, Frantisek Mejzlík, Gregor Kovalcík, Tomás Soucek, Luca Rossetto, Klaus Schöffmann, Werner Bailer, Cathal Gurrin, Loris Sauter, Jaeyub Song, Stefanos Vrochidis, Jiaxin Wu 0001, Björn Þór Jónsson 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 13 |
| 2020 | Interpretable Embedding for Ad-Hoc Video SearchabstractAnswering query with semantic concepts has long been the mainstream approach for video search. Until recently, its performance is surpassed by concept-free approach, which embeds queries in a joint space as videos. Nevertheless, the embedded features as well as search results are not interpretable, hindering subsequent steps in video browsing and query reformulation. This paper integrates feature embedding and concept interpretation into a neural network for unified dual-task learning. In this way, an embedding is associated with a list of semantic concepts as an interpretation of video content. This paper empirically demonstrates that, by using either the embedding features or concepts, considerable search improvement is attainable on TRECVid benchmarked datasets. Concepts are not only effective in pruning false positive videos, but also highly complementary to concept-free search, leading to large margin of improvement compared to state-of-the-art approaches. Jiaxin Wu 0001, Chong-Wah Ngo |
ACM Multimedia | 1 |
| 2020 | VIREO @ Video Browser Showdown 2020
Phuong Anh Nguyen 0002, Jiaxin Wu 0001, Chong-Wah Ngo, Danny Francis, Benoit Huet |
MMM (2) | 2 |
| 2020 | Dynamic graph convolutional network for multi-video summarization
Jiaxin Wu 0001, Shenghua Zhong, Yan Liu 0004 |
Pattern Recognit. | 1 |
| 2019 | MvsGCN: A Novel Graph Convolutional Network for Multi-video SummarizationabstractMulti-video summarization, which tries to generate a single summary for a collection of video, is an important task in dealing with ever-growing video data. In this paper, we are the first to propose a graph convolutional network for multi-video summarization. The novel network measures the importance and relevance of each video shot in its own video as well as in the whole video collection. The important node sampling method is proposed to emphasize the effective features which are more possible to be selected as the final video summary. Two strategies are proposed to integrate into the network to solve the inherent class imbalance problem in the task of video summarization. The loss regularization for diversity is used to encourage a diverse summary to be generated. Extensive experiments are carried out, and in comparison with traditional and recent graph models and the state-of-the-art video summarization methods, our proposed model is effective in generating a representative summary for multiple videos with good diversity. It also achieves state-of-the-art performance on two standard video summarization datasets. Jiaxin Wu 0001, Shenghua Zhong, Yan Liu 0004 |
ACM Multimedia | 1 |
| 2019 | Video summarization via spatio-temporal deep architecture
Shenghua Zhong, Jiaxin Wu 0001, Jianmin Jiang |
Neurocomputing | 2 |
| 2018 | Foveated convolutional neural networks for video summarization
Jiaxin Wu 0001, Shenghua Zhong, Zheng Ma 0004, Stephen J. Heinen, Jianmin Jiang |
Multim. Tools Appl. | 1 |
| 2017 | Adaptive Dehaze Method for Aerial Image Processing
Rong-Qin Xu, Shenghua Zhong, Gaoyang Tang, Jiaxin Wu 0001, Yingying Zhu 0001 |
PSIVT | 4 |
| 2017 | A novel clustering method for static video summarization
Jiaxin Wu 0001, Shenghua Zhong, Jianmin Jiang, Yunyun Yang |
Multim. Tools Appl. | 1 |
| 2016 | Visual Orientation Inhomogeneity Based Convolutional Neural NetworksabstractThe details of oriented visual stimuli are better resolved when they are horizontal or vertical rather than oblique. This "oblique effect" has been researched and confirmed in numerous research studies, including behavioral studies and neurophysiological and neuroimaging findings. Although the "oblique effect" has influence in many fields, little research integrated it into computational models. In this paper, we try to explore this inhomogeneity of visual orientation based on Convolutional neural networks (CNNs) in image recognition. We validate that visual orientation inhomogeneity CNNs can achieve comparable performance with higher computational efficiency on various datasets. We can also get the conclusion that, compared with the cardinal information, oblique information is indeed less useful in natural color image recognition. Through the exploration of the proposed model on image recognition, we gain more understanding of the inhomogeneity of visual orientation. It also illuminates a wide range of opportunities for integrating the inhomogeneity of visual orientation with other computational models. Shenghua Zhong, Jiaxin Wu 0001, Yingying Zhu 0001, Peiqi Liu, Jianmin Jiang, Yan Liu 0004 |
ICTAI | 2 |