EDBT 2026 Demo / reviewers in the wild / expert
Zhixin Ma 0001
dblp:45/524-1
· DBLP profile ↗
11ranked-venue papers
5as first author
9since 2021 · last 2025
0000-0002-9127-4175ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 9 · 5 first-author · 8 since 2021Artificial intelligence and machine learning · 2 · 1 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Seeing Culture: A Benchmark for Visual Reasoning and GroundingabstractBurak Satar, Zhixin Ma, Patrick Amadeus Irawan, Wilfried Ariel Mulyawan, Jing Jiang, Ee-Peng Lim, Chong-Wah Ngo. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Burak Satar, Zhixin Ma 0001, Patrick Amadeus Irawan, Wilfried A. Mulyawan, Jing Jiang 0001, Ee-Peng Lim, Chong-Wah Ngo |
EMNLP | 2 |
| 2025 | Robust Relevance Feedback for Interactive Known-Item Video SearchabstractKnown-item search (KIS) involves only a single search target, making relevance feedback-typically a powerful technique for efficiently identifying multiple positive examples to infer user intent-inapplicable. PicHunter addresses this issue by asking users to select the top-k most similar examples to the unique search target from a displayed set. Under ideal conditions, when the user's perception aligns closely with the machine's perception of similarity, consistent and precise judgments can elevate the target to the top position within a few iterations. However, in practical scenarios, expecting users to provide consistent judgments is often unrealistic, especially when the underlying embedding features used for similarity measurements lack interpretability. To enhance robustness, we first introduce a pairwise relative judgment feedback that improves the stability of top-k selections by mitigating the impact of misaligned feedback. Then, we decompose user perception into multiple sub-perceptions, each represented as an independent embedding space. This approach assumes that users may not consistently align with a single representation but are more likely to align with one or several among multiple representations. We develop a predictive user model that estimates the combination of sub-perceptions based on each user feedback instance. The predictive user model is then trained to filter out the misaligned sub-perceptions. Experimental evaluations on the large-scale open-domain dataset V3C indicate that the proposed model can optimize over 60% search targets to the top rank when their initial ranks at the search depth between 10 and 50. Even for targets initially ranked between 1,000 and 5,000, the model achieves a success rate exceeding 40% in optimizing ranks to the top, demonstrating the enhanced robustness of relevance feedback in KIS despite inconsistent feedback. Zhixin Ma 0001, Chong-Wah Ngo |
ICMR | 1 |
| 2025 | Interactive Video Search with Multi-modal LLM Video Captioning
Yu-Tong Cheng, Jiaxin Wu 0001, Zhixin Ma 0001, Jiangshan He, Xiaoyong Wei, Chong-Wah Ngo |
MMM (5) | 3 |
| 2024 | Leveraging LLMs and Generative Models for Interactive Known-Item Video Search
Zhixin Ma 0001, Jiaxin Wu 0001, Chong-Wah Ngo |
MMM (4) | 1 |
| 2023 | Reinforcement Learning Enhanced PicHunter for Interactive Search
Zhixin Ma 0001, Jiaxin Wu 0001, Weixiong Loo, Chong-Wah Ngo |
MMM (1) | 1 |
| 2023 | Interactive video retrieval in the age of effective joint embedding deep models: lessons from the 11th VBS
Jakub Lokoc, Stelios Andreadis, Werner Bailer, Aaron Duane, Cathal Gurrin, Zhixin Ma 0001, Nicola Messina, Thao-Nhu Nguyen, Ladislav Peska, Luca Rossetto, Loris Sauter, Konstantin Schall, Klaus Schöffmann, Omar Shahbaz Khan, Florian Spiess 0001, Lucia Vadicamo, Stefanos Vrochidis |
Multim. Syst. | 6 |
| 2022 | Interactive Video Corpus Moment Retrieval using Reinforcement LearningabstractKnown-item video search is effective with human-in-the-loop to interactively investigate the search result and refine the initial query. Nevertheless, when the first few pages of results are swamped with visually similar items, or the search target is hidden deep in the ranked list, finding the know-item target usually requires a long duration of browsing and result inspection. This paper tackles the problem by reinforcement learning, aiming to reach a search target within a few rounds of interaction by long-term learning from user feedbacks. Specifically, the system interactively plans for navigation path based on feedback and recommends a potential target that maximizes the long-term reward for user comment. We conduct experiments for the challenging task of video corpus moment retrieval (VCMR) to localize moments from a large video corpus. The experimental results on TVR and DiDeMo datasets verify that our proposed work is effective in retrieving the moments that are hidden deep inside the ranked lists of CONQUER and HERO, which are the state-of-the-art auto-search engines for VCMR. Zhixin Ma 0001, Chong-Wah Ngo |
ACM Multimedia | 1 |
| 2022 | Reinforcement Learning-Based Interactive Video Search
Zhixin Ma 0001, Jiaxin Wu 0001, Zhijian Hou, Chong-Wah Ngo |
MMM (2) | 1 |
| 2021 | SQL-Like Interpretable Interactive Video Search
Jiaxin Wu 0001, Phuong Anh Nguyen 0002, Zhixin Ma 0001, Chong-Wah Ngo |
MMM (2) | 3 |
| 2020 | Re-examining the Role of Schema Linking in Text-to-SQLabstractIn existing sophisticated text-to-SQL models, schema linking is often considered as a simple, minor component, belying its importance.By providing a schema linking corpus based on the Spider text-to-SQL dataset, we systematically study the role of schema linking.We also build a simple BERT-based baseline, called Schema-Linking SQL (SLSQL) to perform a data-driven study.We find when schema linking is done well, SLSQL demonstrates good performance on Spider despite its structural simplicity.Many remaining errors are attributable to corpus noise.This suggests schema linking is the crux for the current textto-SQL task.Our analytic studies provide insights on the characteristics of schema linking for future developments of text-to-SQL tasks. 1 Wenqiang Lei, Zhixin Ma 0001, Tian Gan 0002, Wei Lu 0011, Min-Yen Kan, Tat-Seng Chua |
EMNLP (1) | 3 |
| 2019 | Learn to Gesture: Let Your Body SpeakabstractPresentation is one of the most important and vivid methods to deliver information to audience. Apart from the content of presentation, how the speaker behaves during presentation makes a big difference. In other words, gestures, as part of the visual perception and synchronized with verbal information, express some subtle information that the voice or words alone cannot deliver. One of the most effective ways to improve presentation is to practice through feedback/suggestions by an expert. However, hiring human experts is expensive thus impractical most of the time. Towards this end, we propose a speech to gesture network (POSE) to generate exemplary body language given a vocal behavior speech as input. Specifically, we build an "expert" Speech-Gesture database based on the featured TED talk videos, and design a two-layer attentive recurrent encoder-decoder network to learn the translation from speech to gesture, as well as the hierarchical structure within gestures. Lastly, given a speech audio sequence, the appropriate gesture will be generated and visualized for a more effective communication. Both objective and subjective validation show the effectiveness of our proposed method. Tian Gan 0002, Zhixin Ma 0001, Yuxiao Lu, Xuemeng Song, Liqiang Nie |
MMAsia | 2 |