VLDB 2026 Research / reviewers in the wild / expert
Zhenxin Xiao
dblp:242/8026
· DBLP profile ↗
7ranked-venue papers
1as first author
4since 2021 · last 2026
0000-0003-2762-5097ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 4 · 1 first-author · 3 since 2021Artificial intelligence and machine learning · 2 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Brand crisis and recovery in livestream commerce: A psychological contract violation theory perspective
Zhenxin Xiao, Matthew K. O. Lee |
Decis. Support Syst. | 3 |
| 2026 | The power of language: Other-focused linguistic style and sales performance on home-sharing platforms
Jinming Dang, Chenze Wang, Christy M. K. Cheung, Zhenxin Xiao |
Inf. Manag. | 6 |
| 2025 | Self-disclosure in online social networks: The needs-affordances-features perspective
Zhenxin Xiao, Christy M. K. Cheung |
Inf. Manag. | 1 |
| 2022 | A dedication-constraint model of consumer switching behavior in mobile payment applications
Zhenxin Xiao |
Inf. Manag. | 3 |
| 2020 | Learning Directional Sentence-Pair Embedding for Natural Language Reasoning (Student Abstract)abstractEnabling the models with the ability of reasoning and inference over text is one of the core missions of natural language understanding. Despite deep learning models have shown strong performance on various cross-sentence inference benchmarks, recent work has shown that they are leveraging spurious statistical cues rather than capturing deeper implied relations between pairs of sentences. In this paper, we show that the state-of-the-art language encoding models are especially bad at modeling directional relations between sentences by proposing a new evaluation task: Cause-and-Effect relation prediction task. Back by our curated Cause-and-Effect Relation dataset (Cℰℛ), we also demonstrate that a mutual attention mechanism can guide the model to focus on capturing directional relations between sentences when added to existing transformer-based models. Experiment results show that the proposed approach improves the performance on downstream applications, such as the abductive reasoning task. Zhenxin Xiao, Kai-Wei Chang 0001 |
AAAI | 2 |
| 2019 | Cross-Modal Interaction Networks for Query-Based Moment Retrieval in VideosabstractQuery-based moment retrieval aims to localize the most relevant moment in an untrimmed video according to the given natural language query. Existing works often only focus on one aspect of this emerging task, such as the query representation learning, video context modeling or multi-modal fusion, thus fail to develop a comprehensive system for further performance improvement. In this paper, we introduce a novel Cross-Modal Interaction Network (CMIN) to consider multiple crucial factors for this challenging task, including (1) the syntactic structure of natural language queries; (2) long-range semantic dependencies in video context and (3) the sufficient cross-modal interaction. Specifically, we devise a syntactic GCN to leverage the syntactic structure of queries for fine-grained representation learning, propose a multi-head self-attention to capture long-range semantic dependencies from video context, and next employ a multi-stage cross-modal interaction to explore the potential relations of video and query contents. The extensive experiments demonstrate the effectiveness of our proposed method. Zhijie Lin 0001, Zhou Zhao 0001, Zhenxin Xiao |
SIGIR | 4 |
| 2019 | Long-Form Video Question Answering via Dynamic Hierarchical Reinforced NetworksabstractOpen-ended long-form video question answering is a challenging task in visual information retrieval, which automatically generates a natural language answer from the referenced long-form video contents according to a given question. However, the existing works mainly focus on short-form video question answering, due to the lack of modeling semantic representations from long-form video contents. In this paper, we introduce a dynamic hierarchical reinforced network for open-ended long-form video question answering, which employs an encoder-decoder architecture with a dynamic hierarchical encoder and a reinforced decoder. Concretely, we first propose a frame-level dynamic long-short term memory (LSTM) network with binary segmentation gate to learn frame-level semantic representations according to the given question. We then develop a segment-level highway LSTM network with a question-aware highway gate for segment-level semantic modeling. Furthermore, we devise the reinforced decoder with a hierarchical attention mechanism to generate natural language answers. We construct a large-scale long-form video question answering dataset. The extensive experiments on the long-form dataset and another public short-form dataset show the effectiveness of our method. Zhou Zhao 0001, Shuwen Xiao, Zhenxin Xiao, Jun Yu 0002, Deng Cai 0001, Fei Wu 0001 |
IEEE Trans. Image Process. | 4 |