Fengshun Xiao

dblp:242/7863 · DBLP profile ↗
← Back
5ranked-venue papers
1as first author
3since 2021 · last 2026
0009-0006-2175-2801ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 5 · 1 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
5 papers
Vision and language · 28% Video understanding and tracking · 18% Machine translation · 15%

Topics — the 13 heaviest of 15, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › Video understanding and tracking
video question answering
1.922026
Parse, Align and Aggregate: Graph-Driven Compositional Reasoning for Video Question Answering · IEEE Trans. Pattern Anal. Mach. Intell. 2026
Divide and Conquer: Exploring Language-centric Tree Reasoning for Video Question-Answering · ICML 2025
Computer vision › Vision and language › multimodal reasoning
compositional reasoning
1.012026
Parse, Align and Aggregate: Graph-Driven Compositional Reasoning for Video Question Answering · IEEE Trans. Pattern Anal. Mach. Intell. 2026
Natural language and speech › Question answering and dialogue systems
multimodal question answering
1.012026
Parse, Align and Aggregate: Graph-Driven Compositional Reasoning for Video Question Answering · IEEE Trans. Pattern Anal. Mach. Intell. 2026
Computer vision › Vision and language
multimodal reasoning
1.012026
Parse, Align and Aggregate: Graph-Driven Compositional Reasoning for Video Question Answering · IEEE Trans. Pattern Anal. Mach. Intell. 2026
Natural language and speech › Machine translation
neural machine translation
1.022022
Explicit Alignment Learning for Neural Machine Translation · IJCAI 2022
Lattice-Based Transformer Encoder for Neural Machine Translation · ACL (1) 2019
Computer vision › Vision and language › vision-language model › multimodal large language model
multimodal large language model reasoning
0.912025
Divide and Conquer: Exploring Language-centric Tree Reasoning for Video Question-Answering · ICML 2025
Natural language and speech › Machine translation › statistical machine translation
word alignment
0.612022
Explicit Alignment Learning for Neural Machine Translation · IJCAI 2022
Machine learning › Representation and self-supervised learning › word representation
contextualized word representation
0.412020
Hierarchical Contextualized Representation for Named Entity Recognition · AAAI 2020
Machine learning › Representation and self-supervised learning › word representation
contextual representation
0.412020
Hierarchical Contextualized Representation for Named Entity Recognition · AAAI 2020
Natural language and speech › Information extraction and text analysis
named entity recognition
0.412020
Hierarchical Contextualized Representation for Named Entity Recognition · AAAI 2020
Machine learning › Trustworthy machine learning
interpretability
0.312026
Parse, Align and Aggregate: Graph-Driven Compositional Reasoning for Video Question Answering · IEEE Trans. Pattern Anal. Mach. Intell. 2026
Machine learning › Deep learning architectures and training
recurrent neural network
0.112020
Hierarchical Contextualized Representation for Named Entity Recognition · AAAI 2020
Machine learning › Deep learning architectures and training
transformer
0.112019
Lattice-Based Transformer Encoder for Neural Machine Translation · ACL (1) 2019

Methods — techniques the papers use, named apart from their topics

multimodal large language model · 1.0graph-driven reasoning · 1.0retrieval-augmented generation · 0.9few-shot prompting · 0.9explicit alignment learning · 0.6embedding mixup · 0.6attention weights · 0.6label embedding attention · 0.4key-value memory network · 0.4BiLSTM · 0.4
YearPublicationVenuePosition
2026 Parse, Align and Aggregate: Graph-Driven Compositional Reasoning for Video Question Answering
abstract
Video Question-Answering (VideoQA) enables machines to interpret and respond to complex video content, advancing human-computer interaction. However, existing multimodal large language models (MLLMs) often provide incomplete or opaque explanations and existing benchmarks mainly focus on the correction of final answers, limiting insight into their reasoning processes and hindering both transparency and verifiability. To address this gap, we propose the Question Parsing, Video Alignment and Answer Aggregation framework (QPVA$^{3}$3), which leverages a compositional graph to drive visual and logical reasoning in VideoQA. Specifically, QPVA$^{3}$3 consists of three core components, the planner, executor, and reasoner to generate the compositional graph and conduct graph-driven reasoning. For the original question, the planner parses it into the compositional graph, capturing the underlying reasoning logic and structuring it into a series of interconnected questions. For each question in compositional graph, the executor aligns the video by selecting relevant video clips and generates answers, ensuring accurate, context-specific responses. For each question with its first-order descents, the reasoner aggregates answers by integrating reasoning logic with visual evidence, resolving conflicts to produce a coherent and accurate response. Moreover, to assess the performance of existing MLLMs in the reasoning processes of VideoQA, we introduce novel compositional consistency metrics and construct a VideoQA benchmark (QPVA$^{3}$3 Bench) with 3,492 question-video tuples, each annotated with detailed compositional graphs and fine-grained answers. We evaluate the QPVA$^{3}$3 framework on QPVA$^{3}$3 Bench and 5 other VideoQA benchmarks. Experimental results demonstrate that our framework improves both consistency and accuracy compared to baselines, leading to a more transparent and verifiable VideoQA system. This approach has the potential to advance the field, as supported by our comprehensive evaluation and benchmarking efforts.
Jiangtong Li, Zhaohe Liao, Fengshun Xiao, Tianjiao Li 0001, Qiang Zhang 0055, Haohua Zhao 0001, Li Niu 0002, Guang Chen 0001, Liqing Zhang 0001, Changjun Jiang 0002
IEEE Trans. Pattern Anal. Mach. Intell.3
2025 Divide and Conquer: Exploring Language-centric Tree Reasoning for Video Question-Answering
abstract
Video Question-Answering (VideoQA) remains challenging in achieving advanced cognitive reasoning due to the uncontrollable and opaque reasoning processes in existing Multimodal Large Language Models (MLLMs). To address this issue, we propose a novel Language-centric Tree Reasoning (LTR) framework that targets on enhancing the reasoning ability of models. In detail, it recursively divides the original question into logically manageable parts and conquers them piece by piece, enhancing the reasoning capabilities and interpretability of existing MLLMs. Specifically, in the first stage, the LTR focuses on language to recursively generate a language-centric logical tree, which gradually breaks down the complex cognitive question into simple perceptual ones and plans the reasoning path through a RAG-based few-shot approach. In the second stage, with the aid of video content, the LTR performs bottom-up logical reasoning within the tree to derive the final answer along with the traceable reasoning path. Experiments across 11 VideoQA benchmarks demonstrate that our LTR framework significantly improves both accuracy and interpretability compared to state-of-the-art MLLMs. To our knowledge, this is the first work to implement a language-centric logical tree to guide MLLM reasoning in VideoQA, paving the way for language-centric video understanding from perception to cognition.
Zhaohe Liao, Jiangtong Li, Qingyang Liu 0002, Fengshun Xiao, Tianjiao Li 0001, Qiang Zhang 0055, Guang Chen 0001, Li Niu 0002, Changjun Jiang 0002, Liqing Zhang 0001
ICML5
2022 Explicit Alignment Learning for Neural Machine Translation
abstract
Even though neural machine translation (NMT) has become the state-of-the-art solution for end-to-end translation, it still suffers from a lack of translation interpretability, which may be conveniently enhanced by explicit alignment learning (EAL), as performed in traditional statistical machine translation (SMT). To provide the benefits of both NMT and SMT, this paper presents a novel model design that enhances NMT with an additional training process for EAL, in addition to the end-to-end translation training. Thus, we propose two approaches an explicit alignment learning approach, in which we further remove the need for the additional alignment model, and perform embedding mixup with the alignment based on encoder--decoder attention weights in the NMT model. We conducted experiments on both small-scale (IWSLT14 De->En and IWSLT13 Fr->En) and large-scale (WMT14 En->De, En->Fr, WMT17 Zh->En) benchmarks. Evaluation results show that our EAL methods significantly outperformed strong baseline methods, which shows the effectiveness of EAL. Further explorations show that the translation improvements are due to a better spatial alignment of the source and target language embeddings. Our method improves translation performance without the need to increase model parameters and training data, which verifies that the idea of incorporating techniques of SMT into NMT is worthwhile.
Zuchao Li, Hai Zhao 0001, Fengshun Xiao, Masao Utiyama, Eiichiro Sumita
IJCAI3
2020 Hierarchical Contextualized Representation for Named Entity Recognition
abstract
Named entity recognition (NER) models are typically based on the architecture of Bi-directional LSTM (BiLSTM). The constraints of sequential nature and the modeling of single input prevent the full utilization of global information from larger scope, not only in the entire sentence, but also in the entire document (dataset). In this paper, we address these two deficiencies and propose a model augmented with hierarchical contextualized representation: sentence-level representation and document-level representation. In sentence-level, we take different contributions of words in a single sentence into consideration to enhance the sentence representation learned from an independent BiLSTM via label embedding attention mechanism. In document-level, the key-value memory network is adopted to record the document-aware information for each unique word which is sensitive to similarity of context information. Our two-level hierarchical contextualized representations are fused with each input token embedding and corresponding hidden state of BiLSTM, respectively. The experimental results on three benchmark NER datasets (CoNLL-2003 and Ontonotes 5.0 English datasets, CoNLL-2002 Spanish dataset) show that we establish new state-of-the-art results.
Ying Luo 0012, Fengshun Xiao, Hai Zhao 0001
AAAI2
2019 Lattice-Based Transformer Encoder for Neural Machine Translation
abstract
Neural machine translation (NMT) takes deterministic sequences for source representations.However, either wordlevel or subword-level segmentations have multiple choices to split a source sequence with different word segmentors or different subword vocabulary sizes.We hypothesize that the diversity in segmentations may affect the NMT performance.To integrate different segmentations with the state-of-the-art NMT model, Transformer, we propose lattice-based encoders to explore effective word or subword representation in an automatic way during training.We propose two methods: 1) lattice positional encoding and 2) lattice-aware self-attention.These two methods can be used together and show complementary to each other to further improve translation performance.Experiment results show superiorities of lattice-based encoders in word-level and subword-level representations over conventional Transformer encoder.
Fengshun Xiao, Jiangtong Li, Hai Zhao 0001, Rui Wang 0015, Kehai Chen
ACL (1)1