Qian Yang 0007

dblp:15/3199-7 · DBLP profile ↗
← Back
5ranked-venue papers
3as first author
5since 2021 · last 2025
0009-0009-8975-6982ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 4 · 2 first-author · 4 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
5 papers
Vision and language · 23% Language models and text generation · 20% Question answering and dialogue systems · 16%
Software engineering, system software, and programming languages
1 paper
Program synthesis and code generation · 77% Software testing · 23%
Databases, data mining, and information retrieval
1 paper
Information retrieval · 100%

Topics — the 16 heaviest of 17, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Language models and text generation
hallucination detection
0.912025
CodeHalu: Investigating Code Hallucinations in LLMs via Execution-based Verification · AAAI 2025
Program synthesis and code generation
code generation with language models
0.912025
CodeHalu: Investigating Code Hallucinations in LLMs via Execution-based Verification · AAAI 2025
Machine learning › Trustworthy machine learning
uncertainty estimation
0.812024
Decompose and Compare Consistency: Measuring VLMs' Answer Reliability via Task-Decomposition Consistency Comparison · EMNLP 2024
Computer vision › Vision and language › vision-language model
vision-language model evaluation
0.812024
Decompose and Compare Consistency: Measuring VLMs' Answer Reliability via Task-Decomposition Consistency Comparison · EMNLP 2024
Machine learning › Generative modeling
autoregressive model
0.712023
Fast and Robust Online Handwritten Chinese Character Recognition With Deep Spatial and Contextual Information Fusion Network · IEEE Trans. Multim. 2023
Computer vision › Image recognition and object detection › handwriting recognition
handwritten text recognition
0.712023
Fast and Robust Online Handwritten Chinese Character Recognition With Deep Spatial and Contextual Information Fusion Network · IEEE Trans. Multim. 2023
Natural language and speech › Question answering and dialogue systems › reasoning-based question answering
multi-hop question answering
0.712023
Enhancing Multi-modal Multi-hop Question Answering via Structured Knowledge and Unified Retrieval-Generation · ACM Multimedia 2023
Natural language and speech › Question answering and dialogue systems › multimodal question answering
multimodal multihop question answering
0.712023
Enhancing Multi-modal Multi-hop Question Answering via Structured Knowledge and Unified Retrieval-Generation · ACM Multimedia 2023
Computer vision › Image recognition and object detection › handwriting recognition
online handwritten chinese character recognition
0.712023
Fast and Robust Online Handwritten Chinese Character Recognition With Deep Spatial and Contextual Information Fusion Network · IEEE Trans. Multim. 2023
Knowledge, reasoning and agents › Knowledge representation and reasoning
explanation generation
0.612022
Chunk-aware Alignment and Lexical Constraint for Visual Entailment with Natural Language Explanations · ACM Multimedia 2022
Natural language and speech › Language models and text generation › controllable text generation
lexically constrained generation
0.612022
Chunk-aware Alignment and Lexical Constraint for Visual Entailment with Natural Language Explanations · ACM Multimedia 2022
Computer vision › Vision and language
multimodal reasoning
0.612022
Chunk-aware Alignment and Lexical Constraint for Visual Entailment with Natural Language Explanations · ACM Multimedia 2022
Computer vision › Vision and language › visual reasoning
visual entailment
0.612022
Chunk-aware Alignment and Lexical Constraint for Visual Entailment with Natural Language Explanations · ACM Multimedia 2022
Software testing
execution-based verification
0.312025
CodeHalu: Investigating Code Hallucinations in LLMs via Execution-based Verification · AAAI 2025
Natural language and speech › Language models and text generation
self-consistency
0.212024
Decompose and Compare Consistency: Measuring VLMs' Answer Reliability via Task-Decomposition Consistency Comparison · EMNLP 2024
Knowledge, reasoning and agents › Knowledge representation and reasoning
structured knowledge
0.212023
Enhancing Multi-modal Multi-hop Question Answering via Structured Knowledge and Unified Retrieval-Generation · ACM Multimedia 2023

Methods — techniques the papers use, named apart from their topics

execution-based verification · 1.7dynamic detection · 1.7unified retrieval-generation decoder · 1.3entity-centered fusion encoder · 1.3task decomposition · 0.8consistency comparison · 0.8transformer · 0.7recurrent neural network · 0.7multimodal fusion · 0.7convolutional neural network · 0.7
YearPublicationVenuePosition
2025 CodeHalu: Investigating Code Hallucinations in LLMs via Execution-based Verification
abstract
Large Language Models (LLMs) have made significant progress in code generation, offering developers groundbreaking automated programming support. However, LLMs often generate code that is syntactically correct and even semantically plausible, but may not execute as expected or fulfill specified requirements. This phenomenon of hallucinations in the code domain has not been systematically explored. To advance the community's understanding and research on this issue, we introduce the concept of code hallucinations and propose a classification method for code hallucination based on execution verification. We categorize code hallucinations into four main types: mapping, naming, resource, and logic hallucinations, with each category further divided into different subcategories to understand and address the unique challenges faced by LLMs in code generation with finer granularity. Additionally, we present a dynamic detection algorithm called CodeHalu designed to detect and quantify code hallucinations. We also introduce the CodeHaluEval benchmark, which includes 8,883 samples from 699 tasks, to systematically and quantitatively evaluate code hallucinations. By evaluating 17 popular LLMs using this benchmark, we reveal significant differences in their accuracy and reliability in code generation, offering detailed insights for further improving the code generation capabilities of LLMs.
Weixiang Yan, Qian Yang 0007, Xuandong Zhao, Qian Chen 0003, Wen Wang 0001, Dawn Song
AAAI3
2024 Decompose and Compare Consistency: Measuring VLMs' Answer Reliability via Task-Decomposition Consistency Comparison
abstract
Despite tremendous advancements, current state-of-the-art Vision-Language Models (VLMs) are still far from perfect.They tend to hallucinate and may generate biased responses.In such circumstances, having a way to assess the reliability of a given response generated by a VLM is quite useful.Existing methods, such as estimating uncertainty using answer likelihoods or prompt-based confidence generation, often suffer from overconfidence.Other methods use self-consistency comparison but are affected by confirmation biases.To alleviate these, we propose Decompose and Compare Consistency (DeCC) for reliability measurement.By comparing the consistency between the direct answer generated using the VLM's internal reasoning process, and the indirect answers obtained by decomposing the question into sub-questions and reasoning over the sub-answers produced by the VLM, DeCC measures the reliability of VLM's direct answer.Experiments across six vision-language tasks with three VLMs show DeCC's reliability estimation achieves better correlation with task accuracy compared to the existing methods.The code is publicly available at https: //github.com/MyLittleChange/DeCC.
Qian Yang 0007, Weixiang Yan, Aishwarya Agrawal
EMNLP1
2023 Enhancing Multi-modal Multi-hop Question Answering via Structured Knowledge and Unified Retrieval-Generation
abstract
Multi-modal multi-hop question answering involves answering a question by reasoning over multiple input sources from different modalities. Existing methods often retrieve evidences separately and then use a language model to generate an answer based on the retrieved evidences, and thus do not adequately connect candidates and are unable to model the interdependent relations during retrieval. Moreover, the pipelined approaches of retrieval and generation might result in poor generation performance when retrieval performance is low. To address these issues, we propose a Structured Knowledge and Unified Retrieval-Generation (SKURG) approach. SKURG employs an Entity-centered Fusion Encoder to align sources from different modalities using shared entities. It then uses a unified Retrieval-Generation Decoder to integrate intermediate retrieval results for answer generation and also adaptively determine the number of retrieval steps. Extensive experiments on two representative multi-modal multi-hop QA datasets MultimodalQA and WebQA demonstrate that SKURG outperforms the state-of-the-art models in both source retrieval and answer generation performance with fewer parameters1.
Qian Yang 0007, Qian Chen 0003, Wen Wang 0001, Baotian Hu, Min Zhang 0005
ACM Multimedia1
2023 Fast and Robust Online Handwritten Chinese Character Recognition With Deep Spatial and Contextual Information Fusion Network
abstract
Deep convolutional neuralnetworks have achieved fairly high accuracy for single online handwritten Chinese character recognition (SOLHCCR). However, in real application scenarios, users always write multiple characters to form a complete sentence, and previous contextual information holds significant potential for improving the accuracy, robustness and efficiency of recognition. In this work, we first propose a simple and straightforward model named the vanilla compositional network (VCN) by coupling convolutional neural network with a sequence modeling architecture (i.e., a recurrent neural network or Transformer), which exploits the handwritten character’s previous contextual information. Although VCN performs much better than the previous state-of-the-art SOLHCCR models, it is a two-stage architecture in nature. It suffers from high fragility when confronting with poorly written characters such as sloppy writing, and missing or broken strokes, due to relying heavily on contextual information. To improve the robustness of the OLHCCR model, we further propose a novel deep spatial & contextual information fusion network (DSCIFN). It utilizes an autoregresssive framework pre-trained on a large-scale sentence corpora as the backbone component, and highly integrates the spatial features of handwritten characters and their previous contextual information in a multi-layer fusion module. To verify the effectiveness of models, we reorganize a new form of online Chinese handwritten character with its previous context dataset, named OHCCC. Extensive experimental results demonstrate that DSCIFN achieves state-of-the-art performance and has increased strong robustness compared to VCN and previous SOLHCCR models. The in-depth empirical analysis and case study indicate that DSCIFN can significantly improve the efficiency of handwriting input because it does not need complete strokes to recognize a handwritten Chinese character precisely.
Yunxin Li, Qian Yang 0007, Qingcai Chen, Baotian Hu, Xiaolong Wang 0001, Lin Ma 0002
IEEE Trans. Multim.2
2022 Chunk-aware Alignment and Lexical Constraint for Visual Entailment with Natural Language Explanations
abstract
Visual Entailment with natural language explanations aims to infer the relationship between a text-image pair and generate a sentence to explain the decision-making process. Previous methods rely mainly on a pre-trained vision-language model to perform the relation inference and a language model to generate the corresponding explanation. However, the pre-trained vision-language models mainly build token-level alignment between text and image yet ignore the high-level semantic alignment between the phrases (chunks) and visual contents, which is critical for vision-language reasoning. Moreover, the explanation generator based only on the encoded joint representation does not explicitly consider the critical decision-making points of relation inference. Thus the generated explanations are less faithful to visual-language reasoning. To mitigate these problems, we propose a unified Chunk-aware Alignment and Lexical Constraint based method, dubbed as CALeC. It contains a Chunk-aware Semantic Interactor (arr. CSI), a relation inferrer, and a Lexical Constraint-aware Generator (arr. LeCG). Specifically, CSI exploits the sentence structure inherent in language and various image regions to build chunk-aware semantic alignment. Relation inferrer uses an attention-based reasoning network to incorporate the token-level and chunk-level vision-language representations. LeCG utilizes lexical constraints to expressly incorporate the words or chunks focused by the relation inferrer into explanation generation, improving the faithfulness and informativeness of the explanations. We conduct extensive experiments on three datasets, and experimental results indicate that CALeC significantly outperforms other competitor models on inference accuracy and quality of generated explanations.
Qian Yang 0007, Yunxin Li, Baotian Hu, Lin Ma 0002, Min Zhang 0005
ACM Multimedia1