VLDB 2026 Research / reviewers in the wild / expert
Wei-Yuan Cheng
dblp:95/8727
· DBLP profile ↗
7ranked-venue papers
3as first author
3since 2021 · last 2026
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 5 · 2 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
1 paper |
Trustworthy machine learning · 75% Vision and language · 25% |
Topics — the 4 heaviest of 4, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Trustworthy machine learning › interpretability › explanation evaluation
explanation faithfulness |
0.8 | 1 | 2024 | RAPPER: Reinforced Rationale-Prompted Paradigm for Natural Language Explanation in Visual Question Answering · ICLR 2024 |
Machine learning › Trustworthy machine learning
interpretability |
0.8 | 1 | 2024 | RAPPER: Reinforced Rationale-Prompted Paradigm for Natural Language Explanation in Visual Question Answering · ICLR 2024 |
Machine learning › Trustworthy machine learning › interpretability
natural language explanation |
0.8 | 1 | 2024 | RAPPER: Reinforced Rationale-Prompted Paradigm for Natural Language Explanation in Visual Question Answering · ICLR 2024 |
Computer vision › Vision and language
visual question answering |
0.8 | 1 | 2024 | RAPPER: Reinforced Rationale-Prompted Paradigm for Natural Language Explanation in Visual Question Answering · ICLR 2024 |
Methods — techniques the papers use, named apart from their topics
reinforcement learning from feedback · 0.8large language model · 0.8knowledge distillation · 0.8
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Mitigating Object and Action Hallucinations in Multimodal LLMs via Self-Augmented Contrastive AlignmentabstractRecent advancement in multimodal LLMs (MLLMs) has demonstrated their remarkable capability to generate descriptive captions for input videos. However, these models suffer from factual inaccuracies in the generated descriptions, causing severe hallucination issues. While prior works have explored alleviating hallucinations for static images, jointly mitigating visual object and temporal action hallucinations for dynamic videos remains a challenging and unsolved task. To tackle this challenge, we propose a Self-Augmented Contrastive Alignment (SANTA) framework for enabling object and action faithfulness by exempting the spurious correlations and enforcing the emphasis on visual facts. SANTA employs a hallucinative self-augmentation scheme to identify the potential hallucinations that lie in the MLLM and transform the original captions to the contrasted negatives. Furthermore, we develop a tracklet-phrase contrastive alignment to match the regional objects and relation-guided actions with their corresponding visual and temporal phrases. Extensive experiments demonstrate that SANTA outperforms existing methods in alleviating object and action hallucinations, yielding superior performance on the hallucination examination benchmarks. Kai-Po Chang, Wei-Yuan Cheng, Chi-Pin Huang, Fu-En Yang, Yu-Chiang Frank Wang |
WACV | 2 |
| 2026 | TA-Prompting: Enhancing Video Large Language Models for Dense Video Captioning via Temporal AnchorsabstractDense video captioning aims to interpret and describe all temporally localized events throughout an input video. Recent state-of-the-art methods leverage large language models (LLMs) to provide detailed moment descriptions for video data. However, existing VideoLLMs remain challenging in identifying precise event boundaries in untrimmed videos, causing the generated captions to be not properly grounded. In this paper, we propose TA-Prompting, which enhances VideoLLMs via Temporal Anchors that learn to precisely localize events and prompt the VideoLLMs to perform temporal-aware video event understanding. During inference, in order to properly determine the output caption sequence from an arbitrary number of events presented within a video, we introduce an event coherent sampling strategy to select event captions with sufficient coherence across temporal events and cross-modal similarity with the given video. Through extensive experiments on benchmark datasets, we show that our TA-Prompting is favorable against state-of-the-art VideoLLMs, yielding superior performance on dense video captioning and temporal understanding tasks including moment retrieval and temporalQA. Wei-Yuan Cheng, Kai-Po Chang, Chi-Pin Huang, Fu-En Yang, Yu-Chiang Frank Wang |
WACV | 1 |
| 2024 | RAPPER: Reinforced Rationale-Prompted Paradigm for Natural Language Explanation in Visual Question AnsweringabstractNatural Language Explanation (NLE) in vision and language tasks aims to provide human-understandable explanations for the associated decision-making process. In practice, one might encounter explanations which lack informativeness or contradict visual-grounded facts, known as implausibility and hallucination problems, respectively. To tackle these challenging issues, we consider the task of visual question answering (VQA) and introduce Rapper, a two-stage Reinforced Rationale-Prompted Paradigm. By knowledge distillation, the former stage of Rapper infuses rationale-prompting via large language models (LLMs), encouraging the rationales supported by language-based facts. As for the latter stage, a unique Reinforcement Learning from NLE Feedback (RLNF) is introduced for injecting visual facts into NLE generation. Finally, quantitative and qualitative experiments on two VL-NLE benchmarks show that Rapper surpasses state-of-the-art VQA-NLE methods while providing plausible and faithful NLE. Kai-Po Chang, Chi-Pin Huang, Wei-Yuan Cheng, Fu-En Yang, Chien-Yi Wang, Yung-Hsuan Lai, Yu-Chiang Frank Wang |
ICLR | 3 |
| 2014 | A Fuzzy Model With Online Incremental SVM and Margin-Selective Gradient Descent Learning for Classification ProblemsabstractThis paper proposes a new incremental learning approach to endow a Takagi-Sugeno-type fuzzy classification model with high generalization ability. The proposed fuzzy model is learned through incremental support vector machine (SVM) and margin-selected gradient descent learning and is called FM3. In this learning approach, training samples are fed incrementally one-by-one instead of all in one batch. The FM3evolves from an empty rule set. A one-pass clustering algorithm is used to determine the number of rules and initial fuzzy sets in the rule antecedent part. An online incremental linear SVM is proposed to tune the rule consequent parameters to endow the FM3with high generalization ability. The use of incremental instead of batch SVM enables the FM3to handle online training problems with only one new sample available at a time. For antecedent parameter learning, a margin-selected gradient descent algorithm is proposed to prevent overtraining. Simulation results and comparisons with SVMs and fuzzy classifiers with different learning algorithms demonstrate the advantage of the FM3. Wei-Yuan Cheng, Chia-Feng Juang |
IEEE Trans. Fuzzy Syst. | 1 |
| 2011 | An incremental support vector machine-trained TS-type fuzzy system for online classification problems
Wei-Yuan Cheng, Chia-Feng Juang |
Fuzzy Sets Syst. | 1 |
| 2011 | Speedup of Implementing Fuzzy Neural Networks With High-Dimensional Inputs Through Parallel Processing on Graphic Processing UnitsabstractThis paper proposes the implementation of a zero-order Takagi-Sugeno-Kang (TSK)-type fuzzy neural network (FNN) on graphic processing units (GPUs) to reduce training time. The software platform that this study uses is the compute unified device architecture (CUDA). The implemented FNN uses structure and parameter learning in a self-constructing neural fuzzy inference network because of its admirable learning performance. FNN training is conventionally implemented on a single-threaded CPU, where each input variable and fuzzy rule is serially processed. This type of training is time consuming, especially for a high-dimensional FNN that consists of a large number of rules. The GPU is capable of running a large number of threads in parallel. In a GPU-implemented FNN (GPU-FNN), blocks of threads are partitioned according to parallel and independent properties of fuzzy rules. Large sets of input data are mapped to parallel threads in each block. For memory management, this research suitably divides the datasets in the GPU-FNN into smaller chunks according to fuzzy rule structures to share on-chip memory among multiple thread processors. This study applies the GPU-FNN to different problems to verify its efficiency. The results show that to train an FNN with GPU implementation achieves a speedup of more than 30 times that of CPU implementation for problems with high-dimensional attributes. Chia-Feng Juang, Teng-Chang Chen, Wei-Yuan Cheng |
IEEE Trans. Fuzzy Syst. | 3 |
| 2010 | An Interval Type-2 Fuzzy-Neural Network With Support-Vector Regression for Noisy Regression ProblemsabstractThis paper proposes an interval type-2 fuzzy-neural network with support-vector regression (IT2FNN-SVR) for noisy regression problems. The antecedent part in each fuzzy rule of an IT2FNN-SVR uses interval type-2 fuzzy sets, and the consequent part is of the Takagi-Sugeno-Kang (TSK) type. The use of interval type-2 fuzzy sets helps improve the network's noise resistance. The network inputs may be numerical values or type-1 fuzzy sets, with the latter being used for further improvements in robustness. IT2FNN-SVR learning consists of both structure learning and parameter learning. The structure-learning algorithm is responsible for online rule generation. The parameters are optimized for structural-risk minimization using a two-phase linear SVR algorithm in order to endow the network with high generalization ability. IT2FNN-SVR performance is verified through comparisons with type-1 and type-2 fuzzy-logic systems and other regression models on noisy regression problems. Chia-Feng Juang, Ren-Bo Huang, Wei-Yuan Cheng |
IEEE Trans. Fuzzy Syst. | 3 |