EDBT 2026 Demo / reviewers in the wild / expert
Yuecong Min
dblp:263/3327
· DBLP profile ↗
12ranked-venue papers
4as first author
10since 2021 · last 2026
0000-0002-0696-2468ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 10 · 4 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 10 · 4 first-author · 8 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | INFACT: A Diagnostic Benchmark for Induced Faithfulness and Factuality Hallucinations in Video-LLMsabstractDespite rapid progress, Video Large Language Models (Video-LLMs) remain unreliable due to hallucinations, which are outputs that contradict either video evidence (faithfulness) or verifiable world knowledge (factuality).Existing benchmarks provide limited coverage of factuality hallucinations and predominantly evaluate models only in clean settings.We introduce INFACT, a diagnostic benchmark comprising 9,800 QA instances with fine-grained taxonomies for faithfulness and factuality, spanning real and synthetic videos.INFACT evaluates models in four modes: Base (clean), Visual Degradation, Evidence Corruption, and Temporal Intervention for order-sensitive items.Reliability under induced modes is quantified using Resist Rate (RR) and Temporal Sensitivity Score (TSS).Experiments on 14 representative Video-LLMs reveal that higher Base-mode accuracy does not reliably translate to higher reliability in the induced modes, with evidence corruption reducing stability and temporal intervention yielding the largest degradation.Notably, many open-source baselines exhibit nearzero TSS on factuality, indicating pronounced temporal inertia on order-sensitive questions. Junqi Yang, Yuecong Min, Jie Zhang 0071, Shiguang Shan, Xilin Chen 0001 |
ACL (1) | 2 |
| 2026 | A Survey of Multimodal Hallucination Evaluation and Detection
Yuecong Min, Jie Zhang 0071, Bei Yan, Shiguang Shan |
Int. J. Comput. Vis. | 2 |
| 2025 | SHALE: A Scalable Benchmark for Fine-grained Hallucination Evaluation in LVLMsabstractDespite rapid advances, Large Vision-Language Models (LVLMs) still suffer from hallucinations, i.e., generating content inconsistent with input or established world knowledge, which correspond to faithfulness and factuality hallucinations, respectively. Prior studies primarily evaluate faithfulness hallucination at a rather coarse level (e.g., object-level) and lack fine-grained analysis. Additionally, existing benchmarks often rely on costly manual curation or reused public datasets, raising concerns about scalability and data leakage. To address these limitations, we propose an automated data construction pipeline that produces scalable, controllable, and diverse evaluation data. We also design a hierarchical hallucination induction framework with input perturbations to simulate realistic noisy scenarios. Integrating these designs, we construct SHALE, a Scalable HALlucination Evaluation benchmark designed to assess both faithfulness and factuality hallucinations via a fine-grained hallucination categorization scheme. SHALE comprises over 30K image-instruction pairs spanning 12 representative visual perception aspects for faithfulness and 6 knowledge domains for factuality, considering both clean and noisy scenarios. Extensive experiments on over 20 mainstream LVLMs reveal significant factuality hallucinations and high sensitivity to semantic perturbations. Bei Yan, Yuecong Min, Jie Zhang 0071, Shiguang Shan |
ACM Multimedia | 3 |
| 2024 | Visual Alignment Pre-training for Sign Language Translation
Peiqi Jiao, Yuecong Min, Xilin Chen 0001 |
ECCV (42) | 2 |
| 2023 | CoSign: Exploring Co-occurrence Signals in Skeleton-based Continuous Sign Language RecognitionabstractThe co-occurrence signals (e.g., hand shape, facial expression, and lip pattern) play a critical role in Continuous Sign Language Recognition (CSLR). Compared to RGB data, skeleton data provide a more efficient and concise option, and lay a good foundation for the co-occurrence exploration in CSLR. However, skeleton data are often used as a tool to assist visual grounding and have not attracted sufficient attention. In this paper, we propose a simple yet effective GCN-based approach, named CoSign, to incorporate Co-occurrence Signals and explore the potential of skeleton data in CSLR. Specifically, we propose a group-specific GCN to better exploit the knowledge of each signal and a complementary regularization to prevent complex co-adaptation across signals. Furthermore, we propose a two-stream framework that gradually fuses both static and dynamic information in skeleton data. Experimental results on three public CSLR datasets (PHOENIX14, PHOENIX14-T and CSL-Daily) show that the proposed CoSign achieves competitive performance with recent video-based approaches while reducing the computation cost during training. Peiqi Jiao, Yuecong Min, Xiaotao Wang, Xilin Chen 0001 |
ICCV | 2 |
| 2022 | S2Net: Skeleton-Aware SlowFast Network for Efficient Sign Language Recognition
Yuecong Min, Xilin Chen 0001 |
ACCV (4) | 2 |
| 2022 | Deep Radial Embedding for Visual Sequence Learning
Yuecong Min, Peiqi Jiao, Xiaotao Wang, Xiujuan Chai, Xilin Chen 0001 |
ECCV (6) | 1 |
| 2021 | Self-Mutual Distillation Learning for Continuous Sign Language RecognitionabstractIn recent years, deep learning moves video-based Continuous Sign Language Recognition (CSLR) significantly forward. Currently, a typical network combination for CSLR includes a visual module, which focuses on spatial and short-temporal information, followed by a contextual module, which focuses on long-temporal information, and the Connectionist Temporal Classification (CTC) loss is adopted to train the network. However, due to the limitation of chain rules in back-propagation, the visual module is hard to adjust for seeking optimized visual features. As a result, it enforces that the contextual module focuses on contextual information optimization only rather than balancing efficient visual and contextual information. In this paper, we propose a Self-Mutual Knowledge Distillation (SMKD) method, which enforces the visual and contextual modules to focus on short-term and long-term information and enhances the discriminative power of both modules simultaneously. Specifically, the visual and contextual modules share the weights of their corresponding classifiers, and train with CTC loss simultaneously. Moreover, the spike phenomenon widely exists with CTC loss. Although it can help us choose a few of the key frames of a gloss, it does drop other frames in a gloss and makes the visual feature saturation in the early stage. A gloss segmentation is developed to relieve the spike phenomenon and decrease saturation in the visual module. We conduct experiments on two CSLR bench-marks: PHOENIX14 and PHOENIX14-T. Experimental results demonstrate the effectiveness of the SMKD. Aiming Hao, Yuecong Min, Xilin Chen 0001 |
ICCV | 2 |
| 2021 | Visual Alignment Constraint for Continuous Sign Language RecognitionabstractVision-based Continuous Sign Language Recognition (CSLR) aims to recognize unsegmented signs from image streams. Overfitting is one of the most critical problems in CSLR training, and previous works show that the iterative training scheme can partially solve this problem while also costing more training time. In this study, we revisit the iterative training scheme in recent CSLR works and realize that sufficient training of the feature extractor is critical to solving the overfitting problem. Therefore, we propose a Visual Alignment Constraint (VAC) to enhance the feature extractor with alignment supervision. Specifically, the proposed VAC comprises two auxiliary losses: one focuses on visual features only, and the other enforces prediction alignment between the feature extractor and the alignment module. Moreover, we propose two metrics to reflect overfitting by measuring the prediction inconsistency between the feature extractor and the alignment module. Experimental results on two challenging CSLR datasets show that the proposed VAC makes CSLR networks end-to-end trainable and achieves competitive performance. Yuecong Min, Aiming Hao, Xiujuan Chai, Xilin Chen 0001 |
ICCV | 1 |
| 2021 | Teaching Chinese Sign Language with a SmartphoneabstractThere is a large group of deaf-mutes in the world, and sign language is their major communication tool. Therefore, it’s necessary for the deaf-mutes to communicate with hearing-speech people and the hearing-speech people also have needed to understand sign language, which produces a great demand for sign language teaching. Even though there have already been a large number of books for sign language, it is low efficient to learn sign language with books, even teaching videos. To solve this problem, we develop a smartphone-based interactive Chinese sign language teaching systemfor sign language learning. The system provides a learner with some kinds of learning modes and captures the learner’s actions from its front camera of the smartphone. Right now the system provides a vocabulary set with 1000 frequently used words, and the learner can evaluate his/her sign action by subjective or objective comparison. In the mode of word recognition, the users can play any word within the vocabulary and the system will turn the top three retrieved candidates, so it can remind the learners what the sign is. This system provides interactive learning for a user to learn sign language high efficiently. The systemadopts an algorithm based on point cloud recognition to evaluate a user’s sign and costs about 700ms inference time for each sample, which meets the real-time requirements. This interactive learning system decreases the communication barriers between the deaf-mutes and hearing-speechers. Yanxiao Zhang, Yuecong Min, Xilin Chen 0001 |
Virtual Real. Intell. Hardw. | 2 |
| 2020 | An Efficient PointLSTM for Point Clouds Based Gesture RecognitionabstractPoint clouds contain rich spatial information, which provides complementary cues for gesture recognition. In this paper, we formulate gesture recognition as an irregular sequence recognition problem and aim to capture long-term spatial correlations across point cloud sequences. A novel and effective PointLSTM is proposed to propagate information from past to future while preserving the spatial structure. The proposed PointLSTM combines state information from neighboring points in the past with current features to update the current states by a weight-shared LSTM layer. This method can be integrated into many other sequence learning approaches. In the task of gesture recognition, the proposed PointLSTM achieves state-of-the-art results on two challenging datasets (NVGesture and SHREC'17) and outperforms previous skeleton-based methods. To show its advantages in generalization, we evaluate our method on MSR Action3D dataset, and it produces competitive results with previous skeleton-based methods. Yuecong Min, Yanxiao Zhang, Xiujuan Chai, Xilin Chen 0001 |
CVPR | 1 |
| 2019 | FlickerNet: Adaptive 3D Gesture Recognition from Sparse Point Clouds
Yuecong Min, Xiujuan Chai, Xilin Chen 0001 |
BMVC | 1 |