VLDB 2026 Research / reviewers in the wild / expert
Yuxiao Chen 0002
dblp:158/4934-2
· DBLP profile ↗
13ranked-venue papers
6as first author
8since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 9 · 6 first-author · 6 since 2021Artificial intelligence and machine learning · 8 · 4 first-author · 5 since 2021Databases, data management, data science and information retrieval · 2Applied, interdisciplinary, general and emerging computing · 2Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | The Hidden Life of Tokens: Reducing Hallucination of Large Vision-Language Models Via Visual Information SteeringabstractLarge Vision-Language Models (LVLMs) can reason effectively over both textual and visual inputs, but they tend to hallucinate syntactically coherent yet visually ungrounded contents. In this paper, we investigate the internal dynamics of hallucination by examining the tokens logits rankings throughout the generation process, revealing three key patterns in how LVLMs process information: (1) gradual visual information loss – visually grounded tokens gradually become less favored throughout generation, and (2) early excitation – semantically meaningful tokens achieve peak activation in the layers earlier than the final layer. (3) hidden genuine information – visually grounded tokens though not being eventually decided still retain relatively high rankings at inference. Based on these insights, we propose VISTA (Visual Information Steering with Token-logit Augmentation), a training-free inference-time intervention framework that reduces hallucination while promoting genuine information. VISTA works by combining two complementary approaches: reinforcing visual information in activation space and leveraging early layer activations to promote semantically meaningful decoding. Compared to existing methods, VISTA requires no external supervision and is applicable to various decoding strategies. Extensive experiments show that VISTA on average reduces hallucination by about 40% on evaluated open-ended generation task, and it consistently outperforms existing methods on four benchmarks across four architectures under three decoding strategies. Code is available at https://github.com/LzVv123456/VISTA. Zhuowei Li 0002, Haizhou Shi, Yunhe Gao, Di Liu 0003, Zhenting Wang, Yuxiao Chen 0002, Ting Liu 0005, Long Zhao 0003, Hao Wang 0014, Dimitris N. Metaxas |
ICML | 6 |
| 2025 | Exploiting VLM Localizability and Semantics for Open Vocabulary Action DetectionabstractAction detection aims to detect (recognize and localize) human actions spatially and temporally in videos. Existing approaches focus on the closed-set setting where an action detector is trained and tested on videos from a fixed set of action categories. However, this constrained setting is not viable in an open world where test videos inevitably come beyond the trained action categories. In this paper, we address the practical yet challenging Open-Vocabulary Action Detection (OVAD) problem. It aims to detect any action in test videos while training a model on a fixed set of action categories. To achieve such an open-vocabulary capability, we propose a novel method OpenMixer that exploits the inherent semantics and localizability of large vision-language models (VLM) within the family of query-based detection transformers (DETR). Specifically, the OpenMixer is developed by spatial and temporal OpertMixer blocks (S-OMB and T-OMB), and a dynamically fused alignment (DFA) module. The three components collectively enjoy the merits of strong generalization from pretrained VLMs and end-to-end learning from DETR design. Moreover, we established OVAD benchmarks under various settings, and the experimental results show that the OpenMixer performs the best over baselines for detecting seen and unseen actions. We release the codes, models, and dataset splits at https://github.com/Cogito2012/0penMixer. Wentao Bao, Kai Li 0012, Yuxiao Chen 0002, Deep Patel, Martin Renqiang Min, Yu Kong 0001 |
WACV | 3 |
| 2024 | Learning to Localize Actions in Instructional Videos with LLM-Based Multi-pathway Text-Video Alignment
Yuxiao Chen 0002, Kai Li 0012, Wentao Bao, Deep Patel, Yu Kong 0001, Martin Renqiang Min, Dimitris N. Metaxas |
ECCV (82) | 1 |
| 2024 | ProxEdit: Improving Tuning-Free Real Image Editing with Proximal GuidanceabstractDDIM inversion has revealed the remarkable potential of real image editing within diffusion-based methods. However, the accuracy of DDIM reconstruction degrades as larger classifier-free guidance (CFG) scales being used for enhanced editing. Null-text inversion (NTI) optimizes null embeddings to align the reconstruction and inversion trajectories with larger CFG scales, enabling real image editing with cross-attention control. Negative-prompt inversion (NPI) further offers a training-free closed-form solution of NTI. However, it may introduce artifacts and is still constrained by DDIM reconstruction quality. To overcome these limitations, we propose proximal guidance and incorporate it to NPI with cross-attention control. We enhance NPI with a regularization term and inversion guidance, which reduces artifacts while capitalizing on its training-free nature. Additionally, we extend the concepts to incorporate mutual self-attention control, enabling geometry and layout alterations in the editing process. Our method provides an efficient and straightforward approach, effectively addressing real image editing tasks with minimal computational overhead. Ligong Han, Song Wen 0001, Kunpeng Song, Mengwei Ren, Ruijiang Gao, Anastasis Stathopoulos, Xiaoxiao He, Yuxiao Chen 0002, Di Liu 0003, Qilong Zhangli, Jindong Jiang, Zhaoyang Xia, Akash Srivastava, Dimitris N. Metaxas |
WACV | 10 |
| 2023 | Revisiting Multimodal Representation in Contrastive Learning: From Patch and Token Embeddings to Finite Discrete TokensabstractContrastive learning-based vision-language pretraining approaches, such as CLIP, have demonstrated great success in many vision-language tasks. These methods achieve cross-modal alignment by encoding a matched image-text pair with similar feature embeddings, which are generated by aggregating information from visual patches and language tokens. However, direct aligning cross-modal information using such representations is challenging, as visual patches and text tokens differ in semantic levels and granularities. To alleviate this issue, we propose a Finite Discrete Tokens (FDT) based multimodal representation. FDT is a set of learnable tokens representing certain visualsemantic concepts. Both images and texts are embedded using shared FDT by first grounding multimodal inputs to FDT space and then aggregating the activated FDT representations. The matched visual and semantic concepts are enforced to be represented by the same set of discrete tokens by a sparse activation constraint. As a result, the granularity gap between the two modalities is reduced. Through both quantitative and qualitative analyses, we demonstrate that using FDT representations in CLIP-style models improves cross-modal alignment and performance in visual recognition and vision-language downstream tasks. Furthermore, we show that our method can learn more comprehensive representations, and the learned FDT capture meaningful cross-modal correspondence, ranging from objects to actions and attributes.11The source code can be found at https://github.com/yuxiaochen1103/FDT. Yuxiao Chen 0002, Yu Tian 0003, Shijie Geng, Dimitris N. Metaxas, Hongxia Yang |
CVPR | 1 |
| 2023 | HiCLIP: Contrastive Language-Image Pretraining with Hierarchy-aware Attention
Shijie Geng, Yu Tian 0003, Yuxiao Chen 0002 |
ICLR | 4 |
| 2023 | More Than Just Attention: Improving Cross-Modal Attentions with Contrastive Constraints for Image-Text MatchingabstractCross-modal attention mechanisms have been widely applied to the image-text matching task. They have achieved remarkable improvements thanks to their capability of learning fine-grained relevance across different modalities. However, the cross-modal attention models of existing methods could be sub-optimal and inaccurate because there is no direct supervision provided during the training process. In this work, we propose two novel training strategies, namely Contrastive Content Resourcing (CCR) and Contrastive Content Swapping (CCS) constraints, to address such limitations. These constraints supervise the training of cross-modal attention models in a contrastive learning manner without requiring explicit attention annotations. They are plug-in training strategies and can be generally integrated into existing cross-modal attention models. Additionally, we introduce three metrics, including Attention Precision, Recall, and F1-Score, to quantitatively measure the quality of learned attention models. We evaluate the proposed constraints by incorporating them into four state- of-the-art cross-modal attention-based image-text matching models. Experimental results on both Flickr30k and MS-COCO datasets demonstrate that integrating these constraints generally improves the model performance in terms of both retrieval performance and attention metrics. Yuxiao Chen 0002, Long Zhao 0003, Larry Davis 0001, Dimitris N. Metaxas |
WACV | 1 |
| 2022 | Hierarchically Self-supervised Transformer for Human Skeleton Representation Learning
Yuxiao Chen 0002, Long Zhao 0003, Yu Tian 0003, Zhaoyang Xia, Shijie Geng, Ligong Han, Dimitris N. Metaxas |
ECCV (26) | 1 |
| 2020 | Knowledge As Priors: Cross-Modal Knowledge Generalization for Datasets Without Superior KnowledgeabstractCross-modal knowledge distillation deals with transferring knowledge from a model trained with superior modalities (Teacher) to another model trained with weak modalities (Student). Existing approaches require paired training examples exist in both modalities. However, accessing the data from superior modalities may not always be feasible. For example, in the case of 3D hand pose estimation, depth maps, point clouds, or stereo images usually capture better hand structures than RGB images, but most of them are expensive to be collected. In this paper, we propose a novel scheme to train the Student in a Target dataset where the Teacher is unavailable. Our key idea is to generalize the distilled cross-modal knowledge learned from a Source dataset, which contains paired examples from both modalities, to the Target dataset by modeling knowledge as priors on parameters of the Student. We name our method "Cross-Modal Knowledge Generalization" and demonstrate that our scheme results in competitive performance for 3D hand pose estimation on standard benchmark datasets. Long Zhao 0003, Xi Peng 0005, Yuxiao Chen 0002, Mubbasir Kapadia, Dimitris N. Metaxas |
CVPR | 3 |
| 2019 | Construct Dynamic Graphs for Hand Gesture Recognition via Spatial-Temporal Attention
Yuxiao Chen 0002, Long Zhao 0003, Xi Peng 0005, Dimitris N. Metaxas |
BMVC | 1 |
| 2018 | You Type a Few Words and We Do the Rest: Image Recommendation for Social Multimedia PostsabstractIn this paper, we introduce a new application that can be employed on many social media platforms. We intend to recommend related images from local (e.g. user's local mobile phone storage) and global (e.g. platform's server) image pools while a user is composing the text to post a status. To make the recommendation system applicable to different platforms with or without the support of online computing, we propose two independent frameworks that recommend images at image level and data-driven category level based on a text, respectively. For image-level recommendation, our framework recommends images for a text by predicting the affinity scores of image-text pairs and recommending the images with the highest scores for the text. To improve the ranking performance, we propose a novel patch-level image-text matching framework which strikes a balance between local and global matching of image-text pairs. In particular, it first extracts the affinity between each local word and image patch, then leverages different kinds of attention mechanisms to respectively weight the local words and patches for computing the final image-text affinity scores. For category-level recommendation, we first classify images into categories in an unsupervised way, and then propose a multi-task LSTM-based framework with effective user feature for recommendation. Extensive experiments in two real-world social media datasets demonstrate the effectiveness of the proposed models, which significantly outperform the baselines. We also visualize the patch-word matching details to provide an insight into the image-level recommendation framework and demonstrate the strong capacity of category-level recommendation framework to recommend images in an online fashion. Yuxiao Chen 0002, Jiebo Luo 0001 |
IEEE BigData | 2 |
| 2018 | Mining the Relationship between Emoji Usage Patterns and Personality
Weijian Li 0001, Yuxiao Chen 0002, Tianran Hu, Jiebo Luo 0001 |
ICWSM | 2 |
| 2018 | Twitter Sentiment Analysis via Bi-sense Emoji Embedding and Attention-based LSTMabstractSentiment analysis on large-scale social media data is important to bridge the gaps between social media contents and real world activities including political election prediction, individual and public emotional status monitoring and analysis, and so on. Although textual sentiment analysis has been well studied based on platforms such as Twitter and Instagram, analysis of the role of extensive emoji uses in sentiment analysis remains light. In this paper, we propose a novel scheme for Twitter sentiment analysis with extra attention on emojis. We first learn bi-sense emoji embeddings under positive and negative sentimental tweets individually, and then train a sentiment classifier by attending on these bi-sense emoji embeddings with an attention-based long short-term memory network (LSTM). Our experiments show that the bi-sense embedding is effective for extracting sentiment-aware embeddings of emojis and outperforms the state-of-the-art models. We also visualize the attentions to show that the bi-sense emoji embedding provides better guidance on the attention mechanism to obtain a more robust understanding of the semantics and sentiments. Yuxiao Chen 0002, Quanzeng You, Jiebo Luo 0001 |
ACM Multimedia | 1 |