VLDB 2026 Research / reviewers in the wild / expert
Hongyi Cai
dblp:31/9823
· DBLP profile ↗
11ranked-venue papers
4as first author
10since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 6 · 2 first-author · 6 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | When Vision Meets Texts in Listwise RerankingabstractRecent advancements in information retrieval have highlighted the potential of integrating visual and textual information, yet effective reranking for image-text documents remains challenging due to the modality gap and scarcity of aligned datasets. Meanwhile, existing approaches often rely on large models (7B--32B parameters) with reasoning-based distillation, incurring unnecessary computational overhead while primarily focusing on textual modalities. In this paper, we propose Rank-Nexus, a multimodal image-text document reranker that performs listwise qualitative reranking on retrieved lists incorporating both images and texts. To bridge the modality gap, we introduce a progressive cross-modal training strategy. We first train each modality separately: leveraging abundant text reranking data, we distill ranking knowledge into the text branch using GPT-4o as the teacher model. For images, where labeled data is scarce, we construct distilled pairs from multimodal large language model (MLLM) captions on image retrieval benchmarks. Subsequently, we train on a joint image-text reranking dataset. Rank-Nexus achieves outstanding performance on text reranking benchmarks (TREC, BEIR) and the challenging image reranking benchmarks (INQUIRE, MMDocIR), using only a lightweight 2B pretrained visual-language model. This efficient design ensures strong generalization across diverse multimodal scenarios without excessive parameters or reasoning overhead. Hongyi Cai |
SIGIR | 1 |
| 2026 | Zero-shot Chinese Character Recognition with Radical Tree Positional Embedding and spatial alignment
Anna Zhu, Hongyi Cai |
Eng. Appl. Artif. Intell. | 2 |
| 2026 | Cross-modal local and global alignment for Chinese Character Recognition
Anna Zhu, Hongyi Cai |
Pattern Recognit. | 3 |
| 2025 | AgileIR: Memory-Efficient Group Shifted Windows Attention for Lightweight Image Restoration
Hongyi Cai, Mohammad Mahdinur Rahman, Jingyu Wu, Wenzhen Dong |
ICANN (2) | 1 |
| 2025 | How Effective is In-Context Learning with Large Language Models for Rare Cell Identification in Single-Cell Expression Data?abstractThe recent development of single-cell genomics requires more powerful computational tools to differentiate between different phenotypes, and rare cell identification has been one of the most important problems in this area. Recent data-driven approaches usually adopt feature selection techniques to identify important genes for anomaly detection, which require extensive training data or domain knowledge from experts. In comparison, large language models (LLMs) have achieved certain progress in scientific research with strong generalization ability, which has shown potential in this area. In this paper, we make an attempt to comprehensively evaluate the performance of in-context learning with LLMs in rare cell identification. In particular, we carefully design a chain-of-thought prompt combining token probability analysis and cross-query comparison to generate scores to identify rare cells. From the experimental results on benchmark datasets, we find that LLMs are competitive compared to existing training-based methods, which demonstrates extensive potential for rare cell identification. Our source code and data are available at https://github.com/Hyan-Yao/InContextCells/. Huaiyuan Yao, Zhenxiao Cao, Zhongman Wang, Xiewei Ni, Jinyan Dong, Hongyi Cai, Yuqiang Han, Xiao Luo 0001 |
ICDM | 6 |
| 2025 | A Vision for Access Control in LLM Agent Systems
Hongyi Cai, Xinfeng Li, Yijia Xu |
ICECCS | 3 |
| 2025 | Agent Behavior: The Regulatory Object of the Agent-Centric Online Ecosystem in Digital Age
Qiang Zhang 0057, Pei Yan, Yijia Xu, Xinfeng Li, Hongyi Cai, Chuanpo Fu, Yong Fang 0002, Yang Liu 0003 |
ICECCS | 5 |
| 2025 | To Trust or Not to Trust? Enhancing Large Language Models' Situated Faithfulness to External ContextsabstractLarge Language Models (LLMs) are often augmented with external contexts, such as those used in retrieval-augmented generation (RAG). However, these contexts can be inaccurate or intentionally misleading, leading to conflicts with the model’s internal knowledge. We argue that robust LLMs should demonstrate situated faithfulness, dynamically calibrating their trust in external information based on their confidence in the internal knowledge and the external context to resolve knowledge conflicts. To benchmark this capability, we evaluate LLMs across several QA datasets, including a newly created dataset featuring in-the-wild incorrect contexts sourced from Reddit posts. We show that when provided with both correct and incorrect contexts, both open-source and proprietary models tend to overly rely on external information, regardless of its factual accuracy. To enhance situated faithfulness, we propose two approaches: Self-Guided Confidence Reasoning (SCR) and Rule-Based Confidence Reasoning (RCR). SCR enables models to self-access the confidence of external information relative to their own internal knowledge to produce the most accurate answer. RCR, in contrast, extracts explicit confidence signals from the LLM and determines the final answer using predefined rules.
Our results show that for LLMs with strong reasoning capabilities, such as GPT-4o and GPT-4o mini, SCR outperforms RCR, achieving improvements of up to 24.2\% over a direct input augmentation baseline. Conversely, for a smaller model like Llama-3-8B, RCR outperforms SCR. Fine-tuning SCR with our proposed Confidence Reasoning Direct Preference Optimization (CR-DPO) method improves performance on both seen and unseen datasets, yielding an average improvement of 8.9\% on Llama-3-8B. In addition to quantitative results, we offer insights into the relative strengths of SCR and RCR. Our findings highlight promising avenues for improving situated faithfulness in LLMs. Sanxing Chen, Hongyi Cai, Bhuwan Dhingra |
ICLR | 3 |
| 2025 | CFPFormer: Cross Feature-Pyramid Transformer Decoder for Medical Image SegmentationabstractFeature pyramids have been widely adopted in convolutional neural networks and transformers for tasks in medical image segmentation. However, existing models generally focus on the Encoder-side Transformer for feature extraction. We further explore the potential in improving the feature decoder with a well-designed architecture. We propose Cross Feature Pyramid Transformer decoder (CFPFormer), a novel decoder block that integrates feature pyramids and transformers. Even though transformer-like architecture impress with outstanding performance in segmentation, the concerns to reduce the redundancy and training costs still exist. Specifically, by leveraging patch embedding, cross-layer feature concatenation mechanisms, CFPFormer enhances feature extraction capabilities while complexity issue is mitigated by our Gaussian Attention. Benefiting from Transformer structure and U-shaped connections, our work is capable of capturing long-range dependencies and effectively up-sample feature maps. Experimental results are provided to evaluate CFPFormer on medical image segmentation datasets, demonstrating the efficacy and effectiveness. With a ResNet50 backbone, our method achieves 92.02% Dice Score, highlighting the efficacy of our methods. Notably, our VGG-based model outperformed baselines with more complex ViT and Swin Transformer backbone. Hongyi Cai, Mohammad Mahdinur Rahman, Wenzhen Dong, Jingyu Wu |
IJCNN | 1 |
| 2024 | Cross-Modal Alignment of Local and Global Features for Zero-Shot Chinese Character RecognitionabstractChinese character recognition (CCR) is a pivotal domain in computer vision due to its complexity and diverse applications, especially given the extensive character categories posing challenges in identifying unseen characters. Addressing the zero-shot hurdle, we propose a CLIP-style model, which independently extracts features from aligned Chinese character images and Ideographic Description Sequences (IDS), achieving cross-modal alignment. Our approach encompasses local and global feature alignment. Initially, we introduce learnable discrete tokens to represent shared embeddings for visual and textual modalities, capturing the local context of Chinese characters. Then, encoding each radical extracts local features, mapped to shared discrete tokens via attention mechanisms. Additionally, encoding the entire character obtains global features. Training utilizes contrastive loss to facilitate cross-modal alignment. Experimental results confirm our method’s superiority over conventional approaches, demonstrating remarkable performance on zero-shot Chinese character recognition benchmarks. Hongyi Cai, Anna Zhu |
ICIP | 1 |
| 2020 | Improving High Dynamic Range Image Based Light MeasurementabstractThis study proposes a fast high dynamic range imaging (HDRI) technique for light measurement to shorten the long capturing time of current camera-aided computational photography widely used in lighting practice. In comparison with the conventional meter measurement, HDRI-assisted lighting measurement is a remote, efficient, affordable yet time-consuming method. The fast HDRI technique increases the film speed (ISO) to speed up the process taking a sequence of low dynamic range images. Since increasing camera's film speed may introduce more image noise, the possible error rate of the proposed method is evaluated by applying Gaussian noise estimation and impulsive noise detection on the image with different film speeds. In addition, a new per-pixel calculation process is developed to retrieve the illuminance of a target scene with selected regions of interest, which can be used to assist human-centric lighting tasks. Extensive comparative experiments are also conducted to verify the accuracy and efficiency of the proposed method. Hankun Li, Hongyi Cai, Guanghui Wang 0001 |
SMC | 2 |