EDBT 2026 Demo / reviewers in the wild / expert
Delong Chen
dblp:267/1326
· DBLP profile ↗
20ranked-venue papers
3as first author
20since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 13 · 3 first-author · 13 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 1 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | REVQA: Resource-Efficient MLLM Video Question Answering via Redundant Frame Elimination
Junjie Zhang 0010, Shuxia Wu, Delong Chen, Zhengxin Yu, Zheyi Chen |
ICC | 3 |
| 2026 | Information entropy based evolutionary multitasking optimization
Shuijia Li, Rui Wang 0017, Wenyin Gong, Yanchi Li, Delong Chen, Zuowen Liao |
Expert Syst. Appl. | 5 |
| 2025 | Making Large Vision Language Models to Be Good Few-Shot LearnersabstractFew-shot classification (FSC) is a fundamental yet challenging task in computer vision that involves recognizing novel classes from limited data. While previous methods have focused on enhancing visual features or incorporating additional modalities, Large Vision Language Models (LVLMs) offer a promising alternative due to their rich knowledge and strong visual perception. However, LVLMs risk learning specific response formats rather than effectively extracting useful information from support data in FSC. In this paper, we investigate LVLMs' performance in FSC and identify key issues such as insufficient learning and the presence of severe position biases. To tackle above challenges, we adopt the meta-learning strategy to teach models ``learn to learn". By constructing a rich set of meta-tasks for instruction fine-tuning, LVLMs enhance the ability to extract information from few-shot support data for classification. Additionally, we further boost LVLM's few-shot learning capabilities through label augmentation (LA) and candidate selection (CS) in the fine-tuning and inference stages, respectively. LA is implemented via a character perturbation strategy to ensure the model focuses on support information. CS leverages attribute descriptions to filter out unreliable candidates and simplify the task. Extensive experiments demonstrate that our approach achieves superior performance on both general and fine-grained datasets. Furthermore, our candidate selection strategy has been proven beneficial for training-free LVLMs. Fan Liu 0003, Wenwen Cai, Jian Huo, Chuanyi Zhang, Delong Chen |
AAAI | 5 |
| 2025 | Linguistic Minimal Pairs Elicit Linguistic Similarity in Large Language ModelsabstractWe introduce a novel analysis that leverages linguistic minimal pairs to probe the internal linguistic representations of Large Language Models (LLMs). By measuring the similarity between LLM activation differences across minimal pairs, we quantify the linguistic similarity and gain insight into the linguistic knowledge captured by LLMs. Our large-scale experiments, spanning 100+ LLMs and 150k minimal pairs in three languages, reveal properties of linguistic similarity from four key aspects: consistency across LLMs, relation to theoretical categorizations, dependency to semantic context, and cross-lingual alignment of relevant phenomena. Our findings suggest that 1) linguistic similarity is significantly influenced by training data exposure, leading to higher cross-LLM agreement in higher-resource languages. 2) Linguistic similarity strongly aligns with fine-grained theoretical linguistic categories but weakly with broader ones. 3) Linguistic similarity shows a weak correlation with semantic similarity, showing its context-dependent nature. 4) LLMs exhibit limited cross-lingual alignment in their understanding of relevant linguistic phenomena. This work demonstrates the potential of minimal pairs as a window into the neural representations of language in LLMs, shedding light on the relationship between LLMs and linguistic theory. Delong Chen, Samuel Cahyawijaya, Xufeng Duan, Zhenguang G. Cai |
COLING | 2 |
| 2025 | Chain-of-Talkers (CoTalk): Fast Human Annotation of Dense Image CaptionsabstractWhile densely annotated image captions significantly facilitate the learning of robust visionlanguage alignment, methodologies for systematically optimizing human annotation efforts remain underexplored.We introduce CHAIN-OF-TALKERS (COTALK), an AI-in-the-loop methodology designed to maximize the number of annotated samples and improve their comprehensiveness under fixed budget constraints (e.g., total human annotation time).The framework is built upon two key insights.First, sequential annotation reduces redundant workload compared to conventional parallel annotation, as subsequent annotators only need to annotate the "residual"-the missing visual information that previous annotations have not covered.Second, humans process textual input faster by reading while outputting annotations with much higher throughput via talking; thus a multimodal interface enables optimized efficiency.We evaluate our framework from two aspects: intrinsic evaluations that assess the comprehensiveness of semantic units, obtained by parsing detailed captions into object-attribute trees and analyzing their effective connections; extrinsic evaluation measures the practical usage of the annotated captions in facilitating vision-language alignment.Experiments with eight participants show our CHAIN-OF-TALKERS (CoTalk) improves annotation speed (0.42 vs. 0.30 units/sec) and retrieval performance (41.13% vs. 40.52%)over the parallel method.per minute? a review and meta-analysis of reading rate. Delong Chen, Fan Liu 0003, Chuanyi Zhang, Liang Yao 0001, Yuhui Zheng |
EMNLP | 2 |
| 2025 | Prompting DirectSAM for Semantic Contour Extraction in Remote Sensing ImagesabstractThe Direct Segment Anything Model (DirectSAM) excels in class-agnostic contour extraction. In this paper, we explore its use by applying it to optical remote sensing imagery, where semantic contour extraction—such as identifying buildings, road networks, and coastlines-holds significant practical value. Those applications are currently handled via training specialized small models separately on small datasets in each domain. We introduce a foundation model derived from DirectSAM, termed DirectSAM-RS, which not only inherits the strong segmentation capability acquired from natural images, but also benefits from a large-scale dataset we created for remote sensing semantic contour extraction. This dataset comprises over 34k image-text-contour triplets, making it at least 30 times larger than individual dataset. DirectSAM-RS integrates a prompter module: a text encoder and cross-attention layers attached to the DirectSAM architecture, which allows flexible conditioning on target class labels or referring expressions. We evaluate the DirectSAM-RS in both zero-shot and fine-tuning setting, and demonstrate that it achieves state-of-the-art performance across several downstream benchmarks. Shiyu Miao, Delong Chen, Fan Liu 0003, Chuanyi Zhang, Yanhui Gu, Shengjie Guo, Jun Zhou 0011 |
ICASSP | 2 |
| 2025 | Subobject-level Image TokenizationabstractPatch-based image tokenization ignores the morphology of the visual world, limiting effective and efficient learning of image understanding. Inspired by subword tokenization, we introduce subobject-level adaptive token segmentation and explore several approaches, including superpixel, SAM, and a proposed Efficient and PanOptiC (EPOC) image tokenizer. Our EPOC combines boundary detection–a simple task that can be handled well by a compact model–with watershed segmentation, which inherently guarantees no pixels are left unsegmented. Intrinsic evaluations across 5 datasets demonstrate that EPOC’s segmentation aligns well with human annotations of both object- and part-level visual morphology, producing more monosemantic tokens and offering substantial efficiency advantages. For extrinsic evaluation, we designed a token embedding that handles arbitrary-shaped tokens, and trained VLMs with different tokenizers on 4 datasets of object recognition and detailed captioning. The results reveal that subobject tokenization enables faster convergence and better generalization while using fewer visual tokens. Delong Chen, Samuel Cahyawijaya, Jianfeng Liu 0002, Baoyuan Wang, Pascale Fung |
ICML | 1 |
| 2025 | RemoteSAM: Towards Segment Anything for Earth ObservationabstractWe aim to develop a robust yet flexible visual foundation model for Earth observation. It should possess strong capabilities in recognizing and localizing diverse visual targets while providing compatibility with various input-output interfaces required across different task scenarios. Current systems cannot meet these requirements, as they typically utilize task-specific architecture trained on narrow data domains with limited semantic coverage. Our study addresses these limitations from two aspects: data and modeling. We first introduce an automatic data engine that enjoys significantly better scalability compared to previous human annotation or rule-based approaches. It has enabled us to create the largest dataset of its kind to date, comprising 270K image-text-mask triplets covering an unprecedented range of diverse semantic categories and attribute specifications. Based on this data foundation, we further propose a task unification paradigm that centers around referring expression segmentation. It effectively handles a wide range of vision-centric perception tasks, including classification, detection, segmentation, grounding, etc, using a single model without any task-specific heads. Combining these innovations on data and modeling, we present RemoteSAM, a foundation model that establishes new SoTA on several earth observation perception benchmarks, outperforming other foundation models such as Falcon, GeoChat, and LHRS-Bot with significantly higher efficiency. Models and data are publicly available at https://github.com/1e12Leon/RemoteSAM. Liang Yao 0001, Fan Liu 0003, Delong Chen, Chuanyi Zhang, Ziyun Chen 0004, Shimin Di, Yuhui Zheng |
ACM Multimedia | 3 |
| 2025 | High-Dimension Human Value Representation in Large Language ModelsabstractSamuel Cahyawijaya, Delong Chen, Yejin Bang, Leila Khalatbari, Bryan Wilie, Ziwei Ji, Etsuko Ishii, Pascale Fung. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Samuel Cahyawijaya, Delong Chen, Yejin Bang, Leila Khalatbari, Bryan Wilie, Ziwei Ji 0001, Etsuko Ishii, Pascale Fung |
NAACL (Long Papers) | 2 |
| 2025 | ProtoCLIP: Prototypical Contrastive Language Image PretrainingabstractContrastive language image pretraining (CLIP) has received widespread attention since its learned representations can be transferred well to various downstream tasks. During the training process of the CLIP model, the InfoNCE objective aligns positive image-text pairs and separates negative ones. We show an underlying representation grouping effect during this process: the InfoNCE objective indirectly groups semantically similar representations together via randomly emerged within-modal anchors. Based on this understanding, in this article, prototypical contrastive language image pretraining (ProtoCLIP) is introduced to enhance such grouping by boosting its efficiency and increasing its robustness against the modality gap. Specifically, ProtoCLIP sets up prototype-level discrimination between image and text spaces, which efficiently transfers higher level structural knowledge. Furthermore, prototypical back translation (PBT) is proposed to decouple representation grouping from representation alignment, resulting in effective learning of meaningful representations under a large modality gap. The PBT also enables us to introduce additional external teachers with richer prior language knowledge. ProtoCLIP is trained with an online episodic training strategy, which means it can be scaled up to unlimited amounts of data. We trained our ProtoCLIP on conceptual captions (CCs) and achieved an +5.81% ImageNet linear probing improvement and an +2.01% ImageNet zero-shot classification improvement. On the larger YFCC-15M dataset, ProtoCLIP matches the performance of CLIP with 33% of training time. Delong Chen, Fan Liu 0003, Zaiquan Yang, Shaoqiu Zheng, Ying Tan 0002, Erjin Zhou |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2024 | Visual Instruction Tuning with Polite FlamingoabstractRecent research has demonstrated that the multi-task fine-tuning of multi-modal Large Language Models (LLMs) using an assortment of annotated downstream vision-language datasets significantly enhances their performance. Yet, during this process, a side effect, which we termed as the "multi-modal alignment tax", surfaces. This side effect negatively impacts the model's ability to format responses appropriately - for instance, its "politeness" - due to the overly succinct and unformatted nature of raw annotations, resulting in reduced human preference. In this paper, we introduce Polite Flamingo, a multi-modal response rewriter that transforms raw annotations into a more appealing, "polite" format. Polite Flamingo is trained to reconstruct high-quality responses from their automatically distorted counterparts and is subsequently applied to a vast array of vision-language datasets for response rewriting. After rigorous filtering, we generate the PF-1M dataset and further validate its value by fine-tuning a multi-modal LLM with it. Combined with novel methodologies including U-shaped multi-stage tuning and multi-turn augmentation, the resulting model, Clever Flamingo, demonstrates its advantages in both multi-modal understanding and response politeness according to automated and human evaluations. Code and dataset are available at https://github.com/ChenDelong1999/polite-flamingo Delong Chen, Jianfeng Liu 0002, Wenliang Dai, Baoyuan Wang |
AAAI | 1 |
| 2024 | Measuring Political Bias in Large Language Models: What Is Said and How It Is SaidabstractWe propose to measure political bias in LLMs by analyzing both the content and style of their generated content regarding political issues.Existing benchmarks and measures focus on gender and racial biases.However, political bias exists in LLMs and can lead to polarization and other harms in downstream applications.In order to provide transparency to users, we advocate that there should be fine-grained and explainable measures of political biases generated by LLMs.Our proposed measure looks at different political issues such as reproductive rights and climate change, at both the content (the substance of the generation) and the style (the lexical polarity) of such bias.We measured the political bias in eleven opensourced LLMs and showed that our proposed framework is easily scalable to other topics and is explainable. Yejin Bang, Delong Chen, Nayeon Lee, Pascale Fung |
ACL (1) | 2 |
| 2024 | Few-shot classification guided by generalization error bound
Fan Liu 0003, Sai Yang, Delong Chen, Huaxi Huang, Jun Zhou 0001 |
Pattern Recognit. | 3 |
| 2024 | Few-Shot Classification Model Compression via School LearningabstractFew-shot classification (FSC) is a challenging task due to limitation in accessing training data. Recent methods often employ highly complex networks to obtain high-quality features, but this may not be suitable for resource-limited applications. To tackle this challenge, we introduce Few-Shot Classification Model Compression (FSC-MC), a new task aimed at enhancing the FSC performance of lightweight and low-capacity models by learning from more complex models. We also propose a novel two-level learning strategy called School Learning to accomplish the FSC-MC task by mimicking the real learning process in the social school life. In this new learning paradigm, the first level performs preview learning, in which each student is equipped with a preparer to perform self-learning on the base set. The second level is the team learning, consisting of a complex teacher network and several lightweight student networks organized into a team. One student network is randomly chosen as the leader network, while the remaining student networks serve as member networks. The leader network simultaneously learns knowledge from the teacher network and all member networks. Conversely, each member network receives knowledge from both the teacher network and the leader network. Ultimately, the leader network is deployed for FSC evaluation, resulting in effective model compression. Extensive experiments in the FSC-MC setting demonstrate that School Learning outperforms 17 state-of-the-art knowledge distillation methods including both offline methods and online methods, enabling lightweight models to achieve outstanding FSC performance. Sai Yang, Fan Liu 0003, Delong Chen, Huaxi Huang, Jun Zhou 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2024 | RemoteCLIP: A Vision Language Foundation Model for Remote SensingabstractGeneral-purpose foundation models have led to recent breakthroughs in artificial intelligence. In remote sensing, self-supervised learning (SSL) and Masked Image Modeling (MIM) have been adopted to build foundation models. However, these models primarily learn low-level features and require annotated data for fine-tuning. Moreover, they are inapplicable for retrieval and zero-shot applications due to the lack of language understanding. To address these limitations, we propose RemoteCLIP, the first vision-language foundation model for remote sensing that aims to learn robust visual features with rich semantics and aligned text embeddings for seamless downstream application. To address the scarcity of pre-training data, we leverage data scaling which converts heterogeneous annotations into a unified image-caption data format based on Box-to-Caption (B2C) and Mask-to-Box (M2B) conversion. By further incorporating UAV imagery, we produce a 12 × larger pretraining dataset than the combination of all available datasets. RemoteCLIP can be applied to a variety of downstream tasks, including zero-shot image classification, linear probing,k-NN classification, few-shot classification, image-text retrieval, and object counting in remote sensing images. Evaluation on 16 datasets, including a newly introduced RemoteCount benchmark to test the object counting ability, shows that RemoteCLIP consistently outperforms baseline foundation models across different model scales. Impressively, RemoteCLIP beats the state-of-the-art method by 9.14% mean recall on the RSITMD dataset and 8.92% on the RSICD dataset. For zero-shot classification, our RemoteCLIP outperforms the CLIP baseline by up to 6.39% average accuracy on 12 downstream datasets. Fan Liu 0003, Delong Chen, Zhangqingyun Guan, Xiaocong Zhou, Qiaolin Ye, Liyong Fu, Jun Zhou 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2023 | Few-shot Classification via Ensemble Learning with Multi-Order StatisticsabstractTransfer learning has been widely adopted for few-shot classification. Recent studies reveal that obtaining good generalization representation of images on novel classes is the key to improving the few-shot classification accuracy. To address this need, we prove theoretically that leveraging ensemble learning on the base classes can correspondingly reduce the true error in the novel classes. Following this principle, a novel method named Ensemble Learning with Multi-Order Statistics (ELMOS) is proposed in this paper. In this method, after the backbone network, we use multiple branches to create the individual learners in the ensemble learning, with the goal to reduce the storage cost. We then introduce different order statistics pooling in each branch to increase the diversity of the individual learners. The learners are optimized with supervised losses during the pre-training phase. After pre-training, features from different branches are concatenated for classifier evaluation. Extensive experiments demonstrate that each branch can complement the others and our method can produce a state-of-the-art performance on multiple few-shot classification benchmark datasets. Sai Yang, Fan Liu 0003, Delong Chen, Jun Zhou 0001 |
IJCAI | 3 |
| 2023 | Asymmetric exponential loss function for crack segmentation
Fan Liu 0003, Delong Chen, Chunmei Shen, Feng Xu 0008 |
Multim. Syst. | 3 |
| 2023 | MEP-3M: A large-scale multi-modal E-commerce product dataset
Fan Liu 0003, Delong Chen, Xiaoyu Du 0002, Ruizhuo Gao, Feng Xu 0008 |
Pattern Recognit. | 2 |
| 2022 | A review of driver fatigue detection and its advances on the use of RGB-D camera and deep learning
Fan Liu 0003, Delong Chen, Jun Zhou 0001, Feng Xu 0008 |
Eng. Appl. Artif. Intell. | 2 |
| 2022 | Self-Supervised Music Motion Synchronization Learning for Music-Driven Conducting Motion Generation
Fan Liu 0003, Delong Chen, Ruizhi Zhou, Sai Yang, Feng Xu 0008 |
J. Comput. Sci. Technol. | 2 |