EDBT 2026 Demo / reviewers in the wild / expert
Jiajun Song
dblp:285/5571
· DBLP profile ↗
24ranked-venue papers
5as first author
24since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 14 · 3 first-author · 14 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 1 first-author · 9 since 2021Systems, architecture and hardware · 3 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Computer networks · 2 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Mind the Gap: The Divergence Between Human and LLM-Generated TasksabstractHumans constantly generate a diverse range of tasks guided by internal motivations. While generative agents powered by large language models (LLMs) aim to simulate this complex behavior, it remains uncertain whether they operate on similar cognitive principles. To address this, we conducted a task-generation experiment comparing human responses with those of an LLM agent (GPT-4o). We find that human task generation is consistently influenced by psychological drivers, including personal values (e.g., Openness to Change) and cognitive style. Even when these psychological drivers are explicitly provided to the LLM, it fails to reflect the corresponding behavioral patterns. They produce tasks that are markedly less social, less physical, and thematically biased toward abstraction. Interestingly, while the LLM's tasks were perceived as more fun and novel, this highlights a disconnect between its linguistic proficiency and its capacity to generate human-like, embodied goals. We conclude that there is a core gap between the value-driven, embodied nature of human cognition and the statistical patterns of LLMs, highlighting the necessity of incorporating intrinsic motivation and physical grounding into the design of more human-aligned agents. Yi-Long Lu, Jiajun Song |
AAAI | 2 |
| 2026 | MDK12-Bench: A Multi-Discipline Benchmark for Evaluating Reasoning in Multimodal Large Language ModelsabstractMultimodal large language models (MLLMs), which integrate language and visual cues for problem-solving, are crucial for advancing artificial general intelligence (AGI). However, current benchmarks for measuring the intelligence of MLLMs suffer from limited scale, narrow coverage, and unstructured knowledge, offering only static and undifferentiated evaluations. To bridge this gap, we introduce MDK12-Bench, a large-scale multidisciplinary benchmark built from real-world K–12 exams spanning six disciplines with 141K instances and 6,225 knowledge points organized in a six-layer taxonomy. Covering five question formats with difficulty and year annotations, it enables comprehensive evaluation to capture the extent to which MLLMs perform over four dimensions: 1) difficulty levels, 2) temporal (cross-year) shifts, 3) contextual shifts, and 4) knowledge-driven reasoning. We propose a novel dynamic evaluation framework that introduces unfamiliar visual, textual, and question form shifts to challenge model generalization while improving benchmark objectivity and longevity by mitigating data contamination. We further evaluate knowledge-point reference-augmented generation (KP-RAG) to examine the role of knowledge in reasoning. Key findings reveal limitations in current MLLMs in multiple aspects and provide guidance for enhancing model reasoning, robustness, and AI-assisted education. Xiaopeng Peng 0001, Fanrui Zhang, Zhaopan Xu, Jiaxin Ai, Yansheng Qiu, Wangbo Zhao, Jiajun Song, Chuanhao Li 0001, Weidong Tang, Zhen Li 0026, Haoquan Zhang, Zizhen Li, Xiaofeng Mao, Yukang Feng, Kai Wang 0036, Xiaojun Chang, Wenqi Shao, Yang You 0001, Kaipeng Zhang |
AAAI | 8 |
| 2026 | MARCH: Multi-Agent Reinforced Check for HallucinationabstractZhuo Li, Yupeng Zhang, Pengyu Cheng, Jiajun Song, Mengyu Zhou, Hao Li, Shujie Hu, Yu Qin, Erchao.zec, Xiaoxi Jiang, Guanjunjiang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Pengyu Cheng, Jiajun Song, Mengyu Zhou, Shujie Hu, Erchao Zhao, Xiaoxi Jiang, Guanjun Jiang |
ACL (1) | 4 |
| 2026 | PIDA: Proactive Interactive Dietary Assessment from Food Images via Active Clarification
Donghui Si, Yongcheng Yin, Jiajun Song |
KSEM (6) | 4 |
| 2026 | Single-shot fringe projection profilometry with a TGV-regularized u-net enhanced fourier neural operator
Jiajun Song, Chenxia Wan |
Expert Syst. Appl. | 1 |
| 2025 | What Factors Influence Goal Setting? Insights from Text-Based Task Generation
Yi-Long Lu, Jiajun Song |
CogSci | 2 |
| 2025 | OpenING: A Comprehensive Benchmark for Judging Open-ended Interleaved Image-Text GenerationabstractMultimodal Large Language Models (MLLMs) have made significant strides in visual understanding and generation tasks. However, generating interleaved image-text content remains a challenge, which requires integrated multimodal understanding and generation abilities. While the progress in unified models offers new solutions, existing benchmarks are insufficient for evaluating these methods due to data size and diversity limitations. To bridge this gap, we introduce OpenING, a comprehensive benchmark comprising 5,400 high-quality human-annotated instances across 56 real-world tasks. OpenING covers diverse daily scenarios such as travel guide, design, and brainstorming, offering a robust platform for challenging interleaved generation methods. In addition, we present IntJudge, a judge model for evaluating open-ended multimodal generation methods. Trained with a novel data pipeline, our IntJudge achieves an agreement rate of 82.42% with human judgments, outperforming GPT-based evaluators by 11.34%. Extensive experiments on OpenING reveal that current interleaved generation methods still have substantial room for improvement. Key findings on interleaved image-text generation are further presented to guide the development of next-generation models. Xiaopeng Peng 0001, Jiajun Song, Chuanhao Li 0001, Zhaopan Xu, Ziyao Guo, Hao Zhang 0117, Yuqi Lin, Yefei He, Lirui Zhao, Xiaojun Chang, Yu Qiao 0001, Wenqi Shao, Kaipeng Zhang |
CVPR | 3 |
| 2025 | Beyond A Single AI Cluster: A Survey of Decentralized LLM TrainingabstractThe emergence of large language models (LLMs) has revolutionized AI development, yet their resource demands beyond a single cluster or even datacenter, limiting accessibility to well-resourced organizations.Decentralized training has emerged as a promising paradigm to leverage dispersed resources across clusters, datacenters and even regions, offering the potential to democratize LLM development for broader communities.As the first comprehensive exploration of this emerging field, we present decentralized LLM training as a resource-driven paradigm and categorize existing efforts into community-driven and organizational approaches.We further clarify this through: (1) a comparison with related paradigms, (2) characterization of decentralized resources, and (3) a taxonomy of recent advancements.We also provide up-to-date case studies and outline future directions to advance research in decentralized LLM training. Haotian Dong, Jingyan Jiang, Rongwei Lu, Jiajun Luo, Jiajun Song, Zhi Wang 0001 |
EMNLP | 5 |
| 2025 | SalientFusion: Context-Aware Compositional Zero-Shot Food Recognition
Jiajun Song, Xiaoou Liu |
ICANN (2) | 1 |
| 2025 | DICE: Staleness-Centric Optimizations for Parallel Diffusion MoE Inference
Jiajun Luo, Lizhuo Luo, Jianru Xu, Jiajun Song, Rongwei Lu, Zhi Wang 0001 |
ICCV | 4 |
| 2025 | OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and ReasoningabstractScoring the Optical Character Recognition (OCR) capabilities of Large Multimodal Models (LMMs) has witnessed growing interest. Existing benchmarks have highlighted the impressive performance of LMMs in text recognition; however, their abilities in certain challenging tasks, such as text localization, handwritten content extraction, and logical reasoning, remain underexplored. To bridge this gap, we introduce OCRBench v2, a large-scale bilingual text-centric benchmark with currently the most comprehensive set of tasks ($4\times$ more tasks than the previous multi-scene benchmark OCRBench), the widest coverage of scenarios ($31$ diverse scenarios), and thorough evaluation metrics, with $10,000$ human-verified question-answering pairs and a high proportion of difficult samples. Moreover, we construct a private test set with $1,500$ manually annotated images. The consistent evaluation trends observed across both public and private test sets validate the OCRBench v2's reliability. After carefully benchmarking state-of-the-art LMMs, we find that most LMMs score below $50$ ($100$ in total) and suffer from five-type limitations, including less frequently encountered text recognition, fine-grained perception, layout perception, complex element parsing, and logical reasoning. The benchmark and evaluation scripts are available at https://github.com/Yuliang-Liu/MultimodalOCR. Zhebin Kuang, Jiajun Song, Mingxin Huang, Linghao Zhu, Qidi Luo, Xinyu Wang 0010, Hao Lu 0003, Guozhi Tang, Bin Shan, Chunhui Lin, Binghong Wu, Hao Feng 0009, Hao Liu 0003, Can Huang 0002, Jingqun Tang, Wei Chen 0088, Xiang Bai |
NeurIPS | 3 |
| 2025 | VARMA-Enhanced Transformer for Time Series Forecasting
Jiajun Song, Xiaoou Liu |
PRICAI (5) | 1 |
| 2024 | A Joint Approach to Local Updating and Gradient Compression for Efficient Asynchronous Federated Learning
Jiajun Song, Jiajun Luo, Rongwei Lu, Shuzhao Xie, Bin Chen 0011, Zhi Wang 0001 |
Euro-Par (3) | 1 |
| 2024 | Dataset and Benchmark for Urdu Natural Scenes Text Detection, Recognition and Visual Question Answering
Hiba Maryam, Jiajun Song, Tajrian ABM Shafayet, Qidi Luo, Xiang Bai |
ICDAR (5) | 3 |
| 2024 | A 10-kHz 12-16-bit reconfigurable zoom ADC with pole optimization technique and floating current-starved amplifier
Zhangming Zhu, Jiajun Song, Yuhua Liang |
Sci. China Inf. Sci. | 2 |
| 2024 | Multi-state Ingredient Recognition via Adaptive Multi-centric NetworkabstractIngredient recognition has received significant attention due to its numerous industrial applications, such as intelligent retail terminals and intelligent cooking devices. However, ingredient recognition has the following challenges: 1) dynamic changes in the number of categories; 2) greater diversity and regionality of ingredients; and 3) large visual differences among different states of ingredients. In this article, we propose an adaptive multi-centric network (AdMNet) to solve the problem of ingredient recognition. AdMNet is based on the idea of retrieval, which consists of two main parts, the adaptive multi-centric nearest-neighbor central mean (AdM-NCM) classifier, and the context-aware attentional pooling (CAP) module. The AdM-NCM classifier adaptively establishes category-centric vector groups to recognize ingredients via optimizing the minimum clustering variance, where each state of the ingredient has its corresponding centric vector. The CAP module combines contextual information and multiple attention mechanisms. It captures more focused and discriminative features with higher weights assigned to fine-grained features, which results in better feature representation. In addition, we collect a large-scale ingredient dataset, ISIA Ingredient-201 with 201 classes and 100 442 images. To prove the greater robustness and generalization of our method, we compare the metrics in basic scenarios and realistic scenarios with those of other methods. Specifically, the base scenario is the regular setup, and the real scenario is similar to the class incremental learning setup. The experimental results show that our method reaches the state of the art on both basic scenarios and realistic scenarios with small samples. Jiajun Song, Weiqing Min, Weimin Xiao, Shuqiang Jiang |
IEEE Trans. Ind. Informatics | 2 |
| 2024 | Synthesizing Knowledge-Enhanced Features for Real-World Zero-Shot Food DetectionabstractFood computing brings various perspectives to computer vision like vision-based food analysis for nutrition and health. As a fundamental task in food computing, food detection needs Zero-Shot Detection (ZSD) on novel unseen food objects to support real-world scenarios, such as intelligent kitchens and smart restaurants. Therefore, we first benchmark the task of Zero-Shot Food Detection (ZSFD) by introducing FOWA dataset with rich attribute annotations. Unlike ZSD, fine-grained problems in ZSFD like inter-class similarity make synthesized features inseparable. The complexity of food semantic attributes further makes it more difficult for current ZSD methods to distinguish various food categories. To address these problems, we propose a novel framework ZSFDet to tackle fine-grained problems by exploiting the interaction between complex attributes. Specifically, we model the correlation between food categories and attributes in ZSFDet by multi-source graphs to provide prior knowledge for distinguishing fine-grained features. Within ZSFDet, Knowledge-Enhanced Feature Synthesizer (KEFS) learns knowledge representation from multiple sources (e.g., ingredients correlation from knowledge graph) via the multi-source graph fusion. Conditioned on the fusion of semantic knowledge representation, the region feature diffusion model in KEFS can generate fine-grained features for training the effective zero-shot detector. Extensive evaluations demonstrate the superior performance of our method ZSFDet on FOWA and the widely-used food dataset UECFOOD-256, with significant improvements by 1.8% and 3.7% ZSD mAP compared with the strong baseline RRFS. Further experiments on PASCAL VOC and MS COCO prove that enhancement of the semantic knowledge can also improve the performance on general ZSD. Code and dataset are available at https://github.com/LanceZPF/KEFS. Weiqing Min, Jiajun Song, Yang Zhang 0117, Shuqiang Jiang |
IEEE Trans. Image Process. | 3 |
| 2024 | Towards Food Image Retrieval via Generalization-Oriented Sampling and Loss Function DesignabstractFood computing has increasingly received widespread attention in the multimedia field. As a basic task of food computing, food image retrieval has wide applications, that is, food image retrieval can help users to find the desired food from a large number of food images. Besides, the retrieved information can be applied to establish a richer database for the subsequent food content-related recommendation. Food image retrieval aims to achieve better performance on novel categories. Thus, it is worth studying to transfer the embedding ability from the training set to the unseen test set, that is, the generalization of the model. Food is influenced by various factors, such as culture and geography, leading to great differences between domains, such as Asian food and western food. Therefore, it is challenging to study the generalization of the model in food image retrieval. In this article, we improve the classical metric learning framework and propose a generalization-oriented sampling strategy, which boosts the generalization of the model by maximizing the intra-class distance from a proportion of positive pairs to avoid the excessive distance compression in the embedding space. Considering that the existing optimization process is in an opposite direction to our proposed sampling strategy, we further propose an adaptive gradient assignment policy named gradient-adaptive optimization , which can alleviate the intra-class distance compression during optimization by assigning different gradients to different samples. Extensive evaluation on three popular food image datasets demonstrates the effectiveness of the proposed method. We also experiment on three popular general datasets to prove that solving the problem from the generalization can also improve the performance of general image retrieval. Code is available at https://github.com/Jiajun-ISIA/Generalization-oriented-Sampling-and-Loss . Jiajun Song, Weiqing Min, Shuqiang Jiang |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2023 | DAGC: Data-Aware Adaptive Gradient CompressionabstractGradient compression algorithms are widely used to alleviate the communication bottleneck in distributed ML. However, existing gradient compression algorithms suffer from accuracy degradation in Non-IID scenarios, because a uniform compression scheme is used to compress gradients at workers with different data distributions and volumes, since workers with larger volumes of data are forced to adapt to the same aggressive compression ratios as others. Assigning different compression ratios to workers with different data distributions and volumes is thus a promising solution. In this study, we first derive a function from capturing the correlation between the number of training iterations for a model to converge to the same accuracy, and the compression ratios at different workers; This function particularly shows that workers with larger data volumes should be assigned with higher compression ratios1to guarantee better accuracy. Then, we formulate the assignment of compression ratios to the workers as an n-variables chi-square nonlinear optimization problem under fixed and limited total communication constrain. We propose an adaptive gradient compression strategy called DAGC, which assigns each worker a different compression ratio according to their data volumes. Our experiments confirm that DAGC can achieve better performance facing highly imbalanced data volume distribution and restricted communication. Rongwei Lu, Jiajun Song, Bin Chen 0011, Laizhong Cui, Zhi Wang 0001 |
INFOCOM | 2 |
| 2023 | SeeDS: Semantic Separable Diffusion Synthesizer for Zero-shot Food DetectionabstractFood detection is becoming a fundamental task in food computing that supports various multimedia applications, including food recommendation and dietary monitoring. To deal with real-world scenarios, food detection needs to localize and recognize novel food objects that are not seen during training, demanding Zero-Shot Detection (ZSD). However, the complexity of semantic attributes and intra-class feature diversity poses challenges for ZSD methods in distinguishing fine-grained food classes. To tackle this, we propose the Semantic Separable Diffusion Synthesizer (SeeDS) framework for Zero-Shot Food Detection (ZSFD). SeeDS consists of two modules: a Semantic Separable Synthesizing Module (S3M) and a Region Feature Denoising Diffusion Model (RFDDM). The S3M learns the disentangled semantic representation for complex food attributes from ingredients and cuisines, and synthesizes discriminative food features via enhanced semantic information. The RFDDM utilizes a novel diffusion model to generate diversified region features and enhances ZSFD via fine-grained synthesized features. Extensive experiments show the state-of-the-art ZSFD performance of our proposed method on two food datasets, ZSFooD and UECFOOD-256. Moreover, SeeDS also maintains effectiveness on general ZSD datasets, PASCAL VOC and MS COCO. The code and dataset can be found at https://github.com/LanceZPF/SeeDS https://github.com/LanceZPF/SeeDS. Weiqing Min, Yang Zhang 0117, Jiajun Song, Shuqiang Jiang |
ACM Multimedia | 4 |
| 2023 | A Reconfigurable 12-to-18-Bit Dynamic Zoom ADC With Pole-Optimized TechniqueabstractThis paper presents a resolution-reconfigurable discrete-time dynamic zoom analog-to-digital converter (ADC) with a 20-kHz bandwidth. It employs a coarse 6-bit successive approximation register (SAR) ADC that dynamically updates the references for the post-stage single-bit second-order delta-sigma modulator to obtain a higher resolution with lower power. The sampling rate of the proposed ADC can be configured to be 1.6, 3.2, 6.4 and 10-MS/s such that four resolution modes, including 12, 14, 16 and 18-bit, can be provided accordingly. Note that pole locations of the noise transfer function (NTF) vary with different OSR’s, and noise-shaping effect, together with the power efficiency, can be degraded. To solve this issue, the pole-optimization technique is proposed in this paper. In this way can the pole locations be reconfigured according to the specific OSR. Additionally, the bandwidths of the operational amplifiers are also modulated in different oversampling ratio (OSR) scenarios to further improve the energy efficiency. Fabricated in a 0.18-$\mu \text{m}$CMOS process, the prototype occupies 1.01 mm2. On condition of an OSR of 250, the proposed ADC achieves a signal-to-noise-and-distortion-ratio (SNDR) of 102.8 dB, while dissipating 1.3 mW. Yuhua Liang, Jinyu Ren, Haotian Lan, Jiajun Song, Shida Song, Zhangming Zhu |
IEEE Trans. Circuits Syst. I Regul. Pap. | 5 |
| 2023 | Ingredient Prediction via Context Learning Network With Class-Adaptive Asymmetric LossabstractIngredient prediction has received more and more attention with the help of image processing for its diverse real-world applications, such as nutrition intake management and cafeteria self-checkout system. Existing approaches mainly focus on multi-task food category-ingredient joint learning to improve final recognition by introducing task relevance, while seldom pay attention to making good use of inherent characteristics of ingredients independently. Actually, there are two issues for ingredient prediction. First, compared with fine-grained food recognition, ingredient prediction needs to extract more comprehensive features of the same ingredient and more detailed features of various ingredients from different regions of the food image. Because it can help understand various food compositions and distinguish the differences within ingredient features. Second, the ingredient distributions are extremely unbalanced. Existing loss functions can not simultaneously solve the imbalance between positive-negative samples belonging to each ingredient and significant differences among all classes. To solve these problems, we propose a novel framework named Class-Adaptive Context Learning Network (CACLNet) for ingredient prediction. In order to extract more comprehensive and detailed features, we introduce Ingredient Context Learning (ICL) to reduce the negative impact of complex background in food images and construct internal spatial connections among ingredient regions of food objects in a self-supervised manner, which can strengthen the contacts of the same ingredients through region interactions. In order to solve the imbalance of different classes among ingredients, we propose one novel Class-Adaptive Asymmetric Loss (CAAL) to focus on various ingredient classes adaptively. Besides, considering that the over-suppression of negative samples will over-fit positive samples of those rare ingredients, CAAL alleviates this continuous suppression according to the imbalanced ratios based on gradients while maintaining the contribution of positive samples by lesser suppression. Extensive evaluation on two popular benchmark datasets (Vireo Food-172, UEC Food-100) demonstrates our proposed method achieves the state-of-the-art performance. Further qualitative analysis and visualization show the effectiveness of our method. Code and models are available at https://123.57.42.89/codes/CACLNet/index.html. Mengjiang Luo, Weiqing Min, Jiajun Song, Shuqiang Jiang |
IEEE Trans. Image Process. | 4 |
| 2023 | A Time-Domain Reconfigurable Second-Order Noise Shaping ADC With Single Fan-Out Gated Delay CellsabstractThis brief proposes a time-domain second-order noise shaping analog-to-digital converter (ADC). It employs a voltage-to-time converter (VTC) in tandem with a second-order noise shaping time-to-digital converter (TDC), which realizes the goal of reducing the power consumption and accommodating different resolutions by configuring locations of poles of the transfer function. Note that the time-register constructed TDC, which performs the signal processing function in the time domain, is implemented with digital cells thoroughly. In this way, the proposed architecture can be more friendly to the increasingly scaled process. The reconfigurable architecture enables to obtain the target signal-to-noise-and-distortion ratio (SNDR) with decreased oversampling ratio (OSR), which improves energy efficiency relatively. In addition, compared with the conventional CMOS gated delay cell, the proposed single fan-out (SFO) gated delay cell is featured with lower power dissipation and shorter delaying time. The prototype is implemented in a 0.18-$\mu \text{m}$CMOS technology to demonstrate the validity of this proposed time-domain ADC. It achieves 64.5-, 56.1-, and 51.2-dB SNDR in 0.025-, 0.05-, and 0.1-MHz bandwidth on conditions of the OSR being 200, 100, and 50, respectively. Yuhua Liang, Haotian Lan, Jiajun Song, Shida Song, Zhangming Zhu |
IEEE Trans. Very Large Scale Integr. Syst. | 5 |
| 2022 | Rethinking the Optimization of Average Precision: Only Penalizing Negative Instances before Positive Ones Is EnoughabstractOptimising the approximation of Average Precision (AP) has been widely studied for image retrieval. Limited by the definition of AP, such methods consider both negative and positive instances ranking before each positive instance. However, we claim that only penalizing negative instances before positive ones is enough, because the loss only comes from these negative instances. To this end, we propose a novel loss, namely Penalizing Negative instances before Positive ones (PNP), which can directly minimize the number of negative instances before each positive one. In addition, AP-based methods adopt a fixed and sub-optimal gradient assignment strategy. Therefore, we systematically investigate different gradient assignment solutions via constructing derivative functions of the loss, resulting in PNP-I with increasing derivative functions and PNP-D with decreasing ones. PNP-I focuses more on the hard positive instances by assigning larger gradients to them and tries to make all relevant instances closer. In contrast, PNP-D pays less attention to such instances and slowly corrects them. For most real-world data, one class usually contains several local clusters. PNP-I blindly gathers these clusters while PNP-D keeps them as they were. Therefore, PNP-D is more superior. Experiments on three standard retrieval datasets show consistent results with the above analysis. Extensive evaluations demonstrate that PNP-D achieves the state-of-the-art performance. Code is available at https://github.com/interestingzhuo/PNPloss Weiqing Min, Jiajun Song, Liping Kang, Xiaoming Wei, Xiaolin Wei, Shuqiang Jiang |
AAAI | 3 |