EDBT 2026 Demo / reviewers in the wild / expert
Zishan Xu
dblp:358/6729
· DBLP profile ↗
21ranked-venue papers
6as first author
21since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 8 · 1 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 2 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 3 first-author · 6 since 2021Computer networks · 5 · 3 first-author · 5 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Decoupling Continual Semantic SegmentationabstractContinual Semantic Segmentation (CSS) requires learning new classes without forgetting previously acquired knowledge, addressing the fundamental challenge of catastrophic forgetting in dense prediction tasks. However, existing CSS methods typically employ single-stage encoder-decoder architectures where segmentation masks and class labels are tightly coupled, leading to interference between old and new class learning and suboptimal retention-plasticity balance. We introduce DecoupleCSS, a novel two-stage framework for CSS. By decoupling class-aware detection from class-agnostic segmentation, DecoupleCSS enables more effective continual learning, preserving past knowledge while learning new classes. The first stage leverages pre-trained text and image encoders, adapted using LoRA, to encode class-specific information and generate location-aware prompts. In the second stage, the Segment Anything Model (SAM) is employed to produce precise segmentation masks, ensuring that segmentation knowledge is shared across both new and previous classes. This approach improves the balance between retention and adaptability in CSS, achieving state-of-the-art performance across a variety of challenging tasks. Yifu Guo, Yuquan Lu, Wentao Zhang 0005, Zishan Xu, Dexia Chen, Yizhe Zhang 0001 |
AAAI | 4 |
| 2026 | VideoSeg-R1: Reasoning Video Object Segmentation via Reinforcement LearningabstractTraditional video reasoning segmentation methods rely on supervised fine-tuning, which limits generalization to out-of-distribution scenarios and lacks explicit reasoning. To address this, we propose VideoSeg-R1, the first framework to introduce reinforcement learning into video reasoning segmentation. It adopts a decoupled architecture that formulates the task as joint referring image segmentation and video mask propagation. It comprises three stages: (1) A hierarchical text-guided frame sampler to emulate human attention; (2) A reasoning model that produces spatial cues along with explicit reasoning chains; and (3) A segmentation-propagation stage using SAM2 and XMem. A task difficulty-aware mechanism adaptively controls reasoning length for better efficiency and accuracy. Extensive evaluations on multiple benchmarks demonstrate that VideoSeg-R1 achieves state-of-the-art performance in complex video reasoning and segmentation tasks. Zishan Xu, Yifu Guo, Yuquan Lu, Junxin Li, Lihua Cai |
AAAI | 1 |
| 2026 | ACE-Router: Generalizing History-Aware Routing from MCP Tools to the Agent WebabstractZhiyuan Yao, Zishan Xu, Yifu Guo, Zhiguang Han, Cheng Yang, Shuo Zhang, Weinan Zhang, Xingshan Zeng, Weiwen Liu. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Zishan Xu, Yifu Guo, Zhiguang Han, Weinan Zhang 0001, Xingshan Zeng, Weiwen Liu |
ACL (1) | 2 |
| 2026 | MilleniaGuard: An Event-Driven Edge-AI and AIGC-Based IoT System for Ancient Mural Monitoring and RestorationabstractThis paper addresses the challenges of automatic monitoring and restoration in ancient mural conservation, aiming to enhance the efficiency and quality of heritage preservation. Traditional manual inspection is time-consuming and often misses early damage, while existing digital restoration models struggle with consistent restoration, especially for large-scale damage. To address these issues, we propose an Internet of things (IoT)-based solution combining event-driven edge intelligence and artificial intelligence generated content (AIGC) techniques. A fine-tuned EdgeSAM model, using a Conv-adapter, enables efficient damage segmentation at the edge; an event-driven mechanism reduces resource consumption; and a LoRA-tuned PowerPaint model, aided by Blip2 and Qwen, provides effective restoration of large damaged areas. Cloud-side processing utilizes AIGC techniques to restore damaged mural areas, ensuring high-quality restoration while minimizing communication demands. Experimental results demonstrate that the proposed method achieves accurate damage monitoring on resource-constrained edge devices and generates diverse, contextually appropriate restoration results on cloud servers, providing a deployment-oriented feasibility validation under simulated temporal degradation and real hardware constraints. Zishan Xu, Jiansen Zhang, Wei Chen 0036, Xiaofeng Zhang 0006, Jueting Liu, Zehua Wang 0001, F. Richard Yu, Victor C. M. Leung |
IEEE Internet Things J. | 1 |
| 2026 | TencentLLMEval: A Hierarchical Evaluation of Real-World Capabilities for Human-Aligned LLMsabstractLarge language models (LLMs) have shown impressive capabilities across various natural language tasks. However, evaluating their alignment with human preferences remains a challenge. To this end, we propose a comprehensive human evaluation framework to assess LLMs’ proficiency in following instructions on diverse real-world tasks. We construct a hierarchical task tree encompassing seven major areas covering over 200 categories and over 800 tasks, which covers diverse capabilities such as question answering, reasoning, multi-turn dialogue, and text generation, to evaluate LLMs in a comprehensive and in-depth manner. We also design detailed evaluation standards and processes to facilitate consistent, unbiased judgments from human evaluators. A test set of over 3,000 instances is released, spanning different difficulty levels and knowledge domains. Our work provides a standardized methodology to evaluate human alignment in LLMs for both English and Chinese. We also analyze the feasibility of automating parts of evaluation with a strong LLM (GPT-4). Our framework supports a thorough assessment of LLMs as they are integrated into real-world applications. We have made publicly available the task tree, TencentLLMEval dataset, and evaluation methodology which have been demonstrated as effective in assessing the performance of Tencent Hunyuan LLMs. By doing so, we aim to facilitate the benchmarking of advances in the development of safe and human-aligned LLMs. Shuyi Xie, Wenlin Yao, Yong Dai 0001, Zishan Xu, Fan Lin, Donglin Zhou, Lifeng Jin, Xinhua Feng, Pengzhi Wei, Zhichao Hu, Dong Yu 0001, Zhengyou Zhang |
ACM Trans. Intell. Syst. Technol. | 5 |
| 2026 | Mobiflip: Information-Bottleneck-Guided Minimal Federated Adaptation for Cross-Modal ModelsabstractCross-modal federated learning is constrained by bandwidth and on-device compute. We present Mobiflip: a minimalist strategy that freezes a lightweight backbone and communicates only a channel-wise \(1\times 1\) scaling adapter appended to the image branch. Guided by the Information Bottleneck, we prove that under common distributional and linear-encoder surrogates, per-channel scaling attains the linear optimum; coupled with the directional geometry of (Mobile)CLIP, the adapter is, in first-order approximation, an optimal preconditioner of the cosine-similarity space—preserving discriminative directions while compressing redundancy and suppressing inter-client drift. We adopt MobileCLIP as a mobile-friendly backbone to jointly minimize compute and communication. On CIFAR-10/100 and medical imaging, a single aggregation already yields stable Bacc; each round transmits only about 0.7% of backbone parameters with \(>\!\!92\%\) reduction in communication. Compared with recent federated multimodal/large-model methods, Mobiflip maintains—or even improves—accuracy under ultra-low communication. Zishan Xu, Jiansen Zhang, Wei Chen 0036, Jueting Liu, Zehua Wang 0001, Abdulmotaleb El Saddik |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2025 | EXCGEC: A Benchmark for Edit-Wise Explainable Chinese Grammatical Error CorrectionabstractExisting studies explore the explainability of Grammatical Error Correction (GEC) in a limited scenario, where they ignore the interaction between corrections and explanations and have not established a corresponding comprehensive benchmark. To bridge the gap, this paper first introduces the task of EXplainable GEC (EXGEC), which focuses on the integral role of correction and explanation tasks. To facilitate the task, we propose EXCGEC, a tailored benchmark for Chinese EXGEC consisting of 8,216 explanation-augmented samples featuring the design of hybrid edit-wise explanations. We then benchmark several series of LLMs in multi-task learning settings, including post-explaining and pre-explaining. To promote the development of the task, we also build a comprehensive evaluation suite by leveraging existing automatic metrics and conducting human evaluation experiments to demonstrate the human consistency of the automatic metrics for free-text explanations. Our experiments reveal the effectiveness of evaluating free-text explanations using traditional metrics like METEOR and ROUGE, and the inferior performance of multi-task models compared to the pipeline solution, indicating its challenges to establish positive effects in learning both tasks. Jingheng Ye, Shang Qin, Xuxin Cheng, Libo Qin 0001, Hai-Tao Zheng 0002, Ying Shen 0001, Peng Xing, Zishan Xu |
AAAI | 9 |
| 2025 | CLEME2.0: Towards Interpretable Evaluation by Disentangling Edits for Grammatical Error CorrectionabstractJingheng Ye, Zishan Xu, Yinghui Li, Linlin Song, Qingyu Zhou, Hai-Tao Zheng, Ying Shen, Wenhao Jiang, Hong-Gee Kim, Ruitong Liu, Xin Su, Zifei Shan. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Jingheng Ye, Zishan Xu, Linlin Song, Qingyu Zhou, Hai-Tao Zheng 0002, Ying Shen 0001, Hong-Gee Kim, Zifei Shan |
ACL (1) | 2 |
| 2025 | FocusDet: FocusConv and CLIP Guide Head for Remote Sensing Object Detection
Zilong Wang 0021, Wei Yang 0029, Hongxian Tian, Zishan Xu, Wei Chen 0036, Jueting Liu, Zehua Wang 0001 |
ICIC (9) | 4 |
| 2025 | EdgeSAM-CASD: Lightweight Mural Damage Segmentation via Convolutional Adapter
Jiansen Zhang, Zehua Wang 0001, Wei Chen 0036, Zishan Xu, Jueting Liu |
ICIC (5) | 4 |
| 2025 | Text2Omni: A Text-Only Training Strategy for MLLMs
Junxin Li, Yifu Guo, Zishan Xu, Siyue Chen, Siyan Wu, Lihua Cai |
PRICAI (4) | 3 |
| 2025 | CogMAS: A Cognitively-Grounded Multi-Agent Framework for Explainable and Consistent Open-Ended Student Response ScoringabstractAutomated scoring of open-ended questions continues to face significant challenges in modeling student cognition, ensuring scoring consistency, and providing interpretability. Although large language models (LLMs) have demonstrated substantial potential, single-model architectures exhibit structural limitations in handling complex reasoning and cognitive alignment. To address these issues, we propose CogMAS, a Cognitively-Grounded Multi-Agent Scoring Framework that incorporates three types of agents, i.e., student agents, teacher agents, and evaluation agents, to enable multidimensional and interpretable scoring of open-ended responses. CogMAS leverages Bloom’s taxonomy to construct a mapping between questions and cognitive dimensions, guiding teacher agents to perform dimension-aware scoring. A dual-stage semantic retrieval module is introduced to provide contextually relevant exemplars. Evaluation agents are responsible for detecting explanation path biases and deriving high-confidence reasoning chains and final scores. Teacher agents are further trained using Direct Preference Optimization (DPO) to improve the quality and consistency of scoring explanations. High-confidence score–explanation pairs are stored in a retrievable memory module to support continuous optimization in future tasks. Experiments on three public open-ended question scoring datasets demonstrate that CogMAS achieves state-of-the-art performance in both scoring accuracy and consistency, validating its effectiveness and generalizability. Yixuan Fang, Zishan Xu, Yifu Guo, Yuquan Lu |
SMC | 2 |
| 2025 | SDANet: A Federated Efficient Remote Sensing Object Detection for Space-Air-Ground IoTabstractThe explosive growth of remote-sensing images generated by emerging space–air–ground integrated IoT networks makes centralized detector training infeasible due to limited bandwidth and strict data privacy constraints. While lightweight single-stage object detectors offer efficiency, they suffer significant accuracy degradation for small, dense, and arbitrarily oriented targets. Furthermore, existing federated object detection frameworks typically neglect client heterogeneity. To overcome these limitations, we propose a two-stage personalized federated detection framework. In Stage 1, we independently train a conventional single-stage rotated object detector on each client and aggregate model updates using an adaptive similarity momentum aggregation (ASMA) strategy, effectively pooling knowledge across non-IID client datasets to improve global generalization. In Stage 2, each client is equipped with a private selective depthwise attention convolution (SDAConv) module, leveraging Stage-1 priors to reconstruct fine-grained, client-specific features without additional communication overhead, thus tailoring predictions to local data distributions. Experiments conducted on five non-IID splits derived from DOTA-1.0, along with DIOR and VisDrone datasets, demonstrate improvements of up to +3.5 mAP compared to federated learning baselines under the same communication budget, simultaneously maintaining global robustness and enhancing local detection accuracy. Zilong Wang 0021, Wei Yang 0029, Zishan Xu, Wei Chen 0036, Jueting Liu, Zehua Wang 0001, Victor C. M. Leung |
IEEE Internet Things J. | 3 |
| 2025 | MuralAgent: Enhancing Ancient Mural Outpainting with RAG-Based Texts and Multimodal IntegrationabstractIn the context of the digital age, utilizing cutting-edge technology for the digitization and creative expansion of ancient murals is crucial, aimed at preserving and passing on cultural heritage. Existing image outpainting techniques suffer from a lack of semantic guidance. This article introduces MuralAgent, a multimodal model based on Retrieval-Augmented Generation (RAG) technology. It precisely extracts key information from mural images and integrates it with a constructed ancient texts knowledge base to ensure the cultural and semantic consistency of the expanded images. Moreover, fine-tuning the Stable Diffusion model ensures the fidelity of the generated image styles. Specifically, this study involves constructing an ancient texts knowledge base for accurate matching, designing specific prompts for GPT-4V(ision) to extract key information, and innovatively expanding artworks through Stable Diffusion, providing a novel way for the public to reinterpret ancient murals. Zishan Xu, Xiaofeng Zhang 0006, Wei Chen 0036, Jueting Liu, Zehua Wang 0001, Abdulmotaleb El Saddik |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2025 | Wakeup-Darkness: When Multimodal Meets Unsupervised Low-Light Image EnhancementabstractLow-light image enhancement is a crucial visual task, and many unsupervised methods overlook the degradation of visible information in low-light scenes, adversely affecting the fusion of complementary information and hindering the generation of satisfactory results. To address this, we introduce Wakeup-Darkness, a multimodal enhancement framework that innovatively enriches user interaction through voice and textual commands. This approach signifies a technical leap and represents a paradigm shift in user engagement. We introduce a Cross-Modal Feature Fusion (CMFF) that synergizes semantic and depth context with low-light enhancement operations. Moreover, we propose a Gated Residual Block (GRB) and a channel-aware Look-Up Table (LUT) to adjust the intensity distribution of each channel. Crucially, the proposed Wakeup-Darkness scheme demonstrates remarkable generalization in unsupervised scenarios. The source code can be accessed from https://github.com/zhangbaijin/Wakeup-Dakness . Xiaofeng Zhang 0006, Zishan Xu, Hao Tang 0005, Chaochen Gu, Wei Chen 0036, Abdulmotaleb El Saddik |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2024 | Towards Real-World Writing Assistance: A Chinese Character Checking Benchmark with Faked and Misspelled CharactersabstractYinghui Li, Zishan Xu, Shaoshen Chen, Haojing Huang, Yangning Li, Shirong Ma, Yong Jiang, Zhongli Li, Qingyu Zhou, Hai-Tao Zheng, Ying Shen. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Zishan Xu, Shaoshen Chen, Haojing Huang 0001, Yangning Li, Shirong Ma, Yong Jiang 0001, Zhongli Li, Qingyu Zhou, Hai-Tao Zheng 0002, Ying Shen 0001 |
ACL (1) | 2 |
| 2024 | Harmonizing Stable Diffusion and GPT-4 for Mural Expansion with ArtExtend
Dufeng Chen, Zehua Wang 0001, Zishan Xu, Jueting Liu, Wei Chen 0036 |
ICIC (7) | 4 |
| 2024 | A Large Model Assisted Remote Sensing Image Scene Understanding Algorithm Based on Object Detection
Zilong Wang 0021, Zishan Xu, Wei Yang 0029, Wei Chen 0036, Yuyu Yang |
ICIC (6) | 2 |
| 2024 | MuralRescue: Advancing Blind Mural Restoration via SAM-Adapter Enhanced Damage Segmentation and Integrated Restoration Techniques
Zishan Xu, Dufeng Chen, Qianzhen Fang, Wei Chen 0036, Jueting Liu, Zehua Wang 0001 |
ICIC (7) | 1 |
| 2024 | DeepCRBP: improved predicting function of circRNA-RBP binding sites with deep feature learning
Zishan Xu, Linlin Song, Shichao Liu 0002, Wen Zhang 0008 |
Frontiers Comput. Sci. | 1 |
| 2024 | Shadclips: When Parameter-Efficient Fine-Tuning with Multimodal Meets Shadow RemovalabstractSegment Anything Model (SAM), an advanced universal image segmentation model trained on an expansive visual dataset, has set a new benchmark in image segmentation and computer vision. However, it faced challenges when it came to distinguishing between shadows and their backgrounds. To address this, we proposed ShadClips, which consists of SAM-optimizer and SONet. It has dramatically enhanced SAM’s ability to segment shadow images, differentiating between the background and both soft and hard shadows adeptly. Due to its dependence on pixel point inputs, the SAM-Optimizer interface could do better. This method presents challenges, especially when dealing with long, extended shadows. To make the user experience more intuitive and effective, we incorporated the capabilities of CLIPs. Therefore, simple text descriptions like “A photo of a shadow” can be used to guide the SAM-Optimizer, allowing it to select the most relevant shadow mask from SAM’s comprehensive category list. Meanwhile, we introduce SONet to shadow removal. A large number of experiments on ISTD/SRD prove that the proposed method is effective and satisfactory. The source code of the ShadClips can be accessed from https://github.com/zhangbaijin/SAM-helps-Shadow . Xiaofeng Zhang 0006, Chaochen Gu, Zishan Xu, Hao Tang 0005, Hao Cheng 0004, Kaijie Wu 0002, Shanying Zhu |
Int. J. Pattern Recognit. Artif. Intell. | 3 |