EDBT 2026 Demo / reviewers in the wild / expert
Xinrong Chen
dblp:116/1494
· DBLP profile ↗
27ranked-venue papers
0as first author
25since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 11 · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 9 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 6 since 2021Databases, data management, data science and information retrieval · 4 · 4 since 2021Systems, architecture and hardware · 2 · 2 since 2021Computer networks · 1Security and privacy · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | KineST: A Kinematics-guided Spatiotemporal State Space Model for Human Motion Tracking from Sparse SignalsabstractFull-body motion tracking plays an essential role in AR/VR applications, bridging physical and virtual interactions. However, it is challenging to reconstruct realistic and diverse full-body poses based on sparse signals obtained by head-mounted displays, which are the main devices in AR/VR scenarios. Existing methods for pose reconstruction often incur high computational costs or rely on separately modeling spatial and temporal dependencies, making it difficult to balance accuracy, temporal coherence, and efficiency. To address this problem, we propose KineST, a novel kinematics-guided state space model, which effectively extracts spatiotemporal dependencies while integrating local and global pose perception. The innovation comes from two core ideas. Firstly, in order to better capture intricate joint relationships, the scanning strategy within the State Space Duality framework is reformulated into kinematics-guided bidirectional scanning, which embeds kinematic priors. Secondly, a mixed spatiotemporal representation learning approach is employed to tightly couple spatial and temporal contexts, balancing accuracy and smoothness. Additionally, a geometric angular velocity loss is introduced to impose physically meaningful constraints on rotational variations for further improving motion stability. Extensive experiments demonstrate that KineST has superior performance in both accuracy and temporal consistency within a lightweight framework. Shuting Zhao, Xinrong Chen |
AAAI | 3 |
| 2026 | Towards Order Fairness: Mitigating LLMs Order Sensitivity through Dual Group Advantage OptimizationabstractLarge Language Models (LLMs) suffer from order bias, where their performance is affected by the arrangement order of input elements.This unfairness limits the model's applications in scenarios such as in-context learning and Retrieval-Augmented Generation (RAG).Recent studies attempt to obtain optimal or suboptimal arrangements based on statistical results or using dataset-based search, but these methods increase inference overhead while leaving the model's inherent order bias unresolved.Other studies mitigate order sensitivity through supervised fine-tuning using augmented training sets with multiple order variants, but often at the cost of accuracy, trapping the model in consistent yet incorrect hallucinations.In this paper, we propose Dual Group Advantage Optimization (DGAO), which aims to improve model accuracy and order stability simultaneously.DGAO calculates and balances intragroup relative accuracy advantage and intergroup relative stability advantage, rewarding the policy model for generating order-stable and correct outputs while penalizing ordersensitive or incorrect responses.This marks the first time reinforcement learning has been used to mitigate LLMs' order sensitivity.We also propose two new metrics, Consistency Rate and Overconfidence Rate, to reveal the pseudostability of previous methods and guide more comprehensive evaluation.Extensive experiments demonstrate that DGAO achieves superior order fairness while improving performance on RAG, mathematical reasoning, and classification tasks.Our Zhijie Tan, Xinrong Chen, Tong Mo |
ACL (1) | 4 |
| 2026 | MuSe: Multi-Stage Graph Reasoning via Vision-Language ModelsabstractGraph-related tasks are traditionally addressed with Graph Neural Networks (GNNs) or graph transformers, but their task-specific training limits generalization.Large Language Models (LLMs) offer stronger generalization, yet encoding graphs as one-dimensional text struggles to capture multi-hop dependencies and two-dimensional topology.Vision-Language Models (VLMs) provide an alternative by visualizing graphs, but rendering large graphs in a single image causes clutter, occlusion, and distraction, hindering reasoning.We propose MuSe, a novel multi-stage graph reasoning framework based on VLMs.Instead of processing entire graphs at once, MuSe incrementally samples and visualizes task-relevant subgraphs, enabling progressive reasoning.The framework employs a two-stage training paradigm: supervised fine-tuning to acquire local sampling and reasoning skills, followed by reinforcement learning with GRPO to refine the sampling strategy and control dialog length.To support evaluation, we introduce LGVLQA, a new multimodal dataset with larger and more complex graph structures, addressing the scalability limitations of existing benchmarks.Experiments show that MuSe consistently outperforms leading LLM and VLM baselines, demonstrating improved structural understanding and reasoning ability.Our code and data are available at this url. Guanyu Wang 0002, Xu Chu 0001, Zhijie Tan, Xinrong Chen, Tong Mo, Weiping Li 0002 |
ACL (1) | 4 |
| 2026 | MORE-R1: Guiding LVLM for Multimodal Object-Entity Relation Extraction via Stepwise Reasoning with Reinforcement Learning
Xu Chu 0001, Xinrong Chen, Haochen Li 0001, Zonghong Dai, Hongcheng Fan, Xiaoyue Yuan, Weiping Li 0002, Tong Mo |
DASFAA (6) | 3 |
| 2026 | Extreme cardiac MRI analysis under respiratory motion: Results of the CMRxMotion challenge
Kang Wang 0017, Chen Qin, Zhang Shi, Haoran Wang 0009, Chen Chen 0042, Cheng Ouyang, Chengliang Dai, Yuanhan Mo, Chenchen Dai, Xutong Kuang, Ruizhe Li 0005, Xin Chen 0003, Xiuzheng Yue, Song Tian, Alejandro Mora-Rubio, Kumaradevan Punithakumar, Shizhan Gong, Qi Dou 0001, Sina Amirrajab, Yasmina Alkhalil, Cian M. Scannell, Lexiaozi Fan, Huili Yang, Xiaowu Sun, Rob J. van der Geest, Tewodros Weldebirhan Arega, Fabrice Mériaudeau, Caner Ozer, Amin Ranem, John Kalkhof, Ilkay Öksüz, Anirban Mukhopadhyay 0003, Abdul Qayyum 0002, Moona Mazher, Steven A. Niederer, Carles García-Cabrera, Eric Arazo Sanchez, Michal K. Grzeszczyk, Szymon Plotka, Wanqin Ma, Xiaomeng Li 0001, Rongjun Ge, Yongqing Kou, Xinrong Chen, He Wang 0016, Chengyan Wang, Wenjia Bai, Shuo Wang 0011 |
Medical Image Anal. | 45 |
| 2026 | NEPose: A novel benchmark dataset with an improved framework for vision-based nasal endoscope pose estimation
Liangjing Shao, Benshuang Chen, Xinrong Chen |
Pattern Recognit. | 4 |
| 2026 | EndoLoc: Relative Pose Regression Framework With Transformation and Correlation Features for Visual Localization of EndoscopeabstractReal-time localization of endoscope is significant for the navigation and automation of endoscopic diagnosis and minimally invasive surgery. However, traditional localization based on optical tracking or magnetic tracking is easily influenced by occlusion or electromagnetic interference, while the implementation is complicated. Meanwhile, transformation and correlation information in image pairs are still ignored in existing visual localization methods for endoscopy. In this work, a novel relative pose regression framework is proposed for relative pose estimation and absolute pose tracking of endoscope based on endoscopic videos. Firstly, scene features and transformation features are respectively extracted from endoscopic observations and the corresponding optical flow by the proposed feature encoder based on gated convolution, which can prevent gradient vanishing when training the encoder from scratch on endoscopic data. Furthermore, a novel correlation module based on cross-attention is proposed to extract correlation features from two input images, which can capture more key features in endoscopic frames with more limited vision from local to global. Moreover, a novel pose decoder with upsampling and downsampling on the channel dimension is utilized to extract richer representation from the concatenated feature map for relative transformation vector prediction. The proposed method outperforms the state-of-the-art methods on the datasets from nasal endoscopy and colonoscopy, with an average localization error of less than 5%. The further experiments also demonstrate the efficiency of the proposed method. The demo videos of visual localization can be found on https://endoloc.netlify.app/ Liangjing Shao, Benshuang Chen, Shuting Zhao, Fuming Yang, Xinrong Chen |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2025 | Stepwise Schema-Guided Prompting Framework with Parameter Efficient Instruction Tuning for Multimedia Event ExtractionabstractMultimedia Event Extraction (MEE) has become an important task in information extraction research as news today increasingly prefers to contain multimedia content. Current MEE works mainly face two challenges: (1) Inadequate extraction framework modeling for handling complex and flexible multimedia event structure; (2) The absence of multimodal-aligned training data for effective knowledge transfer to MEE task. In this work, we propose a Stepwise Schema-Guided Prompting Framework (SSGPF) using Multimodal Large Language Model (MLLM) as backbone for adaptive structure capturing to solve MEE task. At the initial step of SSGPF, we design Event Type Schema Guided Prompting (ETSGP) for event detection, then we devise Argument Role Schema Guided Prompting (ARSGP) that contains multi-step prompts with text-bridged grounding technique for argument extraction. We construct a weakly-aligned multimodal event labeled dataset based on existing unimodal event annotations, then conduct parameter efficient instruction tuning with LoRA on LLaVA-v1.5-7B under SSGPF. Experiments on the M2E2 benchmark demonstrate that SSGPF significantly outperforms current SOTA baselines by 5.8 percent F1 on event detection and 8.4 percent F1 on argument extraction. Xinrong Chen, Haochen Li 0001, Guanyu Wang 0002, Weiping Li 0002, Tong Mo |
ICME | 2 |
| 2025 | Remote: Real-Time Ego-Motion Tracking for Various Endoscopes via Multimodal Visual Feature LearningabstractReal-time ego-motion tracking for endoscope is a significant task for efficient navigation and robotic automation of endoscopy. In this paper, a novel framework is proposed to perform real-time ego-motion tracking for endoscope. Firstly, a multi-modal visual feature learning network is proposed to perform relative pose prediction, in which the motion feature from the optical flow, the scene features and the joint feature from two adjacent observations are all extracted for prediction. Due to more correlation information in the channel dimension of the concatenated image, a novel feature extractor is designed based on an attention mechanism to integrate multi-dimensional information from the concatenation of two continuous frames. To extract more complete feature representation from the fused features, a novel pose decoder is proposed to predict the pose transformation from the concatenated feature map at the end of the framework. At last, the absolute pose of endoscope is calculated based on relative poses. The experiment is conducted on three datasets of various endoscopic scenes and the results demonstrate that the proposed method outperforms state-of-the-art methods. Besides, the inference speed of the proposed method is over 30 frames per second, which meets the real-time requirement. The project page is here: remote-bmxs.netlify.app Liangjing Shao, Benshuang Chen, Shuting Zhao, Xinrong Chen |
ICRA | 4 |
| 2025 | EndoMUST: Monocular Depth Estimation for Robotic Endoscopy via End-to-end Multi-step Self-supervised TrainingabstractMonocular depth estimation and ego-motion estimation are significant tasks for scene perception and navigation in stable, accurate and efficient robot-assisted endoscopy. To tackle lighting variations and sparse textures in endoscopic scenes, multiple techniques including optical flow, appearance flow and intrinsic image decomposition have been introduced into the existing methods. However, the effective training strategy for multiple modules are still critical to deal with both illumination issues and information interference for self-supervised depth estimation in endoscopy. Therefore, a novel framework with multistep efficient finetuning is proposed in this work. In each epoch of end-to-end training, the process is divided into three steps, including optical flow registration, multiscale image decomposition and multiple transformation alignments. At each step, only the related networks are trained without interference of irrelevant information. Based on parameter-efficient finetuning on the foundation model, the proposed method achieves state-of-the-art performance on self-supervised depth estimation on SCARED dataset and zero-shot depth estimation on Hamlyn dataset, with 4% ∼ 10% lower error. The evaluation code of this work has been published on https://github.com/BaymaxShao/EndoMUST. Liangjing Shao, Linxin Bai, Chenkang Du, Xinrong Chen |
IROS | 4 |
| 2025 | SSD-Poser: Avatar Pose Estimation with State Space Duality from Sparse Observations
Shuting Zhao, Linxin Bai, Liangjing Shao, Ye Zhang 0015, Xinrong Chen |
ICMR | 5 |
| 2025 | Learn Concepts from Multi-Scale Visual Information for Compositional Zero-Shot LearningabstractCompositional Zero-Shot Learning (CZSL) aims at recognizing novel compositions by combining concepts learned from seen compositions. The key to tackle CZSL is disentangling highly coupled attribute-object compositions and learning exclusive concepts. Previous works mainly design networks to learn visual concepts from top-layer representations provided by visual backbones. As visual backbones progressively integrate information layer by layer, some low-level but critical information for concept learning may be lost, and the coupling between attribute and object features deepens. To address these issues, we propose to extract multi-scale visual features and fuse them in an adaptive way by Mixture of Experts (MoE) networks. We also employ feature-level similarity and a maximum entropy regularization term to constrain the model to effectively disentangle and learn concepts from multi-scale visual information. Comprehensive experiments on three CZSL benchmark datasets demonstrate that our method significantly outperforms previous SOTA methods in both closed-world and open-world settings. Guanyu Wang 0002, Zhijie Tan, Xu Chu 0001, Xinrong Chen, Tong Mo, Weiping Li 0002 |
MMAsia | 4 |
| 2025 | Visual sensing approach for high-precision pressure estimation in robotic manipulation
Linxin Bai, Xinrong Chen |
Eng. Appl. Artif. Intell. | 5 |
| 2025 | Vision-language foundation model for generalizable nasal disease diagnosis using unlabeled endoscopic recordsabstractMedical artificial intelligence (AI) holds significant potential in identifying signs of health conditions in nasal endoscopic images, thereby accelerating the diagnosis of diseases and systemic disorders. However, the performance of AI models heavily relies on expert annotations, and these models are usually task-specific with limited generalization performance across various clinical applications. In this paper, we introduce NasVLM, a Nasal Vision-Language foundation Model designed to extract universal representations from unlabeled nasal endoscopic data. Additionally, we construct a large-scale nasal endoscopic pre-training dataset and three downstream validation datasets from routine diagnostic records. The core strength of NasVLM lies in its ability to learn cross-modal semantic representations and perform multi-granular report-image alignment without depending on expert annotations. Furthermore, to the best of our knowledge, it is the first medical foundation model that effectively aligns medical report with multiple images of different anatomic regions, facilitated by a well-designed hierarchical report-supervised learning framework. The experimental results demonstrate that NasVLM has superior generalization performance across diverse diagnostic tasks and surpasses state-of-the-art self- and report-supervised methods in disease classification and lesion localization, especially in scenarios requiring label-efficient fine-tuning. For instance, NasVLM can distinguish normal nasopharynx (NOR) from abnormalities (benign hyperplasia, BH, and nasopharyngeal carcinoma, NPC) with an accuracy of 91.38% (95% CI, 90.59 to 92.17) and differentiate NPC from BH and NOR with an accuracy of 81.45% (95% CI, 80.21 to 82.67) on the multi-center NPC-Screen dataset using only 1% labeled data, on par with the performance of traditional supervised methods using 100% labeled data. Wentao Gong, Yinlong Liu, Xicai Sun, Xiaofeng Liu 0001, Xinrong Chen, Hongmeng Yu |
Pattern Recognit. | 10 |
| 2025 | Enhancing facial age estimation with local and global multi-attention mechanisms
Mingyan Qiu, Ziqun Zhang, Yinlong Liu, Xinrong Chen, Hongmeng Yu |
Pattern Recognit. Lett. | 8 |
| 2025 | Enhancing Motion Reconstruction From Sparse Tracking Inputs With Kinematic ConstraintsabstractIn virtual reality, there is a growing demand for reconstructing accurate full-body 3D avatars from the sparse motion captured through head-mounted displays and hand-held controllers. However, due to the limited information from sparse inputs, precisely reconstructing full body poses is an ill-posed and challenging task. Existing methods often exhibit notable errors in lower body poses, which results in unrealistic poses and occasional floor penetration artifacts. To address the above issue, a MLP-based model with Kinematic Constraints and Temporal Diversity (KCTD) was proposed for full body poses reconstruction, which incorporates Kinematic Constraints Hierarchical Decoder with Temporal Diversity Awareness Module and a Generative Feedback Module to further improve the accuracy of the reconstruction of the full body poses. Specifically, the potential constraints of human kinematic chain are incorporated into the model through a hierarchical decoder, which elevates overall precision through the interaction of the human kinematic chain. Then, a temporal diversity awareness module is integrated into the hierarchical decoder to help the model capture information at different frequency in the time domain. In addition, a generative feedback module is imposed on leg poses reconstruction to further improve its accuracy without increasing the model’s inference time. Test results on the AMASS dataset demonstrate that, the proposed model effectively improves the reconstruction accuracy of the full-body poses with the mean per joint rotation error and position error of of 2.60 and 3.62 respectively, which surpasses the state-of-the-art methods. Particularly, the proposed model can alleviate irrational poses in the lower body and reduce the floor penetration artifacts. Note to Practitioners—The motivation behind this manuscript stems from the necessity for precise full pose reconstruction in certain virtual reality applications. While existing methods have significantly contributed to the reconstruction of full-body poses based on sparse inputs, they have not taken tailored measures to improve accuracy in the lower body. Consequently, this paper introduces a full human body pose reconstruction model grounded in constraints from the human kinematic chain, featuring a temporal diversity awareness module. This model excels in reconstructing more accurate complete human body poses using sparse inputs from the head and hands, notably addressing issues such as irrational lower-body poses and floor-penetrating artifacts. Benefiting from state-of-the-art precision, the proposed approach has the potential to significantly elevate user experience and immersion in virtual reality applications, particularly in scenarios such as automated task training and simulation of hazardous conditions. Xiaokun Dai, Xinkang Zhang, Shiman Li, Xinrong Chen |
IEEE Trans Autom. Sci. Eng. | 4 |
| 2025 | EndoMODE: A Multimodal Visual Feature-Based Ego-Motion Estimation Framework for Monocular Odometry and Depth Estimation in Various Endoscopic ScenesabstractEgo-motion estimation is a critical task for both monocular odometry and depth estimation, which are significant for navigation and scene perception in endoscopy. Most of existing ego-motion estimation frameworks only utilize a single modality of visual feature. In this work, a novel multimodal visual feature-based framework is proposed to perform real-time vision-based ego-motion estimation for visual odometry and depth estimation in endoscopic scenes. In the framework, to extract correlation information of two adjacent endoscopic frames, a channel attention-based module is proposed to integrate multidimensional features from the concatenated image and a feature interaction module is proposed to fuse scene features from two observations. Moreover, a novel pose decoder based on depthwise separable convolution is proposed to extract multiscale feature representation. Furthermore, a fully supervised pipeline and a self-supervised pipeline with the proposed framework are designed for monocular odometry and depth estimation, respectively. In the experiments, the proposed framework is compared with several advanced methods on five different datasets of various endoscopic scenes. Both quantitative and qualitative results demonstrate the pipeline with the proposed framework provides the most accurate monocular odometry and depth estimation in different kinds of endoscopic scenes. Liangjing Shao, Benshuang Chen, Shuting Zhao, Xinrong Chen |
IEEE Trans. Ind. Informatics | 4 |
| 2024 | SPTESleepNet: Automatic Sleep Staging Model Based On Strip Patch Embeddings And Transformer EncoderabstractAlthough the research on automatic sleep staging has made great progress, there is still a certain distance from its clinical application. For a machine scoring system, in order to work in an interactive and collaborative manner with practitioners, two barriers need to be addressed, which are high accuracy and high efficiency. In this paper, we proposed a sleep staging model, named SPTESleepNet, as a stepping stone towards addressing the two above-mentioned obstacles. SPTESleepNet relies on the concept of Transformer, gives up the traditional convolution and recurrent methods, building a model based on Transformer framework with strip patch embedding. Overall the experimental results of SPTESleepNet on the databases showed that our model outperforms the previous sleep staging methods and improves the stateof-art results on these databases. On the larger database, Sleep-EDF-78, SPTESleep achieved an overall accuracy, macro F1-score, and Cohen’s kappa of 90.3%, 86.8% and 87.0%. Furthermore, on the smaller database, Sleep-EDF-20, SPTESleep surmounted the weakness that the Transformer models depend on large sample data to train, achieving an overall accuracy, macro F1-score, and Cohen’s kappa of 96.6%, 95.5%, 95.0%. Xiaokun Dai, Xinrong Chen |
ICASSP | 4 |
| 2024 | 3D Clothed Human Reconstruction From One In-the-Wild RGB ImageabstractIn recent years, much achievement have been made in the field of 3D clothed human reconstruction. However, most of researches performed not well for reconstruction from in-the-wild images due to the domain gap between the synthetic images of training datasets and the in-the-wild images. In this study, a modular model, including clothes encoder, body encoder and cloth generator, is proposed to perform 3D clothed human reconstruction from one single-view in-the-wild RGB image. In particular, we introduce the adaptive aggregation of convolution and multi-head attention into the cloth encoder and apply the adjustment of the segmentation at the preprocessing stage. According to experiments on MSCOCO and 3DPW datasets, the proposed method achieves state-of-the-art performance on 3D clothed human reconstruction from in-the-wild images compared with previous works. Liangjing Shao, Benshuang Chen, Xinrong Chen |
ICIP | 3 |
| 2024 | NETrack: A Lightweight Attention-Based Network for Real-Time Pose Tracking of Nasal Endoscope Based on Endoscopic Image
Liangjing Shao, Benshuang Chen, Xinrong Chen |
PRICAI (5) | 3 |
| 2024 | Hand-Object Pose Estimation and Reconstruction Based on Signed Distance Field and Multiscale Feature InteractionabstractThe study of reconstruction of hands and objects from color monocular images has garnered considerable attention in recent years. In existing methods, parametric models are constructed at single scale, and the interaction between hands and objects has not fully be explored. As a result, the multiscale information in 2D images cannot be fully exploited. At the same time, the lack of feature fusion and insufficient utilization of labels also have a great impact on the reconstruction accuracy. To address the limitations, a new framework is proposed, which comprises three key modules. Firstly, a multiscale feature extractor, which generates a multiscale representation of feature, is used to capture the interaction between hand and object more effectively. Secondly, a bridge based on attention has been used to establish the connection between hand and object representations, which facilitates the integration of them. Lastly, a module based on token merge is introduced into the framework, which provides the segmentation representation of object. The experimental results on two datasets, named Obman and DexYCB, demonstrated that the proposed method had good performance and achieved a shape error about 0.121$\text{cm}^{2}$on Obman and 0.40$\text{cm}^{2}$on DexYCB, outperforming the state-of-the-art methods. This study will probably provide the human-computer interaction methods with broader application prospects. Xinkang Zhang, Xiaokun Dai, Ziqun Zhang, Xinhan Di, Xinrong Chen |
IEEE Trans. Ind. Informatics | 5 |
| 2023 | A large group hesitant 2-tuple linguistic decision-making trial and evaluation laboratory (DEMATEL) method to evaluate performance indicators
Ziqun Zhang, Shanshan Dai, Yi Zhang 0015, Xinrong Chen |
Inf. Sci. | 5 |
| 2023 | A deep weakly semi-supervised framework for endoscopic lesion segmentation
Hong Wang 0021, Haoqin Ji, Yuexiang Li, Nanjun He, Dong Wei 0004, Yawen Huang, Xinrong Chen, Yefeng Zheng 0001, Hongmeng Yu |
Medical Image Anal. | 11 |
| 2022 | PointCLM: A Contrastive Learning-based Framework for Multi-instance Point Cloud Registration
Mingzhi Yuan, Qiuye Jin, Xinrong Chen, Manning Wang |
ECCV (9) | 4 |
| 2022 | Point Beyond Class: A Benchmark for Weakly Semi-supervised Abnormality Localization in Chest X-Rays
Haoqin Ji, Yuexiang Li, Jinheng Xie, Nanjun He, Yawen Huang, Dong Wei 0004, Xinrong Chen, LinLin Shen, Yefeng Zheng 0001 |
MICCAI (3) | 8 |
| 2020 | Independent Worker Selection In Crowdsourcing
Xueqi Li 0002, Xinrong Chen, Guojun Wang 0001 |
TrustCom | 4 |
| 2020 | Gesture Recognition Through sEMG with Wearable Device Based on Deep Learning
Shu Shen, Kang Gu, Xinrong Chen, Caixia Lv, Ruchuan Wang 0001 |
Mob. Networks Appl. | 3 |