VLDB 2026 Research / reviewers in the wild / expert
Jianxin Yang
dblp:242/4275
· DBLP profile ↗
22ranked-venue papers
4as first author
19since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 16 · 2 first-author · 13 since 2021Systems, architecture and hardware · 8 · 1 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 5 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SIM-Unet: A spatial interaction mamba network for multimodal MRI segmentation of multiple brain tumor types
Jianxin Yang, Mingchen Jiang, Chunna Yang, Shuailin You, Xiran Jiang |
Neurocomputing | 1 |
| 2026 | FBFL: Flexible Byzantine-Robust Federated Learning With Privacy-PreservingabstractPrivacy-preserving federated learning allows many clients to collaboratively train a machine learning model by sharing encrypted models. Among them, mask-based schemes have been widely applied due to their efficiency advantages. Unfortunately, as the invisibility of an individual local model, such schemes are vulnerable to poisoning attacks by Byzantine clients. Current work lacks a practical method to detect Byzantine clients within the mask-based scheme. We propose FBFL, a flexible Byzantine-robust federated learning scheme with privacy-preserving. While protecting the privacy of the client, we implement a defense against Byzantine clients without relying on an individual masked local model. Specifically, we design a secure distance computation method based on the Pedersen commitment. The exponential Manhattan distance we design is utilized to compute the malicious score of the client in order to distinguish the benign model from the abnormal model. Since the malicious score is obtained utilizing only the commitment value and not the masked local model, our Byzantine-robust method can be flexibly combined with other mask-based privacy-preserving methods. Security analysis shows that FBFL is Byzantine-robust and can ensure the data security of clients. Experimental evaluation shows that the robustness and efficiency of FBFL are great. Yanxin Xu, Hua Zhang 0001, Jianjin Zhao, Keke Gai, Zian Tian, Jianxin Yang |
IEEE Trans. Cloud Comput. | 6 |
| 2025 | Edge-Optimized Voice Control with 0.26 M Parameters: Distilling 86M Adaptive Window Audio Transformer for Real-World Variable-Length Inputs
Pinze Ren, Zhen Chen 0001, Yinjun Wu, Weiran Lin, Qilong Shi, Chao Li 0012, Jianxin Yang |
IEEE Big Data | 7 |
| 2025 | Towards Accurate Brain Electrode Implantation via Cross-modality Fusion of White-light and Photoacoustic MicroscopyabstractInvasive flexible neural electrodes are becoming increasingly prevalent in monitoring and modulating brain neural activity, necessitating the precise and minimally invasive implantation of these electrodes to a depth of a few millimeters beneath the cerebral surface. Although Neuralink has pioneered robot-assisted neural electrode implantation guided by microscopy, it currently lacks the ability to detect non-cerebral surface microvessels that are invisible under the white-light microscope, leading to inaccurate implantation planning and a high risk of trauma. To address this limitation, we introduce a vascular-enhanced strategy that fuses intraoperative white-light microscopy and preoperative photoacoustic microscopy and applies the fusion results to our established microsurgical robotic system for brain electrode implantation. Specifically, a multi-modality data preprocessing pipeline is devised to extract representative features, and a 2.5D fusion network that incorporates a depth encoding mechanism is proposed to predict cross-modality correspondence. The enhanced fusion results are utilized for implantation planning and intraoperative guidance during in vivo surgical procedures. Both quantitative and qualitative results are presented to demonstrate the effectiveness of our proposed cross-modality fusion methods. Furthermore, in vivo surgical implementations on mice underscore the potential of the proposed approach for achieving more precise and minimally invasive brain electrode implantation. Yuxuan Liu 0013, Yating Luo, Yunfei Luan, Xinyao Zhou, Jianxin Yang, Yao Guo 0002, Guang-Zhong Yang |
IROS | 5 |
| 2025 | Beyond the 80/20 Rule: High-Entropy Minority Tokens Drive Effective Reinforcement Learning for LLM ReasoningabstractReinforcement Learning with Verifiable Rewards (RLVR) has emerged as a powerful approach to enhancing the reasoning capabilities of Large Language Models (LLMs), yet its underlying mechanisms remain insufficiently understood. In this work, we undertake a pioneering exploration of RLVR through the novel perspective of token entropy patterns, comprehensively analyzing how different tokens influence reasoning performance. By examining token entropy patterns in Chain-of-Thought (CoT) reasoning, we observe that only a small fraction (approximately 20\%) of tokens exhibit high entropy, and these tokens semantically act as critical forks that steer the model toward diverse reasoning pathways. We further demonstrate that moderately increasing the entropy of these high-entropy tokens via decoding temperature adjustments leads to improved performance, quantitatively confirming their role as decision points in reasoning. We ultimately refine RLVR by restricting policy gradient updates to these forking tokens. Despite utilizing only 20\% of tokens, our approach achieves comparable performance to full-gradient updates on the Qwen3-8B base model. Moreover, it demonstrates remarkable improvements on the larger Qwen3-32B base model, boosting AIME'25 scores by 11.04 and AIME'24 scores by 7.71. In contrast, training exclusively on the 80\% lowest-entropy tokens leads to a marked decline in performance. These findings indicate that the efficacy of RLVR primarily arises from optimizing the high-entropy tokens that dictate key reasoning directions. Collectively, our results suggest promising avenues for optimizing RLVR algorithms by strategically leveraging the potential of these high-entropy minority tokens to further enhance the reasoning abilities of LLMs. Shenzhi Wang, Chujie Zheng, Rui Lu 0001, Kai Dang, Xiong-Hui Chen, Jianxin Yang, Zhenru Zhang, Yuqiong Liu, An Yang, Andrew Zhao, Shiji Song, Bowen Yu 0002, Gao Huang 0001, Junyang Lin |
NeurIPS | 9 |
| 2025 | MUSE: an open-access platform for urban expansion simulation with multitype patch generation engine and multilevel morphology regulationabstractUrban expansion models (UEMs) help unfold the mechanism, future pathways, and relevant repercussions of urban landscape dynamics. Although a plethora of UEMs have been developed, the field of open-access modeling platform for urban expansion simulation still presents gaps in terms of representing urban development events, regulating multilevel urban morphologies, and reflecting underlying regularities of physical urban growth. This study develops an open-access Multiengine Urban Expansion Simulation (MUSE) platform together with software that addresses some of these limitations. MUSE employs patch-based concepts and different patch generation engines to represent urban development events of distinctive types, enabling control over spatial morphology at landscape, class and patch levels. Moreover, MUSE incorporates two general regularities on urban development waves and urban diffusion and coalescence to govern physical urban expansion processes. Experiments in two cities examine MUSE’s ability in simulating realistic urban dynamics, and a series of synthetic experiments verify the operational behavior of the patch generation engines. These experiments show that MUSE can simulate a range of urban expansion types, including infill, edge growth and leapfrog expansion, and various morphologies, such as aggregated, sprawl, fragmented and compact. MUSE can empower practitioners to better project, analyze and understand complex urbanization dynamics. Jianxin Yang, Wenwu Tang, Jingruo Wang, Shengbing Yang, Lizhou Wang, Yinkun Chen |
Int. J. Geogr. Inf. Sci. | 1 |
| 2025 | PoseSDF++: Point Cloud-Based 3-D Human Pose Estimation via Implicit Neural RepresentationabstractPredicting accurate human pose from 3-D visual observation presents a formidable challenge in computer vision, with numerous applications across various industries. However, most existing studies tackled this issue by regressing the 3-D pose from depth maps via 2-D convolutional neural networks or parametric human models, with limited development in point cloud-based methods. To this end, we propose PoseSDF++, i.e., a point cloud-based encoder–decoder network utilizing implicit neural representation to perform 3-D human pose estimation (HPE) and nonparametric shape reconstruction simultaneously. Leveraging the representative capacity of the signed distance function (SDF), we conceptualize the 3-D HPE as a multiple-shape reconstruction task and propose a distance-aware regression method to accurately estimate the 3-D joint positions. In specific, our PoseSDF++ consists of three modules: first,a hierarchical encoderwith vector neuron layers extracts the multiscale rotation equivariant features from the point clouds captured from an arbitrary viewpoint, addressing the degradation issue caused by viewpoint variation of implicit representation; second,a shape decodermaps the extracted feature and the query to its corresponding shape SDF; third,a pose decodercomputes the distance between the query and the target keypoints, namely, the pose SDF. Extensive experiments on four publicly available datasets demonstrate that our PoseSDF++ achieves competitive performance against the state-of-the-art point cloud-based methods and covering the human hand (HANDS 2019), lower limbs (ICL-Gait), and full body (DFAUST, LiDARHuman2.6M) pose estimation. Jianxin Yang, Yuxuan Liu 0013, Xiao Gu 0003, Guang-Zhong Yang, Yao Guo 0002 |
IEEE Trans. Ind. Informatics | 1 |
| 2025 | Make PBR Materials Tileable With Latent Diffusion InpaintingabstractPhysically-based-rendering (PBR) materials are crucial in modern rendering pipelines, and many studies have focused on acquiring these materials from reality or images. However, existing methods may result in non-tileable results, since the realistic inputs usually have seams. Compared to non-tileable materials, tileable PBR materials have more universal application scenarios. To address this issue, we introduce MaTi, a novel pipeline that converts non-tileable PBR materials into tileable ones with minimal distortion. MaTi rearranges material patches to align boundaries at the center of the image, and then uses a diffusion model to inpaint the seams. We use scaled gamma correction to reduce the occurrence of collapse when processing special material maps. The color correction and triangular blending are adopt to preserve the original material information. Additionally, we design a division and blending strategy to efficiently handle high resolution materials. Our experiments demonstrate that MaTi can seamlessly modify PBR materials while preserving the original information, outperforming existing synthesis methods. Xiaoyu Zhan, Jianxin Yang, Jun Wang 0039, Yuanqi Li, Jie Guo 0001, Yanwen Guo 0001 |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2024 | Semantic Human Mesh Reconstruction with TexturesabstractThe field of 3D detailed human mesh reconstruction has made significant progress in recent years. However, current methods still face challenges when used in industrial applications due to unstable results, low-quality meshes, and a lack of UV unwrapping and skinning weights. In this paper, we present SHERT, a novel pipeline that can reconstruct semantic human meshes with textures and high-precision details. SHERT applies semantic- and normal-based sampling between the detailed surface (e.g. mesh and SDF) and the corresponding SMPL-X model to obtain a partially sampled semantic mesh and then generates the complete semantic mesh by our specifically designed self-supervised completion and refinement networks. Using the complete semantic mesh as a basis, we employ a texture diffusion model to create human textures that are driven by both images and texts. Our reconstructed meshes have stable UV unwrapping, high-quality triangle meshes, and consistent semantic information. The given SMPL-X model provides semantic information and shape priors, allowing SHERT to perform well even with incorrect and incomplete inputs. The semantic information also makes it easy to substitute and animate different body parts such as the face, body, and hands. Quantitative and qualitative experiments demonstrate that SHERT is capable of producing high-fidelity and robust semantic meshes that outperform state-of-the-art methods. Xiaoyu Zhan, Jianxin Yang, Yuanqi Li, Jie Guo 0001, Yanwen Guo 0001, Wenping Wang 0001 |
CVPR | 2 |
| 2024 | Skill Learning in Robot-Assisted Micro-Manipulation Through Human Demonstrations with Attention GuidanceabstractFor the development of robotic systems for micromanipulation, it is challenging to design appropriate control strategies due to either the lack of sufficient information for feedback or the difficulty in extracting subtle yet critical visual features. With the same system under the teleoperated mode, however, human operators seem to be able to complete the task more successfully with an inherent motion and control strategy. The extraction of implicit human attention during the task and integration of this with robot control could provide crucial guidance in the design of feature extraction and motion control algorithms. In this paper, a micro-assembly task of miniature thin membrane sensors is considered. For human demonstrations, we collected data from repeated tests performed by ten operators following three motion strategies. The human attention during the task is explored according to the coordinates of the eye gaze, and then a neural network with gaze-guided attention is trained to segment the visual Region of Interest (ROI). After quantitative evaluation of operator results in terms of success rate, efficiency, reset time, and the Index of Pupillary Activity (IPA), an optimized motion strategy based on the "palpation" framework was derived. Consequently, we apply this strategy to automated tasks and achieve superior results than human operators, showing an average task completion time of 34.8±5.9s and a success rate of over 90%. Yujian An, Jianxin Yang, Bingze He, Yao Guo 0002, Guang-Zhong Yang |
ICRA | 2 |
| 2024 | Intelligent Disinfection Robot with High-Touch Surface Detection and Dynamic Pedestrian AvoidanceabstractThe increasing awareness of public health issues has highlighted the need for effective disinfection of crowded indoor public areas, leading to the development of automated disinfection robots. However, most of the existing robots spray disinfectant in all areas, and they are still immature to navigate in densely populated environments. Hence, in this paper, we design a new disinfection robotic system consisting of a mobile platform, an RGB-D camera, and a robotic arm with a spray disinfection device. To address the above challenges, we propose a vision-based method for accurately detecting high-touch areas in the surroundings, enabling the disinfection robot to achieve superior disinfection efficiency. In addition, we propose a dynamic pedestrian avoidance method, namely Socially Aware APF (SA-APF), which can predict the movement trend of pedestrians and plan the path in real-time. Both simulated and real-world experiments are conducted to demonstrate the effectiveness of our disinfection robot system, especially highlighting the ability to detect high-touch areas and navigate in the environment while avoiding dynamic pedestrians. Yunfei Luan, Muhang He, Yudong Tian, Chengjie Lin, Yunhan Fang, Zihao Zhao 0005, Jianxin Yang, Yao Guo 0002 |
ICRA | 7 |
| 2023 | FashionSAP: Symbols and Attributes Prompt for Fine-Grained Fashion Vision-Language Pre-TrainingabstractFashion vision-language pre-training models have shown efficacy for a wide range of downstream tasks. However, general vision-language pre-training models pay less attention to fine-grained domain features, while these features are important in distinguishing the specific domain tasks from general tasks. We propose a method for fine-grained fashion vision-language pre-training based on fashion Symbols and Attributes Prompt (FashionSAP) to model fine-grained multi-modalities fashion attributes and characteristics. Firstly, we propose the fashion symbols, a novel abstract fashion concept layer, to represent different fashion items and to generalize various kinds of fine- grained fashion features, making modelling fine-grained attributes more effective. Secondly, the attributes prompt method is proposed to make the model learn specific attributes of fashion items explicitly. We design proper prompt templates according to the format of fashion data. Comprehensive experiments are conducted on two public fashion benchmarks, i.e., FashionGen and FashionIQ, and FashionSAP gets SOTA performances for four popular fashion tasks. The ablation study also shows the proposed abstract fashion symbols, and the attribute prompt method enables the model to acquire fine-grained semantics in the fashion domain effectively. The obvious performance gains from FashionSAP provide a new baseline for future fashion task research.11The source code is available at https://github.com/hssip/FashionSAP Yunpeng Han, Lisai Zhang, Qingcai Chen, Jianxin Yang, Zhao Cao |
CVPR | 6 |
| 2023 | EgoHMR: Egocentric Human Mesh Recovery via Hierarchical Latent Diffusion ModelabstractEgocentric vision has gained increasing popularity in social robotics, demonstrating great potentials for personal assistance and human-centric behavior analysis. Holistic per-ception of human body itself is a prerequisite for downstream applications, including action recognition and anticipation. Extensive research has been performed for human mesh recovery from the exocentric images captured from a third-person view, but limited studies are conducted for heavily distorted yet occluded egocentric images. In this paper, we propose Egocentric Human Mesh Recovery (EgoHMR), a novel hierarchical network based on latent diffusion models. Our method takes a single egocentric frame as the input and it can be trained in an end-to-end manner without supervision of 2D pose. The network is built upon the latent diffusion model by incorporating both global and local features in a hierarchical structure. To train the proposed network, we generate weak labels from synchronized exocentric images. The proposed method can perform human mesh recovery directly from egocentric images and detailed quantitative and qualitative experiments have been conducted to demonstrate the effectiveness of the proposed EgoHMR method. Yuxuan Liu 0013, Jianxin Yang, Xiao Gu 0003, Yao Guo 0002, Guang-Zhong Yang |
ICRA | 2 |
| 2023 | EasyGaze3D: Towards Effective and Flexible 3D Gaze Estimation from a Single RGB CameraabstractEye gaze can convey rich information of human intentions, which enables the social robots to comprehend the cognition and behavior of human targets. However, the existing 3D gaze estimation methods generally have high requirements either on the dedicated hardware or the quantity and quality of training databases, which largely limits their practical application values. This paper proposes EasyGaze3D, an effective 3D gaze estimation framework using a single RGB camera. First, the framework detects the 2D facial landmarks and recovers the 3D facial shape from the input image, and derives the required camera parameters with these features. Then, without loss of generality, the gaze direction can be regarded as the vector pointing from the eyeball center to the pupil center, which are derived respectively from the detected facial landmarks and the spherical fitting performed on the recovered 3D facial shape. Besides, we propose a flexible yet efficient calibration module, namely Easy-Cali, for deriving the subject-specific 3D facial shape and eyeball centers. The features calibrated by Easy-Cali can further boost the performance of EasyGaze3D. Experimental results show that our proposed method, being plug-and-play and without the need of training on large-scale dataset, can achieve superior performance against the existing methods based on deep models. Jianxin Yang, Yuxuan Liu 0013, Zhen Li 0026, Guang-Zhong Yang, Yao Guo 0002 |
IROS | 2 |
| 2023 | Adaptive one-stage generative adversarial network for unpaired image super-resolution
Ming-Wen Shao, Huan Liu 0012, Jianxin Yang, Feilong Cao |
Neural Comput. Appl. | 3 |
| 2023 | EgoFish3D: Egocentric 3D Pose Estimation From a Fisheye Camera via Self-Supervised LearningabstractEgocentric vision has gained increasing popularity recently, opening new avenues for human-centric applications. However, the use of the egocentric fisheye cameras allows wide angle coverage but image distortion is introduced along with strong human body self-occlusion imposing significant challenges in data processing and model reconstruction. Unlike previous work only leveraging synthetic data for model training, this paper presents a new real-world EgoCentric Human Pose (ECHP) dataset. To tackle the difficulty of collecting 3D ground truth using motion capture systems, we simultaneously collect images from a head-mounted egocentric fisheye camera as well as from two third-person-view cameras, circumventing the environmental restrictions. By using self-supervised learning under multi-view constraints, we propose a simple yet effective framework, namely EgoFish3D, for egocentric 3D pose estimation from a single image in different real-world scenarios. The proposed EgoFish3D incorporates three main modules. 1)The third-person-view moduletakes two exocentric images as input and estimates the 3D pose represented in the third-person camera frame; 2)the egocentric modulepredicts the 3D pose in the egocentric camera frame; and 3)the interactive moduleestimates the rotation matrix between the third-person and the egocentric views. Experimental results on our ECHP dataset and existing benchmark datasets demonstrate the effectiveness of the proposed EgoFish3D, which can achieve superior performance to existing methods. Yuxuan Liu 0013, Jianxin Yang, Xiao Gu 0003, Yao Guo 0002, Guang-Zhong Yang |
IEEE Trans. Multim. | 2 |
| 2022 | PoseSDF: Simultaneous 3D Human Shape Reconstruction and Gait Pose Estimation Using Signed Distance FunctionsabstractVision-based 3D human pose estimation and shape reconstruction play important roles in robot-assisted healthcare monitoring and personal assistance. However, 3D data captured from a single viewpoint always encounter occlusions and exhibit substantial heterogeneity across different views, resulting in significant challenges for both tasks. Extensive approaches have been proposed to perform each task separately, but few of them present a unified solution. In this paper, we propose a novel network based on signed distance functions, namely PoseSDF, to simultaneously reconstruct 3D lower limb shape and estimate gait pose by two dedicated branches. To promote multi-task learning, several strategies are developed to ensure that these two branches leverage the same latent shape code while exchanging information between them. More importantly, an auxiliary RotNet is incorporated into the inference phase, overcoming the inherent limitations of implicit neural functions under cross-view scenarios. Experimental results demonstrate that our proposed PoseSDF can achieve both high-quality shape reconstruction and precise pose estimation, generalizing well on the data from novel views, gait patterns, as well as real-world. Jianxin Yang, Yuxuan Liu 0013, Xiao Gu 0003, Guang-Zhong Yang, Yao Guo 0002 |
ICRA | 1 |
| 2022 | Ego+X: An Egocentric Vision System for Global 3D Human Pose Estimation and Social Interaction CharacterizationabstractEgocentric vision is an emerging topic, which has demonstrated great potential in assistive healthcare scenarios, ranging from human-centric behavior analysis to personal social assistance. Within this field, due to the heterogeneity of visual perception from first-person views, egocentric pose estimation is one of the most significant prerequisites for enabling various downstream applications. However, existing methods for egocentric pose estimation mainly focus on predicting the pose represented in the camera coordinates from a single image, which ignores the latent cues in the temporal domain and results in less accuracy. In this paper, we propose Ego+X, an egocentric vision based system for 3D canonical pose estimation and human-centric social interaction characterization. Our system is composed of two head-mounted egocentric cameras, where one is faced downwards and the other looks outwards. By leveraging the global context provided by visual SLAM, we first propose Ego-Glo for spatial-accurate and temporal-consistent egocentric 3D pose estimation in the canonical coordinate system. With the help of an egocentric camera looking outwards, we then propose Ego-Soc by extending Ego-Glo to various social interaction tasks, e.g., object detection and human-human interaction. Quantitative and qualitative experiments have been conducted to demonstrate the effectiveness of our proposed Ego+X. Yuxuan Liu 0013, Jianxin Yang, Xiao Gu 0003, Yao Guo 0002, Guang-Zhong Yang |
IROS | 2 |
| 2022 | An IBC Reference Block Enhancement Model Based on GAN for Screen Content Video Coding
Pengjian Yang, Jun Wang 0015, Guangyu Zhong, Pengyuan Zhang, Lai Zhang, Fan Liang 0001, Jianxin Yang |
MMM (2) | 7 |
| 2020 | A n-Gated Recurrent Unit with review for answer selection
Dongge Tang, Wenge Rong, Shuang Qin, Jianxin Yang, Zhang Xiong 0001 |
Neurocomputing | 4 |
| 2019 | Similarity Based Auxiliary Classifier for Named Entity RecognitionabstractShiyuan Xiao, Yuanxin Ouyang, Wenge Rong, Jianxin Yang, Zhang Xiong. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Shiyuan Xiao, Yuanxin Ouyang, Wenge Rong, Jianxin Yang, Zhang Xiong 0001 |
EMNLP/IJCNLP (1) | 4 |
| 2019 | Syntax Tree Aware Adversarial Question Rewriting for Answer SelectionabstractAnswer selection is an important method to achieve better user experience in question answering (QA) systems and it is essential in ensuring better QA matching performance. Improving mutual information between QA pairs is a useful way to obtain the matching degree improvement and question rewriting has been proven a helpful way in utilizing mutual interaction between questions and answers. In this research, we focus on syntax tree aware question rewriting inspired by the thought of integrating syntactic information into question answering. Besides, to improve the quality of rewriting, we employ the generative adversarial network for rewriting optimization, which consists of a syntax tree aware rewriting model and a discriminator. The quality information given by the discriminator guides the optimizing of the rewriting model in the training phase. The experimental study has shown the effectiveness of syntax tree aware question rewriting and utilizing the generative adversarial network for rewriting. Shuang Qin, Wenge Rong, Libin Shi, Jianxin Yang, Haodong Yang, Zhang Xiong 0001 |
IJCNN | 4 |