Jiayuan Sun

dblp:279/4138 · DBLP profile ↗
← Back
11ranked-venue papers
1as first author
11since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 9 · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 first-author · 4 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Multi-view Invariance Learning for 3D Scene Graph Pre-training via Collaborative Cross-Modal Regularization
abstract
3D scene graph generation is a pivotal task in scene understanding. Its performance is easy to be constrained by the limited availability of annotated data. Currently, the existing solutions on point cloud pre-training usually emphasize on object-centric representations while neglecting the predicate feature learning. This limitation significantly hinders their relational reasoning capabilities, as inter-object relationships are fundamentally governed by predicate features. To enhance 3D Scene Graphs Pre-training, this paper proposes a task-specific Multi-view Invariance Learning framework with Collaborative Cross-modal Regularization. In detail, the inherent horizontal-rotation invariance of 3D objects and their semantic relationships are leveraged to construct a self-supervised paradigm for triplet feature learning. Moreover, our framework harnesses the cross-modal prior knowledge from the vision-language model to regularize model optimization. It could further achieve the semantic discrimination via unsupervised deep clustering. To resolve the knowledge discrepancies arising from the pre-trained model in fine-tuning, a predicate adapter equipped with knowledge filtering gate is devised to selectively aggregate the predicate features of pre-trained model. Extensive experiments demonstrate that our framework is effective in boosting 3D scene graph generation performance, surpassing state-of-the-art ones.
Luping Ji, Ruijie Xiao, Jiayuan Sun
AAAI4
2025 Opposing Stance in Topic Evolution: A Case Analysis of Messi's Visit to Hong Kong
abstract
In February 2024, Lionel Messi's absence from a Hong Kong exhibition match ignited extensive debate across social media platforms, capturing the attention of the public and sparking controversies that extended into realms such as business partnerships and diplomatic implications. This incident not only reflects the public's reaction and pattern of stance changes toward such a complex Incident but also significantly demonstrates the close connection between cyberspace and the real world. Research on this incident has important practical significance for predicting and intervening in the evolution of similar hot incidents. In this article, the case of Messi's absence from the Hong Kong match is delved into. A stance detection method and a quantification approach for stance divergence have been devised to analyze the evolution of the stance surrounding the incident. Furthermore, by examining topic clusters over time, the topical progression of public stances is tracked. Findings reveal that nearly 60% of posts expressed an explicit stance throughout the incident, with the majority (66.9%) taking an opposing stance. Additionally, it was noted that the topics discussed at various stages followed a long-tail distribution, indicating that most discussions revolved around a few dominant themes. Within more segmented and specific topics, stance divergences were often more prominent.
Tao Wang 0172, Yuanhan Xie, Jiayuan Sun, Lifang Li, Weishan Zhang, Fei-Yue Wang 0001
IEEE Trans. Comput. Soc. Syst.4
2024 Hype or Revolution: How GPT4, Claude3 Opus, and GPT3.5 Exaggerate Their Influence on HRM Occupations Exposed in Practice
abstract
This study investigates the actual impact of ChatGPT on human resource management (HRM) occupations and compares it with assessments from large language models (LLMs). Using data from the Occupational Information Network (O*NET), we developed a novel methodology combining real-world evidence from Google search results with evaluations from GPT-4, GPT-3.5, and Claude 3 Opus. Our findings reveal significant variations in the degree of exposure among HRM occupations. Training and Development Managers were the most exposed, with 86.52% of their tasks potentially automated by ChatGPT, while Labor Relations Specialists were the least exposed at 11.08%. Notably, LLM assessments consistently overestimated the exposure levels by an average of 30% compared to real-world evidence. This discrepancy highlights the importance of critically evaluating LLM capabilities in task automation. Our research contributes to understanding AI's impact on HRM practices and provides valuable insights for professionals adapting to the AI era. It also underscores the need for caution when relying solely on LLM assessments for workforce planning and development.
Yuanhan Xie, Jiayuan Sun, Dayong Shen, Zhongshan Zhang
SMC3
2024 Shared Coupling-Bridge Scheme for Weakly Supervised Local Feature Learning
abstract
Local feature learning is believed to be of important significance in classic vision tasks such as visual localization, image matching and 3D reconstruction. Limited by training samples, weakly-supervised strategy has become one of widely-concerned effective schemes for local feature learning. Currently, it still has some weaknesses needing further improvement, mainly including the discrimination power of extracted local descriptors, the localization accuracy of detected keypoints, and the efficiency of weakly-supervised local feature learning. Focusing on promoting the performance of sparse local feature learning with camera pose supervision, this article pertinently proposes a Shared Coupling-bridge scheme with four light-weight yet effective improvements for weakly-supervised local feature (SCFeat) learning. It mainly contains: i) theFeature-Fusion-ResUNet Backbone(F2R-Backbone) for local descriptors learning, ii) a shared coupling-bridge normalization to improve the decoupling training of description network and detection network, iii) an improved detection network with peakiness measurement to detect keypoints and iv) a new reward factor of fundamental matrix error to further optimize feature detection training. Extensive experiments prove that our SCFeat scheme is effective and has wide task adaptability. It could often obtain a state-of-the-art performance on classic image matching and visual localization. Even in terms of 3D reconstruction, it could still achieve competitive results.
Jiayuan Sun, Luping Ji, Jiewen Zhu
IEEE Trans. Multim.1
2023 AirwayFormer: Structure-Aware Boundary-Adaptive Transformers for Airway Anatomical Labeling
Weihao Yu 0004, Hao Zheng 0008, Yun Gu, Fangfang Xie, Jiayuan Sun, Jie Yang 0002
MICCAI (7)5
2023 Multi-site, Multi-domain Airway Tree Modeling
Yangqian Wu, Yulei Qin, Hao Zheng 0008, Wen Tang 0005, Corey W. Arnold, Chenhao Pei, Pengxin Yu, Yang Nan 0002, Guang Yang 0006, Simon Walsh, Dominic C. Marshall, Matthieu Komorowski, Puyang Wang, Dazhou Guo, Dakai Jin, Shuiqing Zhao, Runsheng Chang, Abdul Qayyum 0002, Moona Mazher, Yonghuang Wu, Ying'ao Liu, Jiancheng Yang, Ashkan Pakzad, Bojidar Rangelov, Raúl San José Estépar, Carlos Cano-Espinosa, Jiayuan Sun, Guang-Zhong Yang, Yun Gu
Medical Image Anal.34
2023 Automatic Representative Frame Selection and Intrathoracic Lymph Node Diagnosis With Endobronchial Ultrasound Elastography Videos
abstract
Endobronchial ultrasound (EBUS) elastography videos have shown great potential to supplement intrathoracic lymph node diagnosis. However, it is laborious and subjective for the specialists to select the representative frames from the tedious videos and make a diagnosis, and there lacks a framework for automatic representative frame selection and diagnosis. To this end, we propose a novel deep learning framework that achieves reliable diagnosis by explicitly selecting sparse representative frames and guaranteeing the invariance of diagnostic results to the permutations of video frames. Specifically, we develop a differentiable sparse graph attention mechanism that jointly considers frame-level features and the interactions across frames to select sparse representative frames and exclude disturbed frames. Furthermore, instead of adopting deep learning-based frame-level features, we introduce the normalized color histogram that considers the domain knowledge of EBUS elastography images and achieves superior performance. To our best knowledge, the proposed framework is the first to simultaneously achieve automatic representative frame selection and diagnosis with EBUS elastography videos. Experimental results demonstrate that it achieves an average accuracy of 81.29% and area under the receiver operating characteristic curve (AUC) of 0.8749 on the collected dataset of 727 EBUS elastography videos, which is comparable to the performance of the expert-based clinical methods based on manually-selected representative frames.
Mingxing Xu, Junxiang Chen, Jin Li 0057, Xinxin Zhi, Wenrui Dai, Jiayuan Sun, Hongkai Xiong
IEEE J. Biomed. Health Informatics6
2023 TNN: Tree Neural Network for Airway Anatomical Labeling
abstract
Detailed anatomical labeling of bronchial trees extracted from CT images can be used as fine-grained maps for intra-operative navigation. To cater to the sparse distribution of airway voxels and large class imbalance in 3D image space, a graph-neural-network-based method is proposed to map branches to nodes in a graph space and assign anatomical labels down to subsegmental level. To address the inherent problem of overlapping distribution of positional and morphological features, especially for subsegmental categories, the proposed method focuses on the relative position between sibling subsegments which is fixed in most cases. The hierarchical nomenclature is represented by multi-level labeling and each category is associated with one or two subtrees in the graph. Hyperedges are used to extract the representation of subtrees while a hypergraph neural network is developed to encode their intrinsic relationship through hyperedge interaction. A filter module is further designed to guide feature aggregation between nodes and hyperedges. With the proposed method, the final accuracies for segmental and subsegmental node classification can achieve 93.6% and 82.0% respectively. The corresponding code is publicly available at https://github.com/haozheng-sjtu/airway-labeling.
Weihao Yu 0004, Hao Zheng 0008, Yun Gu, Fangfang Xie, Jie Yang 0002, Jiayuan Sun, Guang-Zhong Yang
IEEE Trans. Medical Imaging6
2022 Vision-Kinematics Interaction for Robotic-Assisted Bronchoscopy Navigation
abstract
Endobronchial intervention is increasingly used as a minimally invasive means for the treatment of pulmonary diseases. In order to acquire the position of bronchoscopy, vision-based localization approaches are clinically preferable but are sensitive to visual variations. The static nature of pre-operative planning makes mapping of intraoperative anatomical features challenging for learning-based methods using visual features alone. In this work, we propose a robust navigation framework based on Vision Kinematic Interaction (VKI) for monocular bronchoscopic videos. To address visual-imbalance between the virtual and real views of bronchoscopy images, a Visual Similarity Network (VSN) is proposed to extract domain-invariant features to represent the lumen structure from endoscopic views, as well as domain-specific features to characterize the surface texture and visual artefacts. To improve the robustness of online estimation of camera pose, we also introduce a Kinematic Refinement Network (KRN) that allows progressive refinement of camera pose estimation based on network prediction and robot control signals. The accuracy of camera localization is validated on phantom and porcine lung datasets from a robotically controlled endobronchial intervention system, with both quantitative and qualitative results demonstrating the performance of the techniques. Results show that the features extracted by the proposed method can preserve the structural information of small airways in the presence of large visual variations along with the much-improved camera localization accuracy. The absolute trajectory errors (ATE) on phantom data and porcine data are 8.01 mm and 8.62 mm respectively.
Yun Gu, Chuanjia Gu, Jie Yang 0002, Jiayuan Sun, Guang-Zhong Yang
IEEE Trans. Medical Imaging4
2021 Refined Local-imbalance-based Weight for Airway Segmentation in CT
Hao Zheng 0008, Yulei Qin, Yun Gu, Fangfang Xie, Jiayuan Sun, Jie Yang 0002, Guang-Zhong Yang
MICCAI (1)5
2021 Alleviating Class-Wise Gradient Imbalance for Pulmonary Airway Segmentation
abstract
Automated airway segmentation is a prerequisite for pre-operative diagnosis and intra-operative navigation for pulmonary intervention. Due to the small size and scattered spatial distribution of peripheral bronchi, this is hampered by a severe class imbalance between foreground and background regions, which makes it challenging for CNN-based methods to parse distal small airways. In this paper, we demonstrate that this problem is arisen by gradient erosion and dilation of the neighborhood voxels. During back-propagation, if the ratio of the foreground gradient to background gradient is small while the class imbalance is local, the foreground gradients can be eroded by their neighborhoods. This process cumulatively increases the noise information included in the gradient flow from top layers to the bottom ones, limiting the learning of small structures in CNNs. To alleviate this problem, we use group supervision and the corresponding WingsNet to provide complementary gradient flows to enhance the training of shallow layers. To further address the intra-class imbalance between large and small airways, we design a General Union loss function that obviates the impact of airway size by distance-based weights and adaptively tunes the gradient ratio based on the learning process. Extensive experiments on public datasets demonstrate that the proposed method can predict the airway structures with higher accuracy and better morphological completeness than the baselines.
Hao Zheng 0008, Yulei Qin, Yun Gu, Fangfang Xie, Jie Yang 0002, Jiayuan Sun, Guang-Zhong Yang
IEEE Trans. Medical Imaging6