Yining Sun

dblp:53/8664 · DBLP profile ↗
← Back
19ranked-venue papers
1as first author
19since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 9 · 1 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 8 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 4 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Dual Graph Disambiguation for Multi-Instance Partial-Label Learning
abstract
In multi-instance partial label learning (MIPL), each sample is a bag of multiple instances linked to a candidate label set containing one true and multiple false labels, yielding inexact supervision in both instance features and label space. However, existing works adopt decoupled approaches that focus exclusively on either instance-level feature fusion or label-level disambiguation, failing to fully exploit the intrinsic dependencies between these two spaces. To overcome this limitation, graph-based methods are widely recognized as a powerful paradigm in weakly supervised learning, yet their success hinges on reliable features—precisely what MIPL lacks due to instance-level noise. To bridge this gap, we propose DualG, a novel framework that simultaneously addresses feature learning and label disambiguation through dual-level graph propagation. Specifically, we construct dual relevance graphs at both the bag and instance levels. At the bag level, we build a similarity graph based on fused feature representations; at the instance level, we employ attention scores to filter out irrelevant instances and construct a reliable instance-level relevance graph. These complementary graphs enable our joint label disambiguation framework to simultaneously address inexact supervision signals in both instance space and label space. Experimental results on five benchmark datasets demonstrate that DualG outperforms existing MIPL and partial label learning methods, validating its effectiveness and superiority.
Zhen Zhu 0007, Songhe Feng, Haobo Wang 0001, Gengyu Lyu, Cheng Peng 0011, Yining Sun
AAAI8
2026 The Paradox of Outcome Optimization: A Causal Information-Theoretic Bound on Reasoning Shortcuts in LLMs
abstract
Large Language Models (LLMs) aligned via outcome-based Reinforcement Learning (RL) frequently exhibit a critical failure mode: they achieve high performance on in-distribution benchmarks while demonstrating brittle reasoning capabilities on out-of-distribution (OOD) tasks.We term this phenomenon Reward-Induced Manifold Collapse.We establish a theoretical framework bridging Structural Causal Models (SCM) and the Information Bottleneck (IB) principle to explain this paradox.We define reasoning as a high-complexity causal process and shortcut learning as the exploitation of low-complexity spurious correlations.Under the implicit inductive bias of Stochastic Gradient Descent (SGD), models optimized for outcome rewards are biased toward shortcut solutions whenever the training distribution allows for a "Markovian Screening" of the true causal mechanism.We derive a new generalization bound based on Semantic Coverage Measure (η) rather than sample size, showing why data scaling on homogeneous distributions may fail to correct reasoning flaws.We also show that Process Reward Models (PRMs) function as Topological Filters, enforcing stepwise mutual information constraints that render the low-complexity shortcut manifold inadmissible.These results provide a mathematical grounding for the role of process supervision beyond simple credit assignment.
Wenxiang Geng, Zenghui Ding, Yining Sun
ACL (1)5
2026 AdaFedRec: Adaptive Heterogeneous Federated Recommender Systems Across Multi-Device Users
Zhenkai Li, Ming Hu 0003, Chentao Jia, Yining Sun, Zhufeng Lu, Yanxin Yang, Xiaofei Xie, Mingsong Chen 0001
ICDE4
2026 DyTimeNet: Dynamic cross-variable dependency network with sparse strategy for multivariate time series forecasting
Ting Chen 0009, Jinzhou Lai, Yining Sun, Wai Kin Chan
Neurocomputing3
2025 Recognizing Diabetic Foot Progression via Multimodal Fusion of Infrared Thermography and Clinical Data
abstract
Diabetic foot (DF), a severe diabetes complication, remains a leading cause of lower-limb amputations, underscoring the urgent need for early-stage detection to enable timely intervention and improve outcomes. However, current approaches often fail to distinguish diabetic patients without foot complications (DM) from those with DF, as single-modality infrared thermography (IRT) struggles with subtle thermal cues and lacks localization of pathological regions. In this paper, we propose DFP-MMNet (Diabetic Foot Progression MultiModal Network), a novel two-stage multimodal framework that addresses this challenge. In the first stage, a deep registration model, guided by corresponding RGB images, accurately localizes lesion Regions of Interest (ROIs) on thermograms. In the second stage, a dual-branch network fuses deep visual features from the ROI (via RegNetY-16GF) with semantic embeddings derived from structured clinical data, encoded using CLIP text branch. Extensive experiments on the collected ITC dataset demonstrate that DFP-MMNet achieves an F1-score of 82.7%, outperforming unimodal baselines by over 15%, offering a robust and accurate solution for early DF recognition.
Huarui Liu, Qiankun Li 0004, Xianjun Yang, Yining Sun
BIBM6
2025 ConcealGS: Concealing Invisible Copyright Information in 3D Gaussian Splatting
abstract
As 3D Gaussian Splatting (3D-GS) emerges as a promising technique for 3D reconstruction and novel view synthesis, offering superior rendering quality and efficiency, it becomes crucial to ensure secure transmission and copyright protection of 3D assets in anticipation of widespread distribution. While steganography has advanced significantly in common 3D media like meshes and Neural Radiance Fields (NeRF), research into steganography for 3D- GS representations remains largely unexplored. To address this gap, we propose ConcealGS, a novel 3D steganography method that embeds implicit information into the explicit 3D representation of Gaussian Splatting. By introducing a consistency strategy for the decoder and a gradient optimization approach, ConcealGS overcomes limitations of NeRF-based models, enhancing both the robustness of implicit information and the quality of 3D reconstruction. Extensive evaluations across various potential application scenarios demonstrate that ConcealGS successfully recovers implicit information with negligible impact on rendering quality, offering a groundbreaking approach for embedding invisible yet recoverable information into 3D models. This work paves the way for advanced copyright protection and secure data transmission in the evolving landscape of 3D content creation and distribution. Code is available at https://github.com/zxk1212/ConcealGS.
Hengyu Liu 0007, Chenxin Li, Yining Sun, Wuyang Li, Yifan Liu 0010, Yiyang Lin, Yixuan Yuan, Nanyang Ye 0001
ICASSP4
2025 Beyond Training: Dynamic Token Merging for Zero-Shot Video Understanding
abstract
Recent advancements in multimodal large language models (MLLMs) have opened new avenues for video understanding. However, achieving high fidelity in zero-shot video tasks remains challenging. Traditional video processing methods rely heavily on fine-tuning to capture nuanced spatial-temporal details, which incurs significant data and computation costs. In contrast, training-free approaches, though efficient, often lack robustness in preserving context-rich features across complex video content. To this end, we propose DYTO, a novel dynamic token merging framework for zero-shot video understanding that adaptively optimizes token efficiency while preserving crucial scene details. DYTO integrates a hierarchical frame selection and a bipartite token merging strategy to dynamically cluster key frames and selectively compress token sequences, striking a balance between computational efficiency with semantic richness. Extensive experiments across multiple benchmarks demonstrate the effectiveness of DYTO, achieving superior performance compared to both fine-tuned and training-free methods and setting a new state-of-the-art for zero-shot video understanding.
Zhuokai Zhao, Zhaorun Chen, Zenghui Ding, Xianjun Yang, Yining Sun
ICCV6
2025 RankCLIP: Ranking-Consistent Language-Image Pretraining
abstract
Self-supervised contrastive learning models, such as CLIP, have set new benchmarks for vision-language models in many downstream tasks. However, their dependency on rigid one-to-one mappings overlooks the complex and often multifaceted relationships between and within texts and images. To this end, we introduce RankCLIP, a novel pre-training method that extends beyond the rigid one-to-one matching framework of CLIP and its variants. By extending the traditional pair-wise loss to list-wise, and leveraging both in-modal and cross-modal ranking consistency, RankCLIP improves the alignment process, enabling it to capture the nuanced many-to-many relationships between and within each modality. Through comprehensive experiments, we demonstrate the effectiveness of RankCLIP in various downstream tasks, notably achieving significant gains in zero-shot classifications over state-of-the-art methods, underscoring the importance of this enhanced learning process.
Zhuokai Zhao, Zhaorun Chen, Zhili Feng, Zenghui Ding, Yining Sun
ICCV6
2025 Stock Price Prediction with Attention-Based Framework by Integrating LLM-Generated Features
Yining Sun, Penglei Gao, Yuyao Yan, Xi Yang 0008
ICONIP (5)1
2025 RBS-YOLO: A Lightweight YOLOv5-Based Surface Defect Detection Model for Castings
abstract
ABSTRACT To ensure precise and rapid identification of casting surface defects and to support the subsequent realisation of high‐precision grinding, this study introduces a method for detecting casting surface defects using a lightweight YOLOv5 framework. The enhanced model integrates the ShuffleNetV2 high‐efficiency CNN architecture into the YOLOv5 foundation, substantially reducing network parameters to achieve a lightweight model. Additionally, the Convolutional Block Attention Module (CBAM) attention mechanism is incorporated to enhance the model's capability to detect defects. The ReLU activation function replaces the SiLU function in the convolutional layer, decreasing the computational load and boosting efficiency. Subsequently, the optimised model is quantised and implemented on the RV1126 embedded development board, successfully performing image inference. To validate the effectiveness of the proposed method, a dataset of casting surface defects was designed and constructed. The optimised model has a file size of 7.6 MB, representing 55.4% of the original model, with about 50.6% of the original model's parameters. The onboard inference speed of the improved model is 50 ms per image, which is 9.1% faster than the traditional YOLOv5 model. These results offer valuable insights for future casting surface defect detection technologies.
Kezhu Wu, Shaoming Sun, Yining Sun, Cunyi Wang
IET Image Process.3
2025 Driving Risk Assessment for Intelligent Vehicles Based on Entropy-Informed Graph Neural Networks and Gaussian Distributions
abstract
This study proposes a novel framework based on an entropy-informed graph neural network (EIGNN) integrated with Gaussian distribution (GD) to assess the driving risk of intelligent vehicles in typical traffic scenarios. Existing research often overlooks comprehensive spatiotemporal modeling of vehicle interaction characteristics and the quantification of uncertainty in dynamic risk assessments. In this work, vehicle speed and acceleration are probabilistically modeled using GD, while entropy theory is introduced to quantify risk uncertainty. A risk assessment model based on graph neural networks (GNNs) is then designed to capture the spatiotemporal dynamics of multivehicle interactions and predict the potential risk levels of driving strategies. The results demonstrate that the framework accurately quantifies collision risks in multivehicle interactions in complex traffic scenarios, with high accuracy and robustness across typical situations such as cruising, cut-ins, lane changes, overtaking, and different density traffic. By thoroughly analyzing traffic risk characteristics and incorporating them into intelligent driving decision-making, this study provides significant technical insights and theoretical support for enhancing the safety and decision-making efficiency of autonomous driving systems.
Hongbo Gao 0001, Chengbo Wang 0001, Runda Niu, Xiaozhao Fang, Jinpeng Chen 0001, Yining Sun, Huiqing Jin, Danwei Wang
IEEE Trans. Neural Networks Learn. Syst.6
2024 Video-text retrieval via multi-modal masked transformer and adaptive attribute-aware graph convolutional network
Yining Sun, Fudong Nian
Multim. Syst.2
2023 VQA-CLPR: Turning a Visual Question Answering Model into a Chinese License Plate Recognizer
Xuhao Jiang, Yining Sun, Weiya Ni, Fudong Nian
ICIG (2)3
2023 Information Bottleneck Revisited: Posterior Probability Perspective with Optimal Transport
abstract
Information bottleneck (IB) is a paradigm to extract information in one target random variable from another relevant random variable, which has aroused great interest due to its potential to explain deep neural networks in terms of information compression and prediction. Despite its great importance, finding the optimal bottleneck variable involves a difficult nonconvex optimization problem due to the nonconvexity of mutual information constraint. The Blahut-Arimoto algorithm and its variants provide an approach by considering its Lagrangian with fixed Lagrange multiplier. However, only the strictly concave IB curve can be fully obtained by the BA algorithm, which strongly limits its application in machine learning and related fields, as strict concavity cannot be guaranteed in those problems. To overcome the above difficulty, we derive an entropy regularized optimal transport (OT) model for IB problem from a posterior probability perspective. Correspondingly, we use the alternating optimization procedure and generalize the Sinkhorn algorithm to solve the above OT model. The effectiveness and efficiency of our approach are demonstrated via numerical experiments.
Lingyi Chen, Shitong Wu, Wenhao Ye, Huihui Wu, Hao Wu 0060, Wenyi Zhang 0001, Bo Bai 0001, Yining Sun
ISIT8
2023 COME: Clip-OCR and Master ObjEct for text image captioning
Yining Sun, Fudong Nian, Maofei Zhu, Wenliang Tang
Image Vis. Comput.2
2023 Hierarchical cross-modal contextual attention network for visual grounding
Yining Sun, Yuxia Hu, Fudong Nian
Multim. Syst.3
2022 Stitching High Resolution Notebook Keyboard Surface Based on Halcon Calibration
Zuchang Ma, Yining Sun, Fudong Nian
ICIC (1)4
2022 Automatic Angle's classification based on the occlusal contact information
abstract
Malocclusion has a high prevalence in the population, which seriously affects patients’ oral and mental health. Angle’s classification is a widely accepted diagnostic standard for malocclusion, either requiring professional intervention and complicated procedures, or increasing radiation risks. This paper proposes a new method of Angle’s classification based on occlusal contact information to realize the automatic Angle’s classification. Firstly, a novel bite force measurement device is used to record the occlusal data of subjects with different occlusal categories, Meta-analysis evaluated several occlusion quantitative evaluation indicators. Then, the imbalance of the data set is improved by oversampling and popular machine learning models are used for training and performance evaluation. The result shows that the accuracy of the random forest model combined with occlusal contact information reaches 87.83%, and the performance of other evaluation indexes is good. It is demonstrated that machine learning models can be applied to Angle’s classification and shows the great potential of occlusal contact information in the aided diagnosis of oral diseases.
Zhiming Yao, Xianjun Yang, Yuanyin Wang, Wenhua Xu, Yining Sun
SMC8
2021 A Hybrid Model for Named Entity Recognition on Chinese Electronic Medical Records
abstract
Electronic medical records (EMRs) contain valuable information about the patients, such as clinical symptoms, diagnostic results, and medications. Named entity recognition (NER) aims to recognize entities from unstructured text, which is the initial step toward the semantic understanding of the EMRs. Extracting medical information from Chinese EMRs could be a more complicated task because of the difference between English and Chinese. Some researchers have noticed the importance of Chinese NER and used the recurrent neural network or convolutional neural network (CNN) to deal with this task. However, it is interesting to know whether the performance could be improved if the advantages of the RNN and CNN can be both utilized. Moreover, RoBERTa-WWM, as a pre-training model, can generate the embeddings with word-level features, which is more suitable for Chinese NER compared with Word2Vec. In this article, we propose a hybrid model. This model first obtains the entities identified by bidirectional long short-term memory and CNN, respectively, and then uses two hybrid strategies to output the final results relying on these entities. We also conduct experiments on raw medical records from real hospitals. This dataset is provided by the China Conference on Knowledge Graph and Semantic Computing in 2019 (CCKS 2019). Results demonstrate that the hybrid model can improve performance significantly.
Yu Wang 0121, Yining Sun, Zuchang Ma, Lisheng Gao, Yang Xu 0041
ACM Trans. Asian Low Resour. Lang. Inf. Process.2