Xiao Yun

dblp:23/1937 · DBLP profile ↗
← Back
20ranked-venue papers
11as first author
11since 2021 · last 2026
0000-0002-1538-5279ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 9 · 6 first-author · 6 since 2021Systems, architecture and hardware · 5 · 3 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 1 since 2021Computer networks · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
Video understanding and tracking · 34% Transfer learning and domain adaptation · 25% Trustworthy machine learning · 22%
Computer graphics and multimedia
1 paper
Image and video processing · 100%

Topics — the 12 heaviest of 12, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Graph learning › graph neural network
graph convolutional network
0.812024
Behavioral Recognition of Skeletal Data Based on Targeted Dual Fusion Strategy · AAAI 2024
Computer vision › Video understanding and tracking
multi-stream fusion
0.812024
Behavioral Recognition of Skeletal Data Based on Targeted Dual Fusion Strategy · AAAI 2024
Computer vision › Video understanding and tracking › action recognition
skeleton-based action recognition
0.812024
Behavioral Recognition of Skeletal Data Based on Targeted Dual Fusion Strategy · AAAI 2024
Machine learning › Transfer learning and domain adaptation
few-shot learning
0.612022
Robust Meta-learning with Sampling Noise and Label Noise via Eigen-Reptile · ICML 2022
Machine learning › Trustworthy machine learning › robustness
learning with noisy labels
0.612022
Robust Meta-learning with Sampling Noise and Label Noise via Eigen-Reptile · ICML 2022
Machine learning › Transfer learning and domain adaptation
meta-learning
0.612022
Robust Meta-learning with Sampling Noise and Label Noise via Eigen-Reptile · ICML 2022
Machine learning › Trustworthy machine learning
robustness
0.612022
Robust Meta-learning with Sampling Noise and Label Noise via Eigen-Reptile · ICML 2022
Machine learning › Deep learning architectures and training
attention mechanism
0.212024
Behavioral Recognition of Skeletal Data Based on Targeted Dual Fusion Strategy · AAAI 2024
Machine learning › Transfer learning and domain adaptation › meta-learning
gradient-based meta-learning
0.212022
Robust Meta-learning with Sampling Noise and Label Noise via Eigen-Reptile · ICML 2022
Computer vision › Video understanding and tracking › object tracking › kernel-based tracking
mean-shift tracking
0.112012
A multi-cue mean-shift target tracking approach based on fuzzified region dynamic image fusion · Sci. China Inf. Sci. 2012
Computer vision › Video understanding and tracking
object tracking
0.112012
A multi-cue mean-shift target tracking approach based on fuzzified region dynamic image fusion · Sci. China Inf. Sci. 2012
Image and video processing
image fusion
0.012012
A multi-cue mean-shift target tracking approach based on fuzzified region dynamic image fusion · Sci. China Inf. Sci. 2012

Methods — techniques the papers use, named apart from their topics

multi-stream fusion · 0.8graph convolutional network · 0.8attention mechanism · 0.8meta-gradient · 0.6introspective self-paced learning · 0.6eigen-reptile · 0.6mean shift · 0.3fuzzy region dynamic image fusion · 0.3
YearPublicationVenuePosition
2026 JIMI: A Hierarchical Partition-Refinement MIQP Framework for Jitter Minimization in PONs
Lihe Liang, Xiao Yun, Kuangxun Huang
ICC2
2026 Unsupervised video stitching via spatiotemporal coupling with temporal propagation and dynamic memory
abstract
Video stitching poses a fundamental challenge: while video frames exhibit strong spatiotemporal coupling, prevailing decoupled paradigms artificially separate spatial and temporal information, leading to limited precision and severe efficiency bottlenecks. Consequently, achieving a joint spatiotemporal optimization remains a critical, unresolved objective in this field. To overcome this, we propose an unsupervised video stitching framework that fundamentally reconceptualizes the task as a joint spatiotemporal state estimation problem, driven by temporal propagation and a dynamic memory mechanism. Building upon spatiotemporally decoupled multi-view inputs, we design a Spatially-Aware Gated Recurrent Unit (SA-GRU) structure that mines intrinsic spatiotemporal feature correlations to establish cross-frame temporal propagation and spatiotemporal mapping, thereby compensating for missing temporal information and improving alignment. To overcome the perceptual constraints of sliding windows, we introduce a Temporally-Gated Recurrent Unit (T-GRU) structure integrated with a Spatiotemporal Attention Fusion (STA-Fusion) module, which adaptively integrates long-term historical context with current spatiotemporal constraints to smooth stitching trajectories and enable long-term stabilization with a minimal memory footprint. Evaluations on the Stabstitch-D dataset demonstrate that our method outperforms existing approaches in alignment and geometric fidelity, achieving a PSNR of 31.18, an SSIM of 0.907, and a distortion measure of 0.011. Compared to traditional sliding-window approaches, our framework drastically reduces the memory footprint.
Xiao Yun, Fanfei Yu, Kaiwen Dong, Yanjing Sun
Knowl. Based Syst.1
2026 Temporal Consistency and Variation-Guided Spatio-Temporal Aggregation for Few-Shot Action Recognition
abstract
Few-shot Action Recognition (FSAR) aims to recognize novel actions from only a few labeled examples, posing challenges due to limited supervision and complex temporal dynamics. Existing methods often adopt a unified motion modeling strategy for both short- and long-term dynamics, overlooking the need to adapt motion pattern extraction to the specific temporal properties inherent to different timescales. This forces models to hedge against multi-scale relevance through exhaustive searches over temporal tuples, followed by heavy spatio-temporal fusion, which substantially increases parameters and computation and ultimately limits efficiency. To this end, we propose the efficient Temporal Consistency and Variation-Guided Spatio-Temporal Aggregation Network (TCV-STA), which comprises four key components: the Temporal Consistency Module (TCM), the Temporal Variation Module (TVM), the Spatio-Temporal Aggregation attention (STA), and the Shifted Window Temporal Attention (SWTA). The TCM captures stable motion patterns to suppress short-term perturbations and enhance temporal consistency for robust motion representation, while the TVM models dynamic motion patterns to highlight long-term variations that improve inter-class discriminability and facilitate intra-class alignment. Built upon these complementary motion cues, the STA selectively aggregates spatial and temporal representations under the guidance of the learned stable and dynamic motion patterns, avoiding global dense fusion. Finally, to address the limited receptive field and discontinuous modeling caused by frame grouping in TCM and TVM, we adapt a SWTA to capture longer-range temporal dependencies and ensure smooth transitions across subaction segments for few-shot action recognition. Experiments demonstrate that TCV-STA achieves competitive accuracy across four widely-used FSAR benchmarks while reducing parameters by up to 27.9% and computational cost by 21.3%, striking a favorable balance between accuracy and efficiency for deployment in resource-constrained scenarios.
Kaiwen Dong, Quanyi Li, Yanjing Sun, Xiao Yun, Yu Zhou 0009, Kévin Riou, Xiaofeng Hou, Patrick Le Callet
IEEE Trans. Circuits Syst. Video Technol.4
2025 DARE: Towards Maximized Bandwidth Utilization of Dynamic Bandwidth Allocation for Next-Generation Passive Optical Networks
abstract
Dynamic Bandwidth Allocation (DBA), which allocates the shared bandwidth among multiple competing network devices at runtime, serves as one of the most fundamental components in Passive Optical Networks (PONs). However, with the evolution of PON standards and the growing demand for high bandwidth, conventional DBA schemes suffer from bandwidth under-utilization and also fail to meet stringent throughput requirements. To overcome these challenges, in this paper, we propose DARE, a high-performance DBA scheme that aims to maximize bandwidth utilization for next-generation PONs. DARE introduces three novel mechanisms: dynamic adjustment of service interval, idle bandwidth reallocation, and grant fragment reuse. (1) The dynamic adjustment of service interval adaptively changes the polling interval based on the real-time workload characteristics, reducing protocol elements and increasing throughput. (2) The idle bandwidth reallocation scheme makes full use of unexploited bandwidth to jointly enhance bandwidth utilization and response speed for bursty traffic. (3) The grant fragment reuse strategy assembles and reassigns the cropped bandwidth fragments induced by the alignment of grant granularity, allowing more precise mapping of grants to actual bandwidth requests. Experimental results demonstrate that DARE achieves 98.4% of the theoretical bound of maximum bandwidth utilization. Moreover, DARE outperforms state-of-the-art DBA algorithms in three key aspects: throughput improvements of$1.14 \times -1.32 \times$across various load conditions, maximum bandwidth utilization increased by$1.07 \times -1.85 \times$, and faster response speed with average delay reduced by$450 \mu$s under bursty traffic.
Kuangxun Huang, Xiao Yun, Lihe Liang
ICPADS2
2025 Advancing toxicity AI-based prediction with multilevel systems biology: a case study on genotoxicity
abstract
The rapid expansion of chemical diversity presents substantial challenges for health and environmental risk assessment, necessitating the development of alternative, high-throughput computational methodologies. A key hurdle in toxicity prediction lies in the heterogeneous nature of adverse health outcomes at the tissue and cellular levels, as biological processes exhibit cell-type-specific and context-dependent responses. Effective prediction of individual-level health effects thus requires the integration of multimodal data, capturing both structural and biological perturbations induced by chemical exposures. We present GenotoxNet, a multimodal deep learning framework that enhances genotoxicity prediction by systematically integrating chemical structures, high-throughput in vitro assay data, and transcriptomics data. By leveraging this multimodal integration, GenotoxNet effectively captures cellular heterogeneity and mechanistic complexity, enabling more comprehensive evaluation of chemical-induced genotoxicity. The model outperformed single-modality approaches, achieving AUCROC of 0.891 ± 0.017 on the internal test set, demonstrating superior predictive capability over models relying solely on chemical structures or individual biological features. The model still performed well on the external chemical set. Beyond classification, GenotoxNet facilitates mechanistic interpretation by aligning multimodal feature representations of genotoxic chemicals with adverse outcome pathway (AOP). This framework not only offers a robust approach for predicting genotoxicity but also aids in the development of preventive strategies and regulatory decisions aimed at mitigating the health risks posed by hazardous chemicals.
Huazhou Zhang, Xiao Yun, Wenxiao Pan, Qiao Xue, Jianjie Fu, Aiqian Zhang
Briefings Bioinform.3
2025 Heterogeneous modal collaborative training network for human action recognition
Xiao Yun, Yanjing Sun, Kaiwen Dong
Knowl. Based Syst.2
2024 Behavioral Recognition of Skeletal Data Based on Targeted Dual Fusion Strategy
abstract
The deployment of multi-stream fusion strategy on behavioral recognition from skeletal data can extract complementary features from different information streams and improve the recognition accuracy, but suffers from high model complexity and a large number of parameters. Besides, existing multi-stream methods using a fixed adjacency matrix homogenizes the model’s discrimination process across diverse actions, causing reduction of the actual lift for the multi-stream model. Finally, attention mechanisms are commonly applied to the multi-dimensional features, including spatial, temporal and channel dimensions. But their attention scores are typically fused in a concatenated manner, leading to the ignorance of the interrelation between joints in complex actions. To alleviate these issues, the Front-Rear dual Fusion Graph Convolutional Network (FRF-GCN) is proposed to provide a lightweight model based on skeletal data. Targeted adjacency matrices are also designed for different front fusion streams, allowing the model to focus on actions of varying magnitudes. Simultaneously, the mechanism of Spatial-Temporal-Channel Parallel Attention (STC-P), which processes attention in parallel and places greater emphasis on useful information, is proposed to further improve model’s performance. FRF-GCN demonstrates significant competitiveness compared to the current state-of-the-art methods on the NTU RGB+D, NTU RGB+D 120 and Kinetics-Skeleton 400 datasets. Our code is available at: https://github.com/sunbeam-kkt/FRF-GCN-master.
Xiao Yun, Kévin Riou, Kaiwen Dong, Yanjing Sun, Song Li 0001, Kévin Subrin, Patrick Le Callet
AAAI1
2024 Asymmetric network pseudo labels mutual refinement for unsupervised domain adaptation person re-identification
Xiao Yun, Kaiwen Dong, Yanjing Sun
Multim. Tools Appl.1
2023 Discrepant mutual learning fusion network for unsupervised domain adaptation on person re-identification
Xiao Yun, Qunqun Wang, Xiaozhou Cheng, Kaili Song, Yanjing Sun
Appl. Intell.1
2022 Robust Meta-learning with Sampling Noise and Label Noise via Eigen-Reptile
abstract
Recent years have seen a surge of interest in meta-learning techniques for tackling the few-shot learning (FSL) problem. However, the meta-learner is prone to overfitting since there are only a few available samples, which can be identified as sampling noise on a clean dataset. Besides, when handling the data with noisy labels, the meta-learner could be extremely sensitive to label noise on a corrupted dataset. To address these two challenges, we present Eigen-Reptile (ER) that updates the meta-parameters with the main direction of historical task-specific parameters. Specifically, the main direction is computed in a fast way, where the scale of the calculated matrix is related to the number of gradient steps for the specific task instead of the number of parameters. Furthermore, to obtain a more accurate main direction for Eigen-Reptile in the presence of many noisy labels, we further propose Introspective Self-paced Learning (ISPL). We have theoretically and experimentally demonstrated the soundness and effectiveness of the proposed Eigen-Reptile and ISPL. Particularly, our experiments on different tasks show that the proposed method is able to outperform or achieve highly competitive performance compared with other gradient-based methods with or without noisy labels. The code and data for the proposed method are provided for research purposes https://github.com/Anfeather/Eigen-Reptile.
Dong Chen 0017, Lingfei Wu 0001, Siliang Tang, Xiao Yun, Bo Long, Yueting Zhuang
ICML4
2022 Triplet attention multiple spacetime-semantic graph convolutional network for skeleton-based action recognition
Yanjing Sun, Xiao Yun, Kaiwen Dong
Appl. Intell.3
2018 Multi-layer convolutional network-based visual tracking via important region selection
Xiao Yun, Yanjing Sun, Yunkai Shi, Nannan Lu
Neurocomputing1
2017 Spiral visual and motional tracking
Xiao Yun
Neurocomputing1
2016 A compressive tracking based on time-space Kalman fusion model
Xiao Yun, Zhongliang Jing, Bo Jin 0003, Canlong Zhang
Sci. China Inf. Sci.1
2016 Kernel joint visual tracking and recognition based on structured sparse representation
Xiao Yun, Zhongliang Jing
Neurocomputing1
2012 A 36V biphasic stimulator with electrode monitoring circuit
abstract
A 36V, 7-bit biphasic stimulator with an electrode monitoring circuit (EMC) were implemented in a high voltage (HV) 0.35µm CMOS process. The HV devices can tolerate a gate-to-source voltage up to ∼18V. A HV switch was proposed for the EMC to monitor a rail-to-rail signal of 36V. A fast settling current mirror for the stimulator was also proposed. The settling time of the output current was reduced significantly even with relatively slow HV devices. The worst-case settling time was measured to be 3.6µs. The current efficiency of the stimulator was ∼96%. The mismatch between the anodic and cathodic outputs was <1.65%.
Edward K. F. Lee, Rongching Dai, Natasha Reeves, Xiao Yun
ISCAS4
2012 A multi-cue mean-shift target tracking approach based on fuzzified region dynamic image fusion
Xiao Yun, Jianmin Wu
Sci. China Inf. Sci.2
2010 Low-power charge sensitive amplifier for semiconductor scintillator
abstract
A design of low-noise charge sensitive amplifier (CSA) for measurement of optical response of photo-detector registering light produced by semiconductor scintillator is presented. Detailed analysis of the CSA suitable for large parasitic detector capacitance is provided, regarding noise, power and stability. Two scenarios where input transistor is biased in strong inversion and weak inversion are compared, with included accurate 1/f noise modeling. The experimental prototype was implemented in 0.5 μm CMOS process with a 5 V power supply.
Xiao Yun, Milutin Stanacevic, Serge Luryi
ISCAS1
2009 An Adaptive Front-end Readout System for Radiation Detection
abstract
A design of adaptive front-end preamplifier for measurement of optical response of epitaxial photodiode, registering light produced by semiconductor scintillator, is presented. A time constant of continuous-time based pulse shaping filter is adaptively determined to achieve optimal sensitivity for variable parameters of photodetector and readout electronics. Experimental prototype was designed in 0.5µm CMOS process.
Xiao Yun, Milutin Stanacevic
ISCAS1
2008 Extended counting ADC for 32-channel neural recording headstage for small animals
abstract
Extended counting analog-to-digital converter (ECADC) combines the accuracy of delta-sigma modulation and the speed of algorithmic conversion. This conversion architecture is shown to be useful in biomedical applications, where both resolution and speed are demanded. This work presents a design of ECADC for 32 neural recording channels. Several power optimizing methods are described. The designed converter achieves a resolution of 13 bits and a sampling frequency of 512 kHz. With 3.3V supply, the total power consumption is estimated to be 7 mW. The whole system including 32 neural recording channels is fitted in an area of 3mm × 3mm in 0.5μm CMOS process.
Xiao Yun, Milutin Stanacevic
ISCAS1