Xiaojian Lin

dblp:194/1586 · DBLP profile ↗
← Back
10ranked-venue papers
4as first author
8since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 5 · 1 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 3 since 2021Computer networks · 2 · 2 first-authorSystems, architecture and hardware · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
Generative modeling · 33% 3D vision · 29% Autonomous driving · 17%
Computer graphics and multimedia
2 papers
Visual content generation and editing · 77% Image and video processing · 23%
Computer networks
1 paper
Cellular and mobile networks · 86% Datacenter networks · 14%

Topics — the 13 heaviest of 14, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Generative modeling
diffusion model
0.912025
GDrag: Towards General-Purpose Interactive Editing with Anti-ambiguity Point Diffusion · ICLR 2025
Computer vision › Image recognition and object detection › object detection
object detection for autonomous driving
0.912025
Butter: Frequency Consistency and Hierarchical Fusion for Autonomous Driving Object Detection · ACM Multimedia 2025
Robotics › Autonomous driving › perception › environment perception
perception for self-driving vehicles
0.912025
Butter: Frequency Consistency and Hierarchical Fusion for Autonomous Driving Object Detection · ACM Multimedia 2025
Machine learning › Generative modeling › diffusion model › image editing
training-free image editing
0.912025
GDrag: Towards General-Purpose Interactive Editing with Anti-ambiguity Point Diffusion · ICLR 2025
Visual content generation and editing › image editing › interactive image editing
drag-based image editing
0.912025
GDrag: Towards General-Purpose Interactive Editing with Anti-ambiguity Point Diffusion · ICLR 2025
Computer vision › 3D vision › 3d human reconstruction
hand mesh reconstruction
0.812024
Monocular 3D Hand Mesh Recovery via Dual Noise Estimation · AAAI 2024
Cellular and mobile networks › radio access networks
cloud radio access network
0.312017
Efficient remote radio head switching scheme in cloud radio access network: A load balancing perspective · INFOCOM 2017
Image and video processing › frequency domain analysis
frequency-domain image processing
0.312025
Butter: Frequency Consistency and Hierarchical Fusion for Autonomous Driving Object Detection · ACM Multimedia 2025
Natural language and speech › Speech recognition and synthesis › speech enhancement
noise estimation
0.212024
Monocular 3D Hand Mesh Recovery via Dual Noise Estimation · AAAI 2024
Cellular and mobile networks › mobility management
handover
0.112017
Efficient remote radio head switching scheme in cloud radio access network: A load balancing perspective · INFOCOM 2017
Datacenter networks
load balancing
0.112017
Efficient remote radio head switching scheme in cloud radio access network: A load balancing perspective · INFOCOM 2017
Cellular and mobile networks
mobility management
0.112017
Efficient remote radio head switching scheme in cloud radio access network: A load balancing perspective · INFOCOM 2017
Cellular and mobile networks
radio access networks
0.112017
Efficient remote radio head switching scheme in cloud radio access network: A load balancing perspective · INFOCOM 2017

Methods — techniques the papers use, named apart from their topics

motion supervision · 1.7low-rank adaptation · 1.7hierarchical feature fusion · 1.7diffusion · 1.7probabilistic modeling · 0.8dual noise estimation · 0.8local search · 0.3approximation algorithm · 0.3
YearPublicationVenuePosition
2025 FreCT: Frequency-Augmented Convolutional Transformer for Robust Time Series Anomaly Detection
Wenxin Zhang 0005, Guangzhen Yao, Xiaojian Lin, Renxiang Guan, Chengze Du 0001, Renda Han, Xi Xuan, Cuicui Luo
ICIC (16)4
2025 GDrag: Towards General-Purpose Interactive Editing with Anti-ambiguity Point Diffusion
abstract
Recent interactive point-based image manipulation methods have gained considerable attention for being user-friendly. However, these methods still face two types of ambiguity issues that can lead to unsatisfactory outcomes, namely, intention ambiguity which misinterprets the purposes of users, and content ambiguity where target image areas are distorted by distracting elements. To address these issues and achieve general-purpose manipulations, we propose a novel task-aware, training-free framework called GDrag. Specifically, GDrag defines a taxonomy of atomic manipulations, which can be parameterized and combined unitedly to represent complex manipulations, thereby reducing intention ambiguity. Furthermore, GDrag introduces two strategies to mitigate content ambiguity, including an anti-ambiguity dense trajectory calculation method (ADT) and a self-adaptive motion supervision method (SMS). Given an atomic manipulation, ADT converts the sparse user-defined handle points into a dense point set by selecting their semantic and geometric neighbors, and calculates the trajectory of the point set. Unlike previous motion supervision methods relying on a single global scale for low-rank adaption, SMS jointly optimizes point-wise adaption scales and latent feature biases. These two methods allow us to model fine-grained target contexts and generate precise trajectories. As a result, GDrag consistently produces precise and appealing results in different editing tasks. Extensive experiments on the challenging DragBench dataset demonstrate that GDrag outperforms state-of-the-art methods significantly. The code of GDrag will be released upon acceptance.
Xiaojian Lin, Yuhao Cheng, Yiqiang Yan, Xiaodan Liang
ICLR1
2025 A-MESS: Anchor-based Multimodal Embedding with Semantic Synchronization for Multimodal Intent Recognition
abstract
In the domain of multimodal intent recognition (MIR), the objective is to recognize human intent by integrating a variety of modalities, such as language text, body gestures, and tones. However, existing approaches face difficulties adequately capturing the intrinsic connections between the modalities and overlooking the corresponding semantic representations of intent. To address these limitations, we present the Anchor-based Multimodal Embedding with Semantic Synchronization (A-MESS) framework. We first design an Anchor-based Multimodal Embedding (A-ME) module that employs an anchor-based embedding fusion mechanism to integrate multimodal inputs. Furthermore, we develop a Semantic Synchronization (SS) strategy with the Triplet Contrastive Learning pipeline, which optimizes the process by synchronizing multimodal representation with label descriptions produced by the large language model. Comprehensive experiments indicate that our A-MESS achieves state-of-the-art and provides substantial insight into multimodal representation and downstream tasks.
Yaomin Shen, Xiaojian Lin
ICME2
2025 DConAD: A Differencing-based Contrastive Representation Learning Framework for Time Series Anomaly Detection
abstract
Time series anomaly detection holds notable importance for risk identification and fault detection across diverse application domains. Unsupervised learning methods have become popular because they have no requirement for labels. However, due to the challenges posed by the multiplicity of abnormal patterns, the sparsity of anomalies, and the growth of data scale and complexity, these methods often fail to capture robust and representative dependencies within the time series for identifying anomalies. To enhance the ability of models to capture normal patterns of time series and avoid the retrogression of modeling ability triggered by the dependencies on high-quality prior knowledge, we propose a differencing-based contrastive representation learning framework for time series anomaly detection (DConAD). Specifically, DConAD generates differential data to provide additional information about time series and utilizes transformer-based architecture to capture spatiotemporal dependencies, which enhances the robustness of unbiased representation learning ability. Furthermore, DConAD implements a novel KL divergence-based contrastive learning paradigm that only uses positive samples to avoid deviation from reconstruction and deploys the stop-gradient strategy to compel convergence. Extensive experiments on five public datasets show the superiority and effectiveness of DConAD compared with nine baselines. The code is available at https://github.com/shaieesss/DConAD.
Wenxin Zhang 0005, Xiaojian Lin, Guangzhen Yao, Jingxing Zhong, Yu Li 0047, Renda Han, Songcheng Xu, Cuicui Luo
IJCNN2
2025 Dual-channel Heterophilic Message Passing for Graph Fraud Detection
abstract
Fraudulent activities have significantly increased across various domains, such as e-commerce, online review platforms, and social networks, making fraud detection a critical task. Spatial Graph Neural Networks (GNNs) have been successfully applied to fraud detection tasks due to their strong inductive learning capabilities. However, existing spatial GNN-based methods often enhance the graph structure by excluding heterophilic neighbors during message passing to align with the homophilic bias of GNNs. Unfortunately, this approach can disrupt the original graph topology and increase uncertainty in predictions. To address these limitations, this paper proposes a novel framework, Dual-channel Heterophilic Message Passing (DHMP), for fraud detection. DHMP leverages a heterophily separation module to divide the graph into homophilic and heterophilic subgraphs, mitigating the low-pass inductive bias of traditional GNNs. It then applies shared weights to capture signals at different frequencies independently and incorporates a customized sampling strategy for training. This allows nodes to adaptively balance the contributions of various signals based on their labels. Extensive experiments on three real-world datasets demonstrate that DHMP outperforms existing methods, highlighting the importance of separating signals with different frequencies for improved fraud detection. The code is available at https://github.com/shaieesss/DHMP.
Wenxin Zhang 0005, Jingxing Zhong, Guangzhen Yao, Renda Han, Xiaojian Lin, Zeyu Zhang 0006, Cuicui Luo
IJCNN5
2025 Reusing Attention for One-stage Lane Topology Understanding
abstract
Understanding lane topology relationships accurately is critical for safe autonomous driving. However, existing two-stage methods suffer from inefficiencies due to error propagations and increased computational overheads. To address these challenges, we propose a one-stage architecture that simultaneously predicts traffic elements, lane centerlines and topology relationship, improving both the accuracy and inference speed of lane topology understanding for autonomous driving. Our key innovation lies in reusing intermediate attention resources within distinct transformer decoders. This approach effectively leverages the inherent relational knowledge within the element detection module to enable the modeling of topology relationships among traffic elements and lanes without requiring additional computationally expensive graph networks. Furthermore, we are the first to demonstrate that knowledge can be distilled from models that utilize standard definition (SD) maps to those operates without using SD maps, enabling superior performance even in the absence of SD maps. Extensive experiments on the OpenLane-V2 dataset show that our approach outperforms baseline methods in both accuracy and efficiency, achieving superior results in lane detection, traffic element identification, and topology reasoning. Our code is available at https://github.com/Yang-Li-2000/one-stage.git.
Yang Li 0178, Zongzheng Zhang, Xuchong Qiu, Xinrun Li, Leichen Wang, Ruikai Li, Zhenxin Zhu, Huan-ang Gao, Xiaojian Lin, Zhiyong Cui, Hang Zhao 0021, Hao Zhao 0002
IROS10
2025 Butter: Frequency Consistency and Hierarchical Fusion for Autonomous Driving Object Detection
Xiaojian Lin, Wenxin Zhang 0005, Yuchu Jiang, Wangyu Wu, Kangxu Wang, Zongzheng Zhang, Guijin Wang, Lei Jin 0003, Hao Zhao 0002
ACM Multimedia1
2024 Monocular 3D Hand Mesh Recovery via Dual Noise Estimation
abstract
Current parametric models have made notable progress in 3D hand pose and shape estimation. However, due to the fixed hand topology and complex hand poses, current models are hard to generate meshes that are aligned with the image well. To tackle this issue, we introduce a dual noise estimation method in this paper. Given a single-view image as input, we first adopt a baseline parametric regressor to obtain the coarse hand meshes. We assume the mesh vertices and their image-plane projections are noisy, and can be associated in a unified probabilistic model. We then learn the distributions of noise to refine mesh vertices and their projections. The refined vertices are further utilized to refine camera parameters in a closed-form manner. Consequently, our method obtains well-aligned and high-quality 3D hand meshes. Extensive experiments on the large-scale Interhand2.6M dataset demonstrate that the proposed method not only improves the performance of its baseline by more than 10% but also achieves state-of-the-art performance. Project page: https://github.com/hanhuili/DNE4Hand.
Xiaojian Lin, Zejun Yang, Zhisheng Wang 0001, Xiaodan Liang
AAAI2
2017 Joint user association and base station switching on/off for green heterogeneous cellular networks
abstract
Heterogeneous cellular network (HCN), which generally consists of small cell base stations (SBSs) and macro base stations (MBSs), is proposed as a promising scheme to improve the capacity of the cellular network. However, huge energy consumption by the densely deployed SBSs arises as a challenging problem that should be addressed properly to fulfill the potential of HCNs. In this paper, we aim to minimize the total power consumption of the HCNs by jointly designing energy-efficient user association and SBS switching schemes. We develop an approximation algorithm to solve the intractable user association problem efficiently, based on which an efficient local search procedure is introduced to minimize the total power consumption of the HCN by controlling active/inactive state of each SBS dynamically. Numerical results demonstrate that our proposed method reduces the energy consumption of the HCN significantly as compared to other representative methods.
Xiaojian Lin, Shaowei Wang 0001
ICC1
2017 Efficient remote radio head switching scheme in cloud radio access network: A load balancing perspective
abstract
Cloud radio access network (C-RAN) is deemed as a promising architecture to meet the exponentially increasing traffic demand in mobile networks, where baseband processing is separated from remote radio heads (RRHs) and performed in a centralized baseband unit (BBU) pool. However, the densely deployed RRHs, as well as the passive optical network which provides high capacity backhauls between the RRHs and the BBU pool, consume a large amount of energy. In this paper, we propose efficient RRH switching schemes to achieve a tradeoff between the system energy saving and the load balance among the RRHs in the C-RAN. We first develop an approximation algorithm to address the intractable user association problem for a given set of RRHs, based on which we introduce efficient local search algorithms to perform RRH selection procedure, which can reduce the load fairness index of the C-RAN by controlling the active/inactive state of each RRH. We also discuss the handover signalling overhead issue and introduce an adaptive trigger mechanism to avoid switching on/off too many RRHs simultaneously so as to keep the signalling overhead of the C-RAN below an acceptable level. Numerical results demonstrate that the proposed RRH switching schemes can improve the system performance of the C-RAN significantly. Moreover, our proposal sheds light on how to design effective and efficient handover schemes for next generation mobile networks.
Xiaojian Lin, Shaowei Wang 0001
INFOCOM1