Xiaohao Chen

dblp:199/2054 · DBLP profile ↗
← Back
12ranked-venue papers
1as first author
9since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 11 · 1 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
Generative modeling · 50% Efficient and distributed learning · 12% Image recognition and object detection · 12%
Computer graphics and multimedia
2 papers
Geometric modeling and processing · 36% Multimedia analysis and retrieval · 36% Rendering · 28%

Topics — the 14 heaviest of 15, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Generative modeling › diffusion model › guided diffusion
classifier-free guidance
0.912025
Teefusion: Blending Text Embeddings to Distill Classifier-Free Guidance · ICCV 2025
Machine learning › Generative modeling › diffusion model
diffusion distillation
0.912025
Teefusion: Blending Text Embeddings to Distill Classifier-Free Guidance · ICCV 2025
Machine learning › Generative modeling
diffusion model
0.912025
Teefusion: Blending Text Embeddings to Distill Classifier-Free Guidance · ICCV 2025
Machine learning › Efficient and distributed learning
model compression
0.912025
Teefusion: Blending Text Embeddings to Distill Classifier-Free Guidance · ICCV 2025
Machine learning › Generative modeling › diffusion model
text-to-image generation
0.912025
Teefusion: Blending Text Embeddings to Distill Classifier-Free Guidance · ICCV 2025
Machine learning › Deep learning architectures and training › convolutional neural network
conditional convolution
0.512021
CondLaneNet: a Top-to-down Lane Detection Framework Based on Conditional Convolution · ICCV 2021
Computer vision › Segmentation and scene understanding
instance segmentation
0.512021
CondLaneNet: a Top-to-down Lane Detection Framework Based on Conditional Convolution · ICCV 2021
Robotics › Autonomous driving › perception › vision-based perception
lane detection
0.512021
CondLaneNet: a Top-to-down Lane Detection Framework Based on Conditional Convolution · ICCV 2021
Geometric modeling and processing › computer-aided design › CAD model processing
CAD drawing analysis
0.512021
FloorPlanCAD: A Large-Scale CAD Drawing Dataset for Panoptic Symbol Spotting · ICCV 2021
Rendering
novel view synthesis
0.412019
Structure-Preserving Stereoscopic View Synthesis With Multi-Scale Adversarial Correlation Matching · CVPR 2019
Computer vision › Image recognition and object detection › object detection
bounding box regression
0.312017
Accurate Single Stage Detector Using Recurrent Rolling Convolution · CVPR 2017
Machine learning › Deep learning architectures and training
multi-scale feature fusion
0.312017
Accurate Single Stage Detector Using Recurrent Rolling Convolution · CVPR 2017
Computer vision › Image recognition and object detection
object detection
0.312017
Accurate Single Stage Detector Using Recurrent Rolling Convolution · CVPR 2017
Computer vision › Image recognition and object detection › object detection
one-stage object detection
0.312017
Accurate Single Stage Detector Using Recurrent Rolling Convolution · CVPR 2017

Methods — techniques the papers use, named apart from their topics

text embedding fusion · 0.9knowledge distillation · 0.9recurrent instance module · 0.5graph convolutional network · 0.5convolutional neural network · 0.5conditional convolution · 0.5multi-scale feature extraction · 0.4correlation matching · 0.4adversarial training · 0.4recurrent rolling convolution · 0.3multi-scale feature map · 0.3
YearPublicationVenuePosition
2026 A spectral heterogeneous graph neural network with multi-sensor fusion for machine fault diagnosis
Zhilin Dong, Weidong Jiao, Xiaohao Chen, Wanxiu Xu, Siyu Liu 0008, Yonghua Jiang 0003
Eng. Appl. Artif. Intell.5
2026 Disentangled image-text classification: Enhancing visual representations with MLLM-driven knowledge transfer
Qianjun Shuai, Xiaohao Chen, Yongqiang Cheng 0001, Fang Miao, Libiao Jin
Expert Syst. Appl.2
2026 Linguistic query-guided mask generation for referring image segmentation
Zhichao Wei, Xiaohao Chen, Mingqiang Chen, Zilong Dong, Siyu Zhu 0001
Pattern Recognit.2
2026 Adaptive Time-Leaping Gradient Aligned Based Unsupervised Domain Adaptation for Bearing Fault Diagnosis Under Imbalanced Samples
abstract
Unsupervised domain adaptation shows promise in fault diagnosis by aligning cross-domain feature distributions. However, class imbalance in the target domain degrades the discriminative feature structure and hinders domain shift reduction. To address the challenge, a novel Time-Leaping Gradient Aligned Dynamics Gramian Angle Feature Pyramid Network (TGA-DGAFPN) is proposed. Firstly, a dynamic time leaping sampling mechanism that adaptively predicts reconstructed time points, replacing fixed intervals in Gramian Angular Difference Fields. Secondly, a gradient alignment loss is designed based on time-leaping parameters and analyzed for its correlation with domain shift. This loss aligns the gradient directions between source and target features with respect to time leaping offsets, thereby promoting the dynamic time-leaping sampling mechanism to reduce domain discrepancies. These innovative modules are incorporated into a Feature Pyramid Network framework to integrate multi-scale feature information, thereby capturing more comprehensive and nuanced fault patterns. Extensive evaluations across nine tasks under three imbalanced scenarios demonstrate the superior performance of TGA-DGAFPN in fault diagnosis. In a slightly imbalanced scenario, it achieves 93.04% accuracy, 89.13% precision, and a 91.85% F1-score. Based on various scenarios, the performance of TGA-DGAFPN is superior to the current state-of-the-art models.
Daxuan Lin, Weidong Jiao, Zhilin Dong, Xiaohao Chen, Wanxiu Xu, Siyu Liu 0008, Yonghua Jiang 0003
IEEE Trans. Reliab.5
2025 Teefusion: Blending Text Embeddings to Distill Classifier-Free Guidance
abstract
Recent advances in text-to-image synthesis largely benefit from sophisticated sampling strategies and classifier-free guidance (CFG) to ensure high-quality generation. However, CFG's reliance on two forward passes, especially when combined with intricate sampling algorithms, results in prohibitively high inference costs. To address this, we introduce TeEFusion (Text Embeddings Fusion), a novel and efficient distillation method that directly incorporates the guidance magnitude into the text embeddings and distills the teacher model's complex sampling strategy. By simply fusing conditional and unconditional text embeddings using linear operations, TeEFusion reconstructs the desired guidance without adding extra parameters, simultaneously enabling the student model to learn from the teacher's output produced via its sophisticated sampling approach. Extensive experiments on state-of-the-art models such as SD3 demonstrate that our method allows the student to closely mimic the teacher's performance with a far simpler and more efficient sampling strategy. Consequently, the student model achieves inference speeds up to 6$\times$ faster than the teacher model, while maintaining image quality at levels comparable to those obtained through the teacher's complex sampling approach. The code is publicly available at https://github.com/AIDC-AI/TeEFusion.
Minghao Fu 0001, Guo-Hua Wang, Xiaohao Chen, Weihua Luo, Kaifu Zhang
ICCV3
2025 SDDA: A progressive self-distillation with decoupled alignment for multimodal image-text classification
Xiaohao Chen, Qianjun Shuai, Yongqiang Cheng 0001
Neurocomputing1
2022 GB-CosFace: Rethinking Softmax-Based Face Recognition from the Perspective of Open Set Classification
Mingqiang Chen, Lizhe Liu, Xiaohao Chen, Siyu Zhu 0001
ACCV (4)3
2021 FloorPlanCAD: A Large-Scale CAD Drawing Dataset for Panoptic Symbol Spotting
abstract
Access to large and diverse computer-aided design (CAD) drawings is critical for developing symbol spotting algorithms. In this paper, we present FloorPlan-CAD, a large-scale real-world CAD drawing dataset containing over 10,000 floor plans, ranging from residential to commercial buildings. CAD drawings in the dataset are all represented as vector graphics, which enable us to provide line-grained annotations of 30 object categories. Equipped by such annotations, we introduce the task of panoptic symbol spotting, which requires to spot not only instances of countable things, but also the semantic of uncountable stuff. Aiming to solve this task, we propose a novel method by combining Graph Convolutional Networks (GCNs) with Convolutional Neural Networks (CNNs), which captures both non-Euclidean and Euclidean features and can be trained end-to-end. The proposed CNN-GCN method achieved state-of-the-art (SOTA) performance on the task of semantic symbol spotting, and help us build a baseline network for the panoptic symbol spotting task. Our contributions are three-fold: 1) to the best of our knowledge, the presented CAD drawing dataset is the first of its kind; 2) the panoptic symbol spotting task considers the spotting of both thing instances and stuff semantic as one recognition problem; and 3) we presented a baseline solution to the panoptic symbol spotting task based on a novel CNN-GCN method, which achieved SOTA performance on semantic symbol spotting. We believe that these contributions will boost research in related areas. The dataset and code is publicly available at https://floorplancad.github.io/.
Zhiwen Fan, Lingjie Zhu, Honghua Li, Xiaohao Chen, Siyu Zhu 0001, Ping Tan 0002
ICCV4
2021 CondLaneNet: a Top-to-down Lane Detection Framework Based on Conditional Convolution
abstract
Modern deep-learning-based lane detection methods are successful in most scenarios but struggling for lane lines with complex topologies. In this work, we propose CondLaneNet, a novel top-to-down lane detection framework that detects the lane instances first and then dynamically predicts the line shape for each instance. Aiming to resolve lane instance-level discrimination problem, we introduce a conditional lane detection strategy based on conditional convolution and row-wise formulation. Further, we design the Recurrent Instance Module(RIM) to overcome the problem of detecting lane lines with complex topologies such as dense lines and fork lines. Benefit from the end-to-end pipeline which requires little post-process, our method has real-time efficiency. We extensively evaluate our method on three benchmarks of lane detection. Results show that our method achieves state-of-the-art performance on all three benchmark datasets. Moreover, our method has the coexistence of accuracy and efficiency, e.g. a 78.14 F1 score and 220 FPS on CULane. Our code is available at https://github.com/aliyun/conditional-lane-detection.
Lizhe Liu, Xiaohao Chen, Siyu Zhu 0001, Ping Tan 0002
ICCV2
2019 Structure-Preserving Stereoscopic View Synthesis With Multi-Scale Adversarial Correlation Matching
abstract
This paper addresses stereoscopic view synthesis from a single image. Various recent works solve this task by reorganizing pixels from the input view to reconstruct the target one in a stereo setup. However, purely depending on such photometric-based reconstruction process, the network may produce structurally inconsistent results. Regarding this issue, this work proposes Multi-Scale Adversarial Correlation Matching (MS-ACM), a novel learning framework for structure-aware view synthesis. The proposed framework does not assume any costly supervision signal of scene structures such as depth. Instead, it models structures as self-correlation coefficients extracted from multi-scale feature maps in transformed spaces. In training, the feature space attempts to push the correlation distances between the synthesized and target images far apart, thus amplifying inconsistent structures. At the same time, the view synthesis network minimizes such correlation distances by fixing mistakes it makes. With such adversarial training, structural errors of different scales and levels are iteratively discovered and reduced, preserving both global layouts and fine-grained details. Extensive experiments on the KITTI benchmark show that MS-ACM improves both visual quality and the metrics over existing methods when plugged into recent view synthesis architectures.
Yu Zhang 0035, Dongqing Zou, Jimmy S. J. Ren, Xiaohao Chen
CVPR5
2018 Learning Selfie-Friendly Abstraction from Artistic Style Images
abstract
Artistic style transfer can be thought as a process to generate different versions of abstraction of the original image. However, most of the artistic style transfer operators are not optimized for human faces thus mainly suffers from two undesirable features when applying them to selfies. First, the edges of human faces may unpleasantly deviate from the ones in the original image. Second, the skin color is far from faithful to the original one which is usually problematic in producing quality selfies. In this paper, we take a different approach and formulate this abstraction process as a gradient domain learning problem. We aim to learn a type of abstraction which not only achieves the specified artistic style but also circumvents the two aforementioned drawbacks thus highly applicable to selfie photography. We also show that our method can be directly generalized to videos with high inter-frame consistency. Our method is also robust to non-selfie images, and the generalization to various kinds of real-life scenes is discussed. We will make our code publicly available.
Yicun Liu, Jimmy S. J. Ren, Xiaohao Chen
ACML4
2017 Accurate Single Stage Detector Using Recurrent Rolling Convolution
abstract
Most of the recent successful methods in accurate object detection and localization used some variants of R-CNN style two stage Convolutional Neural Networks (CNN) where plausible regions were proposed in the first stage then followed by a second stage for decision refinement. Despite the simplicity of training and the efficiency in deployment, the single stage detection methods have not been as competitive when evaluated in benchmarks consider mAP for high IoU thresholds. In this paper, we proposed a novel single stage end-to-end trainable object detection network to overcome this limitation. We achieved this by introducing Recurrent Rolling Convolution (RRC) architecture over multi-scale feature maps to construct object classifiers and bounding box regressors which are deep in context. We evaluated our method in the challenging KITTI dataset which measures methods under IoU threshold of 0.7. We showed that with RRC, a single reduced VGG-16 based model already significantly outperformed all the previously published results. At the time this paper was written our models ranked the first in KITTI car detection (the hard level), the first in cyclist detection and the second in pedestrian detection. These results were not reached by the previous single stage methods. The code is publicly available.
Jimmy S. J. Ren, Xiaohao Chen, Wenxiu Sun, Jiahao Pang, Qiong Yan, Yu-Wing Tai, Li Xu 0001
CVPR2