Xiao Tu

dblp:72/8364 · DBLP profile ↗
← Back
11ranked-venue papers
1as first author
6since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 1 first-author · 5 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 2

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
4 papers
Trustworthy machine learning · 27% Image recognition and object detection · 19% Efficient and distributed learning · 14%
Human-computer interaction and pervasive computing
2 papers
Interaction techniques and input · 49% Ubiquitous computing and smart environments · 25% User interface design and tools · 19%
Computer graphics and multimedia
1 paper
Visual content generation and editing · 100%

Topics — the 14 heaviest of 19, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Language models and text generation
text generation
0.912025
TrInk: Ink Generation with Transformer Network · EMNLP 2025
Computer vision › Vision and language
multimodal reasoning
0.812024
Image as a Language: Revisiting Scene Text Recognition via Balanced, Unified and Synchronized Vision-Language Reasoning Network · AAAI 2024
Computer vision › Image recognition and object detection
scene text recognition
0.812024
Image as a Language: Revisiting Scene Text Recognition via Balanced, Unified and Synchronized Vision-Language Reasoning Network · AAAI 2024
Interaction techniques and input › input modality › multimodal input
pen and touch input
0.722019
Sensing Posture-Aware Pen+Touch Interaction on Tablets · CHI 2019
WritLarge: Ink Unleashed by Unified Scope, Action, & Zoom · CHI 2017
Machine learning › Trustworthy machine learning › robustness
adversarial robustness
0.712023
IPMix: Label-Preserving Data Augmentation Method for Training Robust Classifiers · NeurIPS 2023
Machine learning › Trustworthy machine learning › robustness
corruption robustness
0.712023
IPMix: Label-Preserving Data Augmentation Method for Training Robust Classifiers · NeurIPS 2023
Machine learning › Deep learning architectures and training
data augmentation
0.712023
IPMix: Label-Preserving Data Augmentation Method for Training Robust Classifiers · NeurIPS 2023
Computer vision › Image recognition and object detection
image classification
0.712023
IPMix: Label-Preserving Data Augmentation Method for Training Robust Classifiers · NeurIPS 2023
Machine learning › Trustworthy machine learning
robustness
0.712023
IPMix: Label-Preserving Data Augmentation Method for Training Robust Classifiers · NeurIPS 2023
Machine learning › Optimization for machine learning › sparse learning
group sparsity
0.512021
Only Train Once: A One-Shot Neural Network Training And Pruning Framework · NeurIPS 2021
Machine learning › Efficient and distributed learning
model compression
0.512021
Only Train Once: A One-Shot Neural Network Training And Pruning Framework · NeurIPS 2021
Machine learning › Efficient and distributed learning › model compression › pruning
structured pruning
0.512021
Only Train Once: A One-Shot Neural Network Training And Pruning Framework · NeurIPS 2021
Machine learning › Representation and self-supervised learning
multimodal representation learning
0.212024
Image as a Language: Revisiting Scene Text Recognition via Balanced, Unified and Synchronized Vision-Language Reasoning Network · AAAI 2024
Interaction techniques and input
gesture input
0.112017
WritLarge: Ink Unleashed by Unified Scope, Action, & Zoom · CHI 2017

Methods — techniques the papers use, named apart from their topics

transformer network · 1.7vision-language model · 0.8masked modality modeling · 0.8iterative reasoning · 0.8data augmentation · 0.7convolutional neural network · 0.7structured-sparsity optimization · 0.5half-space stochastic projected gradient · 0.5inertial motion sensing · 0.4electric field sensing · 0.4capacitive touchscreen sensing · 0.4user study · 0.3
YearPublicationVenuePosition
2025 TrInk: Ink Generation with Transformer Network
abstract
Zezhong Jin, Shubhang Desai, Xu Chen, Biyi Fang, Zhuoyi Huang, Zhe Li, Chong-Xin Gan, Xiao Tu, Man-Wai Mak, Yan Lu, Shujie Liu. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025.
Zezhong Jin, Shubhang Desai, Biyi Fang, Zhuoyi Huang, Zhe Li 0030, Chong-Xin Gan, Xiao Tu, Man-Wai Mak, Yan Lu 0001, Shujie Liu 0001
EMNLP8
2024 Image as a Language: Revisiting Scene Text Recognition via Balanced, Unified and Synchronized Vision-Language Reasoning Network
abstract
Scene text recognition is inherently a vision-language task. However, previous works have predominantly focused either on extracting more robust visual features or designing better language modeling. How to effectively and jointly model vision and language to mitigate heavy reliance on a single modality remains a problem. In this paper, aiming to enhance vision-language reasoning in scene text recognition, we present a balanced, unified and synchronized vision-language reasoning network (BUSNet). Firstly, revisiting the image as a language by balanced concatenation along length dimension alleviates the issue of over-reliance on vision or language. Secondly, BUSNet learns an ensemble of unified external and internal vision-language model with shared weight by masked modality modeling (MMM). Thirdly, a novel vision-language reasoning module (VLRM) with synchronized vision-language decoding capacity is proposed. Additionally, BUSNet achieves improved performance through iterative reasoning, which utilizes the vision-language prediction as a new language input. Extensive experiments indicate that BUSNet achieves state-of-the-art performance on several mainstream benchmark datasets and more challenge datasets for both synthetic and real training data compared to recent outstanding methods. Code and dataset will be available at https://github.com/jjwei66/BUSNet.
Jiajun Wei, Hongjian Zhan, Yue Lu 0001, Xiao Tu, Cong Liu 0006, Umapada Pal 0001
AAAI4
2023 Scene Text Recognition with Image-Text Matching-Guided Dictionary
Jiajun Wei, Hongjian Zhan, Xiao Tu, Yue Lu 0001, Umapada Pal 0001
ICDAR (6)3
2023 Table Structure Recognition of Historical Dongba Documents
Hongjian Zhan, Xiao Tu, Yue Lu 0001
ICIG (1)3
2023 IPMix: Label-Preserving Data Augmentation Method for Training Robust Classifiers
abstract
Data augmentation has been proven effective for training high-accuracy convolutional neural network classifiers by preventing overfitting. However, building deep neural networks in real-world scenarios requires not only high accuracy on clean data but also robustness when data distributions shift. While prior methods have proposed that there is a trade-off between accuracy and robustness, we propose IPMix, a simple data augmentation approach to improve robustness without hurting clean accuracy. IPMix integrates three levels of data augmentation (image-level, patch-level, and pixel-level) into a coherent and label-preserving technique to increase the diversity of training data with limited computational overhead. To further improve the robustness, IPMix introduces structural complexity at different levels to generate more diverse images and adopts the random mixing method for multi-scale information fusion. Experiments demonstrate that IPMix outperforms state-of-the-art corruption robustness on CIFAR-C and ImageNet-C. In addition, we show that IPMix also significantly improves the other safety measures, including robustness to adversarial perturbations, calibration, prediction consistency, and anomaly detection, achieving state-of-the-art or comparable results on several benchmarks, including ImageNet-R, ImageNet-A, and ImageNet-O.
Zhenglin Huang, Xiaoan Bao, Qingqi Zhang, Xiao Tu
NeurIPS5
2021 Only Train Once: A One-Shot Neural Network Training And Pruning Framework
abstract
Structured pruning is a commonly used technique in deploying deep neural networks (DNNs) onto resource-constrained devices. However, the existing pruning methods are usually heuristic, task-specified, and require an extra fine-tuning procedure. To overcome these limitations, we propose a framework that compresses DNNs into slimmer architectures with competitive performances and significant FLOPs reductions by Only-Train-Once (OTO). OTO contains two key steps: (i) we partition the parameters of DNNs into zero-invariant groups, enabling us to prune zero groups without affecting the output; and (ii) to promote zero groups, we then formulate a structured-sparsity optimization problem, and propose a novel optimization algorithm, Half-Space Stochastic Projected Gradient (HSPG), to solve it, which outperforms the standard proximal methods on group sparsity exploration, and maintains comparable convergence. To demonstrate the effectiveness of OTO, we train and compress full models simultaneously from scratch without fine-tuning for inference speedup and parameter reduction, and achieve state-of-the-art results on VGG16 for CIFAR10, ResNet50 for CIFAR10 and Bert for SQuAD and competitive result on ResNet50 for ImageNet. The source code is available at https://github.com/tianyic/onlytrainonce.
Bo Ji 0003, Tianyu Ding, Biyi Fang, Guanyi Wang, Zhihui Zhu, Luming Liang, Yixin Shi, Xiao Tu
NeurIPS10
2020 Orthant Based Proximal Stochastic Gradient Method for ℓ 1-Regularized Optimization
Tianyu Ding, Bo Ji 0003, Guanyi Wang, Yixin Shi, Xiao Tu, Zhihui Zhu
ECML/PKDD (3)8
2020 Discriminant sub-dictionary learning with adaptive multiscale superpixel representation for hyperspectral image classification
Xiao Tu, Xiaobo Shen 0001, Peng Fu 0003, Tao Wang 0020, Quan-Sen Sun, Zexuan Ji
Neurocomputing1
2019 Sensing Posture-Aware Pen+Touch Interaction on Tablets
abstract
Many status-quo interfaces for tablets with pen + touch input capabilities force users to reach for device-centric UI widgets at fixed locations, rather than sensing and adapting to the user-centric posture. To address this problem, we propose sensing techniques that transition between various nuances of mobile and stationary use via postural awareness. These postural nuances include shifting hand grips, varying screen angle and orientation, planting the palm while writing or sketching, and detecting what direction the hands approach from. To achieve this, our system combines three sensing modalities: 1) raw capacitance touchscreen images, 2) inertial motion, and 3) electric field sensors around the screen bezel for grasp and hand proximity detection. We show how these sensors enable posture-aware pen+touch techniques that adapt interaction and morph user interface elements to suit fine-grained contexts of body-, arm-, hand-, and grip-centric frames of reference.
Yang Zhang 0041, Michel Pahud, Christian Holz 0001, Haijun Xia, Gierad Laput, Michael J. McGuffin, Xiao Tu, Andrew Mittereder, William Buxton, Ken Hinckley
CHI7
2019 Residual CRNN and Its Application to Handwritten Digit String Recognition
Hongjian Zhan, Shujing Lyu, Xiao Tu, Yue Lu 0001
ICONIP (5)3
2017 WritLarge: Ink Unleashed by Unified Scope, Action, & Zoom
abstract
WritLarge is a freeform canvas for early-stage design on electronic whiteboards with pen+touch input. The system aims to support a higher-level flow of interaction by 'chunking' the traditionally disjoint steps of selection and action into unified selection-action phrases. This holistic goal led us to address two complementary aspects: SELECTION, for which we devise a new technique known as the Zoom-Catcher that integrates pinch-to-zoom and selection in a single gesture for fluidly selecting and acting on content; plus: ACTION, where we demonstrate how this addresses the combined issues of navigating, selecting, and manipulating content. In particular, the designer can transform select ink strokes in flexible and easily-reversible representations via semantic, structural, and temporal axes of movement that are defined as conceptual 'moves' relative to the specified content.
Haijun Xia, Ken Hinckley, Michel Pahud, Xiao Tu, William Buxton
CHI4