EDBT 2026 Demo / reviewers in the wild / expert
Xiao Tu
dblp:72/8364
· DBLP profile ↗
11ranked-venue papers
1as first author
6since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 8 · 1 first-author · 5 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 2
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
4 papers |
Trustworthy machine learning · 27% Image recognition and object detection · 19% Efficient and distributed learning · 14% | |
| Human-computer interaction and pervasive computing
2 papers |
Interaction techniques and input · 49% Ubiquitous computing and smart environments · 25% User interface design and tools · 19% | |
| Computer graphics and multimedia
1 paper |
Visual content generation and editing · 100% |
Topics — the 14 heaviest of 19, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Language models and text generation
text generation |
0.9 | 1 | 2025 | TrInk: Ink Generation with Transformer Network · EMNLP 2025 |
Computer vision › Vision and language
multimodal reasoning |
0.8 | 1 | 2024 | Image as a Language: Revisiting Scene Text Recognition via Balanced, Unified and Synchronized Vision-Language Reasoning Network · AAAI 2024 |
Computer vision › Image recognition and object detection
scene text recognition |
0.8 | 1 | 2024 | Image as a Language: Revisiting Scene Text Recognition via Balanced, Unified and Synchronized Vision-Language Reasoning Network · AAAI 2024 |
Interaction techniques and input › input modality › multimodal input
pen and touch input |
0.7 | 2 | 2019 | Sensing Posture-Aware Pen+Touch Interaction on Tablets · CHI 2019 WritLarge: Ink Unleashed by Unified Scope, Action, & Zoom · CHI 2017 |
Machine learning › Trustworthy machine learning › robustness
adversarial robustness |
0.7 | 1 | 2023 | IPMix: Label-Preserving Data Augmentation Method for Training Robust Classifiers · NeurIPS 2023 |
Machine learning › Trustworthy machine learning › robustness
corruption robustness |
0.7 | 1 | 2023 | IPMix: Label-Preserving Data Augmentation Method for Training Robust Classifiers · NeurIPS 2023 |
Machine learning › Deep learning architectures and training
data augmentation |
0.7 | 1 | 2023 | IPMix: Label-Preserving Data Augmentation Method for Training Robust Classifiers · NeurIPS 2023 |
Computer vision › Image recognition and object detection
image classification |
0.7 | 1 | 2023 | IPMix: Label-Preserving Data Augmentation Method for Training Robust Classifiers · NeurIPS 2023 |
Machine learning › Trustworthy machine learning
robustness |
0.7 | 1 | 2023 | IPMix: Label-Preserving Data Augmentation Method for Training Robust Classifiers · NeurIPS 2023 |
Machine learning › Optimization for machine learning › sparse learning
group sparsity |
0.5 | 1 | 2021 | Only Train Once: A One-Shot Neural Network Training And Pruning Framework · NeurIPS 2021 |
Machine learning › Efficient and distributed learning
model compression |
0.5 | 1 | 2021 | Only Train Once: A One-Shot Neural Network Training And Pruning Framework · NeurIPS 2021 |
Machine learning › Efficient and distributed learning › model compression › pruning
structured pruning |
0.5 | 1 | 2021 | Only Train Once: A One-Shot Neural Network Training And Pruning Framework · NeurIPS 2021 |
Machine learning › Representation and self-supervised learning
multimodal representation learning |
0.2 | 1 | 2024 | Image as a Language: Revisiting Scene Text Recognition via Balanced, Unified and Synchronized Vision-Language Reasoning Network · AAAI 2024 |
Interaction techniques and input
gesture input |
0.1 | 1 | 2017 | WritLarge: Ink Unleashed by Unified Scope, Action, & Zoom · CHI 2017 |
Methods — techniques the papers use, named apart from their topics
transformer network · 1.7vision-language model · 0.8masked modality modeling · 0.8iterative reasoning · 0.8data augmentation · 0.7convolutional neural network · 0.7structured-sparsity optimization · 0.5half-space stochastic projected gradient · 0.5inertial motion sensing · 0.4electric field sensing · 0.4capacitive touchscreen sensing · 0.4user study · 0.3
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | TrInk: Ink Generation with Transformer NetworkabstractZezhong Jin, Shubhang Desai, Xu Chen, Biyi Fang, Zhuoyi Huang, Zhe Li, Chong-Xin Gan, Xiao Tu, Man-Wai Mak, Yan Lu, Shujie Liu. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Zezhong Jin, Shubhang Desai, Biyi Fang, Zhuoyi Huang, Zhe Li 0030, Chong-Xin Gan, Xiao Tu, Man-Wai Mak, Yan Lu 0001, Shujie Liu 0001 |
EMNLP | 8 |
| 2024 | Image as a Language: Revisiting Scene Text Recognition via Balanced, Unified and Synchronized Vision-Language Reasoning NetworkabstractScene text recognition is inherently a vision-language task. However, previous works have predominantly focused either on extracting more robust visual features or designing better language modeling. How to effectively and jointly model vision and language to mitigate heavy reliance on a single modality remains a problem. In this paper, aiming to enhance vision-language reasoning in scene text recognition, we present a balanced, unified and synchronized vision-language reasoning network (BUSNet). Firstly, revisiting the image as a language by balanced concatenation along length dimension alleviates the issue of over-reliance on vision or language. Secondly, BUSNet learns an ensemble of unified external and internal vision-language model with shared weight by masked modality modeling (MMM). Thirdly, a novel vision-language reasoning module (VLRM) with synchronized vision-language decoding capacity is proposed. Additionally, BUSNet achieves improved performance through iterative reasoning, which utilizes the vision-language prediction as a new language input. Extensive experiments indicate that BUSNet achieves state-of-the-art performance on several mainstream benchmark datasets and more challenge datasets for both synthetic and real training data compared to recent outstanding methods. Code and dataset will be available at https://github.com/jjwei66/BUSNet. Jiajun Wei, Hongjian Zhan, Yue Lu 0001, Xiao Tu, Cong Liu 0006, Umapada Pal 0001 |
AAAI | 4 |
| 2023 | Scene Text Recognition with Image-Text Matching-Guided Dictionary
Jiajun Wei, Hongjian Zhan, Xiao Tu, Yue Lu 0001, Umapada Pal 0001 |
ICDAR (6) | 3 |
| 2023 | Table Structure Recognition of Historical Dongba Documents
Hongjian Zhan, Xiao Tu, Yue Lu 0001 |
ICIG (1) | 3 |
| 2023 | IPMix: Label-Preserving Data Augmentation Method for Training Robust ClassifiersabstractData augmentation has been proven effective for training high-accuracy convolutional neural network classifiers by preventing overfitting. However, building deep neural networks in real-world scenarios requires not only high accuracy on clean data but also robustness when data distributions shift. While prior methods have proposed that there is a trade-off between accuracy and robustness, we propose IPMix, a simple data augmentation approach to improve robustness without hurting clean accuracy. IPMix integrates three levels of data augmentation (image-level, patch-level, and pixel-level) into a coherent and label-preserving technique to increase the diversity of training data with limited computational overhead. To further improve the robustness, IPMix introduces structural complexity at different levels to generate more diverse images and adopts the random mixing method for multi-scale information fusion. Experiments demonstrate that IPMix outperforms state-of-the-art corruption robustness on CIFAR-C and ImageNet-C. In addition, we show that IPMix also significantly improves the other safety measures, including robustness to adversarial perturbations, calibration, prediction consistency, and anomaly detection, achieving state-of-the-art or comparable results on several benchmarks, including ImageNet-R, ImageNet-A, and ImageNet-O. Zhenglin Huang, Xiaoan Bao, Qingqi Zhang, Xiao Tu |
NeurIPS | 5 |
| 2021 | Only Train Once: A One-Shot Neural Network Training And Pruning FrameworkabstractStructured pruning is a commonly used technique in deploying deep neural networks (DNNs) onto resource-constrained devices. However, the existing pruning methods are usually heuristic, task-specified, and require an extra fine-tuning procedure. To overcome these limitations, we propose a framework that compresses DNNs into slimmer architectures with competitive performances and significant FLOPs reductions by Only-Train-Once (OTO). OTO contains two key steps: (i) we partition the parameters of DNNs into zero-invariant groups, enabling us to prune zero groups without affecting the output; and (ii) to promote zero groups, we then formulate a structured-sparsity optimization problem, and propose a novel optimization algorithm, Half-Space Stochastic Projected Gradient (HSPG), to solve it, which outperforms the standard proximal methods on group sparsity exploration, and maintains comparable convergence. To demonstrate the effectiveness of OTO, we train and compress full models simultaneously from scratch without fine-tuning for inference speedup and parameter reduction, and achieve state-of-the-art results on VGG16 for CIFAR10, ResNet50 for CIFAR10 and Bert for SQuAD and competitive result on ResNet50 for ImageNet. The source code is available at https://github.com/tianyic/onlytrainonce. Bo Ji 0003, Tianyu Ding, Biyi Fang, Guanyi Wang, Zhihui Zhu, Luming Liang, Yixin Shi, Xiao Tu |
NeurIPS | 10 |
| 2020 | Orthant Based Proximal Stochastic Gradient Method for ℓ 1-Regularized Optimization
Tianyu Ding, Bo Ji 0003, Guanyi Wang, Yixin Shi, Xiao Tu, Zhihui Zhu |
ECML/PKDD (3) | 8 |
| 2020 | Discriminant sub-dictionary learning with adaptive multiscale superpixel representation for hyperspectral image classification
Xiao Tu, Xiaobo Shen 0001, Peng Fu 0003, Tao Wang 0020, Quan-Sen Sun, Zexuan Ji |
Neurocomputing | 1 |
| 2019 | Sensing Posture-Aware Pen+Touch Interaction on TabletsabstractMany status-quo interfaces for tablets with pen + touch input capabilities force users to reach for device-centric UI widgets at fixed locations, rather than sensing and adapting to the user-centric posture. To address this problem, we propose sensing techniques that transition between various nuances of mobile and stationary use via postural awareness. These postural nuances include shifting hand grips, varying screen angle and orientation, planting the palm while writing or sketching, and detecting what direction the hands approach from. To achieve this, our system combines three sensing modalities: 1) raw capacitance touchscreen images, 2) inertial motion, and 3) electric field sensors around the screen bezel for grasp and hand proximity detection. We show how these sensors enable posture-aware pen+touch techniques that adapt interaction and morph user interface elements to suit fine-grained contexts of body-, arm-, hand-, and grip-centric frames of reference. Yang Zhang 0041, Michel Pahud, Christian Holz 0001, Haijun Xia, Gierad Laput, Michael J. McGuffin, Xiao Tu, Andrew Mittereder, William Buxton, Ken Hinckley |
CHI | 7 |
| 2019 | Residual CRNN and Its Application to Handwritten Digit String Recognition
Hongjian Zhan, Shujing Lyu, Xiao Tu, Yue Lu 0001 |
ICONIP (5) | 3 |
| 2017 | WritLarge: Ink Unleashed by Unified Scope, Action, & ZoomabstractWritLarge is a freeform canvas for early-stage design on electronic whiteboards with pen+touch input. The system aims to support a higher-level flow of interaction by 'chunking' the traditionally disjoint steps of selection and action into unified selection-action phrases. This holistic goal led us to address two complementary aspects: SELECTION, for which we devise a new technique known as the Zoom-Catcher that integrates pinch-to-zoom and selection in a single gesture for fluidly selecting and acting on content; plus: ACTION, where we demonstrate how this addresses the combined issues of navigating, selecting, and manipulating content. In particular, the designer can transform select ink strokes in flexible and easily-reversible representations via semantic, structural, and temporal axes of movement that are defined as conceptual 'moves' relative to the specified content. Haijun Xia, Ken Hinckley, Michel Pahud, Xiao Tu, William Buxton |
CHI | 4 |