EDBT 2026 Demo / reviewers in the wild / expert
Zhengzhuo Xu
dblp:250/1076
· DBLP profile ↗
16ranked-venue papers
7as first author
15since 2021 · last 2025
0000-0003-4620-9187ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 12 · 5 first-author · 12 since 2021Artificial intelligence and machine learning · 7 · 4 first-author · 7 since 2021Computer networks · 2 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | ChartPoint: Guiding MLLMs with Grounding Reflection for Chart ReasoningabstractMultimodal Large Language Models (MLLMs) have emerged as powerful tools for chart comprehension. However, they heavily rely on extracted content via OCR, which leads to numerical hallucinations when chart textual annotations are sparse. While existing methods focus on scaling instructions, they fail to address the fundamental challenge, i.e., reasoning with visual perception. In this paper, we identify a critical observation: MLLMs exhibit weak grounding in chart elements and proportional relationships, as evidenced by their inability to localize key positions to match their reasoning. To bridge this gap, we propose PointCoT, which integrates reflective interaction into chain-of-thought reasoning in charts. By prompting MLLMs to generate bounding boxes and re-render charts based on location annotations, we establish connections between textual reasoning steps and visual grounding regions. We further introduce an automated pipeline to construct ChartPoint-SFT-62k, a dataset featuring 19.2K high-quality chart samples with step-by-step CoT, bounding box, and re-rendered visualizations. Leveraging this data, we develop two instruction-tuned models, ChartPointQ2 and ChartPointQ2.5, which outperform state-of-the-art across several chart benchmarks, e.g., +5.04\% on ChartBench. Zhengzhuo Xu, SiNan Du, Yiyan Qi, Siwen Lu, Chengjin Xu, Chun Yuan 0003 |
ICCV | 1 |
| 2025 | ChartMoE: Mixture of Diversely Aligned Expert Connector for Chart UnderstandingabstractAutomatic chart understanding is crucial for content comprehension and document parsing. Multimodal Large Language Models (MLLMs) have demonstrated remarkable capabilities in chart understanding through domain-specific alignment and fine-tuning. However, current MLLMs still struggle to provide faithful data and reliable analysis only based on charts. To address it, we propose ChartMoE, which employs the Mixture of Expert (MoE) architecture to replace the traditional linear projector to bridge the modality gap. Specifically, we train several linear connectors through distinct alignment tasks, which are utilized as the foundational initialization parameters for different experts. Additionally, we introduce ChartMoE-Align, a dataset with nearly 1 million chart-table-JSON-code quadruples to conduct three alignment tasks (chart-table/JSON/code). Combined with the vanilla connector, we initialize different experts diversely and adopt high-quality knowledge learning to further refine the MoE connector and LLM parameters. Extensive experiments demonstrate the effectiveness of the MoE connector and our initialization strategy, e.g., ChartMoE improves the accuracy of the previous state-of-the-art from 80.48% to 84.64% on the ChartQA benchmark. Zhengzhuo Xu, Bowen Qu, Yiyan Qi, Sinan Du, Chengjin Xu, Chun Yuan 0003 |
ICLR | 1 |
| 2025 | Boosting Long-Tailed Recognition With Label Descriptor and BeyondabstractLong-Tailed Recognition (LTR) poses significant challenges due to the heavily imbalanced nature of real-world data, which severely skews data-driven deep neural networks. Despite the rapid progress of Vision-Language Models (VLMs), they still face challenges in effectively learning from long-tailed visual data. In this paper, we present a comprehensive analysis of the reasons behind the underperformance of VLMs and propose a hierarchical inference framework to address this issue. Specifically, we prompt the large language models to generatesentence-leveldescriptors for class labels and conduct the open vocabulary classification by computing the average similarity between the image and each descriptor. Areweightingmechanism is further proposed to filter out uninformative descriptors. To mitigate model bias incurred by the long-tail distribution, we propose a feature adapter with the logit adjustment technique and fine-tune the CLIP model via visual prompt tokens. We introduce the Shared Feature space Mixup (SFM) to enhance the interaction between modalities to address tail visual feature insufficiency. Finally, we propose a hierarchical inference manner to combine the aforementioned proposals. Extensive evaluations demonstrate that our approach achieves state-of-the-art performance by fine-tuning only a few parameters on the Places-LT, ImageNet-LT, and iNaturalist 2018 benchmarks. Zhengzhuo Xu, Ruikang Liu, Zenghao Chai, Yiyan Qi, Lei Li 0051, Haiqin Yang, Chun Yuan 0003 |
IEEE Trans. Multim. | 1 |
| 2024 | Towards Effective Collaborative Learning in Long-Tailed RecognitionabstractReal-world data usually suffers from severe class imbalance and long-tailed distributions, where minority classes are significantly underrepresented compared to the majority ones. Recent research prefers to utilize multi-expert architectures to mitigate the model uncertainty on the minority, where collaborative learning is employed to aggregate the knowledge of experts, i.e., online distillation. In this article, we observe that the knowledge transfer between experts is imbalanced in terms of class distribution, which results in limited performance improvement of the minority classes. To address it, we propose a re-weighted distillation loss by comparing two classifiers' predictions, which are supervised by online distillation and label annotations, respectively. We also emphasize that feature-level distillation will significantly improve model performance and increase feature robustness. Finally, we propose an Effective Collaborative Learning (ECL) framework that integrates a contrastive proxy task branch to further improve feature quality. Quantitative and qualitative experiments on four standard datasets demonstrate that ECL achieves state-of-the-art performance and the detailed ablation studies manifest the effectiveness of each component in ECL. Zhengzhuo Xu, Zenghao Chai, Chengyin Xu, Chun Yuan 0003, Haiqin Yang |
IEEE Trans. Multim. | 1 |
| 2023 | Learning Imbalanced Data with Vision TransformersabstractThe real-world data tends to be heavily imbalanced and severely skew the data-driven deep neural networks, which makes Long-Tailed Recognition (LTR) a massive challenging task. Existing LTR methods seldom train Vision Transformers (ViTs) with Long-Tailed (LT) data, while the off-the-shelf pretrain weight of ViTs always leads to unfair comparisons. In this paper, we systematically investigate the ViTs' performance in LTR and propose LiVT to train ViTs from scratch only with LT data. With the observation that ViTs suffer more severe LTR problems, we conduct Masked Generative Pretraining (MGP) to learn generalized features. With ample and solid evidence, we show that MGP is more robust than supervised manners. Although Binary Cross Entropy (BCE) loss performs well with ViTs, it struggles on the LTR tasks. We further propose the balanced BCE to ameliorate it with strong theoretical groundings. Specially, we derive the unbiased extension of Sigmoid and compensate extra logit margins for deploying it. Our Bal-BCE contributes to the quick convergence of ViTs in just a few epochs. Extensive experiments demonstrate that with MGP and Bal-BCE, LiVT successfully trains ViTs well without any additional data and outperforms comparable state-of-the-art methods significantly, e.g., our ViT-B achieves 81.0% Top-1 accuracy in iNaturalist 2018 without bells and whistles. Code is available at https://github.com/XuZhengzhuo/LiVT. Zhengzhuo Xu, Ruikang Liu, Shuo Yang 0011, Zenghao Chai, Chun Yuan 0003 |
CVPR | 1 |
| 2023 | Rethink Long-Tailed Recognition with Vision TransformsabstractIn the real world, data tends to follow long-tailed distributions w.r.t. class or attribution, motivating the challenging Long-Tailed Recognition (LTR) problem. In this paper, we revisit recent LTR methods with promising Vision Transformers (ViT). We figure out that 1) ViT is hard to train with longtailed data. 2) ViT learns generalized features in an unsupervised manner, like mask generative training, either on longtailed or balanced datasets. Hence, we propose to adopt unsupervised learning to utilize long-tailed data. Furthermore, we propose the Predictive Distribution Calibration (PDC) as a novel metric for LTR, where the model tends to simply classify inputs into common classes. Our PDC can measure the model calibration of predictive preferences quantitatively. On this basis, we find many LTR approaches alleviate it slightly, despite the accuracy improvement. Extensive experiments on benchmark datasets validate that PDC reflects the model’s predictive preference precisely, which is consistent with the visualization. Zhengzhuo Xu, Shuo Yang 0011, Chun Yuan 0003 |
ICASSP | 1 |
| 2023 | A Lightweight Approach for Network Intrusion Detection Based on Self-Knowledge DistillationabstractNetwork Intrusion Detection (NID) works as a kernel technology for the security network environment, obtaining extensive research and application. Despite enormous efforts by researchers, NID still faces challenges in deploying on resource-constrained devices. To improve detection accuracy while reducing computational costs and model storage simultaneously, we propose a lightweight intrusion detection approach based on self-knowledge distillation, namely LNet-SKD, which achieves the trade-off between accuracy and efficiency. Specifically, we carefully design the DeepMax block to extract compact representation efficiently and construct the LNet by stacking DeepMax blocks. Furthermore, considering compensating for performance degradation caused by the lightweight network, we adopt batchwise self-knowledge distillation to provide the regularization of training consistency. Experiments on benchmark datasets demonstrate the effectiveness of our proposed LNet-SKD, which outperforms existing state-of-the-art techniques with fewer parameters and lower computation loads. Shuo Yang 0011, Xinran Zheng, Zhengzhuo Xu |
ICC | 3 |
| 2023 | Accurate 3D Face Reconstruction with Facial Component TokensabstractAccurately reconstructing 3D faces from monocular images and videos is crucial for various applications, such as digital avatar creation. However, the current deep learning-based methods face significant challenges in achieving accurate reconstruction with disentangled facial parameters and ensuring temporal stability in single-frame methods for 3D face tracking on video data. In this paper, we propose TokenFace, a transformer-based monocular 3D face reconstruction model. TokenFace uses separate tokens for different facial components to capture information about different facial parameters and employs temporal transformers to capture temporal information from video data. This design can naturally disentangle different facial components and is flexible to both 2D and 3D training data. Trained on hybrid 2D and 3D data, our model shows its power in accurately reconstructing faces from images and producing stable results for video data. Experimental results on popular benchmarks NoWand Stirling demonstrate that TokenFace achieves state-of-the-art performance, outperforming existing methods on all metrics by a large margin. Tianke Zhang, Xuangeng Chu, Yunfei Liu 0001, Lijian Lin, Zhendong Yang, Zhengzhuo Xu, Chengkun Cao, F. Richard Yu, Changyin Zhou, Chun Yuan 0003, Yu Li 0003 |
ICCV | 6 |
| 2023 | HHF: Hashing-Guided Hinge Function for Deep Hashing RetrievalabstractDeep hashing has shown promising performance in large-scale image retrieval. The hashing process utilizes Deep Neural Networks (DNNs) to embed images into compact continuous latent codes, then map them into binary codes by hashing function for efficient retrieval. Recent approaches perform metric loss and quantization loss to supervise the two procedures that cluster samples with the same categories and alleviate semantic information loss after binarization in the end-to-end training framework. However, we observe the incompatible conflict that the optimal cluster positions are not identical to the ideal hash positions because of the different objectives of the two loss terms, which lead to severe ambiguity and error-hashing after the binarization process. To address the problem, we borrow the Theory of Minimum-Distance Bounds for Binary Linear Codes to design the inflection point that depends on the hash bit length and category numbers and thereby propose Hashing-guided Hinge Function (HHF) to explicitly enforce the termination of metric loss to prevent the negative pairs unlimited alienated. Such modification is proven effective and essential for training, which contributes to proper intra- and inter-distances for clusters and better hash positions for accurate image retrieval simultaneously. Extensive experiments in CIFAR-10, CIFAR-100, ImageNet, and MS-COCO justify that HHF consistently outperforms existing techniques and is robust and flexible to transplant into other methods. Code is available athttps://github.com/JerryXu0129/HHF. Chengyin Xu, Zenghao Chai, Zhengzhuo Xu, Qiruyi Zuo, Lingyu Yang, Chun Yuan 0003 |
IEEE Trans. Multim. | 3 |
| 2022 | Semantic-Sparse Colorization Network for Deep Exemplar-Based Colorization
Yunpeng Bai, Chao Dong 0005, Zenghao Chai, Andong Wang, Zhengzhuo Xu, Chun Yuan 0003 |
ECCV (6) | 5 |
| 2022 | REALY: Rethinking the Evaluation of 3D Face Reconstruction
Zenghao Chai, Haoxian Zhang, Jing Ren 0004, Zhengzhuo Xu, Xuefei Zhe, Chun Yuan 0003, Linchao Bao |
ECCV (8) | 5 |
| 2022 | Modernn: Towards Fine-Grained Motion Details for Spatiotemporal Predictive LearningabstractSpatiotemporal predictive learning (ST-PL) aims at predicting the subsequent frames via limited observed sequences, and it has broad applications in the real world. However, learning representative spatiotemporal features for prediction is challenging. Moreover, chaotic uncertainty among consecutive frames exacerbates the difficulty in long-term prediction. This paper concentrates on improving prediction quality by enhancing the correspondence between the previous context and the current state. We carefully design Detail Context Block (DCB) to extract fine-grained details and improve the isolated correlation between upper context state and current input state. We integrate DCB with standard ConvLSTM and introduce Motion Details RNN (MoDeRNN) to capture fine-grained spatiotemporal features and improve the expression of latent states of RNNs to achieve significant quality. Experiments on Moving MNIST and Typhoon datasets demonstrate the effectiveness of the proposed method. MoDeRNN outperforms existing state-of-the-art techniques qualitatively and quantitatively with lower computation loads. Zenghao Chai, Zhengzhuo Xu, Chun Yuan 0003 |
ICASSP | 2 |
| 2022 | CMS-LSTM: Context Embedding and Multi-Scale Spatiotemporal Expression LSTM for Predictive LearningabstractSpatiotemporal predictive learning (ST-PL) is a hotspot with numerous applications, such as object movement and mete-orological prediction. It aims at predicting the subsequent frames via observed sequences. However, inherent uncer-tainty among consecutive frames exacerbates the difficulty in long-term prediction. To tackle the increasing ambigu-ity during forecasting, we design CMS-LSTM to focus on context correlations and multi-scale spatiotemporal flow with details on fine-grained locals, containing two elaborate de-signed blocks: Context Embedding (CE) and Spatiotemporal Expression (SE) blocks. CE is designed for abundant context interactions, while SE focuses on multi-scale spatiotemporal expression in hidden states. The newly introduced blocks also facilitate other spatiotemporal models (e.g., PredRNN, SA-ConvLSTM) to produce representative implicit features for ST-PL and improve prediction quality. Qualitative and quanti-tative experiments demonstrate the effectiveness and flexibil-ity of our proposed method. With fewer params, CMS-LSTM outperforms state-of-the-art methods in numbers of metrics on two representative benchmarks and scenarios. Code is available at https://github.com/czh-98/CMS-LSTM. Zenghao Chai, Zhengzhuo Xu, Yunpeng Bai, Zhihui Lin, Chun Yuan 0003 |
ICME | 2 |
| 2022 | HyP2 Loss: Beyond Hypersphere Metric Space for Multi-label Image RetrievalabstractImage retrieval has become an increasingly appealing technique with broad multimedia application prospects, where deep hashing serves as the dominant branch towards low storage and efficient retrieval. In this paper, we carried out in-depth investigations on metric learning in deep hashing for establishing a powerful metric space in multi-label scenarios, where the pair loss suffers high computational overhead and converge difficulty, while the proxy loss is theoretically incapable of expressing the profound label dependencies and exhibits conflicts in the constructed hypersphere space. To address the problems, we propose a novel metric learning framework with Hybrid Proxy-Pair Loss (HyP$^2$ Loss) that constructs an expressive metric space with efficient training complexity w.r.t. the whole dataset. The proposed HyP$^2$ Loss focuses on optimizing the hypersphere space by learnable proxies and excavating data-to-data correlations of irrelevant pairs, which integrates sufficient data correspondence of pair-based methods and high-efficiency of proxy-based methods. Extensive experiments on four standard multi-label benchmarks justify the proposed method outperforms the state-of-the-art, is robust among different hash bits and achieves significant performance gains with a faster, more stable convergence speed. Our code is available at https://github.com/JerryXu0129/HyP2-Loss. Chengyin Xu, Zenghao Chai, Zhengzhuo Xu, Chun Yuan 0003, Yanbo Fan, Jue Wang 0001 |
ACM Multimedia | 3 |
| 2021 | Towards Calibrated Model for Long-Tailed Visual Recognition from Prior PerspectiveabstractReal-world data universally confronts a severe class-imbalance problem and exhibits a long-tailed distribution, i.e., most labels are associated with limited instances. The naïve models supervised by such datasets would prefer dominant labels, encounter a serious generalization challenge and become poorly calibrated. We propose two novel methods from the prior perspective to alleviate this dilemma. First, we deduce a balance-oriented data augmentation named Uniform Mixup (UniMix) to promote mixup in long-tailed scenarios, which adopts advanced mixing factor and sampler in favor of the minority. Second, motivated by the Bayesian theory, we figure out the Bayes Bias (Bayias), an inherent bias caused by the inconsistency of prior, and compensate it as a modification on standard cross-entropy loss. We further prove that both the proposed methods ensure the classification calibration theoretically and empirically. Extensive experiments verify that our strategies contribute to a better-calibrated model, and their combination achieves state-of-the-art performance on CIFAR-LT, ImageNet-LT, and iNaturalist 2018. Zhengzhuo Xu, Zenghao Chai, Chun Yuan 0003 |
NeurIPS | 1 |
| 2019 | Backscatter-Assisted Hybrid Relaying Strategy for Wireless Powered IoT CommunicationsabstractIn this work, we consider multiple energy harvesting relays to assist information transmission from a hybrid access point (HAP) to a distant receiver. The multi-antenna HAP also beamforms RF power to the relays by using a power-splitting protocol. We aim to maximize the throughput by jointly optimizing the HAP's beamforming strategy as well as individual relays' energy harvesting and collaborative beamforming strategies. With dense user devices, the throughput maximization takes account of the direct links from the HAP to the receiver as they are short and contribute considerably to the overall throughput. Moreover, we introduce the concept of hybrid relaying communications which allows the energy harvesting relays to switch between two radio modes. In particular, the relays can operate either in RF communications or backscatter communications, depending on their channel conditions and energy status. This results in a non-convex and combinatorial throughput maximization problem. With the fixed relay mode, we can find a feasible lower performance bound via convex approximation, which further motivates our algorithm design to update the relay mode in an iterative manner. Simulation results verify that the proposed hybrid relaying strategy can achieve significant performance improvement compared to the conventional relaying strategy with all relays operating in the RF communications mode. Yutong Xie 0003, Zhengzhuo Xu, Shimin Gong, Jing Xu 0005, Dinh Thai Hoang, Dusit Niyato |
GLOBECOM | 2 |