Yongping Xiong

dblp:61/8573 · DBLP profile ↗
← Back
26ranked-venue papers
4as first author
19since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 10 · 10 since 2021Computer networks · 6 · 4 first-authorGraphics, computer vision, multimedia, augmented reality and games · 6 · 6 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Tele-Gaussian: Real-Time Multi-view Generation from a Single Camera for Autostereoscopic Telepresence
Yongping Xiong
ICIC (6)2
2026 ELF: Edit anything for light field displays
Baolin Liu 0002, Zongyuan Yang, Yingde Song, Yongping Xiong
Pattern Recognit.4
2026 CPG: Contrastive Patch-Graph learning for 3D point cloud
Junjie Zhou 0001, Yingde Song, Chinwai Chiu, Yongping Xiong, Yuxin Luo, Siyang Song
Pattern Recognit.4
2025 MegaPairs: Massive Data Synthesis for Universal Multimodal Retrieval
abstract
Despite the rapidly growing demand for multimodal retrieval, progress in this field remains severely constrained by a lack of training data. In this paper, we introduce MegaPairs, a novel data synthesis method that leverages vision language models (VLMs) and open-domain images, together with a massive synthetic dataset generated from this method. Our empirical analysis shows that MegaPairs generates high-quality data, enabling the multimodal retriever to significantly outperform the baseline model trained on 70\times more data from existing datasets. Moreover, since MegaPairs solely relies on general image corpora and open-source VLMs, it can be easily scaled up, enabling continuous improvements in retrieval performance. In this stage, we produced more than 26 million training instances and trained several models of varying sizes using this data. These new models achieve state-of-the-art zero-shot performance across 4 popular composed image retrieval (CIR) benchmarks and the highest overall performance on the 36 datasets provided by MMEB. They also demonstrate notable performance improvements with additional downstream fine-tuning. Our code, synthesized dataset, and pre-trained models are publicly available at https://github.com/VectorSpaceLab/MegaPairs.
Junjie Zhou 0001, Yongping Xiong, Zheng Liu 0011, Shitao Xiao, Yueze Wang, Bo Zhao 0015, Chen Zhang 0013, Defu Lian
ACL (1)2
2025 MLVU: Benchmarking Multi-task Long Video Understanding
abstract
The evaluation of Long Video Understanding (LVU) performance poses an important but challenging research problem. Despite previous efforts, the existing video understanding benchmarks are severely constrained by several issues, especially the insufficient lengths of videos, a lack of diversity in video types and evaluation tasks, and the inappropriateness for evaluating LVU performances. To address the above problems, we propose a new benchmark called MLVU (Multitask Long Video Understanding Benchmark) for the comprehensive and in-depth evaluation of LVU. MLVU presents the following critical values: 1) The substantial and flexible extension of video lengths, which enables the benchmark to evaluate LVU performance across a wide range of durations. 2) The inclusion of various video genres, such as movies, surveillance, egocentric videos, and cartoons, reflects the models’ LVU performances in different scenarios. 3) The development of diversified evaluation tasks, which enables a comprehensive examination of MLLMs’ key abilities in long-video understanding. The empirical study with 23 latest MLLMs reveals significant room for improvement in today’s technique, as all existing methods struggle with most of the evaluation tasks and exhibit severe performance degradation when handling longer videos. Additionally, it suggests that factors such as context length, image-understanding ability, and the choice of LLM backbone can play critical roles in future advancements. We anticipate that MLVU will advance the research of LVU by providing a comprehensive and in-depth analysis of MLLMs. The code and dataset can be accessed from https://github.com/JUNJIE99/MLVU.
Junjie Zhou 0001, Bo Zhao 0015, Boya Wu, Zhengyang Liang, Shitao Xiao, Minghao Qin, Yongping Xiong, Tiejun Huang 0001, Zheng Liu 0011
CVPR9
2025 RF-PBFT: A Dynamic Consensus Algorithm Based on Reputation Partitioned Clusters
Yongping Xiong
ICA3PP (6)2
2025 EYE3: Turn Anything into Naked-Eye 3D
Yingde Song, Zongyuan Yang, Baolin Liu 0002, Yongping Xiong, Sai Chen, Lan Yi, Zhaohe Zhang, Xunbo Yu
ICCV4
2025 TextDiff: Enhancing scene text image super-resolution with mask-guided residual diffusion models
Baolin Liu 0002, Zongyuan Yang, Chinwai Chiu, Yongping Xiong
Pattern Recognit.4
2024 VISTA: Visualized Text Embedding For Universal Multi-Modal Retrieval
abstract
Multi-modal retrieval becomes increasingly popular in practice.However, the existing retrievers are mostly text-oriented, which lack the capability to process visual information.Despite the presence of vision-language models like CLIP, the current methods are severely limited in representing the text-only and imageonly data.In this work, we present a new embedding model VISTA for universal multimodal retrieval.Our work brings forth threefold technical contributions.Firstly, we introduce a flexible architecture which extends a powerful text encoder with the image understanding capability by introducing visual token embeddings.Secondly, we develop two data generation strategies, which bring highquality composed image-text to facilitate the training of the embedding model.Thirdly, we introduce a multi-stage training algorithm, which first aligns the visual token embedding with the text encoder using massive weakly labeled data, and then develops multi-modal representation capability using the generated composed image-text data.In our experiments, VISTA achieves superior performances across a variety of multi-modal retrieval tasks in both zero-shot and supervised settings.Our model, data, and source code are available at https://github.com/FlagOpen/FlagEmbedding.
Junjie Zhou 0001, Zheng Liu 0011, Shitao Xiao, Bo Zhao 0015, Yongping Xiong
ACL (1)5
2024 LayoutPointer: A Spatial-Context Adaptive Pointer Network for Visual Information Extraction
abstract
Huang Siyuan, Yongping Xiong, Wu Guibin. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Yongping Xiong, Guibin Wu
NAACL-HLT2
2024 GDB: Gated Convolutions-based Document Binarization
Zongyuan Yang, Baolin Liu 0002, Yongping Xiong, Guibin Wu
Pattern Recognit.3
2024 FAT: Field-Aware Transformer for Point Cloud Segmentation With Adaptive Attention Fields
abstract
Point cloud segmentation is crucial for various industrial applications, such as autonomous driving and robotics. Recent developments underscore the significant potential of transformer models in this field. However, existing attention mechanisms apply the same feature learning paradigm for all points equally, ignoring the considerable size differences among objects in a scene. To rectify this, we introduce the field-aware transformer (FAT), engineered to tailor effective receptive fields to objects of varying sizes. Our FAT achieves field-aware learning through two primary components: the multigranularity attention (MGA) scheme and the reattention module. The MGA scheme is proficient in aggregating tokens from distant areas while preserving multiscale features within each attention layer. The reattention module dynamically adjusts the attention scores to the fine- and coarse-grained features output by MGA for each point. Extensive experimental results underscore the effectiveness and efficiency of our FAT, which delivers state-of-the-art performance on both the stanford 3D indoor scene dataset (S3DIS) and ScanNetV2 datasets.
Junjie Zhou 0001, Baolin Liu 0002, Yongping Xiong, Chinwai Chiu, Xiangyang Gong
IEEE Trans. Ind. Informatics3
2024 DirectL: Efficient Radiance Fields Rendering for 3D Light Field Displays
abstract
Autostereoscopic display technology, despite decades of development, has not achieved extensive application, primarily due to the daunting challenge of three-dimensional (3D) content creation for non-specialists. The emergence of Radiance Field as an innovative 3D representation has markedly revolutionized the domains of 3D reconstruction and generation, simplifying 3D content creation for common users and broadening the applicability of Light Field Displays (LFDs). However, the combination of these two technologies remains largely unexplored. The standard paradigm to create optimal content for parallax-based light field displays demands rendering at least 45 slightly shifted views preferably at high resolution per frame, a substantial hurdle for real-time rendering. We introduce DirectL, a novel rendering paradigm for Radiance Fields on autostereoscopic displays with lenticular lens. By thoroughly analyzing the interleaved mapping of spatial rays to screen sub-pixels, we accurately render only the light rays entering the human eye and propose subpixel repurposing to significantly reduce the pixel count required for rendering. Tailored for the two predominant radiance fields---Neural Radiance Fields (NeRFs) and 3D Gaussian Splatting (3DGS), we propose corresponding optimized rendering pipelines that directly render the light field images instead of multi-view images, achieving state-of-the-art rendering speeds on autostereoscopic displays. Extensive experiments across various autostereoscopic displays and user visual perception assessments demonstrate that DirectL accelerates rendering by up to 40 times compared to the standard paradigm without sacrificing visual quality. Its rendering process-only modification allows seamless integration into subsequent radiance field tasks. Finally, we incorporate DirectL into diverse applications, showcasing the stunning visual experiences and the synergy between Light Field Displays and Radiance Fields, which reveals the immense potential for application prospects. DirectL Project Homepage: direct-l.github.io
Zongyuan Yang, Baolin Liu 0002, Yingde Song, Lan Yi, Yongping Xiong, Zhaohe Zhang, Xunbo Yu
ACM Trans. Graph.5
2023 Document Binarization with Multi-Branch Gated Convolutional Generative Adversarial Networks
abstract
Existing document binarization methods can not extract stroke edges finely, mainly due to the fair-treatment nature of vanilla convolutions and the extraction of stroke edges without adequate supervision by boundary-related information. In this paper, we formulate text extraction as the learning of gating values and propose a novel end-to-end gated convolutions-based network (GDB) to solve the problem of imprecise stroke edge extraction. The gated convolutions are applied to selectively extract the features of strokes with different attention. Firstly, a coarse sub-network with an extra edge branch is trained to get more precise feature maps by feeding a priori mask and edge. Secondly, a refinement sub-network is cascaded to refine the output of the first stage by gated convolutions based on the sharp edge. For global information, GDB also contains a multi-scale operation to combine local and global features. Experimental results show that our proposed methods outperform the SOTA methods in terms of all metrics on average over all DIBCO datasets from 2009 to 2019 and achieve top ranking on six benchmark datasets. Available codes: https://github.com/Royalvice/GDB.
Zongyuan Yang, Yongping Xiong, Guibin Wu
ICIP2
2023 Fat: Field-Aware Transformer for 3D Point Cloud Semantic Segmentation
abstract
Transformer models have achieved promising performances in point cloud segmentation. However, most existing attention schemes provide the same feature learning paradigm for all points equally and overlook the enormous difference in size among scene objects. In this paper, we propose the Field-Aware Transformer (FAT) that adjusts the attentive receptive fields for objects of different sizes. Our FAT achieves field-aware learning via two steps: introduce multi-granularity features to each attention layer and allow each point to choose its attentive fields adaptively. It contains two key designs: the Multi-Granularity Attention (MGA) scheme and the Re-Attention module. Extensive experimental results demonstrate that FAT achieves state-of-the-art performances on S3DIS [1] and ScanNetV2 [2] datasets.
Junjie Zhou 0001, Yongping Xiong, Chinwai Chiu, Xiangyang Gong
ICIP2
2023 DocDiff: Document Enhancement via Residual Diffusion Models
abstract
Removing degradation from document images not only improves their visual quality and readability, but also enhances the performance of numerous automated document analysis and recognition tasks. However, existing regression-based methods optimized for pixel-level distortion reduction tend to suffer from significant loss of high-frequency information, leading to distorted and blurred text edges. To compensate for this major deficiency, we propose DocDiff, the first diffusion-based framework specifically designed for diverse challenging document enhancement problems, including document deblurring, denoising, and removal of watermarks and seals. DocDiff consists of two modules: the Coarse Predictor (CP), which is responsible for recovering the primary low-frequency content, and the High-Frequency Residual Refinement (HRR) module, which adopts the diffusion models to predict the residual (high-frequency information, including text edges), between the ground-truth and the CP-predicted image. DocDiff is a compact and computationally efficient model that benefits from a well-designed network architecture, an optimized training loss objective, and a deterministic sampling process with short time steps. Extensive experiments demonstrate that DocDiff achieves state-of-the-art (SOTA) performance on multiple benchmark datasets, and can significantly enhance the readability and recognizability of degraded document images. Furthermore, our proposed HRR module in pre-trained DocDiff is plug-and-play and ready-to-use, with only 4.17M parameters. It greatly sharpens the text edges generated by SOTA deblurring methods without additional joint training. Available codes: https://github.com/Royalvice/DocDiff https://github.com/Royalvice/DocDiff.
Zongyuan Yang, Baolin Liu 0002, Yongping Xiong, Lan Yi, Guibin Wu, Junjie Zhou 0001
ACM Multimedia3
2022 CarveNet: a channel-wise attention-based network for irregular scene text recognition
Guibin Wu, Zheng Zhang 0038, Yongping Xiong
Int. J. Document Anal. Recognit.3
2022 TableRobot: an automatic annotation method for heterogeneous tables
abstract
Abstract Using deep learning networks to recognize the table attracts lots of attention. However, due to the lack of high-quality table datasets, the performance of using deep learning networks is limited. Therefore, TableRobot has been proposed, an automatic annotation method for heterogeneous tables. To be more specific, the annotations of table consist of the coordinates of the item block and the mapping relationship between item blocks and table cells. In order to transform the task, we successfully design an algorithm based on the greedy approach to find the optimum solution. To evaluate the performance of TableRobot, we check the annotation data of 3000 tables collected from the LaTex documents in arXiv.com , and the result shows that TableRobot can generate table annotation datasets with the accuracy of 93.2%. Besides, the table annotation data is feed into GraphTSR which is a state-of-the-art table recognition graph neural network, and the F1 value of the network has increased by nearly 10% compared with before.
Guibin Wu, Yongping Xiong, Chaoyi Zhou
Pers. Ubiquitous Comput.3
2022 A Scalable Graph-Based Framework for Multi-Organ Histology Image Classification
abstract
Graph-based approaches are successful for histology image classification tasks but still face many challenges, such as: 1) the lack of nuclei-level labels and the significant variations between histology images make it extremely difficult to extract discriminative high-level nuclei features like nuclei type, texture and micro-environment; 2) graph-based approaches cannot handle large-scale cell graph nodes typically contained in histology images; and 3) graph neural networks (GNNs) struggle to learn the long-range dependency of cell graphs. To address the above challenges, we propose a scalable graph-based framework for multi-organ histology image classification. We develop a two-step masked nuclei patches supervised training approach to extract discriminative high-level nuclei features for histology images without nuclei-level labels. Additionally, we introduce a nuclei sampling strategy to make our graph-based framework scalable for large-scale cell graphs. Furthermore, we proposeHierArchicalTransformer Graph NeuralNetwork (HAT-Net+) for cell graph classi- fications. HAT-Net+ adopts Transformer to model the long-range dependency of cell graphs and a parameter-free approach to adaptively fuse different hierarchical graph representations of each layer. We achieved the state-of-the-art results on four public histology image classification datasets: CRC dataset (100%), Extended CRC dataset (98%), UZH dataset (96.9%) and BACH dataset (88%). Unlike other methods, our approach can be used in various histology image classification tasks, even for images without nuclei-level labels, indicating its potential in cancer diagnosis. The code is available athttps://github.com/suyouooooo/HAT-Net.
Yu Bai 0020, Yue Mi, Yihan Su, Bo Zhang 0032, Zheng Zhang 0038, Jingyun Wu, Haiwen Huang, Yongping Xiong, Xiangyang Gong, Wendong Wang 0003
IEEE J. Biomed. Health Informatics8
2015 A three-dimensional sub-region query processing mechanism in underwater WSNs
Zhangbing Zhou, Riliang Xing, Walid Gaaloul, Yongping Xiong
Pers. Ubiquitous Comput.4
2014 ZiFi: Exploiting Cross-Technology Interference Signatures for Wireless LAN Discovery
abstract
Wi-Fi networks have enjoyed an unprecedent penetration rate in recent years. However, due to the limited coverage, existing Wi-Fi infrastructure only provides intermittent connectivity for mobile users. Once leaving the current network coverage, Wi-Fi clients must actively discover new Wi-Fi access points (APs), which wastes the precious energy of mobile devices. Although several solutions have been proposed to address this issue, they either require significant modifications to existing network infrastructures or rely on context information that is not available in unknown environments. In this work, we develop a system called ZiFithat utilizes ZigBee radios to identify the existence of Wi-Fi networks through unique interference signatures generated by Wi-Fi beacons. We develop a new digital signal processing algorithm called common multiple folding (CMF) that accurately amplifies periodic beacons in Wi-Fi interference signals. ZiFi also adopts a constant false alarm rate (CFAR) detector that can minimize the false negative (FN) rate of Wi-Fi beacon detection while satisfying the user-specified upper bound on false positive (FP) rate. We have implemented ZiFi on two platforms, a Linux netbook integrating a TelosB mote through the USB interface, and a Nokia N73 smartphone integrating a ZigBee card through the miniSD interface. Our experiments show that, under typical settings, ZiFi can detect Wi-Fi APs with high accuracy (<;5 percent total FP and FN rate), short delay (~780 ms), and little computation overhead.
Yongping Xiong, Ruogu Zhou, Minming Li, Guoliang Xing, Limin Sun 0001, Jian Ma 0001
IEEE Trans. Mob. Comput.1
2012 An efficient and security dynamic identity based authentication protocol for multi-server architecture using smart cards
Xiong Li 0002, Yongping Xiong, Jian Ma 0001, Wendong Wang 0003
J. Netw. Comput. Appl.2
2011 Phoenix: Peer-to-Peer Location Based Notification in Mobile Networks
abstract
Location Based Notification (LBN) aims to alert the users in a target area with the information of interest to them. With a wide range of applications, LBN has been gaining more and more attraction among wireless users and service providers. The mainstream centralized solution based on cellular networks may incur high service cost. In this paper, we present an innovative scheme called Phoenix, which does not rely on any infrastructure, to implement for location based notification service. In our design, devices (users) across the target area form a dynamic peer-to-peer network, where a user can be a message source, a message carrier, or a message subscriber. When a user meets the message carrier, the user can get a copy of the message. Phoenix keeps messages of interest being circulated in the target area, hence users are being notified. To achieve desired notification performance, Phoenix adaptively controls when a user should take the carrier role and help disseminating a message in order to keep the message ``alive", given the fact that message carriers may leave the target area and drop the message. Extensive simulations have been conducted to show the efficacy of Phoenix notification system.
Yongping Xiong, Canfeng Chen, Jian Ma 0001, Limin Sun 0001
MASS1
2010 Anycast routing in mobile opportunistic networks
abstract
A mobile opportunistic network consists of sparsely scattered mobile nodes communicating via short range radios. It is characterized by frequent and unpredictable network partitions and intermittent connectivity. Anycast in opportunistic networks is anticipated in many application scenarios, and deserves great attention. In this paper, we propose an anycast routing algorithm in which each node is associated with a forwarding metric indicating its delivery probability to the destination anycast group and the node with lower value hands over the message to the encountered node with higher metric. The forwarding metric is determined according to historical node encounter information. We use three different forwarding metrics (variables) to guide the transmission of messages. The Group Forwarding Metric (GFM) treats the entire group as a whole, and it is defined as the probability of meeting any member in the anycast group to deliver a message. Similar to GFM, Probability Forwarding Metric (PFM) is defined as the probability of encountering at least one anycast group member, but it relies on the probability of meeting individual group members. The Distance Forwarding Metric (DFM) takes a function of the delivery probability to an anycast group member as the distance to the member. The DFM is the combination of these distances to forward messages towards the higher member density. Different metrics can be adopted for different mobile opportunistic networks based on the connectivity characteristics of the networks. We analyze the control overhead of the anycast algorithm and the message delivery delay of the routing protocol. Extensive simulations are carried out to evaluate the performance of the proposed solution under synthetic and realistic traces. The results show that our algorithm will significantly improve the anycast delivery performance when compared with simple routing algorithms in term of average message delivery delay and transmission overhead.
Yongping Xiong, Limin Sun 0001, Jian Ma 0001
ISCC1
2010 Optimal infostation deployment for spatio-temporal information dissemination
abstract
A growing number of applications require disseminating information around specific geographical areas within a limited valid time. For example, the store in the mall area expects to publish the time-limited sales promotion to all the potential clients in the nearby area, before the discount activity end. In this paper, we study the problem of deploying infostation for geographical information dissemination. It aims to achieve the desired dissemination ratio under the given time constraint and to minimize the infostation deployment cost. Inspired by several observations in recent studies on realistic mobility model, we build a mobility graph to reflect the statistical characteristic of users movement in a area. Based on this graph, we formulate the infostation deployment problem as an optimization problem. Then, we prove it is NP-hard by reducing it to the classical vertex cover problem and then develop a greedy heuristic algorithm DGREEDY with the polynomial time complexity. Extensive simulations based on the real human mobility traces have been carried out to show the efficacy of our approach.
Yongping Xiong, Jian Ma 0001, Yan Liu 0021, Limin Sun 0001
ISCC1
2010 ZiFi: wireless LAN discovery via ZigBee interference signatures
abstract
WiFi networks have enjoyed an unprecedent penetration rate in recent years. However, due to the limited coverage, existing WiFi infrastructure only provides intermittent connectivity for mobile users. Once leaving the current network coverage, WiFi clients must actively discover new WiFi access points (APs), which wastes the precious energy of mobile devices. Although several solutions have been proposed to address this issue, they either require significant modifications to existing network infrastructures or rely on context information that is not available in unknown environments. In this work, we develop a system called ZiFi that utilizes ZigBee radios to identify the existence of WiFi networks through unique interference signatures generated by WiFi beacons. We develop a new digital signal processing algorithm called Common Multiple Folding (CMF) that accurately amplifies periodic beacons in WiFi interference signals. ZiFi also adopts a constant false alarm rate (CFAR) detector that can minimize the false negative (FN) rate of WiFi beacon detection while satisfying the user-specified upper bound on false positive (FP) rate. We have implemented ZiFi on two platforms, a Linux netbook integrating a TelosB mote through the USB interface, and a Nokia N73 smartphone integrating a ZigBee card through the miniSD interface. Our experiments show that, under typical settings, ZiFi can detect WiFi APs with high accuracy (<5% total FP and FN rate), short delay (~780 ms), and little computation overhead
Ruogu Zhou, Yongping Xiong, Guoliang Xing, Limin Sun 0001, Jian Ma 0001
MobiCom2