Xudong Ma

dblp:19/2951 · DBLP profile ↗
← Back
37ranked-venue papers
6as first author
14since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 22 · 3 first-author · 10 since 2021Systems, architecture and hardware · 14 · 2 first-authorGraphics, computer vision, multimedia, augmented reality and games · 7 · 2 first-author · 3 since 2021Databases, data management, data science and information retrieval · 4 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 first-authorSecurity and privacy · 1
YearPublicationVenuePosition
2026 Robust Knowledge Graph Embedding via Denoising
Tengwei Song, Xudong Ma, Yang Liu 0450, Jie Luo 0004, Robert Hoehndorf
ESWC (1)2
2025 BinaryDM: Accurate Weight Binarization for Efficient Diffusion Models
abstract
With the advancement of diffusion models (DMs) and the substantially increased computational requirements, quantization emerges as a practical solution to obtain compact and efficient low-bit DMs. However, the highly discrete representation leads to severe accuracy degradation, hindering the quantization of diffusion models to ultra-low bit-widths. This paper proposes a novel weight binarization approach for DMs, namely BinaryDM, pushing binarized DMs to be accurate and efficient by improving the representation and optimization. From the representation perspective, we present an Evolvable-Basis Binarizer (EBB) to enable a smooth evolution of DMs from full-precision to accurately binarized. EBB enhances information representation in the initial stage through the flexible combination of multiple binary bases and applies regularization to evolve into efficient single-basis binarization. The evolution only occurs in the head and tail of the DM architecture to retain the stability of training. From the optimization perspective, a Low-rank Representation Mimicking (LRM) is applied to assist the optimization of binarized DMs. The LRM mimics the representations of full-precision DMs in low-rank space, alleviating the direction ambiguity of the optimization process caused by fine-grained alignment. Comprehensive experiments demonstrate that BinaryDM achieves significant accuracy and efficiency gains compared to SOTA quantization methods of DMs under ultra-low bit-widths. With 1-bit weight and 4-bit activation (W1A4), BinaryDM achieves as low as 7.74 FID and saves the performance from collapse (baseline FID 10.87). As the first binarization method for diffusion models, W1A4 BinaryDM achieves impressive 15.2x OPs and 29.2x model size savings, showcasing its substantial potential for edge deployment.
Xingyu Zheng, Xianglong Liu 0001, Haotong Qin, Xudong Ma, Haojie Hao, Jiakai Wang, Zixiang Zhao, Jinyang Guo 0002, Michele Magno
ICLR4
2025 Clustering Properties of Self-Supervised Learning
abstract
Self-supervised learning (SSL) methods via joint embedding architectures have proven remarkably effective at capturing semantically rich representations with strong clustering properties, magically in the absence of label supervision. Despite this, few of them have explored leveraging these untapped properties to improve themselves. In this paper, we provide an evidence through various metrics that the encoder's output *encoding* exhibits superior and more stable clustering properties compared to other components. Building on this insight, we propose a novel positive-feedback SSL method, termed **Re**presentation **S**elf-**A**ssignment (ReSA), which leverages the model's clustering properties to promote learning in a self-guided manner. Extensive experiments on standard SSL benchmarks reveal that models pretrained with ReSA outperform other state-of-the-art SSL methods by a significant margin. Finally, we analyze how ReSA facilitates better clustering properties, demonstrating that it effectively enhances clustering performance at both fine-grained and coarse-grained levels, shaping representations that are inherently more structured and semantically meaningful.
Xi Weng, Jianing An, Xudong Ma, Binhang Qi, Jie Luo 0004, Jin Song Dong 0001, Lei Huang 0015
ICML3
2025 BiVM: Accurate Binarized Neural Network for Efficient Video Matting
abstract
Deep neural networks for real-time video matting suffer significant computational limitations on edge devices, hindering their adoption in widespread applications such as online conferences and short-form video production. Binarization emerges as one of the most common neural network compression approaches, substantially curbing computational and memory requirements through compact 1-bit parameters and efficient bitwise operations. However, the empirical observation reveals that accuracy and efficiency limitations exist in the binarized video matting network due to its degenerated encoder and redundant decoder. Following a theoretical analysis based on the information bottleneck principle, the limitations are mainly caused by the degradation of prediction-relevant information in the intermediate features and the redundant computation in prediction-irrelevant areas. We presentBiVM, an accurate and resource-efficientBinarized neural network forVideoMatting, where architecture and optimization allow real-time video matting to proceed on edge hardware. First, we present a series of binarized computation structures with elastic shortcuts and evolvable topologies, enabling the constructed encoder backbone to extract high-quality representation from input videos for accurate prediction. Second, we sparse the intermediate feature of the binarized decoder by masking homogeneous parts, allowing the decoder to focus on representation with diverse details while alleviating the computation burden for efficient inference. Furthermore, we construct a localized binarization-aware mimicking framework with the information-guided strategy, prompting matting-related representation in full-precision counterparts to be accurately and fully utilized. Comprehensive experiments show that the proposed BiVM surpasses alternative binarized video matting networks, including state-of-the-art (SOTA) binarization methods, by a substantial margin. For example, BiVM surpasses 16.67 MAD compared to SOTA binarization on the VM dataset. Notably, our approach can even perform comparably to the full-precision counterpart in terms of visual quality. Moreover, our BiVM achieves significant savings of 14.3× and 21.6× in computation and storage costs, respectively. We also evaluate BiVM on ARM CPU hardware, underscoring its potential for deployment in resource-constrained scenarios.
Haotong Qin, Xianglong Liu 0001, Xudong Ma, Lei Ke, Yulun Zhang 0001, Jie Luo 0004, Michele Magno
IEEE Trans. Pattern Anal. Mach. Intell.3
2025 Deep Hashing with Semantic Hash Centers for Image Retrieval
abstract
Deep hashing presents an effective strategy for large-scale image retrieval. Current hashing methods are generally categorized by their supervision types: point-wise, pairwise, and list-wise. Recent advancements in point-wise methods (e.g., CSQ, MDS) have significantly enhanced retrieval performance across diverse datasets by pre-assigning a hash center to each class, thereby improving the discriminability of the resultant hash codes. However, these methods employ purely data-independent algorithms for generating hash centers, overlooking the semantic connections between different classes, which, we argue, could degrade retrieval performance. To tackle this problem, this article expands on the newly emerged concept of “hash centers” to introduce “ semantic hash centers,” which posits that hash centers of semantically related classes should exhibit closer Hamming distances, while those of unrelated classes should be more distant. Based on this hypothesis, we propose a three-stage framework, termed Semantic Hash Centers (SHC), to produce hash codes that preserve semantics. First, we build a classification network to detect semantic similarities between classes, and utilize a data-dependent approach to similarity calculation that can adapt to varied data distributions. Next, we develop a new optimization algorithm to generate SHC. This algorithm not only maintains semantic relatedness among hash centers but also integrates a constraint to ensure a minimum distance between them, addressing the issue of excessively proximate hash centers potentially impairing retrieval performance. Finally, we train a deep hashing network with the above generated SHC to convert each image into a binary hash code. Experiments on large-scale image retrieval across several public datasets demonstrate that SHC generates more discriminative hash codes, markedly enhancing retrieval performance. Specifically, in terms of the mAP@100, mAP@1000, and mAP@ALL metrics, SHC records average improvements of +6.24%, +6.68%, and +10.39%, respectively, over the most competitive existing methods. The code of our SHC project is available at https://github.com/cc752424640/Deep-Hashing-with-Semantic-Hash-Centers-for-Image-Retrieval .
Rui Liu 0007, Xudong Ma, Yong Chen 0008, Dell Zhang
ACM Trans. Inf. Syst.4
2024 Accurate LoRA-Finetuning Quantization of LLMs via Information Retention
abstract
The LoRA-finetuning quantization of LLMs has been extensively studied to obtain accurate yet compact LLMs for deployment on resource-constrained hardware. However, existing methods cause the quantized LLM to severely degrade and even fail to benefit from the finetuning of LoRA. This paper proposes a novel IR-QLoRA for pushing quantized LLMs with LoRA to be highly accurate through information retention. The proposed IR-QLoRA mainly relies on two technologies derived from the perspective of unified information: (1) statistics-based Information Calibration Quantization allows the quantized parameters of LLM to retain original information accurately; (2) finetuning-based Information Elastic Connection makes LoRA utilizes elastic representation transformation with diverse information. Comprehensive experiments show that IR-QLoRA can significantly improve accuracy across LLaMA and LLaMA2 families under 2-4 bit-widths, e.g., 4-bit LLaMA-7B achieves 1.4% improvement on MMLU compared with the state-of-the-art methods. The significant performance gain requires only a tiny 0.31% additional time consumption, revealing the satisfactory efficiency of our IR-QLoRA. We highlight that IR-QLoRA enjoys excellent versatility, compatible with various frameworks (e.g., NormalFloat and Integer quantization) and brings general accuracy gains. The code is available at https://github.com/htqin/ir-qlora .
Haotong Qin, Xudong Ma, Xingyu Zheng, Yang Zhang 0088, Shouda Liu, Jie Luo 0004, Xianglong Liu 0001, Michele Magno
ICML2
2024 BiDM: Pushing the Limit of Quantization for Diffusion Models
abstract
Diffusion models (DMs) have been significantly developed and widely used in various applications due to their excellent generative qualities. However, the expensive computation and massive parameters of DMs hinder their practical use in resource-constrained scenarios. As one of the effective compression approaches, quantization allows DMs to achieve storage saving and inference acceleration by reducing bit-width while maintaining generation performance. However, as the most extreme quantization form, 1-bit binarization causes the generation performance of DMs to face severe degradation or even collapse. This paper proposes a novel method, namely BiDM, for fully binarizing weights and activations of DMs, pushing quantization to the 1-bit limit. From a temporal perspective, we introduce the Timestep-friendly Binary Structure (TBS), which uses learnable activation binarizers and cross-timestep feature connections to address the highly timestep-correlated activation features of DMs. From a spatial perspective, we propose Space Patched Distillation (SPD) to address the difficulty of matching binary features during distillation, focusing on the spatial locality of image generation tasks and noise estimation networks. As the first work to fully binarize DMs, the W1A1 BiDM on the LDM-4 model for LSUN-Bedrooms 256$\times$256 achieves a remarkable FID of 22.74, significantly outperforming the current state-of-the-art general binarization methods with an FID of 59.44 and invalid generative samples, and achieves up to excellent 28.0 times storage and 52.7 times OPs savings.
Xingyu Zheng, Xianglong Liu 0001, Yichen Bian, Xudong Ma, Yulun Zhang 0001, Jiakai Wang, Jinyang Guo 0002, Haotong Qin
NeurIPS4
2024 Gaze-assisted visual grounding via knowledge distillation for referred object grasping with under-specified object referring
Zhuoyang Zhang, Kun Qian 0005, Bo Zhou 0017, Fang Fang 0006, Xudong Ma
Eng. Appl. Artif. Intell.5
2024 BiFSMNv2: Pushing Binary Neural Networks for Keyword Spotting to Real-Network Performance
abstract
Deep neural networks, such as the deep-FSMN, have been widely studied for keyword spotting (KWS) applications while suffering expensive computation and storage. Therefore, network compression technologies such as binarization are studied to deploy KWS models on edge. In this article, we present a strong yet efficient binary neural network for KWS, namely, BiFSMNv2, pushing it to the real-network accuracy performance. First, we present a dual-scale thinnable 1-bit-architecture (DTA) to recover the representation capability of the binarized computation units by dual-scale activation binarization and liberate the speedup potential from an overall architecture perspective. Second, we also construct a frequency-independent distillation (FID) scheme for KWS binarization-aware training, which distills the high- and low-frequency components independently to mitigate the information mismatch between full-precision and binarized representations. Moreover, we propose the learning propagation binarizer (LPB), a general and efficient binarizer that enables the forward and backward propagation of binary KWS networks to be continuously improved through learning. We implement and deploy BiFSMNv2 on ARMv8 real-world hardware with a novel fast bitwise computation kernel (FBCK), which is proposed to fully use registers and increase instruction throughput. Comprehensive experiments show our BiFSMNv2 outperforms the existing binary networks for KWS by convincing margins across different datasets and achieves comparable accuracy with the full-precision networks (only a tiny 1.51% drop on Speech Commands V1-12). We highlight that benefiting from the compact architecture and optimized hardware kernel, BiFSMNv2 can achieve an impressive 25.1× speedup and 20.2× storage-saving on edge hardware.
Haotong Qin, Xudong Ma, Yifu Ding 0001, Yang Zhang 0088, Zejun Ma 0001, Jiakai Wang, Jie Luo 0004, Xianglong Liu 0001
IEEE Trans. Neural Networks Learn. Syst.2
2023 Dual Channel Knowledge Graph Embedding with Ontology Guided Data Augmentation
Tengwei Song, Long Yin, Xudong Ma, Jie Luo 0004
KSEM (1)3
2023 BiMatting: Efficient Video Matting via Binarization
abstract
Real-time video matting on edge devices faces significant computational resource constraints, limiting the widespread use of video matting in applications such as online conferences and short-form video production. Binarization is a powerful compression approach that greatly reduces computation and memory consumption by using 1-bit parameters and bitwise operations. However, binarization of the video matting model is not a straightforward process, and our empirical analysis has revealed two primary bottlenecks: severe representation degradation of the encoder and massive redundant computations of the decoder. To address these issues, we propose BiMatting, an accurate and efficient video matting model using binarization. Specifically, we construct shrinkable and dense topologies of the binarized encoder block to enhance the extracted representation. We sparsify the binarized units to reduce the low-information decoding computation. Through extensive experiments, we demonstrate that BiMatting outperforms other binarized video matting models, including state-of-the-art (SOTA) binarization methods, by a significant margin. Our approach even performs comparably to the full-precision counterpart in visual quality. Furthermore, BiMatting achieves remarkable savings of 12.4$\times$ and 21.6$\times$ in computation and storage, respectively, showcasing its potential and advantages in real-world resource-constrained scenarios. Our code and models are released at https://github.com/htqin/BiMatting .
Haotong Qin, Lei Ke, Xudong Ma, Martin Danelljan, Yu-Wing Tai, Chi-Keung Tang, Xianglong Liu 0001, Fisher Yu 0001
NeurIPS3
2022 ICIP 2022 Challenge on Parasitic Egg Detection and Classification in Microscopic Images: Dataset, Methods and Results
abstract
Manual examination of faecal smear samples to identify the existence of parasitic eggs is very time-consuming and can only be done by specialists. Therefore, an automated system is required to tackle this problem since it can relate to serious intestinal parasitic infections. This paper reviews the ICIP 2022 Challenge on parasitic egg detection and classification in microscopic images. We describe a new dataset for this application, which is the largest dataset of its kind. The methods used by participants in the challenge are summarised and discussed along with their results.
Nantheera Anantrasirichai, Thanarat H. Chalidabhongse, Duangdao Palasuwan, Korranat Naruenatthanaset, Thananop Kobchaisawat, Nuntiporn Nunthanasup, Kanyarat Boonpeng, Xudong Ma, Alin Achim
ICIP8
2022 Unsupervised Image Fusion Using Deep Image Priors
abstract
A significant number of researchers have applied deep learning methods to image fusion. However, most works require a large amount of training data or depend on pre-trained models or frameworks to capture features from source images. This is inevitably hampered by a shortage of training data or a mismatch between the framework and the actual problem. Deep Image Prior (DIP) has been introduced to exploit convolutional neural networks’ ability to synthesize the ‘prior’ in the input image. However, the original design of DIP is hard to be generalized to multi-image processing problems, particularly for image fusion. Therefore, we propose a new image fusion technique that extends DIP to fusion tasks formulated as inverse problems. Additionally, we apply a multichannel approach to enhance DIP’s effect further. The evaluation is conducted with several commonly used image fusion assessment metrics. The results are compared with state-of-the-art image fusion methods. Our method outperforms these techniques for a range of metrics. In particular, it is shown to provide the best objective results for most metrics when applied to medical images.
Xudong Ma, Paul R. Hill, Nantheera Anantrasirichai, Alin Achim
ICIP1
2022 BiFSMN: Binary Neural Network for Keyword Spotting
abstract
The deep neural networks, such as the Deep-FSMN, have been widely studied for keyword spotting (KWS) applications. However, computational resources for these networks are significantly constrained since they usually run on-call on edge devices. In this paper, we present BiFSMN, an accurate and extreme-efficient binary neural network for KWS. We first construct a High-frequency Enhancement Distillation scheme for the binarization-aware training, which emphasizes the high-frequency information from the full-precision network's representation that is more crucial for the optimization of the binarized network. Then, to allow the instant and adaptive accuracy-efficiency trade-offs at runtime, we also propose a Thinnable Binarization Architecture to further liberate the acceleration potential of the binarized network from the topology perspective. Moreover, we implement a Fast Bitwise Computation Kernel for BiFSMN on ARMv8 devices which fully utilizes registers and increases instruction throughput to push the limit of deployment efficiency. Extensive experiments show that BiFSMN outperforms existing binarization methods by convincing margins on various datasets and is even comparable with the full-precision counterpart (e.g., less than 3% drop on Speech Commands V1-12). We highlight that benefiting from the thinnable architecture and the optimized 1-bit implementation, BiFSMN can achieve an impressive 22.3x speedup and 15.5x storage-saving on real-world edge hardware.
Haotong Qin, Xudong Ma, Yifu Ding 0001, Yang Zhang 0088, Zejun Ma 0001, Jie Luo 0004, Xianglong Liu 0001
IJCAI2
2019 Dual-loop Control for Three-phase Vienna Rectifier with Duty-ratio Feedforward
abstract
As Boost-type topology, three-phase Vienna rectifier has the inherent problem of increasing output voltage at partial load. In this paper, a duty-ratio feedforward strategy based on dual-loop control is proposed to solve this problem. The comparison between a three-phase three-wire system and a three-phase four-wire system reveals that the main feature of the former is the zero-sequence component, which will not affect THD. It provides a theoretical basis for calculating three-phase feedforward respectively and carry it out in abc-domain. The feasibility of duty-ratio feedforward in other coordinate systems has also been demonstrated. Finally, the simulation and experimental results are given, which are compared with the commonly used grid-voltage feedforward in the dq-domain, to verify that the proposed feedforward strategy can not only operate stably at partial load, but also optimize THD to a certain extent.
Chuanhao Ji, Li Zhang 0038, Baolin Chen, Yan Ming, Yan Xing 0001, Xudong Ma
IECON7
2019 Srax: A Low Crosstalk and Insertion Loss 5×5 Optical Router for Optical Network-on-Chip
abstract
Optical network-on-chip is a promising on-chip communication paradigm owing to higher bandwidth and lower latency. However, in the optical network, we have to consider the intrinsic characteristic of crosstalk noise, which has a significant impact on the overall performance of ONoC. In this paper, we proposed a low crosstalk and insertion loss 5×5 optical router called Srax with 15 microring resonators and 12 waveguide crossings, which has less microring resonators and waveguide crossings in comparison to other five-port router architectures. Experiments show that the Srax has lower insertion loss than Cygnus, Crossbar and Optical Router, which are 5.5%, 27.4% and 6.9% lower respectively. Similarly, Srax also has lower crosstalk compared to Cygnus and Optical Router, which are 1.0% and 4.3 % lower respectively.
Xinhao Shi, Fen Ge, Gaizhen Yan, Xudong Ma
IECON6
2018 Industrial Maintenance and Assembly Guidance Using a Markerless AR System with Monocular Camera
abstract
Wearable Assistive System (WAS) has shown promising prospect in industrial inspection and maintenance applications, as these tasks are increasingly complex and flexible and especially when they are infrequent and unfamiliar to workers. In this paper, the development of a WAS system with maintenance and assembly guidance function is presented. In particular, the issue of fast and markerless Augmented Reality (AR) instruction generation using monocular camera is addressed. A learning based module of fast equipment panel detection is implemented to allow easy deployment of the system. Based on the recognition results, different features are extracted according the type of equipment, as s set of co-planar reference points, based on which an computationally efficient Perspective-n-Point camera pose estimator is employed. Consequently, a sequence of enriched egocentric-perspective guidance instructions are generated and delivered to users according to an assembly tree based description of the assembly and maintenance sequence. The system is evaluated by experiments conducted for power equipment inspection guidance. The results validate the practicability and effectiveness of the system.
Kun Qian 0005, Hai Yu 0009, Xudong Ma
ICARCV4
2018 Learning Under-Specified Object Manipulations from Human Demonstrations
abstract
Learning by Demonstration (LbD) allows robots to acquire manipulation skills through human demonstration. In this regard, it is a challenging task to perceive spatial-temporal relations between sub-activities and object affordance in human demonstrations, especially when they are under-specified. This work extends the Probability Graph Model based methods to incorporate high-level demonstration classification. We propose an approach to model the semantics of human demonstration using Programming Domain Description Language (PDDL). Therefore, hidden motion primitives that are impossible to be learned directly from observing human demonstration in noisy video data can be inferred and the robot's plans are refined. Experimental results validate the effectiveness of the proposed method, in which more refined scripts can be generated for robot's execution.
Kun Qian 0005, Fang Fang 0006, Xudong Ma
ICARCV5
2017 Learning Spatial Constraints using Gaussian Process for Shared Control of Semi-autonomous Mobile Robots
Kun Qian 0005, Dan Niu, Fang Fang 0006, Xudong Ma
ICINCO (2)4
2016 Dynamic Android Malware Classification Using Graph-Based Representations
abstract
Malware classification for the Android ecosystem can be performed using a range of techniques. One major technique that has been gaining ground recently is dynamic analysis based on system call invocations recorded during the executions of Android applications. Dynamic analysis has traditionally been based on converting system calls into flat feature vectors and feeding the vectors into machine learning algorithms for classification. In this paper, we implement three traditional feature-vector-based representations for Android system calls. For each feature vector representation, we also propose a novel graph-based representation. We then use graph kernels to compute pair-wise similarities and feed these similarity measures into a Support Vector Machine (SVM) for classification. To speed up the graph kernel computation, we compress the graphs using the Compressed Row Storage format, and then we apply OpenMP to parallelize the computation. Experiments show that the graph-based representations are able to improve the classification accuracy over the corresponding feature-vector-based representations from the same input. Finally we show that different representations can be combined together to further improve classification accuracy.
Lifan Xu, Dong Ping Zhang, Marco A. Alvarez, Jose Andre Morales, Xudong Ma, John Cavazos
CSCloud5
2016 Task planning of mobile robots in distributed service framework
abstract
Complexity and uncertainty of service environment where mobile robots work greatly challenge the performance of robot task planning system. In this paper, an effective and flexible task planning system of mobile robots is designed in distributed service framework with Robot Operating System (ROS). The first part of this paper is the design of system architecture. Then, the logical description tool Planning Domain Definition Language (PDDL) is used to describe the task dynamically based on the designed knowledge base and distributed sensing. In addition, a flexible and stable planning module is designed which combines Forward-Chaining Partial-Order Planning (POPF) planner, post-process, action dispatcher and exceptions handling programs. Lastly, the experiment results in distributed service framework of the task planning system are presented, proving the feasibility of the design.
Chenqiang Ma, Fang Fang 0006, Xudong Ma
IECON3
2016 Embedded control system of dual-module automatic platelet function analyzer
abstract
Platelets play a major role in hemostasis and blood clotting. The measure for coagulation and aggregation characteristics of platelet in human body is widely used in clinical practice. Therefore, quicker and more accurate platelet function analyzer is of urgent requirement recently, especially for guidance of accurate treatment and prevention of thrombus disease. In this paper, design and implementation of a new embedded dual-module automatic platelet analyzer based on sequential platelet counting method (SPCM) is presented. The whole blood sample is assayed several times during the testing process to obtain platelet function parameters. The analyzer is redesigned and enhanced with double modules and four assay modes for different requirements based on conventional single channel product, in which the core of the device is ARM-DSP multi processors to control the system and manage the operations. SPI communication protocol is used among subsystems. Software hierarchy and structure of testing process is introduced and a signal modulation for measuring the weak pulse signal of blood cells is proposed. Mechanisms for dual-module coordination and multi-mode control for fast and accurate assay are given in detail respectively. Experimental results validate the efficiency and feasibility of the system.
Xudong Ma, Fang Fang 0006, Ziquan Dong
IECON2
2016 Development of management software for new Automatic Hematology analyzer based on PC/Windows
abstract
Blood cell analysis is significant for both medical treatment and prevention. For this purpose many vital helps are provided by Automatic Hematology analyzers in layered or distributed framework which is generally composed of management systems and control systems, leading to simplified design process and more convenient maintenance and operation procedures. Users interact with the analyzer through a user friendly human-computer interaction software based on PC/Windows or Embedded systems. In this paper, the development of management software on the PC/Windows system for new automatic hematology analyzer is presented, including functional requirements, system and operation management, communication and workflow. Different from the one under the embedded operating system (Embedded OS) which provides simple interface with better real-time and stability in instrumentation systems, the software under PC/Windows system can operate with a plenty of graphics processing tasks, excellent visual effects and better expandability. In order to improve the performance of real-time, it's indeed to ensure a well-designed reliable workflow and dataflow management. In addition, the human-computer interaction (HCI) interface is designed from the user's point of view.
Xudong Ma, Fang Fang 0006, Ziquan Dong
IECON2
2015 Mobile robot navigation in unknown corridors using line and dense features of point clouds
abstract
This paper addresses the problem of mobile robot navigation in unknown corridors using RGB-Depth cameras. Instead of building a full and global 3D map of the environment, the approach exploits line and dense features extracted from RGB-D sensors. Wall-floor boundary lines are extracted from pre-processed point clouds, which ensure reliable line segmentation results compared with monocular based methods. A strategy is then proposed to compute the reference tracking points along the corridor for a wall-following behaviour. Meanwhile, dense 3D point clouds with ground-plane removed are projected which provide occupancy information, so that existing obstacle avoidance algorithms can be reused. A goal-directed navigation function is also developed by constructing a Nearness Diagram based obstacle avoidance behaviour guided by a wall-following behaviour. Experiment results validate the practicability and effectiveness of the approach.
Kun Qian 0005, Xudong Ma, Bo Zhou 0017
IECON3
2014 Combining supervised and unsupervised models via unconstrained probabilistic embedding
Xiang Ao 0001, Ping Luo 0001, Xudong Ma, Fuzhen Zhuang, Qing He 0003, Zhongzhi Shi, Zhiyong Shen
Inf. Sci.3
2012 A non-isolated three-port converter for stand-alone renewable power system
abstract
A novel non-isolated three-port converter (NI-TPC) is proposed interfacing one PV port, one bidirectional battery port and one load port. Single stage power conversion between any two of the three ports is achieved. The topology is derived by decoupling the bidirectional power flow path of the conventional structure into two unidirectional ones. Two of the three ports can be tightly regulated to achieve maximum power harvesting for PV or charge control for battery, and maintain the load voltage constant at the same time, while the third port is left flexible to compensate the power imbalance of the converter. Operation states are analyzed. The multi-regulator competition control strategy is presented to achieve autonomous and smooth state switching when the PV input power fluctuates. The analysis is verified by the experimental results.
Zihu Zhou, Hongfei Wu, Xudong Ma, Yan Xing 0001
IECON3
2011 Combining Supervised and Unsupervised Models via Unconstrained Probabilistic Embedding
Xudong Ma, Ping Luo 0001, Fuzhen Zhuang, Qing He 0003, Zhongzhi Shi, Zhiyong Shen
IJCAI1
2008 Sensor integration for person tracking and following with mobile robot
abstract
The ability to detect and follow a specific person is a key prerequisite of service mobile robots in indoor environments. A novel method for tracking and following a specific person with a mobile robot by combining vision and laser range information is proposed. The vision module which integrates the color matching method using the Horizontal-Projecting Probability Histogram (HPPH) of upper body color clothes region with the Unscented Particle Filter (UPF) tracking method is developed and, at the same time, the laser module detects the pairs of legs to measure the distance from mobile robot. The information is integrated in a detection procedure that returns the direction and the distance of the target person. The proportional compensation method is utilized to enable the robot to follow the target person steadily. Experimental results validate the robust performance of the approach.
Xudong Ma, Xianzhong Dai, Kun Qian 0005
IROS1
2008 Simultaneous robot Localization and Person Tracking using Rao-Blackwellised Particle Filters with multi-modal sensors
abstract
A probabilistic approach is proposed for Simultaneous robot Localization and Person-Tracking using Rao-Blackwellised Particle Filters (RBPF). Such filters represent posteriors over the person location by a mixture of Kalman Filters, where each is conditioned on a sample of robot pose. Furthermore, information collected via multi-modal sensors is utilized in the RBPFs framework to improve the performance of both localization and tracking. This method is capable of tracking human in situations with sensor noise and global uncertainties over the observer’s pose, whilst outperforms the Conditional Particle Filters (CPF) in computational efficiency. Implementation with collaboration of multi-modal sensors is described, and the experimental results illustrate the accuracy in tracking, as well as the performance of sensor collaboration in accelerating global localization and providing more robustness against occlusions.
Kun Qian 0005, Xudong Ma, Xianzhong Dai
IROS2
2007 Constructing LDPC Codes by 2-Lifts
abstract
With the motivation of constructing Low-Density Parity-Check (LDPC) codes with low error floors, we propose a new code construction scheme based on random 2-lifts. An analysis on stopping set distributions of the proposed code ensembles is presented. The analysis shows that low-weight stopping sets in the resulting codes are typically the results of stopping sets with weak graph expansion properties in the base graphs. Based on this analysis, we propose a set of design criteria for constructing codes with low error floors. According to the design criteria, stopping sets with weak graph expansion properties should be avoided in the base graphs. We present numerical results on constructing capacity-approaching codes over Binary Erasure Channels (BEC). The numerical results show that the codes by the proposed scheme have significantly lower error floors compared with the codes by the standard construction.
Xudong Ma, En-Hui Yang
ISIT1
2006 Mobile Robot Localization Based on Improved Model Matching in Hough Space
abstract
Perceiving the position and orientation of the mobile robot in environment is an important element for an autonomous robot. This paper presents a novel method in which the classical Hough transform is introduced into localization of the mobile robot. To reduce ambiguity significantly, an improved more detailed sonar model is utilized. Firstly a local geometric map in the Hough space is built via the sonar system. Then the matching between a known map of the environment and a local map is performed in the Hough space. Finally this matching result is fused with odometry information by means of the extended Kalman filtering. The technique is especially adapted to indoor polygonal environments. Experimental results validate the favorable performance of this approach
Fang Fang 0006, Xudong Ma, Xianzhong Dai
IROS2
2006 A Visual Based Extended Monte Carlo Localization for Autonomous Mobile Robots
abstract
As a probabilistic localization algorithm, Monte Carlo localization (MCL) method has been widely used for mobile robot localization over the past decade. In this paper, an extended MCL method (EMCL) is developed by incorporating two different resampling processes, namely importance resampling and sensor-based resampling, to conventional MCL for improvement of localization performance. Different resampling processes are utilized based on a matching of sample distribution and observations. Two additional processes for validating over-convergence and uniformity are introduced for examination of such matching. A visual based EMCL is further implemented using a triangulation-based resampling from visual features recognized by Bayesian networks. Experiments are conducted to demonstrate the validity of the proposed approach
Wen Shang, Dong Sun 0001, Xudong Ma, Xianzhong Dai
IROS3
2005 Utilizing external perception and intelligence in a networked robotic system
abstract
Contrary to limited local functions and machine intelligence of a single mobile robot, a networked mobile robot can utilize abundant resources over networks, especially the Internet that extend the network and robot applications to a new field. Potential online assistance from network might be computers, digital sensors, database, even operators. It is currently an important issue of the Internet-based robotic system and can often lead to significant improvement in machine intelligence and system performance. In this paper a new approach utilizing network resources is proposed for indoor applications of our Internet-based robotic system. Developed with this strategy, the robot, instead of only being considered as a passive remote tool, can actively seek for help from assistant network resources for perception or intelligence enhancement. The layered framework of the Internet-based robotic system is outlined as application background, and the definition, implementation and utilizing strategy of network resources based on OAA (open agent architecture) discussed, with two practical examples as speech recognition and global vision perception given to demonstrate the potential applications. Finally, the integration to the framework of the Intranet/Internet-based robotic system with the example utilizations is presented.
Xudong Ma, Xianzhong Dai, Dongyao Wang
IROS1
2004 3D objects detection with Bayesian networks for vision-guided mobile robot navigation
abstract
Recognition of environmental features, which is now the central research topic both in computer vision and mobile robot fields, is prerequisite for vision-guided mobile robot navigation. Perceptual organization is a powerful tool for object recognition through grouping low-level features into objects. In this paper, a perceptual organization algorithm based-on Bayesian networks is proposed to recognize 3D polyhedrons, e.g. compartments and doors, from 2D image in office environment. The algorithm makes full use of knowledge representation and probability inference characteristics of Bayesian networks, thus generating robust recognition results. Moreover, mobile robot active ability is developed to enhance recognition effects. Experimental results demonstrate the validity of the algorithm.
Wen Shang, Xudong Ma, Xianzhong Dai
ICARCV2
2004 Web-based Robotic Control System with Flexible Framework
abstract
Successful applications of Web-based robotic systems must be tolerant of restricted and variable bandwidth, arbitrary delay of the Internet and limited robotic autonomy. As a first step to building such a system, this paper presents a framework in which the technologies of distributed object, multicast communication, proxy and leasing patterns are introduced to make it flexible and extendable. Implementation of a primary robotic system based on the proposed framework is given in detail. Using a Web browser, remote operators can control the Pioneer-2DX mobile robot to navigate and operate in the indoor environment with visual feedback and a virtual environment map through the Internet.
Dongyao Wang, Xudong Ma, Xianzhong Dai
ICRA2
2004 Multimedia transmission strategy in Web-based robotic system
abstract
Multimedia transmission is an important component of the Web-based robotic applications which often serve multiple users concurrently. However, the multimedia transmission in large networks such as Internet is suffered from the hugeness and heterogeneity of the users. In this paper, an adaptive transmission strategy based on MLDA is designed to accommodate with the multicast environment. Improvement to MLDA is proposed for applications of the Web-based robotic system, as in the behaviors of the multimedia sender and receiver, synchronization mechanism and estimation of TCP-friendly transmission rate. Finally, a preliminary simulation for the estimation of TCP-friendly transmission rate is demonstrated.
Dongyao Wang, Xudong Ma, Xianzhong Dai
IROS2
2004 Low-density parity-check codes with fast decoding convergence speed
abstract
In this paper, the design of low-density parity-check (LDPC) codes for binary erasure channels (BEC) with less number of decoding iterations in the parallel message-passing decoding is presented. We consider only the cases where the capacity is approached with sufficiently low parity-check matrix density.
Xudong Ma, En-Hui Yang
ISIT1