Hongwei Xie

dblp:37/1678 · DBLP profile ↗
← Back
26ranked-venue papers
7as first author
11since 2021 · last 2026
0009-0006-8905-6909ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 9 · 2 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 4 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 1 since 2021Human-computer interaction and ubiquitous computing · 3 · 1 first-authorComputer networks · 2 · 1 first-authorSystems, architecture and hardware · 1 · 1 since 2021Software engineering, systems software and programming languages · 1Databases, data management, data science and information retrieval · 1
YearPublicationVenuePosition
2026 Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving
abstract
Large Vision-Language Models (LVLMs) have significantly advanced image understanding. Their comprehension and reasoning capabilities enable promising applications in autonomous driving scenarios. However, existing research typically focuses on partial objects within scenes and simple question-answer pair annotations, struggling to achieve comprehensive scene understanding. Meanwhile, existing LVLMs suffer from the lack of mapping relationship between 2D and 3D and insufficient integration of 3D spatial understanding and instruction following. To tackle these limitations, we first introduce NuInteract, a large-scale dataset with over 1.5M multi-view image-language pairs spanning dense scene captions and diverse interactive tasks. Furthermore, we propose DriveMonkey, a simple yet effective framework that seamlessly integrates LVLMs with a spatial processor using a series of learnable queries. The spatial processor, designed as a plug-and-play component, can be initialized with pre-trained 3D detectors to provide structured geometric priors for language-conditioned 3D grounding. Our experiments show that DriveMonkey outperforms general LVLMs, especially achieving a notable 9.86% improvement on the 3D visual grounding task. The dataset and code are released at https://github.com/zc-zhao/DriveMonkey.
Zongchuang Zhao, Haoyu Fu, Dingkang Liang, Xin Zhou 0013, Dingyuan Zhang, Hongwei Xie, Xiang Bai
IEEE Trans. Image Process.6
2025 Orion: A Holistic End-To-End Autonomous Driving Framework by Vision-Language Instructed Action Generation
Haoyu Fu, Diankun Zhang, Zongchuang Zhao, Jianfeng Cui, Dingkang Liang, Dingyuan Zhang, Hongwei Xie, Xiang Bai
ICCV8
2025 A secure and privacy-preserving technique based on coupled chaotic system and plaintext encryption for multimodal medical images
Hongwei Xie, Jing Bian, Hao Zhang 0061
Multim. Tools Appl.1
2024 SurroundSDF: Implicit 3D Scene Understanding Based on Signed Distance Field
abstract
Vision-centric 3D environment understanding is both vi-tal and challenging for autonomous driving systems. Re-cently, object-free methods have attracted considerable at-tention. Such methods perceive the world by predicting the semantics of discrete voxel grids but fail to construct continuous and accurate obstacle surfaces. To this end, in this paper, we propose SurroundSDF to implicitly predict the signed distance field (SDF) and semantic field for the continuous perception from surround images. Specifically, we introduce a query-based approach and utilize SDF con-strained by the Eikonal formulation to accurately describe the surfaces of obstacles. Furthermore, considering the absence of precise SDF ground truth, we propose a novel weakly supervised paradigm for SDF, referred to as the Sandwich Eikonal formulation, which emphasizes applying correct and dense constraints on both sides of the surface, thereby enhancing the perceptual accuracy of the surface. Experiments suggest that our method achieves SOTA for both occupancy prediction and 3D scene reconstruction tasks on the nuScenes dataset.
Lizhe Liu, Bohua Wang, Hongwei Xie, Daqi Liu, Kuiyuan Yang
CVPR3
2024 RenderOcc: Vision-Centric 3D Occupancy Prediction with 2D Rendering Supervision
abstract
3D occupancy prediction holds significant promise in the fields of robot perception and autonomous driving, which quantifies 3D scenes into grid cells with semantic labels. Recent works mainly utilize complete occupancy labels in 3D voxel space for supervision. However, the expensive annotation process and sometimes ambiguous labels have severely constrained the usability and scalability of 3D occupancy models. To address this, we present RenderOcc, a novel paradigm for training 3D occupancy models only using 2D labels. Specifically, we extract a NeRF-style 3D volume representation from multi-view images, and employ volume rendering techniques to establish 2D renderings, thus enabling direct 3D supervision from 2D semantics and depth labels. Additionally, we introduce an Auxiliary Ray method to tackle the issue of sparse viewpoints in autonomous driving scenarios, which leverages sequential frames to construct comprehensive 2D rendering for each object. To our best knowledge, RenderOcc is the first attempt to train multi-view 3D occupancy models only using 2D labels, reducing the dependence on costly 3D occupancy annotations. Extensive experiments demonstrate that RenderOcc achieves comparable performance to models fully supervised with 3D labels, underscoring the significance of this approach in real-world applications. Our code is available at https://github.com/pmj110119/RenderOcc.
Mingjie Pan, Jiaming Liu 0003, Renrui Zhang, Peixiang Huang, Xiaoqi Li 0020, Hongwei Xie, Bing Wang 0013, Li Liu 0069, Shanghang Zhang
ICRA6
2023 An image encryption algorithm based on novel block scrambling scheme and Josephus sequence generator
Hongwei Xie, Ya-jun Gao, Hao Zhang 0061
Multim. Tools Appl.1
2023 Novel medical image cryptogram technology based on segmentation and DNA encoding
Hongwei Xie
Multim. Tools Appl.1
2023 A Fixed-Wing UAV Formation Algorithm Based on Vector Field Guidance
abstract
The vector field method was originally proposed to guide a single fixed-wing Unmanned Aerial Vehicle (UAV) towards a desired path. In this work, a non-uniform vector field method is proposed that changes in both magnitude and direction, for the purpose of achieving formations of UAVs. As compared to related work in the literature, the proposed formation control law does not need to assume absence of wind. That is, due to the effect of the wind on the UAV, one can handle the UAV air speed being different from its ground speed, and the UAV heading angle being different from its course angle. Stability of the proposed formation method is analyzed via Lyapunov stability theory, and validations are carried out in software-in-the-loop and hardware-in-the-loop comparative experiments. Note to Practitioners—The software-in-the-loop and hardware-in-the-loop experiments, which are done with PX4 autopilot software and hardware, show that the proposed method can be implemented on board of UAVs and integrated with the control architecture of existing autopilot suites. Comparisons with standard formation algorithms show that the proposed method is effective in achieving formation in different path scenarios.
Ximan Wang, Simone Baldi, Xuewei Feng, Changwei Wu, Hongwei Xie, Bart De Schutter
IEEE Trans Autom. Sci. Eng.5
2022 Knowledge Distillation via the Target-aware Transformer
abstract
Knowledge distillation becomes a de facto standard to improve the performance of small neural networks. Most of the previous works propose to regress the representational features from the teacher to the student in a one-to-one spatial matching fashion. However, people tend to overlook the fact that, due to the architecture differences, the semantic information on the same spatial location usually vary. This greatly undermines the underlying assumption of the one-to-one distillation approach. To this end, we propose a novel one-to-all spatial matching knowledge distillation approach. Specifically, we allow each pixel of the teacher feature to be distilled to all spatial locations of the student features given its similarity, which is generated from a target-aware transformer. Our approach surpasses the state-of-the-art methods by a significant margin on various computer vision benchmarks, such as ImageNet, Pascal VOC and COCOStuff10k. Code is available at https://github.com/sihaoevery/TaT.
Sihao Lin, Hongwei Xie, Bing Wang 0013, Kaicheng Yu, Xiaojun Chang, Xiaodan Liang
CVPR2
2022 BEVFusion: A Simple and Robust LiDAR-Camera Fusion Framework
abstract
Fusing the camera and LiDAR information has become a de-facto standard for 3D object detection tasks. Current methods rely on point clouds from the LiDAR sensor as queries to leverage the feature from the image space. However, people discovered that this underlying assumption makes the current fusion framework infeasible to produce any prediction when there is a LiDAR malfunction, regardless of minor or major. This fundamentally limits the deployment capability to realistic autonomous driving scenarios. In contrast, we propose a surprisingly simple yet novel fusion framework, dubbed BEVFusion, whose camera stream does not depend on the input of LiDAR data, thus addressing the downside of previous methods. We empirically show that our framework surpasses the state-of-the-art methods under the normal training settings. Under the robustness training settings that simulate various LiDAR malfunctions, our framework significantly surpasses the state-of-the-art methods by 15.7% to 28.9% mAP. To the best of our knowledge, we are the first to handle realistic LiDAR malfunction and can be deployed to realistic scenarios without any post-processing procedure.
Tingting Liang, Hongwei Xie, Kaicheng Yu, Zhongyu Xia, Yongtao Wang, Zhi Tang 0001
NeurIPS2
2021 Exploring Inter-Channel Correlation for Diversity-preserved Knowledge Distillation
abstract
Knowledge Distillation has shown very promising ability in transferring learned representation from the larger model (teacher) to the smaller one (student). Despite many efforts, prior methods ignore the important role of retaining inter-channel correlation of features, leading to the lack of capturing intrinsic distribution of the feature space and sufficient diversity properties of features in the teacher network. To solve the issue, we propose the novel Inter-Channel Correlation for Knowledge Distillation (ICKD), with which the diversity and homology of the feature space of the student network can align with that of the teacher network. The correlation between these two channels is interpreted as diversity if they are irrelevant to each other, otherwise homology. Then the student is required to mimic the correlation within its own embedding space. In addition, we introduce the grid-level inter-channel correlation, making it capable of dense prediction tasks. Extensive experiments on two vision tasks, including ImageNet classification and Pascal VOC segmentation, demonstrate the superiority of our ICKD, which consistently outperforms many existing methods, advancing the state-of-the-art in the fields of Knowledge Distillation. To our knowledge, we are the first method based on knowledge distillation boosts ResNet18 beyond 72% Top-1 accuracy on ImageNet classification. Code is available at: https://github.com/ADLab-AutoDrive/ICKD.
Li Liu 0069, Qingle Huang, Sihao Lin, Hongwei Xie, Bing Wang 0013, Xiaojun Chang, Xiaodan Liang
ICCV4
2020 Overflow Aware Quantization: Accelerating Neural Network Inference by Low-bit Multiply-Accumulate Operations
abstract
The inherent heavy computation of deep neural networks prevents their widespread applications. A widely used method for accelerating model inference is quantization, by replacing the input operands of a network using fixed-point values. Then the majority of computation costs focus on the integer matrix multiplication accumulation. In fact, high-bit accumulator leads to partially wasted computation and low-bit one typically suffers from numerical overflow. To address this problem, we propose an overflow aware quantization method by designing trainable adaptive fixed-point representation, to optimize the number of bits for each input tensor while prohibiting numeric overflow during the computation. With the proposed method, we are able to fully utilize the computing power to minimize the quantization loss and obtain optimized inference performance. To verify the effectiveness of our method, we conduct image classification, object detection, and semantic segmentation tasks on ImageNet, Pascal VOC, and COCO datasets, respectively. Experimental results demonstrate that the proposed method can achieve comparable performance with state-of-the-art quantization methods while accelerating the inference process by about 2 times.
Hongwei Xie, Yafei Song 0002, Mingyang Li 0001
IJCAI1
2018 Double Attention Mechanism for Sentence Embedding
Miguel Kakanakou, Hongwei Xie, Yan Qiang 0001
WISA2
2017 A High Accurate Vision Algorithm on Measuring Arbitrary Contour
Hongwei Xie, Yu Liu 0014, Jiaxiang Luo
ICONIP (6)1
2016 The analyses of human inherited disease and tissue-specific proteins in the interaction network
Hongwei Xie
J. Biomed. Informatics3
2016 A Reliability-Augmented Particle Filter for Magnetic Fingerprinting Based Indoor Localization on Smartphone
abstract
Using magnetic field data as fingerprints for smartphone indoor positioning has become popular in recent years. Particle filter is often used to improve accuracy. However, most of existing particle filter based approaches either are heavily affected by motion estimation errors, which result in unreliable systems, or impose strong restrictions on smartphone such as fixed phone orientation, which are not practical for real-life use. In this paper, we present a novel indoor positioning system for smartphones, which is built on our proposed reliability-augmented particle filter. We create several innovations on the motion model, the measurement model, and the resampling model to enhance the basic particle filter. To minimize errors in motion estimation and improve the robustness of the basic particle filter, we propose a dynamic step length estimation algorithm and a heuristic particle resampling algorithm. We use a hybrid measurement model, combining a new magnetic fingerprinting model and the existing magnitude fingerprinting model, to improve system performance, and importantly avoid calibrating magnetometers for different smartphones. In addition, we propose an adaptive sampling algorithm to reduce computation overhead, which in turn improves overall usability tremendously. Finally, we also analyze the “Kidnapped Robot Problem” and present a practical solution. We conduct comprehensive experimental studies, and the results show that our system achieves an accuracy of 1~2 m on average in a large building.
Hongwei Xie, Tao Gu 0001, XianPing Tao, Haibo Ye, Jian Lu 0001
IEEE Trans. Mob. Comput.1
2014 MaLoc: a practical magnetic fingerprinting approach to indoor localization using smartphones
abstract
Using magnetic field data as fingerprints for localization in indoor environment has become popular in recent years. Particle filter is often used to improve accuracy. However, most of existing particle filter based approaches either are heavily affected by motion estimation errors, which makes the system unreliable, or impose strong restrictions on smartphone such as fixed phone orientation, which is not practical for real-life use. In this paper, we present an indoor localization system named MaLoc, built on our proposed augmented particle filter. We create several innovations on the motion model, the measurement model and the resampling model to enhance the traditional particle filter. To minimize errors in motion estimation and improve the robustness of particle filter, we augment the particle filter with a dynamic step length estimation algorithm and a heuristic particle resampling algorithm. We use a hybrid measurement model which combines a new magnetic fingerprinting model and the existing magnitude fingerprinting model to improve the system performance and avoid calibrating different smartphone magnetometers. In addition, we present a novel localization quality estimation method and a localization failure detection method to address the "Kidnapped Robot Problem" and improve the overall usability. Our experimental studies show that MaLoc achieves a localization accuracy of 1~2.8m on average in a large building.
Hongwei Xie, Tao Gu 0001, XianPing Tao, Haibo Ye, Jian Lu 0001
UbiComp1
2014 FuAET: a tool for developing fuzzy self-adaptive software systems
abstract
Handling uncertainty in software self-adaptation has become an important and challenging issue. In our previous work, we proposed a fuzzy control based approach named Software Fuzzy Self-Adaptation (SFSA) to address fuzziness, a kind of uncertainty in software self-adaptation. However, our SFSA approach still lacks a tool to efficiently support the implementation process of SFSA. Existing tools for realizing self-adaptive applications does not directly deal with fuzziness in self-adaption loops. In this paper, we present the FuAET, a tool designed for building fuzzy self-adaptive software systems. The novelty of the tool is that it can not only provide a friendly GUI for editing and testing fuzzy self-adaptation strategies in intelligible domain specific language (DSL), but also automatically convert DSL-based fuzzy selfadaptation strategies into aspect-based programming code in native-language (NL), e.g., C++. This paper describes the design framework and implementation principles of FuAET, and then proposes a general development process using FuAET for reference to developers. Finally, we conduct an empirical study for evaluation of FuAET using an industrial control application. The results show that FuAET can automate the development of SFSA and ease the burden of the software engineers.
Qiliang Yang, XianPing Tao, Hongwei Xie, Jianchun Xing, Wei Song 0003
Internetware3
2014 Research on the Forecasting Model for Hypertension Drug Efficacy Based on Medical Data
abstract
Hypertension is a common disease endangering human health, and every year the number of cases increases at a speed of ten millions. So far there are about 230 million hypertension patients in China, and the long-term drug treatment for hypertension patients is an effective measure to control blood pressure. In this paper, based on the evaluation period of patients with Medication Possession Ratio (MPR) and blood pressure analysis, the beta distribution model which is used to find the relationship between them is established. Then genetic algorithms (GA) and cross validation are carried out on the model and we make comparative analysis with linear distribution model. The experimental results show that the beta distribution model can well forecast the drug efficacy of patients. Only through long-term drug therapy, can the blood pressure be effectively controlled.
Hongwei Xie, Jiancheng An 0002
MSN2
2014 SILVER: an efficient tool for stable isotope labeling LC-MS data quantitative analysis with quality control methods
abstract
SUMMARY: With the advance of experimental technologies, different stable isotope labeling methods have been widely applied to quantitative proteomics. Here, we present an efficient tool named SILVER for processing the stable isotope labeling mass spectrometry data. SILVER implements novel methods for quality control of quantification at spectrum, peptide and protein levels, respectively. Several new quantification confidence filters and indices are used to improve the accuracy of quantification results. The performance of SILVER was verified and compared with MaxQuant and Proteome Discoverer using a large-scale dataset and two standard datasets. The results suggest that SILVER shows high accuracy and robustness while consuming much less processing time. Additionally, SILVER provides user-friendly interfaces for parameter setting, result visualization, manual validation and some useful statistics analyses. AVAILABILITY AND IMPLEMENTATION: SILVER and its source codes are freely available under the GNU General Public License v3.0 at http://bioinfo.hupo.org.cn/silver.
Jiyang Zhang 0001, Mingfei Han 0001, Songfeng Wu, Kehui Liu, Hongwei Xie, Fuchu He
Bioinform.8
2014 Screening drug target proteins based on sequence information
Tengjiao Wang 0002, Hongwei Xie
J. Biomed. Informatics4
2013 A Wearable RFID System for Real-Time Activity Recognition Using Radio Patterns
Liang Wang 0006, Tao Gu 0001, Hongwei Xie, XianPing Tao, Jian Lu 0001, Yu Huang 0002
MobiQuitous3
2013 Reconstruction of Signaling Network from Protein Interactions Based on Function Annotations
abstract
The directionality of protein interactions is the prerequisite of forming various signaling networks, and the construction of signaling networks is a critical issue in the discovering the mechanism of the life process. In this paper, we proposed a novel method to infer the directionality in protein-protein interaction networks and furthermore construct signaling networks. Based on the functional annotations of proteins, we proposed a novel parameter GODS and established the prediction model. This method shows high sensitivity and specificity to predict the directionality of protein interactions, evaluated by fivefold cross validation. By taking the threshold value of GODS as 2, we achieved accuracy 95.56 percent and coverage 74.69 percent in the human test set. Also, this method was successfully applied to reconstruct the classical signaling pathways in human. This study not only provided an effective method to unravel the unknown signaling pathways, but also the deeper understanding for the signaling networks, from the aspect of protein function.
Dong Li 0018, Hongwei Xie, Fuchu He
IEEE ACM Trans. Comput. Biol. Bioinform.4
2010 Tmod: toolbox of motif discovery
abstract
SUMMARY: Motif discovery is an important topic in computational transcriptional regulation studies. In the past decade, many researchers have contributed to the field and many de novo motif-finding tools have been developed, each may have a different strength. However, most of these tools do not have a user-friendly interface and their results are not easily comparable. We present a software called Toolbox of Motif Discovery (Tmod) for Windows operating systems. The current version of Tmod integrates 12 widely used motif discovery programs: MDscan, BioProspector, AlignACE, Gibbs Motif Sampler, MEME, CONSENSUS, MotifRegressor, GLAM, MotifSampler, SeSiMCMC, Weeder and YMF. Tmod provides a unified interface to ease the use of these programs and help users to understand the tuning parameters. It allows plug-in motif-finding programs to run either separately or in a batch mode with predetermined parameters, and provides a summary comprising of outputs from multiple programs. Tmod is developed in C++ with the support of Microsoft Foundation Classes and Cygwin. Tmod can also be easily expanded to include future algorithms. AVAILABILITY: Tmod is available for download at http://www.fas.harvard.edu/~junliu/Tmod/.
Hanchang Sun, Jun S. Liu, Hongwei Xie
Bioinform.6
2008 Generating GO Slim Using Relational Database Management Systems to Support Proteomics Analysis
abstract
The Gene Ontology Consortium built the Gene Ontology database (GO) to address the need for a common standard in naming genes and gene products. Using different names for the same concepts and different concepts with the same name makes it effectively impossible for humans and computers alike to analyze biological processes across different organisms. The consortium addresses this need by defining terms for categorizing genes and gene products. A convention in GO is that each gene or gene product is annotated to the most specific GO term in the GO database. It is, however, also useful for researchers to be able to group genesor gene products into broad biological categories that give a higher-level view of their function when analyzing results of an experiment. A GO Slim is a subset of the GO ontology that provides such a higher-level view of functions. Existing GO Slim generation tools have two important limitations: programming language dependence, and an inability to dynamically generate a GO Slim while analyzing. We have extended the relational database engine to dynamically generate a GO Slim overcoming this limitations. Using this extension, we have developed a tool (Dynamic GOSlim) that dynamically generates a GO Slim and uses the generated GO Slim to categorize genes or gene products. This tool is being used in an ongoing proteomics project aimed at identifying possible oral cancer biomarkers in saliva.
Getiria Onsongo, Hongwei Xie, Timothy J. Griffin, John V. Carlis
CBMS2
2008 A nonparametric model for quality control of database search results in shotgun proteomics
abstract
BACKGROUND: Analysis of complex samples with tandem mass spectrometry (MS/MS) has become routine in proteomic research. However, validation of database search results creates a bottleneck in MS/MS data processing. Recently, methods based on a randomized database have become popular for quality control of database search results. However, a consequent problem is the ignorance of how to combine different database search scores to improve the sensitivity of randomized database methods. RESULTS: In this paper, a multivariate nonlinear discriminate function (DF) based on the multivariate nonparametric density estimation technique was used to filter out false-positive database search results with a predictable false positive rate (FPR). Application of this method to control datasets of different instruments (LCQ, LTQ, and LTQ/FT) yielded an estimated FPR close to the actual FPR. As expected, the method was more sensitive when more features were used. Furthermore, the new method was shown to be more sensitive than two commonly used methods on 3 complex sample datasets and 3 control datasets. CONCLUSION: Using the nonparametric model, a more flexible DF can be obtained, resulting in improved sensitivity and good FPR estimation. This nonparametric statistical technique is a powerful tool for tackling the complexity and diversity of datasets in shotgun proteomics.
Jiyang Zhang 0001, Hongwei Xie, Fuchu He
BMC Bioinform.4