Hongfei Xu

dblp:26/7840 · DBLP profile ↗
← Back
33ranked-venue papers
11as first author
20since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 22 · 8 first-author · 15 since 2021Computer networks · 4 · 4 since 2021Databases, data management, data science and information retrieval · 4 · 2 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
YearPublicationVenuePosition
2026 Paraphrasing as Zero-shot Translation with Feature-guided Diversity Enhancement
abstract
Paraphrasing uses different words, sentence structures, or expressions to convey similar semantics.It is an effective training data augmentation method to improve low-resource Natural Language Processing (NLP) tasks.Existing studies normally leverage parallel corpora to construct parabanks, regarding the Machine Translation (MT) results of source sentences as the paraphrases of the corresponding target sentences.As MT models are usually trained on the same parallel corpus, translation of the training set may suffer from overfitting, which leads to less diverse paraphrases.Training paraphrasers on the parabank generated via MT may also suffer from the information loss issue, as the parabank is derived from the parallel corpora, and the knowledge inside the parabank is a subset of that inside the parallel corpora.In this paper, we train bidirectional Multilingual Neural Machine Translation (MNMT) on the bi-directional bilingual parallel corpus, and use the MNMT model directly as a paraphrasing model by asking it to generate "translations" of the input language.As some source tokens also appear in the translation in the parallel corpus, we introduce "copy"/"not-copy" tags to indicate the existence/non-existence of source tokens in the target translation during training, and use the "not-copy" tag to encourage paraphrasing during inference.Manual and automatic evaluation results show that our ParaMNMT method can generate paraphrases of higher semantic consistency, literal fluency and sentential diversity compared to existing parabanks and LLMs.Our data augmentation experiments verify the effectiveness of ParaM-NMT on improving low-resource NLP tasks.
Ziyue Yan, Hongying Zan, Xinglin Lyu, Hongfei Xu
ACL (1)4
2026 Semantic-Aware Co-Design of Spatio-Temporal Focusing and Adaptive Modulation for Wearable Brain-Body Neuro-Interfaces
Tianyi Yan, Chenxuan Wu, Hongzhen Pan, Zhaolan He, Hongfei Xu, Yongyi Dou
ICC6
2026 Harnessing Poiseuille Flow for ISI-Resilient Molecular Communication: A Spatio-Temporal Focusing Approach
Chenxuan Wu, Hongzhen Pan, Zhaolan He, Hongfei Xu, Yongyi Dou
WCNC6
2025 COLD-QA: A Complex Long-Distance Numerical Reasoning Dataset for Hybrid Tabular-Textual Question Answering
abstract
Hybrid tabular-textual question answering (QA) typically requires numerical reasoning over heterogeneous data, with the reasoning program first generated and then executed to obtain the final answer. However, in most existing hybrid QA benchmarks, each step only relies on the numbers in the input or the calculation result of the preceding step. Questions that require long-distance numerical reasoning (where the calculation of a step relies on the results calculated a number of steps previously) are rare. However, they may be required to solve some complicated financial problems; these questions require a more complex model capable of capturing long-distance dependencies. To study more challenging hybrid tabular-textual QA, we construct a new large-scale hybrid tabular-textual dataset, COLD-QA, COmplex Long-Distance numerical reasoning Question Answering dataset. We also conduct extensive experiments with multiple baselines. The COLD-QA dataset is significantly more difficult than previous work, according to experiment results.
Xiaoqing Cheng, Hongying Zan, Tengxun Zhang, Hongfei Xu
IJCNN4
2025 Overview of NLPCC 2025 Shared Task 5: Chinese Government Text Correction with Knowledge Bases
Yuxiang Jia, Jiajia Cui, Lingling Mu, Hongfei Xu
NLPCC (4)6
2025 Very-Long-Distance Dependency Capturing Evaluation via Language Modeling Based on Gender Consistency
Hongfei Xu, Zhuofei Liang, Josef van Genabith, Deyi Xiong, Hongying Zan, Qiuhui Liu, Tengxun Zhang
NLPCC (4)1
2025 Mitigating Scalability Walls of RDMA-based Container Networks
Wei Liu 0148, Kun Qian 0021, Zhenhua Li 0001, Feng Qian 0001, Tianyin Xu, Yunhao Liu 0001, Yu Guan 0005, Shuhong Zhu, Hongfei Xu, Lanlan Xi, Ennan Zhai
NSDI9
2025 SkeletonHunter: Diagnosing and Localizing Network Failures in Containerized Large Model Training
abstract
The flexibility and portability characteristics have made containers a popular serverless environment for large model training in recent years. Unfortunately, these advantages render the network support for containerized large model training extremely challenging, due to the high dynamics of containers, the complex interplay between underlay and overlay networks, and the stringent requirements on failure detection and localization. Existing data center network debugging tools, which rely on comprehensive or opportunistic monitoring, are either inefficient or inaccurate in this setting.
Wei Liu 0148, Kun Qian 0021, Zhenhua Li 0001, Tianyin Xu, Yunhao Liu 0001, Jiakang Li, Shuhong Zhu, Xue Li 0024, Hongfei Xu, Ennan Zhai
SIGCOMM11
2024 Zhongjing: Enhancing the Chinese Medical Capabilities of Large Language Model through Expert Feedback and Real-World Multi-Turn Dialogue
abstract
Recent advances in Large Language Models (LLMs) have achieved remarkable breakthroughs in understanding and responding to user intents. However, their performance lag behind general use cases in some expertise domains, such as Chinese medicine. Existing efforts to incorporate Chinese medicine into LLMs rely on Supervised Fine-Tuning (SFT) with single-turn and distilled dialogue data. These models lack the ability for doctor-like proactive inquiry and multi-turn comprehension and cannot align responses with experts' intentions. In this work, we introduce Zhongjing, the first Chinese medical LLaMA-based LLM that implements an entire training pipeline from continuous pre-training, SFT, to Reinforcement Learning from Human Feedback (RLHF). Additionally, we construct a Chinese multi-turn medical dialogue dataset of 70,000 authentic doctor-patient dialogues, CMtMedQA, which significantly enhances the model's capability for complex dialogue and proactive inquiry initiation. We also define a refined annotation rule and evaluation criteria given the unique characteristics of the biomedical domain. Extensive experimental results show that Zhongjing outperforms baselines in various capacities and matches the performance of ChatGPT in some abilities, despite the 100x parameters. Ablation studies also demonstrate the contributions of each component: pre-training enhances medical knowledge, and RLHF further improves instruction-following ability and safety. Our code, datasets, and models are available at https://github.com/SupritYoung/Zhongjing.
Songhua Yang, Hanjie Zhao, Senbin Zhu, Hongfei Xu, Yuxiang Jia, Hongying Zan
AAAI5
2024 Rewiring the Transformer with Depth-Wise LSTMs
abstract
Stacking non-linear layers allows deep neural networks to model complicated functions, and including residual connections in Transformer layers is beneficial for convergence and performance. However, residual connections may make the model “forget” distant layers and fail to fuse information from previous layers effectively. Selectively managing the representation aggregation of Transformer layers may lead to better performance. In this paper, we present a Transformer with depth-wise LSTMs connecting cascading Transformer layers and sub-layers. We show that layer normalization and feed-forward computation within a Transformer layer can be absorbed into depth-wise LSTMs connecting pure Transformer attention layers. Our experiments with the 6-layer Transformer show significant BLEU improvements in both WMT 14 English-German / French tasks and the OPUS-100 many-to-many multilingual NMT task, and our deep Transformer experiments demonstrate the effectiveness of depth-wise LSTM on the convergence and performance of deep Transformers.
Hongfei Xu, Qiuhui Liu, Josef van Genabith, Deyi Xiong
LREC/COLING1
2024 Multi-pass Decoding for Grammatical Error Correction
abstract
Sequence-to-sequence (seq2seq) models achieve comparable or better grammatical error correction performance compared to sequenceto-edit (seq2edit) models.Seq2edit models normally iteratively refine the correction result, while seq2seq models decode only once without aware of subsequent tokens.Iteratively refining the correction results of seq2seq models via Multi-Pass Decoding (MPD) may lead to better performance.However, MPD increases the inference costs.Deleting or replacing corrections in previous rounds may lose useful information in the source input.We present an early-stop mechanism to alleviate the efficiency issue.To address the source information loss issue, we propose to merge the source input with the previous round correction result into one sequence.Experiments on the CoNLL-14 test set and BEA-19 test set show that our approach can lead to consistent and significant improvements over strong BART and T5 baselines (+1.80, +1.35, and +2.02 F0.5 for BART 12-2, large and T5 large respectively on CoNLL-14 and +2.99, +1.82, and +2.79 correspondingly on BEA-19), obtaining F0.5 scores of 68.41 and 75.36 on CoNLL-14 and BEA-19 respectively.We brought go to the orchard and apples but , forget pears .We bring go to the orchard and apples but , forget pears We bring go to the orchard and apples but , forget pears . .
Lingling Mu, Jingyi Zhang 0002, Hongfei Xu
EMNLP4
2024 Knowledge-injected Prompt Learning for Chinese Biomedical Entity Normalization
abstract
The Biomedical Entity Normalization (BEN) task aims to align raw, unstructured medical entities to standard entities, thus promoting data coherence and facilitating better downstream medical applications. Recently, prompt learning methods have shown promising results in the natural language processing field. However, existing research falls short in tackling the more complex Chinese BEN task, especially in the few-shot scenario with limited medical data, and the vast potential of the external medical knowledge base has not yet been fully exploited. To address these challenges, this article proposes a novel Knowledge-injected Prompt Learning (PL-Knowledge) method. Specifically, the approach consists of five stages: candidate entity matching, knowledge extraction, knowledge encoding, knowledge injection, and prediction output. By effectively encoding the knowledge items contained in medical entities and incorporating them into tailor-made knowledge-injected templates, the additional knowledge enhances the model’s ability to capture latent relationships between medical entities, thus achieving a better match with the standard entities. Comprehensive experiments are conducted on a benchmark dataset in both few-shot and full-scale settings. This method outperforms existing baselines, with an average accuracy improvement of 12.96 percentage points in few-shot and 0.94 percentage points in full-data cases, showcasing its excellence in the BEN task.
Songhua Yang, Chenghao Zhang 0005, Chenyuan He, Hongfei Xu, Hongying Zan, Yuxiang Jia
ACM Trans. Asian Low Resour. Lang. Inf. Process.4
2023 A Novel Method for Maneuvering Extended Vehicle Tracking with Automotive Radar
abstract
In high-resolution automotive radar tracking systems, vehicle targets are often regarded as extended targets, which means multiple measurements originated from scattering centers of vehicle targets can be detected at each scan and thus the traditional point target tracking schemes are unsuitable. Meanwhile, vehicle maneuvers, e.g., braking and swerving, cause serious degradation of the classical extended target tracking methods. In this paper, a novel method is proposed for maneuvering extended vehicle tracking with automotive radar. The data-region association (DRA) strategy is adopted to handle the vehicle extension effect, which is superior in describing the complex spatial distribution of vehicle target measurements. The interacting multiple model (IMM) method is combined with this DRA strategy to describe the evolution of target motion models. Accordingly, the proposed DRA-IMM method achieves satisfying tracking performance of extended vehicles and also guarantees the robustness in case of maneuvers. Furthermore, in view of the correlation between vehicle extension and its kinematic state, a ray-based strategy is devised to improve the prior distribution of the data-region association of the basic DRA-IMM, and accordingly an enhanced DRA-IMM (EDRA-IMM) method is proposed. Simulation result validates the effectiveness of the proposed DRA-IMM method for maneuvering extended vehicle tracking and the further improvement of the proposed EDRA-IMM method.
Hongfei Xu, Yaowen Li, Yuxin Ke, Zhizhuo Jiang, Yu Liu 0005
FUSION1
2023 Improving Aspect Sentiment Triplet Extraction with Perturbed Masking and Edge-Enhanced Sentiment Graph Attention Network
abstract
Aspect Sentiment Triplet Extraction (ASTE) aims to extract sentiment triplets from texts. Previous studies attempted to use an unsupervised perturbed masking technique to derive syntax trees from pre-trained language models (PLMs) to extract aspect sentiment simply. However, existing methods neglect further exploration of the technique for more complex tasks like ASTE. In this paper, we propose a novel Edge-enhanced Sentiment Graph Attention Network model (ES-GAT), which transforms the impact matrix generated by the technique into a new linguistic feature, and uses a novel effective fusion strategy to incorporate multiple features, which can better capture the implicit connections among sentiment elements. In the graph attention module, we simultaneously consider edge and node attention weight calculations and updates to solve the edge-sensitive end to end ASTE task. Experiments show that our model outperforms the strong baselines. The F1 scores are improved by 2.52% and 1.56% on average on the two versions of the benchmark datasets.11Code and datasets are available at https://github.com/SupritYoung/ESGAT
Songhua Yang, Tengxun Zhang, Hongfei Xu, Yuxiang Jia
IJCNN3
2023 NAPG: Non-Autoregressive Program Generation for Hybrid Tabular-Textual Question Answering
Tengxun Zhang, Hongfei Xu, Josef van Genabith, Deyi Xiong, Hongying Zan
NLPCC (1)2
2023 JoinER-BART: Joint Entity and Relation Extraction With Constrained Decoding, Representation Reuse and Fusion
abstract
Joint Entity and Relation Extraction (JERE) is an important research direction in Information Extraction (IE). Given the surprising performance with fine-tuning of pre-trained BERT in a wide range of NLP tasks, nowadays most studies for JERE are based on the BERT model. Rather than predicting a simple tag for each word, these approaches are usually forced to design complex tagging schemes, as they may have to extract entity-relation pairs which may overlap with others from the same sequence of word representations in a sentence. Recently, sequence-to-sequence (seq2seq) pre-trained BART models show better performance than BERT models in many NLP tasks. Importantly, a seq2seq BART model can simply generate sequences of (many) entity-relation triplets with its decoder, rather than just tag input words. In this paper, we present a new generative JERE framework based on pre-trained BART. Different from the basic seq2seq BART architecture: 1) our framework employs a constrained classifier which only predicts either a token of the input sentence or a relation in each decoding step, and 2) we reuse representations from the pre-trained BART encoder in the classifier instead of a newly trained weight matrix, as this better utilizes the knowledge of the pre-trained model and context-aware representations for classification, and empirically leads to better performance. In our experiments on the widely studied NYT and WebNLG datasets, we show that our approach outperforms previous studies and establishes a new state-of-the-art (92.91 and 91.37 F1 respectively in exact match evaluation).
Hongyang Chang, Hongfei Xu, Josef van Genabith, Deyi Xiong, Hongying Zan
IEEE ACM Trans. Audio Speech Lang. Process.2
2022 ParaZh-22M: A Large-Scale Chinese Parabank via Machine Translation
abstract
Paraphrasing, i.e., restating the same meaning in different ways, is an important data augmentation approach for natural language processing (NLP). Zhang et al. (2019b) propose to extract sentence-level paraphrases from multiple Chinese translations of the same source texts, and construct the PKU Paraphrase Bank of 0.5M sentence pairs. However, despite being the largest Chinese parabank to date, the size of PKU parabank is limited by the availability of one-to-many sentence translation data, and cannot well support the training of large Chinese paraphrasers. In this paper, we relieve the restriction with one-to-many sentence translation data, and construct ParaZh-22M, a larger Chinese parabank that is composed of 22M sentence pairs, based on one-to-one bilingual sentence translation data and machine translation (MT). In our data augmentation experiments, we show that paraphrasing based on ParaZh-22M can bring about consistent and significant improvements over several strong baselines on a wide range of Chinese NLP tasks, including a number of Chinese natural language understanding benchmarks (CLUE) and low-resource machine translation.
Wenjie Hao, Hongfei Xu, Deyi Xiong, Hongying Zan, Lingling Mu
COLING2
2021 Multi-Head Highly Parallelized LSTM Decoder for Neural Machine Translation
abstract
Hongfei Xu, Qiuhui Liu, Josef van Genabith, Deyi Xiong, Meng Zhang. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Hongfei Xu, Qiuhui Liu, Josef van Genabith, Deyi Xiong, Meng Zhang 0019
ACL/IJCNLP (1)1
2021 Self-Supervised Curriculum Learning for Spelling Error Correction
abstract
Spelling Error Correction (SEC) that requires high-level language understanding is a challenging but useful task.Current SEC approaches normally leverage a pre-training then fine-tuning procedure that treats data equally.By contrast, Curriculum Learning (CL) utilizes training data differently during training and has shown its effectiveness in improving both performance and training efficiency in many other NLP tasks.In NMT, a model's performance has been shown sensitive to the difficulty of training examples, and CL has been shown effective to address this.In SEC, the data from different language learners are naturally distributed at different difficulty levels (some errors made by beginners are obvious to correct while some made by fluent speakers are hard), and we expect that designing a curriculum correspondingly for model learning may also help its training and bring about better performance.In this paper, we study how to further improve the performance of the state-of-the-art SEC method with CL, and propose a Self-Supervised Curriculum Learning (SSCL) approach.Specifically, we directly use the cross-entropy loss as criteria for: 1) scoring the difficulty of training data, and 2) evaluating the competence of the model.In our approach, CL improves the model training, which in return improves the CL measurement.In our experiments on the SIGHAN 2015 Chinese spelling check task, we show that SSCL is superior to previous norm-based and uncertainty-aware approaches, and establish a new state of the art (74.38%F1).
Zifa Gan, Hongfei Xu, Hongying Zan
EMNLP (1)2
2021 Probing Word Translations in the Transformer and Trading Decoder for Encoder Layers
abstract
Hongfei Xu, Josef van Genabith, Qiuhui Liu, Deyi Xiong. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021.
Hongfei Xu, Josef van Genabith, Qiuhui Liu, Deyi Xiong
NAACL-HLT1
2020 Dynamically Adjusting Transformer Batch Size by Monitoring Gradient Direction Change
abstract
The choice of hyper-parameters affects the performance of neural models.While much previous research (Sutskever et al., 2013;Duchi et al., 2011;Kingma and Ba, 2015) focuses on accelerating convergence and reducing the effects of the learning rate, comparatively few papers concentrate on the effect of batch size.In this paper, we analyze how increasing batch size affects gradient direction, and propose to evaluate the stability of gradients with their angle change.Based on our observations, the angle change of gradient direction first tends to stabilize (i.e.gradually decrease) while accumulating mini-batches, and then starts to fluctuate.We propose to automatically and dynamically determine batch sizes by accumulating gradients of mini-batches and performing an optimization step at just the time when the direction of gradients starts to fluctuate.To improve the efficiency of our approach for large models, we propose a sampling approach to select gradients of parameters sensitive to the batch size.Our approach dynamically determines proper and efficient batch sizes during training.In our experiments on the WMT 14 English to German and English to French tasks, our approach improves the Transformer with a fixed 25k batch size by +0.73 and +0.82 BLEU respectively.
Hongfei Xu, Josef van Genabith, Deyi Xiong, Qiuhui Liu
ACL1
2020 Learning Source Phrase Representations for Neural Machine Translation
abstract
The Transformer translation model (Vaswani et al., 2017) based on a multi-head attention mechanism can be computed effectively in parallel and has significantly pushed forward the performance of Neural Machine Translation (NMT).Though intuitively the attentional network can connect distant words via shorter network paths than RNNs, empirical analysis demonstrates that it still has difficulty in fully capturing long-distance dependencies (Tang et al., 2018).Considering that modeling phrases instead of words has significantly improved the Statistical Machine Translation (SMT) approach through the use of larger translation blocks ("phrases") and its reordering ability, modeling NMT at phrase level is an intuitive proposal to help the model capture long-distance relationships.In this paper, we first propose an attentive phrase representation generation mechanism which is able to generate phrase representations from corresponding token representations.In addition, we incorporate the generated phrase representations into the Transformer translation model to enhance its ability to capture long-distance relationships.In our experiments, we obtain significant improvements on the WMT 14 English-German and English-French tasks on top of the strong Transformer baseline, which shows the effectiveness of our approach.Our approach helps Transformer Base models perform at the level of Transformer Big models, and even significantly better for long sentences, but with substantially fewer parameters and training steps.The fact that phrase representations help even in the big setting further supports our conjecture that they make a valuable contribution to long-distance relations.
Hongfei Xu, Josef van Genabith, Deyi Xiong, Qiuhui Liu, Jingyi Zhang 0002
ACL1
2020 Lipschitz Constrained Parameter Initialization for Deep Transformers
abstract
The Transformer translation model employs residual connection and layer normalization to ease the optimization difficulties caused by its multi-layer encoder/decoder structure. Previous research shows that even with residual connection and layer normalization, deep Transformers still have difficulty in training, and particularly Transformer models with more than 12 encoder/decoder layers fail to converge. In this paper, we first empirically demonstrate that a simple modification made in the official implementation, which changes the computation order of residual connection and layer normalization, can significantly ease the optimization of deep Transformers. We then compare the subtle differences in computation order in considerable detail, and present a parameter initialization method that leverages the Lipschitz constraint on the initialization of Transformer parameters that effectively ensures training convergence. In contrast to findings in previous research we further demonstrate that with Lipschitz parameter initialization, deep Transformers with the original computation order can converge, and obtain significant BLEU improvements with up to 24 layers. In contrast to previous research which focuses on deep encoders, our approach additionally enables Transformers to also benefit from deep decoders.
Hongfei Xu, Qiuhui Liu, Josef van Genabith, Deyi Xiong, Jingyi Zhang 0002
ACL1
2020 The Transference Architecture for Automatic Post-Editing
abstract
In automatic post-editing (APE) it makes sense to condition post-editing (pe) decisions on both the source (src) and the machine translated text (mt) as input.This has led to multi-encoder based neural APE approaches.A research challenge now is the search for architectures that best support the capture, preparation and provision of src and mt information and its integration with pe decisions.In this paper we present an efficient multi-encoder based APE model, called transference.Unlike previous approaches, it (i) uses a transformer encoder block for src, (ii) followed by a decoder block, but without masking for self-attention on mt, which effectively acts as second encoder combining src → mt, and (iii) feeds this representation into a final decoder block generating pe.Our model outperforms the best performing systems by 1 BLEU point on the WMT 2016, 2017, and 2018 English-German APE shared tasks (PBSMT and NMT).Furthermore, the results of our model on the WMT 2019 APE task using NMT data shows performance at the level of the state-of-the-art.The inference time of our model is similar to the vanilla transformer-based NMT system although our model deals with two separate encoders.We further investigate the importance of our newly introduced second encoder and find that decreasing the number of layers hurts performance, while reducing the number of layers of the decoder does not matter much.
Santanu Pal, Hongfei Xu, Nico Herbig 0001, Sudip Kumar Naskar, Antonio Krüger, Josef van Genabith
COLING2
2020 Efficient Context-Aware Neural Machine Translation with Layer-Wise Weighting and Input-Aware Gating
abstract
Existing Neural Machine Translation (NMT) systems are generally trained on a large amount of sentence-level parallel data, and during prediction sentences are independently translated, ignoring cross-sentence contextual information. This leads to inconsistency between translated sentences. In order to address this issue, context-aware models have been proposed. However, document-level parallel data constitutes only a small part of the parallel data available, and many approaches build context-aware models based on a pre-trained frozen sentence-level translation model in a two-step training manner. The computational cost of these approaches is usually high. In this paper, we propose to make the most of layers pre-trained on sentence-level data in contextual representation learning, reusing representations from the sentence-level Transformer and significantly reducing the cost of incorporating contexts in translation. We find that representations from shallow layers of a pre-trained sentence-level encoder play a vital role in source context encoding, and propose to perform source context encoding upon weighted combinations of pre-trained encoder layers' outputs. Instead of separately performing source context and input encoding, we propose to iteratively and jointly encode the source input and its contexts and to generate input-aware context representations with a cross-attention layer and a gating mechanism, which resets irrelevant information in context encoding. Our context-aware Transformer model outperforms the recent CADec [Voita et al., 2019c] on the English-Russian subtitle data and is about twice as fast in training and decoding.
Hongfei Xu, Deyi Xiong, Josef van Genabith, Qiuhui Liu
IJCAI1
2020 CMeIE: Construction and Evaluation of Chinese Medical Information Extraction Dataset
Tongfeng Guan, Hongying Zan, Xiabing Zhou, Hongfei Xu, Kunli Zhang
NLPCC (1)4
2020 Efficient processing of moving collective spatial keyword queries
Hongfei Xu, Yu Gu 0002, Yu Sun 0021, Jianzhong Qi 0001, Ge Yu 0001, Rui Zhang 0003
VLDB J.1
2019 Diversifying Top-k Routes with Spatial Constraints
Hongfei Xu, Yu Gu 0002, Jianzhong Qi 0001, Estrid He, Ge Yu 0001
J. Comput. Sci. Technol.1
2017 The Moving K Diversified Nearest Neighbor Query
abstract
We study result diversification in continuous spatial query processing and formulate a new type of queries, the moving k diversified nearest neighbor query (MkDNN). Given a moving query object, an MkDNN query maintains continuously the k diversified nearest neighbors of the query object. Here, how diversified the nearest neighbors are is defined on the distance between the nearest neighbors. We propose an algorithm to maintain incrementally the k diversified nearest neighbors to reduce the costs of continuous query processing. We further propose two approximate algorithms to obtain even higher query efficiency with precision bounds. We verify the effectiveness and efficiency of the proposed algorithms empirically. The results confirm the superiority of the proposed algorithms.
Yu Gu 0002, Guanli Liu, Jianzhong Qi 0001, Hongfei Xu, Ge Yu 0001, Rui Zhang 0003
ICDE4
2017 Improving Chinese-English Neural Machine Translation with Detected Usages of Function Words
Kunli Zhang, Hongfei Xu, Deyi Xiong, Qiuhui Liu, Hongying Zan
NLPCC2
2017 Reconfigurable interlocking furniture
abstract
Reconfigurable assemblies consist of a common set of parts that can be assembled into different forms for use in different situations. Designing these assemblies is a complex problem, since it requires a compatible decomposition of shapes with correspondence across forms, and a planning of well-matched joints to connect parts in each form. This paper presents computational methods as tools to assist the design and construction of reconfigurable assemblies, typically for furniture. There are three key contributions in this work. First, we present the compatible decomposition as a weakly-constrained dissection problem, and derive its solution based on a dynamic bipartite graph to construct parts across multiple forms; particularly, we optimize the parts reuse and preserve the geometric semantics. Second, we develop a joint connection graph to model the solution space of reconfigurable assemblies with part and joint compatibility across different forms. Third, we formulate the backward interlocking and multi-key interlocking models, with which we iteratively plan the joints consistently over multiple forms. We show the applicability of our approach by constructing reconfigurable furniture of various complexities, extend it with recursive connections to generate extensible and hierarchical structures, and fabricate a number of results using 3D printing, 2D laser cutting, and woodworking.
Peng Song 0001, Chi-Wing Fu, Yueming Jin, Hongfei Xu, Ligang Liu 0001, Pheng-Ann Heng, Daniel Cohen-Or
ACM Trans. Graph.4
2017 Computational design of wind-up toys
abstract
Wind-up toys are mechanical assemblies that perform intriguing motions driven by a simple spring motor. Due to the limited motor force and small body size, wind-up toys often employ higher pair joints of less frictional contacts and connector parts of nontrivial shapes to transfer motions. These unique characteristics make them hard to design and fabricate as compared to other automata. This paper presents a computational system to aid the design of wind-up toys, focusing on constructing a compact internal wind-up mechanism to realize user-requested part motions. Our key contributions include an analytical modeling of a wide variety of elemental mechanisms found in common wind-up toys, including their geometry and kinematics, conceptual design of wind-up mechanisms by computing motion transfer trees to realize the requested part motions, automatic construction of wind-up mechanisms by connecting multiple elemental mechanisms, and an optimization on the part and joint geometry with an objective of compacting the mechanism, reducing its weight, and avoiding collision. We use our system to design wind-up toys of various forms, fabricate a number of them using 3D printing, and show the functionality of various results.
Peng Song 0001, Xiao Tang 0005, Chi-Wing Fu, Hongfei Xu, Ligang Liu 0001, Niloy J. Mitra
ACM Trans. Graph.5
2016 The Moving K Diversified Nearest Neighbor Query
abstract
As a major type of continuous spatial queries, the moving$k$nearest neighbor ($k$NN) query has been studied extensively. However, most existing studies have focused on only the query efficiency. In this paper, we consider further the usability of the query results, in particular the diversification of the returned data points. We thereby formulate a new type of query named themoving$k$diversified nearest neighbor query (M$k$DNN). This type of query continuously reports the$k$diversified nearest neighbors while the query object is moving. Here, the degree of diversity of the$k$NN set is defined on the distance between the objects in the$k$NN set. Computing the$k$diversified nearest neighbors is an NP-hard problem. We propose an algorithm to maintain incrementally the$k$diversified nearest neighbors to reduce the query processing costs. We further propose two approximate algorithms to obtain even higher query efficiency with precision bounds. We verify the effectiveness and efficiency of the proposed algorithms both theoretically and empirically. The results confirm the superiority of the proposed algorithms over the baseline algorithm.
Yu Gu 0002, Guanli Liu, Jianzhong Qi 0001, Hongfei Xu, Ge Yu 0001, Rui Zhang 0003
IEEE Trans. Knowl. Data Eng.4