VLDB 2026 Research / reviewers in the wild / expert
Xiaojiang Liu
dblp:21/4323
· DBLP profile ↗
41ranked-venue papers
5as first author
15since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 27 · 1 first-author · 4 since 2021Computer networks · 9 · 4 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5Applied, interdisciplinary, general and emerging computing · 4 · 3 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Competitive Learning-Based Clock Parameters Estimation for PTP Synchronization With Unknown Delay DistributionsabstractClock synchronization is a crucial requirement for coordinated activities in distributed networks. As an effective solution, Precision Time Protocol (PTP) is tailored to provide tight synchronization which, however, suffers from packet delay variation (PDV). Therefore, it is a challenging task to mitigate the uncertainties of PDV so that clock skew and offset can be estimated with higher accuracy and robustness. Under the assumption of delay symmetry, PDV is caused by unavailable prior information about the statistical distributions of stochastic delay. This paper investigates a robust estimator that can achieve joint estimation for clock skew and offset under delay with unknown distributions by employing the competitive learning-based rival penalized expectation maximization (RPEM) algorithm to learn the unknown probability density function (pdf) of stochastic delay. Moreover, for the asymmetric delay scenario, besides the unknown delay distributions, delay asymmetry is another important influence on synchronization accuracy. Therefore, the optimal invariant robust estimator is developed to simultaneously handle the performance degradation caused by delay asymmetry and unknown delay distributions. The estimator can even cope with the scenario where the random delays in the uplink and downlink follow different distributions. The effectiveness of the estimation schemes is validated by the computer simulations. Heng Wang 0003, Wenqiao Ma, Xiaojiang Liu, Xiong Zhu, Min Li 0005 |
IEEE Trans. Commun. | 4 |
| 2026 | Resilient Clock Offset Estimation for PTP Synchronization With Unknown and Unmodeled Delay Distributions in Industrial NetworksabstractPrecise clock synchronization is a critical requirement in industrial networks. To meet this need, the Precision Time Protocol (PTP) has been widely adopted as the standard solution. Nevertheless, the protocol’s accuracy is often compromised by unknown and unmodeled random delays, presenting a significant obstacle to reliable network operation. In this article, we propose a robust clock offset estimation scheme to enable a much higher synchronization accuracy for PTP in the presence of unknown and unmodeled delay distributions. To make this possible, we first derive the probability density function (pdf) of the delays via a maximum entropy approach, which is constrained by fractional moments and an unbiased likelihood estimation. Based on the derived pdfs, we further develop a low computational cost L-estimator for the clock offset, which is a linear function of the order statistics and robust against unknown and unmodeled random queueing delays. Simulation results illustrate that the proposed clock offset estimation scheme has much lower computational complexity and yields much more accurate estimation results than the compared existing methods in the presence of unknown and unmodeled disturbances. Heng Wang 0003, Xiaojiang Liu, Min Li 0005 |
IEEE Trans. Ind. Informatics | 4 |
| 2025 | MR. Judge: Multimodal Reasoner as a JudgeabstractThe paradigm of using Large Language Models (LLMs) and Multimodal Large Language Models (MLLMs) as evaluative judges has emerged as an effective approach in RLHF and inference-time scaling. In this work, we propose Multimodal Reasoner as a Judge (MR. Judge), a paradigm for empowering general-purpose MLLMs judges with strong reasoning capabilities. Instead of directly assigning scores for each response, we formulate the judgement process as a reasoning-inspired multiple-choice problem. Specifically, the judge model first conducts deliberate reasoning covering different aspects of the responses and eventually selects the best response from them. This reasoning process not only improves the interpretibility of the judgement, but also greatly enhances the performance of MLLM judges. To cope with the lack of questions with scored responses, we propose the following strategy to achieve automatic annotation: 1) Reverse Response Candidates Synthesis: starting from a supervised fine-tuning (SFT) dataset, we treat the original response as the best candidate and prompt the MLLM to generate plausible but flawed negative candidates. 2) Text-based reasoning distillation: we carefully design a data synthesis pipeline for distilling the reasoning capability from a text-based reasoning model, which is adopted to enable the MLLM judges to regain complex reasoning ability via warm up supervised fine-tuning. Experiments demonstrate that our MR. Judge is effective across a wide range of tasks. Specifically, our MR. Judge-7B surpasses GPT-4o by 9.9% on VL-RewardBench, and improves performance on MM-Vet during inference-time scaling by up to 7.7%. Renjie Pi, Haoping Bai, Xiaoming Simon Wang, Jiulong Shan, Xiaojiang Liu |
EMNLP | 6 |
| 2025 | TIS-DPO: Token-level Importance Sampling for Direct Preference Optimization With Estimated WeightsabstractDirect Preference Optimization (DPO) has been widely adopted for preference alignment of Large Language Models (LLMs) due to its simplicity and effectiveness.
However, DPO is derived as a bandit problem in which the whole response is treated as a single arm, ignoring the importance differences between tokens, which may affect optimization efficiency and make it difficult to achieve optimal results.
In this work, we propose that the optimal data for DPO has equal expected rewards for each token in winning and losing responses, as there is no difference in token importance.
However, since the optimal dataset is unavailable in practice, we propose using the original dataset for importance sampling to achieve unbiased optimization.
Accordingly, we propose a token-level importance sampling DPO objective named TIS-DPO that assigns importance weights to each token based on its reward.
Inspired by previous works, we estimate the token importance weights using the difference in prediction probabilities from a pair of contrastive LLMs. We explore three methods to construct these contrastive LLMs: (1) guiding the original LLM with contrastive prompts, (2) training two separate LLMs using winning and losing responses, and (3) performing forward and reverse DPO training with winning and losing responses.
Experiments show that TIS-DPO significantly outperforms various baseline methods on harmlessness and helpfulness alignment and summarization tasks. We also visualize the estimated weights, demonstrating their ability to identify key token positions. Aiwei Liu, Haoping Bai, Zhiyun Lu, Yanchao Sun, Xiang Kong, Xiaoming Simon Wang, Jiulong Shan, Albin Madappally Jose, Xiaojiang Liu, Lijie Wen 0001, Philip S. Yu |
ICLR | 9 |
| 2025 | A Rapid Time Synchronization Scheme Using Virtual Links and Maximum Consensus for Wireless Sensor NetworksabstractConsensus-based time synchronization, which combines multiagent consensus and distributed network techniques, becomes an essential approach to achieve accurate time synchronization in wireless sensor networks. In actual synchronization process, random communication delays will have a negative impact on information exchange, which leads to a significant degradation in the synchronization performance. Besides, the convergence speed of synchronization errors in existing average consensus synchronization algorithms poses challenges due to the iterative update. To deal with these issues, this article proposes a maximum consensus time synchronization scheme based on the multihop virtual links. Primarily, a multihop virtual links mechanism is proposed, which shares clock information among nodes to shorten the convergence time. Meanwhile, a noniterative clock parameters compensation approach based on maximum consensus is developed to further speed up the synchronization process. In addition, a moving average estimator is designed to minimize the adverse impact of delays. The convergence performance of the proposed scheme is proved theoretically. Simulation results verify the correctness of theoretical analysis and also demonstrate that the proposed scheme outperforms existing algorithms in terms of synchronization accuracy and convergence speed under delays. Heng Wang 0003, Yan Zou, Xiaojiang Liu, Zhenya Meng |
IEEE Internet Things J. | 3 |
| 2025 | Robust Clock Parameters Tracking for IEEE 1588 With Asymmetric Packet Delays in Industrial NetworksabstractClock synchronization is a prerequisite for the proper operation of industrial networks. IEEE 1588 Precision Time Protocol (PTP) can provide tight synchronization for a vast number of industrial applications. However, the joint tracking of clock skew and offset for IEEE 1588 under the more realistic delay asymmetries scenario is a challenging problem in sophisticated industrial networks. In this paper, we first investigate the biased estimation problem for clock parameters tracking in the presence of delay asymmetries, which significantly degrade the synchronization accuracy. Based on the investigation, a state-space model is developed that is capable of overcoming the uncertainty caused by the unknown packet delay distribution. A recursive joint clock skew and offset tracking scheme that employs the robust three-step recursive Kalman filter (R3SRKF) is proposed to estimate the time-varying clock parameters in a minimum-variance unbiased manner under the asymmetric packet delays scenario. Also, a variant of the R3SRKF is derived to joint track clock skew and offset with asymmetric packet delays in multi-hop networks. Simulation results indicate the effectiveness and performance enhancement of the presented robust tracking scheme. Xiaojiang Liu, Heng Wang 0003 |
IEEE Trans. Commun. | 1 |
| 2025 | Accurate and Robust Clock Parameters Estimation for Energy-Efficient Time Synchronization With Multi-Link Overhearing in Wireless NetworksabstractTime synchronization is a significant premise for the proper functioning of wireless networks. Several energy-efficient protocols targeting time synchronization have been proposed to maximize accuracy while minimizing energy consumption. Since no packet transmission is required, receiver-only synchronization has been extensively studied in wireless networks. However, there is a performance loss in synchronization accuracy and robustness because inactive nodes are synchronized by passively receiving messages over a single link. In wireless environments, an inactive node inevitably overhears paired synchronization messages from multiple communication links due to the broadcast characteristic. Based on this consideration, an accurate and robust clock parameters estimation algorithm for energy-efficient synchronization with multi-link overhearing is developed. Specifically, a robust joint estimator of clock skew and offset based on invariant decision rules is derived with the assumption that prior information regarding the availability of overhearing links is known. In the absence of such information, as in harsh wireless channel conditions, we also propose a linear aggregation estimation mechanism with outlier diagnostic, which further enhances the robustness and resilience of time synchronization over wireless networks. Finally, numerical simulations are implemented and simulation results demonstrate that the synchronization accuracy and robustness are significantly enhanced. Xiaojiang Liu, Heng Wang 0003 |
IEEE Trans. Wirel. Commun. | 1 |
| 2025 | Multi-Hop Timestamp-Free Synchronization With Arbitrary Distributed Delays in Wireless NetworksabstractTimestamp-free synchronization protocol is tailored to provide a global time understanding for resource-limited wireless networks as it eliminates timestamp interaction, thereby minimizing additional resource overheads. However, the existing two-hop based timestamp-free protocols are not suitable for synchronizing all nodes in multi-hop networks, as they necessitate multiple response times to establish the timestamp relationship between each pair of neighboring nodes. To this end, we introduce a novel multi-hop timestamp-free synchronization protocol. The proposed protocol allows any two nodes to be synchronized using only local timestamps and the skew estimates embedded within packets traversing the reverse path. Furthermore, considering that synchronization accuracy suffers from delay variation resulting from packet loss or retransmission in wireless networks, we derive a Pitman estimator to estimate the clock skew under arbitrary delay models, given known information. To further target unknown arbitrary delay distributions, we approximate the probability density function (pdf) of stochastic delays using a Gaussian mixture model, and then learn the pdf using the rival penalized expectation maximization algorithm. With the aid of the learned pdf, the robustness-enhanced Pitman estimator is derived, which is robust against arbitrary distributed delays without known knowledge. The effectiveness and performance enhancement of estimators are validated by simulations. Heng Wang 0003, Wenqiao Ma, Xiaojiang Liu, Min Li 0005 |
IEEE Trans. Wirel. Commun. | 3 |
| 2024 | A Simple Event-Based Average Consensus Clock Synchronization Scheme in Industrial Wireless Sensor Networks Under Communication DelaysabstractIn industrial wireless sensor networks, the existing event-based consensus clock synchronization schemes can achieve network-wide clock synchronization without considering the communication delays. However, communication delays are a significant factor restricting the achievement of clock synchronization in realistic scenarios. Besides, a constraint to be considered in the design of synchronization scheme is the limited communication resources of sensor nodes. Therefore, in this article, we investigate an event-based consensus clock synchronization scheme under delays. On the one hand, a new low-pass filter is used to estimate the relative skew to resist the influence of communication delays. On the other hand, a new event-triggered scheme is utilized to reduce unnecessary information transmission. Finally, we provide the convergence proof and simulation results of the proposed scheme, which further proves that it can achieve clock synchronization under communication delays while decreasing communication overhead. Heng Wang 0003, Xiaojiang Liu, Yan Zou, Min Li 0005 |
IEEE Trans. Ind. Informatics | 3 |
| 2024 | Scheduling for Minimizing the Age of Information in Multisensor Multiserver Industrial Internet of Things SystemsabstractReal-time data delivery is significant for the Industrial Internet of Things (IIoT). Age of information (AoI), a popular real-time metric, is usually used to measure the data freshness of the IIoT systems. If the data most recently received by the destination at time $t$ was generated at time $t_{1}$ , then the AoI is $t-t_{1}$ . In this paper, we consider a multi-sensor multi-server IIoT system and develop scheduling algorithms to minimize the average AoI. The challenge lies in the strong coupling between link scheduling, server selection, and service preemption. To address this issue, we propose a guided exploration-based deep Q-Network (GE-DQN) algorithm utilizing a fixed advantage policy, which has a faster learning speed compared to classical deep Q-Network. Moreover, we use a shared decision module followed by several network branches to transform the structure of GE-DQN and propose a guided exploration-based Branching Dueling Q-Network (GE-BDQN) algorithm. Since the branch structure of GE-BDQN can decompose the high-dimensional action, GE-BDQN can reduce the approximate exponential growth of the number of output neurons with the increase of the number of sensors to linear growth compared to GE-DQN, ensuring the applicability of the algorithm under large-scale systems. From the simulation results, it can be found that the proposed two algorithms can achieve better average AoI compared to the advanced algorithms, and the GE-BDQN algorithm can achieve up to 36% performance gain. Xin Xie 0004, Heng Wang 0003, Xiaojiang Liu |
IEEE Trans. Ind. Informatics | 3 |
| 2023 | Robust Clock Skew and Offset Estimation for PTP Synchronization With Unknown Delay DistributionsabstractClock synchronization is crucial requirement for coordinated activities in distributed networks. Precision Time Protocol (PTP) is a popular synchronization scheme to provide tight synchronization which, however, suffers from packet delay variation (PDV). In real networks, prior information on the statistical distribution of stochastic delays is always unavailable. Therefore, it is significant to estimate the clock skew and offset in the presence of stochastic delays with unknown statistical distributions. This paper investigates a robust estimation approach applicable to scenarios where the delay distributions are unknown, which utilizes the Gaussian mixture model to approximate the probability density function (pdf) of random delays, and then employs the rival penalized expectation maximization (RPEM) algorithm to learn mixture model parameters. The optimum invariant estimators for the clock skew and offset are derived based on the learned pdfs of random delays. The effectiveness of the estimation scheme is validated by the computer simulations. Heng Wang 0003, Wenqiao Ma, Xiaojiang Liu |
GLOBECOM | 3 |
| 2023 | Improving Accuracy and Robustness of Clock Parameters Estimation Using Multi-Link Overhearing in Wireless Sensor NetworksabstractReceiver-only synchronization (ROS) is a strategy in which nodes only need to receive packets to achieve clock synchronization. Since no packet transmission is required, it has been extensively studied in wireless sensor networks (WSNs) where energy resources are rare. Nevertheless, there is a performance loss in synchronization accuracy and robustness of the inactive nodes in the ROS scenario because these nodes are synchronized with a reference node only by overhearing. In addition, it is observed that in a practical network, the execution of ROS inevitably overhears paired synchronization messages from multiple communication paths. With the help of this feature, on the basis of the clock parameters estimation method and information aggregation technique, a linear aggregation estimation mechanism with outlier diagnostic is presented to improve the synchronization accuracy and robustness in ROS with multi-link overhearing scenario. Finally, numerical simulations are implemented based on two single path clock parameters estimation algorithms of inactive nodes, and simulation results demonstrate that the synchronization accuracy and robustness are significantly enhanced. Xiaojiang Liu, Heng Wang 0003 |
ICC | 1 |
| 2023 | A Robust and Low-Complexity Estimation Scheme for Clock Skew Without Timestamp Exchange in Wireless Sensor NetworksabstractWireless sensor networks (WSNs) have a broad range of applications, and time synchronization is essential to ensure their correct operation. In this paper, a clock skew estimation scheme is presented for timestamp-free synchronization, which greatly simplifies computation while economizing energy consumption. Considering that no single delay model can fit all cases in WSNs, a robust estimator with low complexity based on this scheme is developed that can estimate clock skew accurately without prior statistical characteristics of delays. Simulation results demonstrate the effectiveness of the proposed scheme. Min Li 0005, Fangshi Wang, Xiaojiang Liu, Heng Wang 0003 |
VTC Fall | 3 |
| 2022 | Learning to Break the Loop: Analyzing and Mitigating Repetitions for Neural Text GenerationabstractWhile large-scale neural language models, such as GPT2 and BART,have achieved impressive results on various text generation tasks, they tend to get stuck in undesirable sentence-level loops with maximization-based decoding algorithms (\textit{e.g.}, greedy search). This phenomenon is counter-intuitive since there are few consecutive sentence-level repetitions in the human corpus (e.g., 0.02\% in Wikitext-103). To investigate the underlying reasons for generating consecutive sentence-level repetitions, we study the relationship between the probability of repetitive tokens and their previous repetitions in context. Through our quantitative experiments, we find that 1) Models have a preference to repeat the previous sentence; 2) The sentence-level repetitions have a \textit{self-reinforcement effect}: the more times a sentence is repeated in the context, the higher the probability of continuing to generate that sentence; 3) The sentences with higher initial probabilities usually have a stronger self-reinforcement effect. Motivated by our findings, we propose a simple and effective training method \textbf{DITTO} (Pseu\underline{D}o-Repet\underline{IT}ion Penaliza\underline{T}i\underline{O}n), where the model learns to penalize probabilities of sentence-level repetitions from synthetic repetitive data. Although our method is motivated by mitigating repetitions, our experiments show that DITTO not only mitigates the repetition issue without sacrificing perplexity, but also achieves better generation quality. Extensive experiments on open-ended text generation (Wikitext-103) and text summarization (CNN/DailyMail) demonstrate the generality and effectiveness of our method. Jin Xu 0010, Xiaojiang Liu, Jianhao Yan, Deng Cai 0002, Jian Li 0015 |
NeurIPS | 2 |
| 2022 | Interpretable Real-Time Win Prediction for Honor of Kings - A Popular Mobile MOBA EsportabstractWith the rapid prevalence and explosive development of Multiplayer Online Battle Arena electronic sports (MOBA esports), much research effort has been devoted to automatically predicting game results (win predictions). While this task has great potential in various applications, such as esports live streaming and game commentator artificial intelligence systems, previous studies fail to investigate the methods tointerpretthese win predictions. To mitigate this issue, we collected a large-scale dataset that contains real-time game records with rich input features of the popular MOBA gameHonor of Kings. For interpretable predictions, we proposed a two-stage spatial–temporal network (TSSTN) that can not only provide accurate real-time win predictions but also attribute the ultimate prediction results to the contributions of different features for interpretability. Experiment results and applications in real-world live streaming scenarios showed that the proposed TSSTN model is effective in both prediction accuracy and interpretability. Zelong Yang 0002, Zhufeng Pan, Yan Wang 0060, Deng Cai 0002, Shuming Shi 0001, Shao-Lun Huang, Wei Bi, Xiaojiang Liu |
IEEE Trans. Games | 8 |
| 2020 | Learning to Select Bi-Aspect Information for Document-Scale Text Content Manipulation
Yawei Sun, Bing Qin 0001, Heng Gong, Wei Bi, Xiaojiang Liu, Ting Liu 0001 |
AAAI | 7 |
| 2020 | Relevance-Promoting Language Model for Short-Text ConversationabstractDespite the effectiveness of sequence-to-sequence framework on the task of Short-Text Conversation (STC), the issue of under-exploitation of training data (i.e., the supervision signals from query text is ignored) still remains unresolved. Also, the adopted maximization-based decoding strategies, inclined to generating the generic responses or responses with repetition, are unsuited to the STC task. In this paper, we propose to formulate the STC task as a language modeling problem and tailor-make a training strategy to adapt a language model for response generation. To enhance generation performance, we design a relevance-promoting transformer language model, which performs additional supervised source attention after the self-attention to increase the importance of informative query tokens in calculating the token-level representation. The model further refines the query representation with relevance clues inferred from its multiple references during training. In testing, we adopt a randomization-over-maximization strategy to reduce the generation of generic responses. Experimental results on a large Chinese STC dataset demonstrate the superiority of the proposed model on relevance metrics and diversity metrics.1 Xin Li 0056, Piji Li, Wei Bi, Xiaojiang Liu, Wai Lam |
AAAI | 4 |
| 2020 | Improving Knowledge-Aware Dialogue Generation via Knowledge Base Question AnsweringabstractNeural network models usually suffer from the challenge of incorporating commonsense knowledge into the open-domain dialogue systems. In this paper, we propose a novel knowledge-aware dialogue generation model (called TransDG), which transfers question representation and knowledge matching abilities from knowledge base question answering (KBQA) task to facilitate the utterance understanding and factual knowledge selection for dialogue generation. In addition, we propose a response guiding attention and a multi-step decoding strategy to steer our model to focus on relevant features for response generation. Experiments on two benchmark datasets demonstrate that our model has robust superiority over compared methods in generating informative and fluent dialogues. Our code is available at https://github.com/siat-nlp/TransDG. Jian Wang 0054, Junhao Liu 0001, Wei Bi, Xiaojiang Liu, Kejing He 0001, Ruifeng Xu 0001, Min Yang 0007 |
AAAI | 4 |
| 2020 | Rigid Formats Controlled Text GenerationabstractNeural text generation has made tremendous progress in various tasks.One common characteristic of most of the tasks is that the texts are not restricted to some rigid formats when generating.However, we may confront some special text paradigms such as Lyrics (assume the music score is given), Sonnet, SongCi (classical Chinese poetry of the Song dynasty), etc.The typical characteristics of these texts are in three folds: (1) They must comply fully with the rigid predefined formats.(2) They must obey some rhyming schemes.(3) Although they are restricted to some formats, the sentence integrity must be guaranteed.To the best of our knowledge, text generation based on the predefined rigid formats has not been well investigated.Therefore, we propose a simple and elegant framework named SongNet to tackle this problem.The backbone of the framework is a Transformer-based auto-regressive language model.Sets of symbols are tailor-designed to improve the modeling performance especially on format, rhyme, and sentence integrity.We improve the attention mechanism to impel the model to capture some future information on the format.A pre-training and fine-tuning framework is designed to further improve the generation quality.Extensive experiments conducted on two collected corpora demonstrate that our proposed framework generates significantly better results in terms of both automatic metrics and the human evaluation. 1 Piji Li, Haisong Zhang, Xiaojiang Liu, Shuming Shi 0001 |
ACL | 3 |
| 2020 | Generate, Delete and Rewrite: A Three-Stage Framework for Improving Persona Consistency of Dialogue GenerationabstractMaintaining a consistent personality in conversations is quite natural for human beings, but is still a non-trivial task for machines.The persona-based dialogue generation task is thus introduced to tackle the personalityinconsistent problem by incorporating explicit persona text into dialogue generation models.Despite the success of existing personabased models on generating human-like responses, their one-stage decoding framework can hardly avoid the generation of inconsistent persona words.In this work, we introduce a three-stage framework that employs a generate-delete-rewrite mechanism to delete inconsistent words from a generated response prototype and further rewrite it to a personality-consistent one.We carry out evaluations by both human and automatic metrics.Experiments on the Persona-Chat dataset show that our approach achieves good performance. Haoyu Song 0002, Yan Wang 0060, Weinan Zhang 0003, Xiaojiang Liu, Ting Liu 0001 |
ACL | 4 |
| 2020 | Response-Anticipated Memory for On-Demand Knowledge Integration in Response GenerationabstractNeural conversation models are known to generate appropriate but non-informative responses in general.A scenario where informativeness can be significantly enhanced is Conversing by Reading (CbR), where conversations take place with respect to a given external document.In previous work, the external document is utilized by (1) creating a contextaware document memory that integrates information from the document and the conversational context, and then (2) generating responses referring to the memory.In this paper, we propose to create the document memory with some anticipated responses in mind.This is achieved using a teacher-student framework.The teacher is given the external document, the context, and the ground-truth response, and learns how to build a response-aware document memory from three sources of information.The student learns to construct a response-anticipated document memory from the first two sources, and the teacher's insight on memory creation.Empirical results show that our model outperforms the previous stateof-the-art for the CbR task. Zhiliang Tian, Wei Bi, Lanqing Xue, Yiping Song, Xiaojiang Liu, Nevin Lianwen Zhang |
ACL | 6 |
| 2020 | A Batch Normalized Inference Network Keeps the KL Vanishing AwayabstractVariational Autoencoder (VAE) is widely used as a generative model to approximate a model's posterior on latent variables by combining the amortized variational inference and deep neural networks.However, when paired with strong autoregressive decoders, VAE often converges to a degenerated local optimum known as "posterior collapse".Previous approaches consider the Kullback-Leibler divergence (KL) individual for each datapoint.We propose to let the KL follow a distribution across the whole dataset, and analyze that it is sufficient to prevent posterior collapse by keeping the expectation of the KL's distribution positive.Then we propose Batch Normalized-VAE (BN-VAE), a simple but effective approach to set a lower bound of the expectation by regularizing the distribution of the approximate posterior's parameters.Without introducing any new model component or modifying the objective, our approach can avoid the posterior collapse effectively and efficiently.We further show that the proposed BN-VAE can be extended to conditional VAE (CVAE).Empirically, our approach surpasses strong autoregressive baselines on language modeling, text classification and dialogue generation, and rivals more complex approaches while keeping almost the same training time as VAE. Qile Zhu, Wei Bi, Xiaojiang Liu, Xiyao Ma, Xiaolin Li 0001, Dapeng Oliver Wu |
ACL | 3 |
| 2020 | TableGPT: Few-shot Table-to-Text Generation with Table Structure Reconstruction and Content MatchingabstractAlthough neural table-to-text models have achieved remarkable progress with the help of largescale datasets, they suffer insufficient learning problem with limited training data.Recently, pretrained language models show potential in few-shot learning with linguistic knowledge learnt from pretraining on large-scale corpus.However, benefiting table-to-text generation in few-shot setting with the powerful pretrained language model faces three challenges, including (1) the gap between the task's structured input and the natural language input for pretraining language model.(2) The lack of modeling for table structure and ( 3) improving text fidelity with less incorrect expressions that are contradicting to the table.To address aforementioned problems, we propose TableGPT for table-to-text generation.At first, we utilize table transformation module with template to rewrite structured table in natural language as input for GPT-2.In addition, we exploit multi-task learning with two auxiliary tasks that preserve table's structural information by reconstructing the structure from GPT-2's representation and improving the text's fidelity with content matching task aligning the table and information in the generated text.By experimenting on Humans, Songs and Books, three few-shot table-to-text datasets in different domains, our model outperforms existing systems on most few-shot settings. Heng Gong, Yawei Sun, Bing Qin 0001, Wei Bi, Xiaojiang Liu, Ting Liu 0001 |
COLING | 6 |
| 2020 | Dual Dynamic Memory Network for End-to-End Multi-turn Task-oriented Dialog SystemsabstractExisting end-to-end task-oriented dialog systems struggle to dynamically model long dialog context for interactions and effectively incorporate knowledge base (KB) information into dialog generation. To conquer these limitations, we propose a Dual Dynamic Memory Network (DDMN) for multi-turn dialog generation, which maintains two core components: dialog memory manager and KB memory manager. The dialog memory manager dynamically expands the dialog memory turn by turn and keeps track of dialog history with an updating mechanism, which encourages the model to filter irrelevant dialog history and memorize important newly coming information. The KB memory manager shares the structural KB triples throughout the whole conversation, and dynamically extracts KB information with a memory pointer at each turn. Experimental results on three benchmark datasets demonstrate that DDMN significantly outperforms the strong baselines in terms of both automatic evaluation and human evaluation. Our code is available at https://github.com/siat-nlp/DDMN. Jian Wang 0054, Junhao Liu 0001, Wei Bi, Xiaojiang Liu, Kejing He 0001, Ruifeng Xu 0001, Min Yang 0007 |
COLING | 4 |
| 2020 | The World is Not Binary: Learning to Rank with Grayscale Data for Dialogue Response SelectionabstractResponse selection plays a vital role in building retrieval-based conversation systems.Despite that response selection is naturally a learning-to-rank problem, most prior works take a point-wise view and train binary classifiers for this task: each response candidate is labeled either relevant (one) or irrelevant (zero).On the one hand, this formalization can be sub-optimal due to its ignorance of the diversity of response quality.On the other hand, annotating grayscale data for learning-to-rank can be prohibitively expensive and challenging.In this work, we show that grayscale data can be automatically constructed without human effort.Our method employs off-the-shelf response retrieval models and response generation models as automatic grayscale data generators.With the constructed grayscale data, we propose multi-level ranking objectives for training, which can (1) teach a matching model to capture more fine-grained context-response relevance difference and (2) reduce the traintest discrepancy in terms of distractor strength.Our method is simple, effective, and universal.Experiments on three benchmark datasets and four state-of-the-art matching models show that the proposed approach brings significant and consistent performance improvements. Zibo Lin, Deng Cai 0002, Yan Wang 0060, Xiaojiang Liu, Hai-Tao Zheng 0002, Shuming Shi 0001 |
EMNLP (1) | 4 |
| 2020 | Event Extraction as Machine Reading ComprehensionabstractEvent extraction (EE) is a crucial information extraction task that aims to extract event information in texts.Previous methods for EE typically model it as a classification task, which are data-hungry and suffer from the data scarcity problem.In this paper, we propose a new learning paradigm of EE, by explicitly casting it as a machine reading comprehension problem (MRC).Our approach includes an unsupervised question generation process, which can transfer event schema into a set of natural questions, followed by a BERTbased question-answering process to retrieve answers as EE results.This learning paradigm enables us to strengthen the reasoning process of EE, by introducing sophisticated models in MRC, and relieve the data scarcity problem, by introducing the large-scale datasets in MRC.The empirical results show that: i) our approach attains state-of-the-art performance by considerable margins over previous methods.ii) Our model is excelled in the data-scarce scenario, for example, obtaining 49.8% in F1 for event argument extraction with only 1% data, compared with 2.2% of the previous method.iii) Our model also fits with zero-shot scenarios, achieving 37.0% and 16% in F1 on two datasets without using any EE training data. Jian Liu 0032, Yubo Chen 0001, Kang Liu 0001, Wei Bi, Xiaojiang Liu |
EMNLP (1) | 5 |
| 2020 | Profile Consistency Identification for Open-domain Dialogue AgentsabstractMaintaining a consistent attribute profile is crucial for dialogue agents to naturally converse with humans.Existing studies on improving attribute consistency mainly explored how to incorporate attribute information in the responses, but few efforts have been made to identify the consistency relations between response and attribute profile.To facilitate the study of profile consistency identification, we create a large-scale human-annotated dataset with over 110K single-turn conversations and their key-value attribute profiles.Explicit relation between response and profile is manually labeled.We also propose a key-value structure information enriched BERT model to identify the profile consistency, and it gained improvements over strong baselines.Further evaluations on downstream tasks demonstrate that the profile consistency identification model is conducive for improving dialogue consistency. Haoyu Song 0002, Yan Wang 0060, Weinan Zhang 0003, Zhengyu Zhao 0003, Ting Liu 0001, Xiaojiang Liu |
EMNLP (1) | 6 |
| 2019 | Generating Multiple Diverse Responses for Short-Text ConversationabstractNeural generative models have become popular and achieved promising performance on short-text conversation tasks. They are generally trained to build a 1-to-1 mapping from the input post to its output response. However, a given post is often associated with multiple replies simultaneously in real applications. Previous research on this task mainly focuses on improving the relevance and informativeness of the top one generated response for each post. Very few works study generating multiple accurate and diverse responses for the same post. In this paper, we propose a novel response generation model, which considers a set of responses jointly and generates multiple diverse responses simultaneously. A reinforcement learning algorithm is designed to solve our model. Experiments on two short-text conversation tasks validate that the multiple responses generated by our model obtain higher quality and larger diversity compared with various state-ofthe-art generative models. Wei Bi, Xiaojiang Liu, Junhui Li 0001, Shuming Shi 0001 |
AAAI | 3 |
| 2019 | Better Fine-Tuning via Instance Weighting for Text ClassificationabstractTransfer learning for deep neural networks has achieved great success in many text classification applications. A simple yet effective transfer learning method is to fine-tune the pretrained model parameters. Previous fine-tuning works mainly focus on the pre-training stage and investigate how to pretrain a set of parameters that can help the target task most. In this paper, we propose an Instance Weighting based Finetuning (IW-Fit) method, which revises the fine-tuning stage to improve the final performance on the target domain. IW-Fit adjusts instance weights at each fine-tuning epoch dynamically to accomplish two goals: 1) identify and learn the specific knowledge of the target domain effectively; 2) well preserve the shared knowledge between the source and the target domains. The designed instance weighting metrics used in IW-Fit are model-agnostic, which are easy to implement for general DNN-based classifiers. Experimental results show that IW-Fit can consistently improve the classification accuracy on the target domain. Wei Bi, Yan Wang 0060, Xiaojiang Liu |
AAAI | 4 |
| 2019 | Fine-Grained Sentence Functions for Short-Text ConversationabstractSentence function is an important linguistic feature referring to a user's purpose in uttering a specific sentence.The use of sentence function has shown promising results to improve the performance of conversation models.However, there is no large conversation dataset annotated with sentence functions.In this work, we collect a new Short-Text Conversation dataset with manually annotated SEntence FUNctions (STC-Sefun).Classification models are trained on this dataset to (i) recognize the sentence function of new data in a large corpus of short-text conversations; (ii) estimate a proper sentence function of the response given a test query.We later train conversation models conditioned on the sentence functions, including information retrieval-based and neural generative models.Experimental results demonstrate that the use of sentence functions can help improve the quality of the returned responses. Wei Bi, Xiaojiang Liu, Shuming Shi 0001 |
ACL (1) | 3 |
| 2019 | Retrieval-guided Dialogue Response Generation via a Matching-to-Generation FrameworkabstractDeng Cai, Yan Wang, Wei Bi, Zhaopeng Tu, Xiaojiang Liu, Shuming Shi. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Deng Cai 0002, Yan Wang 0060, Wei Bi, Zhaopeng Tu, Xiaojiang Liu, Shuming Shi 0001 |
EMNLP/IJCNLP (1) | 5 |
| 2019 | A Discrete CVAE for Response Generation on Short-Text ConversationabstractJun Gao, Wei Bi, Xiaojiang Liu, Junhui Li, Guodong Zhou, Shuming Shi. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Wei Bi, Xiaojiang Liu, Junhui Li 0001, Guodong Zhou 0001, Shuming Shi 0001 |
EMNLP/IJCNLP (1) | 3 |
| 2019 | Improving Open-Domain Dialogue Systems via Multi-Turn Incomplete Utterance RestorationabstractZhufeng Pan, Kun Bai, Yan Wang, Lianqiang Zhou, Xiaojiang Liu. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Zhufeng Pan, Yan Wang 0060, Lianqiang Zhou, Xiaojiang Liu |
EMNLP/IJCNLP (1) | 5 |
| 2018 | Towards Less Generic Responses in Neural Conversation Models: A Statistical Re-weighting MethodabstractSequence-to-sequence neural generation models have achieved promising performance on short text conversation tasks.However, they tend to generate generic/dull responses, leading to unsatisfying dialogue experience.We observe that in conversation tasks, each query could have multiple responses, which forms a 1-to-n or m-to-n relationship in the view of the total corpus.The objective function used in standard sequence-to-sequence models will be dominated by loss terms with generic patterns.Inspired by this observation, we introduce a statistical re-weighting method that assigns different weights for the multiple responses of the same query, and trains the standard neural generation model with the weights.Experimental results on a large Chinese dialogue corpus show that our method improves the acceptance rate of generated responses compared with several baseline models and significantly reduces the number of generated generic responses. Wei Bi, Xiaojiang Liu, Jian Yao 0002, Shuming Shi 0001 |
EMNLP | 4 |
| 2018 | Translating Math Word Problem to Expression TreeabstractSequence-to-sequence (SEQ2SEQ) models have been successfully applied to automatic math word problem solving.Despite its simplicity, a drawback still remains: a math word problem can be correctly solved by more than one equations.This non-deterministic transduction harms the performance of maximum likelihood estimation.In this paper, by considering the uniqueness of expression tree, we propose an equation normalization method to normalize the duplicated equations.Moreover, we analyze the performance of three popular SEQ2SEQ models on the math word problem solving.We find that each model has its own specialty in solving problems, consequently an ensemble model is then proposed to combine their advantages.Experiments on dataset Math23K show that the ensemble model with equation normalization significantly outperforms the previous state-of-the-art methods. Lei Wang 0185, Yan Wang 0060, Deng Cai 0002, Dongxiang Zhang, Xiaojiang Liu |
EMNLP | 5 |
| 2017 | Deep Neural Solver for Math Word ProblemsabstractThis paper presents a deep neural solver to automatically solve math word problems.In contrast to previous statistical learning approaches, we directly translate math word problems to equation templates using a recurrent neural network (RNN) model, without sophisticated feature engineering.We further design a hybrid model that combines the RNN model and a similarity-based retrieval model to achieve additional performance improvement.Experiments conducted on a large dataset show that the RNN model and the hybrid model significantly outperform stateof-the-art statistical learning methods for math word problem solving. 1 We plan to make the dataset publicly available when the paper is published Yan Wang 0060, Xiaojiang Liu, Shuming Shi 0001 |
EMNLP | 2 |
| 2015 | Automatically Solving Number Word Problems by Semantic Parsing and ReasoningabstractThis paper presents a semantic parsing and reasoning approach to automatically solving math word problems.A new meaning representation language is designed to bridge natural language text and math expressions.A CFG parser is implemented based on 9,600 semi-automatically created grammar rules.We conduct experiments on a test set of over 1,500 number word problems (i.e., verbally expressed number problems) and yield 95.4% precision and 60.2% recall. Shuming Shi 0001, Yuehui Wang, Chin-Yew Lin, Xiaojiang Liu, Yong Rui |
EMNLP | 4 |
| 2010 | BioSnowball: automated population of WikisabstractInternet users regularly have the need to find biographies and facts of people of interest. Wikipedia has become the first stop for celebrity biographies and facts. However, Wikipedia can only provide information for celebrities because of its neutral point of view (NPOV) editorial policy. In this paper we propose an integrated bootstrapping framework named BioSnowball to automatically summarize the Web to generate Wikipedia-style pages for any person with a modest web presence. In BioSnowball, biography ranking and fact extraction are performed together in a single integrated training and inference process using Markov Logic Networks (MLNs) as its underlying statistical model. The bootstrapping framework starts with only a small number of seeds and iteratively finds new facts and biographies. As biography paragraphs on the Web are composed of the most important facts, our joint summarization model can improve the accuracy of both fact extraction and biography ranking compared to decoupled methods in the literature. Empirical results on both a small labeled data set and a real Web-scale data set show the effectiveness of BioSnowball. We also empirically show that BioSnowball outperforms the decoupled methods. Xiaojiang Liu, Zaiqing Nie, Nenghai Yu, Ji-Rong Wen |
KDD | 1 |
| 2010 | Novel dynamic delay allocation adjustment for improving bandwidth efficiency
Yingfei Dong, Xiaojiang Liu |
Comput. Commun. | 2 |
| 2009 | StatSnowball: a statistical approach to extracting entity relationshipsabstractTraditional relation extraction methods require pre-specified relations and relation-specific human-tagged examples. Bootstrapping systems significantly reduce the number of training examples, but they usually apply heuristic-based methods to combine a set of strict hard rules, which limit the ability to generalize and thus generate a low recall. Furthermore, existing bootstrapping methods do not perform open information extraction (Open IE), which can identify various types of relations without requiring pre-specifications. In this paper, we propose a statistical extraction framework called Statistical Snowball (StatSnowball), which is a bootstrapping system and can perform both traditional relation extraction and Open IE. Jun Zhu 0001, Zaiqing Nie, Xiaojiang Liu, Bo Zhang 0010, Ji-Rong Wen |
WWW | 3 |
| 2008 | Intelligently Balancing Per-Hop Delay Allocation to Improve Network UtilizationabstractDelay guarantees are critical to real-time applications. To ensure the end-to-end delay guarantee of a flow, the previous approaches usually statically divide the end-to-end delay requirement into per-hop delay requirements. However, such static resource allocations may cause unbalanced usage in a network, limit the total number of sessions that can be accommodated, and result in low network utilization. To address this issue, we propose a dynamic allocation adjustment (DAA) approach to evenly spread traffic over flow paths and related links, and improve the system utilization. Our simulation results show that DAA is able to significantly increase the total number of flows admitted; it reduces the average link reservations, decreases the total resource usage per flow, and balances link loads. Xiaojiang Liu, Yingfei Dong |
ICC | 1 |