EDBT 2026 Demo / reviewers in the wild / expert
Zhenxin Fu
dblp:209/8408
· DBLP profile ↗
17ranked-venue papers
6as first author
4since 2021 · last 2023
0009-0009-9830-913XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 11 · 4 first-author · 1 since 2021Systems, architecture and hardware · 5 · 1 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 1 first-authorDatabases, data management, data science and information retrieval · 4 · 4 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | Critique of "A Parallel Framework for Constraint-Based Bayesian Network Learning via Markov Blanket Discovery" by SCC Team From Peking UniversityabstractAnkit Srivastava et al. (Srivastava et al. 2020) proposed a parallel framework for Constraint-Based Bayesian Network (BN) Learning via Markov Blanket Discovery (referred to as ramBLe) and implemented it over three existing BN learning algorithms, namely, GS, IAMB and Inter-IAMB. As part of the Student Cluster Competition at SC21, we reproduce the computational efficiency of ramBLe on our assigned Oracle cluster. The cluster has 4x36 cores in total with 100 Gbps RoCE v2 support and is equipped with CentOS-compatible Oracle Linux. Our experiments, covering the same three algorithms of the original ramBLe article (Srivastava et al. 2020), evaluate the strong and weak scalability of the algorithms using real COVID-19 data sets. We verify part of the conclusions from the original article and propose our explanation of the differences obtained in our results.Author: Please confirm or add details for any funding or financial support for the research of this article. ?> Jiaqi Si, Junyi Guo, Zhewen Hao, Wenyang He, Yueyang Pan, Zhenxin Fu, Chun Fan 0001 |
IEEE Trans. Parallel Distributed Syst. | 7 |
| 2022 | ProphetChat: Enhancing Dialogue Generation with Simulation of Future ConversationabstractTypical generative dialogue models utilize the dialogue history to generate the response.However, since one dialogue utterance can often be appropriately answered by multiple distinct responses, generating a desired response solely based on the historical information is not easy.Intuitively, if the chatbot can foresee in advance what the user would talk about (i.e., the dialogue future) after receiving its response, it could possibly provide a more informative response.Accordingly, we propose a novel dialogue generation framework named ProphetChat that utilizes the simulated dialogue futures in the inference phase to enhance response generation.To enable the chatbot to foresee the dialogue future, we design a beam-search-like roll-out strategy for dialogue future simulation using a typical dialogue generation model and a dialogue selector.With the simulated futures, we then utilize the ensemble of a history-to-response generator and a future-to-response generator to jointly generate a more informative response.Experiments on two popular open-domain dialogue datasets demonstrate that ProphetChat can generate better responses over strong baselines, which validates the advantages of incorporating the simulated dialogue futures. Chang Liu 0076, Xu Tan 0003, Chongyang Tao, Zhenxin Fu, Dongyan Zhao 0001, Tie-Yan Liu, Rui Yan 0001 |
ACL (1) | 4 |
| 2022 | Critique of "MemXCT: Memory-Centric X-Ray CT Reconstruction With Massive Parallelization" by SCC Team From Peking UniversityabstractHidayetoluet al.(2019) proposed a novel memory-centric computation system, MemXCT. As a challenge at SC20, we reproduce the computational efficiency of MemXCT on our Azure cloud cluster. Our experiments evaluate the overall performance and the strong scalability with real datasets and verify part of the conclusions in the original article. Zejia Fan, Zhewen Hao, Yueyang Pan, Pengcheng Xu 0005, Yuxuan Yan, Fangyuan Yang, Zhenxin Fu, Yun Liang 0001 |
IEEE Trans. Parallel Distributed Syst. | 8 |
| 2021 | Critique of "Planetary Normal Mode Computation: Parallel Algorithms, Performance, and Reproducibility" by SCC Team From Peking UniversityabstractShi et al. (2018) proposed a highly parallel polynomial filtering eigensolver for the computation of planetary normal modes. As a challenge at the Student Cluster Competition in The International Conference for High Performance Computing, Networking, Storage and Analysis (SC19), we reproduce the computational efficiency of the polynomial filtering eigensolver on our Intel Xeon machine. We present the weak scalability, scaling of runtime with model size (in a fixed interval) and the strong scalability results in this report. Yihua Cheng, Zejia Fan, Jing Mai, Yifan Wu 0005, Pengcheng Xu 0005, Yuxuan Yan, Zhenxin Fu, Yun Liang 0001 |
IEEE Trans. Parallel Distributed Syst. | 7 |
| 2020 | Learning Sense Representation from Word Representation for Unsupervised Word Sense Disambiguation (Student Abstract)abstractUnsupervised WSD methods do not rely on annotated training datasets and can use WordNet. Since each ambiguous word in the WSD task exists in WordNet and each sense of the word has a gloss, we propose SGM and MGM to learn sense representations for words in WordNet using the glosses. In the WSD task, we calculate the similarity between each sense of the ambiguous word and its context to select the sense with the highest similarity. We evaluate our method on several benchmark WSD datasets and achieve better performance than the state-of-the-art unsupervised WSD systems. Zhenxin Fu, Moxin Li, Haisong Zhang, Dongyan Zhao 0001, Rui Yan 0001 |
AAAI | 2 |
| 2020 | Query-to-Session Matching: Do NOT Forget History and Future during Response Selection for Multi-Turn Dialogue SystemsabstractGiven a user query, traditional multi-turn retrieval-based dialogue systems first retrieve a set of candidate responses from the historical dialogue sessions. Then the response selection models select the most appropriate response to the given query. However, previous work only considers the matching between the query and the response but ignores the informative dialogue session in which the response is located. Nevertheless, this session, composed of the response, the response's history and the response's future, always contains valuable contextual information which can help the response selection task. More specifically, if the current query and a response's history both refer to the same question, we can conclude that this response is quite likely to answer this query. As for the response's future, it can always provide contextual hints and supplementary information that might be omitted in the response. Inspired by such motivation, we propose a query-to-session matching (QSM) framework to make full use of the session information: matching the query with the candidate session instead of the response only. Different from the previous work which ranks response directly, the response in the session with the highest query-to-session matching score will be selected as the desired response. In our proposed framework, the query, history, and future are all sequences of utterances, which makes it necessary to model the relationships among the utterances. So we propose a novel dialogue flow aware query-to-session matching (DF-QSM) model. The dialogue flows model the relationships among the utterances through a memory network. To our best knowledge, our paper is the first work to utilize both the response's history and future in the response selection task. The experimental results on three multi-turn response selection benchmarks show that our proposed model outperforms existing state-of-the-art methods by a large margin. Zhenxin Fu, Shaobo Cui 0001, Ji Zhang 0011, Haiqing Chen, Dongyan Zhao 0001, Rui Yan 0001 |
CIKM | 1 |
| 2020 | Context-to-Session Matching: Utilizing Whole Session for Response Selection in Information-Seeking Dialogue SystemsabstractWe study the retrieval-based multi-turn information-seeking dialogue systems, which are widely used in many scenarios. Most of the previous works select the response according to the matching degree between the query's context and the candidate responses. Though great progress has been made, existing works ignore the contexts of the responses, which could provide rich information for selecting the most appropriate response. The more similar the query's context and certain response's context are, the more likely they are to indicate the same question, and thus, the more likely this response is to answer the query. In this paper, we consider the response and its context as a whole session and explore the task of matching the query's context with the sessions. More specifically, we propose to match between the query's context and response's context and integrate the context-to-context matching with context-to-response matching. Experiment results prove that our proposed context-to-session method outperforms the strong baselines significantly. Zhenxin Fu, Shaobo Cui 0001, Mingyue Shang, Dongyan Zhao 0001, Haiqing Chen, Rui Yan 0001 |
KDD | 1 |
| 2020 | Be Aware of the Hot Zone: A Warning System of Hazard Area Prediction to Intervene Novel Coronavirus COVID-19 OutbreakabstractDating back from late December 2019, the Chinese city of Wuhan has reported an outbreak of atypical pneumonia, now known as lung inflammation caused by novel coronavirus (COVID-19). Cases have spread to other cities in China and more than 180 countries and regions internationally. World Health Organization (WHO) officially declares the coronavirus outbreak a pandemic and the public health emergency is perhaps one of the top concerns in the year of 2020 for governments all over the world. Till today, the coronavirus outbreak is still raging and has no sign of being under control in many countries. In this paper, we aim at drawing lessons from the COVID-19 outbreak process in China and using the experiences to help the interventions against the coronavirus wherever in need. To this end, we have built a system predicting hazard areas on the basis of confirmed infection cases with location information. The purpose is to warn people to avoid of such hot zones and reduce risks of disease transmission through droplets or contacts. We analyze the data from the daily official information release which are publicly accessible. Based on standard classification frameworks with reinforcements incrementally learned day after day, we manage to conduct thorough feature engineering from empirical studies, including geographical, demographic, temporal, statistical, and epidemiological features. Compared with heuristics baselines, our method has achieved promising overall performance in terms of precision, recall, accuracy, F1 score, and AUC. We expect that our efforts could be of help in the battle against the virus, the common opponent of human kind. Zhenxin Fu, Yu Wu 0024, Hailei Zhang, Yichuan Hu, Dongyan Zhao 0001, Rui Yan 0001 |
SIGIR | 1 |
| 2019 | Find a Reasonable Ending for Stories: Does Logic Relation Help the Story Cloze Test?abstractNatural language understanding is a challenging problem that covers a wide range of tasks. While previous methods generally train each task separately, we consider combining the cross-task features to enhance the task performance. In this paper, we incorporate the logic information with the help of the Natural Language Inference (NLI) task to the Story Cloze Test (SCT). Previous work on SCT considered various semantic information, such as sentiment and topic, but lack the logic information between sentences which is an essential element of stories. Thus we propose to extract the logic information during the course of the story to improve the understanding of the whole story. The logic information is modeled with the help of the NLI task. Experimental results prove the strength of the logic information. Mingyue Shang, Zhenxin Fu, Hongzhi Yin, Bo Tang 0016, Dongyan Zhao 0001, Rui Yan 0001 |
AAAI | 2 |
| 2019 | Query-bag Matching with Mutual Coverage for Information-seeking Conversations in E-commerceabstractInformation-seeking conversation system aims at satisfying the information needs of users through conversations. Text matching between a user query and a pre-collected question is an important part of the information-seeking conversation in E-commerce. In the practical scenario, a sort of questions always correspond to a same answer. Naturally, these questions can form a bag. Learning the matching between user query and bag directly may improve the conversation performance, denoted as query-bag matching. Inspired by such opinion, we propose a query-bag matching model which mainly utilizes the mutual coverage between query and bag and measures the degree of the content in the query mentioned by the bag, and vice verse. In addition, the learned bag representation in word level helps find the main points of a bag in a fine grade and promotes the query-bag matching performance. Experiments on two datasets show the effectiveness of our model. Zhenxin Fu, Wenpeng Hu, Dongyan Zhao 0001, Haiqing Chen, Rui Yan 0001 |
CIKM | 1 |
| 2019 | Semi-supervised Text Style Transfer: Cross Projection in Latent SpaceabstractMingyue Shang, Piji Li, Zhenxin Fu, Lidong Bing, Dongyan Zhao, Shuming Shi, Rui Yan. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Mingyue Shang, Piji Li, Zhenxin Fu, Lidong Bing, Dongyan Zhao 0001, Shuming Shi 0001, Rui Yan 0001 |
EMNLP/IJCNLP (1) | 3 |
| 2019 | Multilingual Dialogue Generation with Shared-Private Memory
Lisong Qiu, Zhenxin Fu, Rui Yan 0001 |
NLPCC (1) | 3 |
| 2018 | Style Transfer in Text: Exploration and EvaluationabstractThe ability to transfer styles of texts or images, is an important measurement of the advancement of artificial intelligence (AI). However, the progress in language style transfer is lagged behind other domains, such as computer vision, mainly because of the lack of parallel data and reliable evaluation metrics. In response to the challenge of lacking parallel data, we explore learning style transfer from non-parallel data. We propose two models to achieve this goal. The key idea behind the proposed models is to learn separate content representations and style representations using adversarial networks. Considering the problem of lacking principle evaluation metrics, we propose two novel evaluation metrics that measure two aspects of style transfer: transfer strength and content preservation. We benchmark our models and the evaluation metrics on two style transfer tasks: paper-news title transfer, and positive-negative review transfer. Results show that the proposed content preservation metric is highly correlate to human judgments, and the proposed models are able to generate sentences with similar content preservation score but higher style transfer strength comparing to auto-encoder. Zhenxin Fu, Xiaoye Tan, Nanyun Peng 0001, Dongyan Zhao 0001, Rui Yan 0001 |
AAAI | 1 |
| 2018 | Learning to Converse with Noisy Data: Generation with CalibrationabstractThe availability of abundant conversational data on the Internet brought prosperity to the generation-based open domain conversation systems. In the training of the generation models, existing methods generally treat all the training data equivalently. However, the data crawled from the websites may contain many noises. Blindly training with the noisy data could harm the performance of the final generation model. In this paper, we propose a generation with calibration framework, that allows high- quality data to have more influences on the generation model and reduces the effect of noisy data. Specifically, for each instance in training set, we employ a calibration network to produce a quality score for it, then the score is used for the weighted update of the generation model parameters. Experiments show that the calibrated model outperforms baseline methods on both automatic evaluation metrics and human annotations. Mingyue Shang, Zhenxin Fu, Nanyun Peng 0001, Yansong Feng 0002, Dongyan Zhao 0001, Rui Yan 0001 |
IJCAI | 2 |
| 2018 | One "Ruler" for All Languages: Multi-Lingual Dialogue Evaluation with Adversarial Multi-Task LearningabstractAutomatic evaluating the performance of Open-domain dialogue system is a challenging problem. Recent work in neural network-based metrics has shown promising opportunities for automatic dialogue evaluation. However, existing methods mainly focus on monolingual evaluation, in which the trained metric is not flexible enough to transfer across different languages. To address this issue, we propose an adversarial multi-task neural metric (ADVMT) for multi-lingual dialogue evaluation, with shared feature extraction across languages. We evaluate the proposed model in two different languages. Experiments show that the adversarial multi-task neural metric achieves a high correlation with human annotation, which yields better performance than monolingual ones and various existing metrics. Xiaowei Tong, Zhenxin Fu, Mingyue Shang, Dongyan Zhao 0001, Rui Yan 0001 |
IJCAI | 2 |
| 2018 | Student Cluster Competition 2017, Team Peking University: Reproducing vectorization of the Tersoff multi-body potential on the Intel Broadwell architecture
Zhenxin Fu, Lei Yang 0031, Wenbin Hou, Yifan Wu 0005, Yihua Cheng, Yun Liang 0001 |
Parallel Comput. | 1 |
| 2017 | ParConnect reproducibility report
Lei Yang 0031, Zhenxin Fu, Wenbin Hou, Haoze Wu 0002, Yun Liang 0001 |
Parallel Comput. | 3 |