VLDB 2026 Research / reviewers in the wild / expert
Chen Xing
dblp:69/4881
· DBLP profile ↗
27ranked-venue papers
7as first author
15since 2021 · last 2027
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 22 · 7 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 3 first-author · 5 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2027 | QDRF: Quantized and distilled robust fusion for incomplete data in multimodal sentiment analysis
Jiachen Hou, Zheyan Cheng, Chen Xing, Yushi Li, Shaoyu Cai, Yushan Pan |
Inf. Sci. | 3 |
| 2025 | LineDi2Vec: An Edge-Based Graph Embedding on Signed Social Networks
Chen Xing, Masoud Makrehchi |
ASONAM (2) | 1 |
| 2025 | ReGenesis: LLMs can Grow into Reasoning Generalists via Self-ImprovementabstractPost-training Large Language Models (LLMs) with explicit reasoning trajectories can enhance their reasoning abilities. However, acquiring such high-quality trajectory data typically demands meticulous supervision from humans or superior models, which can be either expensive or license-constrained. In this paper, we explore how far an LLM can improve its reasoning by self-synthesizing reasoning paths as training data without any additional supervision. Existing self-synthesizing methods, such as STaR, suffer from poor generalization to out-of-domain (OOD) reasoning tasks. We hypothesize it is due to that their self-synthesized reasoning paths are too task-specific, lacking general task-agnostic reasoning guidance. To address this, we propose **Reasoning Generalist via Self-Improvement (ReGenesis)**, a method to *self-synthesize reasoning paths as post-training data by progressing from abstract to concrete*. More specifically, ReGenesis self-synthesizes reasoning paths by converting general reasoning guidelines into task-specific ones, generating reasoning structures, and subsequently transforming these structures into reasoning paths, without the need for human-designed task-specific examples used in existing methods. We show that ReGenesis achieves superior performance on all in-domain and OOD settings tested compared to existing methods. For six OOD tasks specifically, while previous methods exhibited an average performance decrease of approximately 4.6% after post training, ReGenesis delivers around 6.1% performance improvement. We also conduct an in-depth analysis of our framework and show ReGenesis is effective across various language models and design choices. Congying Xia, Xinyi Yang 0002, Caiming Xiong, Chien-Sheng Wu, Chen Xing |
ICLR | 6 |
| 2025 | A Multi-branch Independent Masking and Dirichlet-Based Fusion Method for Multiple Instance Learning on Whole-Slide Images
Kaiwen Tan 0001, Chen Xing, Honghao Zhu, Xinxiang Fan, Jianqiu Kong |
PRCV (13) | 3 |
| 2024 | FOFO: A Benchmark to Evaluate LLMs' Format-Following CapabilityabstractCongying Xia, Chen Xing, Jiangshu Du, Xinyi Yang, Yihao Feng, Ran Xu, Wenpeng Yin, Caiming Xiong. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Congying Xia, Chen Xing, Jiangshu Du, Xinyi Yang 0002, Yihao Feng, Ran Xu 0001, Wenpeng Yin 0001, Caiming Xiong |
ACL (1) | 2 |
| 2024 | LLMs Assist NLP Researchers: Critique Paper (Meta-)ReviewingabstractJiangshu Du, Yibo Wang, Wenting Zhao, Zhongfen Deng, Shuaiqi Liu, Renze Lou, Henry Peng Zou, Pranav Narayanan Venkit, Nan Zhang, Mukund Srinath, Haoran Ranran Zhang, Vipul Gupta, Yinghui Li, Tao Li, Fei Wang, Qin Liu, Tianlin Liu, Pengzhi Gao, Congying Xia, Chen Xing, Cheng Jiayang, Zhaowei Wang, Ying Su, Raj Sanjay Shah, Ruohao Guo, Jing Gu, Haoran Li, Kangda Wei, Zihao Wang, Lu Cheng, Surangika Ranathunga, Meng Fang, Jie Fu, Fei Liu, Ruihong Huang, Eduardo Blanco, Yixin Cao, Rui Zhang, Philip S. Yu, Wenpeng Yin. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024. Jiangshu Du, Yibo Wang 0001, Wenting Zhao 0006, Zhongfen Deng, Shuaiqi Liu 0002, Renze Lou, Henry Peng Zou, Pranav Venkit, Mukund Srinath, Ranran Haoran Zhang, Tao Li 0039, Fei Wang 0060, Qin Liu 0010, Tianlin Liu, Pengzhi Gao, Congying Xia, Chen Xing, Cheng Jiayang, Zhaowei Wang 0003, Raj Sanjay Shah, Ruohao Guo, Haoran Li 0003, Kangda Wei, Zihao Wang 0001, Lu Cheng 0001, Surangika Ranathunga, Fei Liu 0004, Ruihong Huang, Eduardo Blanco 0002, Yixin Cao 0002, Rui Zhang 0037, Philip S. Yu, Wenpeng Yin 0001 |
EMNLP | 20 |
| 2024 | Lemur: Harmonizing Natural Language and Code for Language AgentsabstractWe introduce Lemur and Lemur-Chat, openly accessible language models optimized
for both natural language and coding capabilities to serve as the backbone
of versatile language agents. The evolution from language chat models to
functional language agents demands that models not only master human interaction,
reasoning, and planning but also ensure grounding in the relevant environments.
This calls for a harmonious blend of language and coding capabilities
in the models. Lemur and Lemur-Chat are proposed to address this necessity,
demonstrating balanced proficiencies in both domains, unlike existing
open-source models that tend to specialize in either. Through meticulous pretraining
using a code-intensive corpus and instruction fine-tuning on text and code
data, our models achieve state-of-the-art averaged performance across diverse
text and coding benchmarks. Comprehensive experiments demonstrate Lemur’s
superiority over existing open-source models and its proficiency across various
agent tasks involving human communication, tool usage, and interaction under
fully- and partially- observable environments. The harmonization between natural
and programming languages enables Lemur-Chat to significantly narrow the
gap with proprietary models on agent abilities, providing key insights into developing
advanced open-source agents adept at reasoning, planning, and operating
seamlessly across environments. Our model and code have been open-sourced at
https://github.com/OpenLemur/Lemur. Yiheng Xu, Hongjin Su, Chen Xing, Boyu Mi, Qian Liu 0033, Binyuan Hui, Yitao Liu, Tianbao Xie, Zhoujun Cheng, Siheng Zhao, Lingpeng Kong, Bailin Wang, Caiming Xiong, Tao Yu 0009 |
ICLR | 3 |
| 2023 | Mask-Free OVIS: Open-Vocabulary Instance Segmentation without Manual Mask AnnotationsabstractExisting instance segmentation models learn task-specific information using manual mask annotations from base (training) categories. These mask annotations require tremendous human effort, limiting the scalability to annotate novel (new) categories. To alleviate this problem, Open-Vocabulary (OV) methods leverage large-scale image-caption pairs and vision-language models to learn novel categories. In summary, an OV method learns task-specific information using strong supervision from base annotations and novel category information using weak supervision from image-captions pairs. This difference between strong and weak supervision leads to overfitting on base categories, resulting in poor generalization towards novel categories. In this work, we overcome this issue by learning both base and novel categories from pseudo-mask annotations generated by the vision-language model in a weakly supervised manner using our proposed Mask-free OVIS pipeline. Our method automatically generates pseudo-mask annotations by leveraging the localization ability of a pre-trained vision-language model for objects present in image-caption pairs. The generated pseudo-mask annotations are then used to supervise an instance segmentation model, freeing the entire pipeline from any labour-expensive instance-level annotations and overfitting. Our extensive experiments show that our method trained with just pseudo-masks significantly improves the mAP scores on the MS-COCO dataset and OpenImages dataset compared to the recent state-of-the-art methods trained with manual masks. Codes and models are provided in https://vibashan.github.io/ovis-web/. Vibashan VS, Ning Yu 0006, Chen Xing, Can Qin, Mingfei Gao, Juan Carlos Niebles, Vishal M. Patel, Ran Xu 0001 |
CVPR | 3 |
| 2023 | ULIP: Learning a Unified Representation of Language, Images, and Point Clouds for 3D UnderstandingabstractThe recognition capabilities of current state-of-the-art 3D models are limited by datasets with a small number of annotated data and a pre-defined set of categories. In its 2D counterpart, recent advances have shown that similar problems can be significantly alleviated by employing knowledge from other modalities, such as language. Inspired by this, leveraging multimodal information for 3D modality could be promising to improve 3D understanding under the restricted data regime, but this line of research is not well studied. Therefore, we introduce ULIP to learn a unified representation of image, text, and 3D point cloud by pre-training with object triplets from the three modalities. To overcome the shortage of training triplets, ULIP leverages a pre-trained vision-language model that has already learned a common visual and textual space by training with massive image-text pairs. Then, ULIP learns a 3D representation space aligned with the common image-text space, using a small number of automatically synthesized triplets. ULIP is agnostic to 3D backbone networks and can easily be integrated into any 3D architecture. Experiments show that ULIP effectively improves the performance of multiple recent 3D backbones by simply pre-training them on ShapeNet55 using our framework, achieving state-of-the-art performance in both standard 3D classification and zero-shot 3D classification on ModelNet40 and ScanObjectNN. ULIP also improves the performance of PointMLP by around 3% in 3D classification on ScanObjectNN, and outperforms PointCLIP by 28.8% on top-1 accuracy for zero-shot 3D classification on ModelNet40. Our code and pre-trained models will be released. Le Xue, Mingfei Gao, Chen Xing, Roberto Martin Martin, Jiajun Wu 0001, Caiming Xiong, Ran Xu 0001, Juan Carlos Niebles, Silvio Savarese |
CVPR | 3 |
| 2023 | GlueGen: Plug and Play Multi-modal Encoders for X-to-image GenerationabstractText-to-image (T2I) models based on diffusion processes have achieved remarkable success in controllable image generation using user-provided captions. However, the tight coupling between the current text encoder and image decoder in T2I models makes it challenging to replace or upgrade. Such changes often require massive fine-tuning or even training from scratch with the prohibitive expense. To address this problem, we propose GlueGen, which applies a newly proposed GlueNet model to align features from single-modal or multi-modal encoders with the latent space of an existing T2I model. The approach introduces a new training objective that leverages parallel corpora to align the representation spaces of different encoders. Empirical results show that GlueNet can be trained efficiently and enables various capabilities beyond previous state-of-the-art models: 1) multilingual language models such as XLM-Roberta can be aligned with existing T2I models, allowing for the generation of high-quality images from captions beyond English; 2) GlueNet can align multi-modal encoders such as AudioCLIP with the Stable Diffusion model, enabling sound-to-image generation; 3) it can also upgrade the current text encoder of the latent diffusion model for challenging case generation. By the alignment of various feature representations, the GlueNet allows for flexible and efficient integration of new functionality into existing T2I models and sheds light on X-to-image (X2I) generation.1 Can Qin, Ning Yu 0006, Chen Xing, Shu Zhang 0007, Zeyuan Chen 0001, Stefano Ermon, Yun Fu 0001, Caiming Xiong, Ran Xu 0001 |
ICCV | 3 |
| 2023 | A Nested Differential Evolution Algorithm for Optimal Designs of Quantile Regression Models
Zhenyang Xia, Chen Xing |
ICIC (1) | 2 |
| 2023 | Model ensemble instead of prompt fusion: a sample-specific knowledge transfer method for few-shot prompt tuning
Chen Xing, Prafulla Kumar Choubey, Chien-Sheng Wu, Caiming Xiong |
ICLR | 2 |
| 2022 | DocQueryNet: Value Retrieval with Arbitrary Queries for Form-like DocumentsabstractWe propose, DocQueryNet, a value retrieval method with arbitrary queries for form-like documents to reduce human effort of processing forms. Unlike previous methods that only address a fixed set of field items, our method predicts target value for an arbitrary query based on the understanding of the layout and semantics of a form. To further boost model performance, we propose a simple document language modeling (SimpleDLM) strategy to improve document understanding on large-scale model pre-training. Experimental results show that DocQueryNet outperforms previous designs significantly and the SimpleDLM further improves our performance on value retrieval by around 17% F1 score compared with the state-of-the-art pre-training method. Code is available here, https://github.com/salesforce/QVR-SimpleDLM. Mingfei Gao, Le Xue, Chetan Ramaiah, Chen Xing, Ran Xu 0001, Caiming Xiong |
COLING | 4 |
| 2022 | Open Vocabulary Object Detection with Pseudo Bounding-Box Labels
Mingfei Gao, Chen Xing, Juan Carlos Niebles, Junnan Li 0001, Ran Xu 0001, Wenhao Liu 0003, Caiming Xiong |
ECCV (10) | 2 |
| 2021 | Taking Notes on the Fly Helps Language Pre-Training
Qiyu Wu 0001, Chen Xing, Yatao Li, Guolin Ke, Di He 0001, Tie-Yan Liu |
ICLR | 2 |
| 2020 | Distance-Based Learning from Errors for Confidence Calibration
Chen Xing, Sercan Ö. Arik, Tomas Pfister |
ICLR | 1 |
| 2020 | On Layer Normalization in the Transformer ArchitectureabstractThe Transformer is widely used in natural language processing tasks. To train a Transformer however, one usually needs a carefully designed learning rate warm-up stage, which is shown to be crucial to the final performance but will slow down the optimization and bring more hyper-parameter tunings. In this paper, we first study theoretically why the learning rate warm-up stage is essential and show that the location of layer normalization matters. Specifically, we prove with mean field theory that at initialization, for the original-designed Post-LN Transformer, which places the layer normalization between the residual blocks, the expected gradients of the parameters near the output layer are large. Therefore, using a large learning rate on those gradients makes the training unstable. The warm-up stage is practically helpful for avoiding this problem. On the other hand, our theory also shows that if the layer normalization is put inside the residual blocks (recently proposed as Pre-LN Transformer), the gradients are well-behaved at initialization. This motivates us to remove the warm-up stage for the training of Pre-LN Transformers. We show in our experiments that Pre-LN Transformers without the warm-up stage can reach comparable results with baselines while requiring significantly less training time and hyper-parameter tuning on a wide range of applications. Ruibin Xiong, Yunchang Yang, Di He 0001, Kai Zheng 0007, Shuxin Zheng, Chen Xing, Huishuai Zhang, Yanyan Lan, Liwei Wang 0001, Tie-Yan Liu |
ICML | 6 |
| 2019 | Adaptive Cross-Modal Few-shot LearningabstractMetric-based meta-learning techniques have successfully been applied to few-shot classification problems. In this paper, we propose to leverage cross-modal information to enhance metric-based few-shot learning methods. Visual and semantic feature spaces have different structures by definition. For certain concepts, visual features might be richer and more discriminative than text ones. While for others, the inverse might be true. Moreover, when the support from visual information is limited in image classification, semantic representations (learned from unsupervised text corpora) can provide strong prior knowledge and context to help learning. Based on these two intuitions, we propose a mechanism that can adaptively combine information from both modalities according to new image categories to be learned. Through a series of experiments, we show that by this adaptive combination of the two modalities, our model outperforms current uni-modality few-shot learning methods and modality-alignment methods by a large margin on all benchmarks and few-shot scenarios tested. Experiments also show that our model can effectively adjust its focus on the two modalities. The improvement in performance is particularly large when the number of shots is very small. Chen Xing, Negar Rostamzadeh, Boris N. Oreshkin, Pedro O. Pinheiro |
NeurIPS | 1 |
| 2019 | A Sequential Matching Framework for Multi-Turn Response Selection in Retrieval-Based ChatbotsabstractWe study the problem of response selection for multi-turn conversation in retrieval-based chatbots. The task involves matching a response candidate with a conversation context, the challenges for which include how to recognize important parts of the context, and how to model the relationships among utterances in the context. Existing matching methods may lose important information in contexts as we can interpret them with a unified framework in which contexts are transformed to fixed-length vectors without any interaction with responses before matching. This motivates us to propose a new matching framework that can sufficiently carry important information in contexts to matching and model relationships among utterances at the same time. The new framework, which we call a sequential matching framework (SMF), lets each utterance in a context interact with a response candidate at the first step and transforms the pair to a matching vector. The matching vectors are then accumulated following the order of the utterances in the context with a recurrent neural network (RNN) that models relationships among utterances. Context-response matching is then calculated with the hidden states of the RNN. Under SMF, we propose a sequential convolutional network and sequential attention network and conduct experiments on two public data sets to test their performance. Experiment results show that both models can significantly outperform state-of-the-art matching methods. We also show that the models are interpretable with visualizations that provide us insights on how they capture and leverage important information in contexts for matching. Yu Wu 0012, Wei Wu 0014, Chen Xing, Can Xu 0002, Zhoujun Li 0001, Ming Zhou 0001 |
Comput. Linguistics | 3 |
| 2018 | Hierarchical Recurrent Attention Network for Response GenerationabstractWe study multi-turn response generation in chatbots where a response is generated according to a conversation context. Existing work has modeled the hierarchy of the context, but does not pay enough attention to the fact that words and utterances in the context are differentially important. As a result, they may lose important information in context and generate irrelevant responses. We propose a hierarchical recurrent attention network (HRAN) to model both the hierarchy and the importance variance in a unified framework. In HRAN, a hierarchical attention mechanism attends to important parts within and among utterances with word level attention and utterance level attention respectively. Chen Xing, Yu Wu 0012, Wei Wu 0014, Yalou Huang, Ming Zhou 0001 |
AAAI | 1 |
| 2017 | Topic Aware Neural Response GenerationabstractWe consider incorporating topic information into a sequence-to-sequence framework to generate informative and interesting responses for chatbots. To this end, we propose a topic aware sequence-to-sequence (TA-Seq2Seq) model. The model utilizes topics to simulate prior human knowledge that guides them to form informative and interesting responses in conversation, and leverages topic information in generation by a joint attention mechanism and a biased generation probability. The joint attention mechanism summarizes the hidden vectors of an input message as context vectors by message attention and synthesizes topic vectors by topic attention from the topic words of the message obtained from a pre-trained LDA model, with these vectors jointly affecting the generation of words in decoding. To increase the possibility of topic words appearing in responses, the model modifies the generation probability of topic words by adding an extra probability item to bias the overall distribution. Empirical studies on both automatic evaluation metrics and human annotations show that TA-Seq2Seq can generate more informative and interesting responses, significantly outperforming state-of-the-art response generation models. Chen Xing, Wei Wu 0014, Yu Wu 0012, Jie Liu 0007, Yalou Huang, Ming Zhou 0001, Wei-Ying Ma |
AAAI | 1 |
| 2017 | Sequential Matching Network: A New Architecture for Multi-turn Response Selection in Retrieval-Based ChatbotsabstractWe study response selection for multiturn conversation in retrieval-based chatbots.Existing work either concatenates utterances in context or matches a response with a highly abstract context vector finally, which may lose relationships among utterances or important contextual information.We propose a sequential matching network (SMN) to address both problems.SMN first matches a response with each utterance in the context on multiple levels of granularity, and distills important matching information from each pair as a vector with convolution and pooling operations.The vectors are then accumulated in a chronological order through a recurrent neural network (RNN) which models relationships among utterances.The final matching score is calculated with the hidden states of the RNN.An empirical study on two public data sets shows that SMN can significantly outperform stateof-the-art methods for response selection in multi-turn conversation. Yu Wu 0012, Wei Wu 0014, Chen Xing, Ming Zhou 0001, Zhoujun Li 0001 |
ACL (1) | 3 |
| 2016 | Hashtag-Based Sub-Event Discovery Using Mutually Generative LDA in TwitterabstractSub-event discovery is an effective method for social event analysis in Twitter. It can discover sub-events from large amount of noisy event-related information in Twitter and semantically represent them. The task is challenging because tweets are short, informal and noisy. To solve this problem, we consider leveraging event-related hashtags that contain many locations, dates and concise sub-event related descriptions to enhance sub-event discovery. To this end, we propose a hashtag-based mutually generative Latent Dirichlet Allocation model(MGe-LDA). In MGe-LDA, hashtags and topics of a tweet are mutually generated by each other. The mutually generative process models the relationship between hashtags and topics of tweets, and highlights the role of hashtags as a semantic representation of the corresponding tweets. Experimental results show that MGe-LDA can significantly outperform state-of-the-art methods for sub-event discovery. Chen Xing, Jie Liu 0007, Yalou Huang, Wei-Ying Ma |
AAAI | 1 |
| 2016 | Detecting Context Dependent Messages in a Conversational EnvironmentabstractWhile automatic response generation for building chatbot systems has drawn a lot of attention recently, there is limited understanding on when we need to consider the linguistic context of an input text in the generation process. The task is challenging, as messages in a conversational environment are short and informal, and evidence that can indicate a message is context dependent is scarce. After a study of social conversation data crawled from the web, we observed that some characteristics estimated from the responses of messages are discriminative for identifying context dependent messages. With the characteristics as weak supervision, we propose using a Long Short Term Memory (LSTM) network to learn a classifier. Our method carries out text representation and classifier learning in a unified framework. Experimental results show that the proposed method can significantly outperform baseline methods on accuracy of classification. Chaozhuo Li, Yu Wu 0012, Wei Wu 0014, Chen Xing, Zhoujun Li 0001, Ming Zhou 0001 |
COLING | 4 |
| 2013 | Feature based image registration using non-degenerate pixels
Peihua Qiu, Chen Xing |
Signal Process. | 2 |
| 2011 | Intensity-Based Image Registration by Nonparametric Local SmoothingabstractImage registration is used widely in applications for mapping one image to another. Existing image registration methods are either feature-based or intensity-based. Feature-based methods first extract relevant image features and then find the geometrical transformation that best matches the two corresponding sets of features extracted from the two images. Because identification and extraction of image features is often a challenging and time-consuming process, intensity-based image registration, by which the mapping transformation is estimated directly from the observed image intensities of the two images, has received much attention recently. In the literature, most existing intensity-based image registration methods estimate the mapping transformation globally by solving a minimization/maximization problem defined by the two entire images to register. To this end, it needs to be assumed that the mapping transformation has a certain type of parametric form or it is a continuous bivariate function satisfying certain regularity conditions. In this paper, we propose a novel intensity-based image registration method using nonparametric local smoothing. By this method, the mapping transformation at a given pixel is estimated locally in a neighborhood after certain image features are accommodated in the estimation. Due to the flexibility of local smoothing, this method does not require any parametric form for the mapping transformation. It even allows the transformation to be a discontinuous function. Numerical examples show that it is effective in various applications. Chen Xing, Peihua Qiu |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2010 | The Study and Realization of Virtual Organization File System Based on DHT Technology
Jiqing Liu, Jinhua Huang, Chen Xing |
ICIC (1) | 3 |