VLDB 2026 Research / reviewers in the wild / expert
Chen Qian 0003
dblp:70/3604-3
· DBLP profile ↗
22ranked-venue papers
12as first author
9since 2021 · last 2025
0000-0003-0155-8009ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 7 · 5 first-author · 4 since 2021Databases, data management, data science and information retrieval · 6 · 4 first-author · 2 since 2021Computer networks · 5 · 4 first-authorSoftware engineering, systems software and programming languages · 5 · 2 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | The Tug of War Within: Mitigating the Fairness-Privacy Conflicts in Large Language ModelsabstractEnsuring awareness of fairness and privacy in Large Language Models (LLMs) is critical.Interestingly, we discover a counterintuitive trade-off phenomenon that enhancing an LLM's privacy awareness through Supervised Fine-Tuning (SFT) methods significantly decreases its fairness awareness with thousands of samples.To address this issue, inspired by the information theory, we introduce a training-free method to Suppress the Privacy and faIrness coupled Neurons (SPIN), which theoretically and empirically decrease the mutual information between fairness and privacy awareness.Extensive experimental results demonstrate that SPIN eliminates the tradeoff phenomenon and significantly improves LLMs' fairness and privacy awareness simultaneously without compromising general capabilities, e.g., improving Qwen-2-7B-Instruct's fairness awareness by 12.2% and privacy awareness by 14.0%.More crucially, SPIN remains robust and effective with limited annotated data or even when only malicious fine-tuning data is available, whereas SFT methods may fail to perform properly in such scenarios.Furthermore, we show that SPIN could generalize to other potential trade-off dimensions.We hope this study provides valuable insights into concurrently addressing fairness and privacy concerns in LLMs and can be integrated into comprehensive frameworks to develop more ethical and responsible AI systems.Our code is available at https://github.com/ChnQ/SPIN. Chen Qian 0003, Dongrui Liu, Jie Zhang 0121, Yong Liu 0018 |
ACL (1) | 1 |
| 2025 | Multi-View Trace Clustering Based on Graph Convolutional Networks in Process MiningabstractProcess mining techniques can extract process models from event logs produced by information systems. However, in flexible environments, simply using existing methods often leads to complex process models that are hard to understand, due to the less structured and greater complexity of processes in real life. Trace clustering is a pre-processing technique that enhances the effectiveness of model mining by partitioning similar behaviors in logs. In this paper, we present a Multi-view Trace Clustering method, named MTC, that improves the homogeneity of trace subclusters. Our method consists of three parts: (1) We use trace profiles to depict the traces from different views, and each profile would be transformed into a graph based on the k-nearest neighbor algorithm; (2) A fusion graph is designed to capture the information among these graphs based on an attention coefficient matrix, and then the graph convolutional networks are used to encode all graphs for obtaining the common representation; (3) We also enhance the characterization of the common representation with an inner decoder. Finally, we adopt k-means to cluster the traces in the log based on the common representation. Extensive experiments using multiple datasets illustrate that MTC significantly surpasses state-of-the-art trace clustering methods Leilei Lin, Yunuo Cao, Zan Zong, Chen Qian 0003, Lijie Wen 0001 |
IEEE Trans. Serv. Comput. | 5 |
| 2025 | MAO: A Framework for Process Model Generation With Multi-Agent OrchestrationabstractProcess models are frequently used in software engineering to describe business requirements, guide software testing and control system improvement. However, traditional process modeling methods often require the participation of numerous experts, which is expensive and time-consuming. Therefore, the exploration of a more efficient and cost-effective automated modeling method has emerged as a focal point in current research. This article explores a framework for automatically generating process models with multi-agent orchestration (MAO), aiming to enhance the efficiency of process modeling and offer valuable insights for domain experts. Our framework MAO leverages large language models as the cornerstone for multi-agent, employing an innovative prompt strategy to ensure efficient collaboration among multi-agent. Specifically, 1) Generation: The first phase of MAO is to generate a slightly rough process model from the text description; 2) Refinement: The agents would continuously refine the initial process model through multiple rounds of dialogue; 3) Reviewing: Large language models are prone to hallucination phenomena among multi-turn dialogues, so the agents need to review and repair semantic hallucinations in process models; 4) Modifying: The agents utilize external tools to test whether the generated process model contains format errors, and then adjust the process model to conform to the output paradigm. The experiments demonstrate that the process models generated by our framework outperform existing methods and surpass manual modeling by 89%, 61%, 52% and 75% on four different processes, respectively. Leilei Lin, Yumeng Jin, Yingming Zhou, Chen Qian 0003 |
IEEE Trans. Serv. Comput. | 5 |
| 2024 | Self-Improving Teacher Cultivates Better Student: Distillation Calibration for Multimodal Large Language ModelsabstractMultimodal content generation, which leverages visual information to enhance the comprehension of cross-modal understanding, plays a critical role in Multimodal Information Retrieval. With the development of large language models (LLMs), recent research has adopted visual instruction tuning to inject the knowledge of LLMs into downstream multimodal tasks. The high complexity and great demand for resources urge researchers to study efficient distillation solutions to transfer the knowledge from pre-trained multimodal models.(teachers) to more compact student models. However, the instruction tuning for knowledge distillation in multimodal LLMs is resource-intensive and capability-restricted. The comprehension of students is highly reliant on the teacher models. To address this issue, we propose a novel Multimodal Distillation Calibration framework (MmDC). The main idea is to generate high-quality training instances that challenge student models to comprehend and prompt the teacher to calibrate the knowledge transferred to students, ultimately cultivating a better student model in downstream tasks. This framework comprises two stages: (1) multimodal alignment and (2) knowledge distillation calibration. In the first stage, parameter-efficient fine-tuning is used to enhance feature alignment between different modalities. In the second stage, we develop a calibration strategy to assess the student model's capability and generate high-quality instances to calibrate knowledge distillation from teacher to student. The experiments on diverse datasets show that our framework efficiently improves the student model's capabilities. Our 7B-size student model, after three iterations of distillation calibration, outperforms the current state-of-the-art LLaVA-13B model on the ScienceQA and LLaVA Test datasets and also exceeds other strong baselines in a zero-shot setting. Li Lin 0011, Shuai Wang 0008, Chen Qian 0003 |
SIGIR | 4 |
| 2022 | MespaConfig: Memory-Sparing Configuration Auto-Tuning for Co-Located In-Memory Cluster Computing JobsabstractDistributed in-memory computing frameworks usually have lots of parameters (e.g., the buffer size of shuffle) to form a configuration for each execution. A well-tuned configuration can bring large improvements of performance. However, to improve resource utilization, jobs are often share the same cluster, which causes dynamic cluster load conditions. According to our observation, the variation of cluster load reduces effectiveness of configuration tuning. Besides, as a common problem of cluster computing jobs, overestimation of resources also occurs during configuration tuning. It is challenging to efficiently find the optimal configuration in a shared cluster with the consideration of memory-sparing. In this article, we introduce MespaConfig, a job-level configuration optimizer for distributed in-memory computing jobs. Advancements of MespaConfig over previous work are features including memory-sparing and load-sensitive. We evaluate MespaConfig by 6 typical Spark programs under different load conditions. The evaluation results show that MespaConfig improves the performance of six typical programs by up to 12× compared with default configurations. MespaConfig also achieves at most 41 percent reduction of configuration memory usage and reduces the optimization time overhead by 10.8× compared with the state-of-the-art approach. Zan Zong, Lijie Wen 0001, Xuming Hu, Rui Han 0001, Chen Qian 0003, Li Lin 0011 |
IEEE Trans. Serv. Comput. | 5 |
| 2021 | Knowledge-aware Named Entity Recognition with Alleviating HeterogeneityabstractNamed Entity Recognition (NER) is a fundamental and important research topic for many downstream NLP tasks, aiming at detecting and classifying named entities (NEs) mentioned in unstructured text into pre-defined categories. Learning from labeled data only is far from enough when it comes to domain-specific or temporally-evolving entities (medical terminologies or restaurant names). Luckily, open-source Knowledge Bases (KBs) (Wikidata and Freebase) contain NEs that are manually labeled with predefined types in different domains, which is potentially beneficial to identify entity boundaries and recognize entity types more accurately. However, the type system of a domain-specific NER task is typically independent of that of current KBs and thus exhibits heterogeneity issue inevitably, which makes matching between the original NER and KB types (Person in NER potentially matches President in KBs) less likely, or introduces unintended noises without considering domain-specific knowledge (Band in NER should be mapped to Out_of_Entity_Types in the restaurant-related task). To better incorporate and denoise the abundant knowledge in KBs, we propose a new KB-aware NER framework (KaNa), which utilizes type-heterogeneous knowledge to improve NER. Specifically, for an entity mention along with a set of candidate entities that are linked from KBs, KaNa first uses a type projection mechanism that maps the mention type and entity types into a shared space to homogenize the heterogeneous entity types. Then, based on projected types, a noise detector filters out certain less-confident candidate entities in an unsupervised manner. Finally, the filtered mention-entity pairs are injected into a NER model as a graph to predict answers. The experimental results demonstrate KaNa's state-of-the-art performance on five public benchmark datasets from different domains. Binling Nie, Ruixue Ding, Pengjun Xie, Fei Huang 0002, Chen Qian 0003, Luo Si |
AAAI | 5 |
| 2021 | Conceptualized and Contextualized Gaussian EmbeddingabstractWord embedding can represent a word as a point vector or a Gaussian distribution in high-dimensional spaces. Gaussian distribution is innately more expressive than point vector owing to the ability to additionally capture semantic uncertainties of words, and thus can express asymmetric relations among words more naturally (e.g., animal entails cat but not the reverse. However, previous Gaussian embedders neglect inner-word conceptual knowledge and lack tailored Gaussian contextualizer, leading to inferior performance on both intrinsic (context-agnostic) and extrinsic (context-sensitive) tasks. In this paper, we first propose a novel Gaussian embedder which explicitly accounts for inner-word conceptual units (sememes) to represent word semantics more precisely; during learning, we propose Gaussian Distribution Attention over Gaussian representations to adaptively aggregate multiple sememe distributions into a word distribution, which guarantees the Gaussian linear combination property. Additionally, we propose a Gaussian contextualizer to utilize outer-word contexts in a sentence, producing contextualized Gaussian representations for context-sensitive tasks. Extensive experiments on intrinsic and extrinsic tasks demonstrate the effectiveness of the proposed approach, achieving state-of-the-art performance with near 5.00% relative improvement. Chen Qian 0003, Fuli Feng, Lijie Wen 0001, Tat-Seng Chua |
AAAI | 1 |
| 2021 | Counterfactual Inference for Text Classification DebiasingabstractChen Qian, Fuli Feng, Lijie Wen, Chunping Ma, Pengjun Xie. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Chen Qian 0003, Fuli Feng, Lijie Wen 0001, Chunping Ma, Pengjun Xie |
ACL/IJCNLP (1) | 1 |
| 2021 | MM-CPred: A Multi-task Predictive Model for Continuous-Time Event Sequences with Mixture Learning Losses
Li Lin 0011, Zan Zong, Lijie Wen 0001, Chen Qian 0003, Shuang Li 0015, Jianmin Wang 0001 |
DASFAA (1) | 4 |
| 2020 | Solving Sequential Text Classification as Board-Game Playing
Chen Qian 0003, Fuli Feng, Lijie Wen 0001, Zhenpeng Chen 0001, Li Lin 0011, Yanan Zheng, Tat-Seng Chua |
AAAI | 1 |
| 2020 | An Approach for Process Model Extraction by Multi-grained Text Classification
Chen Qian 0003, Lijie Wen 0001, Akhil Kumar 0001, Leilei Lin, Li Lin 0011, Zan Zong, Shuang Li 0015, Jianmin Wang 0001 |
CAiSE | 1 |
| 2020 | GAHNE: Graph-Aggregated Heterogeneous Network EmbeddingabstractThe real-world networks often compose of different types of nodes and edges with rich semantics, widely known as heterogeneous information network (HIN). Heterogeneous network embedding aims to embed nodes into low-dimensional vectors which capture rich intrinsic information of heterogeneous networks. However, existing models either depend on manually designing meta-paths, ignore mutual effects between different semantics, or omit some aspects of information from global networks. To address these limitations, we propose a novel Graph-Aggregated Heterogeneous Network Embedding (GAHNE), which is designed to extract the semantics of HINs as comprehensively as possible to improve the results of downstream tasks based on graph convolutional neural networks. In GAHNE model, we develop several mechanisms that can aggregate semantic representations from different single-type sub-networks as well as fuse the global information into final embeddings. Extensive experiments on three real-world HIN datasets show that our proposed model consistently outperforms the existing state-of-the-art methods. Xiaohe Li, Lijie Wen 0001, Chen Qian 0003, Jianmin Wang 0001 |
ICTAI | 3 |
| 2020 | Enhancing Text Classification via Discovering Additional Semantic Clues from LogogramsabstractText classification in low-resource languages (eg Thai) is of great practical value for some information retrieval applications (eg sentiment-analysis-based restaurant recommendation). Due to lacking large-scale corpus for learning comprehensive text representation, bilingual text classification which borrows the linguistics knowledge from a rich-resource language becomes a promising solution. Despite the success of bilingual methods, they largely ignore another source of semantic information---the writing system. Noting that most low-resource languages are phonographic languages, we argue that a logographic language (eg Chinese) can provide helpful information for improving some phonographic languages' text classification, since a logographic character (ie logogram) could represent a sememe or a whole concept, not only a phoneme or a sound. In this paper, by using a phonographic labeled corpus and its machine-translated logographic corpus both, we devise a framework to explore the central theme of utilizing logograms as a "semantic detection assistant''. Specifically, from a logographic labeled corpus, we first devise a statistical-significance-based module to pick out informative text pieces. To represent them and further reduce the effects of translation errors, our approach is equipped with Gaussian embedding whose covariances serve as reliable signals of translation errors. For a test document, all seeds' Gaussian representations are used to convolute the document and produce a logographic embedding, before being fused with its phonographic embedding for final prediction. Extensive experiments validate the effectiveness of our approach and further investigations show its generalizability and robustness. Chen Qian 0003, Fuli Feng, Lijie Wen 0001, Li Lin 0011, Tat-Seng Chua |
SIGIR | 1 |
| 2019 | BePT: A Behavior-based Process Translator for Interpreting and Understanding Process ModelsabstractSharing process models on the web has emerged as a common practice. Users can collect and share their experimental process models with others. However, some users always feel confused about the shared process models for lack of necessary guidelines or instructions. Therefore, several process translators have been proposed to explain the semantics of process models in natural language (NL). We find that previous studies suffer from information loss and generate semantically erroneous descriptions that diverge from original model behaviors. In this paper, we propose a novel process translator named BePT (Behavior-based Process Translator) based on the encoder-decoder paradigm, encoding a process model into a middle representation and decoding the representation into NL descriptions. Our theoretical analysis demonstrates that BePT satisfies behavior correctness, behavior completeness and description minimality. The qualitative and quantitative experiments show that BePT outperforms the state-of-the-art baselines. Chen Qian 0003, Lijie Wen 0001, Akhil Kumar 0001 |
CIKM | 1 |
| 2017 | Structural Descriptions of Process Models Based on Goal-Oriented Unfolding
Chen Qian 0003, Lijie Wen 0001, Jianmin Wang 0001, Akhil Kumar 0001 |
CAiSE | 1 |
| 2015 | Shuffled iterative receiver for LDPC-coded MIMO systemsabstractIn this paper, we consider the low density parity check (LDPC) coded multi-input multi-output (MIMO) system with iterative detection and decoding (IDD). Since the traditional frame-by-frame receiver scheme suffers from a huge decoding delay, we propose an efficient scheme with a shuffled structure between the demapper and decoder, which adopts group vertical shuffled belief propagation (BP) algorithm. The proposed shuffled iterative receiver converges faster and significantly reduces the delay introduced by the IDD process. Simulation results demonstrate that our proposed shuffled iterative receiver exhibits several tenths dB of signal-to-noise ratio gain in comparison to the existing schemes, while imposing a much lower average number of iterations for the IDD process. Chen Qian 0003, Zhaocheng Wang 0001, Linglong Dai, Sheng Chen 0001 |
ICC | 2 |
| 2015 | Location-Based Channel Estimation for Massive Full-Dimensional MIMO SystemsabstractIn this paper, a two-dimensional location-based channel estimation algorithm is proposed for massive full-dimensional multiple-input multiple-output (FD-MIMO) systems with intracell pilot reuse, which serves more users using the same timefrequency resource and significantly improves the sum capacity of users compared with conventional schemes. By utilizing the two-dimensional antenna arrays, different users assigned with the same pilot could be distinguished by their non-overlapping azimuth angle-of-arrivals (A-AOAs) and elevation angle-of-arrivals (EAOAs), and thus the interference caused by pilot reuse could be removed. Simulation results show that the sum capacity of the proposed algorithm could be improved as the number of antennas grows large. Wendong Liu, Chen Qian 0003, Zhaocheng Wang 0001 |
VTC Fall | 2 |
| 2015 | Performance optimisation for bit-interleaved coded modulation with iterative demapping with max-log- maximum a posterior detectionabstractMax‐log‐maximum a posterior (MAP) detection is preferred in practical systems rather than log‐MAP because of its lower complexity, but also suffers from a considerable performance loss. In this study, the authors focus on the performance optimisation for (doped) bit‐interleaved coded modulation with iterative demapping schemes where max‐log‐MAP detection is employed. First, they study the effects of max‐log‐MAP detection to the extrinsic information transfer curves of the (doped) demapper and decoder, and reselect the constellation labelling. Second, they find that the small difference between max‐log‐MAP and log‐MAP detection in each iteration would be accumulated during the whole iterative procedure referred to as error accumulation, which causes a large performance loss, and consequently they propose some methods to control the error accumulation. Owing to the labelling reselection and error‐accumulation control, the proposed scheme exhibits several tenths to 1 dB gains, compared with the conventional schemes, while maintaining low complexity. Chen Qian 0003, Qiuliang Xie, Zhaocheng Wang 0001 |
IET Commun. | 1 |
| 2014 | Low complexity detection algorithm for under-determined MIMO systemsabstractA low complexity detection algorithm based on list sphere decoding (LSD) is proposed for under-determined multiple-input multiple-output (UD-MIMO) systems with N transmit antennas and M <; N receive antennas. The proposed algorithm utilizes the unique structure of UD-MIMO systems by dividing the N detection layers into two groups. Group 1 contains layers 1 to M that have similar structures as a symmetric MIMO system; while Group 2 contains layers M + 1 to N that contribute to the rank deficiency of the channel Gram matrix. Tree search algorithms are used for both groups, but with different search radii. A new method is proposed to adaptively adjust the tree search radius of Group 2 based on the statistical properties of the received signals. The employment of the adaptive tree search can significantly reduce the computational complexity. Simulation results show that the proposed algorithm can reduce the complexity by one orders of magnitude with less than 0.01 dB degradation in the Bit-Error-Rate (BER) performance. Chen Qian 0003, Jingxian Wu 0001, Yahong Rosa Zheng, Zhaocheng Wang 0001 |
ICC | 1 |
| 2013 | Two-Stage List Sphere Decoding for Under-Determined Multiple-Input Multiple-Output SystemsabstractA two-stage list sphere decoding (LSD) algorithm is proposed for under-determined multiple-input multiple-output (UD-MIMO) systems that employ N transmit antennas and M<;N receive antennas. The two-stage LSD algorithm exploits the unique structure of UD-MIMO systems by dividing the N detection layers into two groups. Group 1 contains layers 1 to M that have similar structures as a symmetric MIMO system; while Group 2 contains layers M+1 to N that contribute to the rank deficiency of the channel Gram matrix. Tree search algorithms are used for both groups, but with different search radii. A new method is proposed to adaptively adjust the tree search radius of Group 2 based on the statistical properties of the received signals. The employment of the adaptive tree search can significantly reduce the computation complexity. We also propose a modified channel Gram matrix to combat the rank deficiency problem, and it provides better performance than the generalized Gram matrix used in the Generalized Sphere-Decoding (GSD) algorithm. Simulation results show that the proposed two-stage LSD algorithm can reduce the complexity by one to two orders of magnitude with less than 0.1 dB degradation in the Bit-Error-Rate (BER) performance. Chen Qian 0003, Jingxian Wu 0001, Yahong Rosa Zheng, Zhaocheng Wang 0001 |
IEEE Trans. Wirel. Commun. | 1 |
| 2012 | A modified fixed sphere decoding algorithm for under-determined MIMO systemsabstractA modified FSD algorithm is proposed for under-determined (UD) multiple-input multiple-output (MIMO) systems with N transmit antennas and M < N receive antennas. This paper focuses on the low-complexity detection of coded UD-MIMO systems with iterative turbo detection, where a soft-input soft-output (SISO) MIMO detector exchanges soft information with a SISO decoder. In the first iteration, a modified fixed complexity sphere decoding (FSD) method is developed by utilizing the structure of a UD-MIMO system. The modified FSD employs a new detection ordering scheme that has a lower complexity but a better performance compared to the conventional ordering scheme. From the second iteration and beyond, the MIMO detector is implemented with a generalized serial interference cancelation (GSIC) scheme and a block decision feedback equalizer (BDFE) to further reduce the complexity. Simulation results show that the newly proposed FSD-GSIC-BDFE structure can achieve significant performance gains over existing schemes, especially for systems with high level modulations. Chen Qian 0003, Jingxian Wu 0001, Yahong Rosa Zheng, Zhaocheng Wang 0001 |
GLOBECOM | 1 |
| 2012 | Rate-compatible QC-LDPC codes design based on EXIT chart analysisabstractThis paper proposes a "column extension" design method for rate-compatible quasi-cyclic (QC) low-density parity-check (LDPC) codes, which can provide efficient multi-rate coding schemes for communication systems with flexible spectrum efficiency. The optimization of degree distributions among multiple code rates makes the designed QC-LDPC codes achieve excellent performance at each code rate, while the nested parity-check matrix structure of the designed codes facilitates the implementation of multi-rate encoder/decoder and offers the advantage of low-complexity. A comparison with the multi-rate LDPC codes in DVB-S2 specification is addressed. Extrinsic information transfer (EXIT) chart analysis predicts the superiorities of the optimized degree distributions in Eb/N0thresholds, and the bit error rate (BER) simulation results show that the designed codes perform 0.03-0.1 dB better than the DVB-S2 64K codes at each typical code rate, with even shorter code length. Zaishuang Liu, Kewu Peng, Weilong Lei, Chen Qian 0003, Zhaocheng Wang 0001 |
IWCMC | 4 |