VLDB 2026 Research / reviewers in the wild / expert
Kan Li 0001
dblp:21/2083-1
· DBLP profile ↗
87ranked-venue papers
7as first author
40since 2021 · last 2026
0000-0003-3528-4739ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 70 · 4 first-author · 37 since 2021Graphics, computer vision, multimedia, augmented reality and games · 16 · 9 since 2021Databases, data management, data science and information retrieval · 12 · 2 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 first-authorSystems, architecture and hardware · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Leveraging Dissimilarity Invariance as a Robust Anchor for Learning with Noisy LabelsabstractDeep learning models excel in visual recognition but suffer severe performance drops when training labels are corrupted by noise. Under label noise prior work cannot learn accurate similarities and thus misguide the learning process. In this paper, we uncover a complementary and novel phenomenon, Dissimilarity Invariance, whereby semantic dissimilarity between unrelated samples remains stable despite label noise. Leveraging this insight, we propose NegScale, a plug-and-play framework that shifts focus from fragile similarity to robust dissimilarity. NegScale integrates: (1) Structured Negative Orthogonality Penalty (SNOP), enforcing subspace orthogonality for unrelated samples; and (2) Dissimilarity-Calibrated Similarity Adjustment (DCSA), suppressing spurious similarity using dissimilarity anchors. We also give theoretical analysis that proves Dissimilarity Invariance and the effectiveness of NegScale. Empirical results demonstrate that NegScale consistently outperforms state-of-the-art baselines, establishing new benchmarks on CIFAR with synthetic noise and real-world datasets. Wenxiao Fan, Kan Li 0001 |
AAAI | 2 |
| 2026 | LLM-Powered Benchmark Factory: Reliable, Generic, and EfficientabstractPeiwen Yuan, Shaoxiong Feng, Yiwei Li, Xinglin Wang, Yueqi Zhang, Jiayi Shi, Chuyi Tan, Boyuan Pan, Yao Hu, Kan Li. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Peiwen Yuan, Shaoxiong Feng, Yiwei Li 0001, Xinglin Wang, Chuyi Tan, Boyuan Pan, Yao Hu 0002, Kan Li 0001 |
ACL (1) | 10 |
| 2026 | PCF-MAR: Pose- and calibration-free 3D mesh-aligned surface attribute reconstruction from sparse-view images
Rama Bastola Neupane, Zhuqing Mao, Bayu Nadya Kusuma, Kan Li 0001 |
Neurocomputing | 4 |
| 2026 | MODE+: A benchmark and a probe into multimodal open-domain dialogue evaluation
Hang Yin 0007, Xinglin Wang, Pinren Lu, Bin Sun 0004, Peiwen Yuan, Kan Li 0001 |
Neurocomputing | 7 |
| 2026 | High-fidelity 3D reconstruction via unified NeRF-mesh optimization with geometric and color consistency
Rama Bastola Neupane, Kan Li 0001, Zhuqing Mao |
Pattern Recognit. | 2 |
| 2025 | Combating Semantic Contamination in Learning with Label NoiseabstractNoisy labels can negatively impact the performance of deep neural networks. One common solution is label refurbishment, which involves reconstructing noisy labels through predictions and distributions. However, these methods may introduce problematic semantic associations, a phenomenon that we identify as Semantic Contamination. Through an analysis of Robust LR, a representative label refurbishment method, we found that utilizing the logits of views for refurbishment does not adequately balance the semantic information of individual classes. Conversely, using the logits of models fails to maintain consistent semantic relationships across models, which explains why label refurbishment methods frequently encounter issues related to Semantic Contamination. To address this issue, we propose a novel method called Collaborative Cross Learning, which utilizes semi-supervised learning on refurbished labels to extract appropriate semantic associations from embeddings across views and models. Experimental results show that our method outperforms existing approaches on both synthetic and real-world noisy datasets, effectively mitigating the impact of label noise and Semantic Contamination. Wenxiao Fan, Kan Li 0001 |
AAAI | 2 |
| 2025 | From Sub-Ability Diagnosis to Human-Aligned Generation: Bridging the Gap for Text Length Control via MarkerGenabstractPeiwen Yuan, Chuyi Tan, Shaoxiong Feng, Yiwei Li, Xinglin Wang, Yueqi Zhang, Jiayi Shi, Boyuan Pan, Yao Hu, Kan Li. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Peiwen Yuan, Chuyi Tan, Shaoxiong Feng, Yiwei Li 0001, Xinglin Wang, Boyuan Pan, Yao Hu 0002, Kan Li 0001 |
ACL (1) | 10 |
| 2025 | Beyond One-Size-Fits-All: Tailored Benchmarks for Efficient EvaluationabstractEvaluating models on large benchmarks can be very resource-intensive, especially during a period of rapid model evolution. Existing efficient evaluation methods estimate the performance of target models by testing them on a small, static coreset derived from the publicly available evaluation results of source models, which are separate from the target models. However, these approaches rely on the assumption that target models have high prediction consistency with source models, which doesn’t generalize well in practice. To fill this gap, we propose TailoredBench, a method that conducts customized evaluation tailored to each target model. Specifically, a Global-coreset is first constructed as a probe to identify the most consistent source models for each target model with an adaptive source model selection strategy. Afterwards, a scalable K-Medoids clustering algorithm is proposed to extend the Global-coreset to a tailored Native-coreset for each target model. According to the predictions on respective Native-coreset, we estimate the overall performance of target models with a calibrated estimation strategy. Comprehensive experiments on five benchmarks across over 300 models demonstrate that compared to best performing baselines, TailoredBench achieves an average reduction of 31.4% in MAE of accuracy estimates under the same inference budgets, showcasing strong effectiveness and generalizability. Peiwen Yuan, Shaoxiong Feng, Yiwei Li 0001, Xinglin Wang, Chuyi Tan, Boyuan Pan, Yao Hu 0002, Kan Li 0001 |
ACL (1) | 10 |
| 2025 | Do Multimodal Language Models Really Understand Direction? A Benchmark for Compass Direction ReasoningabstractDirection reasoning is essential for intelligent systems to understand the real world. While existing work focuses primarily on spatial reasoning, compass direction reasoning remains underexplored. To address this, we propose the Compass Direction Reasoning (CDR) benchmark, designed to evaluate the direction reasoning capabilities of multimodal language models (MLMs). CDR includes three types images to test spatial (up, down, left, right) and compass (north, south, east, west) directions. Our evaluation reveals that most MLMs struggle with direction reasoning, often performing at random guessing levels. Experiments show that training directly with CDR data yields limited improvements, as it requires an understanding of real-world physical rules. We explore the impact of mixdata and CoT fine-tuning methods, which significantly enhance MLM performance in compass direction reasoning by incorporating diverse data and step-by-step reasoning, improving the model’s ability to understand direction relationships. Hang Yin 0007, Zhifeng Lin, Xin Liu 0132, Bin Sun 0004, Kan Li 0001 |
ICASSP | 5 |
| 2025 | UniCBE: An Uniformity-driven Comparing Based Evaluation Framework with Unified Multi-Objective OptimizationabstractHuman preference plays a significant role in measuring large language models and guiding them to align with human values. Unfortunately, current comparing-based evaluation (CBE) methods typically focus on a single optimization objective, failing to effectively utilize scarce yet valuable preference signals. To address this, we delve into key factors that can enhance the accuracy, convergence, and scalability of CBE: suppressing sampling bias, balancing descending process of uncertainty, and mitigating updating uncertainty.
Following the derived guidelines, we propose UniCBE, a unified uniformity-driven CBE framework which simultaneously optimize these core objectives by constructing and integrating three decoupled sampling probability matrices, each designed to ensure uniformity in specific aspects. We further ablate the optimal tuple sampling and preference aggregation strategies to achieve efficient CBE.
On the AlpacaEval benchmark, UniCBE saves over 17% of evaluation budgets while achieving a Pearson correlation with ground truth exceeding 0.995, demonstrating excellent accuracy and convergence. In scenarios where new models are continuously introduced, UniCBE can even save over 50% of evaluation costs, highlighting its improved scalability. Peiwen Yuan, Shaoxiong Feng, Yiwei Li 0001, Xinglin Wang, Chuyi Tan, Boyuan Pan, Yao Hu 0002, Kan Li 0001 |
ICLR | 10 |
| 2025 | CogLM: Tracking Cognitive Development of Large Language ModelsabstractXinglin Wang, Peiwen Yuan, Shaoxiong Feng, Yiwei Li, Boyuan Pan, Heda Wang, Yao Hu, Kan Li. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Xinglin Wang, Peiwen Yuan, Shaoxiong Feng, Yiwei Li 0001, Boyuan Pan, Heda Wang, Yao Hu 0002, Kan Li 0001 |
NAACL (Long Papers) | 8 |
| 2025 | Every Rollout Counts: Optimal Resource Allocation for Efficient Test-Time ScalingabstractTest-Time Scaling (TTS) improves the performance of Large Language Models (LLMs) by using additional inference-time computation to explore multiple reasoning paths through search. Yet how to allocate a fixed rollout budget most effectively during search remains underexplored, often resulting in inefficient use of compute at test time. To bridge this gap, we formulate test-time search as a resource allocation problem and derive the optimal allocation strategy that maximizes the probability of obtaining a correct solution under a fixed rollout budget. Within this formulation, we reveal a core limitation of existing search methods: solution-level allocation tends to favor reasoning directions with more candidates, leading to theoretically suboptimal and inefficient use of compute. To address this, we propose Direction-Oriented Resource Allocation (DORA), a provably optimal method that mitigates this bias by decoupling direction quality from candidate count and allocating resources at the direction level. To demonstrate DORA’s effectiveness, we conduct extensive experiments on challenging mathematical reasoning benchmarks including MATH500, AIME2024, and AIME2025. The empirical results show that DORA consistently outperforms strong baselines with comparable computational cost, achieving state-of-the-art accuracy. We hope our findings contribute to a broader understanding of optimal TTS for LLMs. Xinglin Wang, Yiwei Li 0001, Shaoxiong Feng, Peiwen Yuan, Chuyi Tan, Boyuan Pan, Yao Hu 0002, Kan Li 0001 |
NeurIPS | 10 |
| 2025 | Stitch and Tell: A Structured Data Augmentation Method for Spatial UnderstandingabstractExisting vision-language models often suffer from spatial hallucinations, i.e., generating incorrect descriptions about the relative positions of objects in an image. We argue that this problem mainly stems from the asymmetric properties between images and text. To enrich the spatial understanding ability of vision-language models, we propose a simple, annotation-free, plug-and-play method named Stitch and Tell (abbreviated as SiTe), which injects structured spatial supervision into multimodal data. It constructs stitched image–text pairs by stitching images along a spatial axis and generating spatially-aware captions or question answer pairs based on the layout of stitched image, without relying on costly advanced models or human involvement. We evaluate SiTe across three architectures including LLaVA-v1.5-7B, LLaVA-Qwen2-1.5B and HALVA-7B, two training datasets, and thirteen benchmarks. Experiments show that SiTe improves spatial understanding tasks such as $\text{MME}_{\text{Position}}$ (+5.50\%) and Spatial-MM (+4.19\%), while maintaining or improving performance on general vision-language benchmarks. Our findings suggest that explicitly injecting spatially-aware structure into training data offers an effective way to mitigate spatial hallucinations and improve spatial understanding, while preserving general vision-language capabilities. Hang Yin 0007, Xiaomin He, Peiwen Yuan, Yiwei Li 0001, Wenxiao Fan, Shaoxiong Feng, Kan Li 0001 |
NeurIPS | 8 |
| 2025 | Silencer: From Discovery to Mitigation of Self-Bias in LLM-as-Benchmark-GeneratorabstractLLM-as-Benchmark-Generator methods have been widely studied as a supplement to human annotators for scalable evaluation, while the potential biases within this paradigm remain underexplored.
In this work, we systematically define and validate the phenomenon of inflated performance in models evaluated on their self-generated benchmarks, referred to as self-bias, and attribute it to sub-biases arising from question domain, language style, and wrong labels.
On this basis, we propose Silencer, a general framework that leverages the heterogeneity between multiple generators at both the sample and benchmark levels to neutralize bias and generate high-quality, self-bias-silenced benchmark. Experimental results across various settings demonstrate that Silencer can suppress self-bias to near zero, significantly improve evaluation effectiveness of the generated benchmark (with an average improvement from 0.655 to 0.833 in Pearson correlation with high-quality human-annotated benchmark), while also exhibiting strong generalizability. Peiwen Yuan, Yiwei Li 0001, Shaoxiong Feng, Xinglin Wang, Chuyi Tan, Boyuan Pan, Yao Hu 0002, Kan Li 0001 |
NeurIPS | 10 |
| 2025 | Mind the Quote: Enabling Quotation-Aware Dialogue in LLMs via Plug-and-Play ModulesabstractHuman–AI conversation frequently relies on quoting earlier text—“check it with the formula I just highlighted”—yet today’s large language models (LLMs) lack an explicit mechanism for locating and exploiting such spans. We formalize the challenge as span-conditioned generation, decomposing each turn into the dialogue history, a set of token-offset quotation spans, and an intent utterance. Building on this abstraction, we introduce a quotation-centric data pipeline that automatically synthesizes task-specific dialogues, verifies answer correctness through multi-stage consistency checks, and yields both a heterogeneous training corpus and the first benchmark covering five representative scenarios. To meet the benchmark’s zero-overhead and parameter-efficiency requirements, we propose QuAda, a lightweight training-based method that attaches two bottleneck projections to every attention head, dynamically amplifying or suppressing attention to quoted spans at inference time while leaving the prompt unchanged and updating < 2.8% of backbone weights. Experiments across models show that QuAda is suitable for all scenarios and generalizes to unseen topics, offering an effective, plug-and-play solution for quotation-aware dialogue. Peiwen Yuan, Yiwei Li 0001, Shaoxiong Feng, Xinglin Wang, Chuyi Tan, Boyuan Pan, Yao Hu 0002, Kan Li 0001 |
NeurIPS | 10 |
| 2025 | A client-level dynamic federated learning reweighting strategy for long-tailed classification
Yang Li 0082, Ji Zhang 0017, Kan Li 0001 |
Expert Syst. Appl. | 4 |
| 2025 | Federated ensemble learning on long-tailed data with prototypesabstractFederated learning enables multiple participants to train models without sharing their raw data. However, long-tailed data with imbalanced sample sizes among clients deteriorates the model’s performance in federated learning. Additionally, existing studies on class prototypes are less effective for federated long-tailed issues, as the difference in class prototype representation between the head classes and tail classes exacerbates global model updates and instability among clients. Therefore, we propose a Federated Ensemble Prototypes Learning (FedEP) approach that employs ensemble class prototypes instead of local class prototypes to alleviate class representation bias. Specifically, each client partitions its local dataset into multiple subsets to derive subset class prototypes and filters biased subset class prototypes using a threshold to obtain ensemble class prototypes. The server then aggregates these ensemble prototypes to develop novel global ones, which guide local training without extra data. Concurrently, we track category probability differences to assess the degree of deviation among class prototypes during the iterative process. Furthermore, our method has proven effective and outperforms baseline approaches on long-tailed data across various experimental settings. Yang Li 0082, Xin Liu 0132, Kan Li 0001 |
Intell. Data Anal. | 3 |
| 2025 | 3D mesh colorization from a single image via geometry prior modulation
Rama Bastola Neupane, Kan Li 0001, Zhuqing Mao |
Knowl. Based Syst. | 2 |
| 2024 | Turning Dust into Gold: Distilling Complex Reasoning Capabilities from LLMs by Leveraging Negative DataabstractLarge Language Models (LLMs) have performed well on various reasoning tasks, but their inaccessibility and numerous parameters hinder wide application in practice. One promising way is distilling the reasoning ability from LLMs to small models by the generated chain-of-thought reasoning paths. In some cases, however, LLMs may produce incorrect reasoning chains, especially when facing complex mathematical problems. Previous studies only transfer knowledge from positive samples and drop the synthesized data with wrong answers. In this work, we illustrate the merit of negative data and propose a model specialization framework to distill LLMs with negative samples besides positive ones. The framework consists of three progressive steps, covering from training to inference stages, to absorb knowledge from negative data. We conduct extensive experiments across arithmetic reasoning tasks to demonstrate the role of negative data in distillation from LLM. Yiwei Li 0001, Peiwen Yuan, Shaoxiong Feng, Boyuan Pan, Bin Sun 0004, Xinglin Wang, Heda Wang, Kan Li 0001 |
AAAI | 8 |
| 2024 | Integrate the Essence and Eliminate the Dross: Fine-Grained Self-Consistency for Free-Form Language GenerationabstractXinglin Wang, Yiwei Li, Shaoxiong Feng, Peiwen Yuan, Boyuan Pan, Heda Wang, Yao Hu, Kan Li. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Xinglin Wang, Yiwei Li 0001, Shaoxiong Feng, Peiwen Yuan, Boyuan Pan, Heda Wang, Yao Hu 0002, Kan Li 0001 |
ACL (1) | 8 |
| 2024 | BatchEval: Towards Human-like Text EvaluationabstractPeiwen Yuan, Shaoxiong Feng, Yiwei Li, Xinglin Wang, Boyuan Pan, Heda Wang, Yao Hu, Kan Li. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Peiwen Yuan, Shaoxiong Feng, Yiwei Li 0001, Xinglin Wang, Boyuan Pan, Heda Wang, Yao Hu 0002, Kan Li 0001 |
ACL (1) | 8 |
| 2024 | Generative Dense Retrieval: Memory Can Be a BurdenabstractPeiwen Yuan, Xinglin Wang, Shaoxiong Feng, Boyuan Pan, Yiwei Li, Heda Wang, Xupeng Miao, Kan Li. Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Peiwen Yuan, Xinglin Wang, Shaoxiong Feng, Boyuan Pan, Yiwei Li 0001, Heda Wang, Xupeng Miao, Kan Li 0001 |
EACL (1) | 8 |
| 2024 | Focused Large Language Models are Stable Many-Shot LearnersabstractPeiwen Yuan, Shaoxiong Feng, Yiwei Li, Xinglin Wang, Yueqi Zhang, Chuyi Tan, Boyuan Pan, Heda Wang, Yao Hu, Kan Li. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024. Peiwen Yuan, Shaoxiong Feng, Yiwei Li 0001, Xinglin Wang, Chuyi Tan, Boyuan Pan, Heda Wang, Yao Hu 0002, Kan Li 0001 |
EMNLP | 10 |
| 2024 | Escape Sky-high Cost: Early-stopping Self-Consistency for Multi-step ReasoningabstractSelf-consistency (SC) has been a widely used decoding strategy for chain-of-thought reasoning. Despite bringing significant performance improvements across a variety of multi-step reasoning tasks, it is a high-cost method that requires multiple sampling with the preset size. In this paper, we propose a simple and scalable sampling process, Early-Stopping Self-Consistency (ESC), to greatly reduce the cost of SC without sacrificing performance. On this basis, one control scheme for ESC is further derivated to dynamically choose the performance-cost balance for different tasks and models. To demonstrate ESC's effectiveness, we conducted extensive experiments on three popular categories of reasoning tasks: arithmetic, commonsense and symbolic reasoning over language models with varying scales. The empirical results show that ESC reduces the average number of sampling of chain-of-thought reasoning by a significant margin on six benchmarks, including MATH (-33.8%), GSM8K (-80.1%), StrategyQA (-76.8%), CommonsenseQA (-78.5%), Coin Flip (-84.2%) and Last Letters (-67.4%), while attaining comparable performances. Yiwei Li 0001, Peiwen Yuan, Shaoxiong Feng, Boyuan Pan, Xinglin Wang, Bin Sun 0004, Heda Wang, Kan Li 0001 |
ICLR | 8 |
| 2024 | Instruction Embedding: Latent Representations of Instructions Towards Task IdentificationabstractInstruction data is crucial for improving the capability of Large Language Models (LLMs) to align with human-level performance. Recent research LIMA demonstrates that alignment is essentially a process where the model adapts instructions' interaction style or format to solve various tasks, leveraging pre-trained knowledge and skills. Therefore, for instructional data, the most important aspect is the task it represents, rather than the specific semantics and knowledge information. The latent representations of instructions play roles for some instruction-related tasks like data selection and demonstrations retrieval. However, they are always derived from text embeddings, encompass overall semantic information that influences the representation of task categories. In this work, we introduce a new concept, instruction embedding, and construct Instruction Embedding Benchmark (IEB) for its training and evaluation. Then, we propose a baseline Prompt-based Instruction Embedding (PIE) method to make the representations more attention on tasks. The evaluation of PIE, alongside other embedding methods on IEB with two designed tasks, demonstrates its superior performance in accurately identifying task categories. Moreover, the application of instruction embeddings in four downstream tasks showcases its effectiveness and suitability for instruction-related tasks. Yiwei Li 0001, Shaoxiong Feng, Peiwen Yuan, Xinglin Wang, Boyuan Pan, Heda Wang, Yao Hu 0002, Kan Li 0001 |
NeurIPS | 9 |
| 2024 | MODE: a multimodal open-domain dialogue dataset with explanation
Hang Yin 0007, Pinren Lu, Bin Sun 0004, Kan Li 0001 |
Appl. Intell. | 5 |
| 2024 | Federated deep long-tailed learning: A survey
Kan Li 0001, Yang Li 0082, Ji Zhang 0017, Xin Liu 0132, Zhichao Ma 0002 |
Neurocomputing | 1 |
| 2024 | Tackling confusion among actions for action segmentation with adaptive margin and energy-driven refinement
Zhichao Ma 0002, Kan Li 0001 |
Mach. Vis. Appl. | 2 |
| 2023 | Heterogeneous-Branch Collaborative Learning for Dialogue GenerationabstractWith the development of deep learning, advanced dialogue generation methods usually require a greater amount of computational resources. One promising approach to obtaining a high-performance and lightweight model is knowledge distillation, which relies heavily on the pre-trained powerful teacher. Collaborative learning, also known as online knowledge distillation, is an effective way to conduct one-stage group distillation in the absence of a well-trained large teacher model. However, previous work has a severe branch homogeneity problem due to the same training objective and the independent identical training sets. To alleviate this problem, we consider the dialogue attributes in the training of network branches. Each branch learns the attribute-related features based on the selected subset. Furthermore, we propose a dual group-based knowledge distillation method, consisting of positive distillation and negative distillation, to further diversify the features of different branches in a steadily and interpretable way. The proposed approach significantly improves branch heterogeneity and outperforms state-of-the-art collaborative learning methods on two widely used open-domain dialogue datasets. Yiwei Li 0001, Shaoxiong Feng, Bin Sun 0004, Kan Li 0001 |
AAAI | 4 |
| 2023 | Towards Diverse, Relevant and Coherent Open-Domain Dialogue Generation via Hybrid Latent VariablesabstractConditional variational models, using either continuous or discrete latent variables, are powerful for open-domain dialogue response generation. However, previous works show that continuous latent variables tend to reduce the coherence of generated responses. In this paper, we also found that discrete latent variables have difficulty capturing more diverse expressions. To tackle these problems, we combine the merits of both continuous and discrete latent variables and propose a Hybrid Latent Variable (HLV) method. Specifically, HLV constrains the global semantics of responses through discrete latent variables and enriches responses with continuous latent variables. Thus, we diversify the generated responses while maintaining relevance and coherence. In addition, we propose Conditional Hybrid Variational Transformer (CHVT) to construct and to utilize HLV with transformers for dialogue generation. Through fine-grained symbolic-level semantic information and additive Gaussian mixing, we construct the distribution of continuous variables, prompting the generation of diverse expressions. Meanwhile, to maintain the relevance and coherence, the discrete latent variable is optimized by self-separation training. Experimental results on two dialogue generation datasets (DailyDialog and Opensubtitles) show that CHVT is superior to traditional transformer-based variational mechanism w.r.t. diversity, relevance and coherence metrics. Moreover, we also demonstrate the benefit of applying HLV to fine-tuning two pre-trained dialogue models (PLATO and BART-base). Bin Sun 0004, Fei Mi, Weichao Wang, Yiwei Li 0001, Kan Li 0001 |
AAAI | 6 |
| 2023 | Denoised Temporal Relation Network for Temporal Action Segmentation
Zhichao Ma 0002, Kan Li 0001 |
PRCV (6) | 2 |
| 2023 | Self-supervised method for 3D human pose estimation with consistent shape and viewpoint factorization
Zhichao Ma 0002, Kan Li 0001, Yang Li 0082 |
Appl. Intell. | 2 |
| 2023 | Stop filtering: Multi-view attribute-enhanced dialogue learning
Yiwei Li 0001, Bin Sun 0004, Shaoxiong Feng, Kan Li 0001 |
Knowl. Based Syst. | 4 |
| 2022 | Diversifying Neural Dialogue Generation via Negative DistillationabstractGenerative dialogue models suffer badly from the generic response problem, limiting their applications to a few toy scenarios.Recently, an interesting approach, namely negative training, has been proposed to alleviate this problem by reminding the model not to generate high-frequency responses during training.However, its performance is hindered by two issues, ignoring low-frequency but generic responses and bringing low-frequency but meaningless responses.In this paper, we propose a novel negative training paradigm, called negative distillation, to keep the model away from the undesirable generic responses while avoiding the above problems.First, we introduce a negative teacher model that can produce querywise generic responses, and then the student model is required to maximize the distance with multi-level negative knowledge.Empirical results show that our method outperforms previous negative training methods significantly. 1 Yiwei Li 0001, Shaoxiong Feng, Bin Sun 0004, Kan Li 0001 |
NAACL-HLT | 4 |
| 2022 | THINK: A novel conversation model for generating grammatically correct and coherent responses
Bin Sun 0004, Shaoxiong Feng, Yiwei Li 0001, Jiamou Liu, Kan Li 0001 |
Knowl. Based Syst. | 5 |
| 2022 | Magnitude Bounded Matrix Factorisation for Recommender SystemsabstractLow rank matrix factorisation is often used in recommender systems as a way of extracting latent features. When dealing with large and sparse datasets, traditional recommendation algorithms face the problem of acquiring large, unrestrained, fluctuating values over predictions. Imposing bounding constraints has been proven an effective solution. However, existing bounding algorithms can only deal with one pair of fixed bounds, and are very time-consuming when applied on large-scale datasets. In this paper, we propose a novel algorithm named Magnitude Bounded Matrix Factorisation (MBMF), which allows different bounds for individual users/items and performs very quickly on large scale datasets. The key idea of our algorithm is to construct a model by constraining the magnitudes of each individual user/item feature vector. By converting coordinate system with radii set as the corresponding magnitudes, MBMF allows the above constrained optimisation problem to become an unconstrained one, which can be solved by unconstrained optimisation algorithms such as the stochastic gradient descent. We also explore an acceleration approach and the choice of magnitudes are given in detail as well. Experiments on synthetic and real datasets demonstrate that in most cases the proposed MBMF is superior over all existing algorithms in terms of accuracy and time complexity. Kan Li 0001 |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2021 | Collaborative Group LearningabstractCollaborative learning has successfully applied knowledge transfer to guide a pool of small student networks towards robust local minima. However, previous approaches typically struggle with drastically aggravated student homogenization when the number of students rises. In this paper, we propose Collaborative Group Learning, an efficient framework that aims to diversify the feature representation and conduct an effective regularization. Intuitively, similar to the human group study mechanism, we induce students to learn and exchange different parts of course knowledge as collaborative groups. First, each student is established by randomly routing on a modular neural network, which facilitates flexible knowledge communication between students due to random levels of representation sharing and branching. Second, to resist the student homogenization, students first compose diverse feature sets by exploiting the inductive bias from sub-sets of training data, and then aggregate and distill different complementary knowledge by imitating a random sub-group of students at each time step. Overall, the above mechanisms are beneficial for maximizing the student population to further improve the model generalization without sacrificing computational efficiency. Empirical evaluations on both image and text tasks indicate that our method significantly outperforms various state-of-the-art collaborative approaches whilst enhancing computational efficiency. Shaoxiong Feng, Hongshen Chen, Xuancheng Ren, Zhuoye Ding, Kan Li 0001, Xu Sun 0001 |
AAAI | 5 |
| 2021 | Multi-View Feature Representation for Dialogue Generation with Bidirectional DistillationabstractNeural dialogue models suffer from low-quality responses when interacted in practice, demonstrating difficulty in generalization beyond training data. Recently, knowledge distillation has been used to successfully regularize the student by transferring knowledge from the teacher. However, the teacher and the student are trained on the same dataset and tend to learn similar feature representations, whereas the most general knowledge should be found through differences. The finding of general knowledge is further hindered by the unidirectional distillation, as the student should obey the teacher and may discard some knowledge that is truly general but refuted by the teacher. To this end, we propose a novel training framework, where the learning of general knowledge is more in line with the idea of reaching consensus, i.e., finding common knowledge that is beneficial to different yet all datasets through diversified learning partners. Concretely, the training task is divided into a group of subtasks with the same number of students. Each student assigned to one subtask not only is optimized on the allocated subtask but also imitates multi-view feature representation aggregated from other students (i.e., student peers), which induces students to capture common knowledge among different subtasks and alleviates the over-fitting of students on the allocated subtasks. To further enhance generalization, we extend the unidirectional distillation to the bidirectional distillation that encourages the student and its student peers to co-evolve by exchanging complementary knowledge with each other. Empirical results and analysis demonstrate that our training framework effectively improves the model generalization without sacrificing training efficiency. Shaoxiong Feng, Xuancheng Ren, Kan Li 0001, Xu Sun 0001 |
AAAI | 3 |
| 2021 | Generating Relevant and Coherent Dialogue Responses using Self-Separated Conditional Variational AutoEncodersabstractBin Sun, Shaoxiong Feng, Yiwei Li, Jiamou Liu, Kan Li. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Bin Sun 0004, Shaoxiong Feng, Yiwei Li 0001, Jiamou Liu, Kan Li 0001 |
ACL/IJCNLP (1) | 5 |
| 2021 | SRCNN-PIL: Side Road Convolution Neural Network Based on Pseudoinverse Learning Algorithm
Mohammed A. B. Mahmoud, Ping Guo 0002, Ahmed Fathy, Kan Li 0001 |
Neural Process. Lett. | 4 |
| 2020 | Posterior-GAN: Towards Informative and Coherent Response Generation with Posterior Generative Adversarial NetworkabstractNeural conversational models learn to generate responses by taking into account the dialog history. These models are typically optimized over the query-response pairs with a maximum likelihood estimation objective. However, the query-response tuples are naturally loosely coupled, and there exist multiple responses that can respond to a given query, which leads the conversational model learning burdensome. Besides, the general dull response problem is even worsened when the model is confronted with meaningless response training instances. Intuitively, a high-quality response not only responds to the given query but also links up to the future conversations, in this paper, we leverage the query-response-future turn triples to induce the generated responses that consider both the given context and the future conversations. To facilitate the modeling of these triples, we further propose a novel encoder-decoder based generative adversarial learning framework, Posterior Generative Adversarial Network (Posterior-GAN), which consists of a forward and a backward generative discriminator to cooperatively encourage the generated response to be informative and coherent by two complementary assessment perspectives. Experimental results demonstrate that our method effectively boosts the informativeness and coherence of the generated response on both automatic and human evaluation, which verifies the advantages of considering two assessment perspectives. Shaoxiong Feng, Hongshen Chen, Kan Li 0001, Dawei Yin 0001 |
AAAI | 3 |
| 2020 | Geometry-Driven Self-Supervised Method for 3D Human Pose EstimationabstractThe neural network based approach for 3D human pose estimation from monocular images has attracted growing interest. However, annotating 3D poses is a labor-intensive and expensive process. In this paper, we propose a novel self-supervised approach to avoid the need of manual annotations. Different from existing weakly/self-supervised methods that require extra unpaired 3D ground-truth data to alleviate the depth ambiguity problem, our method trains the network only relying on geometric knowledge without any additional 3D pose annotations. The proposed method follows the two-stage pipeline: 2D pose estimation and 2D-to-3D pose lifting. We design the transform re-projection loss that is an effective way to explore multi-view consistency for training the 2D-to-3D lifting network. Besides, we adopt the confidences of 2D joints to integrate losses from different views to alleviate the influence of noises caused by the self-occlusion problem. Finally, we design a two-branch training architecture, which helps to preserve the scale information of re-projected 2D poses during training, resulting in accurate 3D pose predictions. We demonstrate the effectiveness of our method on two popular 3D human pose datasets, Human3.6M and MPI-INF-3DHP. The results show that our method significantly outperforms recent weakly/self-supervised approaches. Yang Li 0082, Kan Li 0001, Congzhentao Huang |
AAAI | 2 |
| 2020 | Regularizing Dialogue Generation by Imitating Implicit ScenariosabstractHuman dialogues are scenario-based and appropriate responses generally relate to the latent context knowledge entailed by the specific scenario.To enable responses that are more meaningful and context-specific, we propose to improve generative dialogue systems from the scenario perspective, where both dialogue history and future conversation are taken into account to implicitly reconstruct the scenario knowledge.More importantly, the conversation scenarios are further internalized using imitation learning framework, where the conventional dialogue model that has no access to future conversations is effectively regularized by transferring the scenario knowledge contained in hierarchical supervising signals from the scenario-based dialogue model, so that the future conversation is not required in actual inference.Extensive evaluations show that our approach significantly outperforms state-of-theart baselines on diversity and relevance, and expresses scenario-specific knowledge. Shaoxiong Feng, Xuancheng Ren, Hongshen Chen, Bin Sun 0004, Kan Li 0001, Xu Sun 0001 |
EMNLP (1) | 5 |
| 2020 | TemporalGAT: Attention-Based Dynamic Graph Representation Learning
Ahmed Fathy, Kan Li 0001 |
PAKDD (1) | 2 |
| 2020 | DONE: Enhancing network embedding via greedy vertex domination
Ahmed Fathy, Kan Li 0001 |
Neurocomputing | 2 |
| 2020 | Minimizing rumor influence in multiplex online social networks based on human individual and social behaviors
Adil Imad Eddine Hosni, Kan Li 0001, Sadique Ahmad |
Inf. Sci. | 2 |
| 2020 | Minimizing the influence of rumors during breaking news events in online social networks
Adil Imad Eddine Hosni, Kan Li 0001 |
Knowl. Based Syst. | 2 |
| 2020 | Recognizing actions in images by fusing multiple body structure cues
Yang Li 0082, Kan Li 0001 |
Pattern Recognit. | 2 |
| 2020 | Exploring temporal consistency for human pose estimation in videos
Yang Li 0082, Kan Li 0001 |
Pattern Recognit. | 2 |
| 2019 | ComNE: Reinforcing Network Embedding with Community Learning
Ahmed Fathy, Kan Li 0001 |
ICONIP (4) | 2 |
| 2019 | DARIM: Dynamic Approach for Rumor Influence Minimization in Online Social Networks
Adil Imad Eddine Hosni, Kan Li 0001, Sadique Ahmad |
ICONIP (2) | 2 |
| 2019 | RsyGAN: Generative Adversarial Network for Recommender SystemsabstractMany recommender systems rely on the information of user-item interactions to generate recommendations. In real applications, the interaction matrix is usually very sparse, as a result, the model cannot be optimised stably with different initial parameters and the recommendation performance is unsatisfactory. Many works attempted to solve this problem, however, the parameters in their models may not be trained effectively due to the sparse nature of the dataset which results in a lower quality local optimum. In this paper, we propose a generative network for making user recommendations and a discriminative network to guide the training process. An adversarial training strategy is also applied to train the model. Under the guidance of a discriminative network, the generative network converges to an optimal solution and achieves better recommendation performance on a sparse dataset. We also show that the proposed method significantly improves the precision of the recommendation performance on several datasets. Ruiping Yin, Kan Li 0001, Jie Lu 0001, Guangquan Zhang 0001 |
IJCNN | 2 |
| 2019 | Enhancing Fashion Recommendation with Visual Compatibility RelationshipabstractWith the increasing of online shopping services, fashion recommendation plays an important role in daily online shopping scenes. A lot of recommender systems have been developed with visual information. However, few works take into account compatibility relationship when they are generating recommendations. The challenge is that fashion concept is often subtle and subjective for different customers. In this paper, we propose a fashion compatibility knowledge learning method that incorporates visual compatibility relationships as well as style information. We also propose a fashion recommendation method with domain adaptation strategy to alleviate the distribution gap between the items in target domain and the items of external compatible outfits. Our results indicate that the proposed method is capable of learning visual compatibility knowledge and outperforms all the baselines. Ruiping Yin, Kan Li 0001, Jie Lu 0001, Guangquan Zhang 0001 |
WWW | 2 |
| 2019 | A deeper graph neural network for recommender systems
Ruiping Yin, Kan Li 0001, Guangquan Zhang 0001, Jie Lu 0001 |
Knowl. Based Syst. | 2 |
| 2019 | Relative Pairwise Relationship Constrained Non-Negative Matrix FactorisationabstractNon-negative Matrix Factorisation (NMF) has been extensively used in machine learning and data analytics applications. Most existing variations of NMF only consider how each row/column vector of factorised matrices should be shaped, and ignore the relationship among pairwise rows or columns. In many cases, such pairwise relationship enables better factorisation, for example, image clustering and recommender systems. In this paper, we propose an algorithm named, Relative Pairwise Relationship constrained Non-negative Matrix Factorisation (RPR-NMF), which places constraints over relative pairwise distances amongst features by imposing penalties in a triplet form. Two distance measures, squared Euclidean distance and Symmetric divergence, are used, and exponential and hinge loss penalties are adopted for the two measures, respectively. It is well known that the so-called “multiplicative update rules” result in a much faster convergence than gradient descend for matrix factorisation. However, applying such update rules to RPR-NMF and also proving its convergence is not straightforward. Thus, we use reasonable approximations to relax the complexity brought by the penalties, which are practically verified. Experiments on both synthetic datasets and real datasets demonstrate that our algorithms have advantages on gaining close approximation, satisfying a high proportion of expected constraints, and achieving superior performance compared with other algorithms. Kan Li 0001 |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2018 | Disease gene identification by walking on multilayer heterogeneous networksabstractNOTE FROM ACM: It has been determined that this article plagiarized the contents of a previously published paper. Therefore ACM has shut off access to this paper. Cangfeng Ding, Kan Li 0001 |
CF | 2 |
| 2018 | A Hierarchy Based Influence Maximization Algorithm in Social Networks
Kan Li 0001, Chao Xiang |
ICANN (2) | 2 |
| 2018 | HISBmodel: A Rumor Diffusion Model Based on Human Individual and Social Behaviors in Online Social Networks
Adil Imad Eddine Hosni, Kan Li 0001, Sadique Ahmad |
ICONIP (2) | 2 |
| 2018 | Least Cost Rumor Influence Minimization in Multiplex Social Networks
Adil Imad Eddine Hosni, Kan Li 0001, Cangfeng Ding, Sadique Ahmad |
ICONIP (6) | 2 |
| 2018 | Deeply-Supervised CNN Model for Action Recognition with Trainable Feature AggregationabstractIn this paper, we propose a deeply-supervised CNN model for action recognition that fully exploits powerful hierarchical features of CNNs. In this model, we build multi-level video representations by applying our proposed aggregation module at different convolutional layers. Moreover, we train this model in a deep supervision manner, which brings improvement in both performance and efficiency. Meanwhile, in order to capture the temporal structure as well as preserve more details about actions, we propose a trainable aggregation module. It models the temporal evolution of each spatial location and projects them into a semantic space using the Vector of Locally Aggregated Descriptors (VLAD) technique. This deeply-supervised CNN model integrating the powerful aggregation module provides a promising solution to recognize actions in videos. We conduct experiments on two action recognition datasets: HMDB51 and UCF101. Results show that our model outperforms the state-of-the-art methods. Yang Li 0082, Kan Li 0001 |
IJCAI | 2 |
| 2018 | Centrality Ranking via Topologically Biased Random Walks in Multiplex NetworksabstractCharacterizing the statistically significant centrality of nodes is one of the main objectives of multiplex networks. However, current centrality rankings concentrate only on either the topological structure of the network or diffusion processes based on random walks. A pressing challenge is how to measure centralities of nodes in multiplex networks, depending both on network topology and on diffusion processes (the type of biases in the walks). In the paper, considering these two aspects, we propose a mathematical framework based on topologically biased random walk, called topologically biased multiplex PageRank, which allows to calculate centrality and accordingly rank nodes in multiplex networks. In particular, depending on the nature of biases and the interaction of nodes between different layers, we distinguish additive, multiplicative and combined cases of topologically biased multiplex PageRank. Each case by tuning the bias parameters reflects how the centrality ranking of a node in one layer affects the ranking its replica can gain in the other layers, and captures to which extent the walkers preferentially visit hubs or poorly connected nodes. Experiments on two real-world multiplex networks show that the topologically biased multiplex PageRank outperforms both its corresponding unbiased case and the current ranking methods, and it can efficiently capture the significantly top-ranked nodes in multiplex networks by means of a proper tuning of the biases in the walks. Cangfeng Ding, Kan Li 0001 |
IJCNN | 2 |
| 2018 | Centrality ranking in multiplex networks using topologically biased random walks
Cangfeng Ding, Kan Li 0001 |
Neurocomputing | 2 |
| 2018 | Fuzzy competence model drift detection for data-driven decision support systems
Guangquan Zhang 0001, Jie Lu 0001, Kan Li 0001 |
Knowl. Based Syst. | 4 |
| 2017 | Detecting overlapping protein complexes in dynamic protein-protein interaction networks by developing a fuzzy clustering algorithmabstractProtein complexes play important roles in proteinprotein interaction networks. Recent studies reveal that many proteins have multiple functions and belong to more than one different complexes. To get better complex division, we need to consider time-dependent information of networks. However, only few studies can be found to concentrate on detecting overlapping clusters in time-dependent networks. To solve this problem, we propose integrated model of time-dependent network (IM-TDN) to describe time-dependent networks. On the base of this model, we propose similarity based dynamic fuzzy clustering (SDFC) algorithm to detect overlapping clusters. We apply the algorithm to synthetic data and real world protein-protein interaction network dataset. The results showed that our algorithm by using the model which we proposed achieved better results over the state-of-the-art baseline algorithms. Ruiping Yin, Kan Li 0001, Guangquan Zhang 0001, Jie Lu 0001 |
FUZZ-IEEE | 2 |
| 2017 | Knowledge Graph Based Question Routing for Community Question Answering
Kan Li 0001, Dacheng Qu |
ICONIP (5) | 2 |
| 2017 | A Deep Model Combining Structural Features and Context Cues for Action Recognition in Static Images
Kan Li 0001, Yang Li 0082 |
ICONIP (6) | 2 |
| 2017 | A general description generator for human activity images based on deep understanding framework
Kan Li 0001 |
Neural Comput. Appl. | 2 |
| 2016 | A Generative Model for Recognizing Mixed Group Activities in Still Images
Kan Li 0001, Xiangjian He |
IJCAI | 2 |
| 2016 | Estimating multilateral trade behaviors on the world trade web with limited information
Cangfeng Ding, Kan Li 0001 |
Neurocomputing | 2 |
| 2016 | Main objects interaction activity recognition in real images
Kan Li 0001, JianMeng Pei |
Neural Comput. Appl. | 2 |
| 2015 | Predicting image caption by a unified hierarchical modelabstractAutomatically describing the content of an image is a challenging task in artificial intelligence. The difficulty is particularly pronounced in activity recognition and the image caption revealed by the relationship analysis of the activities involved in the image. This paper presents a unified hierarchical model to model the interaction activity between human and nearby object, and then speculates the image content by analyzing the logical relationship among the interaction activities. In our model, the first-layer factored three-way interaction machine models the 3D spatial context between human and the relevant object to straightly aid the prediction of human-object interaction activities. Then, the activities are further processed through the top-layer factored three-way interaction machine to learn the image content with the help of 3D spatial context among the activities. Experiments on joint dataset show that our unified hierarchical model outperforms state-of-the-arts in predicting human-object interaction activities and describing the image caption. Kan Li 0001 |
ICME | 2 |
| 2015 | Recognizing visual composite in real imagesabstractAutomatically discovering and recognizing the main structured visual pattern of an image is a challenging problem. The most difficulties are how to find the component objects and how to recognize the interaction among these objects. The component objects of the structured visual pattern have consistent 3D spatial co-occurrence layout across images, which manifest themselves as a predictable pattern called visual composite. In this paper, we propose a visual composite recognition model to automatically discover and recognize the visual composite of an image. Our model firstly learns 3D spatial co-occurrence statistics among objects to discover the potential structured visual pattern of an image so that it captures the component objects of visual composite. Secondly, we construct a feedforward architecture using the proposed factored three-way interaction machine to recognize the visual composite, which casts the recognition problem as a structured prediction task. It predicts the visual composite by maximizing the probability of the correct structured label given the component objects and their 3D spatial context. Experiments conducted on a six-class sports dataset and a phrasal recognition dataset respectively demonstrate the encouraging performance of our model in discovery precision and recognition accuracy compared with competing approaches. Kan Li 0001 |
IJCNN | 2 |
| 2015 | Generating image description by modeling spatial context of an imageabstractGenerating the descriptive sentences of a real image is a challenging task in image understanding. The difficulty mainly lies in recognizing the interaction activities between objects, and predicting the relationship between objects and stuff/scene. In this paper, we propose a framework for improving image description generation by addressing the above problems. Our framework mainly includes two models: a unified spatial context model and an image description generation model. The former, as the centerpiece of our framework, models 3D spatial context to learn the human-object interaction activities and predict the semantic relationship between these activities and stuff/scene. The spatial context model casts the problems as latent structured labeling problems, and can be resolved by a unified mathematical optimization. Then based on the semantic relationship, the image description generation model generates image descriptive sentences through the proposed lexicalized tree-based algorithm. Experiments on a joint dataset show that our framework outperforms state-of-the-art methods in spatial co-occurrence context analysis, the human-object interaction recognition, and the image description generation. Kan Li 0001 |
IJCNN | 1 |
| 2015 | Recognizing Human Activity in Still Images by Integrating Group-Based Contextual CuesabstractImages with wider angles usually capture more persons in wider scenes, and recognizing individuals' activities in these images based on existing contextual cues usually meet difficulties. We instead construct a novel group-based cue to utilize the context carried by suitable surrounding persons. We propose a global-local cue integration model (GLCIM) to find a suitable group of local cues extracted from individuals and form a corresponding global cue. A fusion restricted Boltzmann machine, a focal subspace measurement and a cue integration algorithm based on entropy are proposed to enable the GLCIM to integrate most of the relevant local cues and least of the irrelevant ones into the group. Our experiments demonstrate how integrating group-based cues improves the activity recognition accuracies in detail and show that all of the key parts of GLCIM make positive contributions to the increases of the accuracies. Kan Li 0001, Xiangjian He |
ACM Multimedia | 2 |
| 2015 | 3D Depth Perception from Single Monocular Images
Kan Li 0001, Fuyu Lv, JianMeng Pei |
MMM (1) | 2 |
| 2015 | Detecting hierarchical structure of community members in social networks
Fengjiao Chen, Kan Li 0001 |
Knowl. Based Syst. | 2 |
| 2014 | Active Query Selection for Constraint-Based Clustering Algorithms
Walid Atwa, Kan Li 0001 |
DEXA (1) | 2 |
| 2014 | Clustering Evolving Data Stream with Affinity Propagation Algorithm
Walid Atwa, Kan Li 0001 |
DEXA (1) | 2 |
| 2014 | Detecting Hierarchical Structure of Community Members by Link Pattern Expansion Method
Fengjiao Chen, Kan Li 0001 |
WISE (1) | 2 |
| 2014 | A Unified Model for Community Detection of Multiplex Networks
Guangyao Zhu, Kan Li 0001 |
WISE (1) | 2 |
| 2014 | A unified community detection algorithm in complex network
Kan Li 0001, Yin Pang |
Neurocomputing | 1 |
| 2013 | A Hybrid-Sorting Semantic Matching Method
Kan Li 0001, Wensi Mu, Yong Luan, Shaohua An |
ADMA (2) | 1 |
| 2013 | An Energy Model for Network Community Structure Detection
Yin Pang, Kan Li 0001 |
ADMA (1) | 2 |
| 2013 | Detecting core protein complexes based on link patternsabstractProtein complexes (or communities) exhibit specific topological structures in Protein-protein interaction (PPI) networks. As the central parts of the protein complexes, core protein complexes (also called seed communities) can represent functions of the protein complexes and reveal significant features of their topological structures. However, PPI networks are proved to be mixed and overlapped, and current methods have limitation in detecting the mixed structure. In this paper, a core protein complexes detection algorithm based on link patterns is proposed that can detect both clique structure and star structure. For validation, we detect core protein complexes in the real PPI networks, and results show our method can find both the clique structure and the star structure. A comparison with the state-of-the-art methods on the performance of Gene Ontology demonstrates that our method detects core protein complexes with higher similarity of the gene fragments. Kan Li 0001, Fengjiao Chen |
BIBM | 1 |
| 2013 | Communities analysis in protein-protein interaction networksabstractNo common definition of community has been agreed upon till now. One topology can be of different types, uni-partite or bipartite/multipartite. Most of the former literatures are proposed for one type of community. As the understanding of the community definition is different, the grouping results always applicable to the specified network. If the definition changes, the grouping result will no longer be "good". To do the community detection in mixed protein-protein interaction (PPI) networks, we propose a community detection method with two steps. Firstly, group vertices "must be" in the same community by properties; secondly, find overlapping vertices by functions. We apply the energy model to find community structure in PPI networks. The results show that our method is applicable to PPI network, unipartite, bipartite or mixed. It groups vertices with similar property/roles in the same community and finds overlapping vertices in the network. Kan Li 0001, Yin Pang |
BIBM | 1 |
| 2012 | A Vertex Similarity Probability Model for Finding Network Community Structure
Kan Li 0001, Yin Pang |
PAKDD (1) | 1 |
| 2008 | SLUP: A Semantic-Based and Location-Aware Unstructured P2P NetworkabstractThe topological properties of P2P overlay networks are the key factor that dominates the performance of search. Most systems construct logical overlay topologies without considering the underlying physical topology, which causes a serious topology mismatch between the P2P overlay network and the physical network. Thus location of the underlying network should be taken into account during the construction of P2P networks. It can shorten the length of route in network layer and reduce the bandwidth consumed. Furthermore, if nodes with semantic similarity of shared resources are clustered together, queries are routed to the semantically related peers, which can increase the chances of finding the matching resources quickly. Based on these two characters, we propose a Semantic-based and location-aware unstructured P2P Model (SLUP), in which nodes are clustered into domains according to physical distance, and the nodes in a domain are clustered into groups according to similar resources. Simulation experiments show that SLUP can significantly shorten the latency of the searching process and reduce the searching overhead. Xin Sun 0004, Kan Li 0001, Yushu Liu |
HPCC | 2 |