VLDB 2026 Research / reviewers in the wild / expert
Chenyi Zhuang
dblp:15/10220
· DBLP profile ↗
29ranked-venue papers
9as first author
19since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 18 · 5 first-author · 12 since 2021Databases, data management, data science and information retrieval · 14 · 4 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 14 · 4 first-author · 10 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 first-authorHuman-computer interaction and ubiquitous computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | RAG-R1: Incentivizing the Search and Reasoning Capabilities of LLMs Through Multi-Query ParallelismabstractLarge Language Models (LLMs), despite their remarkable capabilities, are prone to generating hallucinated or outdated content due to their static internal knowledge. While Retrieval-Augmented Generation (RAG) integrated with Reinforcement Learning (RL) offers a solution, these methods are fundamentally constrained by a single-query mode, leading to prohibitive latency and inherent brittleness. To overcome these limitations, we introduce RAG-R1, a novel two-stage training framework centered around multi-query parallelism. Our framework enables LLMs to adaptively leverage internal and external knowledge during the reasoning process while transitioning from the single-query mode to multi-query parallelism. This architectural shift bolsters reasoning robustness while significantly reducing inference latency. Extensive experiments on seven question-answering benchmarks confirm the superiority of our method, which outperforms the strongest baseline by up to 13.7% and decreases inference time by 11.1%. Zhiwen Tan, Qintong Wu, Hongxuan Zhang, Chenyi Zhuang, Jinjie Gu |
AAAI | 5 |
| 2026 | CWPS: Efficient Channel-Wise Parameter Sharing for Knowledge TransferabstractKnowledge transfer aims to apply existing knowledge to different tasks or new data, and it has extensive applications in multi-domain and Multi-Task Learning. The key to this task is quickly identifying a fine-grained object for knowledge sharing and efficiently transferring knowledge. Current methods, such as fine-tuning, layer-wise parameter sharing, and task-specific adapters, only offer coarse-grained sharing solutions and struggle to effectively search for shared parameters, thus hindering the performance and efficiency of knowledge transfer. To address these issues, we propose Channel-Wise Parameter Sharing (CWPS), a novel fine-grained parameter-sharing method for knowledge transfer, which is efficient for parameter sharing, comprehensive, and plug-and-play. For the coarse-grained problem, we first achieve fine-grained parameter sharing by refining the granularity of shared parameters from the level of layers to the level of neurons. The knowledge learned from previous tasks can be utilized through the explicit composition of the model neurons. Besides, we promote an effective search strategy to minimize computational costs, simplifying the selection of shared weights. In addition, our CWPS has strong composability and generalization ability, which theoretically can be applied to any network consisting of linear and convolution layers. We introduce several datasets in both Incremental Learning and Multi-Task Learning scenarios. Our method has achieved state-of-the-art precision-to-parameter ratio performance with various backbones, demonstrating its efficiency and versatility. Mingxuan Cui, Xuewei Li 0003, Cunzheng Wang, Gaoang Wang, Chenyi Zhuang, Jinjie Gu, Xiubo Liang, Xi Li 0001 |
IEEE Trans. Image Process. | 6 |
| 2025 | CSR: Achieving 1 Bit Key-Value Cache via Sparse RepresentationabstractThe emergence of long-context text applications utilizing large language models (LLMs) has presented significant scalability challenges, particularly in memory footprint. The linear growth of the Key-Value (KV) cache, which stores attention keys and values to reduce redundant computations, can significantly increase memory usage and may prevent models from functioning properly in memory-constrained environments. To address this issue, we propose a novel approach called Cache Sparse Representation (CSR), which converts the KV cache by transforming the dense Key-Value cache tensor into sparse indexes and weights, offering a more memory-efficient representation during LLM inference. Furthermore, we introduce NeuralDict, a novel neural network-based method to automatically generate the dictionary used in our sparse representation. Our extensive experiments demonstrate that CSR matches the performance of state-of-the-art KV cache quantization algorithms while ensuring robust functionality in memory-constrained environments. Hongxuan Zhang, Yao Zhao 0011, Jiaqi Zheng 0001, Chenyi Zhuang, Jinjie Gu, Guihai Chen |
AAAI | 4 |
| 2025 | Exploring Heterogeneity and Uncertainty for Graph-based Cognitive Diagnosis Models in Intelligent EducationabstractGraph-based Cognitive Diagnosis (CD) has attracted much research interest due to its strong ability on inferring students' proficiency levels on knowledge concepts. While graph-based CD models have demonstrated remarkable performance, we contend that they still cannot achieve optimal performance due to the neglect of edge heterogeneity and uncertainty. Edges involve both correct and incorrect response logs, indicating heterogeneity. Meanwhile, a response log can have uncertain semantic meanings, e.g., a correct log can indicate true mastery or fortunate guessing, and a wrong log can indicate a lack of understanding or a careless mistake. In this paper, we propose an Informative Semantic-aware Graph-based Cognitive Diagnosis model (ISG-CD), which focuses on how to utilize the heterogeneous graph in CD and minimize effects of uncertain edges. Specifically, to explore heterogeneity, we propose a semantic-aware graph neural networks based CD model. To minimize effects of edge uncertainty, we propose an Informative Edge Differentiation layer from an information bottleneck perspective, which suggests keeping a minimal yet sufficient reliable graph for CD in an unsupervised way. We formulate this process as maximizing mutual information between the reliable graph and response logs, while minimizing mutual information between the reliable graph and the original graph. After that, we prove that mutual information maximization can be theoretically converted to the classic binary cross entropy loss function, while minimizing mutual information can be realized by the Hilbert-Schmidt Independence Criterion.Finally, we adopt an alternating training strategy for optimizing learnable parameters of both the semantic-aware graph neural networks based CD model and the edge differentiation layer. Extensive experiments on three real-world datasets have demonstrated the effectiveness of ISG-CD. Pengyang Shao, Yonghui Yang 0001, Chen Gao 0001, Lei Chen 0051, Kun Zhang 0015, Chenyi Zhuang, Le Wu 0001, Yong Li 0008, Meng Wang 0001 |
KDD (1) | 6 |
| 2025 | AlignCAT: Visual-Linguistic Alignment of Category and Attribute for Weakly Supervised Visual GroundingabstractWeakly supervised visual grounding (VG) aims to locate objects in images based on text descriptions. Despite significant progress, existing methods lack strong cross-modal reasoning to distinguish subtle semantic differences in text expressions due to category-based and attribute-based ambiguity. To address these challenges, we introduce AlignCAT, a novel query-based semantic matching framework for weakly supervised VG. To enhance visual-linguistic alignment, we propose a coarse-grained alignment module that utilizes category information and global context, effectively mitigating interference from category-inconsistent objects. Subsequently, a fine-grained alignment module leverages descriptive information and captures word-level text features to achieve attribute consistency. By exploiting linguistic cues to their fullest extent, our proposed AlignCAT progressively filters out misaligned visual queries and enhances contrastive learning efficiency. Extensive experiments on three VG benchmarks, namely RefCOCO, RefCOCO+, and RefCOCOg, verify the superiority of AlignCAT against existing weakly supervised methods on two VG tasks. Our code is available at: https://github.com/I2-Multimedia-Lab/AlignCAT. Chenyi Zhuang, Wutao Liu, Pan Gao 0001, Nicu Sebe |
ACM Multimedia | 2 |
| 2025 | ReGA: Reasoning and Grounding Decoupled GUI Navigation Agents
Feiyue Ni, Yanchu Guan, Yuchong Sun, Dong Wang 0062, Chenyi Zhuang, Jinjie Gu, Ruihua Song |
NLPCC (1) | 5 |
| 2024 | MoDE: A Mixture-of-Experts Model with Mutual Distillation among the ExpertsabstractThe application of mixture-of-experts (MoE) is gaining popularity due to its ability to improve model's performance. In an MoE structure, the gate layer plays a significant role in distinguishing and routing input features to different experts. This enables each expert to specialize in processing their corresponding sub-tasks. However, the gate's routing mechanism also gives rise to "narrow vision": the individual MoE's expert fails to use more samples in learning the allocated subtask, which in turn limits the MoE to further improve its generalization ability. To effectively address this, we propose a method called Mixture-of-Distilled-Expert (MoDE), which applies moderate mutual distillation among experts to enable each expert to pick up more features learned by other experts and gain more accurate perceptions on their allocated sub-tasks. We conduct plenty experiments including tabular, NLP and CV datasets, which shows MoDE's effectiveness, universality and robustness. Furthermore, we develop a parallel study through innovatively constructing "expert probing", to experimentally prove why MoDE works: moderate distilling knowledge from other experts can improve each individual expert's test performances on their assigned tasks, leading to MoE's overall performance improvement. Zhitian Xie, Yinger Zhang, Chenyi Zhuang, Qitao Shi, Zhining Liu 0001, Jinjie Gu |
AAAI | 3 |
| 2024 | CDFormer: When Degradation Prediction Embraces Diffusion Model for Blind Image Super-ResolutionabstractExisting Blind image Super-Resolution (BSR) methods focus on estimating either kernel or degradation infor-mation, but have long overlooked the essential content details. In this paper, we propose a novel BSR approach, Content-aware Degradation-driven Transformer (CDFormer), to capture both degradation and content rep-resentations. However, low-resolution images cannot pro-vide enough content details, and thus we introduce a diffusion-based module CD Former dif f to first learn Con-tent Degradation Prior (CDP) in both low- and high-resolution images, and then approximate the real distribution given only low-resolution information. Moreover, we apply an adaptive SR network CDFormersR that effectively utilizes CDP to refine features. Compared to previous diffusion-based SR methods, we treat the diffusion model as an estimator that can overcome the limitations of expensive sampling time and excessive diversity. Experiments show that CDFormer can outperform existing methods, establishing a new state-of-the-art performance on various bench-marks under blind settings. Codes and models will be avail-able at https://github.com/I2-Multimedia-Lab/CDFormer. Qingguo Liu, Chenyi Zhuang, Pan Gao 0001, Jie Qin 0004 |
CVPR | 2 |
| 2024 | Intelligent Agents with LLM-based Process AutomationabstractWhile intelligent virtual assistants like Siri, Alexa, and Google Assistant have become ubiquitous in modern life, they still face limitations in their ability to follow multi-step instructions and accomplish complex goals articulated in natural language. However, recent breakthroughs in large language models (LLMs) show promise for overcoming existing barriers by enhancing natural language processing and reasoning capabilities. Though promising, applying LLMs to create more advanced virtual assistants still faces challenges like ensuring robust performance and handling variability in real-world user commands. This paper proposes a novel LLM-based virtual assistant that can automatically perform multi-step operations within mobile apps based on high-level user requests. The system represents an advance in assistants by providing an end-to-end solution for parsing instructions, reasoning about goals, and executing actions. LLM-based Process Automation (LLMPA) has modules for decomposing instructions, generating descriptions, detecting interface elements, predicting next actions, and error checking. Experiments demonstrate the system completing complex mobile operation tasks in Alipay based on natural language instructions. This showcases how large language models can enable automated assistants to accomplish real-world tasks. The main contributions are the novel LLMPA architecture optimized for app process automation, the methodology for applying LLMs to mobile apps, and demonstrations of multi-step task completion in a real-world environment. Notably, this work represents the first real-world deployment and extensive evaluation of a large language model-based virtual assistant in a widely used mobile application with an enormous user base numbering in the hundreds of millions. Yanchu Guan, Dong Wang 0062, Zhixuan Chu, Shiyu Wang 0001, Feiyue Ni, Ruihua Song, Chenyi Zhuang |
KDD | 7 |
| 2024 | Lookahead: An Inference Acceleration Framework for Large Language Model with Lossless Generation AccuracyabstractAs Large Language Models (LLMs) have made significant advancements across various tasks, such as question answering, translation, text summarization, and dialogue systems, the need for accuracy in information becomes crucial, especially for serious financial products serving billions of users like Alipay. However, for a real-world product serving millions of users, the inference speed of LLMs becomes a critical factor compared to a mere experimental model. Yao Zhao 0011, Zhitian Xie, Chenyi Zhuang, Jinjie Gu |
KDD | 4 |
| 2024 | DiffuseST: Unleashing the Capability of the Diffusion Model for Style Transfer
Chenyi Zhuang, Pan Gao 0001 |
MMAsia | 2 |
| 2024 | Magnet: We Never Know How Text-to-Image Diffusion Models Work, Until We Learn How Vision-Language Models FunctionabstractText-to-image diffusion models particularly Stable Diffusion, have revolutionized the field of computer vision. However, the synthesis quality often deteriorates when asked to generate images that faithfully represent complex prompts involving multiple attributes and objects. While previous studies suggest that blended text embeddings lead to improper attribute binding, few have explored this in depth. In this work, we critically examine the limitations of the CLIP text encoder in understanding attributes and investigate how this affects diffusion models. We discern a phenomenon of attribute bias in the text space and highlight a contextual issue in padding embeddings that entangle different concepts. We propose Magnet, a novel training-free approach to tackle the attribute binding problem. We introduce positive and negative binding vectors to enhance disentanglement, further with a neighbor strategy to increase accuracy. Extensive experiments show that Magnet significantly improves synthesis quality and binding accuracy with negligible computational cost, enabling the generation of unconventional and unnatural concepts. Chenyi Zhuang |
NeurIPS | 1 |
| 2024 | DNS-Rec: Data-aware Neural Architecture Search for Recommender SystemsabstractIn the era of data proliferation, efficiently sifting through vast information to extract meaningful insights has become increasingly crucial. This paper addresses the computational overhead and resource inefficiency prevalent in existing Sequential Recommender Systems (SRSs). We introduce an innovative approach combining pruning methods with advanced model designs. Furthermore, we delve into resource-constrained Neural Architecture Search (NAS), an emerging technique in recommender systems, to optimize models in terms of FLOPs, latency, and energy consumption while maintaining or enhancing accuracy. Our principal contribution is the development of a Data-aware Neural Architecture Search for Recommender System (DNS-Rec). DNS-Rec is specifically designed to tailor compact network architectures for attention-based SRS models, thereby ensuring accuracy retention. It incorporates data-aware gates to enhance the performance of the recommendation network by learning information from historical user-item interactions. Moreover, DNS-Rec employs a dynamic resource constraint strategy, stabilizing the search process and yielding more suitable architectural solutions. We demonstrate the effectiveness of our approach through rigorous experiments conducted on three benchmark datasets, which highlight the superiority of DNS-Rec in SRSs. Our findings set a new standard for future research in efficient and accurate recommendation systems, marking a significant step forward in this rapidly evolving field. Sheng Zhang 0028, Maolin Wang 0001, Xiangyu Zhao 0001, Ruocheng Guo, Yao Zhao 0011, Chenyi Zhuang, Jinjie Gu, Zijian Zhang 0009, Hongzhi Yin |
RecSys | 6 |
| 2024 | Tensorized Hypergraph Neural NetworksabstractHypergraph neural networks (HGNN) have recently become attractive and received significant attention due to their excellent performance in various domains. However, most existing HGNNs rely on first-order approximations of hypergraph connectivity patterns, which ignores important high-order information. To address this issue, we propose a novel adjacency-tensor-based Tensorized Hypergraph Neural Network (THNN). THNN is a faithful hypergraph modeling framework through high-order outer product feature message passing and is a natural tensor extension of the adjacency-matrix-based graph neural networks. The proposed THNN is equivalent to a high-order polynomial regression scheme, which enables THNN with the ability to efficiently extract high-order information from uniform hypergraphs. Moreover, in consideration of the exponential complexity of directly processing high-order outer product features, we propose using a partially symmetric CP decomposition approach to reduce model complexity to a linear degree. Additionally, we propose two simple yet effective extensions of our method for non-uniform hypergraphs commonly found in real-world applications. Results from experiments on two widely used hypergraph datasets for 3-D visual object classification show the model's promising performance. Maolin Wang 0001, Yaoming Zhen, Yu Pan 0005, Yao Zhao 0011, Chenyi Zhuang, Zenglin Xu, Ruocheng Guo, Xiangyu Zhao 0001 |
SDM | 5 |
| 2023 | What Wikipedia Misses About Yuriko Nakamura? Predicting Missing Biography Content by Learning Latent Life PatternsabstractAction-related KnowledGe (AKG) is important for facilitating deeper understanding of people’s life patterns, objectives and motivations. In this study, we present a novel framework for automatically predicting missing human biography records in Wikipedia by generating such knowledge. The generation method, which is based on a neural network matrix factorization model, is capable of encoding action semantics from diverse perspectives and discovering latent inter-action relations. By correctly predicting missing information and correcting errors, our work can effectively improve the quality of data about the behavioral records of historical figures in the knowledge base (e.g., biographies in Wikipedia), thus contributing to the understanding and study of human actions by the general public on the one hand, and can be considered as a new paradigm for managing action-related knowledge in digital libraries on the other. Extensive experiments demonstrate that the AKG we generate can capture well missing or “forgotten” human biography related information in Wikipedia. Yijun Duan, Xin Liu 0020, Adam Jatowt, Chenyi Zhuang, Hai-Tao Yu 0003, Steven J. Lynden, Kyoung-Sook Kim 0001, Akiyoshi Matono |
ECAI | 4 |
| 2023 | GreenFlow: A Computation Allocation Framework for Building Environmentally Sound Recommendation SystemabstractGiven the enormous number of users and items, industrial cascade recommendation systems (RS) are continuously expanded in size and complexity to deliver relevant items, such as news, services, and commodities, to the appropriate users. In a real-world scenario with hundreds of thousands requests per second, significant computation is required to infer personalized results for each request, resulting in a massive energy consumption and carbon emission that raises concern. This paper proposes GreenFlow, a practical computation allocation framework for RS, that considers both accuracy and carbon emission during inference. For each stage (e.g., recall, pre-ranking, ranking, etc.) of a cascade RS, when a user triggers a request, we define two actions that determine the computation: (1) the trained instances of models with different computational complexity; and (2) the number of items to be inferred in the stage. We refer to the combinations of actions in all stages as action chains. A reward score is estimated for each action chain, followed by dynamic primal-dual optimization considering both the reward and computation budget. Extensive experiments verify the effectiveness of the framework, reducing computation consumption by 41% in an industrial mobile application while maintaining commercial revenue. Moreover, the proposed framework saves approximately 5000kWh of electricity and reduces 3 tons of carbon emissions per day. Xingyu Lu 0004, Zhining Liu 0001, Yanchu Guan, Hongxuan Zhang, Chenyi Zhuang, Wenqi Ma, Yize Tan, Jinjie Gu |
IJCAI | 5 |
| 2023 | StylePrompter: All Styles Need Is AttentionabstractGAN inversion aims at inverting given images into corresponding latent codes for Generative Adversarial Networks (GANs), especially StyleGAN where exists a disentangled latent space that allows attribute-based image manipulation. As most inversion methods build upon Convolutional Neural Networks (CNNs), we transfer a hierarchical vision Transformer backbone innovatively to predict W+ latent codes at token level. We further apply a Style-driven Multi-scale Adaptive Refinement Transformer (SMART) in ℱ space to refine the intermediate style features of the generator. By treating style features as queries to retrieve lost identity information from the encoder's feature maps, SMART can not only produce high-quality inverted images but also surprisingly adapt to editing tasks. We then prove that StylePrompter lies in a more disentangled W+ and show the controllability of SMART. Finally, quantitative and qualitative experiments demonstrate that Style Prompter can achieve desirable performance in balancing reconstruction quality and editability, and is "smart" enough to fit into most edits, outperforming other ℱ -involved inversion methods. Our code is available at: https://github.com/I2-Multimedia-Lab/StylePrompter. Chenyi Zhuang, Pan Gao 0001, Aljoscha Smolic |
ACM Multimedia | 1 |
| 2023 | Towards Robust Fairness-aware RecommendationabstractDue to the progressive advancement of trustworthy machine learning algorithms, fairness in recommender systems is attracting increasing attention and is often considered from the perspective of users. Conventional fairness-aware recommendation models assume that user preferences remain the same between the training set and the testing set. However, this assumption is arguable in reality, where user preference can shift in the testing set due to the natural spatial or temporal heterogeneity. It is concerning that conventional fairness-aware models may be unaware of such distribution shifts, leading to a sharp decline in the model performance. To address the distribution shift problem, we propose a robust fairness-aware recommendation framework based on Distributionally Robust Optimization (DRO) technique. In specific, we assign learnable weights for each sample to approximate the distributions that leads to the worst-case model performance, and then optimize the fairness-aware recommendation model to improve the worst-case performance in terms of both fairness and recommendation accuracy. By iteratively updating the weights and the model parameter, our framework can be robust to unseen testing sets. To ease the learning difficulty of DRO, we use a hard clustering technique to reduce the number of learnable sample weights. To optimize our framework in a full differentiable manner, we soften the above clustering strategy. Empirically, we conduct extensive experiments based on four real-world datasets to verify the effectiveness of our proposed framework. Hao Yang 0045, Zhining Liu 0001, Zeyu Zhang 0007, Chenyi Zhuang, Xu Chen 0017 |
RecSys | 4 |
| 2022 | Non-stationary Time-aware Kernelized Attention for Temporal Event PredictionabstractModeling sequential data is essential to many applications such as natural language processing, recommendation systems, time series predictions, anomaly detection, etc. When processing sequential data, one of the critical issues is how to capture the temporal-correlation among events. Though prevalent and effective in many applications, conventional approaches such as RNNs and Transformers, struggle with handling the non-stationary characteristics (i.e., such temporal-correlation among events would change over time), which is indeed encountered in many real-world scenarios. In this paper, we present a non-stationary time-aware kernelized attention approach for input sequences of neural networks. By constructing the Generalized Spectral Mixture Kernel (GSMK), and integrating it to the attention mechanism, we mathematically reveal its representation capability in terms of the time-dependent temporal-correlation. Following that, a novel neural network structure is proposed, which would enable us to encode both stationary and non-stationary time event series. Finally, we demonstrate the performance of the proposed method on both synthetic data which presents the theoretical insights, and a variety of real-world datasets which shows its competitive performance against related work. Zhining Liu 0001, Chenyi Zhuang, Yize Tan, Leon Wenliang Zhong, Jinjie Gu |
KDD | 3 |
| 2020 | Two-Stage Audience Expansion for Financial Targeting in MarketingabstractWith the revolution of mobile internet, online finance has grown explosively. In this new area, one challenge of significant importance is how to effectively deliver the financial products or services to a set of target users by marketing. Given a product or service to be promoted and a set of users as seeds, audience expansion is such a targeting technique, which aims to find potential audience among a large number of users. However, in the context of finance, financial products and services are dynamic in nature as they co-vary with the socio-economic environment. Moreover, marketing campaigns for promoting products or services always consist of different rules of play, even for the same type of products or services. As a result, there is a strong demand for the timeliness of seeds in financial targeting. Conventional one-stage audience expansion methods, which generate expanded users by expanding over seeds, would encounter two problems under this setting: (1) the seeds would inevitably involve a number of users that are not representative for expansion, and direct expansion over these noisy seeds would dramatically deteriorate the performance; (2) one-stage expansion over fixed seeds cannot timely and accurately capture users' preferences over the currently running campaign due to the lack of timeliness of seeds. Zhining Liu 0001, Xiao-Fan Niu, Chenyi Zhuang, Yize Tan, Yixiang Mu, Jinjie Gu |
CIKM | 3 |
| 2020 | Hubble: An Industrial System for Audience Expansion in Mobile MarketingabstractRecently, in order to take a preemptive opportunity in the mobile economy, the Internet companies conduct thousands of marketing campaigns every day, to promote their mobile products and services. In the mobile marketing scenario, one of the fundamental issues is the audience expansion task for marketing campaigns. Given a set of seed users, audience expansion aims to seek more users (audiences), who are similar to the seeds and will finish the business goal of the targeted campaign (ie convert). However, the problem is challenging in three aspects. First, a company will run hundreds of campaigns to serve massive users every day. The requirements of scalability and timeliness make training model for each campaign extremely resource-consuming thus impractical. Therefore, we proposed to solve the problem in a two-stage manner, in which the offline stage employs heavyweight user representation learning and the online stage performs embedding-based lightweight audience expansion. Second, conventional two-stage audience expansion systems neglect the high-order user-campaign interactions and usually generate entangled user embeddings, thus fail to achieve high-quality user representation. Third, the seeds, which are usually provided by experts or collected from users' feedbacks, could be noisy and cannot cover the entire actual audiences, thus introduce coverage bias. Unfortunately, to our best knowledge, none of the related literatures tackle this crucial issue of audience expansion. Chenyi Zhuang, Zhiqiang Zhang 0012, Yize Tan, Zhengwei Wu, Zhining Liu 0001, Jianping Wei, Jinjie Gu, Jun Zhou 0011, Yuan Qi 0001 |
KDD | 1 |
| 2019 | A General View for Network Embedding as Matrix FactorizationabstractWe propose a general view that demonstrates the relationship between network embedding approaches and matrix factorization. Unlike previous works that present the equivalence for the approaches from a skip-gram model perspective, we provide a more fundamental connection from an optimization (objective function) perspective. We demonstrate that matrix factorization is equivalent to optimizing two objectives: one is for bringing together the embeddings of similar nodes; the other is for separating the embeddings of distant nodes. The matrix to be factorized has a general form: S-β. The elements of $\mathbfS $ indicate pairwise node similarities. They can be based on any user-defined similarity/distance measure or learned from random walks on networks. The shift number β is related to a parameter that balances the two objectives. More importantly, the resulting embeddings are sensitive to β and we can improve the embeddings by tuning β. Experiments show that matrix factorization based on a new proposed similarity measure and β-tuning strategy significantly outperforms existing matrix factorization approaches on a range of benchmark networks. Xin Liu 0020, Tsuyoshi Murata, Kyoung-Sook Kim 0001, Chatchawan Kotarasu, Chenyi Zhuang |
WSDM | 5 |
| 2019 | Robust visual object clustering and its application to sightseeing spot assessment
Min Ge, Chenyi Zhuang, Qiang Ma 0001 |
Multim. Tools Appl. | 2 |
| 2018 | Dual Graph Convolutional Networks for Graph-Based Semi-Supervised ClassificationabstractThe problem of extracting meaningful data through graph analysis spans a range of different fields, such as the internet, social networks, biological networks, and many others. The importance of being able to effectively mine and learn from such data continues to grow as more and more structured data become available. In this paper, we present a simple and scalable semi-supervised learning method for graph-structured data in which only a very small portion of the training data are labeled. To sufficiently embed the graph knowledge, our method performs graph convolution from different views of the raw data. In particular, a dual graph convolutional neural network method is devised to jointly consider the two essential assumptions of semi-supervised learning: (1) local consistency and (2) global consistency. Accordingly, two convolutional neural networks are devised to embed the local-consistency-based and global-consistency-based knowledge, respectively. Given the different data transformations from the two networks, we then introduce an unsupervised temporal loss function for the ensemble. In experiments using both unsupervised and supervised loss functions, our method outperforms state-of-the-art techniques on different datasets. Chenyi Zhuang, Qiang Ma 0001 |
WWW | 1 |
| 2017 | Understanding People Lifestyles: Construction of Urban Movement Knowledge Graph from GPS TrajectoryabstractTechnologies are increasingly taking advantage of the explosion in the amount of data generated by social multimedia (e.g., web searches, ad targeting, and urban computing). In this paper, we propose a multi-view learning framework for presenting the construction of a new urban movement knowledge graph, which could greatly facilitate the research domains mentioned above. In particular, by viewing GPS trajectory data from temporal, spatial, and spatiotemporal points of view, we construct a knowledge graph of which nodes and edges are their locations and relations, respectively. On the knowledge graph, both nodes and edges are represented in latent semantic space. We verify its utility by subsequently applying the knowledge graph to predict the extent of user attention (high or low) paid to different locations in a city. Experimental evaluations and analysis of a real-world dataset show significant improvements in comparison to state-of-the-art methods. Chenyi Zhuang, Nicholas Jing Yuan, Ruihua Song, Xing Xie 0001, Qiang Ma 0001 |
IJCAI | 1 |
| 2017 | SNS user classification and its application to obscure POI discovery
Chenyi Zhuang, Qiang Ma 0001, Masatoshi Yoshikawa |
Multim. Tools Appl. | 1 |
| 2015 | Discovering Obscure Sightseeing Spots by Analysis of Geo-tagged Social ImagesabstractIn contrast to conventional studies of discovering hot spots, by analyzing geo-tagged images on Flickr, we introduce novel methods to discover obscure sightseeing spots that are less well-known while still worth visiting. To this end, we face two new challenges that the classical authority analysis based methods do not encounter: how to discover and rank spots on the basis of 1) popularity (obscurity level) and 2) scenery quality. For the first challenge, we estimate the obscurity level of a spot in accordance with the visiting asymmetry between photographers who are familiar with a target city and those who are not. For the second challenge, the behavior of both viewers who browsed the images and photographers are analyzed per each spot. We also develop an application system to help users to explore sightseeing spots with different geographical granularities. Experimental evaluations and analysis on a real dataset well demonstrate the effectiveness of the proposed methods. Chenyi Zhuang, Qiang Ma 0001, Xuefeng Liang, Masatoshi Yoshikawa |
ASONAM | 1 |
| 2015 | Location familiarity based flickr photographer classification for POI miningabstractIn this paper, we propose and compare three ways of modeling photographers' location familiarity: a social network driven model, a time driven model and a location driven model. Then, the integration of the three models is further discussed. Experimental evaluations and analysis on a real data set consisting of 14,112 images collected from three cities well demonstrate the performance of the proposed classification methods. Many applications could benefit from information about the location familiarity, such as personalized geo-social recommendation, epidemic dispersion, urban computing, and so on. Chenyi Zhuang, Qiang Ma 0001, Masatoshi Yoshikawa |
SIGSPATIAL/GIS | 1 |
| 2014 | Anaba: An obscure sightseeing spots discovering systemabstractDiscovering and recommending points of interest (POI) are drawing more attention to meet the increasing demand from personalized tours. Unlike conventional systems focusing on popular sightseeing locations, we develop a system, Anaba, to discover the obscure sightseeing spots that are less well-known while still worth visiting. By analyzing geo-tagged images on image hosting websites (Flickr, etc.), Anaba discovers and ranks sightseeing spots based on their obscurity levels and scenery quality. Anaba first selects obscure candidates in accordance with the asymmetry between visitors who are familiar with a target area and those who are not. Then, it evaluates the scenery quality of each candidate by considering both social appreciation and the content of images shot around there. The experiments on a newly retrieved dataset demonstrate the effectiveness of the proposed system. Chenyi Zhuang, Qiang Ma 0001, Xuefeng Liang, Masatoshi Yoshikawa |
ICME | 1 |