Ziming Wu

dblp:49/3584 · DBLP profile ↗
← Back
21ranked-venue papers
6as first author
11since 2021 · last 2026
0000-0003-3348-7727ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Human-computer interaction and ubiquitous computing · 8 · 4 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 1 first-author · 2 since 2021Artificial intelligence and machine learning · 5 · 3 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 MemBuilder: Reinforcing LLMs for Long-Term Memory Construction via Attributed Dense Rewards
abstract
Maintaining consistency in long-term dialogues remains a fundamental challenge for LLMs, as standard retrieval mechanisms often fail to capture the temporal evolution of historical states.While memory-augmented frameworks offer a structured alternative, current systems rely on static prompting of closed-source models or suffer from ineffective training paradigms with sparse rewards.We introduce MemBuilder, a reinforcement learning framework that trains models to orchestrate multi-dimensional memory construction with attributed dense rewards.MemBuilder addresses two key challenges: (1) Sparse Trajectory-Level Rewards: we employ synthetic session-level question generation to provide dense intermediate rewards across extended trajectories; and (2) Multi-Dimensional Memory Attribution: we introduce contribution-aware gradient weighting that scales policy updates based on each component's downstream impact.Experimental results show that MemBuilder enables a 4Bparameter model to outperform state-of-the-art closed-source baselines, exhibiting strong generalization across long-term dialogue benchmarks. 1
Zhiyu Shen, Ziming Wu, Fuming Lai, Shaobing Lian, Yanghui Rao
ACL (1)2
2025 DART: Distilling Autoregressive Reasoning to Silent Thought
abstract
Chain-of-Thought (CoT) reasoning has significantly advanced Large Language Models (LLMs) in solving complex tasks.However, its autoregressive paradigm leads to significant computational overhead, hindering its deployment in latency-sensitive applications.To address this, we propose DART (Distilling Autoregressive Reasoning to Silent Thought), a self-distillation framework that enables LLMs to replace autoregressive CoT with non-autoregressive Silent Thought (ST).Specifically, DART introduces two training pathways: the CoT pathway for traditional reasoning and the ST pathway for generating answers directly from a few ST tokens.The ST pathway utilizes a lightweight Reasoning Evolvement Module (REM) to align its hidden states with the CoT pathway, enabling the ST tokens to evolve into informative embeddings.During inference, only the ST pathway is activated, leveraging evolving ST tokens to deliver the answer directly.Extensive experimental results demonstrate that DART offers significant performance gains compared with existing non-autoregressive baselines without extra inference latency, serving as a feasible alternative for efficient reasoning.
Ziming Wu, De-Chuan Zhan, Fuming Lai, Shaobing Lian
EMNLP2
2025 FROG: Effective Friend Recommendation in Online Games via Modality-aware User Preferences
abstract
Due to the convenience of mobile devices, the online games have become an important part for user entertainments in reality, creating a demand for friend recommendation in online games.However, none of existing approaches can effectively incorporate the multi-modal user features (e.g., images and texts) with the structural information in the friendship graph, due to the following limitations: (1) some of them ignore the high-order structural proximity between users, (2) some fail to learn the pairwise relevance between users at modality-specific level, and (3) some cannot capture both the local and global user preferences on different modalities.By addressing these issues, in this paper, we propose an end-to-end model FROG that better models the user preferences on potential friends.Comprehensive experiments on both offline evaluation and online deployment at Tencent have demonstrated the superiority of FROG over existing approaches.The source code of this paper can be found at https://github.com/socialalgo/FROG.
Dandan Lin, Wenqing Lin, Ziming Wu
SIGIR4
2025 Deciphering Explicit and Implicit Features for Reliable, Interpretable, and Actionable User Churn Prediction in Online Video Games
abstract
The burgeoning online video game industry has sparked intense competition among providers to both expand their user base and retain existing players, particularly within social interaction genres. To anticipate player churn, there is an increasing reliance on machine learning (ML) models that focus on social interaction dynamics. However, the prevalent opacity of most ML algorithms poses a significant hurdle to their acceptance among domain experts, who often view them as "opaque models". Despite the availability of eXplainable Artificial Intelligence (XAI) techniques capable of elucidating model decisions, their adoption in the gaming industry remains limited. This is primarily because non-technical domain experts, such as product managers and game designers, encounter substantial challenges in deciphering the "explicit" and "implicit" features embedded within computational models. This study proposes a reliable, interpretable, and actionable solution for predicting player churn by restructuring model inputs into explicit and implicit features. It explores how establishing a connection between explicit and implicit features can assist experts in understanding the underlying implicit features. Moreover, it emphasizes the necessity for XAI techniques that not only offer implementable interventions but also pinpoint the most crucial features for those interventions. Two case studies, including expert feedback and a within-subject user study, demonstrate the efficacy of our approach.
Laixin Xie, He Wang 0053, Xingxing Xing, Ziming Wu, Xiaojuan Ma, Quan Li 0002
IEEE Trans. Vis. Comput. Graph.6
2024 DAG: Deep Adaptive and Generative K-Free Community Detection on Attributed Graphs
abstract
Community detection on attributed graphs with rich semantic and topological information offers great potential for real-world network analysis, especially user matching in online games. Graph Neural Networks (GNNs) have recently enabled Deep Graph Clustering (DGC) methods to learn cluster assignments from semantic and topological information. However, their success depends on the prior knowledge related to the number of communities K, which is unrealistic due to the high costs and privacy issues of acquisition. In this paper, we investigate the community detection problem without prior K, referred to as K-Free Community Detection problem. To address this problem, we propose a novel Deep Adaptive and Generative model~(DAG) for community detection without specifying the prior K. DAG consists of three key components, i.e., a node representation learning module with masked attribute reconstruction, a community affiliation readout module, and a community number search module with group sparsity. These components enable DAG to convert the process of non-differentiable grid search for the community number, i.e., a discrete hyperparameter in existing DGC methods, into a differentiable learning process. In such a way, DAG can simultaneously perform community detection and community number search end-to-end. To alleviate the cost of acquiring community labels in real-world applications, we design a new metric, EDGE, to evaluate community detection methods even when the labels are not feasible. Extensive offline experiments on five public datasets and a real-world online mobile game dataset demonstrate the superiority of our DAG over the existing state-of-the-art (SOTA) methods. DAG has a relative increase of 7.35% in teams in a Tencent online game compared with the best competitor.
Chang Liu 0078, Yuwen Yang, Yue Ding 0001, Hongtao Lu 0001, Wenqing Lin, Ziming Wu, Wendong Bi
KDD6
2024 Towards Better Modeling With Missing Data: A Contrastive Learning-Based Visual Analytics Perspective
abstract
Missing data can pose a challenge for machine learning (ML) modeling. To address this, current approaches are categorized into feature imputation and label prediction and are primarily focused on handling missing data to enhance ML performance. These approaches rely on the observed data to estimate the missing values and therefore encounter three main shortcomings in imputation, including the need for different imputation methods for various missing data mechanisms, heavy dependence on the assumption of data distribution, and potential introduction of bias. This study proposes a Contrastive Learning (CL) framework to model observed data with missing values, where the ML model learns the similarity between an incomplete sample and its complete counterpart and the dissimilarity between other samples. Our proposed approach demonstrates the advantages of CL without requiring any imputation. To enhance interpretability, we introduce CIVis, a visual analytics system that incorporates interpretable techniques to visualize the learning process and diagnose the model status. Users can leverage their domain knowledge through interactive sampling to identify negative and positive pairs in CL. The output of CIVis is an optimized model that takes specified features and predicts downstream tasks. We provide two usage scenarios in regression and classification tasks and conduct quantitative experiments, expert interviews, and a qualitative user study to demonstrate the effectiveness of our approach. In short, this study offers a valuable contribution to addressing the challenges associated with ML modeling in the presence of missing data by providing a practical solution that achieves high predictive accuracy and model interpretability.
Laixin Xie, Yang Ouyang, Ziming Wu, Quan Li 0002
IEEE Trans. Vis. Comput. Graph.4
2022 RoleSeer: Understanding Informal Social Role Changes in MMORPGs via Visual Analytics
abstract
Massively multiplayer online role-playing games create virtual communities that support heterogeneous “social roles” determined by gameplay interaction behaviors under a specific social context. For all social roles, formal roles are pre-defined, obvious, and explicitly ascribed to the people holding the roles, whereas informal roles are not well-defined and unspoken. Identifying the informal roles and understanding their subtle changes are critical to designing sociability mechanisms. However, it is nontrivial to understand the existence and evolution of such roles due to their loosely defined, interconvertible, and dynamic characteristics. We propose a visual analytics system, RoleSeer, to investigate informal roles from the perspectives of behavioral interactions and depict their dynamic interconversions and transitions. Two cases, experts’ feedback, and a user study suggest that RoleSeer helps interpret the identified informal roles and explore the patterns behind role changes. We see our approach’s potential in investigating informal roles in a broader range of social games.
Laixin Xie, Ziming Wu, Wei Li 0094, Xiaojuan Ma, Quan Li 0002
CHI2
2021 MetaMap: Supporting Visual Metaphor Ideation through Multi-dimensional Example-based Exploration
abstract
Visual metaphors, which are widely used in graphic design, can deliver messages in creative ways by fusing different objects. The keys to creating visual metaphors are diverse exploration and creative combinations, which is challenging with conventional methods like image searching. To streamline this ideation process, we propose to use a mind-map-like structure to recommend and assist users to explore materials. We present MetaMap, a supporting tool which inspires visual metaphor ideation through multi-dimensional example-based exploration. To facilitate the divergence and convergence of the ideation process, MetaMap provides 1) sample images based on keyword association and color filtering; 2) example-based exploration in semantics, color, and shape dimensions; and 3) thinking path tracking and idea recording. We conduct a within-subject study with 24 design enthusiasts by taking a Pinterest-like interface as the baseline. Our evaluation results suggest that MetaMap provides an engaging ideation process and helps participants create diverse and creative ideas.
Youwen Kang, Zhida Sun, Sitong Wang 0001, Ziming Wu, Xiaojuan Ma
CHI5
2021 Exploring Designers' Practice of Online Example Management for Supporting Mobile UI Design
abstract
The use of digital examples plays a critical role in mobile UI design. Yet, it remains unclear how UX/UI designers manage (i.e., collect, archive, and utilize) examples to facilitate their design processes at different stages, and what possible challenges are imposed on the design of proper tools to support these practices. In this paper, we conduct a qualitative interview study with mobile UI/UX designers (12 experts and 12 novices), deriving the commonality in practices and analyzing possible differences across four design phases (Discover, Define, Develop, and Deliver) and expertise. In brief, we find that there is more diverse and frequent use of examples in the Discover and Develop phases, and that experts take more diverse advantage of the information from examples compared to novices. We further identify the challenges faced by designers when using existing example management services, and propose potential design implications for the development of more supportive design tools in the future.
Ziming Wu, Qianyao Xu, Zhenhui Peng, Ying-Qing Xu, Xiaojuan Ma
MobileHCI1
2021 Deep Music Retrieval for Fine-Grained Videos by Exploiting Cross-Modal-Encoded Voice-Overs
abstract
Recently, the witness of the rapidly growing popularity of short videos on different Internet platforms has intensified the need for a background music (BGM) retrieval system. However, existing video-music retrieval methods only based on the visual modality cannot show promising performance regarding videos with fine-grained virtual contents. In this paper, we also investigate the widely added voice-overs in short videos and propose a novel framework to retrieve BGM for fine-grained short videos. In our framework, we use the self-attention (SA) and the cross-modal attention (CMA) modules to explore the intra- and the inter-relationships of different modalities respectively. For balancing the modalities, we dynamically assign different weights to the modal features via a fusion gate. For paring the query and the BGM embeddings, we introduce a triplet pseudo-label loss to constrain the semantics of the modal embeddings. As there are no existing virtual-content video-BGM retrieval datasets, we build and release two virtual-content video datasets HoK400 and CFM400. Experimental results show that our method achieves superior performance and outperforms other state-of-the-art methods with large margins.
Tingtian Li, Zixun Sun, Haoruo Zhang, Ziming Wu, Hui Zhan, Yipeng Yu, Hengcan Shi
SIGIR5
2021 Exploration of text matching methods in Chinese disease Q&A systems: A method using ensemble based on BERT and boosted tree models
Ziming Wu, Zhongan Zhang, Jianbo Lei
J. Biomed. Informatics1
2020 Predicting and Diagnosing User Engagement with Mobile UI Animation via a Data-Driven Approach
abstract
Animation, a common design element in user interfaces (UI), can impact user engagement (UE) with mobile applications. To avoid impairing UE due to improper design of animation, designers rely on resource-intensive evaluation methods like user studies or expert reviews. To alleviate this burden, we propose a data-driven approach to assisting designers in examining UE issues with their animation designs. We first crowdsource UE assessments of mobile UI animations. Based on the collected data, we then build a novel deep learning model that captures both spatial and temporal features of animations to predict their UE levels. Evaluations show that our model achieves a reasonable accuracy. We further leverage the animation feature encoded by our model and a sample set of expert reviews to derive potential UE issues of a particular animation. Finally, we develop a proof-of-concept tool and evaluate its potential usage in actual design practices with experts
Ziming Wu, Yulun Jiang, Xiaojuan Ma
CHI1
2020 WeSeer: Visual Analysis for Better Information Cascade Prediction of WeChat Articles
abstract
Social media, such as Facebook and WeChat, empowers millions of users to create, consume, and disseminate online information on an unprecedented scale. The abundant information on social media intensifies the competition of WeChat Public Official Articles (i.e., posts) for gaining user attention due to the zero-sum nature of attention. Therefore, only a small portion of information tends to become extremely popular while the rest remains unnoticed or quickly disappears. Such a typical "long-tail" phenomenon is very common in social media. Thus, recent years have witnessed a growing interest in predicting the future trend in the popularity of social media posts and understanding the factors that influence the popularity of the posts. Nevertheless, existing predictive models either rely on cumbersome feature engineering or sophisticated parameter tuning, which are difficult to understand and improve. In this paper, we study and enhance a point process-based model by incorporating visual reasoning to support communication between the users and the predictive model for a better prediction result. The proposed system supports users to uncover the working mechanism behind the model and improve the prediction accuracy accordingly based on the insights gained. We use realistic WeChat articles to demonstrate the effectiveness of the system and verify the improved model on a large scale of WeChat articles. We also elicit and summarize the feedback from WeChat domain experts.
Quan Li 0002, Ziming Wu, Lingling Yi, Kristanto Sean Njotoprawiro, Huamin Qu, Xiaojuan Ma
IEEE Trans. Vis. Comput. Graph.2
2019 Design and Evaluation of Service Robot's Proactivity in Decision-Making Support Process
abstract
As service robots are envisioned to provide decision-making support (DMS) in public places, it is becoming essential to design the robot's manner of offering assistance. For example, robot shop assistants that proactively or reactively give product recommendations may impact customers' shopping experience. In this paper, we propose an anticipation-autonomy policy framework that models three levels of proactivity (high, medium and low) of service robots in DMS contexts. We conduct a within-subject experiment with 36 participants to evaluate the effects of DMS robot's proactivity on user perceptions and interaction behaviors. Results show that a highly proactive robot is deemed inappropriate though people can get rich information from it. A robot with medium proactivity helps reduce the decision space while maintaining users' sense of engagement. The least proactive robot grants users more control but may not realize its full capability. We conclude the paper with design considerations for service robot's manner.
Zhenhui Peng, Yunhwan Kwon, Jiaan Lu, Ziming Wu, Xiaojuan Ma
CHI4
2019 Understanding and Modeling User-Perceived Brand Personality from Mobile Application UIs
abstract
Designers strive to make their mobile apps stand out in a competitive market by creating a distinctive brand personality. However, it is unclear whether users can form a consistent impression of brand personality by looking at a few user interface (UI) screenshots in the app store, and if this process can be modeled computationally. To bridge this gap, we first collect crowd assessment on brand personalities depicted by the UIs of 318 applications, and statistically confirm that users can reach substantial agreement. To further model how users process mobile UI visually, we compute UI descriptors including Color, Organization, and Texture at both element and page levels. We feed these descriptors to a computational model, achieving a high accuracy of predicting perceived brand personality (MSE = 0.035 and R^2 = 0.78). This work could benefit designers by highlighting contributing visual factors to brand personality creation and providing quick, low-cost design feedback.
Ziming Wu, Quan Li 0002, Xiaojuan Ma
CHI1
2019 MMAct: A Large-Scale Dataset for Cross Modal Human Action Understanding
abstract
Unlike vision modalities, body-worn sensors or passive sensing can avoid the failure of action understanding in vision related challenges, e.g. occlusion and appearance variation. However, a standard large-scale dataset does not exist, in which different types of modalities across vision and sensors are integrated. To address the disadvantage of vision-based modalities and push towards multi/cross modal action understanding, this paper introduces a new large-scale dataset recorded from 20 distinct subjects with seven different types of modalities: RGB videos, keypoints, acceleration, gyroscope, orientation, Wi-Fi and pressure signal. The dataset consists of more than 36k video clips for 37 action classes covering a wide range of daily life activities such as desktop-related and check-in-based ones in four different distinct scenarios. On the basis of our dataset, we propose a novel multi modality distillation model with attention mechanism to realize an adaptive knowledge transfer from sensor-based modalities to vision-based modalities. The proposed model significantly improves performance of action recognition compared to models trained with only RGB information. The experimental results confirm the effectiveness of our model on cross-subject, -view, -scene and -session evaluation criteria. We believe that this new large-scale multimodal dataset will contribute the community of multimodal based action understanding.
Quan Kong, Ziming Wu, Ziwei Deng, Martin Klinkigt, Bin Tong, Tomokazu Murakami
ICCV2
2019 An AR Benchmark System for Indoor Planar Object Tracking
abstract
Planar object tracking (POT) is the basis of many indoor AR applications. However, there still lacks a systematic way to assess AR trackers. Existing benchmarks usually focus on the tracking accuracy of an algorithm without sufficient details about its sensitivity to various object properties and user behaviors, shedding limited light on possible resulting usability issues. We therefore propose a comprehensive POT benchmark system to understand the weakness of a tracker and derive cues for system improvement. We first identify a set of objects that are commonly used as indoor mobile AR markers and specify their vision-related properties. We then construct a video collection to record typical user interactions with these markers, and statistically quantify the consequent changes as a result of individual or multiple basic manipulations. Evaluation shows that this work can expose a tracker's sensitiveness to different object properties and user behaviors, drawing insights for system improvement and algorithm design.
Ziming Wu, Jiabin Guo, Shuangli Zhang, Xiaojuan Ma
ICME1
2019 DeepDeblur: text image recovery from blur to sharp
Jianhan Mei, Ziming Wu, Yu Qiao 0001, Henghui Ding, Xudong Jiang 0001
Multim. Tools Appl.2
2018 A Multi-Phased Co-design of an Interactive Analytics System for MOBA Game Occurrences
abstract
To ensure the playability of Multiplayer Online Battle Arena (MOBA) games, designers strive to balance different game occurrences. Although machine learning (ML) can help classify matches into different occurrence categories, designers demand more flexible input, interpretable output, and interactive collaboration with ML to facilitate analysis in breadth and depth. To this end, we work closely with a game company to design a visual occurrence analytics system through a stepwise co-design process. We first identify bottlenecks in game designers' conventional practices and their concerns about ML via an observational study. Then, we develop the single-match module of the visualization system to familiarize users with interactive analytics. Next, we incorporate ML models to recommend match segments of interest during occurrence classification and streamline the cross-match analysis. Empirical studies confirm the efficacy of our system. Experts' feedback suggests that our stepwise co-design process indeed helps them better embrace collaboration with machines.
Quan Li 0002, Ziming Wu, Huamin Qu, Xiaojuan Ma
Conference on Designing Interactive Systems2
2018 Coloring with Words: Guiding Image Colorization Through Text-Based Palette Generation
Hyojin Bahng, Seungjoo Yoo, Wonwoong Cho, David Keetae Park, Ziming Wu, Xiaojuan Ma, Jaegul Choo
ECCV (12)5
2018 Mediating Color Filter Exploration with Color Theme Semantics Derived from Social Curation Data
abstract
Despite the popularity of photo editors used to improve image attractiveness and expressiveness on social media, many users have trouble making sense of color filter effects and locating a preferred filter among a set of designer-crafted candidates. The problem gets worse when more computer-generated filters are introduced. To enhance filter findability, we semantically name and organize color effects leveraging data curated by creative communities online. We first model semantic mappings between color themes and keywords in everyday language. Next, we index and organize each filter by the derived semantic information. We conduct three separate studies to investigate the benefit of the semantic features on filter exploration. Our results indicate that color theme semantics constructed through social curation enhances filter findability, providing important implications into how to use the wisdom of the crowd to improve user experience with image editors.
Ziming Wu, Zhida Sun, Manuele Reani, Caroline Jay, Xiaojuan Ma
Proc. ACM Hum. Comput. Interact.1