Qingyu Guo

dblp:74/4904 · DBLP profile ↗
← Back
34ranked-venue papers
14as first author
24since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Human-computer interaction and ubiquitous computing · 11 · 4 first-author · 10 since 2021Systems, architecture and hardware · 8 · 3 first-author · 8 since 2021Databases, data management, data science and information retrieval · 7 · 3 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 4 first-author · 2 since 2021Artificial intelligence and machine learning · 6 · 4 first-authorApplied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 2 since 2021Computer networks · 1Software engineering, systems software and programming languages · 1 · 1 since 2021
YearPublicationVenuePosition
2026 ConSearcher: Supporting Conversational Information Seeking in Online Communities with Member Personas
abstract
Many people browse online communities to learn from others’ experiences and opinions, e.g., for constructing travel plans. Conversational search powered by large language models (LLMs) could ease this information-seeking task, but it remains under-investigated within the online community. In this paper, we first conducted an exploratory study (N=10) that indicated the helpfulness of a classic conversational search tool and identified room for improvement. Then, we proposed ConSearcher, an LLM-powered tool with dynamically generated member personas based on user queries to facilitate conversational search in the community. In ConSearcher, users can clarify their interests by checking what a simulated member similar to them may ask and get responses from diverse members’ perspectives. A within-subjects study (N=27) showed that compared to two conversational search baselines, ConSearcher led to significantly higher information-seeking outcome and user engagement but raised concerns about over-personalization. We discuss implications for supporting conversational information seeking in online communities.
Xingbo Wang 0001, Qingyu Guo, Chuhan Shi, Zhenhui Peng
DIS5
2026 Characterizing Cloud-Native LLM Inference at Bytedance and Exposing Optimization Challenges and Opportunities for Future AI Accelerators
abstract
As a major provider of LLM inference services, ByteDance has continuously explored diverse accelerator options to meet the rapidly growing inference demands of various heterogeneous LLM scenarios with higher cost-effectiveness, thereby enabling LLMs to serve more people worldwide. However, during this process, we have found that the complexity and opacity of cloud scenarios and corresponding cloud accelerators make it difficult for academia and many innovative chip startups to fully understand the real demands and challenges of these scenarios, which in turn severely restricts innovation and application potential in this field. To bridge this gap, we first present and analyze the data and characteristics of the ByteDance Doubao LLM app across multiple dimensions, helping the community understand real-world cloud scenarios, and detail the challenges and opportunities we have identified. Second, we propose and plan to open-source our multi-level evaluation framework, XPU-Perf, which includes benchmarks spanning instructions, operators, and models. This framework improves interpretability and trustworthiness, and helps promising new accelerator architectures gain wider adoption and development. Finally, we present comparative results of four typical accelerators, summarize their shortcomings and challenges, conduct in-depth analysis, and highlight numerous architectural and scheduling innovation opportunities we have observed.
Jingwei Cai, Dehao Kong, Hantao Huang, Zishan Jiang, Zixuan Ma, Qingyu Guo, Guiming Shi, Mingyu Gao 0001, Kaisheng Ma, Minghui Yu
HPCA6
2026 EMINDS: Understanding User Behavior Progression for Mental Health Exploration on Social Media
abstract
Mental health is an urgent societal issue, and social scientists are increasingly turning to online mental health communities (OMHCs) to analyze user behavior data for early intervention. However, existing sequence mining techniques fall short of the urgent need to explore the behavior progression of different groups (e.g., recovery or deterioration groups) and track the potential long-term impact of behaviors on mental health status. To address this issue, we introduce EMINDS, a visual analytics system built on a novel automatic mining pipeline that extracts distinct behavior stages and assesses the potential impact of frequent stage patterns on mental health status over time. The system includes a set of interactive visualizations that summarize the meaning of each behavior stage and the evolution of different stage patterns. We feature a pattern-centric Sankey diagram to reveal contextual information about the impact of stage patterns on mental health, helping experts understand the specific changes in sequences before and after a stage pattern. We evaluated the effectiveness and usability of EMINDS through two case studies and expert interviews, which examined the potential stage patterns impacting long-term mental health by analyzing user behaviors on Reddit.
Rui Sheng, Yifang Wang 0001, Xingbo Wang 0001, Shun Dai, Qingyu Guo, Tai-Quan Peng, Huamin Qu, Dongyu Liu
IEEE Trans. Vis. Comput. Graph.5
2025 LightMamba: Efficient Mamba Acceleration on FPGA with Quantization and Hardware Co-design
abstract
State space models (SSMs) like Mamba have recently attracted much attention. Compared to Transformer-based large language models (LLMs), Mamba achieves linear computation complexity with the sequence length and demonstrates superior performance. However, Mamba is hard to accelerate due to the scattered activation outliers and the complex computation dependency, rendering existing LLM accelerators inefficient. In this paper, we propose LightMamba that co-designs the quantization algorithm and FPGA accelerator architecture for efficient Mamba inference. We first propose an FPGA-friendly post-training quantization algorithm that features rotation-assisted quantization and power-of-two SSM quantization to reduce the majority of computation to 4-bit. We further design an FPGA accelerator that partially unrolls the Mamba computation to balance the efficiency and hardware costs. Through computation reordering as well as fine-grained tiling and fusion, the hardware utilization and memory efficiency of the accelerator get drastically improved. We implement LightMamba on Xilinx Versal VCK190 FPGA and achieve 4.65~6.06 x higher energy efficiency over the GPU baseline. When evaluated on Alveo U280 FPGA, LightMamba reaches 93 tokens/s, which is 1.43 x that of the GPU baseline.
Renjie Wei, Songqiang Xu, Linfeng Zhong, Qingyu Guo, Runsheng Wang
DATE5
2025 SpecMamba: Accelerating Mamba Inference on FPGA with Speculative Decoding
abstract
The growing demand for efficient long-sequence modeling on edge devices has propelled widespread adoption of State Space Models (SSMs) like Mamba, due to their superior computational efficiency and scalability. As its autoregressive generation process remains memory-bound, speculative decoding has been proposed that incorporates draft model generation and target model verification. However, directly applying speculative decoding to SSMs faces three key challenges: (1) hidden state backtracking difficulties, (2) tree-based parallel verification incompatibility, and (3) hardware workload mismatch. To address these challenges, we propose SpecMamba, the first FPGA-based accelerator for Mamba with speculative decoding, which features system, algorithm, and hardware co-design. At the system level, we present a memory-aware hybrid backtracking strategy to coordinate both models. At the algorithm level, we propose first-in-first-out (FIFO)-based tree verification with tiling to minimize memory access. At the hardware level, we customize a dataflow that computes linear layers in parallel and SSM layers in series to enable maximal overlapping. Implemented on AMD FPGA platforms (VHK158 and VCK190), SpecMamba achieves a 2.27× speedup over GPU baselines and a 2.85× improvement compared to prior FPGA solutions, while demonstrating 5.41× and 1.26× higher energy efficiency, respectively.
Linfeng Zhong, Songqiang Xu, Huifeng Wen, Tong Xie, Qingyu Guo, Yuan Wang 0001, Meng Li 0004
ICCAD5
2025 Exploring the Evolvement of User Engagement in Online Creative Community under the Surge of Generative AI: A Case Study of DeviantArt
abstract
The rise of AI-generated content (AIGC) is transforming online creative communities (OCCs) and posing challenges to their regulation. The interacting behaviors, such as sharing artworks with descriptions, commenting on creations, and creators' subsequent replying are the essential components of user engagement in these communities. Understanding the influence of AIGC on the evolving user engagement could be helpful for community regulation. In this work, we collect 235K posts and their associated 255K comments from DeviantArt, a large creative community allowing uploading AIGC. Through open coding, we identify five categories of practices in describing and commenting on artworks, respectively. A set of deep learning models are applied to classify the posts and comments. We then combine time series regression analysis, causal inference analysis, and logistic regression analysis, to examine the impact of the surge of AIGC on user engagement. Results suggest that AI-generated artworks show a decreasing emphasis on the content of creations but an increasing trend toward commercial and promotion purposes. AI-generated artworks emphasize less on IP issues than human-created ones, while the awareness of IP issues drops for human-created artworks with the growth of AIGC as well. Although comments with high sentiment valence, for peer bonding or for requesting usage positively predict the reply behavior for human-created artworks, community members are less likely to maintain these interactions as AIGC rises. Finally, we discuss insights and design implications for OCCs.
Qingyu Guo, Kangyu Yuan, Changyang He, Zhenhui Peng, Xiaojuan Ma
Proc. ACM Hum. Comput. Interact.1
2025 MentalImager: Exploring Generative Images for Assisting Support-Seekers' Self-Disclosure in Online Mental Health Communities
abstract
Support-seekers' self-disclosure of their suffering experiences, thoughts, and feelings in the post can help them get needed peer support in online mental health communities (OMHCs). However, such mental health self-disclosure could be challenging. Images can facilitate the manifestation of relevant experiences and feelings in the text; yet, relevant images are not always available. In this paper, we present a technical prototype named MentalImager and validate in a human evaluation study that it can generate topical- and emotional-relevant images based on the seekers' drafted posts or specified keywords. Two user studies demonstrate that MentalImager not only improves seekers' satisfaction with their self-disclosure in their posts but also invokes support-providers' empathy for the seekers and willingness to offer help. Such improvements are credited to the generated images, which help seekers express their emotions and inspire them to add more details about their experiences and feelings. We report concerns on MentalImager and discuss insights for supporting self-disclosure in OMHCs.
Han Zhang 0062, Ryan Louie, Taewook Kim 0001, Qingyu Guo, Shuailin Li, Zhenhui Peng
Proc. ACM Hum. Comput. Interact.6
2024 HG-PIPE: Vision Transformer Acceleration with Hybrid-Grained Pipeline
abstract
Vision Transformer (ViT) acceleration with field programmable gate array (FPGA) is promising but challenging. Existing FPGA-based ViT accelerators mainly rely on temporal architectures, which process different operators by reusing the same hardware blocks and suffer from extensive memory access overhead. Pipelined architectures, either coarse-grained or fine-grained, unroll the ViT computation spatially for memory access efficiency. However, they usually suffer from significant hardware resource constraints and pipeline bubbles induced by the global computation dependency of ViT. In this paper, we introduce HG-PIPE, a pipelined FPGA accelerator for high-throughput and low-latency ViT processing. HG-PIPE features a hybrid-grained pipeline architecture to reduce on-chip buffer cost and couples the computation dataflow and parallelism design to eliminate the pipeline bubbles. HG-PIPE further introduces careful approximations to implement both linear and non-linear operators with abundant Lookup Tables (LUTs), thus alleviating resource constraints. With a VCK190 FPGA, HG-PIPE realizes end-to-end ViT acceleration on a single device and achieves 7118 images/s, which is 2.81× faster than a V100 GPU.
Qingyu Guo, Jiayong Wan, Songqiang Xu, Meng Li 0004, Yuan Wang 0001
ICCAD1
2024 Understanding the Features of Text-Image Posts and Their Received Social Support in Online Grief Support Communities
abstract
People in grief can create posts with text and images to disclose themselves and seek social support in online grief support communities. Existing work largely focuses on understanding the received social support of a post in pure text but often overlooks the post that attaches an image in grief communities. In this paper, we first computationally characterize the textual (e.g., theme), visual (e.g., color), and text-image coherence (i.e., semantic and sentiment coherence) features of text-image posts in a grief support community. Then, we conduct regression analyses to systematically examine the effects of these features on their received informational, emotional, esteem, and network support. We find that attaching a selfie image in the post positively predicts received informational and emotional support, while the social image of a post is a positive predictor of network and esteem support. A post is also likely to get more social support if its text is describing the visible content or telling a story depicted in the image or the perceived emotions in the text and image are not conflict. These results supplement existing research on mental health communities and provide actionable insights into assisting grief people to seek social support online.
Shuailin Li, Han Zhang 0062, Qingyu Guo, Zhenhui Peng
ICWSM5
2024 Engage Wider Audience or Facilitate Quality Answers? a Mixed-methods Analysis of Questioning Strategies for Research Sensemaking on a Community Q&A Site
abstract
Discussing research-sensemaking questions on Community Question and Answering (CQA) platforms has been an increasingly common practice for the public to participate in science communication. Nonetheless, how users strategically craft research-sensemaking questions to engage public participation and facilitate knowledge construction is a significant yet less understood problem. To fill this gap, we collected 837 science-related questions and 157,684 answers from Zhihu, and conducted a mixed-methods study to explore user-developed strategies in proposing research-sensemaking questions, and their potential effects on public engagement and knowledge construction. Through open coding, we captured a comprehensive taxonomy of question-crafting strategies, such as eyecatching narratives with counter-intuitive claims and rigorous descriptions with data use. Regression analysis indicated that these strategies correlated with user engagement and answer construction in different ways (e.g., emotional questions attracted more views and answers), yet there existed a general divergence between wide participation and quality knowledge establishment, when most questioning strategies could not ensure both. Based on log analysis, we further found that collaborative editing afforded unique values in refining research-sensemaking questions regarding accuracy, rigor, comprehensiveness and attractiveness. We propose design implications to facilitate accessible, accurate and engaging science communication on CQA platforms.
Changyang He, Yue Deng 0003, Qingyu Guo, Yu Zhang 0097, Zhicong Lu, Bo Li 0001
Proc. ACM Hum. Comput. Interact.4
2024 CASCADE: A Framework for CNN Accelerator Synthesis With Concatenation and Refreshing Dataflow
abstract
Layer Pipeline (LP) represents an innovative architecture for neural network accelerators, which implements task-level pipelining at the granularity of layers. Despite improvements in throughput, LP architectures face challenges due to complicated dataflow design, intricate design space and high resource requirements. In this paper, we introduce an accelerator synthesis framework, CASCADE. CASCADE leverages a novel dataflow, CARD, to efficiently manage convolutional operations’ irregular memory access patterns using simplified logic and minimal buffers. It also employs advanced design space exploration methods to optimize unrolling parallelism and FIFO depth settings automatically for each layer. Finally, to further enhance resource efficiency, CASCADE leverages Lookup Table-based multiplication and accumulation units. With extensive experimental results, we demonstrate that CASCADE significantly outperforms existing works, achieving a$3\times $improvement in resource efficiency and a$4\times $improvement in power efficiency. It achieves over$1.5\times 10^{4}$frames per second throughput and 71.9% accuracy on ImageNet.
Qingyu Guo, Haoyang Luo, Meng Li 0004, Xiyuan Tang, Yuan Wang 0001
IEEE Trans. Circuits Syst. I Regul. Pap.1
2024 A 16.38TOPS and 4.55POPS/W SRAM Computing-in-Memory Macro for Signed Operands Computation and Batch Normalization Implementation
abstract
Edge artificial intelligence applications impose rigorous demands on local hardware to improve throughput and energy efficiency. Computing-in-memory (CIM) architectures provide high parallel and energy-efficient solutions to accelerate the multiply-and-accumulate (MAC) operations in neural networks (NNs). While SRAM-based charge-domain CIM is achieving thousands of TOPS/W energy efficiency, it encounters limitations when dealing with full NN model deployments where both activations and weights are signed. This paper proposes an SRAM-based signed batch normalization (BN) CIM macro for supporting efficient bitwise sparse MAC computation with signed operands and BN operations in deep neural networks. The key features of this macro encompass: 1) a multibit weight unit for the optimization of bitstream sparsity and the sign bit computation, 2) a 2b-serial input configuration to increase throughput and the ADC energy amortization, and 3) a quantization-hardware co-design for the BN implementation. Measurement results show that the proposed 28 nm 64 Kb CIM macro achieves 16.38 TOPS throughput and 4.55 POPS/W energy efficiency, both normalized to 1b operands. The test accuracy of CIFAR10 is 92%, based on the ResNet18 model with co-design BN implementation at signed-8b precision activations and weights.
Qingyu Guo, Xiyuan Tang, Renjie Wei, Meng Li 0004, Runsheng Wang, Yuan Wang 0001
IEEE Trans. Circuits Syst. I Regul. Pap.2
2023 What Makes Creators Engage with Online Critiques? Understanding the Role of Artifacts' Creation Stage, Characteristics of Community Comments, and their Interactions
abstract
Online critique communities (OCCs) provide a convenient space for creators to solicit feedback on their artifacts and improve skills. Creators’ behavioral, emotional, and cognitive engagement with comments on their works contribute to their skill development. However, what kinds of critique creators feel engaging may change with the creation stage of their shared artifacts. In this paper, we first model three dimensions of engagement expressed in creators’ replies to peer comments. Then we quantitatively examine how their engagement is affected by artifacts’ stage and feedback characteristics via regression analysis. Results show that creators sharing works-in-progress tend to exhibit lower behavioral and emotional engagement, but higher cognitive engagement than those sharing complete works. The increase in the valence of the feedback is associated with a stronger increase in behavior engagement for seekers sharing complete works than works-in-progress. Finally, we discuss how our insights could benefit OCCs and other online help-seeking platforms.
Qingyu Guo, Chao Zhang 0082, Hanfang Lyu, Zhenhui Peng, Xiaojuan Ma
CHI1
2023 A Survey on Knowledge Graph-Based Recommender Systems : Extended Abstract
abstract
To solve the information explosion problem and enhance user experience in various online applications, recommender systems have been developed to model users’ preferences. Although numerous efforts have been made toward more personalized recommendations, recommender systems still suffer from several challenges, such as data sparsity and cold-start problems. In recent years, generating recommendations with the knowledge graph as side information has attracted considerable interest. Such an approach can not only alleviate the above mentioned issues for a more accurate recommendation, but also provide explanations for recommended items. In this paper, we conduct a systematical survey of knowledge graph-based recommender systems. We collect recently published papers in this field, and group them into three categories, i.e., embedding-based methods, connection-based methods, and propagation-based methods. Also, we further subdivide each category according to the characteristics of these approaches. Moreover, we investigate the proposed algorithms by focusing on how the papers utilize the knowledge graph for accurate and explainable recommendation. Finally, we propose several potential research directions in this field.
Qingyu Guo, Fuzhen Zhuang, Chuan Qin 0002, Hengshu Zhu, Xing Xie 0001, Hui Xiong 0001, Qing He 0003
ICDE1
2023 Surgical Video Captioning with Mutual-Modal Concept Alignment
Zhen Chen 0018, Qingyu Guo, Leo K. T. Yeung, Danny T. M. Chan, Zhen Lei 0001, Hongbin Liu 0001, Jinqiao Wang
MICCAI (9)2
2023 CriTrainer: An Adaptive Training Tool for Critical Paper Reading
abstract
Learning to read scientific papers critically, which requires first grasping their main ideas and then raising critical thoughts, is important yet challenging for novice researchers. The traditional ways to develop critical paper reading (CPR) skills, e.g., checking general tutorials or taking reading courses, often can not provide individuals with adaptive and accessible support. In this paper, we first derive user requirements of a CPR training tool based on literature and a survey study (N=52). Then, we develop CriTrainer , an interactive tool for CPR training. It leverages text summarization techniques to train readers’ skills in grasping the paper’s main ideas. It further utilizes template-based generated questions to help them learn how to raise critical thoughts. A mixed-design study (N=24) shows that compared to a baseline tool with general CPR guidance, students trained by CriTrainer perform better in independently raising critical thinking questions on a new paper. We conclude with design considerations for CPR training tools.
Kangyu Yuan, Hehai Lin, Shilei Cao 0005, Zhenhui Peng, Qingyu Guo, Xiaojuan Ma
UIST5
2023 Exploring the Effects of Event-induced Sudden Influx of Newcomers to Online Pop Music Fandom Communities: Content, Interaction, and Engagement
abstract
Online fandom communities (OFCs) provide a convenient space for fans to create, collect, and discuss the content of their mutual interest (e.g., music artists). Real-world events could frequently attract outsiders to join OFCs, providing both the opportunity to expand the fan base and challenges to manage the community. However, it is unclear that how influxes of newcomers would influence the development of OFCs and what user behaviors may be correlated with their future engagement. To fill this gap, we took the music OFCs as the focus, and quantitatively analyzed user behaviors and their correlations with users' future engagement in the community. Results suggested that 1) event-induced newcomers expressed more hate speech and negative sentiment, praised less celebrity-related content (e.g., song, album), and interacted with narrower cohorts than existing members; 2) Although existing members tended to receive more upvotes during the events than before and after the events, newcomers showed an opposite trend; 3) keeping users' activeness, expressing positive sentiments, and having diverse interactions during periods of influx were helpful when maintaining members' future levels of engagement. This work deepened the understanding of fan behaviors in the dynamic period, and we discussed how our insights could benefit OFCs.
Qingyu Guo, Chuhan Shi, Zhuohao Yin, Chengzhong Liu, Xiaojuan Ma
Proc. ACM Hum. Comput. Interact.1
2023 Characterizing and Forecasting Urban Vibrancy Evolution: A Multi-View Graph Mining Perspective
abstract
Urban vibrancy describes the prosperity, diversity, and accessibility of urban areas, which is vital to a city’s socio-economic development and sustainability. While many efforts have been made for statically measuring and evaluating urban vibrancy, there are few studies on the evolutionary process of urban vibrancy, yet we know little about the relationship between urban vibrancy evolution and sophisticated spatiotemporal dynamics. In this article, we make use of multi-sourced urban data to develop a data-driven framework, U-Evolve , to investigate urban vibrancy evolution. Specifically, we first exploit the spatiotemporal characteristics of urban areas to create multi-view time-dependent graphs. Then, we analyze the contextual features and graph patterns of multi-view time-dependent graphs in terms of informing future urban vibrancy variations. Our analysis validates the informativeness of multi-view time-dependent graphs for characterizing and informing future urban vibrancy evolution. After that, we construct a feature based model to forecast future urban vibrancy evolution and quantify each feature’s importance. Moreover, to further enhance the forecasting effectiveness, we propose a graph learning based model to capture spatiotemporal autocorrelation of urban areas based on multi-view time-dependent graphs in an end-to-end manner. Finally, extensive experiments on two metropolises, Beijing and Shanghai, demonstrate the effectiveness of our forecasting models. The U-Evolve framework has also been deployed in the production environment to deliver real-world urban development and planning insights for various cities in China.
Hao Liu 0026, Qingyu Guo, Hengshu Zhu, Yanjie Fu, Fuzhen Zhuang, Xiaojuan Ma, Hui Xiong 0001
ACM Trans. Knowl. Discov. Data2
2022 Understanding and Modeling Viewers' First Impressions with Images in Online Medical Crowdfunding Campaigns
abstract
Online medical crowdfunding campaigns (OMCCs) help patients seek financial support. First impressions (FIs) of an OMCC, including perceived empathy, credibility, justice, impact, and attractiveness, could affect viewers’ donation decisions. Images play a crucial role in manifesting FIs, and it is beneficial for fundraisers to understand how viewers may judge their selected images for OMCCs beforehand. This work proposes a data-driven approach to assessing whether an OMCC image conveys appropriate FIs. We first crowdsource viewers’ perception of OMCC images. Statistical analysis confirms that agreement on all five dimensions of FIs exists, and these FIs positively correlate with donation intention. We compute image content, color, texture, and composition features, then analyze the correlation between these visual features and FIs. We further predict FIs based on these features, and the best model achieves an overall F1-score of 0.727. Finally, we discuss how our insights could benefit fundraisers and possible ethical concerns.
Qingyu Guo, Zhenhui Peng, Xiaojuan Ma
CHI1
2022 A 4-bit Integer-Only Neural Network Quantization Method Based on Shift Batch Normalization
abstract
Neural networks are powerful, but at the cost of huge amounts of computation. Deploying neural networks on edge devices is especially challenging. Quantization is a possible solution to alleviate the huge cost, while most quantization methods are not sufficiently hardware-friendly. In this paper, we proposed an integer-only quantization method. With no division or big integer multiplication, this quantization method is suitable to be deployed on co-designed hardware platforms. We applied 4-bit quantization on some classical networks and corresponding datasets. On MNIST, CIFAR10 and CFAR100, quantization networks perform as well as original networks. On SpeechCommands, accuracy error induced by quantization is 0.16%. We also deployed quantized networks under OpenCL framework and on a flash-based in-memory-computing chip to verify this method’s feasibility.
Qingyu Guo, Xiaoxin Cui, Aifei Zhang, Xinjie Guo, Yuan Wang 0001
ISCAS1
2022 A 28nm 64Kb SRAM based Inference-Training Tri-Mode Computing-in-Memory Macro
abstract
Many computing-in-memory (CIM) macros achieve local inference with forward propagation (FP), and some CIM macros also support backward propagation (BP) computation. However, they can not calculate the weight change related to the learning rate and forward propagation input. these macros can not support backward propagation training algorithm completely. In this paper, we proposed a 28nm 64Kb SRAM based CIM macro, which supports a more complete backward propagation training algorithm. This macro supports three computing modes. A multiply unit (MU) supports FP and BP modes. A multiply circuit (MC) supports three-inputs-multiplication (TIM) mode for the weight change analog computing. MC uses the principle of charge sharing which has a high resistance to process variation and perfect linearity. In FP and BP modes, this macro achieves an energy efficiency of 42.1TOPS/W with 2-bit input, 8-bit weight and 14-bit output multiplication and accumulation operations (MAC). In TIM mode, this macro achieves an energy efficiency of 59.4 - 2222TOPS/W with multiplication of 3 inputs and 1 output.
Nanbing Pan, Xiaoxin Cui, Kanglin Xiao, Qingyu Guo, Yuan Wang 0001
ISCAS5
2022 Who will Win the Data Science Competition? Insights from KDD Cup 2019 and Beyond
abstract
Data science competitions are becoming increasingly popular for enterprises collecting advanced innovative solutions and allowing contestants to sharpen their data science skills. Most existing studies about data science competitions have a focus on improving task-specific data science techniques, such as algorithm design and parameter tuning. However, little effort has been made to understand the data science competition itself. To this end, in this article, we shed light on the team’s competition performance, and investigate the team’s evolving performance in the crowd-sourcing competitive innovation context. Specifically, we first acquire and construct multi-sourced datasets of various data science competitions, including the KDD Cup 2019 machine learning competition and beyond. Then, we conduct an empirical analysis to identify and quantify a rich set of features that are significantly correlated with teams’ future performances. By leveraging team’s rank as a proxy, we observe “the stronger, the stronger” rule; that is, top-ranked teams tend to keep their advantages and dominate weaker teams for the rest of the competition. Our results also confirm that teams with diversified backgrounds tend to achieve better performances. After that, we formulate the team’s future rank prediction problem and propose the Multi-Task Representation Learning (MTRL) framework to model both static features and dynamic features. Extensive experimental results on four real-world data science competitions demonstrate the team’s future performance can be well predicted by using MTRL. Finally, we envision our study will not only help competition organizers to understand the competition in a better way, but also provide strategic implications to contestants, such as guiding the team formation and designing the submission strategy.
Hao Liu 0026, Qingyu Guo, Hengshu Zhu, Fuzhen Zhuang, Shenwen Yang, Dejing Dou, Hui Xiong 0001
ACM Trans. Knowl. Discov. Data2
2022 A Survey on Knowledge Graph-Based Recommender Systems
abstract
To solve the information explosion problem and enhance user experience in various online applications, recommender systems have been developed to model users’ preferences. Although numerous efforts have been made toward more personalized recommendations, recommender systems still suffer from several challenges, such as data sparsity and cold-start problems. In recent years, generating recommendations with the knowledge graph as side information has attracted considerable interest. Such an approach can not only alleviate the above mentioned issues for a more accurate recommendation, but also provide explanations for recommended items. In this paper, we conduct a systematical survey of knowledge graph-based recommender systems. We collect recently published papers in this field, and group them into three categories, i.e., embedding-based methods, connection-based methods, and propagation-based methods. Also, we further subdivide each category according to the characteristics of these approaches. Moreover, we investigate the proposed algorithms by focusing on how the papers utilize the knowledge graph for accurate and explainable recommendation. Finally, we propose several potential research directions in this field.
Qingyu Guo, Fuzhen Zhuang, Chuan Qin 0002, Hengshu Zhu, Xing Xie 0001, Hui Xiong 0001, Qing He 0003
IEEE Trans. Knowl. Data Eng.1
2021 Effects of Support-Seekers' Community Knowledge on Their Expressed Satisfaction with the Received Comments in Mental Health Communities
abstract
Online mental health communities (OMHCs) are prominent resources for improving people’s mental wellbeing. An immediate cue of such improvement is support-seekers’ satisfaction expressed in their replies to the received comments. However, the comments that seekers find satisfying may change with their community knowledge, e.g., measured by tenure and posting experience in that community. In this paper, we first model the amount of satisfaction conveyed in the support-seekers’ replies to the received comments. Then we quantitatively examine how seekers’ expressed satisfaction is affected by their community knowledge, sought and received support in an OMHC. Results show that support-seekers with more posting experience generally display less contentment to the received comments. Compared to newcomers, higher tenured members express less satisfaction when receiving informational support. We also found that support matching positively predicts seekers’ satisfaction regardless of their community knowledge. Our findings have implications for OMHCs to satisfy support-seekers through their community knowledge.
Zhenhui Peng, Xiaojuan Ma, Diyi Yang, Ka Wing Tsang, Qingyu Guo
CHI5
2020 Exploring the Effects of Technological Writing Assistance for Support Providers in Online Mental Health Community
abstract
Textual comments from peers with informational and emotional support are beneficial to members of online mental health communities (OMHCs). However, many comments are not of high quality in reality. Writing support technologies that assess (AS) the text or recommend (RE) writing examples on the fly could potentially help support providers to improve the quality of their comments. However, how providers perceive and work with such technologies are under-investigated. In this paper, we present a technological prototype MepsBot which offers providers in-situ writing assistance in either AS or RE mode. Results of a mixed-design study with 30 participants show that both types of MepsBots improve users' confidence in and satisfaction with their comments. The AS-mode MepsBot encourages users to refine expressions and is deemed easier to use, while the RE-mode one stimulates more support-related content re-editions. We report concerns on MepsBot and propose design considerations for writing support technologies in OMHCs.
Zhenhui Peng, Qingyu Guo, Ka Wing Tsang, Xiaojuan Ma
CHI2
2020 Contextual User Browsing Bandits for Large-Scale Online Mobile Recommendation
abstract
Online recommendation services recommend multiple commodities to users. Nowadays, a considerable proportion of users visit e-commerce platforms by mobile devices. Due to the limited screen size of mobile devices, positions of items have a significant influence on clicks: 1) Higher positions lead to more clicks for one commodity. 2) The ‘pseudo-exposure’ issue: Only a few recommended items are shown at first glance and users need to slide the screen to browse other items. Therefore, some recommended items ranked behind are not viewed by users and it is not proper to treat this kind of items as negative samples. While many works model the online recommendation as contextual bandit problems, they rarely take the influence of positions into consideration and thus the estimation of the reward function may be biased. In this paper, we aim at addressing these two issues to improve the performance of online mobile recommendation. Our contributions are four-fold. First, since we concern the reward of a set of recommended items, we model the online recommendation as a contextual combinatorial bandit problem and define the reward of a recommended set. Second, we propose a novel contextual combinatorial bandit method called UBM-LinUCB to address two issues related to positions by adopting the User Browsing Model (UBM), a click model for web search. Third, we provide a formal regret analysis and prove that our algorithm achieves sublinear regret independent of the number of items. Finally, we evaluate our algorithm on two real-world datasets by a novel unbiased estimator. An online experiment is also implemented in Taobao, one of the most popular e-commerce platforms in the world. Results on two CTR metrics show that our algorithm outperforms the other contextual bandit algorithms.
Bo An 0001, Yanghua Li, Haikai Chen, Qingyu Guo, Zhirong Wang
RecSys5
2019 On the Inducibility of Stackelberg Equilibrium for Security Games
abstract
Strong Stackelberg equilibrium (SSE) is the standard solution concept of Stackelberg security games. As opposed to the weak Stackelberg equilibrium (WSE), the SSE assumes that the follower breaks ties in favor of the leader and this is widely acknowledged and justified by the assertion that the defender can often induce the attacker to choose a preferred action by making an infinitesimal adjustment to her strategy. Unfortunately, in security games with resource assignment constraints, the assertion might not be valid; it is possible that the defender cannot induce the desired outcome. As a result, many results claimed in the literature may be overly optimistic. To remedy, we first formally define the utility guarantee of a defender strategy and provide examples to show that the utility of SSE can be higher than its utility guarantee. Second, inspired by the analysis of leader’s payoff by Von Stengel and Zamir (2004), we provide the solution concept called the inducible Stackelberg equilibrium (ISE), which owns the highest utility guarantee and always exists. Third, we show the conditions when ISE coincides with SSE and the fact that in general case, SSE can be extremely worse with respect to utility guarantee. Moreover, introducing the ISE does not invalidate existing algorithmic results as the problem of computing an ISE polynomially reduces to that of computing an SSE. We also provide an algorithmic implementation for computing ISE, with which our experiments unveil the empirical advantage of the ISE over the SSE.
Qingyu Guo, Jiarui Gan, Fei Fang 0001, Long Tran-Thanh, Milind Tambe, Bo An 0001
AAAI1
2019 Optimal Interdiction of Urban Criminals with the Aid of Real-Time Information
abstract
Most violent crimes happen in urban and suburban cities. With emerging tracking techniques, law enforcement officers can have real-time location information of the escaping criminals and dynamically adjust the security resource allocation to interdict them. Unfortunately, existing work on urban network security games largely ignores such information. This paper addresses this omission. First, we show that ignoring the real-time information can cause an arbitrarily large loss of efficiency. To mitigate this loss, we propose a novel NEtwork purSuiT game (NEST) model that captures the interaction between an escaping adversary and a defender with multiple resources and real-time information available. Second, solving NEST is proven to be NP-hard. Third, after transforming the non-convex program of solving NEST to a linear program, we propose our incremental strategy generation algorithm, including: (i) novel pruning techniques in our best response oracle; and (ii) novel techniques for mapping strategies between subgames and adding multiple best response strategies at one iteration to solve extremely large problems. Finally, extensive experiments show the effectiveness of our approach, which scales up to realistic problem sizes with hundreds of nodes on networks including the real network of Manhattan.
Youzhi Zhang 0001, Qingyu Guo, Bo An 0001, Long Tran-Thanh, Nicholas R. Jennings
AAAI2
2019 Manipulating a Learning Defender and Ways to Counteract
abstract
In Stackelberg security games when information about the attacker's payoffs is uncertain, algorithms have been proposed to learn the optimal defender commitment by interacting with the attacker and observing their best responses. In this paper, we show that, however, these algorithms can be easily manipulated if the attacker responds untruthfully. As a key finding, attacker manipulation normally leads to the defender learning a maximin strategy, which effectively renders the learning attempt meaningless as to compute a maximin strategy requires no additional information about the other player at all. We then apply a game-theoretic framework at a higher level to counteract such manipulation, in which the defender commits to a policy that specifies her strategy commitment according to the learned information. We provide a polynomial-time algorithm to compute the optimal such policy, and in addition, a heuristic approach that applies even when the attacker's payoff space is infinite or completely unknown. Empirical evaluation shows that our approaches can improve the defender's utility significantly as compared to the situation when attacker manipulation is ignored.
Jiarui Gan, Qingyu Guo, Long Tran-Thanh, Bo An 0001, Michael J. Wooldridge
NeurIPS2
2019 Securing the Deep Fraud Detector in Large-Scale E-Commerce Platform via Adversarial Machine Learning Approach
abstract
Fraud transactions are one of the major threats faced by online e-commerce platforms. Recently, deep learning based classifiers have been deployed to detect fraud transactions. Inspired by findings on adversarial examples, this paper is the first to analyze the vulnerability of deep fraud detector to slight perturbations on input transactions, which is very challenging since the sparsity and discretization of transaction data result in a non-convex discrete optimization. Inspired by the iterative Fast Gradient Sign Method (FGSM) for the L8 attack, we first propose the Iterative Fast Coordinate Method (IFCM) for discrete L1 and L2 attacks which is efficient to generate large amounts of instances with satisfactory effectiveness. We then provide two novel attack algorithms to solve the discrete optimization. The first one is the Augmented Iterative Search (AIS) algorithm, which repeatedly searches for effective “simple” perturbation. The second one is called the Rounded Relaxation with Reparameterization (R3), which rounds the solution obtained by solving a relaxed and unconstrained optimization problem with reparameterization tricks. Finally, we conduct extensive experimental evaluation on the deployed fraud detector in TaoBao, one of the largest e-commerce platforms in the world, with millions of real-world transactions. Results show that (i) The deployed detector is highly vulnerable to attacks as the average precision is decreased from nearly 90% to as low as 20% with little perturbations; (ii) Our proposed attacks significantly outperform the adaptions of the state-of-the-art attacks. (iii) The model trained with an adversarial training process is significantly robust against attacks and performs well on the unperturbed data.
Qingyu Guo, Zhao Li 0007, Bo An 0001, Pengrui Hui, Mengchen Zhao
WWW1
2017 Comparing Strategic Secrecy and Stackelberg Commitment in Security Games
abstract
The Strong Stackelberg Equilibrium (SSE) has drawn extensive attention recently in several security domains. However, the SSE concept neglects the advantage of defender's strategic revelation of her private information, and overestimates the observation ability of the adversaries. In this paper, we overcome these restrictions and analyze the tradeoff between strategic secrecy and commitment in security games. We propose a Disguised-resource Security Game (DSG) where the defender strategically disguises some of her resources. We compare strategic information revelation with public commitment and formally show that they have different advantages depending the payoff structure. To compute the Perfect Bayesian Equilibrium (PBE), several novel approaches are provided, including a novel algorithm based on support set enumeration, and an approximation algorithm for \epsilon-PBE. Extensive experimental evaluation shows that both strategic secrecy and Stackelberg commitment are critical measures in security domain, and our approaches can efficiently solve PBEs for realistic-sized problems.
Qingyu Guo, Bo An 0001, Branislav Bosanský, Christopher Kiekintveld
IJCAI1
2017 Playing Repeated Network Interdiction Games with Semi-Bandit Feedback
abstract
We study repeated network interdiction games with no prior knowledge of the adversary and the environment, which can model many real world network security domains. Existing works often require plenty of available information for the defender and neglect the frequent interactions between both players, which are unrealistic and impractical, and thus, are not suitable for our settings. As such, we provide the first defender strategy, that enjoys nice theoretical and practical performance guarantees, by applying the adversarial online learning approach. In particular, we model the repeated network interdiction game with no prior knowledge as an online linear optimization problem, for which a novel and efficient online learning algorithm, SBGA, is proposed, which exploits the unique semi-bandit feedback in network security domains. We prove that SBGA achieves sublinear regret against adaptive adversary, compared with both the best fixed strategy in hindsight and a near optimal adaptive strategy. Extensive experiments also show that SBGA significantly outperforms existing approaches with fast convergence rate.
Qingyu Guo, Bo An 0001, Long Tran-Thanh
IJCAI1
2016 Optimal Interdiction of Illegal Network Flow
Qingyu Guo, Bo An 0001, Yair Zick, Chunyan Miao
IJCAI1
2008 Broadcast Routing and Channel Selection in Multi-Radio Wireless Mesh Networks
abstract
In multi-radio wireless mesh networks, each node can be equipped with multiple network interface cards tuned to different channels. In this paper, we present a routing and channel selection algorithm for reducing the broadcast redundancy in multi-radio wireless mesh networks. In our approach, the concept of Relaying Channel Redundancy is proposed, which is the sum of the number of different channels selected by each forward node in a broadcast tree. Our aim is to build a broadcast tree with minimum Relaying Channel Redundancy. We prove that building such a broadcast tree is a NP-hard problem, and propose an approximate algorithm for it. Our algorithm has an approximation ratio of at most 20delta5+2, wheredeltais the number of available non-overlapping channels. Finally, the distributed implementation of our algorithm is also presented.
Kai Han 0003, Qingyu Guo, Mingjun Xiao
WCNC3