VLDB 2026 Research / reviewers in the wild / expert
Zhuojun Li
dblp:43/5826
· DBLP profile ↗
15ranked-venue papers
4as first author
13since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 6 · 1 first-author · 4 since 2021Human-computer interaction and ubiquitous computing · 6 · 2 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SGoT-R1: Social Graph of Thought Reasoning-Enhanced Multimodal Large Language Model for Harmful Meme DetectionabstractInternet memes serve as widely distributed multimodal social content that conveys complex ideas through metaphorical expressions, often containing harmful implications that make accurate harmful meme detection an important problem. Reasoning knowledge extracted from large language models plays a crucial role in recent advances in harmful meme detection. However, these methods only perform reasoning analysis on memes from a single opinion, ignoring that memes are essentially products of group consensus, where their true meaning interpretation highly depends on the collision and aggregation process of diverse user viewpoints. To address this problem, we propose a Social Graph of Thought Reasoning Enhancement (SGoTRE) framework for harmful meme detection. The SGoTRE contains three key steps: First, through multi-agent simulation technology, we obtain diverse chains of thought that represent the parsing logic of users from different backgrounds toward memes, authentically restoring the diversity characteristics of group cognition. Second, we construct a Social Graph of Thought (SGoT) that effectively integrates multi-chain reasoning processes and structurally expresses the consensus and diversity of viewpoints among users. Finally, we utilize the SGoT for cognitive distillation, internalizing multi-opinion reasoning logic into a single multimodal large model SGoT-R1 to achieve efficient and interpretable harmful meme detection. Experimental results show that SGoT-R1 significantly improves detection performance on mainstream datasets. Particularly on the most challenging FHM dataset, SGoT-R1 achieves an 8.9% improvement over state-of-the-art models. Xiuxian Wang, Yuting Su 0001, Wenhui Li 0001, Zhuojun Li, Anan Liu |
AAAI | 5 |
| 2026 | KeySense: LLM-Powered Hands-Down, Ten-Finger Typing on Commodity TouchscreensabstractExisting touchscreen software keyboards prevent users from resting their hands, forcing slow and fatiguing index-finger tapping (“chicken typing”) instead of familiar hands-down ten-finger typing. We present KeySense, a purely software solution that preserves physical keyboard motor skills. KeySense isolates intentional taps from resting-finger noise with cognitive–motor timing patterns, and then uses a fine-tuned LLM decoder to turn the resulting noisy letter sequence into the intended word. In controlled component tests, this decoder substantially outperforms 2 statistical baselines (top-1 accuracy 84.8% vs 75.7% and 79.3%). A 12-participant study shows clear ergonomic and performance benefits: compared with the conventional hover-style keyboard, users rated KeySense as markedly less physically demanding (NASA-TLX median 1.5 vs 4.0), and after brief practice, typed significantly faster (WPM 28.3 vs 26.2, p <0.01). These results indicate that KeySense enables accurate, efficient and comfortable ten-finger text entry on commodity touchscreens, without any extra hardware. Tony Li, Yan Ma 0006, Zhuojun Li, Chun Yu, I. V. Ramakrishnan, Xiaojun Bi 0001 |
CHI | 3 |
| 2026 | 3DRing: Enabling Low-Cost 3D Hand Position Tracking by Fusing Inertial and Low-Framerate Optical SensingabstractCurrent mobile hand tracking systems primarily rely on high-framerate (HFR) optical sensors to capture hand positions, resulting in high computational cost and limiting the applicability in end devices. We propose 3DRing, a 3D hand position tracking method that requires only low-framerate (LFR, <10 FPS) optical data and a single IMU ring. It consists of two stages: (1) a Deep Extended Kalman Filter module that predicts high-framerate hand positions from LFR optical measurements and a single IMU; (2) a Reinforcement Learning module that adaptively selects minimal keyframes for calibration, further reducing the average optical framerate. Using only 6.61 FPS optical data, 3DRing achieves an average real-time tracking error of 1.75 cm and an interaction efficiency of 86.0% in a 3D target selection task, compared to the 67 FPS hand tracking system of Meta Quest Pro, demonstrating a strong potential to reduce the reliance on optical data in mobile hand tracking tasks. Zhuojun Li, Chun Yu, Chang Liu 0150, Mingyuan Du, Weinan Shi, Yuanchun Shi |
CHI | 1 |
| 2026 | HiSync: Spatio-Temporally Aligning Hand Motion from Wearable IMU and On-Robot Camera for Command Source Identification in Long-Range HRIabstractLong-range Human-Robot Interaction (HRI) remains underexplored. Within it, Command Source Identification (CSI) — determining who issued a command — is especially challenging due to multi-user and distance-induced sensor ambiguity. We introduce HiSync, an optical-inertial fusion framework that treats hand motion as binding cues by aligning robot-mounted camera optical flow with hand-worn IMU signals. We first elicit a user-defined (N=12) gesture set and collect a multimodal command gesture dataset (N=38) in long-range multi-user HRI scenarios. Next, HiSync extracts frequency-domain hand motion features from both camera and IMU data, and a learned CSINet denoises IMU readings, temporally aligns modalities, and performs distance-aware multi-window fusion to compute cross-modal similarity of subtle, natural gestures, enabling robust CSI. In three-person scenes up to 34 m, HiSync achieves 92.32% CSI accuracy, outperforming the prior SOTA by 48.44%. HiSync is also validated on real-robot deployment. By making CSI reliable and natural, HiSync provides a practical primitive and design guidance for public-space HRI. Chun Yu, Borong Zhuang, Haopeng Jin, Qingyang Wan, Zhuojun Li, Zhoutong Ye, Chang Liu 0150, Weinan Shi, Yuanchun Shi |
CHI | 6 |
| 2026 | CFLight: Enhancing Safety with Traffic Signal Control through Counterfactual LearningabstractTraffic accidents result in millions of injuries and fatalities globally, with a significant number occurring at intersections each year. Traffic Signal Control (TSC) is an effective strategy for enhancing safety at these urban junctures. Despite the growing popularity of Reinforcement Learning (RL) methods in optimizing TSC, these methods often prioritize driving efficiency over safety, thus failing to address the critical balance between these two aspects. Additionally, these methods usually need more interpretability. CounterFactual (CF) learning is a promising approach for various causal analysis fields. In this study, we introduce a novel framework to improve RL for safety aspects in TSC. This framework introduces a novel method based on CF learning to address the question: ``What if, when an unsafe event occurs, we backtrack to perform alternative actions, and will this unsafe event still occur in the subsequent period?'' To answer this question, we propose a new structure causal model to predict the result after executing different actions, and we propose a new CF module that integrates with additional ``X'' modules to promote safe RL practices. Our new algorithm, CFLight, which is derived from this framework, effectively tackles challenging safety events and significantly improves safety at intersections through a near-zero collision control strategy. Through extensive numerical experiments on both real-world and synthetic datasets, we demonstrate that CFLight reduces collisions and improves overall traffic performance compared to conventional RL methods and the recent safe RL model. Moreover, our method represents a generalized and safe framework for RL methods, opening possibilities for applications in other domains. The data and code are available in the github https://github.com/AdvancedAI-ComplexSystem/SmartCity/tree/main/CFLight. Mingyuan Li 0006, Zhuojun Li, Xiao Liu 0037, Guangsheng Yu, Bo Du 0004, Jun Shen 0001, Qiang Wu 0010 |
KDD (1) | 3 |
| 2026 | The gains do not make up for the losses: a comprehensive evaluation for safety alignment of large language models via machine unlearningabstractAbstract Machine Unlearning (MU) has emerged as a promising technique for aligning large language models (LLMs) with safety requirements to steer them forgetting specific harmful contents. Despite the significant progress in previous studies, we argue that the current evaluation criteria, which solely focus on safety evaluation, are actually impractical and biased , leading to concerns about the true effectiveness of MU techniques. To address this, we propose to comprehensively evaluate LLMs after MU from three aspects: safety, over-safety, and general utility. Specifically, a novel benchmark M u B ench with 18 related datasets is first constructed, where the safety is measured with both vanilla harmful inputs and 10 types of jailbreak attacks. Furthermore, we examine whether MU introduces side effects, focusing on over-safety and utility-loss. Extensive experiments are performed on 3 popular LLMs with 7 recent MU methods. The results highlight a challenging trilemma in safety alignment without side effects, indicating that there is still considerable room for further exploration. M u B ench serves as a comprehensive benchmark, fostering future research on MU for safety alignment of LLMs. Weixiang Zhao, Yulin Hu, Xingyu Sui, Zhuojun Li, Yang Deng 0002, Bing Qin 0001, Wanxiang Che |
Frontiers Comput. Sci. | 4 |
| 2025 | NES-DM: Representative Attention-based CNN Network for Medical Image ClassificationabstractVarious lesions in different body parts have different sizes and, in particular, different representations, which leads to great challenges for medical image classification tasks. To avoid incorrect classification, it is critical to design a model that has the ability to extract descriptive and discriminatory features. The attention mechanism is a method that can utilize layer weight adjustment to enhance feature extraction capability for various deep learning models. In this study, an attention-based CNN framework, NES-DM, is proposed to adapt to different lesions. Specifically, for the channel attention, we propose to leverage a neuron energy separability module, which aims to aggregate more representative features in a channel; for the spatial attention, we design to decompose the feature map along with two dimensions and then input them into a multiple convolution module, respectively, which aims to reduce the feature redundancy by the decomposition and capture long-range dependencies of features by the multiple convolution. The effectiveness of the proposed NES-DM based CNN framework is evaluated on three medical image datasets and NES-DM achieves at most 7.8% improvement on sensitivity compared to baselines. Zhuojun Li, Lanjun Wang, Anan Liu |
ISCAS | 1 |
| 2025 | InterQuest: A Mixed-Initiative Framework for Dynamic User Interest Modeling in Conversational Search
Yuanxi Wang, Qingyang Wan, Zhuojun Li, Chun Yu, Weinan Shi, Yuanchun Shi |
UIST | 5 |
| 2025 | A parental emotion coaching dialogue assistant for better parent-child interaction
Weixiang Zhao, Shilong Wang 0003, Yanpeng Tong, Zhuojun Li, Chenxue Wang, Bing Qin 0001 |
Sci. China Inf. Sci. | 5 |
| 2024 | PoseAugment: Generative Human Pose Data Augmentation with Physical Plausibility for IMU-Based Motion Capture
Zhuojun Li, Chun Yu, Yuanchun Shi |
ECCV (32) | 1 |
| 2023 | Knowledge-Bridged Causal Interaction Network for Causal Emotion EntailmentabstractCausal Emotion Entailment aims to identify causal utterances that are responsible for the target utterance with a non-neutral emotion in conversations. Previous works are limited in thorough understanding of the conversational context and accurate reasoning of the emotion cause. To this end, we propose Knowledge-Bridged Causal Interaction Network (KBCIN) with commonsense knowledge (CSK) leveraged as three bridges. Specifically, we construct a conversational graph for each conversation and leverage the event-centered CSK as the semantics-level bridge (S-bridge) to capture the deep inter-utterance dependencies in the conversational context via the CSK-Enhanced Graph Attention module. Moreover, social-interaction CSK serves as emotion-level bridge (E-bridge) and action-level bridge (A-bridge) to connect candidate utterances with the target one, which provides explicit causal clues for the Emotional Interaction module and Actional Interaction module to reason the target emotion. Experimental results show that our model achieves better performance over most baseline models. Our source code is publicly available at https://github.com/circle-hit/KBCIN. Weixiang Zhao, Zhuojun Li, Bing Qin 0001 |
AAAI | 3 |
| 2023 | ResType: Invisible and Adaptive Tablet Keyboard Leveraging Resting FingersabstractText entry on tablet touchscreens is a basic need nowadays. Tablet keyboards require visual attention for users to locate keys, thus not supporting efficient touch typing. They also take up a large proportion of screen space, which affects the access to information. To solve these problems, we propose ResType, an adaptive and invisible keyboard on three-state touch surfaces (e.g. tablets with unintentional touch prevention). ResType allows users to rest their hands on it and automatically adapts the keyboard to the resting fingers. Thus, users do not need visual attention to locate keys, which supports touch typing. We quantitatively explored users’ resting finger patterns on ResType, based on which we proposed an augmented Bayesian decoding algorithm for ResType, with 96.3% top-1 and 99.0% top-3 accuracies. After a 5-day evaluation, ResType achieved 41.26 WPM, outperforming normal tablet keyboards by 13.5% and reaching 86.7% of physical keyboards. It solves the occlusion problem while maintaining comparable typing speed with current methods on visible tablet keyboards. Zhuojun Li, Chun Yu, Yizheng Gu, Yuanchun Shi |
CHI | 1 |
| 2021 | TypeBoard: Identifying Unintentional Touch on Pressure-Sensitive Touchscreen KeyboardsabstractText input is essential in tablet computer interaction. However, tablet software keyboards face the problem of misrecognizing unintentional touch, which affects efficiency and usability [29, 49]. In this paper, we proposed TypeBoard, a pressure-sensitive touchscreen keyboard that prevents unintentional touches. The TypeBoard allows users to rest their fingers on the touchscreen, which changes the user behavior: on average, users generate 40.83 unintentional touches every 100 keystrokes. The TypeBoard prevents unintentional touch with an accuracy of 98.88%. A typing study showed that the TypeBoard reduced fatigue (p < 0.005) and typing errors (p < 0.01), and improved the touchscreen keyboard’ typing speed by 11.78% (p < 0.005). As users could touch the screen without triggering responses, we added tactile landmarks on the TypeBoard, allowing users to locate the keys by the sense of touch. This feature further improves the typing speed, outperforming the ordinary tablet keyboard by 21.19% (p < 0.001). Results show that pressure-sensitive touchscreen keyboards can prevent unintentional touch, improving usability from many aspects, such as avoiding fatigue, reducing errors, and mediating touch typing on tablets. Yizheng Gu, Chun Yu, Xuanzhong Chen, Zhuojun Li, Yuanchun Shi |
UIST | 4 |
| 2017 | The Time Ontology of Allen's Interval AlgebraabstractAllen's interval algebra is a set of thirteen jointly exhaustive and pairwise disjoint binary relations representing temporal relationships between pairs of timeintervals. Despite widespread use, there is still the question of which time ontology actually underlies Allen's algebra. Early work specified a first-order ontology that can interpret Allen's interval algebra; in this paper, we identify the first-order ontology that is logically synonymous with Allen's interval algebra, so that there is a one-to-one correspondence between models of the ontology and solutions to temporal constraints that are specified using the temporal relations. We further prove a representation theorem for the ontology, thus characterizing its models up to isomorphism. Michael Grüninger, Zhuojun Li |
TIME | 2 |
| 2016 | Virus image classification using multi-scale completed local binary pattern features extracted from filtered images by multi-scale principal component analysis
Zhijie Wen, Zhuojun Li, Yaxin Peng, Shihui Ying |
Pattern Recognit. Lett. | 2 |