VLDB 2026 Research / reviewers in the wild / expert
Wei Li 0002
dblp:l/WeiLi2
· DBLP profile ↗
46ranked-venue papers
11as first author
25since 2021 · last 2026
0000-0001-7174-9228ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 24 · 7 first-author · 13 since 2021Human-computer interaction and ubiquitous computing · 18 · 4 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 14 · 4 first-author · 8 since 2021Databases, data management, data science and information retrieval · 3 · 1 since 2021Software engineering, systems software and programming languages · 2 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Multi-Granularity Semantic Revision for Large Language Model DistillationabstractXiaoyu Liu, Yun Zhang, Wei Li, Simiao Li, Xudong Huang, Hanting Chen, Yehui Tang, Jie Hu, Zhiwei Xiong, Yunhe Wang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Xiaoyu Liu 0006, Wei Li 0002, Simiao Li, Hanting Chen, Yehui Tang 0001, Jie Hu 0021, Zhiwei Xiong, Yunhe Wang 0001 |
ACL (1) | 3 |
| 2025 | GIM: A Million-scale Benchmark for Generative Image Manipulation Detection and LocalizationabstractThe extraordinary ability of generative models emerges as a new trend in image editing and generating realistic images, posing a serious threat to the trustworthiness of multimedia data and driving the research of image manipulation detection and location (IMDL). However, the lack of a large-scale data foundation makes the IMDL task unattainable. In this paper, we build a local manipulation data generation pipeline that integrates the powerful capabilities of SAM, LLM, and generative models. Upon this basis, we propose the GIM dataset, which has the following advantages: 1) Large scale, GIM includes over one million pairs of AI-manipulated images and real images. 2) Rich image content, GIM encompasses a broad range of image classes. 3) Diverse generative manipulation, the images are manipulated images with state-of-the-art generators and various manipulation tasks. The aforementioned advantages allow for a more comprehensive evaluation of IMDL methods, extending their applicability to diverse images. We introduce the GIM benchmark with two settings to evaluate existing IMDL methods. In addition, we propose a novel IMDL framework, termed GIMFormer, which consists of a ShadowTracer, Frequency-Spatial block (FSB), and a Multi-Window Anomalous Modeling (MWAM) module. Extensive experiments on the GIM demonstrate that GIMFormer surpasses the previous state-of-the-art approach on two different benchmarks. Yirui Chen, Wei Li 0002, Mingjian Zhu, Qiangyu Yan, Simiao Li, Hanting Chen, Hailin Hu 0002, Jie Yang 0002, Jie Hu 0021 |
AAAI | 4 |
| 2025 | Dynamic Contrastive Knowledge Distillation for Efficient Image RestorationabstractKnowledge distillation (KD) is a valuable yet challenging approach that enhances a compact student network by learning from a high-performance but cumbersome teacher model. However, previous KD methods for image restoration overlook the state of the student during the distillation, adopting a fixed solution space that limits the capability of KD. Additionally, relying solely on L1-type loss struggles to leverage the distribution information of images. In this work, we propose a novel dynamic contrastive knowledge distillation (DCKD) framework for image restoration. Specifically, we introduce dynamic contrastive regularization to perceive the student's learning state and dynamically adjust the distilled solution space using contrastive learning. Additionally, we also propose a distribution mapping module to extract and align the pixel-level category distribution of the teacher and student models. Note that the proposed DCKD is a structure-agnostic distillation framework, which can adapt to different backbones and can be combined with methods that optimize upper-bound constraints to further enhance model performance. Extensive experiments demonstrate that DCKD significantly outperforms the state-of-the-art KD methods across various image restoration tasks and backbones. Yunshuai Zhou, Junbo Qiao, Jincheng Liao, Wei Li 0002, Simiao Li, Jiao Xie, Yunhang Shen, Jie Hu 0021, Shaohui Lin |
AAAI | 4 |
| 2025 | Weakly Supervised Semantic Segmentation via Progressive Confidence Region ExpansionabstractWeakly supervised semantic segmentation (WSSS) has garnered considerable attention due to its effective reduction of annotation costs. Most approaches utilize Class Activation Maps (CAM) to produce pseudo-labels, thereby localizing target regions using only image-level annotations. However, the prevalent methods relying on vision transformers (ViT) encounter an "over-expansion" issue, i.e., CAM incorrectly expands high activation value from the target object to the background regions, as it is difficult to learn pixel-level local intrinsic inductive bias in ViT from weak supervisions. To solve this problem, we propose a Progressive Confidence Region Expansion (PCRE) framework for WSSS, it gradually learns a faithful mask over the target region and utilizes this mask to correct the confusion in CAM. PCRE has two key components: Confidence Region Mask Expansion (CRME) and Class-Prototype Enhancement (CPE). CRME progressively expands the mask in the small region with the highest confidence, eventually encompassing the entire target, thereby avoiding unintended coverage of background areas. CPE aims to enhance mask generation in CRME by leveraging the similarity between the learned, dataset-level class prototypes and patch features as supervision to optimize the mask output from CRME. Extensive experiments demonstrate that our method outperforms the existing single-stage and multi-stage approaches on the PASCAL VOC and MS COCO benchmark. Our code is available at https://github.com/xxf011/WSSS-PCRE. Xiangfeng Xu, Pinyi Zhang, Wenxuan Huang 0001, Yunhang Shen, Jingzhong Lin, Wei Li 0002, Gaoqi He, Jiao Xie, Shaohui Lin |
CVPR | 7 |
| 2025 | CBQ: Cross-Block Quantization for Large Language ModelsabstractPost-training quantization (PTQ) has played a pivotal role in compressing large language models (LLMs) at ultra-low costs. Although current PTQ methods have achieved promising results by addressing outliers and employing layer- or block-wise loss optimization techniques, they still suffer from significant performance degradation at ultra-low bits precision. To dissect this issue, we conducted an in-depth analysis of quantization errors specific to LLMs and surprisingly discovered that, unlike traditional sources of quantization errors, the growing number of model parameters, combined with the reduction in quantization bits, intensifies inter-layer and intra-layer dependencies, which severely impact quantization accuracy. This finding highlights a critical challenge in quantizing LLMs. To address this, we propose CBQ, a cross-block reconstruction-based PTQ method for LLMs. CBQ leverages a cross-block dependency to establish long-range dependencies across multiple blocks and integrates an adaptive LoRA-Rounding technique to manage intra-layer dependencies. To further enhance performance, CBQ incorporates a coarse-to-fine pre-processing mechanism for processing weights and activations. Extensive experiments show that CBQ achieves superior low-bit quantization (W4A4, W4A8, W2A16) and outperforms existing state-of-the-art methods across various LLMs and datasets. Notably, CBQ only takes 4.3 hours to quantize a weight-only quantization of a 4-bit LLAMA1-65B model, achieving a commendable trade off between performance and efficiency. Xiaoyu Liu 0006, Zhijun Tu, Wei Li 0002, Jie Hu 0021, Hanting Chen, Yehui Tang 0001, Zhiwei Xiong, Baoqun Yin, Yunhe Wang 0001 |
ICLR | 5 |
| 2025 | Knowledge Distillation with Multi-granularity Mixture of Priors for Image Super-ResolutionabstractKnowledge distillation (KD) is a promising yet challenging model compression approach that transmits rich learning representations from robust but resource-demanding teacher models to efficient student models. Previous methods for image super-resolution (SR) are often tailored to specific teacher-student architectures, limiting their potential for improvement and hindering broader applications. This work presents a novel KD framework for SR models, the multi-granularity Mixture of Priors Knowledge Distillation (MiPKD), which can be universally applied to a wide range of architectures at both feature and block levels. The teacher’s knowledge is effectively integrated with the student's feature via the Feature Prior Mixer, and the reconstructed feature propagates dynamically in the training phase with the Block Prior Mixer. Extensive experiments illustrate the significance of the proposed MiPKD technique. Simiao Li, Wei Li 0002, Hanting Chen, Wenjia Wang 0005, Bing-Yi Jing, Shaohui Lin, Jie Hu 0021 |
ICLR | 3 |
| 2025 | AugKD: Ingenious Augmentations Empower Knowledge Distillation for Image Super-ResolutionabstractKnowledge distillation (KD) compresses deep neural networks by transferring task-related knowledge from cumbersome pre-trained teacher models to more compact student models. However, vanilla KD for image super-resolution (SR) networks yields only limited improvements due to the inherent nature of SR tasks, where the outputs of teacher models are noisy approximations of high-quality label images. In this work, we show that the potential of vanilla KD has been underestimated and demonstrate that the ingenious application of data augmentation methods can close the gap between it and more complex, well-designed methods. Unlike conventional training processes typically applying image augmentations simultaneously to both low-quality inputs and high-quality labels, we propose AugKD utilizing unpaired data augmentations to 1) generate auxiliary distillation samples and 2) impose label consistency regularization. Comprehensive experiments show that the AugKD significantly outperforms existing state-of-the-art KD methods across a range of SR tasks. Wei Li 0002, Simiao Li, Hanting Chen, Zhijun Tu, Bing-Yi Jing, Shaohui Lin, Jie Hu 0021, Wenjia Wang 0005 |
ICLR | 2 |
| 2025 | DCS-RISR: Dynamic channel splitting for efficient real-world image super-resolution
Junbo Qiao, Shaohui Lin, Yulun Zhang 0001, Wei Li 0002, Jie Hu 0021, Gaoqi He, Changbo Wang, Lizhuang Ma |
Neural Networks | 4 |
| 2025 | Hi-Mamba: Hierarchical Mamba for Efficient Image Super-ResolutionabstractDespite Transformers have achieved significant success in low-level vision tasks, they are constrained by computing self-attention with a quadratic complexity and limited-size windows. This limitation results in a lack of global receptive field across the entire image. Recently, State Space Models (SSMs) have gained widespread attention due to their global receptive field and linear complexity with respect to input length. However, integrating SSMs into low-level vision tasks presents two major challenges: 1) Relationship degradation of long-range tokens with a long-range forgetting problem by encoding pixel-by-pixel high-resolution images. 2) Significant redundancy in the existing multi-direction scanning strategy. To this end, we propose Hi-Mamba for image super-resolution (SR) to address these challenges, which unfolds the image with only a single scan. Specifically, the Global Hierarchical Mamba Block (GHMB) enables token interactions across the entire image, providing a global receptive field while leveraging a multi-scale structure to facilitate long-range dependency learning. Additionally, the Direction Alternation Module (DAM) adjusts the scanning patterns of GHMB across different layers to enhance spatial relationship modeling. Extensive experiments demonstrate that our Hi-Mamba achieves 0.2-0.27dB PSNR gains on the Urban100 dataset across different scaling factors compared to the state-of-the-art MambaIRv2 for SR. Moreover, our lightweight Hi-Mamba also outperforms lightweight SRFormer by 0.39dB PSNR for $\times 2$ SR. Junbo Qiao, Jincheng Liao, Wei Li 0002, Yulun Zhang 0001, Jiao Xie, Jie Hu 0021, Shaohui Lin |
IEEE Trans. Image Process. | 3 |
| 2024 | Distilling Semantic Priors from SAM to Efficient Image Restoration ModelsabstractIn image restoration (IR), leveraging semantic priors from segmentation models has been a common approach to improve performance. The recent segment anything model (SAM) has emerged as a powerful tool for extracting advanced semantic priors to enhance IR tasks. However, the computational cost of SAM is prohibitive for IR, compared to existing smaller IR models. The incorporation of SAMfor extracting semantic priors considerably hampers the model inference efficiency. To address this issue, we propose a general framework to distill SAM's semantic knowledge to boost exiting IR models without interfering with their inference process. Specifically, our proposed framework consists of the semantic priors fusion (SPF) scheme and the semantic priors distillation (SPD) scheme. SPF fuses two kinds of information between the restored image predicted by the original IR model and the semantic mask predicted by SAM for the refined restored image. SPD leverages a self-distillation manner to distill the fused semantic priors to boost the performance of original IR models. Additionally, we design a semantic-guided relation (SGR) module for SPD, which ensures semantic feature representation space consistency to fully distill the priors. We demonstrate the effectiveness of our framework across multiple IR models and tasks, including deraining, deblurring, and denoising. Xiaoyu Liu 0006, Wei Li 0002, Hanting Chen, Junchao Liu, Jie Hu 0021, Zhiwei Xiong, Chun Yuan 0003, Yunhe Wang 0001 |
CVPR | 3 |
| 2024 | PQ-SAM: Post-training Quantization for Segment Anything Model
Xiaoyu Liu 0006, Yuanyuan Xi, Wei Li 0002, Zhijun Tu, Jie Hu 0021, Hanting Chen, Baoqun Yin, Zhiwei Xiong |
ECCV (10) | 5 |
| 2024 | Wavelet-based network for high dynamic range imagingabstractHigh dynamic range (HDR) imaging from multiple low dynamic range (LDR) images has been suffering from ghosting artifacts caused by scene and objects motion. Existing methods, such as optical flow based and end-to-end deep learning based solutions, are error-prone either in detail restoration or ghosting artifacts removal. Comprehensive empirical evidence shows that ghosting artifacts caused by large foreground motion are mainly low-frequency signals and the details are mainly high-frequency signals. In this work, we propose a novel frequency-guided end-to-end deep neural network (FHDRNet) to conduct HDR fusion in the frequency domain, and Discrete Wavelet Transform (DWT) is used to decompose inputs into different frequency bands. The low-frequency signals are used to avoid specific ghosting artifacts, while the high-frequency signals are used for preserving details. Using a U-Net as the backbone, we propose two novel modules: merging module and frequency-guided upsampling module. The merging module applies the attention mechanism to the low-frequency components to deal with the ghost caused by large foreground motion. The frequency-guided upsampling module reconstructs details from multiple frequency-specific components with rich details. In addition, a new RAW dataset is created for training and evaluating multi-frame HDR imaging algorithms in the RAW domain. Extensive experiments are conducted on public datasets and our RAW dataset, showing that the proposed FHDRNet achieves state-of-the-art performance. Tianhong Dai, Wei Li 0002, Xilei Cao, Jianzhuang Liu, Xu Jia 0012, Ales Leonardis, Youliang Yan, Shanxin Yuan |
Comput. Vis. Image Underst. | 2 |
| 2023 | T-Force: Exploring the Use of Typing Force for Three State Virtual KeyboardsabstractThree state virtual keyboards which differentiate contact events between released, touched, and pressed states have the potential to improve overall typing experience and reduce the gap between virtual keyboards and physical keyboards. Incorporating force sensitivity, three-state virtual keyboards can utilize a force threshold to better classify a contact event. However, our limited knowledge of how force plays a role during typing on virtual keyboards limits further progress. Through a series of studies we observe that using a uniform threshold is not an optimal approach. Furthermore, the force being applied while typing varies significantly across the keys and among participants. As such, we propose three different approaches to further improve the uniform threshold. We show that a carefully selected non-uniform threshold function could be sufficient in delineating typing events on a three-state keyboard. Finally, we conclude our work with lessons learned, suggestion for future improvements, and comparisons with current methods available. Shariff A. M. Faleel, Yishuo Liu, Roya Allison Cody, Bradley Rey, Linghao Du, Jiangyue Yu, Da-Yuan Huang, Pourang Irani, Wei Li 0002 |
CHI | 9 |
| 2023 | Evaluating Across-Hinge Dragging with Pen and Touch on Curved and Foldable DisplaysabstractFoldable touch screens are increasingly popular, but little research has explored how the hinge impacts usability and performance. We evaluate across- and along-hinge drag gestures on a series of prototypes emulating foldable all-screen laptops with a curved hinge radius ranging from 1mm to 24mm. Results show that using a large 24mm hinge radius instead of a small 1mm hinge radius can decrease drag time by 13% and movement variability by 7% for touch input. However, hinge radius had no effect on performance for pen input. Further, we found that dragging along the hinge was up to 30% faster than dragging across the hinge, especially when dragging across at an acute angle to the hinge. Using these results, we demonstrate use cases for across- and along-hinge gestures. Our findings provide guidance for hardware and interaction designers seeking to create foldable touchscreen devices and their accompanying software. Graeme Zinck, Roya Allison Cody, Che Yan, Da-Yuan Huang, Wei Li 0002, Daniel Vogel 0001 |
CHI | 5 |
| 2023 | RefSR-NeRF: Towards High Fidelity and Super Resolution View SynthesisabstractWe present Reference-guided Super-Resolution Neural Radiance Field (RefSR-NeRF) that extends NeRF to super resolution and photorealistic novel view synthesis. Despite NeRF's extraordinary success in the neural rendering field, it suffers from blur in high resolution rendering because its inherent multilayer perceptron struggles to learn high frequency details and incurs a computational explosion as resolution increases. Therefore, we propose RefSR-NeRF, an end-to-end framework that first learns a low resolution NeRF representation, and then reconstructs the high frequency details with the help of a high resolution reference image. We observe that simply introducing the pre-trained models from the literature tends to produce unsatisfied artifacts due to the divergence in the degradation model. To this end, we design a novel lightweight RefSR model to learn the inverse degradation process from NeRF renderings to target HR ones. Extensive experiments on multiple benchmarks demonstrate that our method exhibits an impressive trade-off among rendering quality, speed, and memory usage, outperforming or on par with NeRF and its variants while being$52\times$speedup with minor extra memory usage. Code will be available at: Mindspore and Pytorch Wei Li 0002, Jie Hu 0021, Hanting Chen, Yunhe Wang 0001 |
CVPR | 2 |
| 2023 | GenImage: A Million-Scale Benchmark for Detecting AI-Generated ImageabstractThe extraordinary ability of generative models to generate photographic images has intensified concerns about the spread of disinformation, thereby leading to the demand for detectors capable of distinguishing between AI-generated fake images and real images. However, the lack of large datasets containing images from the most advanced image generators poses an obstacle to the development of such detectors. In this paper, we introduce the GenImage dataset, which has the following advantages: 1) Plenty of Images, including over one million pairs of AI-generated fake images and collected real images. 2) Rich Image Content, encompassing a broad range of image classes. 3) State-of-the-art Generators, synthesizing images with advanced diffusion models and GANs. The aforementioned advantages allow the detectors trained on GenImage to undergo a thorough evaluation and demonstrate strong applicability to diverse images. We conduct a comprehensive analysis of the dataset and propose two tasks for evaluating the detection method in resembling real-world scenarios. The cross-generator image classification task measures the performance of a detector trained on one generator when tested on the others. The degraded image classification task assesses the capability of the detectors in handling degraded images such as low-resolution, blurred, and compressed images. With the GenImage dataset, researchers can effectively expedite the development and evaluation of superior AI-generated image detectors in comparison to prevailing methodologies. Mingjian Zhu, Hanting Chen, Qiangyu Yan, Guanyu Lin, Wei Li 0002, Zhijun Tu, Hailin Hu 0002, Jie Hu 0021, Yunhe Wang 0001 |
NeurIPS | 6 |
| 2023 | Modeling User Reviews through Bayesian Graph Attention Networks for RecommendationabstractRecommender systems relieve users from cognitive overloading by predicting preferred items for users. Due to the complexity of interactions between users and items, graph neural networks (GNN) use graph structures to effectively model user–item interactions. However, existing GNN approaches have the following limitations: (1) User reviews are not adequately modeled in graphs. Therefore, user preferences and item properties that are described in user reviews are lost for modeling users and items; and (2) GNNs assume deterministic relations between users and items, which lack the stochastic modeling to estimate the uncertainties in neighbor relations. To mitigate the limitations, we build tripartite graphs to model user reviews as nodes that connect with users and items. We estimate neighbor relations with stochastic variables and propose a Bayesian graph attention network (i.e., ContGraph) to accurately predict user ratings. ContGraph incorporates the prior knowledge of user preferences to regularize the posterior inference of attention weights. Our experimental results show that ContGraph significantly outperforms 13 state-of-the-art models and improves the best performing baseline (i.e., ANR) by 5.23% on 25 datasets in the five-core version. Moreover, we show that correctly modeling the semantics of user reviews in graphs can help express the semantics of users and items. Yu Zhao 0041, Qiang Xu 0005, Ying Zou 0001, Wei Li 0002 |
ACM Trans. Inf. Syst. | 4 |
| 2022 | Switching Between Standard Pointing Methods with Current and Emerging Computer Form FactorsabstractWe investigate performance characteristics when switching between four pointing methods: absolute touch, absolute pen, relative mouse, and relative trackpad. The established “subtraction method” protocol used in mode-switching studies is extended to test pairs of methods and accommodate switch direction, multiple baselines, and controlling relative cursor position. A first experiment examines method switching on and around the horizontal surface of a tablet. Results find switching between pen and touch is fastest, and switching between relative and absolute methods incurs additional time penalty. A second experiment expands the investigation to an emerging foldable all-screen laptop form factor where switching also occurs on an angled surface and along a smoothly curved hinge. Results find switching between trackpad and touch is fastest, with all switching times generally higher. Our work contributes missing empirical evidence for switching performance using modern input methods, and our results can inform interaction design for current and emerging device form factors. Margaret Jean Foley, Quentin Roy, Da-Yuan Huang, Wei Li 0002, Daniel Vogel 0001 |
CHI | 4 |
| 2021 | Elbow-Anchored Interaction: Designing Restful Mid-Air InputabstractWe designed a mid-air input space for restful interactions on the couch. We observed people gesturing in various postures on a couch and found that posture affects the choice of arm motions when no constraints are imposed by a system. Study participants that sat with the arm rested were more likely to use the forearm and wrist, as opposed to the whole arm. We investigate how a spherical input space, where forearm angles are mapped to screen coordinates, can facilitate restful mid-air input in multiple postures. We present two controlled studies. In the first, we examine how a spherical space compares with a planar space in an elbow-anchored setup, with a shoulder-level input space as baseline. In the second, we examine the performance of a spherical input space in four common couch postures that set unique constraints to the arm. We observe that a spherical model that captures forearm movement facilitates comfortable input across different seated postures. Rafael Veras, Gaganpreet Singh, Farzin Farhadi-Niaki, Ritesh Udhani, Parth Pradeep Patekar, Wei Zhou 0021, Pourang Irani, Wei Li 0002 |
CHI | 8 |
| 2021 | Users, Tasks, and Conversational Agents: A Personality StudyabstractConversational Agents (CA) have become one of the common user interfaces in many online domains. In this paper, we ask whether users have a preference about the personality of CAs, and whether this preference changes depending on the length and type of the tasks CAs are used for. In an online study (N = 410), we investigated three different CA personalities (introvert, extrovert, and non-personified) in four different tasks with different natures and lengths (teaching, booking, todo, and weather). Most of the participants preferred to interact with a conversational agent (introvert or extrovert) as opposed to a non-personified interface, regardless of their own personality. Results suggested that this preference may be task dependent: when CA’s goal was to provide information, participants preferred an extrovert agent. We did not observe a difference between the preference for introvert and extrovert agents when the task’s goal was to complete an assignment. Quentin Roy, Moojan Ghafurian, Wei Li 0002, Jesse Hoey |
HAI | 3 |
| 2021 | Empirical Evaluation of Moving Target Selection in Virtual Reality Using Egocentric Metaphors
Yuan Chen 0011, Junwei Sun 0001, Qiang Xu 0005, Edward Lank, Pourang Irani, Wei Li 0002 |
INTERACT (4) | 6 |
| 2021 | Global Scene Filtering, Exploration, and Pointing in Occluded Virtual Space
Yuan Chen 0011, Junwei Sun 0001, Qiang Xu 0005, Edward Lank, Pourang Irani, Wei Li 0002 |
INTERACT (5) | 6 |
| 2021 | Leveraging CD Gain for Precise Barehand Video Timeline Browsing on Smart Displays
Futian Zhang, Sachi Mizobuchi, Wei Zhou 0021, Taslim Arefin Khan, Wei Li 0002, Edward Lank |
INTERACT (4) | 5 |
| 2021 | ARO: Exploring the Design of Smart-Ring Interactions for Encumbered HandsabstractFingertip computing has seen increased interest through miniaturized smart-rings for augmenting digital peripherals. One key advantages of such always-available input devices is the non-necessity to hold a device for interaction, as it remains affixed to a finger for access when needed. Such a wearable device makes it possible to interaction with content even when the hand is encumbered, by grasping or holding objects. Our investigation aims at understanding the properties of this fundamental smart-ring advantage. We designed a smart-ring prototype, ARO (in-Air, on-Ring, on-Object interaction), which facilitates input while grasping objects. To better identify interaction possibilities, we present the results of an elicitation study through which we grouped various forms of micro-gestures possible with ARO while holding objects under different grasp requirements. We then explored the ability for users to perform different navigation tasks (i.e. zooming and panning) using the smart-ring with encumbered hands. In our studies, users were most efficient when using either In-air or On-ring interactions, in comparisons to gestures detected On-object. Furthermore, In-air was the most preferred by our participants. Based on our findings, we conclude with recommendations for the design of future smart-rings and fingertip devices at large, to allow efficient interaction while hands are encumbered. Sandra Bardot, Surya Rawat, Duy Thai Nguyen, Sawyer Rempel, Huizhe Zheng, Bradley Rey, Jun Li 0067, Kevin Fan, Da-Yuan Huang, Wei Li 0002, Pourang Irani |
MobileHCI | 10 |
| 2021 | HPUI: Hand Proximate User Interfaces for One-Handed Interactions on Head Mounted DisplaysabstractWe explore the design of Hand Proximate User Interfaces (HPUIs) for head-mounted displays (HMDs) to facilitate near-body interactions with the display directly projected on, or around the user's hand. We focus on single-handed input, while taking into consideration the hand anatomy which distorts naturally when the user interacts with the display. Through two user studies, we explore the potential for discrete as well as continuous input. For discrete input, HPUIs favor targets that are directly on the fingers (as opposed to off-finger) as they offer tactile feedback. We demonstrate that continuous interaction is also possible, and is as effective on the fingers as in the off-finger space between the index finger and thumb. We also find that with continuous input, content is more easily controlled when the interaction occurs in the vertical or horizontal axes, and less with diagonal movements. We conclude with applications and recommendations for the design of future HPUIs. Shariff A. M. Faleel, Michael Gammon, Kevin Fan, Da-Yuan Huang, Wei Li 0002, Pourang Irani |
IEEE Trans. Vis. Comput. Graph. | 5 |
| 2020 | Weight Excitation: Built-in Attention Mechanisms in Convolutional Neural Networks
Niamul Quader, Md Mafijul Islam Bhuiyan, Juwei Lu, Peng Dai 0002, Wei Li 0002 |
ECCV (30) | 5 |
| 2020 | Towards Efficient Coarse-to-Fine Networks for Action and Gesture Recognition
Niamul Quader, Juwei Lu, Peng Dai 0002, Wei Li 0002 |
ECCV (30) | 4 |
| 2020 | Tent Mode Interactions: Exploring Collocated Multi-User Interaction on a Foldable DeviceabstractFoldable handheld displays have the potential to offer a rich interaction space, particularly as they fold into a convex form factor, for collocated multi-user interactions. In this paper, we explore Tent mode, a convex configuration of a foldable device partitioned into a primary and a secondary display, as well as a tertiary, Edge display that sits at the intersection of the two. We specifically explore the design space for a wide range of scenarios, such as co-browsing a gallery or co-planning a trip. Through a first collection of interviews, end-users identified a suite of apps that could leverage Tent mode for multi-user interactions. Based on these results we propose an interaction design space that builds on unique Tent mode properties, such as folding, flattening or tilting the device, and the interplay between the three sub-displays. We examine how end-users exploit this rich interaction space when presented with a set of collaborative tasks through a user study, and elicit potential interaction techniques. We implemented these interaction techniques and report on the preliminary user feedback we collected. Finally, we discuss the design implications for collocated interaction in Tent mode configurations. Gazelle Saniee-Monfared, Kevin Fan, Qiang Xu 0005, Sachi Mizobuchi, Lewis Zhou, Pourang Irani, Wei Li 0002 |
MobileHCI | 7 |
| 2019 | Exploring Cross-Modal Training via Touch to Learn a Mid-Air Marking Menu Gesture SetabstractWhile mid-air gestures are an attractive modality with an extensive research history, one challenge with their usage is that the gestures are not self-revealing. Scaffolding techniques to teach these gestures are difficult to implement since the input device, e.g. a hand, wand or arm, cannot present the gestures to the user. In contrast, for touch gestures, feedforward mechanisms (such as Marking Menus or OctoPocus) have been shown to effectively support user awareness and learning. In this paper, we explore whether touch gesture input can be leveraged to teach users to perform mid-air gestures. We show that marking menu touch gestures transfer directly to knowledge of mid-air gestures, allowing performance of these gestures without intervention. We argue that cross-modal learning can be an effective mechanism for introducing users to mid-air gestural input. Jay Henderson, Sachi Mizobuchi, Wei Li 0002, Edward Lank |
MobileHCI | 3 |
| 2019 | Hand-Over-Face Input Sensing for Interaction with Smartphones through the Built-in CameraabstractThis paper proposes using face as a touch surface and employing hand-over-face (HOF) gestures as a novel input modality for interaction with smartphones, especially when touch input is limited. We contribute InterFace, a general system framework that enables the HOF input modality using advanced computer vision techniques. As an examplar of the usage of this framework, we demonstrate the feasibility and usefulness of HOF with an Android application for improving single-user and group selfie-taking experience through providing appearance customization in real-time. In a within-subjects study comparing HOF against touch input for single-user interaction, we found that HOF input led to significant improvements in accuracy and perceived workload, and was preferred by the participants. Qualitative results of an observational study also demonstrated the potential of HOF input modality to improve the user experience in multi-user interactions. Based on the lessons learned from our studies, we propose a set of potential applications of HOF to support smartphone interaction. We envision that the affordances provided by the this modality can expand the mobile interaction vocabulary and facilitate scenarios where touch input is limited or even not possible. Mona Hosseinkhani Loorak, Wei Zhou 0021, Ha Trinh, Jian Zhao 0010, Wei Li 0002 |
MobileHCI | 5 |
| 2016 | Four-Bar Linkage Synthesis Using Non-convex Optimization
Vincent Goulet, Wei Li 0002, Hyunmin Cheong, Francesco Iorio, Claude-Guy Quimper |
CP | 2 |
| 2016 | Convolutional Neural Networks for Steady Flow ApproximationabstractIn aerodynamics related design, analysis and optimization problems, flow fields are simulated using computational fluid dynamics (CFD) solvers. However, CFD simulation is usually a computationally expensive, memory demanding and time consuming iterative process. These drawbacks of CFD limit opportunities for design space exploration and forbid interactive design. We propose a general and flexible approximation model for real-time prediction of non-uniform steady laminar flow in a 2D or 3D domain based on convolutional neural networks (CNNs). We explored alternatives for the geometry representation and the network architecture of CNNs. We show that convolutional neural networks can estimate the velocity field two orders of magnitude faster than a GPU-accelerated CFD solver and four orders of magnitude faster than a CPU-based CFD solver at a cost of a low error rate. This approach can provide immediate feedback for real-time design iterations at the early stage of design. Compared with existing approximation models in the aerodynamics domain, CNNs enable an efficient estimation for the entire velocity field. Furthermore, designers and engineers can directly apply the CNN approximation model in their design space exploration algorithms without training extra lower-dimensional surrogate models. Wei Li 0002, Francesco Iorio |
KDD | 2 |
| 2016 | Automated transformation of design text ROM diagram into SysML models
Hyunmin Cheong, Wei Li 0002, Yong Zeng 0004, Francesco Iorio |
Adv. Eng. Informatics | 3 |
| 2014 | Deploying CommunityCommands: A Software Command Recommender System Case StudyabstractIn 2009 we presented the idea of using collaborative filtering within a complex software application to help users learn new and relevant commands (Matejka et al. 2009). This project continued to evolve and we explored the design space of a contextual software command recommender system and completed a four-week user study (Li et al. 2011). We then expanded the scope of our project by implementing CommunityCommands, a fully functional and deployable recommender system. CommunityCommands was made available as a publically available plug-in download for Autodesk‟s flagship software application AutoCAD. During a one-year period, the recommender system was used by more than 1100 AutoCAD users. In this paper, we present our system usage data and payoff. We also provide an in-depth discussion of the challenges and design issues associated with developing and deploying the front end AutoCAD plug-in and its back end system. This includes a detailed description of the issues surrounding cold start and privacy. We also discuss how our practical system architecture was designed to leverage Autodesk‟s existing Customer Involvement Program (CIP) data to deliver in-product contextual recommendations to endusers. Our work sets important groundwork for the future development of recommender systems within the domain of end-user software learning assistance. Wei Li 0002, Justin Matejka, Tovi Grossman, George W. Fitzmaurice |
AAAI | 1 |
| 2014 | CADament: a gamified multiplayer software tutorial systemabstractWe present CADament, a gamified multiplayer tutorial system for learning AutoCAD. Compared with existing gamified software tutorial systems, CADament generates engaging learning experience through competitions. We investigate two variations of our game, where over-the-shoulder learning was simulated by providing viewports into other player's screens. We introduce an empirical lab study methodology where participants compete with one another, and we study knowledge transfer effects by tracking the migration of strategies between players during the study session. Our study shows that CADament has an advantage over pre-authored tutorials for improving learners' performance, increasing motivation, and stimulating knowledge transfer. Wei Li 0002, Tovi Grossman, George W. Fitzmaurice |
CHI | 1 |
| 2013 | TutorialPlan: Automated Tutorial Generation from CAD Drawings
Wei Li 0002, Yuanlin Zhang 0002, George W. Fitzmaurice |
IJCAI | 1 |
| 2012 | GamiCAD: a gamified tutorial system for first time autocad usersabstractWe present GamiCAD, a gamified in-product, interactive tutorial system for first time AutoCAD users. We introduce a software event driven finite state machine to model a user's progress through a tutorial, which allows the system to provide real-time feedback and recognize success and failures. GamiCAD provides extensive real-time visual and audio feedback that has not been explored before in the context of software tutorials. We perform an empirical evaluation of GamiCAD, comparing it to an equivalent in-product tutorial system without the gamified components. In an evaluation, users using the gamified system reported higher subjective engagement levels and performed a set of testing tasks faster with a higher completion ratio. Wei Li 0002, Tovi Grossman, George W. Fitzmaurice |
UIST | 1 |
| 2011 | Searching for software learning resources using application contextabstractUsers of complex software applications frequently need to consult documentation, tutorials, and support resources to learn how to use the software and further their understand-ing of its capabilities. Existing online help systems provide limited context awareness through "what's this?" and simi-lar techniques. We examine the possibility of making more use of the user's current context in a particular application to provide useful help resources. We provide an analysis and taxonomy of various aspects of application context and how they may be used in retrieving software help artifacts with web browsers, present the design of a context-aware augmented web search system, and describe a prototype implementation and initial user study of this system. We conclude with a discussion of open issues and an agenda for further research. Michael D. Ekstrand, Wei Li 0002, Tovi Grossman, Justin Matejka, George W. Fitzmaurice |
UIST | 2 |
| 2011 | TwitApp: in-product micro-blogging for design sharingabstractWe describe TwitApp, an enhanced micro-blogging system integrated within AutoCAD for design sharing. TwitApp integrates rich content and still keeps the sharing transaction cost low. In TwitApp, tweets are organized by their project, and users can follow or unfollow each individual project. We introduce the concept of automatic tweet drafting and other novel features such as enhanced real-time search and integrated live video streaming. The TwitApp system leverages the existing Twitter micro-blogging system. We also contribute a study which provides insights on these concepts and associated designs, and demonstrates potential user excitement of such tools. Wei Li 0002, Tovi Grossman, Justin Matejka, George W. Fitzmaurice |
UIST | 1 |
| 2011 | Exploiting Structure in Weighted Model Counting Approaches to Probabilistic InferenceabstractPrevious studies have demonstrated that encoding a Bayesian network into a SAT formula and then performing weighted model counting using a backtracking search algorithm can be an effective method for exact inference. In this paper, we present techniques for improving this approach for Bayesian networks with noisy-OR and noisy-MAX relations---two relations that are widely used in practice as they can dramatically reduce the number of probabilities one needs to specify. In particular, we present two SAT encodings for noisy-OR and two encodings for noisy-MAX that exploit the structure or semantics of the relations to improve both time and space efficiency, and we prove the correctness of the encodings. We experimentally evaluated our techniques on large-scale real and randomly generated Bayesian networks. On these benchmarks, our techniques gave speedups of up to two orders of magnitude over the best previous approaches for networks with noisy-OR/MAX relations and scaled up to larger networks. As well, our techniques extend the weighted model counting approach for exact inference to networks that were previously intractable for the approach. Wei Li 0002, Pascal Poupart, Peter van Beek |
J. Artif. Intell. Res. | 1 |
| 2011 | Design and evaluation of a command recommendation system for software applicationsabstractWe examine the use of modern recommender system technology to aid command awareness in complex software applications. We first describe our adaptation of traditional recommender system algorithms to meet the unique requirements presented by the domain of software commands. A user study showed that our item-based collaborative filtering algorithm generates 2.1 times as many good suggestions as existing techniques. Motivated by these positive results, we propose a design space framework and its associated algorithms to support both global and contextual recommendations. To evaluate the algorithms, we developed the CommunityCommands plug-in for AutoCAD. This plug-in enabled us to perform a 6-week user study of real-time, within-application command recommendations in actual working environments. We report and visualize command usage behaviors during the study, and discuss how the recommendations affected users behaviors. In particular, we found that the plug-in successfully exposed users to new commands, as unique commands issued significantly increased. Wei Li 0002, Justin Matejka, Tovi Grossman, Joseph A. Konstan, George W. Fitzmaurice |
ACM Trans. Comput. Hum. Interact. | 1 |
| 2009 | CommunityCommands: command recommendations for software applicationsabstractWe explore the use of modern recommender system technology to address the problem of learning software applications. Before describing our new command recommender system, we first define relevant design considerations. We then discuss a 3 month user study we conducted with professional users to evaluate our algorithms which generated customized recommendations for each user. Analysis shows that our item-based collaborative filtering algorithm generates 2.1 times as many good suggestions as existing techniques. In addition we present a prototype user interface to ambiently present command recommendations to users, which has received promising initial user feedback. Justin Matejka, Wei Li 0002, Tovi Grossman, George W. Fitzmaurice |
UIST | 2 |
| 2008 | Exploiting Causal Independence Using Weighted Model Counting
Wei Li 0002, Pascal Poupart, Peter van Beek |
AAAI | 1 |
| 2006 | Performing Incremental Bayesian Inference by Dynamic Model Counting
Wei Li 0002, Peter van Beek, Pascal Poupart |
AAAI | 1 |
| 2004 | A Hypergraph Separator Based Variable Ordering Heuristic for Solving Real World SAT
Wei Li 0002 |
CP | 1 |
| 2004 | Guiding Real-World SAT Solving with Dynamic Hypergraph Separator DecompositionabstractThe general solution of satisfiability problems is NP-complete. Although state-of-the-art SAT solvers can efficiently obtain the solutions of many real-world instances, there are still a large number of real-world SAT families which cannot be solved in reasonable time. Much effort has been spent to take advantage of the internal structure of SAT instances. Existing decomposition techniques are based on preprocessing the static structure of the original problem. We present a dynamic decomposition method based on hypergraph separators. Integrating the separator decomposition into the variable ordering of a modern SAT solver leads to speedups on large real-world satisfiability problems. Compared with a static decomposition based variable ordering, such as Dtree (Huang and Darwiche, 2003), our approach does not need time to construct the full tree decomposition, which sometimes needs more time than the solving process itself. Our primary focus is to achieve speedups on large real-world satisfiability problems. Our results show that the new solver often outperforms both regular zChaff and zChaff integrated with Dtree decomposition. The dynamic separator decomposition shows promise in that it significantly decreases the number of decisions for some real-world problems. Wei Li 0002, Peter van Beek |
ICTAI | 1 |