Shu Wei

dblp:59/4109 · DBLP profile ↗
← Back
16ranked-venue papers
3as first author
15since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 9 · 1 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 2 first-author · 5 since 2021Human-computer interaction and ubiquitous computing · 5 · 2 first-author · 5 since 2021Systems, architecture and hardware · 2 · 1 since 2021
YearPublicationVenuePosition
2026 Implicit Neural Representation with Multi-Scale Sine Activation
abstract
Implicit Neural Representations (INRs) have become a powerful paradigm for modeling continuous signals in computer vision, graphics, and scientific computing. However, multilayer perceptrons (MLPs) generally suffer from severe spectral bias, which limits their ability to accurately model high-frequency details and multi-scale structures. To address this challenge, we propose a novel Multi-Scale Sine Activation (MSA), which explicitly introduces multi-scale frequency responses by incorporating multiple sets of sine activations with logarithmically spaced frequencies in parallel at each layer. MSA is further combined with an amplitude modulation mechanism to ensure numerical stability and robust optimization across different frequency channels. We conduct extensive experiments on a series of challenging tasks, including 1D multi-scale function fitting, image representation, video representation, 3D shape representation, and PDEs solving. Experimental results show that MSA outperforms existing state-of-the-art methods in terms of reconstruction accuracy, detail preservation, and training stability.
Jufeng Han, Shu Wei, Weijun Li 0002, Linjun Sun, Hong Qin 0007
AAAI2
2026 Eyes Can't Always Tell: Fusing Eye Tracking and User Priors for User Modeling under AI Advice Conditions
abstract
Modeling users' cognitive states (e.g., cognitive load and decision confidence) is essential for building adaptive AI in high-stakes decision-making. While eye tracking provides non-invasive behavioral signals correlated with cognitive effort, prior work has not systematically examined how AI assistance contexts, specifically varying advice reliability and user heterogeneity, can alter the mapping between gaze signals and cognitive states. We conducted a within-subject lab eye-tracking study (N=54) on factual verification tasks under three conditions: No-AI, Correct-AI advice, and Incorrect-AI advice. We analyze condition-dependent changes in self-reports and eye-tracking patterns and evaluate the robustness of eye-tracking-based user modeling. Results show that AI advice increases decision confidence compared to No-AI, while Correct-AI is associated with lower perceived cognitive load and more efficient gaze behavior. Crucially, predictive modeling is context-sensitive: the relationship between eye-tracking signals and cognitive states shifts across AI conditions. Finally, fusing eye-tracking features with user priors (demographics, AI literacy/experience, and propensity to trust technology) improves cross-participant generalization. These findings support condition-aware and personalized user modeling for cognitively aligned adaptive AI systems.
Xin Sun 0016, Shu Wei, Jos A. Bosch, Isao Echizen, Abdallah El Ali, Saku Sugawara
UMAP2
2026 Autoencoder: An efficient inverse design method for gallium nitride high electron mobility transistor structures
Meilan Hao, Shu Wei, Jufeng Han, Hong Qin 0007, Weijun Li 0002
Eng. Appl. Artif. Intell.3
2026 Understanding trust toward human versus AI-generated health information through behavioral and physiological sensing
abstract
As AI-generated health information proliferates online and becomes increasingly indistinguishable from human-sourced information, it becomes critical to understand how people trust and label such content, especially when the information is inaccurate. We conducted two complementary studies: (1) a mixed-methods survey (N=142) employing a 2 (source: Human vs. LLM) × 2 (label: Human vs. AI) × 3 (type: General, Symptom, Treatment) design, and (2) a within-subjects lab study (N=40) incorporating eye-tracking and physiological sensing (ECG, EDA, skin temperature). Participants were presented with health information varying by source-label combinations and asked to rate their trust, while their gaze behavior and physiological signals were recorded. We found that LLM-generated information was trusted more than human-generated content, whereas information labeled as human was trusted more than that labeled as AI. Trust remained consistent across information types. Eye-tracking and physiological responses varied significantly by source and label. Machine learning models trained on these behavioral and physiological features predicted binary self-reported trust levels with 73 % accuracy and information source with 65 % accuracy. Our findings demonstrate that adding transparency labels to online health information modulates trust. Behavioral and physiological features show potential to verify trust perceptions and indicate if additional transparency is needed.
Xin Sun 0016, Rongjun Ma, Shu Wei, Pablo César, Jos A. Bosch, Abdallah El Ali
Int. J. Hum. Comput. Stud.3
2026 FIRE: Fourier-series Implicit Neural Representations for high-fidelity continuous signal modeling
Jufeng Han, Shu Wei, Xin Ning 0001, Lusi Li, Hong Qin 0007, Weijun Li 0002
Pattern Recognit.3
2025 MetaSymNet: A Tree-like Symbol Network with Adaptive Architecture and Activation Functions
abstract
Mathematical formulas are the language of communication between humans and nature. Discovering latent formulas from observed data is an important challenge in artificial intelligence, commonly known as symbolic regression(SR). The current mainstream SR algorithms regard SR as a combinatorial optimization problem and use Genetic Programming (GP) or Reinforcement Learning (RL) to solve the SR problem. These methods perform well on simple problems, but poorly on slightly more complex tasks. In addition, this class of algorithms ignores an important aspect: in SR tasks, symbols have explicit numerical meaning. So can we take full advantage of this important property and try to solve the SR problem with more efficient numerical optimization methods? Extrapolation and Learning Equation (EQL) replaces activation functions in neural networks with basic symbols and sparsifies connections to derive a simplified expression from a large network. However, EQL's fixed network structure can't adapt to the complexity of different tasks, often resulting in redundancy or insufficient, limiting its effectiveness. Based on the above analysis, we propose MetaSymNet, a tree-like network that employs the PANGU meta-function as its activation function. PANGU meta-function can evolve into various candidate functions during training. The network structure can also be adaptively adjusted according to different tasks. Then the symbol network evolves into a concise, interpretable mathematical expression. To evaluate the performance of MetaSymNet and five baseline algorithms, we conducted experiments across more than ten datasets, including SRBench. The experimental results show that MetaSymNet has achieved relatively excellent results on various evaluation metrics.
Yanjie Li 0005, Weijun Li 0002, Shu Wei, Yusong Deng, Meilan Hao
AAAI6
2025 Closed-form Solutions: A New Perspective on Solving Differential Equations
abstract
The quest for analytical solutions to differential equations has traditionally been constrained by the need for extensive mathematical expertise. Machine learning methods like genetic algorithms have shown promise in this domain, but are hindered by significant computational time and the complexity of their derived solutions. This paper introduces **SSDE** (Symbolic Solver for Differential Equations), a novel reinforcement learning-based approach that derives symbolic closed-form solutions for various differential equations. Evaluations across a diverse set of ordinary and partial differential equations demonstrate that SSDE outperforms existing machine learning methods, delivering superior accuracy and efficiency in obtaining analytical solutions.
Shu Wei, Yanjie Li 0005, Weijun Li 0002, Linjun Sun, Hong Qin 0007, Yusong Deng, Jufeng Han
ICML1
2025 Demonstrating the Screenless Optical Theremin with Tremolo (ScOTT)
abstract
Figure 1: A demonstration of the Screenless Optical Theremin with Tremolo (ScOTT) in a collaborative musical setting.
Michael Gancz, Justin Berry, Shu Wei, Kimberly Hieftje, Asher Marks
IMX3
2025 The Arborist: A Collective Bloom Through Physiological Data in Mixed Reality
abstract
This mixed reality installationThe Arborist transforms physiological data (heart rate-HR, galvanic skin conductance-GSR, temperature) into collaborative digital flora using Meta Quest 3 and Shimmer3 sensors.Each user's biometrics generate unique flowers-color tied to HR, bloom size to arousal-that populate a Tree of Shared Breath.
Shu Wei, Barnabas Lee, Michael Gancz, Asher Marks, Kimberly Hieftje
IMX1
2025 CaMo: Capturing the modularity by end-to-end models for Symbolic Regression
Weijun Li 0002, Yanjie Li 0005, Meilan Hao, Yusong Deng, Shu Wei
Knowl. Based Syst.9
2024 A Neural-Guided Dynamic Symbolic Network for Exploring Mathematical Expressions from Data
abstract
Symbolic regression (SR) is a powerful technique for discovering the underlying mathematical expressions from observed data. Inspired by the success of deep learning, recent deep generative SR methods have shown promising results. However, these methods face difficulties in processing high-dimensional problems and learning constants due to the large search space, and they don’t scale well to unseen problems. In this work, we propose DySymNet, a novel neural-guided Dynamic Symbolic Network for SR. Instead of searching for expressions within a large search space, we explore symbolic networks with various structures, guided by reinforcement learning, and optimize them to identify expressions that better-fitting the data. Based on extensive numerical experiments on low-dimensional public standard benchmarks and the well-known SRBench with more variables, DySymNet shows clear superiority over several representative baseline models. Open source code is available at https://github.com/AILWQ/DySymNet.
Weijun Li 0002, Linjun Sun, Yanjie Li 0005, Shu Wei, Yusong Deng, Meilan Hao
ICML8
2024 TabPedia: Towards Comprehensive Visual Table Understanding with Concept Synergy
abstract
Tables contain factual and quantitative data accompanied by various structures and contents that pose challenges for machine comprehension. Previous methods generally design task-specific architectures and objectives for individual tasks, resulting in modal isolation and intricate workflows. In this paper, we present a novel large vision-language model, TabPedia, equipped with a concept synergy mechanism. In this mechanism, all the involved diverse visual table understanding (VTU) tasks and multi-source visual embeddings are abstracted as concepts. This unified framework allows TabPedia to seamlessly integrate VTU tasks, such as table detection, table structure recognition, table querying, and table question answering, by leveraging the capabilities of large language models (LLMs). Moreover, the concept synergy mechanism enables table perception-related and comprehension-related tasks to work in harmony, as they can effectively leverage the needed clues from the corresponding source perception embeddings. Furthermore, to better evaluate the VTU task in real-world scenarios, we establish a new and comprehensive table VQA benchmark, ComTQA, featuring approximately 9,000 QA pairs. Extensive quantitative and qualitative experiments on both table perception and comprehension tasks, conducted across various public benchmarks, validate the effectiveness of our TabPedia. The superior performance further confirms the feasibility of using LLMs for understanding visual tables when all concepts work in synergy. The benchmark ComTQA has been open-sourced at https://huggingface.co/datasets/ByteDance/ComTQA. The source code and model also have been released at https://github.com/zhaowc-ustc/TabPedia.
Weichao Zhao, Hao Feng 0009, Jingqun Tang, Binghong Wu, Shu Wei, Yongjie Ye, Hao Liu 0003, Wengang Zhou 0001, Houqiang Li, Can Huang 0002
NeurIPS7
2024 Harmonizing Visual Text Comprehension and Generation
abstract
In this work, we present TextHarmony, a unified and versatile multimodal generative model proficient in comprehending and generating visual text. Simultaneously generating images and texts typically results in performance degradation due to the inherent inconsistency between vision and language modalities. To overcome this challenge, existing approaches resort to modality-specific data for supervised fine-tuning, necessitating distinct model instances. We propose Slide-LoRA, which dynamically aggregates modality-specific and modality-agnostic LoRA experts, partially decoupling the multimodal generation space. Slide-LoRA harmonizes the generation of vision and language within a singular model instance, thereby facilitating a more unified generative process. Additionally, we develop a high-quality image caption dataset, DetailedTextCaps-100K, synthesized with a sophisticated closed-source MLLM to enhance visual text generation capabilities further. Comprehensive experiments across various benchmarks demonstrate the effectiveness of the proposed approach. Empowered by Slide-LoRA, TextHarmony achieves comparable performance to modality-specific fine-tuning results with only a 2% increase in parameters and shows an average improvement of 2.5% in visual text comprehension tasks and 4.0% in visual text generation tasks. Our work delineates the viability of an integrated approach to multimodal generation within the visual text domain, setting a foundation for subsequent inquiries. Code is available at https://github.com/bytedance/TextHarmony.
Jingqun Tang, Binghong Wu, Chunhui Lin, Shu Wei, Hao Liu 0003, Xin Tan 0002, Zhizhong Zhang 0001, Can Huang 0002, Yuan Xie 0006
NeurIPS5
2023 A Preliminary Study of the Eye Tracker in the Meta Quest Pro
abstract
This paper presents the preliminary results of an accuracy testing of the Meta Quest Pro’s eye tracker. We conducted user testing to evaluate the spatial accuracy, spatial precision and subjective performance under head-free and head-restrained conditions. Our measurements indicated an average accuracy of 1.652° with a precision of 0.699° (standard deviation) and 0.849° (root mean square) for a visual field spanning 15° during head-free. The signal quality of Quest Pro’s eye-tracker is comparable to existing AR/VR eye-tracking headsets. Notably, careful considerations are required when designing the size of scene objects, mapping areas of interest, and determining the interaction flow. Researchers should also be cautious about interpreting the fixation results when multiple targets are within close proximity. Further investigation and better specification information transparency are needed to establish its capabilities and limitations.
Shu Wei, Desmond Bloemers, Aitor Rovira
IMX1
2022 Total Ionization Dose Measurement onboard a 1U CubeSat in Low Earth Orbit
abstract
The TMCR (Total Ionizing Dose Measurement of COTS and Onboard Rad-Hard Components Mission) mission onboard BIRDS-4, a 1U CubeSat, is attempting to measure changes in the electrical properties of semiconductors due to total ionizing dose effects. The mission has two objectives. The mission has two objectives: the first is to verify the accuracy of the TID tests performed on the ground. The second is to validate a system to measure radiation dose from the change in drain current of commercially available MOSFET devices. Based on the measurement of 237 days in LEO of 51degree inclination and 400km altitude, no significant changes in drain current have yet been observed, indicating a deviation from the results of the ground tests.
Akihiro Oboshi, Tomoaki Murase, Hirokazu Masui, Shu Wei, Mengu Cho
ISCAS4
2020 SS27 Radiation Protection and Monitoring Experiment on-Board a 1U CubeSat and its Ground Verification
abstract
Commercial-Off-The-Shelf(COTS) Integrated Circuits (ICs) are mostly sensitive to irradiations (e.g., protons, heavy-ions, etc.)in space, leading to various Single-Event Effects (SEEs). Among different SEEs, Single-Event Latch-up (SEL) is one of the most critical because it is detrimental and often damages ICs. This paper reports a circuit implementation that serves to provide protection for COTS ICs from SEL. The on-ground hardware verification on the basis of Californium 252 depicts that the circuit successfully cuts off the SEL current and subsequently resets the COTS IC. The reported circuit has been adopted in a 1U CubeSat that will be launched early 2020.
Tomoaki Murase, Hirokazu Masui, Mengu Cho, Shu Wei
ISCAS4