Sijin Chen

dblp:96/9616 · DBLP profile ↗
← Back
24ranked-venue papers
8as first author
17since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 14 · 7 first-author · 14 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 2 first-author · 5 since 2021Human-computer interaction and ubiquitous computing · 4Applied, interdisciplinary, general and emerging computing · 3 · 1 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 LetheVR: A First-Person Serious Game for Empathy and Public Understanding of Dementia
abstract
Dementia, one of the leading causes of neurodegenerative mortality in older adults, remains widely misunderstood by the general public—not only stigmatized socially, but also subject to persistent misconceptions about its symptoms, progression, and lived experience. These misunderstandings hinder timely care, empathy, and social support. To address this, we introduce LetheVR, a first-person serious game designed to promote both empathic understanding of individuals living with dementia and cognitive awareness of the disease itself. Targeted at general audiences, the system adopts experiential methods embedded within a game-based structure. Unlike traditional media and static educational tools, LetheVR integrates immersive symptom simulation, narrative-driven gameplay, and guided pedagogical reflection to engage users in the lived experience of dementia. In a controlled study with 60 participants, LetheVR significantly outperformed conventional interventions in improving measured empathy levels and symptom understanding. These findings highlight the potential of Virtual Reality combining with serious games and experiential methods as effective public health interventions for reshaping attitudes and correcting public misunderstandings about dementia.
Cheng Nie, Ding Ding 0002, Chenjun Wu, Sijin Chen, Zhuying Li 0001
IEEE Trans. Games4
2025 MeshAnything: Artist-Created Mesh Generation with Autoregressive Transformers
abstract
Recently, 3D assets created via reconstruction and generation have matched the quality of manually crafted assets, highlighting their potential for replacement. However, this potential is largely unrealized because these assets always need to be converted to meshes for 3D industry applications, and the meshes produced by current mesh extraction methods are significantly inferior to Artist-Created Meshes (AMs), i.e., meshes created by human artists. Specifically, current mesh extraction methods rely on dense faces and ignore geometric features, leading to inefficiencies, complicated post-processing, and lower representation quality. To address these issues, we introduce MeshAnything, a model that treats mesh extraction as a generation problem, producing AMs aligned with specified shapes. By converting 3D assets in any 3D representation into AMs, MeshAnything can be integrated with various 3D asset production methods, thereby enhancing their application across the 3D industry. The architecture of MeshAnything comprises a VQ-VAE and a shape-conditioned decoder-only transformer. We first learn a mesh vocabulary using the VQ-VAE, then train the shape-conditioned decoder-only transformer on this vocabulary for shape-conditioned autoregressive mesh generation. Our extensive experiments show that our method generates AMs with hundreds of times fewer faces, significantly improving storage, rendering, and simulation efficiencies, while achieving precision comparable to previous methods.
Tong He 0001, Weicai Ye, Sijin Chen, Jiaxiang Tang, Zhongang Cai, Lei Yang 0045, Gang Yu 0002, Guosheng Lin, Chi Zhang 0007
ICLR5
2025 Decoding Game: On Minimax Optimality of Heuristic Text Generation Strategies
abstract
Decoding strategies play a pivotal role in text generation for modern language models, yet a puzzling gap divides theory and practice. Surprisingly, strategies that should intuitively be optimal, such as Maximum a Posteriori (MAP), often perform poorly in practice. Meanwhile, popular heuristic approaches like Top-$k$ and Nucleus sampling, which employ truncation and normalization of the conditional next-token probabilities, have achieved great empirical success but lack theoretical justifications. In this paper, we propose Decoding Game, a comprehensive theoretical framework which reimagines text generation as a two-player zero-sum game between Strategist, who seeks to produce text credible in the true distribution, and Nature, who distorts the true distribution adversarially. After discussing the decomposibility of multi-step generation, we derive the optimal strategy in closed form for one-step Decoding Game. It is shown that the adversarial Nature imposes an implicit regularization on likelihood maximization, and truncation-normalization methods are first-order approximations to the optimal strategy under this regularization. Additionally, by generalizing the objective and parameters of Decoding Game, near-optimal strategies encompass diverse methods such as greedy search, temperature scaling, and hybrids thereof. Numerical experiments are conducted to complement our theoretical analysis.
Sijin Chen, Omar Hagrass, Jason M. Klusowski
ICLR1
2025 OmniSVG: A Unified Scalable Vector Graphics Generation Model
abstract
Scalable Vector Graphics (SVG) is an important image format widely adopted in graphic design because of their resolution independence and editability. The study of generating high-quality SVG has continuously drawn attention from both designers and researchers in the AIGC community. However, existing methods either produces unstructured outputs with huge computational cost or is limited to generating monochrome icons of over-simplified structures. To produce high-quality and complex SVG, we propose OmniSVG, a unified framework that leverages pre-trained Vision-Language Models (VLMs) for end-to-end multimodal SVG generation. By parameterizing SVG commands and coordinates into discrete tokens, OmniSVG decouples structural logic from low-level geometry for efficient training while maintaining the expressiveness of complex SVG structure. To further advance the development of SVG synthesis, we introduce MMSVG-2M, a multimodal dataset with two million richly annotated SVG assets, along with a standardized evaluation protocol for conditional SVG generation tasks. Extensive experiments show that OmniSVG outperforms existing methods and demonstrates its potential for integration into professional SVG design workflows.
Sijin Chen, Xianfang Zeng, Fukun Yin, Gang Yu 0002, Xingjun Ma, Yu-Gang Jiang 0001
NeurIPS3
2025 WI3D: Weakly Incremental 3D Detection via Vision Foundation Models
abstract
Class-incremental 3D object detection demands a 3D detector tolocateandrecognizenovel categories in a stream fashion while preserving its base detection ability. However, existing methods require delicate 3D annotations for learning novel categories, resulting in significant labeling costs. To this end, we explore a label-efficient approach calledWeaklyIncremental3DDetection (WI3D), which teaches a 3D detector to learn incrementally with off-the-shelf vision foundation models. We propose a novel dual-teaching framework incorporating both intra-modal and inter-modal knowledge from pseudo labels and feature space. Specifically, our framework features a class-agnostic pseudo-label refinement module, designed for the generation of high-quality 3D pseudo labels. This module is built on a lightweight transformer that models the spatial relationships between pseudo labels and their interactions with rich contextual information in point clouds. Additionally, we introduce a cross-modal knowledge transfer module to enhance the representation learning of novel classes, along with a reweighting knowledge distillation strategy that dynamically assesses and distills knowledge from previously learned categories. Extensive experiments show that our approach can efficiently learn novel concepts while preserving knowledge of base classes in WI3D scenarios, and surpass baseline approaches on both SUN-RGBD and ScanNet.
Mingsheng Li, Sijin Chen, Shengji Tang, Hongyuan Zhu 0002, Yanyan Fang, Xin Chen 0040, Zhuoyuan Li 0006, Fukun Yin, Tao Chen 0003
IEEE Trans. Multim.2
2024 Non-Convex Joint Community Detection and Group Synchronization via Generalized Power Method
abstract
This paper proposes a Generalized Power Method (GPM) to simultaneously solve the joint problem of community detection and group synchronization in a direct non-convex manner, in contrast to the existing method of semidefinite programming (SDP). Under a natural extension of stochastic block model (SBM), our theoretical analysis proves that the proposed algorithm is able to exactly recover the ground truth in $O(n\log^2 n)$ time for problems of size $n$, sharply outperforming the $O(n^{3.5})$ runtime of SDP. Moreover, we give a lower bound of model parameters as a sufficient condition for the exact recovery of GPM. The new bound breaches the information-theoretic limit for pure community detection under SBM, thus demonstrating the superiority of our simultaneous optimization algorithm over any two-stage method that performs the two tasks in succession. We also conduct numerical experiments on GPM and SDP to corroborate our theoretical analysis.
Sijin Chen, Xiwei Cheng, Anthony Man-Cho So
AISTATS1
2024 Escaping Saddle Points in Heterogeneous Federated Learning via Distributed SGD with Communication Compression
abstract
We consider the problem of finding second-order stationary points in the optimization of heterogeneous federated learning (FL). Previous works in FL mostly focus on first-order convergence guarantees, which do not rule out the scenario of unstable saddle points. Meanwhile, it is a key bottleneck of FL to achieve communication efficiency without compensating the learning accuracy, especially when local data are highly heterogeneous across different clients. Given this, we propose a novel algorithm PowerEF-SGD that only communicates compressed information via a novel error-feedback scheme. To our knowledge, PowerEF-SGD is the first distributed and compressed SGD algorithm that provably escapes saddle points in heterogeneous FL without any data homogeneity assumptions. In particular, PowerEF-SGD improves to second-order stationary points after visiting first-order (possibly saddle) points, using additional gradient queries and communication rounds only of almost the same order required by first-order convergence, and the convergence rate shows a linear-speedup pattern in terms of the number of workers. Our theory improves/recovers previous results, while extending to much more tolerant settings on the local data. Numerical experiments are provided to complement the theory.
Sijin Chen, Zhize Li 0001, Yuejie Chi
AISTATS1
2024 LL3DA: Visual Interactive Instruction Tuning for Omni-3D Understanding, Reasoning, and Planning
abstract
Recent progress in Large Multimodal Models (LMM) has opened up great possibilities for various applications in the field of human-machine interactions. However, developing LMMs that can comprehend, reason, and plan in complex and diverse 3D environments remains a challenging topic, especially considering the demand for understanding permutation-invariant point cloud representations of the 3D scene. Existing works seek help from multi-view images by projecting 2D features to 3D space, which inevitably leads to huge computational overhead and performance degradation. In this paper, we present LL3DA, a Large Language 3D Assistant that takes point cloud as the direct input and responds to both text instructions and visual interactions. The additional visual interaction enables LMMs to better comprehend human interactions with the 3D environment and further remove the ambiguities within plain texts. Experiments show that LL3DA achieves remarkable results and surpasses various 3D vision-language models on both 3D Dense Captioning and 3D Question Answering.
Sijin Chen, Xin Chen 0040, Chi Zhang 0007, Mingsheng Li, Gang Yu 0002, Hao Fei 0001, Hongyuan Zhu 0002, Jiayuan Fan 0001, Tao Chen 0003
CVPR1
2024 M3DBench: Towards Omni 3D Assistant with Interleaved Multi-modal Instructions
Mingsheng Li, Xin Chen 0040, Chi Zhang 0007, Sijin Chen, Hongyuan Zhu 0002, Fukun Yin, Zhuoyuan Li 0006, Gang Yu 0002, Tao Chen 0003
ECCV (58)4
2024 Hybrid Explainable Network Intrusion Detection Framework Based on Shapley Additive Explanations
abstract
With the rapid advancements in network technology and automation processes, the threats posed by cyberattacks have become increasingly significant. To address these threats, numerous researchers have developed various network intrusion detection systems (NIDS) to monitor network traffic. However, with the continuous complexification of 5G networks and the exponential increase in network traffic, the emergence of new attacks alongside the lack of interpretability in NIDS posed challenges to the performance and efficiency of network intrusion detection. To tackle these issues, this paper proposes a hybrid explainable network intrusion detection framework that combines the strengths of supervised and unsupervised learning, enabling effective detection of emerging attacks within the network. Specifically, we utilize a Light Gradient Boosting Machine (LightGBM) model for supervised learning, followed by the SHapley Additive exPlanations (SHAP) method for model explanation and feature selection. Additionally, we employ a network utilizing Convolutional Neural Network and Long short term memory for encoding and decoding (ACLNet) purposes in unsupervised learning. Finally, the results from these two learning processes are integrated for anomaly detection. The simulation experimental results on the NSL-KDD dataset verify that our approach generates detection performance on par with state-of-the-art methods while offering significantly enhanced interpretability and improving detection efficiency.
Sijin Chen, Jing Liu 0032, Cen Chen 0002, Songyu Xie, Zhongyao Cheng
ISPA1
2024 MeshXL: Neural Coordinate Field for Generative 3D Foundation Models
abstract
The polygon mesh representation of 3D data exhibits great flexibility, fast rendering speed, and storage efficiency, which is widely preferred in various applications. However, given its unstructured graph representation, the direct generation of high-fidelity 3D meshes is challenging. Fortunately, with a pre-defined ordering strategy, 3D meshes can be represented as sequences, and the generation process can be seamlessly treated as an auto-regressive problem. In this paper, we validate Neural Coordinate Field (NeurCF), an explicit coordinate representation with implicit neural embeddings, is a simple-yet-effective representation for large-scale sequential mesh modeling. After that, we present MeshXL, a family of generative pre-trained auto-regressive models that addresses 3D mesh generation with modern large language model approaches. Extensive experiments show that MeshXL is able to generate high-quality 3D meshes, and can also serve as foundation models for various down-stream applications.
Sijin Chen, Xin Chen 0040, Anqi Pang, Xianfang Zeng, Yijun Fu, Fukun Yin, Billzb Wang, Jingyi Yu 0001, Gang Yu 0002, Tao Chen 0003
NeurIPS1
2024 3DET-Mamba: Causal Sequence Modelling for End-to-End 3D Object Detection
abstract
Transformer-based architectures have been proven successful in detecting 3D objects from point clouds. However, the quadratic complexity of the attention mechanism struggles to encode rich information as point cloud resolution increases. Recently, state space models (SSM) such as Mamba have gained great attention due to their linear complexity and long sequence modeling ability for language understanding. To exploit the potential of Mamba on 3D scene-level perception, for the first time, we propose 3DET-Mamba, which is a novel SSM-based model designed for indoor 3d object detection. Specifically, we divide the point cloud into different patches and use a lightweight yet effective Inner Mamba to capture local geometric information. To observe the scene from a global perspective, we introduce a novel Dual Mamba module that models the point cloud in terms of spatial distribution and continuity. Additionally, we design a Query-aware Mamba module that decodes context features into object sets under the guidance of learnable queries. Extensive experiments demonstrate that 3DET-Mamba surpasses previous 3DETR on indoor 3D detection benchmarks such as ScanNet, improving AP25/AP50 from 65.0\%/47.0\% to 70.4\%/54.4\%, respectively.
Mingsheng Li, Jiakang Yuan, Sijin Chen, Lin Zhang 0055, Anyu Zhu, Tao Chen 0003
NeurIPS3
2024 HEN: a novel hybrid explainable neural network based framework for robust network intrusion detection
Wei Wei 0006, Sijin Chen, Cen Chen 0002, Heshi Wang, Jing Liu 0032, Zhongyao Cheng, Xiaofeng Zou
Sci. China Inf. Sci.2
2024 Vote2Cap-DETR++: Decoupling Localization and Describing for End-to-End 3D Dense Captioning
abstract
3D dense captioning requires a model to translate its understanding of an input 3D scene into several captions associated with different object regions. Existing methods adopt a sophisticated "detect-then-describe" pipeline, which builds explicit relation modules upon a 3D detector with numerous hand-crafted components. While these methods have achieved initial success, the cascade pipeline tends to accumulate errors because of duplicated and inaccurate box estimations and messy 3D scenes. In this paper, we first propose Vote2Cap-DETR, a simple-yet-effective transformer framework that decouples the decoding process of caption generation and object localization through parallel decoding. Moreover, we argue that object localization and description generation require different levels of scene understanding, which could be challenging for a shared set of queries to capture. To this end, we propose an advanced version, Vote2Cap-DETR++, which decouples the queries into localization and caption queries to capture task-specific features. Additionally, we introduce the iterative spatial refinement strategy to vote queries for faster convergence and better localization performance. We also insert additional spatial information to the caption head for more accurate descriptions. Without bells and whistles, extensive experiments on two commonly used datasets, ScanRefer and Nr3D, demonstrate Vote2Cap-DETR and Vote2Cap-DETR++ surpass conventional "detect-then-describe" methods by a large margin.
Sijin Chen, Hongyuan Zhu 0002, Mingsheng Li, Xin Chen 0040, Peng Guo 0011, Yinjie Lei, Gang Yu 0002, Taihao Li, Tao Chen 0003
IEEE Trans. Pattern Anal. Mach. Intell.1
2023 End-to-End 3D Dense Captioning with Vote2Cap-DETR
abstract
3D dense captioning aims to generate multiple captions localized with their associated object regions. Existing methods follow a sophisticated “detect-then-describe” pipeline equipped with numerous hand-crafted components. However, these hand-crafted components would yield sub-optimal performance given cluttered object spatial and class distributions among different scenes. In this paper, we propose a simple-yet-effective transformer framework Vote2Cap-DETR based on recent popular DEtection TRansformer (DETR). Compared with prior arts, our framework has several appealing advantages: 1) Without resorting to numerous hand-crafted components, our method is based on a full transformer encoder-decoder architecture with a learnable vote query driven object decoder, and a caption decoder that produces the dense captions in a set-prediction manner. 2) In contrast to the two-stage scheme, our method can perform detection and captioning in one-stage. 3) Without bells and whistles, extensive experiments on two commonly used datasets, ScanRefer and Nr3D, demonstrate that our Vote2Cap-DETR surpasses current state-of-the-arts by 11.13% and 7.11% in [email protected], respectively. Codes will be released soon.
Sijin Chen, Hongyuan Zhu 0002, Xin Chen 0040, Yinjie Lei, Gang Yu 0002, Tao Chen 0003
CVPR1
2022 Aspect-level sentiment classification based on attention-BiLSTM model and transfer learning
Zixin Zhang 0008, Shaona Yu, Yueting Meng, Sijin Chen
Knowl. Based Syst.6
2021 CIA-SSD: Confident IoU-Aware Single-Stage Object Detector From Point Cloud
abstract
Existing single-stage detectors for locating objects in point clouds often treat object localization and category classification as separate tasks, so the localization accuracy and classification confidence may not well align. To address this issue, we present a new single-stage detector named the Confident IoU-Aware Single-Stage object Detector (CIA-SSD). First, we design the lightweight Spatial-Semantic Feature Aggregation module to adaptively fuse high-level abstract semantic features and low-level spatial features for accurate predictions of bounding boxes and classification confidence. Also, the predicted confidence is further rectified with our designed IoU-aware confidence rectification module to make the confidence more consistent with the localization accuracy. Based on the rectified confidence, we further formulate the Distance-variant IoU-weighted NMS to obtain smoother regressions and avoid redundant predictions. We experiment CIA-SSD on 3D car detection in the KITTI test set and show that it attains top performance in terms of the official ranking metric (moderate AP 80.28%) and above 32 FPS inference speed, outperforming all prior single-stage detectors. The code is available at https://github.com/Vegeta2020/CIA-SSD.
Wu Zheng, Weiliang Tang, Sijin Chen, Li Jiang 0009, Chi-Wing Fu
AAAI3
2018 Why doesn't it work?: voice-driven interfaces and young children's communication repair strategies
abstract
In this study, we examine the conversational repair strategies that preschoolers use to correct communication breakdowns with a voice-driven interface. We conducted a two-week deployment in the homes of 14 preschoolers of a tablet game that included a broken voice-driven mini-game. We collected 107 audio samples of these children's (unsuccessful) attempts to communicate with the mini-game. We found that children tried a common set of repair strategies, including repeating themselves and experimenting with the tone and pronunciation of their words. Children were persistent, rarely giving up on the interaction, asking for help, or showing frustration. When parents participated in the interaction, they moved through four phases of engagement: first making suggestions, then intervening, then making statements of resignation, and finally pronouncing that the interaction could not be repaired. Designers should anticipate that in this context, children will borrow behaviors from person-to-person communication, such as pivoting strategies to probe the source of failed communication and structuring communication into turn-taking attempts.
Kate Yen, Yeqi Chen, Sijin Chen, Alexis Hiniker
IDC4
2018 Empowering Families Facing English Literacy Challenges to Jointly Engage in Computer Programming
abstract
Research suggests that parental engagement through Joint Media Engagement (JME) is an important factor in children's learning for coding and programming. Unfortunately, parents with limited technology background may have difficulty supporting their children's access to programming. English-language learning (ELL) families from marginalized communities face particular challenges in understanding and supporting programming, as code is primarily authored using English text. We present BlockStudio, a programming tool for empowering ELL families to jointly engage in introductory coding, using an environment embodying two design principles, text-free and visually concrete. We share a case study involving three community centers serving immigrant and refugee populations. Our findings show ELL families can jointly engage in programming without text, via co-creation and flexible roles, and can create a range of artifacts, indicating understanding of aspects of programming within this environment. We conclude with implications for coding together in ELL families and design ideas for text-free programming research.
Rahul Banerjee, Leanne Liu, Kiley Sobel, Caroline Pitt, Kung Jin Lee, Sijin Chen, Lydia Davison, Jason C. Yip 0001, Amy J. Ko, Zoran Popovic
CHI7
2018 Joint Media Engagement between Parents and Preschoolers in the U.S., China, and Taiwan
abstract
Global app marketplaces make families in foreign countries easily accessible to developers, but most scholarship on joint media engagement (JME) between parents and children reports on data from participants in Western contexts. We conducted an observational lab study to examine how preschoolers (age 3-5) and parents (N=74) from three different regions of the world (communities in China, Taiwan, and the United States) engage with two types of tablet games: an instructional game with goals and an exploratory, open-ended game. We found systematic differences among groups and between games. For example, parents from China and Taiwan frequently picked up their child's hand and used it as a tool to engage with the screen, a practice parents in our U.S. sample did not employ. Dyads from all three samples exhibited more warmth when playing an instructional game than an exploratory one. Our results suggest that characteristics of the populations we sampled interact with design features, that is, the same design prompted opposing behaviors in different groups. We conclude that it may be useful to examine goal-free and goal-oriented JME as separate constructs, that design choices influence the roles parents adopt during JME, and that the range of behaviors we observed complicate the prevailing research narrative of what positive and productive JME looks like.
Kate Yen, Yeqi Chen, Sijin Chen, Ying-Yu Chen, Yiran Ni, Alexis Hiniker
Proc. ACM Hum. Comput. Interact.4
2017 Examining Adult-Child Interactions in Intergenerational Participatory Design
abstract
Prior studies have focused on child interactions in participatory design (PD) with adults and children, but less is known about what specific adult-child interactions constitute a partnership. In this study, we unpack what constitutes an "equal partnership" in PD between adults and children. On the basis of prior literature, we created a new framework that examines the complementary roles between children and adults. Next, we analyzed a case study of a year-long intergenerational design team of children (ages 7-11) and adults. From this analysis, we determined that design partnerships are composed of four dimensions that span from unbalanced to balanced interactions: facilitation, relationship building, design-by-doing, and elaborating together. Finally, to demonstrate its utility, we analyzed two focal co-design sessions using our framework. Our analysis suggests that equal partnership in PD is not a single static interaction but a development over time of design interactions influenced by context, experience, and participants.
Jason C. Yip 0001, Kiley Sobel, Caroline Pitt, Kung Jin Lee, Sijin Chen, Kari Nasu, Laura R. Pina
CHI5
2004 The application of dyadic wavelet in the RS building image edge detection
Qiming Qin, Sijin Chen
ICIP3
2004 The founding and application of pattern database for building recognition
abstract
The founding of building pattern database is the key technique of building recognition of high resolution remote sensing imagery. This paper founded a simple pattern database by extracting the characteristics of building in remote sensing imagery and summing up several building patterns. Next, on the basis of patterns from the pattern database, this paper used the wavelet descriptor to describe the characteristics of buildings and utilized its affine invariant to recognize them. Then an image of Peking University was taken as an example to do the experiment. The result proved the method of founding pattern database was feasible.
Qiming Qin, Sijin Chen
IGARSS2
2004 Research of digital semi-fragile watermarking of remote sensing image based on wavelet analysis
abstract
In this paper, we present a novel semi-fragile watermarking scheme based on wavelet packet. The method in the paper includes four parts: first, to produce watermark; second, to scramble watermarking image; third, to embed watermark; last, to inspect and locate tampered marked image. To inspect whether including watermark in an image with the key attained from process of embedding watermarking. If it is a marked image, then extracting watermarking. At last to validate the degree of robustness by compression and noise, to locate tamper by cutting and altering.
Qiming Qin, Sijin Chen, Dezhi Chen
IGARSS3