EDBT 2026 Demo / reviewers in the wild / expert
Yu Sheng
dblp:77/2258
· DBLP profile ↗
49ranked-venue papers
19as first author
28since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 15 · 5 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 13 · 9 first-author · 6 since 2021Artificial intelligence and machine learning · 12 · 4 first-author · 10 since 2021Human-computer interaction and ubiquitous computing · 8 · 4 first-author · 3 since 2021Systems, architecture and hardware · 7 · 5 since 2021Computer networks · 2 · 1 first-author · 2 since 2021Security and privacy · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Learning Theorem Rationale for Improving the Mathematical Reasoning Capability of Large Language ModelsabstractLarge language models (LLMs) have achieved significant progress in mathematical reasoning, especially in elementary math. However, they remain indisposed on tackling complex questions at high-school or college levels, which put forward a more advanced requirement of mastering relevant mathematical theorems. For we humans, whether selecting the appropriate theorems according to the provided question is a crucial factor affecting the quality of the ultimate solutions, yet which has been neglected by previous research in the field of LLM reasoning. In this paper, we propose a novel approach to enhance the LLM's capability of utilizing the mathematical theorems to specific problems, which we refer to as Theorem Rationale (TR). To this end, a new dataset encompassing problem-theorem-solution triples is deliberately established for transferring principles of TR. Furthermore, we develop an evolving strategy to boost hierarchical instructions oriented on the theorems to alleviate difficulty in acquiring the curated data and facilitate the digestion of theorem application from various perspectives. Evaluations on a wide range of public datasets exhibit that the model fine-tuned with our dataset achieves consistent improvements at varying mathematical levels compared to the backbone. And further ablation studies illustrate the effectiveness of our proposed evolutionary strategies on enhancing the model's capability of math problem-solving. Overall, extensive experiments reveal the potential of our proposed method which highlights the significance of aligning the problems with the concrete theorems for LLMs to alleviate hallucination and improve the models' mathematical reasoning capabilities. Yu Sheng, Linjing Li, Daniel Dajun Zeng |
AAAI | 1 |
| 2025 | Evaluating Generalization Capability of Language Models across Abductive, Deductive and Inductive Logical ReasoningabstractTransformer-based language models (LMs) have demonstrated remarkable performance on many natural language tasks, yet to what extent LMs possess the capability of generalizing to unseen logical rules remains not explored sufficiently. In classical logic category, abductive, deductive and inductive (ADI) reasoning are defined as the fundamental reasoning types, sharing the identical reasoning primitives and properties, and some research have proposed that there exists mutual generalization across them. However, in the field of natural language processing, previous research generally study LMs’ ADI reasoning capabilities separately, overlooking the generalization across them. To bridge this gap, we propose UniADILR, a novel logical reasoning dataset crafted for assessing the generalization capabilities of LMs across different logical rules. Based on UniADILR, we conduct extensive investigations from various perspectives of LMs’ performance on ADI reasoning. The experimental results reveal the weakness of current LMs in terms of extrapolating to unseen rules and inspire a new insight for future research in logical reasoning. Yu Sheng, Wanting Wen, Linjing Li, Daniel Dajun Zeng |
COLING | 1 |
| 2025 | SpatialSplat: Efficient Semantic 3D from Sparse Unposed ImagesabstractA major breakthrough in 3D reconstruction is the feedforward paradigm to generate pixel-wise 3D points or Gaussian primitives from sparse, unposed images. To further incorporate semantics while avoiding the significant memory and storage costs of high-dimensional semantic features, existing methods extend this paradigm by associating each primitive with a compressed semantic feature vector. However, these methods have two major limitations: (a) the naively compressed feature compromises expressiveness, affecting the model's ability to capture fine-grained semantics, and (b) the pixel-wise primitive prediction introduces redundancy in overlapping areas, causing unnecessary memory overhead. To this end, we introduce \textbf{SpatialSplat}, a feedforward framework that produces redundancy-aware Gaussians and capitalizes on a dual-field semantic representation. Particularly, with the insight that primitives within the same instance exhibit high semantic consistency, we decompose the semantic representation into a coarse feature field that encodes uncompressed semantics with minimal primitives, and a fine-grained yet low-dimensional feature field that captures detailed inter-instance relationships. Moreover, we propose a selective Gaussian mechanism, which retains only essential Gaussians in the scene, effectively eliminating redundant primitives. Our proposed Spatialsplat learns accurate semantic information and detailed instances prior with more compact 3D Gaussians, making semantic 3D reconstruction more applicable. We conduct extensive experiments to evaluate our method, demonstrating a remarkable 60\% reduction in scene representation parameters while achieving superior performance over state-of-the-art methods. The code is available at https://github.com/shengyuuu/SpatialSplat.git Yu Sheng, Jiajun Deng, Yu Zhang 0086, Bei Hua, Yanyong Zhang, Jianmin Ji |
ICCV | 1 |
| 2025 | ChatSeek: Open-Vocabulary Object Seeking with LLM-Informed Belief Field and Vehicle-Arm CooperationabstractObject Target Search (OTS) tasks require robots to navigate to objects specified by semantic labels, (e.g., find a fire extinguisher). Many existing OTS methods strongly rely on semantic co-occurrence relations among objects to search and localize the object targets in a closed set. However, simplistic domestic or chaotic rescue scenes often fail to provide rich semantics, which in some cases even require the robot to find previously unseen object instances. In addition, most of the existing OTS methods adopt a fixed Field of View (FoV) setting relative to the robot base. When the limited FoV only allows the observation of incomplete objects, this can result in incorrect categorical information and lead to wrong navigation actions. To address the above issues, we propose an open-vocabulary object-seeking method named ChatSeek based on the Large Language Model (LLM)-informed Object Belief Field (OBF) and vehicle-arm cooperation. In particular, our method prompts LLM to generate target objects’ affordance and geometric-part attributes to enhance object localization. By projecting CLIP-based object recognition likelihoods into 3D reconstructions, the OBF is updated to maintain the robot’s cognition of the surrounding scene. During OTS, the robot achieves a flexible vision by moving the vehicle-mounted robotic arm intentionally to translate and rotate the robot’s FoV to look around or even inspect hidden corners. Sufficient comparative and ablation studies demonstrate that our method can significantly improve OTS performance. Furthermore, real-world experiments show that our approach can find novel objects without requiring semantic priors. Bolei Chen, Liangbai Liu, Haonan Yang 0001, Yongzheng Cui, Shengsheng Yan, Ping Zhong 0002, Yu Sheng |
IJCNN | 7 |
| 2025 | LiteSCTransNet: Lightweight CNN-Transformer for 3D Medical Image Segmentation
Yu Sheng, Yiyi Hong, Yongchang Jia, Guihua Duan |
ISBRA (2) | 1 |
| 2025 | Unbiased Embodied Visual Representation Learning with Causal Inference and Cross-Modality AlignmentabstractObject Goal Navigation (ObjectNav) in novel environments relies on comprehensive scene understanding, including precise visual perception and accurate modeling of spatial-semantic regularities. However, excessive attention to the hand-crafted scene representation in prevailing approaches leads to the neglect of the negative influence of the perception bias hidden in the visual observations. The hand-crafted semantic distribution in domestic environments causes the spurious association bias, while the semantic conflict bias arises due to the dynamic perspective changes. Biased visual perception significantly limits the generalization of the navigation strategy. In this article, we propose the U nbiased E mbodied V isual R epresentation ( UEVR ), which overcomes the perception biases using causal inference and cross-modality alignment. Specifically, we establish reasonable assumptions about confounders for multi-object features through our proposed Unbiased Causal R-CNN framework and eliminate the spurious associations bias through B ackdoor I ntervention C ausal A djustment ( BICA ) module during navigation. To overcome the dynamic-view bias hidden in 2D image features, we propose to employ the cross-modality alignment mechanism with the Geometric Consistency ( GeoCon ) to encode 3D geometry prior into the 2D representations. Finally, we design a modular ObjectNav framework integrated with UEVR named Causal-ObjectNav , which consists of the corner-based scene exploration module and target object discrimination module. Extensive experiments on the MP3D and HM3D datasets demonstrate the superiority of the unbiased navigation model over existing ObjectNav methods. Jiaxu Kang, Bolei Chen, Ping Zhong 0002, Yifei Wang 0006, Haonan Yang 0001, Yu Sheng |
ACM Trans. Multim. Comput. Commun. Appl. | 6 |
| 2024 | Finding and Grasping: The Last-Mile of Object Goal Navigation
Yongzheng Cui, Bolei Chen, Haonan Yang 0001, Ping Zhong 0002, Yu Sheng |
CGI (3) | 5 |
| 2024 | Robot Autonomous Exploration System Base on Arm-Chassis CollaborationabstractToday, robots are finding more and more applications in areas such as industrial and agricultural production, environmental exploration, and disaster relief. To be effective in these roles, robots must be able to autonomously navigate complex and unstructured environments. Manipulative robotic arms can navigate through confined spaces and have the ability to explore diverse scenes. Therefore, the exploration of unknown environments, facilitated by the synergy between robotic arms and wheeled chassis, provides a more efficient and accurate approach. We have designed a comprehensive arm-chassis collaboration system that uses an information gain-based utility function to support exploration. By coordinating the motion of the robotic arm with the mobility of the chassis, this system efficiently performs complex exploration tasks. It enables autonomous exploration of unknown environments with reduced time and energy consumption. The code is published at: https://github.com/Southyang/Arm-Chassis. Haonan Yang 0001, Shengsheng Yan, Bolei Chen, Ping Zhong 0002, Yongzheng Cui, Yu Sheng |
CSCWD | 6 |
| 2024 | Unveiling Factual Recall Behaviors of Large Language Models through Knowledge NeuronsabstractIn this paper, we investigate whether Large Language Models (LLMs) actively recall or retrieve their internal repositories of factual knowledge when faced with reasoning tasks.Through an analysis of LLMs' internal factual recall at each reasoning step via Knowledge Neurons, we reveal that LLMs fail to harness the critical factual associations under certain circumstances.Instead, they tend to opt for alternative, shortcut-like pathways to answer reasoning questions.By manually manipulating the recall process of parametric knowledge in LLMs, we demonstrate that enhancing this recall process directly improves reasoning performance whereas suppressing it leads to notable degradation.Furthermore, we assess the effect of Chain-of-Thought (CoT) prompting, a powerful technique for addressing complex reasoning tasks.Our findings indicate that CoT can intensify the recall of factual knowledge by encouraging LLMs to engage in orderly and reliable reasoning.Furthermore, we explored how contextual conflicts affect the retrieval of facts during the reasoning process to gain a comprehensive understanding of the factual recall behaviors of LLMs. Wanting Wen, Yu Sheng, Linjing Li, Daniel Dajun Zeng |
EMNLP | 4 |
| 2024 | Integrating Language Models with Symbolic Formulas for First-Order Logic ReasoningabstractPerforming logical reasoning based on prior knowledge is a crucial human cognitive ability and has been a long-standing objective in the field of artificial intelligence. Large language models based on transformer architecture have been a common approach for logical reasoning over text. However, the current language models often struggle to learn semantic information from logical expressions, resulting in underwhelming performance on logical reasoning tasks. In this paper, we propose a novel method to convert first-order logic (FOL) expressions to the form of a graph and integrate it with embeddings from language models to enhance their reasoning ability. The proposed method is designed to learn directly from FOL formulas and is able to generalize to any scenarios involving logical expressions. Experimental results demonstrate that the proposed method enhances the model’s ability of learning logical semantic representations, and thus it brings a significant improvement on the performance of complex reasoning tasks. The code is available at https://github.com/FOL-GNN. Yu Sheng, Linjing Li, Daniel Dajun Zeng |
ICASSP | 1 |
| 2024 | HSPNav: Hierarchical Scene Prior Learning for Visual Semantic Navigation Towards Real SettingsabstractVisual Semantic Navigation (VSN) aims at navigating a robot to a given target object in a previously unseen scene. To tackle this task, the robot must learn a nimble navigation policy by utilizing spatial patterns and semantic co-occurrence relations among objects in the scene. Prevailing approaches extract scene priors from the instant visual observations and solidify them in neural episodic memory to achieve flexible navigation. However, due to the oblivion and underuse of the scene priors, these methods are plagued by repeated exploration, effective-knowledge sparsity, and wrong decisions. To alleviate these issues, we propose a novel VSN policy, HSPNav, based on Hierarchical Scene Priors (HSP) and Deep Reinforcement Learning (DRL). The HSP contains two components, i.e., the egocentric semantic map-based Local Scene Priors (LSP) and the commonsense relational graph-based Global Scene Priors (GSP). Then, efficient semantic navigation is achieved by employing an immediate LSP to retrieve conducive contextual memories from the GSP. By utilizing the MP3D dataset, the experimental results in the Habitat simulator demonstrate that our HSP brings a significant boost over the baselines. Furthermore, we take an essential step from simulation to reality by bridging the gap from Habitat to ROS. The migration evaluations show that HSPNav can generalize to realistic settings well and achieve promising performance. Jiaxu Kang, Bolei Chen, Ping Zhong 0002, Haonan Yang 0001, Yu Sheng, Jianxin Wang 0001 |
ICRA | 5 |
| 2024 | PGDepth: Roadside Long-Range Depth Estimation Guided by Priori Geometric InformationabstractLong-range depth is crucial for roadside perception, which helps vehicles detect potential threats earlier, respond promptly, and avoid collisions. However, a notable challenge with existing roadside perception methods is their difficulty in accurately perceiving objects at long-range depth. To this end, we propose a Priori Geometric-Guided long-range Depth estimation framework, named PGDepth. First, inspired by the human ability to perceive depth by referencing objects, we utilize priori Geometric as reference information for road objects, assigning each pixel a predefined depth range. Second, a coarse-to-fine approach is introduced to continuously refine the accuracy of depth distribution for pixels. Furthermore, we propose a new loss function to effectively supervise the depth distributions of road objects. Extensive experimental results on the DAIR dataset demonstrate that the proposed method surpasses previous state-of-the-art competitors. Wanrui Chen, Yu Sheng, Hui Feng 0001, Tao Yang 0008 |
INDIN | 2 |
| 2024 | SocialNav-FTI: Field-Theory-Inspired Social-aware Navigation Framework based on Human Behavior and Social NormsabstractSocial navigation is a key consideration for integrating robots into human environments. Concurrently, it imposes heightened requisites: tasks must not only be executed succesfully without collisions, but also adhere to principles encompassing comprehensibility, courtesy, social compliance, comprehension, foresight, and scenario compliance. In this paper, we present the incorporation of social norms as a guiding framework for robot navigation within social contexts. We adopt field theory to provide a formal elucidation of the social norms, using Physical-Informed Neural Network (PINN) to predict pedestrian movement under the influence of social norms, respectively, and using Reinforcement Learning (RL) for navigation. We use supervised learning to train the pedestrian velocity field prediction model and reinforcement learning to train the navigation policy. We conduct three parts of experiments: (1) analyzing the spatiotemporal characteristics of the velocity field in the walking pedestrians dataset; (2) evaluating the accuracy of the vector field prediction in the pedestrian dataset; (3) using Gazebo simulation and the PEDSIM library to evaluate the improvement of navigation performance under constraints of social norms. Experiments have confirmed that the pedestrian motion data set indeed satisfies the Gaussian divergence theorem and can be described by the concept of field. The performance of navigation strategies incorporating social rules has been improved to a certain extent. Siyi Lu, Ping Zhong 0002, Shuqi Ye, Bolei Chen, Yu Sheng, Run Liu 0001 |
IROS | 5 |
| 2024 | MM-Gaussian: 3D Gaussian-based Multi-modal Fusion for Localization and Reconstruction in Unbounded ScenesabstractLocalization and mapping are critical tasks for various applications such as autonomous vehicles and robotics. The challenges posed by outdoor environments present particular complexities due to their unbounded characteristics. In this work, we present MM-Gaussian, a LiDAR-camera multimodal fusion system for localization and mapping in unbounded scenes. Our approach is inspired by the recently developed 3D Gaussians, which demonstrate remarkable capabilities in achieving high rendering quality and fast rendering speed. Specifically, our system fully utilizes the geometric structure information provided by solid-state LiDAR to address the problem of inaccurate depth encountered when relying solely on visual solutions in unbounded, outdoor scenarios. Additionally, we utilize 3D Gaussian point clouds, with the assistance of pixel-level gradient descent, to fully exploit the color information in photos, thereby achieving realistic rendering effects. To further bolster the robustness of our system, we designed a relocalization module, which assists in returning to the correct trajectory in the event of a localization failure. Experiments conducted in multiple scenarios demonstrate the effectiveness of our method. Yifan Duan, Yu Sheng, Jianmin Ji, Yanyong Zhang |
IROS | 4 |
| 2024 | Embodied Contrastive Learning with Geometric Consistency and Behavioral Awareness for Object NavigationabstractObject Navigation (ObjcetNav), which enables an agent to seek any instance of an object category specified by a semantic label, has shown great advances. However, current agents are built upon occlusion-prone visual observations or compressed 2D semantic maps, which hinder their embodied perception of 3D scene geometry and easily lead to ambiguous object localization and blind exploration. To address these limitations, we present an Embodied Contrastive Learning (ECL) method with Geometric Consistency (GC) and Behavioral Awareness (BA), which motivates agents to actively encode 3D scene layouts and semantic cues. Driven by our embodied exploration strategy, BA is modeled by predicting navigational actions based on multi-frame visual images, as behaviors that cause differences between adjacent visual sensations are crucial for learning correlations among continuous visions. The GC is modeled as the alignment of behavior-aware visual stimulus with 3D semantic shapes by employing unsupervised contrastive learning. The aligned behavior-aware visual features and geometric invariance priors are injected into a modular ObjectNav framework to enhance object recognition and exploration capabilities. As expected, our ECL method performs well on object detection and instance segmentation tasks. Our ObjectNav strategy outperforms state-of-the-art methods on MP3D and Gibson datasets, showing the potential of our ECL in embodied navigation. Bolei Chen, Jiaxu Kang, Ping Zhong 0002, Yixiong Liang, Yu Sheng, Jianxin Wang 0001 |
ACM Multimedia | 5 |
| 2024 | Ph.D. Forum: Intelligent Home Energy Management: Developing AI-Driven Systems for Sustainable LivingabstractThe MAI-HOME project, funded by the Interreg initiative, addresses energy poverty and CO2 emission reduction through an AI-driven framework tailored for vulnerable populations. This research spans three years of data collection from multiple sensors installed in every room of sixteen houses across the Netherlands and Belgium. It aims to predict and promote energy-saving behaviors effectively. Utilizing an innovative blend of digital twins and robust data privacy measures, this project explores four critical areas: real-time data collection, predictive AI model development, data privacy enhancement, and behavioral intervention strategies. Initial findings suggest promising avenues for technological advancements and societal benefits in sustainable energy practices. Yu Sheng |
SenSys | 1 |
| 2024 | Judicial intelligent assistant system: Extracting events from Chinese divorce cases to detect disputes for the judgeabstractAbstract In the formal procedure of Chinese civil cases, the textual materials provided by different parties describe the development process of the cases. It is a difficult but necessary task to extract the key information for the cases from these textual materials and to clarify the dispute focus of related parties. Currently, officers read the materials manually and use methods, such as keyword searching and regular matching, to get the target information. These approaches are time‐consuming and heavily depend on prior knowledge and the carefulness of the officers. To assist the officers in enhancing working efficiency and accuracy, we conduct a case study of detecting disputes from Chinese divorce cases based on proposing a Two‐Round‐Labeling (TRL) event extracting technique in this article. We implement the Judicial Intelligent Assistant (JIA) system according to the proposed approach to (1) automatically extract focus events from divorce case materials, (2) align events by identifying co‐reference among them, and (3) detect conflicts among events brought by the plaintiff and the defendant. With the JIA system, it is convenient for judges to determine the disputed issues in Chinese divorce cases. Experimental results demonstrate that the proposed approach and system can obtain the focus of Chinese divorce cases and detect conflicts more effectively and efficiently compared with the existing method. Chuanyi Li, Yu Sheng, Jidong Ge, Bin Luo 0003 |
Expert Syst. J. Knowl. Eng. | 3 |
| 2023 | Knowledge-Enhanced Difference-Aware Clinical Time Series Representation Learning for Diagnosis PredictionabstractPredicting future health status based on historical patient visits is one of the essential tasks in healthcare. Many existing approaches attempt to enhance the representation learning capability of models by incorporating relevant medical knowledge, but their effectiveness is severely affected by the incompleteness and noise of the knowledge graphs. Moreover, due to the inability to capture temporal features at a fine-grained level, most existing methods also have limitations in learning the temporal development of patients’ health status. To address these issues, we propose a Knowledge-Enhanced Difference-Aware clinical time series representation learning model (KEDA) for diagnosis prediction. In this model, we first combine the medical ontology graph and co-occurrence graph, and use hierarchical graph convolution and contrastive learning methods to enhance the semantic representation of medical entities. After that, a task-specific difference-aware temporal module is designed to improve the accuracy of patient representation, which adds two novel gated units in the original GRU to fuse multi-type clinical information based on the relationship between different types of data and prediction tasks and capture fine-grained temporal evolution of patient health status. We validate our model on two publicly available datasets, and the experimental results demonstrate that KEDA outperforms the state-of-the-art methods. Ying An, Yinghong Shi, Lin Guo 0014, Yu Sheng, Xianlai Chen |
BIBM | 4 |
| 2023 | SCTransNet: 3D Medical Image Segmentation Model Based on the Fusion of CNN and TransformerabstractDeep learning has played an important role in medical image segmentation of liver and liver tumors, but existing models are still insufficient in accuracy and efficiency. In this paper, We proposed SCTransNet, a 3D image segmentation model based on the fusion of CNN and Transformer model. SCTransNet combines the feature extraction and expression capabilities of convolutional neural network(CNN) and the long-distance dependency modeling capability of Transformer model. SCTransNet improves the embedding layer and position encoding layer of the Transformer model to enhance the global contextual feature extraction ability. Meanwhile, a spatial attention module and a channel attention module based on improved Transformer model are designed in SCTransNet to enhance the feature extraction ability of low-dimensional pixel information and high-dimensional semantic information. Experimental results on the public dataset show that SCTransNet achieve relatively better performance than state-of-the-art methods. Yongchang Jia, Guihua Duan, Yu Sheng |
BIBM | 3 |
| 2023 | JARAD: An Approach for Java API Mention Recognition and Disambiguation in Stack Overflow
Qingmi Liang, Qi Xie 0010, Li Kuang, Yu Sheng |
CollaborateCom (1) | 5 |
| 2023 | CrowdNav-HERO: Pedestrian Trajectory Prediction Based Crowded Navigation with Human-Environment-Robot Ternary Fusion
Siyi Lu, Bolei Chen, Ping Zhong 0002, Yu Sheng, Yongzheng Cui, Run Liu 0001 |
ICONIP (4) | 4 |
| 2023 | CATAD: exploring topologically associating domains from an insight of core-attachment structureabstractIdentifying topologically associating domains (TADs), which are considered as the basic units of chromosome structure and function, can facilitate the exploration of the 3D-structure of chromosomes. Methods have been proposed to identify TADs by detecting the boundaries of TADs or identifying the closely interacted regions as TADs, while the possible inner structure of TADs is seldom investigated. In this study, we assume that a TAD is composed of a core and its surrounding attachments, and propose a method, named CATAD, to identify TADs based on the core-attachment structure model. In CATAD, the cores of TADs are identified based on the local density and cosine similarity, and the surrounding attachments are determined based on boundary insulation. CATAD was applied to the Hi-C data of two human cell lines and two mouse cell lines, and the results show that the boundaries of TADs identified by CATAD are significantly enriched by structural proteins, histone modifications, transcription start sites and enzymes. Furthermore, CATAD outperforms other methods in many cases, in terms of the average peak, boundary tagged ratio and fold change. In addition, CATAD is robust and rarely affected by the different resolutions of Hi-C matrices. Conclusively, identifying TADs based on the core-attachment structure is useful, which may inspire researchers to explore TADs from the angles of possible spatial structures and formation process. Xiaoqing Peng, Mengxi Zou, Xiangyan Kong, Yu Sheng |
Briefings Bioinform. | 5 |
| 2022 | Research on the Prediction Method of Disease Classification Based on Imaging Features
Yu Sheng, Shengyi Yang, Huirong Hu, Guihua Duan |
ISBRA | 1 |
| 2021 | MAIN: Multimodal Attention-based Fusion Networks for Diagnosis PredictionabstractPredicting the future diagnoses from patients’ historical Electronic Health Records (EHR) is a significant task in healthcare. EHR consist of multiple modal data, each modality has different features and contains a wealth of information of patients. However, most of the existing EHR-based prediction methods either only use unimodal data, or fail to fully explore the correlation between different modalities when fusing multimodal data. To address these challenges, we propose a Multimodal Attention-based fusIon Networks (MAIN) for diagnosis prediction. In this model, we first design different feature extraction modules for each modality. Then, an inter-modal correlation module which contains two layers is applied to capture the intermodal correlation. Finally, a multimodal fusion module based on weighted averaging is utilized to integrate the representations derived from different modalities and their correlation to obtain the patient representation for diagnosis prediction. We evaluate our proposed model on two medical datasets, and the experimental results demonstrate the effectiveness of MAIN. Ying An, Haojia Zhang, Yu Sheng, Jianxin Wang 0001, Xianlai Chen |
BIBM | 3 |
| 2021 | Backdoor Attack of Graph Neural Networks Based on Subgraph Trigger
Yu Sheng, Guanyu Cai, Li Kuang |
CollaborateCom (2) | 1 |
| 2021 | Space-Heuristic Navigation and Occupancy Map Prediction for Robot Autonomous Exploration
Ping Zhong 0002, Bolei Chen, Yongzheng Cui, Hanchen Song, Yu Sheng |
ICA3PP (1) | 5 |
| 2021 | EnvFaker: A Method to Reinforce Linux Sandbox Based on Tracer, Filter and Emulator against Environmental-Sensitive MalwareabstractSandbox is an excellent tool for dynamic malware analysis. However, the sandbox detection techniques are increasingly adopted to develop malwares, which has been a significant threat to sandbox analysis. These malwares can detect the running environment and show different behaviors in corresponding environments. So far, there have been several studies about countermeasures, but most of them concentrate on Windows OS. Environmental features in Linux sandbox have not been summarized yet. Besides, existing popular sandboxes can hardly combat against sandbox detecting techniques. In this paper, we focus on Linux sandbox. We firstly propose Linux environmental features from six aspects and implement an effective tool to collect features from running environment to tell the discrepancy among physical machine, virtual machine and sandbox. More importantly, we present EnvFaker, an effective method to reinforce Linux sandbox against environmental-sensitive malware. This method uses tracer to track child process and injected process, filters to intercept sandbox detecting behaviors, and emulator to disguise wear-and-tear and network environment. The experimental results further demonstrate that our method is effective against detecting techniques for Linux sandbox. Chenglin Xie, Shaosen Shi, Yu Sheng, Xiarun Chen, Weiping Wen |
TrustCom | 4 |
| 2021 | Virtual Traffic Signals: Safe, Rapid, Efficient and Autonomous Driving Without Traffic ControlabstractConnected and autonomous vehicles (CAV) will open the future to vast possibilities in transportation, most of which have yet to be imagined. As these new technologies emerge, they also have the potential to render many of today’s most common transportation control methods irrelevant. In this paper, concepts of connected and autonomous are applied to coordinate and synchronize the arrival and departure of vehicles at intersections; effectively ending the need for traffic signals. Under a set of “autonomous rules of operation”, vehicles can be kept in near-continuous motion while also maintaining safe separation, creating a system of “virtual traffic signal” (VTS). Assessments of these rules shows that delay and stopping reductions of more than 50 to 97 percent will be possible over actuated control and fixed time signal control – and without the risk of collisions. Not will these rules reduce congestion, travel time, and prevent crashes; they will also lower fuel consumption and exhaust emissions. Zhao Zhang 0014, Feng Liu 0058, Brian Wolshon, Yu Sheng |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2020 | Improved ASD classification using dynamic functional connectivity and multi-task feature selection
Jin Liu 0012, Yu Sheng, Wei Lan 0001, Rui Guo 0009, Jianxin Wang 0001 |
Pattern Recognit. Lett. | 2 |
| 2015 | A Similarity Detection Platform for Programming LearningabstractCode similarity detection has been studied for several decades, which are prevailing categorized into attributecounting
and structure-metric. Due to the one fold validity of attribute-counting for full replication, mature
systems usually use the GST string matching algorithm to detect code structure. However, the accuracy of
GST is vulnerable to interference in code similarity detection. This paper presents a code similarity detection
method combining string matching and sub-graph isomorphism. The similarity is calculated with the GST
algorithm. Then according to the similarity, the system determines whether further processing with the sub-graph
iIsomorphism algorithm is required. Extensive experimental results illustrate that our method significantly
enhances the efficiency of string matching as well as the accuracy of code similarity detecting. Yu Sheng |
CSEDU (1) | 2 |
| 2015 | Virtual laboratory platform for computer science curriculaabstractExperimenting in computer science course is challenging due to the limitation of site, equipment and special experiment tools. In this paper, based on the analysis of the experiments features of computer science curricula, such as, Principle of Computer Organization, Digital Image Process, Digital Signal Process etc., we design two kinds of virtual lab platforms and develop corresponding virtual lab systems for the courses in computer science curricula. In the first virtual lab, every experiment instrument in real lab is visualized as a Java component. In the other kind of virtual lab, the algorithm students learn in the course is packed as a Web Service component by C, C++ or Java. In both platforms, those components are listed in the system. Students is able to select Java components or Web Service components as the experiment they want to do need and integrate them by building connections between them. Then, the students can set input for the experiment in the input component. After click the button of rum, the result will be display for the students in the platform. In the two kinds of virtual lab, students can also write the code in the platform. After it is submitted, those codes will be transferred as a component by the platform and be added in the component list. Then they can use them as the components provide by the platform. So, they can test if the algorithm they write is correct. Both platforms are developed by Java Applet and can be run by a browse. By using our system, students can experiment at any time and any place via the Internet. Teacher is also able to do experiment in the classroom. Base on the platforms, 6 virtual lab systems have been developed and used by more than 5 universities in China. Those virtual lab systems have been received favorably by teachers and students. Yu Sheng |
FIE | 3 |
| 2014 | Style-based human motion segmentationabstractThis paper presents a method for segmenting human motion based on a notion of quality and the movement of a user such that the exact segmentation is tailored for different subjects. The problem is solved via an inverse optimal control problem where the parameter of optimization is a time along the movement trajectory that splits the longer trajectory into distinct “moves.” First, trajectories are generated using a “forward” optimal control problem; then, the match of these generated trajectories is optimized via a second, “inverse” optimization, which determines the appropriate point of segmentation. An analytical solution to this set up, its numerical implementation, and an application to real data are presented. A key novel contribution of this paper is the analytical derivation of first order necessary conditions for optimality. The segmented movements may populate a library of movement primitives in order for robots and automated systems to perform and interpret novel tasks. Yu Sheng, Amy LaViers |
SMC | 1 |
| 2014 | Translucent Radiosity: Efficiently CombiningDiffuse Inter-Reflection andSubsurface ScatteringabstractIt is hard to efficiently model the light transport in scenes with translucent objects for interactive applications. The inter-reflection between objects and their environments and the subsurface scattering through the materials intertwine to produce visual effects like color bleeding, light glows, and soft shading. Monte-Carlo based approaches have demonstrated impressive results but are computationally expensive, and faster approaches model either only inter-reflection or only subsurface scattering. In this paper, we present a simple analytic model that combines diffuse inter-reflection and isotropic subsurface scattering. Our approach extends the classical work in radiosity by including a subsurface scattering matrix that operates in conjunction with the traditional form factor matrix. This subsurface scattering matrix can be constructed using analytic, measurement-based or simulation-based models and can capture both homogeneous and heterogeneous translucencies. Using a fast iterative solution to radiosity, we demonstrate scene relighting and dynamically varying object translucencies at near interactive rates. Yu Sheng, Yulong Shi, Lili Wang 0006, Srinivasa G. Narasimhan |
IEEE Trans. Vis. Comput. Graph. | 1 |
| 2013 | A practical analytic model for the radiosity of translucent scenesabstractLight propagation in scenes with translucent objects is hard to model efficiently for interactive applications. The inter-reflections between objects and their environments and the subsurface scattering through the materials intertwine to produce visual effects like color bleeding, light glows and soft shading. Monte-Carlo based approaches have demonstrated impressive results but are computationally expensive, and faster approaches model either only inter-reflections or only subsurface scattering. In this paper, we present a simple analytic model that combines diffuse inter-reflections and isotropic subsurface scattering. Our approach extends the classical work in radiosity by including a subsurface scattering matrix that operates in conjunction with the traditional form-factor matrix. This subsurface scattering matrix can be constructed using analytic, measurement-based or simulation-based models and can capture both homogeneous and heterogeneous translucencies. Using a fast iterative solution to radiosity, we demonstrate scene relighting and dynamically varying object translucencies at near interactive rates. Yu Sheng, Yulong Shi, Lili Wang 0006, Srinivasa G. Narasimhan |
I3D | 1 |
| 2013 | Non-polynomial Galerkin projection on deforming meshesabstractThis paper extends Galerkin projection to a large class of non-polynomial functions typically encountered in graphics. We demonstrate the broad applicability of our approach by applying it to two strikingly different problems: fluid simulation and radiosity rendering, both using deforming meshes. Standard Galerkin projection cannot efficiently approximate these phenomena. Our approach, by contrast, enables the compact representation and approximation of these complex non-polynomial systems, including quotients and roots of polynomials. We rely on representing each function to be model-reduced as a composition of tensor products, matrix inversions, and matrix roots. Once a function has been represented in this form, it can be easily model-reduced, and its reduced form can be evaluated with time and memory costs dependent only on the dimension of the reduced space. Matt Stanton, Yu Sheng, Martin Wicke, Federico Perazzi, Amos Yuen, Srinivasa G. Narasimhan, Adrien Treuille |
ACM Trans. Graph. | 2 |
| 2011 | Modeling permafrost distribution using remote sensing-derived vegetation data in the source region of the Datong River in the northwestern ChinaabstractThe source region of the Datong River is an inland watershed and located at the northeastern edge of the Qinghai-Tibetan Plateau. There is relatively plentiful rainfall in the source region. Vegetation, mainly composed of alpine meadow and alpine swamp meadow, developed well and covered mostly the ground surface of the source region. The equivalent-elevation approach has been proved a valid method to evaluate the double effects of latitude and altitude on permafrost in the Qilianshan Mountains, northwestern China. This method was used in this research. The field investigation indicated that vegetation class was an important local factor determining permafrost development and distribution in the study area. According to the vegetation samplings, the Landsat TM images of the source region were interpreted using the maximum likelihood algorism. Ground surface of the study area were classified into three classes, alpine meadow, alpine swamp meadow and the bare ground. According to the classification, boreholes were also divided into three classes correspondingly. As the small amount of boreholes drilled in the bare ground, permafrost distribution in the vegetated areas was studied. By setting up different datum point suitable for different vegetated covering areas, equivalent elevations of each borehole were calculated. As far as the ground temperatures of permafrost are concerned, the ground temperatures at 15m depth were usually chosen as the proxy of mean annual ground temperature (MAGT) because of the index has a tiny seasonal variation although the large seasonal changes of ambient temperatures. Permafrost models, between permafrost MAGT and the calculated equivalent-elevations, in different vegetated areas were constructed. Using the GIS software combined with the DEM data, permafrost MAGT of the vegetated areas in the source region were computed and mapped. The calculation results indicated that permafrost MAGT varied from -3.1°C to 1.1°C and from -4.6°C to 1.7°C in the alpine swampy meadow and alpine meadow vegetation areas, respectively. The vegetated areas had a variation of permafrost MAGT from -4.6°C to 1.7°C. In the vegetated areas, the areal percentages of permafrost and the seasonally frozen ground were 94% and 6%. Yu Sheng, Xiumin Zhang, Jichun Wu, Yuanbing Cao |
IGARSS | 2 |
| 2011 | Changes of vegetation biomass, species diversity and NDVI across the ground temperature of permafrost in the Datong river source region on the northeastern edge of the Qinghai-Tibet plateauabstractA primary analysis was applied to analyze the effects of degradation of frozen soil on the alpine vegetation ecosystem based on the investigation data from 92 vegetation plots sampled in 2010 and 30m- normalized difference vegetation index (NDVI) data of these plots across the ground temperature of permafrost in the Datong river source region, on the northeastern edge of the Qinghai-Tibet plateau in China. The results show that: with the increase of the ground temperature, the species diversity index of plant communities had an ascendant trend and the succession of vegetation type had a change, like the wet plant species in communities decreased gradually and mesophyte plants expanded rapidly. When the permafrost type from sub-stable permafrost type to transitional type and then to unstable permafrost type, the phytomass had decrease trend; NDVI and soil moisture had the biggest value in the transitional permafrost region. Xiumin Zhang, Yu Sheng, Jichun Wu |
IGARSS | 2 |
| 2011 | Perceptual Global Illumination Cancellation in Complex Projection EnvironmentsabstractAbstract The unintentional scattering of light between neighboring surfaces in complex projection environments increases the brightness and decreases the contrast, disrupting the appearance of the desired imagery. To achieve satisfactory projection results, the inverse problem of global illumination must be solved to cancel this secondary scattering. In this paper, we propose a global illumination cancellation method that minimizes the perceptual difference between the desired imagery and the actual total illumination in the resulting physical environment. Using Gauss‐Newton and active set methods, we design a fast solver for the bound constrained nonlinear least squares problem raised by the perceptual error metrics. Our solver is further accelerated with a CUDA implementation and multi‐resolution method to achieve 1–2 fps for problems with approximately 3000 variables. We demonstrate the global illumination cancellation algorithm with our multi‐projector system. Results show that our method preserves the color fidelity of the desired imagery significantly better than previous methods. Yu Sheng, Barbara Cutler, Chao Chen 0012, Joshua D. Nasman |
Comput. Graph. Forum | 1 |
| 2011 | A Spatially Augmented Reality Sketching Interface for Architectural Daylighting DesignabstractWe present an application of interactive global illumination and spatially augmented reality to architectural daylight modeling that allows designers to explore alternative designs and new technologies for improving the sustainability of their buildings. Images of a model in the real world, captured by a camera above the scene, are processed to construct a virtual 3D model. To achieve interactive rendering rates, we use a hybrid rendering technique, leveraging radiosity to simulate the interreflectance between diffuse patches and shadow volumes to generate per-pixel direct illumination. The rendered images are then projected on the real model by four calibrated projectors to help users study the daylighting illumination. The virtual heliodon is a physical design environment in which multiple designers, a designer and a client, or a teacher and students can gather to experience animated visualizations of the natural illumination within a proposed design by controlling the time of day, season, and climate. Furthermore, participants may interactively redesign the geometry and materials of the space by manipulating physical design elements and see the updated lighting simulation. Yu Sheng, Theodore C. Yapo, Christopher Young, Barbara Cutler |
IEEE Trans. Vis. Comput. Graph. | 1 |
| 2010 | Primary analysis on distribution characteristics of permafrost in the upper area of the Shule River Watershed, on the northeastern edge of the Qinghai-Tibetan PlateauabstractTaking the upstream of the Shule River Watershed, on the northeastern edge of the Qinghai-Tibetan Plateau in China as the study area, a primary analysis on the effects of influencing factors on the distribution of permafrost was carried out. Relationships between the ground temperatures of permafrost and the geographical and topographic factors longitude, latitude, elevation, slope and aspect were analyzed using the statistical method. An empirical-statistical model to be used to modeling the distribution of permafrost was developed, which took the factors latitude, elevation, slope and aspect as input variables, and the ground temperatures of permafrost as the output variable. Results of the calculation indicated that 83% of the upstream was underlain by the perennially frozen ground and the rest 17% by the seasonally frozen ground. Classify permafrost based on the calculated temperatures into four types: low-temperature permafrost (GT ≤-2°C), middle-temperature permafrost (-2°C<; GT ≤-1°C), high-temperature permafrost (-1°C<; GT ≤-5°C) and extremely-high-temperature permafrost (-0.5°C<; GT ≤0°C). Areal percentages of each type were 38%, 23%, 14%, and 8% in order. Yu Sheng, Huijun Jin, Jichun Wu, Baisheng Ye |
IGARSS | 1 |
| 2010 | Analysis of time-series modis 250M vegetation index data for vegetation classifiation in the wenquan area over the qinghai-tibet plateauabstractA multi-temporal analysis and statistical analysis were applied to examine vegetation index (VI) based vegetation separability based on investigation data from 99 vegetation plots sampled in 2009 and corresponding 10-year timeseries of MODIS EVI and NDVI data of these plots in the Wenquan area, a transition area between permafrost and seasonally frozen soil on the eastern Qinghai-Tibet plateau. The analyses of multi-temporal VI dataset show similar phenological characteristics of all vegetation types during the entire growth season, while they have distinguishable VI values at different periods, namely germination, maturity and senescence periods. Spectral separability between every two vegetation types identified by the Jeffries-Matusita distance statistic is not significant in the growing season, indicating the vegetation classes of this study area not spectrally separable. High intra-class variability of each vegetation type is the primary cause of low separability in different growth periods. Xiumin Zhang, Zhuotong Nan, Yu Sheng, Lin Zhao 0013, Guoying Zhou, Guangyang Yue, Jichun Wu |
IGARSS | 3 |
| 2010 | Global Illumination Compensation for Spatially Augmented RealityabstractAbstract When projectors are used to display images on complex, non‐planar surface geometry, indirect illumination between the surfaces will disrupt the final appearance of this imagery, generally increasing brightness, decreasing contrast, and washing out colors. In this paper we predict through global illumination simulation this unintentional indirect component and solve for the optimal compensated projection imagery that will minimize the difference between the desired imagery and the actual total illumination in the resulting physical scene. Our method makes use of quadratic programming to minimize this error within the constraints of the physical system, namely, that negative light is physically impossible. We demonstrate our compensation optimization in both computer simulation and physical validation within a table‐top spatially augmented reality system. We present an application of these results for visualization of interior architectural illumination. To facilitate interactive modifications to the scene geometry and desired appearance, our system is accelerated with a CUDA implementation of the QP optimization method. Yu Sheng, Theodore C. Yapo, Barbara Cutler |
Comput. Graph. Forum | 1 |
| 2009 | Polymorphic Worm Detection Using Signatures Based on Neighborhood RelationabstractIn recent years, worm signatures suffer from difficulties to detect polymorphic worms because these worms can change their patterns dynamically. In this paper, a class of neighborhood-relation signatures (NRS) are proposed, including 1-NRS, 2-NRS and (1,2)-NRS. NRS can be used for detecting polymorphic worms since these worms often remain the same relationship between bytes in changing their patterns. Two signature generation algorithm based on expectation-maximization (EM) and Gibbs Sampling are designed to generate NRS. We perform extensive experiments to demonstrate the effectiveness of NRS and the correctness of the process of signatures generation. Experiment results show that our approach of defending polymorphic worm based on NRS is more effective than other approach based on existed signatures. Jie Wang 0067, Jianxin Wang 0001, Yu Sheng, Jianer Chen |
HPCC | 3 |
| 2009 | VL-DSC: A Dynamic Service Composition Based Model for Virtual Laboratory Platform and Its ImplementationabstractWeb-based Virtual Laboratory is emerging as a promising teaching assistant tool for modern distance education and attracts many attentions from both of academia and industry. In this paper, a novel virtual laboratory platform based on dynamic service composition, which is named as VL-DSC, is presented. In this platform, experimental process can be customized by service virtualization. A domain-oriented service composition description language is designed to describe the relationship between heterogeneous virtual instruments. Furthermore, based on the description language, dynamic message routing is adopted to realize dynamic schedule and replacement of virtual instruments. VL-DSC can significantly improve the customizability and dynamic interoperability of virtual lab, which can be used to develop virtual lab easily and quickly. Jianxin Wang 0001, Yu Sheng, Songqiao Chen, Zhaohui Xie |
HPCC | 3 |
| 2009 | Analysis on Factors Affecting the Development of Alpine Permafrost in Central-Eastern Qilianshan Mountains, Northwest ChinaabstractUsing data from 190 boreholes drilled in 2004, an analysis of the factors affecting the development of alpine permafrost was carried out using the statistical analysis software of SPSS. The factors considered in this case were the terrain factors of elevation, slope, aspect, curvature, plan curve and profile curve, the climatic factors of latitude, longitude, and the potential incoming solar radiation, as well as the topographic wetness index and the vegetation abundance indicating factor NDVI. The results indicated that longitude had the most significant negative effect on the presence of permafrost. Elevation, the potential incoming solar radiation and latitude had significant positive correlations with the occurrence of permafrost. Yu Sheng, Shixing Jiao, Guojing Yang |
IGARSS (2) | 2 |
| 2009 | Virtual Heliodon: Spatially Augmented Reality for Architectural Daylighting DesignabstractWe present an application of interactive global illumination and spatially augmented reality to architectural daylight modeling that allows designers to explore alternative designs and new technologies for improving the sustainability of their buildings. Images of a model in the real world, captured by a camera above the scene, are processed to construct a virtual 3D model. To achieve interactive rendering rates, we use a hybrid rendering technique, leveraging radiosity to simulate the inter-reflectance between diffuse patches and shadow volumes to generate per-pixel direct illumination. The rendered images are then projected on the real model by four calibrated projectors to help users study the daylighting illumination. The virtual heliodon is a physical design environment in which multiple designers, a designer and a client, or a teacher and students can gather to experience animated visualizations of the natural illumination within a proposed design by controlling the time of day, season, and climate. Furthermore, participants may interactively redesign the geometry and materials of the space by manipulating physical design elements and see the updated lighting simulation. Yu Sheng, Theodore C. Yapo, Christopher Young, Barbara Cutler |
VR | 1 |
| 2008 | Analysis of Novel User Detection Scheme Based on Polling for E-MBMS NetworksabstractMultimedia broadcast/multicast service (MBMS) is an important part of the UTRAN evolution and supports downlink streaming and download-and-play type services to large groups of users. For enhanced MBMS (E-MBMS) under 3GPP long term evolution (LTE) system, there are two ways to transmissions being performed: multi-cell transmissions and single-cell transmissions. One requirement identified to be supported is that the capability of the network to detect at least one MBMS user interested receiving one given MBMS service in the cell which belongs to above scenarios. It is significant to avoid unnecessary MBMS transmission in a cell where there is no MBMS user especially for single-cell transmissions mode. In this investigation, a low complex method is discussed to solve the detection on MBMS interested users in one cell and we also introduce code diversity strategy into the feedback signal transmission. Theoretical analysis and simulation results all show that the efficiency of proposed detection scheme is obvious which can dramatically reduce the average MBMS service polling time and control overhead. Yu Sheng, Mugen Peng, Wenbo Wang 0007 |
VTC Fall | 1 |
| 2005 | Awareness Scheduling and Algorithm Implementation for Collaborative Virtual Environment
Yu Sheng, Dongming Lu, Yifeng Hu, Qingshu Yuan |
ICCSA (1) | 1 |
| 2005 | Motion Prediction in a High-Speed, Dynamic EnvironmentabstractThe immanent existence of system latency greatly affects the control behavior of a closed-loop system. In order to reduce the influence induced by latency, this paper proposes a systematic method based on neural network to predict the motion of objects in a high-speed, dynamic, and competitive environment. We apply this method to the competition of RoboCup Small Size League, which greatly improves the performance of our control system. Yu Sheng, Yonghai Wu |
ICTAI | 1 |