VLDB 2026 Research / reviewers in the wild / expert
Xinyu Zhang 0001
dblp:58/4582-1
· DBLP profile ↗
45ranked-venue papers
6as first author
35since 2021 · last 2026
0000-0003-0034-9037ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 22 · 3 first-author · 17 since 2021Applied, interdisciplinary, general and emerging computing · 14 · 1 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 8 since 2021Systems, architecture and hardware · 6 · 2 first-author · 5 since 2021Computer networks · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | V 2 -Fusion: Virtual voxel enhanced 4D radar-image feature fusion for 3D object detection
Li Wang 0092, Xinyu Zhang 0001, Yuxuan Fan, Tao Xie 0010, Lei Yang 0060, Bin Xu 0003 |
Expert Syst. Appl. | 3 |
| 2026 | Multimodal Large Language Models for Perception in Autonomous Driving: Architecture, Taxonomy, and ChallengesabstractAutonomous vehicles rely on continuous environmental perception to assess obstacle distribution to ensure safe driving. However, current perception technologies face substantial challenges, particularly under adverse weather conditions and when encountering long-tail scenarios. With the advent of transformer-based attention mechanisms, large language models (LLMs), exemplified by GPT, have exhibited emergent intelligence, offering new possibilities for achieving high-performance perception. This technological advancement has led to the development of multi-modal large language models (MLLMs), which incorporate multi-modal encoders. These models enable a single LLM to process multi-source data while performing advanced understanding and reasoning tasks, enhancing complex environmental perception capabilities. Despite significant progress in MLLMs, there remains a notable gap in systematic research on their optimal application to perception tasks. Therefore, this paper presents a comprehensive survey of recent advancements in MLLM-based perception. First, we introduce the mainstream vision–language perception tasks, widely adopted evaluation metrics, and existing language-enhanced autonomous driving datasets. Next, we outline the general architectural design principles of MLLMs. Subsequently, we provide a taxonomy and indepth analysis of current MLLMs, focusing on three dimensions: input modality, alignment technique, and scene representation, elucidating their underlying implementation paradigms. Finally, we summarize the key challenges and emerging research directions in MLLM-driven perception. This survey aims to facilitate further progress in MLLMs by synthesizing these insights. Xinyu Zhang 0001, Yuchuan Ji, Yanchao Ding, Jialun Yin, Ruizhi Jia, Yijin Xiong, Jun Li 0082, Huaping Liu 0001 |
IEEE Internet Things J. | 2 |
| 2026 | GADet: Geometry-Aware oriented object detection for remote sensing
Xinyu Zhang 0001, Ziying Song, Lei Yang 0060, Haicheng Qu |
Knowl. Based Syst. | 3 |
| 2026 | Diffusion Models for Autonomous Driving in Smart Parking: A Data Synthesis FrameworkabstractThe accelerating urbanization process and the increase in vehicle ownership have increasingly highlighted the issue of parking difficulties. Intelligent parking systems, as an emerging solution to this problem, are centered around the perception tasks of autonomous driving technology within parking lots. However, the complexity of parking environments poses significant challenges to the perception systems of autonomous vehicles. To address this, this paper proposes a data synthesis framework based on diffusion models, aimed at generating high-quality synthetic parking lot datasets to support the development of intelligent parking systems. By constructing standardized parking lot scenarios in simulators, simulated images containing original images and real labels are automatically generated to address the high cost and low efficiency of real data collection. The innovative application of diffusion models to the synthesis of intelligent parking perception data results in the generation of highly realistic images that closely resemble real parking lot scenes, while preserving accurate labeling data. This approach effectively reduces the discrepancy between traditional simulation data and real-world data. Experimental evaluations have demonstrated the superior performance of this approach in terms of visual quality, generation diversity, and object detection performance. This research provides a substantial resource for training autonomous driving perception algorithms and offers an effective means to verify and improve system safety and reliability. Xinyu Zhang 0001, Jialun Yin, Jun Li 0082 |
IEEE Trans Autom. Sci. Eng. | 2 |
| 2026 | IGTPT: Intent-Caption Guided Trajectory Prediction TransformerabstractTrajectory prediction is crucial for autonomous vehicles to make safe and informed decisions. However, the lack of transparency in current trajectory prediction models introduces significant security risks, because their output contains almost no explanatory details. To address these challenges and bridge this research gap, we propose a novel approach IGTPT that not only predicts vehicle trajectories but also generates textual descriptions of the vehicle’s intent. The vehicle’s intent can be broadcast to nearby vehicles to enhance decision-making transparency and increase user trust. Our approach directly confronts the opaque nature of existing systems by providing clear, understandable explanations for autonomous decisions, thereby enhancing both the security and reliability of these systems. We enhanced the BDD-X dataset to create the BDD-XE, a specialized dataset for trajectory prediction and behavior description, which our framework uses to achieve superior results in both prediction accuracy and behavioral interpretation compared to established baseline methods. We demonstrate the practical applicability of our framework through a complete system that processes past raw driving videos and trajectory observations to deliver real-time predictions along with insightful behavioral narrations and reasoning. Boqi Li 0001, Xin Gao 0028, Yiguo Lu, Xingang Wu, Xinyu Zhang 0001 |
IEEE Trans. Intell. Transp. Syst. | 6 |
| 2025 | V2X-Radar: A Multi-modal Dataset with 4D Radar for Cooperative PerceptionabstractModern autonomous vehicle perception systems often struggle with occlusions and limited perception range. Previous studies have demonstrated the effectiveness of cooperative perception in extending the perception range and overcoming occlusions, thereby enhancing the safety of autonomous driving. In recent years, a series of cooperative perception datasets have emerged; however, these datasets primarily focus on cameras and LiDAR, neglecting 4D Radar—a sensor used in single-vehicle autonomous driving to provide robust perception in adverse weather conditions. In this paper, to bridge the gap created by the absence of 4D Radar datasets in cooperative perception, we present V2X-Radar, the first large-scale, real-world multi-modal dataset featuring 4D Radar. V2X-Radar dataset is collected using a connected vehicle platform and an intelligent roadside unit equipped with 4D Radar, LiDAR, and multi-view cameras. The collected data encompasses sunny and rainy weather conditions, spanning daytime, dusk, and nighttime, as well as various typical challenging scenarios. The dataset consists of 20K LiDAR frames, 40K camera images, and 20K 4D Radar data, including 350K annotated boxes across five categories. To support various research domains, we have established V2X-Radar-C for cooperative perception, V2X-Radar-I for roadside perception, and V2X-Radar-V for single-vehicle perception. Furthermore, we provide comprehensive benchmarks across these three sub-datasets. Lei Yang 0060, Xinyu Zhang 0001, Jun Li 0082, Jiaqi Ma 0003, Zhiying Song, Ziying Song, Li Wang 0092, Yang Shen 0005, Chen Lv 0001 |
NeurIPS | 2 |
| 2025 | Cooperative Camera-LiDAR Extrinsic Calibration for Vehicle-Infrastructure Systems in Urban IntersectionsabstractCameras and LiDARs are fundamental sensors for autonomous driving, with wide applicability. This paper comprehensively studies and analyzes the calibration of the multi-terminal camera-LiDAR devices from the perspectives of vehicle, roadside, and road cooperation, and outlines its related applications and far-reaching significance. Different from previous studies that focused on sensor calibration of a single platform, ignoring the differences between vehicle-end and roadside, this paper particularly emphasizes the difference between vehicle-end and roadside, as well as heterogeneous calibration challenges emerging in vehicle-road cooperation scenarios. These challenges include cross-sensor temporal synchronization and spatial alignment in dynamic environments. It reveals the advantages and differences of different methods in dealing with constraints such as the fixed installation postures of vehicle-mounted and roadside devices and the large-scale observation domain. This paper further points out key research gaps, such as the adaptability of online calibration and the maintenance of calibration consistency among multiagents, providing important references for the technological evolution from isolated calibration to a sensor-collaborative calibration framework in traffic scenarios. Yijin Xiong, Xinyu Zhang 0001, Xin Gao 0028, Qianxin Qu, Chun Duan, Jun Li 0082 |
IEEE Internet Things J. | 2 |
| 2025 | BEVHeight++: Toward Robust Visual Centric 3D Object DetectionabstractWhile most recent autonomous driving system focuses on developing perception methods on ego-vehicle sensors, people tend to overlook an alternative approach to leverage intelligent roadside cameras to extend the perception ability beyond the visual range. We discover that the state-of-the-art vision-centric detection methods perform poorly on roadside cameras. This is because these methods mainly focus on recovering the depth regarding the camera center, where the depth difference between the car and the ground quickly shrinks while the distance increases. In this paper, we propose a simple yet effective approach, dubbed BEVHeight++, to address this issue. In essence, we regress the height to the ground to achieve a distance-agnostic formulation to ease the optimization process of camera-only perception methods. By incorporating both height and depth encoding techniques, we achieve a more accurate and robust projection from 2D to BEV spaces. On popular 3D detection benchmarks of roadside cameras, our method surpasses all previous vision-centric methods by a significant margin. In terms of the ego-vehicle scenario, BEVHeight++ surpasses depth-only methods with increases of +2.8% NDS and +1.7% mAP on the nuScenes test set, and even higher gains of +9.3% NDS and +8.8% mAP on the nuScenes-C benchmark with object-level distortion. Consistent and substantial performance improvements are achieved across the KITTI, KITTI-360, and Waymo datasets as well. Lei Yang 0060, Jun Li 0082, Kun Yuan 0001, Li Wang 0092, Yi Huang 0038, Xinyu Zhang 0001, Kaicheng Yu |
IEEE Trans. Pattern Anal. Mach. Intell. | 10 |
| 2025 | FMRT: Learning Accurate Feature Matching With Reconciliatory TransformerabstractLocal Feature Matching, a pivotal component of numerous computer vision tasks (e.g., structure from motion and visual localization), has been effectively addressed by Transformer-based methods. Nevertheless, these methods solely incorporate long-range context information among keypoints with a fixed receptive field, which constrains the network from appropriately reconciling the importance of features with diverse receptive fields to realize complete image perception, hence limiting feature matching accuracy. In addition, these methods employ a conventional handcrafted encoding approach to incorporate positional information of keypoints into visual descriptors, which limits the capability of networks to extract effective positional encoding message. In this study, we propose FMRT, a novel detector-free method that reconciles local features with diverse receptive fields adaptively and utilizes parallel networks to realize reliable positional encoding. Specifically, FMRT proposes a dedicated reconciliatory transformer (RecFormer) that contains a global perception attention layer to identify visual descriptors with different receptive fields and integrate global context information under various scales, a perception weight layer to measure the importance of various receptive fields adaptively, and a local perception feed-forward network to extract deep aggregated multi-scale local feature representation. Moreover, we introduce a novel axis-wise position encoder (AWPE) that views positional encoding as two keypoints encoding tasks along the row and column dimensions, decouples the x- and y-coordinates of keypoints into two independent 1D vectors, and designs two parallel network branches to explicitly encodes geometric correlations among keypoints, hence realizing reliable positional encoding. Extensive experiments indicate that FMRT yields impressive performance on multiple tasks, including relative pose estimation, visual localization, homography estimation, and image matching. Besides, we integrate FMRT into a localization framework and conduct a visual localization experiment in a real scene, which further demonstrate the superiority of FMRT. Note to Practitioners—This paper presents a novel approach to enhancing the performance of local feature matching in computer vision tasks. Traditional methods often rely on fixed receptive fields for integrating context among keypoints, which can limit the perception of the complete image and, consequently, the precision of feature matching. Our work introduces a Reconciliatory Transformer that not only addresses these limitations by effectively reconciling the importance of features across varying receptive fields but also improves the integration of positional information into visual descriptors. The techniques developed here can be adapted to a wide range of systems, e.g., image matching for computer vision and visual localization for autonomous driving, offering practitioners a tool to significantly improve the fidelity of feature matching, which is foundational for accurate interaction with the surrounding environment. Li Wang 0092, Xinyu Zhang 0001, Tao Xie 0010, Lei Yang 0060, Wenhao Yu 0006, Yang Shen 0005, Bin Xu 0003, Jun Li 0082 |
IEEE Trans Autom. Sci. Eng. | 2 |
| 2025 | GF-SLAM: A Novel Hybrid Localization Method Incorporating Global and Arc FeaturesabstractA global and feature-based hybrid algorithm, which integrates global information and feature-based simultaneous localization and mapping (GF-SLAM). This system can operate adaptively when external signals are unstable, thereby avoiding cumulative errors produced by local methods. In agricultural planting bases with abundant circular arc features, the focus is on efficiently exploring the correlations between these features to optimize the robust real-time positioning system for work vehicles. In this process, feature-based SLAM (F-SLAM) is applied to partial positioning using a particle filter. Available global information is then fused using an extended Kalman filter (EKF) for precise positioning and deviation correction in the mapping process, thereby achieving an effective combination of two positioning modes. The proposed model was evaluated using two simulation environments and a comparison with representative techniques. Results showed that GF-SLAM was competitive in normal conditions while requiring fewer computations and significantly reducing the drift in F-SLAM for stable global signals. Switching between these two algorithms eliminated positioning errors to within 1 cm for a test case in which global localization was lost in 54.7% of the route, producing an error within 3 cm. The code will be open source. Note to Practitioners—Our adaptive fusion strategy aims to address challenges in real-world agricultural scenarios where global information may be lacking or unstable. This approach enhances the robustness and reliability of robot positioning in practical applications. We invite practitioners to consider the adaptability of our system to diverse environments, particularly those with limited or fluctuating global information. In this paper, we first introduce the EKF module for global localization, then illustrate the establishment of feature maps and particle filter positioning in F-SLAM. We conduct comparison experiments in two virtual environments and real agricultural planting bases, where approximately half of the route lacks global information. The results demonstrate the benefits of our adaptive fusion strategy in practical positioning applications. We also plan to explore the integration of additional sensors, such as cameras, combined with deep learning, to further improve the efficiency and quality of feature extraction. This extension is aimed at mitigating the issue of positioning failure caused by crop and equipment occlusion in agricultural scenarios. Yijin Xiong, Xinyu Zhang 0001, Wenju Gao, Qianxin Qu, Shichun Guo, Yang Shen 0005, Jun Li 0082 |
IEEE Trans Autom. Sci. Eng. | 2 |
| 2025 | PDDepth: Pose Decoupled Monocular Depth Estimation for Roadside Perception SystemabstractAccurate depth information is crucial for roadside perception in Cooperative Vehicle Infrastructure Systems. Beyond existing radar and LiDAR solutions, monocular depth estimation using surveillance cameras is emerging as a superior approach due to its cost-effectiveness and dense depth output. Unlike onboard cameras, roadside cameras are relatively fixed in position. Many existing monocular depth estimation methods, which do not independently model camera pose, tend to overfit to training data and produce suboptimal results when confronted with slight variations in camera poses, which may be caused by external forces within the same camera or across different cameras. To address this issue, a pose decoupled monocular depth estimation method specifically designed for roadside perception systems is proposed. This method separates depth estimation into two components: a pose-dependent modeling portion that recovers ground depth based on the current camera pose, and a pose-agnostic portion that estimates pixel height relative to the ground plane. Additionally, a knowledge distillation framework is introduced to improve the robustness of the proposed method against variations in roadside cameras. To validate the method, we propose the first open source dataset for roadside monocular depth estimation, DAIR-MDE, and a roadside instance segmentation dataset, DAIR-Ins, both derived from the DAIR dataset. The proposed method demonstrates significant advances over the state-of-the-art methods on DAIR-MDE. The proposed dataset and source code are publicly available athttps://github.com/441599828/PDDepth. Huanan Wang, Xinyu Zhang 0001, Zhengxian Chen, Jun Li 0082, Huaping Liu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2025 | Steering Angle-Guided Multimodal Fusion Lane Detection for Autonomous DrivingabstractLane detection is a critical part of autonomous driving technology. When difficult situations are encountered (i.e., adverse light, severe occlusion), the lane detection task is still challenging. However, previous methods strongly depend on the extracted image features and ignore other features. It is necessary to consider the information from other modalities to assist the model for lane detection, especially in the task of curved lane detection. In this paper, considering that the vehicle steering angle is closely related to the visual feature of lane lines, we propose a novel model named Image-Angle Fusion Network (IAFNet) to solve the lane detection problem by fusing vehicle steering angle features with image features. To make the steering angle features better match the image features, we use the tensor outer product to extend the dimensionality of the steering angle information. A lightweight Image-Angle cross-attention module (LIA-CAM) is proposed to learn the implicit relationship between steering angles and visual features of lane lines, aimed at improving the performance of our model in difficult situations. To guide the network to retain the correct steering angle information, we introduced regression prediction loss of steering angle. Besides, we also released a new dataset based on the Udacity dataset: ImageAngle-Udacity (IA-Udacity) dataset. Extensive experiments on the IA-Udacity dataset show that our method outperforms the current state-of-the-art methods showing both higher efficiency and accuracy. Code and data are available onhttps://github.com/gongyan1/LIA-CAM. Xinyu Zhang 0001, Jianli Lu, Xinmin Jiang, Hao Liu 0114, Zhiwei Li 0011, Li Wang 0092, Qingshan Yang, Xingang Wu |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2025 | SGV3D: Toward Scenario Generalization for Vision-Based Roadside 3D Object DetectionabstractRoadside perception can significantly enhance the safety of autonomous vehicles by extending their perceptual capabilities beyond the visual range and addressing occluded regions. However, current state-of-the-art vision-based roadside detection methods exhibit high accuracy on labeled scenes but perform poorly on new scenes. This limitation arises because roadside cameras remain stationary after installation and can only gather data from a single scene, leading the algorithm to overfit these roadside backgrounds and camera positions. To tackle this issue, we propose an innovativeScenarioGeneralization Framework forVision-based Roadside3DObject Detection, calledSGV3D. Specifically, we utilize a Background-suppressed Module (BSM) to reduce background overfitting in vision-centric pipelines by diminishing background features during the 2D to bird’s-eye-view projection. Furthermore, by introducing the Semi-supervised Data Generation Pipeline (SSDG) that employs unlabeled images from new scenes, we generate diverse foreground instances with varying camera poses, mitigating the risk of overfitting to specific camera positions. Experiments conducted on two large-scale roadside benchmarks demonstrate that SGV3D, with only a minimal increase in latency, effectively improves the scenario generalization capabilities of vision-based roadside 3D object detectors. The code is available here (https://github.com/yanglei18/SGV3D). Lei Yang 0060, Xinyu Zhang 0001, Jun Li 0082, Li Wang 0092, Zhiwei Li 0011, Yang Shen 0005, Chen Lv 0001, Hong Wang 0014 |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2025 | V2X-Reg++: A Real-Time Global Registration Method for Multi-End Sensing System in Urban IntersectionsabstractUrban intersections, dense with pedestrian and vehicular traffic and compounded by positioning signal obstructions, are among the most challenging areas in urban traffic systems. Traditional single-vehicle intelligence systems often perform poorly in such environments due to a lack of global scene observations and the inherent uncertainty in predicting other agents’ intentions. Vehicle-to-Everything (V2X) technology, through real-time communication between vehicles (V2V) and vehicles to infrastructure (V2I), offers a robust solution. However, practical applications still face numerous challenges. Spatial registration among vehicle and infrastructure endpoints with different configurations in multi-end sensing systems is crucial for ensuring the accuracy of perception system data. Most existing multi-end spatial registration methods rely on initial extrinsic values provided by positioning systems, but the instability of GNSS signals due to high buildings in urban canyons poses severe challenges to these methods. To address this issue, this paper proposes a novel multi-end spatial registration method that does not require positioning priors to determine initial external parameters and meets real-time requirements. Our method introduces an innovative multi-end perception object association technique that leverages a newOverall Distance(oDist) metric to measure the spatial association between perception objects, subsequently using this metric as the foundation for an optimal transport formulation. By this means, we can extract co-observed targets from object association results for further external parameter computation and optimization. Extensive comparative and ablation experiments conducted on the simulated dataset V2X-Sim and the real dataset DAIR-V2X confirm the effectiveness and efficiency of our method. The code for this method can be accessed at:https://github.com/MassimoQu/v2i-calib. Xinyu Zhang 0001, Qianxin Qu, Yijin Xiong, Chen Xia, Ziqiang Song, Kang Liu 0008, Jun Li 0082, Keqiang Li 0002 |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2024 | Efficient Multi-Scale Network with Learnable Discrete Wavelet Transform for Blind Motion DeblurringabstractCoarse-to-fine schemes are widely used in traditional single-image motion deblur; however, in the context of deep learning, existing multi-scale algorithms not only require the use of complex modules for feature fusion of low-scale RGB images and deep semantics, but also manually generate low-resolution pairs of images that do not have sufficient confidence. In this work, we propose a multi-scale network based on single-input and multiple-outputs(SIMO) for motion deblurring. This simplifies the complexity of algorithms based on a coarse-to-fine scheme. To alleviate restoration defects impacting detail information brought about by using a multi-scale architecture, we combine the characteristics of real-world blurring trajectories with a learnable wavelet transform module to focus on the directional continuity and frequency features of the step-by-step transitions between blurred images to sharp images. In conclusion, we propose a multi-scale network with a learnable discrete wavelet transform (MLWNet), which exhibits state-of-the-art performance on multiple real-world deblurred datasets, in terms of both subjective and objective quality as well as computational efficiency. Our code is available on https://github.com/thqiu0419/MLWNet. Xin Gao 0028, Tianheng Qiu, Xinyu Zhang 0001, Hanlin Bai, Kang Liu 0008, Hu Wei, Guoying Zhang, Huaping Liu 0001 |
CVPR | 3 |
| 2024 | Stimulate the Potential of Robots via CompetitionabstractIt is common for us to feel pressure in a competition environment, which arises from the desire to obtain success comparing with other individuals or opponents. Although we might get anxious under the pressure, it could also be a drive for us to stimulate our potentials to the best in order to keep up with others. Inspired by this, we propose a competitive learning framework which is able to help individual robot to acquire knowledge from the competition, fully stimulating its dynamics potential in the race. Specifically, the competition information among competitors is introduced as the additional auxiliary signal to learn advantaged actions. We further build a Multiagent-Race environment, and extensive experiments are conducted, demonstrating that robots trained in competitive environments outperform ones that are trained with SoTA algorithms in single robot environment. Kangyao Huang, Di Guo 0002, Xinyu Zhang 0001, Xiangyang Ji, Huaping Liu 0001 |
ICRA | 3 |
| 2024 | CompetEvo: Towards Morphological Evolution from Competition
Kangyao Huang, Di Guo 0002, Xinyu Zhang 0001, Xiangyang Ji, Huaping Liu 0001 |
IJCAI | 3 |
| 2024 | Si-GAIS: Siamese Generalizable-Attention Instance Segmentation for Intersection Perception SystemabstractInstance segmentation of traffic participants using vision-based techniques serves as a cornerstone for numerous intelligent transportation systems. Although existing deep learning-based methods have made significant advancements in this field, these algorithms still present considerable challenges with regard to generalizability for commercial deployment. Specifically, the mean average perception (mAP) of these algorithms degrades rapidly when dealing with intersections that are not included in the training set. To address this limitation, a novel instance segmentation approach named Si-GAIS is proposed, which incorporates a siamese structure for the first time with the proposed Generalizable-Attention Encoder (GA). Through the proposed Foreground-Background Fusion Unit (FBF) within GA, efficient feature-level fusion for the foreground and background images is achieved. Additionally, the Interpretable Attention Neck (IA) in GA enables the feature encoder to focus exclusively on the foreground traffic participants while ignoring various backgrounds. To utilize Si-GAIS, an unsupervised method named P-DBSCAN is proposed to obtain high-quality background image for each intersection with slow-moving traffic and camera jitters. Finally, the first multi-intersection multi-category instance segmentation datasets named RopeIns is proposed for validation. Si-GAIS achieves a 7.7% mAP (All APs used in this paper are abbreviations of AP$_{\textit {50}}$the same with PASCAL VOC.) accuracy improvement compared to the state-of-the-art (SOTA) methods while using fewer parameters, with only a 6.4% decline in AP for car segmentation in unseen intersections and weather conditions, whereas all other SOTA methods decline more than 10%. The proposed dataset and source code are publicly available athttps://github.com/441599828/SiGAISand we hope Si-GAIS will be a new baseline for IPS instance segmentation research. Huanan Wang, Xinyu Zhang 0001, Hong Wang 0014, Jun Li 0082 |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2024 | MonoGAE: Roadside Monocular 3D Object Detection With Ground-Aware EmbeddingsabstractAlthough the majority of recent autonomous driving systems concentrate on developing perception methods based on ego-vehicle sensors, there is an overlooked alternative approach that involves leveraging intelligent roadside cameras to help extend the ego-vehicle perception ability beyond the visual range. We discover that most existing monocular 3D object detectors rely on the ego-vehicle prior assumption that the optical axis of the camera is parallel to the ground. However, the roadside camera is installed on a pole with a pitched angle, which makes the existing methods not optimal for roadside scenes. In this paper, we introduce a novel framework for Roadside Monocular 3D object detection with ground-aware embeddings, named MonoGAE. Specifically, the ground plane is a stable and strong prior knowledge due to the fixed installation of cameras in roadside scenarios. In order to reduce the domain gap between the ground geometry information and high-dimensional image features, we employ a supervised training paradigm with a ground plane to predict high-dimensional ground-aware embeddings. These embeddings are subsequently integrated with image features through cross-attention mechanisms. Furthermore, to improve the detector’s robustness to the divergences in cameras’ installation poses, we replace the ground plane depth map with a novel pixel-level refined ground plane equation map. Our approach demonstrates a substantial performance advantage over all previous monocular 3D object detectors on widely recognized 3D detection benchmarks for roadside cameras. The code and pre-trained models will be released soon. Lei Yang 0060, Xinyu Zhang 0001, Jun Li 0082, Li Wang 0092, Yi Huang 0038, Hong Wang 0014 |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2024 | Auto-Points: Automatic Learning for Point Cloud Analysis With Neural Architecture SearchabstractPure point-based neural networks have recently shown tremendous promise for point cloud tasks, including 3D object classification, 3D object part segmentation, 3D semantic segmentation, and 3D object detection. Nevertheless, it is a laborious process to construct a network for each task due to the artificial parameters and hyperparameters involved, e.g., the depths and widths of the network and the number of sampled points at each stage. In this work, we propose Auto-Points, a novel one-shot search framework that automatically seeks the optimal architecture configuration for point cloud tasks. Technically, we introduce a set abstraction mixer (SAM) layer that is capable of scaling up flexibly along the depth and width of the network. Each SAM layer consists of numerous child candidates, which simplifies architecture search and enables us to discover the optimum design for each point cloud task pursuant to resource constraint from an enormous search space. To fully optimize the child candidates, we develop a weight-entwinement neural architecture search (NAS) technique that entwines the weights of different candidates in the same layer during supernet training such that all candidates can be extremely optimized. Benefiting from the proposed techniques, the trained supernet allows the searched subnets to be exceptionally well-optimized without further retraining or finetuning. In particular, the searched models deliver superior performances on multiple extensively employed benchmarks, 93.9% overall accuracy (OA) on ModelNet40, 89.1% OA on ScanObjectNN, 87.1% instance average IoU on ShapeNetPart, 69.1% mIoU on S3DIS, 70.4% [email protected] on ScanNet V2, and 64.4% [email protected] on SUN RGB-D. Li Wang 0092, Tao Xie 0010, Xinyu Zhang 0001, Linqi Yang, Yilong Ren, Haiyang Yu 0002, Jun Li 0082, Huaping Liu 0001 |
IEEE Trans. Multim. | 3 |
| 2024 | FARP-Net: Local-Global Feature Aggregation and Relation-Aware Proposals for 3D Object DetectionabstractIn this work, we introduce FARP-Net, an adaptive local-global feature aggregation and relation-aware proposal network for high-quality 3D object detection from pure point clouds. Our key insight is that learning adaptive local-global feature aggregation from an irregular yet sparse point cloud and generating superb proposals are both pivotal for detection. Technically, we propose a novel local-global feature aggregation layer (LGFAL) that fully exploits the complementary correlation between local features and global features, and fuses their strengths adaptively via an attention-based fusion module. Furthermore, we incorporate a lightweight feature affine module (LFAM) into LGFAL to map the local features into a normal distribution, thus acquiring fine-grained features of each local region in a weight-sharing manner. During object proposal generation, we propose a weighted relation-aware proposal module (WRPM) that uses an objectness-aware formalism to weigh the relation importance among object candidates for a clear and principal context, thereby facilitating the generation of high-quality proposals. The WRPM challenges the traditional practice of extracting contextual information among all object candidates, which is inefficient as object candidates are always noisy and redundant. Experimentally, FARP-Net delivers superior performance on two widely used benchmarks with fewer parameters, 64.0% [email protected] on the SUN RGB-D dataset and 70.9% [email protected] on the ScanNet V2 dataset. We further validate that the proposed LGFAL and WRPM can be integrated into both indoor and outdoor detectors to boost performance. Tao Xie 0010, Li Wang 0092, Ke Wang 0028, Ruifeng Li 0001, Xinyu Zhang 0001, Linqi Yang, Huaping Liu 0001, Jun Li 0082 |
IEEE Trans. Multim. | 5 |
| 2024 | Informative Data Selection With Uncertainty for Multimodal Object DetectionabstractNoise has always been nonnegligible trouble in object detection by creating confusion in model reasoning, thereby reducing the informativeness of the data. It can lead to inaccurate recognition due to the shift in the observed pattern, that requires a robust generalization of the models. To implement a general vision model, we need to develop deep learning models that can adaptively select valid information from multimodal data. This is mainly based on two reasons. Multimodal learning can break through the inherent defects of single-modal data, and adaptive information selection can reduce chaos in multimodal data. To tackle this problem, we propose a universal uncertainty-aware multimodal fusion model. It adopts a multipipeline loosely coupled architecture to combine the features and results from point clouds and images. To quantify the correlation in multimodal information, we model the uncertainty, as the inverse of data information, in different modalities and embed it in the bounding box generation. In this way, our model reduces the randomness in fusion and generates reliable output. Moreover, we conducted a completed investigation on the KITTI 2-D object detection dataset and its derived dirty data. Our fusion model is proven to resist severe noise interference like Gaussian, motion blur, and frost, with only slight degradation. The experiment results demonstrate the benefits of our adaptive fusion. Our analysis on the robustness of multimodal fusion will provide further insights for future research. Xinyu Zhang 0001, Zhiwei Li 0011, Zhenhong Zou, Xin Gao 0028, Yijin Xiong, Dafeng Jin, Jun Li 0082, Huaping Liu 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2023 | BEVHeight: A Robust Framework for Vision-based Roadside 3D Object DetectionabstractWhile most recent autonomous driving system focuses on developing perception methods on ego-vehicle sensors, people tend to overlook an alternative approach to leverage intelligent roadside cameras to extend the perception ability beyond the visual range. We discover that the state-of-the-art vision-centric bird's eye view detection methods have inferior performances on roadside cameras. This is because these methods mainly focus on recovering the depth regarding the camera center, where the depth difference between the car and the ground quickly shrinks while the distance increases. In this paper, we propose a simple yet effective approach, dubbed BEVHeight, to address this issue. In essence, instead of predicting the pixel-wise depth, we regress the height to the ground to achieve a distance-agnostic formulation to ease the optimization process of camera-only perception methods. On popular 3D detection benchmarks of roadside cameras, our method surpasses all previous vision-centric methods by a significant margin. The code is available at https://github.com/ADLab-AutoDrive/BEVHeight. Lei Yang 0060, Kaicheng Yu, Jun Li 0082, Kun Yuan 0001, Li Wang 0092, Xinyu Zhang 0001 |
CVPR | 7 |
| 2023 | CO-Net: Learning Multiple Point Cloud Tasks at Once with A Cohesive NetworkabstractWe present CO-Net, a cohesive framework that optimizes multiple point cloud tasks collectively across heterogeneous dataset domains. CO-Net maintains the characteristics of high storage efficiency since models with the preponderance of shared parameters can be assembled into a single model. Specifically, we leverage residual MLP (Res-MLP) block for effective feature extraction and scale it gracefully along the depth and width of the network to meet the demands of different tasks. Based on the block, we propose a novel nested layer-wise processing policy, which identifies the optimal architecture for each task while provides partial sharing parameters and partial non-sharing parameters inside each layer of the block. Such policy tackles the inherent challenges of multi-task learning on point cloud, e.g., diverse model topologies resulting from task skew and conflicting gradients induced by heterogeneous dataset domains. Finally, we propose a sign-based gradient surgery to promote the training of CO-Net, thereby emphasizing the usage of task-shared parameters and guaranteeing that each task can be thoroughly optimized. Experimental results reveal that models optimized by CO-Net jointly for all point cloud tasks maintain much fewer computation cost and overall storage cost yet outpace prior methods by a significant margin. We also demonstrate that CO-Net allows incremental learning and prevents catastrophic amnesia when adapting to a new point cloud task. Tao Xie 0010, Ke Wang 0028, Siyi Lu, Jie Xu 0066, Li Wang 0092, Lijun Zhao 0003, Xinyu Zhang 0001, Ruifeng Li 0001 |
ICCV | 10 |
| 2023 | SAT-GCN: Self-attention graph convolutional network-based 3D object detection for autonomous driving
Li Wang 0092, Ziying Song, Xinyu Zhang 0001, Jun Li 0082, Huaping Liu 0001 |
Knowl. Based Syst. | 3 |
| 2023 | Lite-FPN for keypoint-based monocular 3D object detection
Lei Yang 0060, Xinyu Zhang 0001, Jun Li 0082, Li Wang 0092, Minghan Zhu |
Knowl. Based Syst. | 2 |
| 2023 | Mix-Teaching: A Simple, Unified and Effective Semi-Supervised Learning Framework for Monocular 3D Object DetectionabstractSemi-supervised learning (SSL) has promising potential for improving model performance using both labelled and unlabelled data. Since recovering 3D information from 2D images is an ill-posed problem, the current state-of-the-art methods of monocular 3D object detection (Mono3D) have relatively low precision and recall, making semi-supervised learning for Mono3D tasks challenging and understudied. In this work, we propose a unified and effective semi-supervised learning framework called Mix-Teaching that can be applied to most monocular 3D object detectors. Based on the idea of decomposition and recombination, unlabelled samples are firstly decomposed into collections of image patches with high-quality predictions and collections of background images containing no objects. The student model is then trained on the mixed images containing dense instances with high-quality pseudo-labels generated by the recombination operation. In addition, we propose an uncertainty-based filter to distinguish high-quality pseudo-labels from noisy predictions during the decomposition process. As results in KITTI and nuScenes benchmarks, Mix-Teaching consistently improves MonoFlex and GUPNet by significant margins under various labeling ratios. Our method achieves around +6.34%$AP_{3D}$improvement against the GUPNet on the validation set when using only 10% labelled data. Using the full training set and the additional 38K raw images from KITTI, it can further improve the MonoFlex by +4.65% absolute improvement on$AP_{3D}$for car detection, reaching 18.54%$AP_{3D}$, which ranks the 1st place among all monocular based methods on the KITTI test leaderboard. Lei Yang 0060, Xinyu Zhang 0001, Jun Li 0082, Li Wang 0092, Minghan Zhu, Huaping Liu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2023 | CAMO-MOT: Combined Appearance-Motion Optimization for 3D Multi-Object Tracking With Camera-LiDAR Fusionabstract3D Multi-object tracking (MOT) ensures consistency during continuous dynamic detection, conducive to subsequent motion planning and navigation tasks in autonomous driving. However, camera-based methods suffer in the case of occlusions and it can be challenging to track the irregular motion of objects for LiDAR-based methods accurately. Some fusion methods work well but do not consider the untrustworthy issue of appearance features under occlusion. At the same time, the false detection problem also significantly affects tracking. As such, we propose a novel camera-LiDAR fusion 3D MOT framework based on Combined Appearance-Motion Optimization (CAMO-MOT), which uses both camera and LiDAR data and significantly reduces tracking failures caused by occlusion and false detection. For occlusion problems, we are the first to propose an occlusion head to select the best object appearance features multiple times effectively, reducing the influence of occlusions. To decrease the impact of false detection in tracking, we design a motion cost matrix based on confidence scores which improve the positioning and object prediction accuracy in 3D space. As existing multi-object tracking methods always evaluate each category separately and do not consider the mismatch between objects of different categories, we also propose to build a multi-category cost to implement multi-object tracking in multi-category scenes. A series of validation experiments are conducted on the KITTI and nuScenes tracking benchmarks. Our proposed method achieves state-of-the-art performance with 79.99% HOTA and the lowest identity switches (IDS) value (23 for Car and 137 for Pedestrian) among all multi-modal MOT methods on the KITTI test dataset. And our method achieves state-of-the-art performance among all algorithms on the nuScenes test dataset with 75.3% AMOTA. Li Wang 0092, Xinyu Zhang 0001, Wenyuan Qin, Jinghan Gao, Lei Yang 0060, Zhiwei Li 0011, Jun Li 0082, Hong Wang 0014, Huaping Liu 0001 |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2022 | IPS300+: a Challenging multi-modal data sets for Intersection Perception SystemabstractDue to high complexity and occlusion, insufficient perception in the crowded urban intersection can be a serious safety risk for both human drivers and autonomous algorithms, whereas CVIS (Cooperative Vehicle Infrastructure System) is a proposed solution for full-participants perception under this scenario. However, the research on roadside multi-modal perception is still in its infancy, and there is no open-source data sets for such scene. Accordingly, this paper fills the gap. Through an IPS (Intersection Perception System) installed at the diagonal of the intersection, this paper proposes a high-quality multi-modal data sets for the intersection perception task. The center of the experimental intersection covers an area of 3000m2, and the extended distance reaches 300m, which is typical for CVIS. The first batch of open-source data includes 14198 frames, and each frame has an average of 319.84 labels, which is 9.6 times larger than the most crowded data sets (H3D data sets in 2019) by now. Our data sets is available at: http://www.openmpd.com/column/IPS300. Huanan Wang, Xinyu Zhang 0001, Zhiwei Li 0011, Jun Li 0082, Zhu Lei, Haibing Ren |
ICRA | 2 |
| 2022 | InterFusion: Interaction-based 4D Radar and LiDAR Fusion for 3D Object DetectionabstractMany recent works detect 3D objects by several sensor modalities for autonomous driving, where high-resolution cameras and high-line LiDARs are mostly used but relatively expensive. To achieve a balance between overall cost and detection accuracy, many multi-modal fusion techniques have been suggested. In recent years, the fusion of LiDAR and Radar has gained ever-increasing attention, especially 4D Radar, which can adapt to bad weather conditions due to its penetrability. Although features have been fused from multiple sensing modalities, most methods cannot learn interactions from different modalities, which does not make for their best use. Inspired by the self-attention mechanism, we present InterFusion, an interaction-based fusion framework, to fuse 16-line LiDAR with 4D Radar. It aggregates features from two modalities and identifies cross-modal relations between Radar and LiDAR features. In experimental evaluations on the Astyx HiRes 2019 dataset, our method outperformed the baseline by 4.20% mAP in 3D and 10.76% BEV mAP for the car class at the moderate level. Li Wang 0092, Xinyu Zhang 0001, Baowei Xv, Jinzhao Zhang, Haibing Ren, Pingping Lu, Jun Li 0082, Huaping Liu 0001 |
IROS | 2 |
| 2021 | Line-based Automatic Extrinsic Calibration of LiDAR and CameraabstractReliable real-time extrinsic parameters of 3D Light Detection and Ranging (LiDAR) and camera are a key component of multi-modal perception systems. However, extrinsic transformation may drift gradually during operation, which can result in decreased accuracy of perception system. To solve this problem, we propose a line-based method that enables automatic online extrinsic calibration of LiDAR and camera in real-world scenes. Herein, the line feature is selected to constrain the extrinsic parameters for its ubiquity. Initially, the line features are extracted and filtered from point clouds and images. Afterwards, an adaptive optimization is utilized to provide accurate extrinsic parameters. We demonstrate that line features are robust geometric features that can be extracted from point clouds and images, thus contributing to the extrinsic calibration. To demonstrate the benefits of this method, we evaluate it on KITTI benchmark with ground truth value. The experiments verify the accuracy of the calibration approach. In online experiments on hundreds of frames, our approach automatically corrects miscalibration errors and achieves an accuracy of 0.2 degrees, which verifies its applicability in various scenarios. This work can provide basis for perception systems and further improve the performance of other algorithms that utilize these sensors. Xinyu Zhang 0001, Shifan Zhu, Shichun Guo, Jun Li 0082, Huaping Liu 0001 |
ICRA | 1 |
| 2021 | Lifelong Localization in Semi-Dynamic EnvironmentabstractMapping and localization in non-static environments are fundamental problems in robotics. Most of previous methods mainly focus on static and highly dynamic objects in the environment, which may suffer from localization failure in semi-dynamic scenarios without considering objects with lower dynamics, such as parked cars and stopped pedestrians. In this paper, we introduce semantic mapping and lifelong localization approaches to recognize semi-dynamic objects in non-static environments. We also propose a generic framework that can integrate mainstream object detection algorithms with mapping and localization algorithms. The mapping method combines an object detection algorithm and a SLAM algorithm to detect semi-dynamic objects and constructs a semantic map that only contains semi-dynamic objects in the environment. During navigation, the localization method can classify observation corresponding to static and non-static objects respectively and evaluate whether those semi-dynamic objects have moved, to reduce the weight of invalid observation and localization fluctuation. Real-world experiments show that the proposed method can improve the localization accuracy of mobile robots in non-static scenarios. Shifan Zhu, Xinyu Zhang 0001, Shichun Guo, Jun Li 0082, Huaping Liu 0001 |
ICRA | 2 |
| 2021 | Channel Attention in LiDAR-camera Fusion for Lane Line Segmentation
Xinyu Zhang 0001, Zhiwei Li 0011, Xin Gao 0028, Dafeng Jin, Jun Li 0082 |
Pattern Recognit. | 1 |
| 2021 | Toward Image-to-Tactile Cross-Modal Perception for Visually Impaired PeopleabstractIt is still a great challenge for the visually impaired people to perceive their surroundings from a global perspective, which makes it difficult for them to interact with unfamiliar environments. The reason is that these conventional assisting devices only address the obstacle avoidance problem. They do not provide visually impaired people with a global perception of the surrounding environment. In this article, a new generative adversarial network (GAN) model is developed to effectively transform the ground images into the tactile signal, which can be displayed by an off-the-shelf vibration device. The algorithm module and the hardware are integrated into a portable device, which provides visually impaired people with effective surrounding perception capability. In addition, a visual-tactile cross-modal data set is constructed to train the proposed deep-learning architecture. Experimental results show that the proposed system can help visually impaired people sense the ground and bring a better traveling experience for them. Note to Practitioners-This article presents a portable device that provides tactile recognition assistance for visually impaired people. Such a technology can be extensively used in tactile mouse and white cane. The developed technology can be extensively used for various industrial applications, such as surrounding monitoring and manipulation. The proposed work demonstrates the promising ability of artificial intelligence in healthcare applications. The generated tactile signals are expected to be used in many human-centered systems, and we believe that our contribution is an important step toward the development of a more comprehensive assisting technology for visually impaired people. Huaping Liu 0001, Di Guo 0002, Xinyu Zhang 0001, Wenlin Zhu, Bin Fang 0003, Fuchun Sun 0001 |
IEEE Trans Autom. Sci. Eng. | 3 |
| 2021 | Active Object Discovery and Localization Using Sound-Induced AttentionabstractIndustrial intelligent devices are usually equipped with both microphones and cameras to perceive and understand the physical world. Though visual object detection technology has achieved a great success, its combination with other sensing modalities remains unsolved. In this article, we establish a novel sound-induced attention framework for the visual object detection, and develop a two-stream weakly supervised deep learning architecture to combine the visual and audio modalities for localizing the sounding object. A dataset is constructed from the Audio Set to validate the proposed method and some realistic experiments are conducted to demonstrate the effectiveness of the proposed system. Huaping Liu 0001, Feng Wang 0034, Di Guo 0002, Xinzhu Liu, Xinyu Zhang 0001, Fuchun Sun 0001 |
IEEE Trans. Ind. Informatics | 5 |
| 2020 | Weakly-paired deep dictionary learning for cross-modal retrieval
Huaping Liu 0001, Feng Wang 0034, Xinyu Zhang 0001, Fuchun Sun 0001 |
Pattern Recognit. Lett. | 3 |
| 2019 | The Architecture of the Intended Safety System for Intelligent DrivingabstractAs the development direction of intelligent driving in the future, the research of key technologies has made significant progress. However, due to the recent unmanned accidents, there are concerns about safety performance. To solve the safety problem, an intended safety systems for intelligent driving was proposed. This system provides security analysis and monitoring services in real time for intended problems with smart car perception, decision and control modules. Based on the concept of safety of the intended functionality, the driving scene and system safety are analyzed and evaluated to improve the safety of intelligent driving, which may help the development of intelligent driving. Xinyu Zhang 0001, Wenbo Shao, Jun Li 0082 |
ISCAS | 1 |
| 2019 | Interactive video summarization with human intentions
Huaping Liu 0001, Fuchun Sun 0001, Xinyu Zhang 0001, Bin Fang 0003 |
Multim. Tools Appl. | 3 |
| 2019 | Active Object Detection With Multistep Action Prediction Using Deep Q-NetworkabstractIn recent years, great success has been achieved in visual object detection, which is one of the fundamental tasks in the field of industrial intelligence. Most of existing methods have been proposed to deal with single well-captured still images, while in practical robotic applications, due to nuisances, such as tiny scale, partial view, or occlusion, one still image may not contain enough information for object detection. However, an intelligent robot has the capability to adjust its viewpoint to get better images for detection. Therefore, active object detection becomes a very important perception strategy for intelligent robots. In this paper, by formulating active object detection as a sequential action decision process, a deep reinforcement learning framework is established to resolve it. Furthermore, a novel deep Q-learning network (DQN) with a dueling architecture is proposed, the network has two separate output channels, one predicts action type and the other predicts action range. By combining the two output channels, the action space is explored more efficiently. Several methods are extensively validated and the results show that the proposed one obtains the best results and predicts action in real time. Xiaoning Han, Huaping Liu 0001, Fuchun Sun 0001, Xinyu Zhang 0001 |
IEEE Trans. Ind. Informatics | 4 |
| 2018 | Adaptive Adversarial Transfer Learning for Electroencephalography ClassificationabstractInsufficient training data is a serious problem in all domains related to bioinformatics. Large-scale annotated electroencephalography (EEG) datasets are almost impossible to acquire because biological data acquisition is challenging and quality annotation is costly. Transfer learning relaxes the hypothesis that the training data must be independent and identically distributed (i.i. d.) with the test data, which motivates us to use transfer learning to solve the problem of insufficient training data in bioinformatics. We propose a new approach to transfer knowledge via a deep transfer learning framework, which includes an adaptive sample selection algorithm and a joint adversarial training algorithm. The adaptive sample selection algorithm dynamically adjusts the sample weights during the training process according to the distance between the source domain and the target domain. The joint adversarial training algorithm forces the network to learn a feature extractor suitable for the target domain based on a dataset from the source domain by using an adversarial network and a specific loss function. The experiments demonstrate that our approach has many advantages, such as robustness and accuracy, when applied to EEG classification tasks. Chuanqi Tan, Fuchun Sun 0001, Tao Kong, Chao Yang 0026, Xinyu Zhang 0001 |
IJCNN | 6 |
| 2018 | Accurately Detecting Community with Large Attribute in Partial Networks
Wei Han 0005, Xinyu Zhang 0001 |
PRICAI (1) | 3 |
| 2018 | Multi-view clustering based on graph-regularized nonnegative matrix factorization for object recognition
Xinyu Zhang 0001, Hongbo Gao 0001, Jianghao Huo, Jialun Yin |
Inf. Sci. | 1 |
| 2017 | Semi-supervised convex nonnegative matrix factorizations with graph regularized for image representation
Xinyu Zhang 0001, Siyi Zheng, Deyi Li |
Neurocomputing | 2 |
| 2013 | Design and Implement a Knowledge Management System to Support Web-based Learning in Higher EducationabstractToday most colleges have web-based learning system to keep a large number of course resources during higher education process. However, these systems aiming to display course resources often neglect users’ knowledge management requirement. Traditional course-based learning confines knowledge to one course. But both teachers and learners often require extracting useful knowledge from course for themselves. Besides, with development of mobile technology, learning with convenient devices was required by college students. Lack of knowledge services and limited clients constrict the development of web-based learning. Taking Tsinghua University as a typical case and according to the knowledge management theory, this paper designed and implemented a knowledge management system (KMS-THU) to support knowledge service for Tsinghua Web School (THU-WS), which is a web-based learning platform of Tsinghua. KMS-THU focuses on knowledge management by people and also abstracts distinctive knowledge services for courses. With campus cloud service and varies mobile clients, it brings a ubiquitous learning style to optimize the learning experience. Besides illustrating design of knowledge service and framework of KMS for web-based learning, this paper set forth the technical details for KMS-THU implementation. Jinyue Peng, Dongxing Jiang, Xinyu Zhang 0001 |
KES | 3 |
| 2002 | Human Communication and Interaction in Web Based Learning: A Case Study of the digital Media Web CourseabstractThe case of web based course include a WBLS (Web based learning system) and face-to-face three times during the semester. Based on the WBLS development and the investigation of the course, the paper shows the student attitudes to the course and WBLS, analyzes the interactive functions and the impact of human communication in web learning. and the impact of the Web based learning. Although most students prefer Web based learning, the current Web learning patterns can not completely replace instructor's supervision and in person communication with students. There are ways to increase human communication and to improve the interactive functions further more, so as to increase the educational effectiveness. Huifen Liu, Xinyu Zhang 0001 |
ICCE | 3 |