Victor Y. Chen

dblp:12/5359 · also Yingjie Chen 0001, Yingjie Victor Chen · DBLP profile ↗
← Back
50ranked-venue papers
4as first author
23since 2021 · last 2026
0000-0001-6705-3535ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 21 · 1 first-author · 17 since 2021Artificial intelligence and machine learning · 16 · 1 first-author · 11 since 2021Computer networks · 8 · 1 first-authorHuman-computer interaction and ubiquitous computing · 7 · 1 first-author · 3 since 2021Systems, architecture and hardware · 5Applied, interdisciplinary, general and emerging computing · 4 · 2 since 2021Security and privacy · 2 · 1 first-author
YearPublicationVenuePosition
2026 Make LVLMs Focus: Context-Aware Attention Modulation for Better Multimodal In-Context Learning
abstract
Multimodal in-context learning (ICL) is becoming a key capability that allows large vision-language models (LVLMs) to adapt to novel tasks without parameter updates, which expands their usefulness in many real-world applications. However, ICL performance remains unstable even when the in-context demonstrations (ICDs) are well matched, showing that LVLMs still struggle to make full use of the provided context. While existing work mainly focuses on prompt engineering or post-hoc logit calibration, we study the attention mechanisms inside LVLMs to address their inherent limitations. We identify two important weaknesses in their self-attention that hinder effective ICL. To address these weaknesses, we propose Context-Aware Modulated Attention (CAMA), a training-free and plug-and-play method that dynamically adjusts attention logits based on the input in-context sequence. CAMA uses a two-stage modulation process that strengthens attention to semantically important tokens, especially visual ones. Across four LVLMs and seven benchmarks, CAMA consistently outperforms vanilla models and baselines, showing clear effectiveness and generalization. It can also activate the intended benefits of prompt engineering methods and remains robust across different sequence configurations. Therefore, CAMA opens up new directions for improving multimodal reasoning through a deeper understanding of attention dynamics.
Yanshu Li, Jianjiang Yang, Ziteng Yang, Bozheng Li, Ligong Han, Hongyang He, Zhengtao Yao, Victor Y. Chen, Songlin Fei, Dongfang Liu, Ruixiang Tang
AAAI8
2025 SplatFlow: Self-Supervised Dynamic Gaussian Splatting in Neural Motion Flow Field for Autonomous Driving
abstract
Most existing Dynamic Gaussian Splatting methods for complex dynamic urban scenarios rely on accurate object-level supervision from expensive manual labeling, limiting their scalability in real-world applications. In this paper, we introduce SplatFlow, a Self-Supervised Dynamic Gaussian Splatting within Neural Motion Flow Fields (NMFF) to learn 4D space-time representations without requiring tracked 3D bounding boxes, enabling accurate dynamic scene reconstruction and novel view RGB/depth/flow synthesis. SplatFlow designs a unified framework to seamlessly integrate time-dependent 4D Gaussian representation within NMFF, where NMFF is a set of implicit functions to model temporal motions of both LiDAR points and Gaussians as continuous motion flow fields. Leveraging NMFF, SplatFlow effectively decomposes static background and dynamic objects, representing them with 3D and 4D Gaussian primitives, respectively. NMFF also models the correspondences of each 4D Gaussian across time, which aggregates temporal features to enhance cross-view consistency of dynamic components. SplatFlow further improves dynamic object identification by distilling features from 2D foundation models into 4D space-time representation. Comprehensive evaluations conducted on the Waymo and KITTI Datasets validate SplatFlow’s state-of-the-art (SOTA) performance for both image reconstruction and novel view synthesis in dynamic urban scenarios.
Su Sun, Cheng Zhao 0002, Zhuoyang Sun, Victor Y. Chen
CVPR4
2025 Probabilistic Token Alignment for Large Language Model Fusion
abstract
Training large language models (LLMs) from scratch can yield models with unique functionalities and strengths, but it is costly and often leads to redundant capabilities. A more cost-effective alternative is to fuse existing pre-trained LLMs with different architectures into a more powerful model. However, a key challenge in existing model fusion is their dependence on manually predefined vocabulary alignment, which may not generalize well across diverse contexts, leading to performance degradation in several evaluation. To solve this, we draw inspiration from distribution learning and propose the probabilistic token alignment method as a general and soft mapping for alignment, named as PTA-LLM. Our approach innovatively reformulates token alignment into a classic mathematical problem: optimal transport, seamlessly leveraging distribution-aware learning to facilitate more coherent model fusion. Apart from its inherent generality, PTA-LLM exhibits interpretability from a distributional perspective, offering insights into the essence of the token alignment. Empirical results demonstrate that probabilistic token alignment enhances the target model's performance across multiple capabilities.
Runjia Zeng, James Liang, Cheng Han 0001, Zhiwen Cao, Xiaojun Quan, Victor Y. Chen, Lifu Huang, Tong Geng, Qifan Wang 0001, Dongfang Liu
NeurIPS7
2025 Dynamic Translational Gains Manipulation for Tiny Object Interaction
abstract
ABSTRACT Interacting with small objects in virtual reality (VR) can be challenging due to the physical limitations of controllers and headsets, which often lead to unintended collisions and tracking loss when devices come too close, thereby disrupting the user's immersive experience. While researchers have developed techniques like translational gain, hand remapping, and specialized interaction to address these challenges, these approaches are often task‐specific or insufficient for precise, detailed interactions or observations. To address these challenges, in this paper, we introduce a novel interaction technique called dynamic translational gains manipulation (DTGM), which adjusts scaling in real‐time based on the user's proximity to objects. We conducted a user study to evaluate the effectiveness of improving precision during object manipulation and understanding the subjective mental workload of the proposed DTGM technique. Our results revealed that the DTGM technique improved interaction efficiency, making it suitable for various VR applications where precision and space optimization are crucial.
Jiahui Dong, Tansi Zhang, Shengyang Luo, Christos Mousas, Victor Y. Chen
Comput. Animat. Virtual Worlds5
2024 ProMotion: Prototypes as Motion Learners
abstract
In this work, we introduce PRoMoTION, a unified proto-typical transformer-based framework engineered to model fundamental motion tasks. PRoMoTION offers a range of compelling attributes that set it apart from current task-specific paradigms. (1) We adopt a prototypical perspective, establishing a unified paradigm that harmonizes disparate motion learning approaches. This novel paradigm stream-lines the architectural design, enabling the simultaneous assimilation of diverse motion information. (2) We capitalize on a dual mechanism involving the feature denoiser and the prototypical learner to decipher the intricacies of motion. This approach effectively circumvents the pitfalls of ambiguity in pixel-wise feature matching, significantly bolstering the robustness of motion representation. (3)) We demon-strate a profound degree of transferability across distinct motion patterns. This inherent versatility reverberates robustly across a comprehensive spectrum of both 2D and 3D downstream tasks. Empirical results demonstrate that PRoMOTION outperforms various well-known specialized architectures, achieving 0.54 and 0.054$AbsRel$error on the Sintel and KITTI depth datasets, 1.04 and 2.01 average endpoint error on the clean and final pass of Sintel flow benchmark, and 4.30 F1-all error on the KITTI flow bench-mark. For its efficacy, we hope our work can catalyze a paradigm shift in universal models in computer vision.
Yawen Lu, Dongfang Liu, Qifan Wang 0001, Cheng Han 0001, Yiming Cui 0002, Zhiwen Cao, Xueling Zhang, Victor Y. Chen, Heng Fan 0001
CVPR8
2024 Behind the Veil: Enhanced Indoor 3D Scene Reconstruction with Occluded Surfaces Completion
abstract
In this paper, we present a novel indoor 3D reconstruction method with occluded surface completion, given a sequence of depth readings. Prior state-of-the-art (SOTA) methods only focus on the reconstruction of the visible areas in a scene, neglecting the invisible areas due to the occlusions, e.g., the contact surface between furniture, occluded wall and floor. Our method tackles the task of completing the occluded scene surfaces, resulting in a complete 3D scene mesh. The core idea of our method is learning 3D geometry prior from various complete scenes to infer the occluded geometry of an unseen scene from solely depth measurements. We design a coarse-fine hierarchical octree representation coupled with a dual-decoder architecture, i.e., Geo-decoder and 3D Inpainter, which jointly reconstructs the complete 3D scene geometry. The Geo-decoder with detailed representation at fine levels is optimized online for each scene to reconstruct visible surfaces. The 3D Inpainter with abstract representation at coarse levels is trained offline using various scenes to complete occluded surfaces. As a result, while the Geo-decoder is specialized for an individual scene, the 3D Inpainter can be generally applied across different scenes. We evaluate the proposed method on the 3D Completed Room Scene (3D-CRS) and iTHOR datasets, significantly outperforming the SOTA methods by a gain of 16.8% and 24.2% in terms of the completeness of 3D reconstruction. 3D-CRS dataset including a complete 3D mesh of each scene is provided on project webpage11https://github.com/BoschRHI3NA/3D-CRS-dataset.
Su Sun, Cheng Zhao 0002, Yuliang Guo, Ruoyu Wang 0012, Xinyu Huang 0001, Victor Y. Chen, Liu Ren 0001
CVPR6
2024 TCLC-GS: Tightly Coupled LiDAR-Camera Gaussian Splatting for Autonomous Driving: Supplementary Materials
Cheng Zhao 0002, Su Sun, Ruoyu Wang 0012, Yuliang Guo, Jun-Jun Wan, Xinyu Huang 0001, Victor Y. Chen, Liu Ren 0001
ECCV (63)8
2024 Understanding Pitfalls and Opportunities of Applying Heuristic Evaluation Methods to VR Training Systems: An Empirical Study
abstract
The usability of virtual reality (VR) training applications is crucial for their success, but examining the usability in the early development stages remains challenging. A realistic and plausible solution would be revisiting and reconciling Heuristics Evaluation (HE) methods among the most widely used usability inspection methods in the human-computer interaction (HCI) domain. While research on studying and using HE methods is growing within the VR domain, few studies have considered the novel VR environment challenges new requirements for fitting HE methods to the context and applying them effectively. To this end, we conducted a user study with 14 evaluators using the standard HE methods to complete two HE sessions for a VR training application. We identified five critical challenges that evaluators encountered in the HE process by observing and interviewing them. Based on our findings, we discuss the importance of considering an easy-to-use heuristic set, how we can facilitate the HE procedures in the VR context, and the opportunities for developing HE-supporting tools.
Kushal Kumar Nerella, Jiahui Dong, Cheryl Z. Qian, Victor Y. Chen
Int. J. Hum. Comput. Interact.5
2024 Effectiveness of Haptic Modality in an Intelligent Bicycle Safety Driving Device Supporting Bicycle Delivery Service
abstract
Deliverymen have widely adopted smart phone apps for better delivery performance. With the development of the delivery industry and the increase of its employees, the personal safety of deliverymen in completing orders has attracted the same attention as task performance. In addition to the smartphone apps, we designed an intelligent delivery assistant (IDA) system with “multi-modal navigation” to improve the safety and efficiency of bicycle deliveries. The IDA system was equipped with a pair of driving controllers that allowed deliverymen to interact with the IDA system and get navigation information during riding. This 60-human-participant study was conducted to investigate the performance of the IDA system, including usability and safety. The participants were divided into two groups for a comparative experiment, one with the IDA system and the other with only a smartphone. Navigation was conducted in a real-world environment, in which eye movements and device interaction were recorded using an eye tracker. Results indicate that the IDA system better supports delivery usability and safety. Higher perceived usability and better situational awareness with a more significant number of hazards acknowledged were detected with the IDA system. The NASA-TLX and UEQ results demonstrated the participants reported having a better user experience with the IDA system. We further discussed the benefits and drawbacks of each condition of the IDA system.
Jinyao Zhang, Weizhuan Hu, Victor Y. Chen
Int. J. Hum. Comput. Interact.5
2024 TrajVis: a visual clinical decision support system to translate artificial intelligence trajectory models in the precision management of chronic kidney disease
abstract
OBJECTIVE: Our objective is to develop and validate TrajVis, an interactive tool that assists clinicians in using artificial intelligence (AI) models to leverage patients' longitudinal electronic medical records (EMRs) for personalized precision management of chronic disease progression. MATERIALS AND METHODS: We first perform requirement analysis with clinicians and data scientists to determine the visual analytics tasks of the TrajVis system as well as its design and functionalities. A graph AI model for chronic kidney disease (CKD) trajectory inference named DisEase PrOgression Trajectory (DEPOT) is used for system development and demonstration. TrajVis is implemented as a full-stack web application with synthetic EMR data derived from the Atrium Health Wake Forest Baptist Translational Data Warehouse and the Indiana Network for Patient Care research database. A case study with a nephrologist and a user experience survey of clinicians and data scientists are conducted to evaluate the TrajVis system. RESULTS: The TrajVis clinical information system is composed of 4 panels: the Patient View for demographic and clinical information, the Trajectory View to visualize the DEPOT-derived CKD trajectories in latent space, the Clinical Indicator View to elucidate longitudinal patterns of clinical features and interpret DEPOT predictions, and the Analysis View to demonstrate personal CKD progression trajectories. System evaluations suggest that TrajVis supports clinicians in summarizing clinical data, identifying individualized risk predictors, and visualizing patient disease progression trajectories, overcoming the barriers of AI implementation in healthcare. DISCUSSION: The TrajVis system provides a novel visualization solution which is complimentary to other risk estimators such as the Kidney Failure Risk Equations. CONCLUSION: TrajVis bridges the gap between the fast-growing AI/ML modeling and the clinical use of such models for personalized and precision management of chronic diseases.
Zuotian Li, Xiang Liu 0016, Ziyang Tang, Nanxin Jin, Pengyue Zhang, Michael Eadon, Qianqian Song 0002, Victor Y. Chen, Jing Su 0003
J. Am. Medical Informatics Assoc.8
2024 Optical Flow as Spatial-Temporal Attention Learners
abstract
Optical flow is an indispensable building block for various important computer vision tasks, including motion estimation, object tracking, and disparity measurement. To date, the dominant methods are CNN-based, leaving plenty of room for improvement. In this work, we propose TransFlow, a transformer architecture for optical flow estimation. Compared to dominant CNN-based methods, TransFlow demonstrates three advantages. First, it provides more accurate correlation and trustworthy matching in flow estimation by utilizing spatial self-attention and cross-attention mechanisms between adjacent frames to effectively capture global dependencies; Second, it recovers more compromised information (e.g., occlusion and motion blur) in flow estimation through long-range temporal association in dynamic scenes; Third, it introduces a concise self-learning paradigm, eliminating the need for complex and laborious multi-stage pre-training procedures. The versatility and superiority of TransFlow extend seamlessly to 3D scene motion, yielding competitive outcomes in 3D scene flow estimation. Our approach attains state-of-the-art results on benchmark datasets such as Sintel and KITTI-15, while also exhibiting exceptional performance on downstream tasks, including video object detection using the ImageNet VID dataset, video frame interpolation using the GoPro dataset, and video stabilization using the DeepStab dataset. We believe that the effectiveness of TransFlow positions it as a flexible baseline for both optical flow and scene flow estimation, offering promising avenues for future research and development.
Yawen Lu, Cheng Han 0001, Qifan Wang 0001, Heng Fan 0001, Zhaodan Kong, Dongfang Liu, Victor Y. Chen
IEEE Trans. Pattern Anal. Mach. Intell.7
2024 Label-Efficient Video Object Segmentation With Motion Clues
abstract
Video object segmentation (VOS) plays an important role in video analysis and understanding, which in turn facilitates a number of diverse applications, including video editing, video rendering, and augmented reality / virtual reality. However, existing deep learning-based approaches rely heavily on a large number of pixel-wise annotated video frames to achieve promising results, which is notoriously laborious and costly. To address this, in this paper, we formulate unsupervised video object detection by exploring simulated dense labels and explicit motion clues. Specifically, we first propose an effective video label generator network based on the sparsely annotated frames and the flow motion between them. It can largely alleviate our dependence and limitation on the sparse labels. Furthermore, we propose a transformer-based architecture to model the appearance and motion clues simultaneously with the cross-attention module, in order to maximally overcome non-linear motion with potential occlusions. Extensive experiments show that the proposed method outperforms recent VOS methods on four popular benchmarks (i.e., DAVIS-16, FBMS, Youtube-VOS and SegTrack-v2). Moreover, the proposed method can be further applied to a wide range of wild scenes such as wild forests and animals. Because of its effectiveness and generalization, we believe that our method could serve as a useful basis for alleviating the dependence on dense annotation in video data.
Yawen Lu, Jie Zhang 0066, Su Sun, Zhiwen Cao, Songlin Fei, Baijian Yang 0001, Victor Y. Chen
IEEE Trans. Circuits Syst. Video Technol.8
2023 TransFlow: Transformer as Flow Learner
abstract
Optical flow is an indispensable building block for various important computer vision tasks, including motion estimation, object tracking, and disparity measurement. In this work, we propose TransFlow, a pure transformer architecture for optical flow estimation. Compared to dominant CNN-based methods, TransFlow demonstrates three advantages. First, it provides more accurate correlation and trustworthy matching in flow estimation by utilizing spatial self-attention and crossattention mechanisms between adjacent frames to effectively capture global dependencies; Second, it recovers more compromised information (e.g., occlusion and motion blur) in flow estimation through long-range temporal association in dynamic scenes; Third, it enables a concise self-learning paradigm and effectively eliminate the complex and laborious multi-stage pre-training procedures. We achieve the state-of-the-art results on the Sintel, KITTI-15, as well as several downstream tasks, including video object detection, interpolation and stabilization. For its efficacy, we hope TransFlow could serve as a flexible baseline for optical flow estimation.
Yawen Lu, Qifan Wang 0001, Siqi Ma 0005, Tong Geng, Victor Y. Chen, Huaijin G. Chen, Dongfang Liu
CVPR5
2022 The Alcoholic Hepatitis Network Research Data Commons (ARDaC): Design and Development
Jing Su 0003, Nanxin Jin, Zuotan Li, Carla Kettler, Bruce Barton, Greg Puetz, Chi Mai Nguyen, Donna McGrath, Victor Y. Chen, Baijian Yang 0001, Vijay Shah, Svetlana Radaeva, Samer Gawrieh, Wanzhu Tu
AMIA9
2022 A Comparative Study of Four 3D Facial Animation Methods: Skeleton, Blendshape, Audio-Driven, and Vision-Based Capture
Mingzhu Wei, Nicoletta Adamo-Villani, Nandhini Giri, Victor Y. Chen
ArtsIT4
2022 Towards Unbiased Label Distribution Learning for Facial Pose Estimation Using Anisotropic Spherical Gaussian
Zhiwen Cao, Dongfang Liu, Qifan Wang 0001, Victor Y. Chen
ECCV (12)4
2022 DG-Labeler and DGL-MOTS Dataset: Boost the Autonomous Driving Perception
abstract
Multi-object tracking and segmentation (MOTS) is a critical task for autonomous driving applications. The existing MOTS studies face two critical challenges: 1) the published datasets inadequately capture the real-world complexity for network training to address various driving settings; 2) the working pipeline annotation tool is under-studied in the literature to improve the quality of MOTS learning examples. In this work, we introduce the DG-Labeler and DGL-MOTS dataset to facilitate the training data annotation for the MOTS task and accordingly improve network training accuracy and efficiency. DG-Labeler uses the novel Depth-Granularity Module to depict the instance spatial relations and produce fine-grained instance masks. Annotated by DG-Labeler, our DGL-MOTS dataset exceeds the prior effort (i.e., KITTI MOTS and BDD100K) in data diversity, annotation quality, and temporal representations. Results on extensive cross-dataset evaluations indicate significant performance improvements for several state-of-the-art methods trained on our DGL-MOTS dataset. We believe our DGL-MOTS Dataset and DG-Labeler hold the valuable potential to boost the visual perception of future transportation. Our dataset and code are available here1.
Yiming Cui 0002, Zhiwen Cao, Chloe Yixin Xie, Xingyu Jiang 0001, Feng Tao 0002, Victor Y. Chen, Dongfang Liu
WACV6
2022 Video Captioning Using Global-Local Representation
abstract
Video captioning is a challenging task as it needs to accurately transform visual understanding into natural language description. To date, state-of-the-art methods inadequately model global-local vision representation for sentence generation, leaving plenty of room for improvement. In this work, we approach the video captioning task from a new perspective and propose a GLR framework, namely a global-local representation granularity. Our GLR demonstrates three advantages over the prior efforts. First, we propose a simple solution, which exploits extensive vision representations from different video ranges to improve linguistic expression. Second, we devise a novel global-local encoder, which encodes different video representations including long-range, short-range and local-keyframe, to produce rich semantic vocabulary for obtaining a descriptive granularity of video contents across frames. Finally, we introduce the progressive training strategy which can effectively organize feature learning to incur optimal captioning behavior. Evaluated on the MSR-VTT and MSVD dataset, we outperform recent state-of-the-art methods including a well-tuned SA-LSTM baseline by a significant margin, with shorter training schedules. Because of its simplicity and efficacy, we hope that our GLR could serve as a strong baseline for many video understanding tasks besides video captioning. Code will be available.
Liqi Yan, Siqi Ma 0005, Qifan Wang 0001, Victor Y. Chen, Xiangyu Zhang 0001, Andreas E. Savakis, Dongfang Liu
IEEE Trans. Circuits Syst. Video Technol.4
2021 DenserNet: Weakly Supervised Visual Localization Using Multi-Scale Feature Aggregation
abstract
In this work, we introduce a Denser Feature Network(DenserNet) for visual localization. Our work provides three principal contributions. First, we develop a convolutional neural network (CNN) architecture which aggregates feature maps at different semantic levels for image representations. Using denser feature maps, our method can produce more key point features and increase image retrieval accuracy. Second, our model is trained end-to-end without pixel-level an-notation other than positive and negative GPS-tagged image pairs. We use a weakly supervised triplet ranking loss to learn discriminative features and encourage keypoint feature repeatability for image representation. Finally, our method is computationally efficient as our architecture has shared features and parameters during forwarding propagation. Our method is flexible and can be crafted on a light-weighted backbone architecture to achieve appealing efficiency with a small penalty on accuracy. Extensive experiment results indicate that our method sets a new state-of-the-art on four challenging large-scale localization benchmarks and three image retrieval benchmarks with the same level of supervision. The code is available at https://github.com/goodproj13/DenserNet
Dongfang Liu, Yiming Cui 0002, Liqi Yan, Christos Mousas, Baijian Yang 0001, Victor Y. Chen
AAAI6
2021 SG-Net: Spatial Granularity Network for One-Stage Video Instance Segmentation
abstract
Video instance segmentation (VIS) is a new and critical task in computer vision. To date, top-performing VIS methods extend the two-stage Mask R-CNN by adding a tracking branch, leaving plenty of room for improvement. In contrast, we approach the VIS task from a new perspective and propose a one-stage spatial granularity network (SG-Net). Compared to the conventional two-stage methods, SG-Net demonstrates four advantages: 1) Our method has a one-stage compact architecture and each task head (detection, segmentation, and tracking) is crafted interdependently so they can effectively share features and enjoy the joint optimization; 2) Our mask prediction is dynamically performed on the sub-regions of each detected instance, leading to high-quality masks of fine granularity; 3) Each of our task predictions avoids using expensive proposal-based RoI features, resulting in much reduced runtime complexity per instance; 4) Our tracking head models objects’ centerness movements for tracking, which effectively enhances the tracking robustness to different object appearances. In evaluation, we present state-of-the-art comparisons on the YouTube-VIS dataset. Extensive experiments demonstrate that our compact one-stage method can achieve improved performance in both accuracy and inference speed. We hope our SG-Net could serve as a strong and flexible base-line for the VIS task. Our code will be available here1.
Dongfang Liu, Yiming Cui 0002, Wenbo Tan, Victor Y. Chen
CVPR4
2021 Hierarchical Attention Fusion for Geo-Localization
abstract
Geo-localization is a critical task in computer vision. In this work, we cast the geo-localization as a 2D image retrieval task. Current state-of-the-art methods for 2D geo-localization are not robust to locate a scene with drastic scale variations because they only exploit features from one semantic level for image representations. To address this limitation, we introduce a hierarchical attention fusion network using multi-scale features for geo-localization. We extract the hierarchical feature maps from a convolutional neural network (CNN) and organically fuse the extracted features for image representations. Our training is self-supervised using adaptive weights to control the attention of feature emphasis from each hierarchical level. Evaluation results on the image retrieval and the large-scale geo-localization benchmarks indicate that our method outperforms the existing state-of-the-art methods. Code is available here: https://github.com/YanLiqi/HAF.
Liqi Yan, Yiming Cui 0002, Victor Y. Chen, Dongfang Liu
ICASSP3
2021 A Vector-based Representation to Enhance Head Pose Estimation
abstract
This paper proposes to use the three vectors in a rotation matrix as the representation in head pose estimation and develops a new neural network based on the characteristic of such representation. We address two potential issues existed in current head pose estimation works: 1. Public datasets for head pose estimation use either Euler angles or quaternions to annotate data samples. However, both of these annotations have the issue of discontinuity and thus could result in some performance issues in neural network training. 2. Most research works report Mean Absolute Error (MAE) of Euler angles as the measurement of performance. We show that MAE may not reflect the actual behavior especially for the cases of profile views. To solve these two problems, we propose a new annotation method which uses three vectors to describe head poses and a new measurement Mean Absolute Error of Vectors (MAEV) to assess the performance. We also train a new neural network to predict the three vectors with the constraints of orthogonality. Our proposed method achieves state-of-the-art results on both AFLW2000 and BIWI datasets. Experiments show our vector-based annotation method can effectively reduce prediction errors for large pose angles.
Zhiwen Cao, Zongcheng Chu, Dongfang Liu, Victor Y. Chen
WACV4
2021 Phoenixmap: An Abstract Approach to Visualize 2D Spatial Distributions
abstract
The multidimensional nature of spatial data poses a challenge for visualization. In this paper, we introduce Phoenixmap, a simple abstract visualization method to address the issue of visualizing multiple spatial distributions at once. The Phoenixmap approach starts by identifying the enclosed outline of the point collection, then assigns different widths to outline segments according to the segments' corresponding inside regions. Thus, one 2D distribution is represented as an outline with varied thicknesses. Phoenixmap is capable of overlaying multiple outlines and comparing them across categories of objects in a 2D space. We chose heatmap as a benchmark spatial visualization method and conducted user studies to compare performances among Phoenixmap, heatmap, and dot distribution map. Based on the analysis and participant feedback, we demonstrate that Phoenixmap 1) allows users to perceive and compare spatial distribution data efficiently; 2) frees up graphics space with a concise form that can provide visualization design possibilities like overlapping; and 3) provides a good quantitative perceptual estimating capability given the proper legends. Finally, we discuss several possible applications of Phoenixmap and present one visualization of multiple species of birds' active regions in a nature preserve.
Junhan Zhao, Xiang Liu 0016, Cheryl Z. Qian, Victor Y. Chen
IEEE Trans. Vis. Comput. Graph.5
2020 Visual Localization for Autonomous Driving: Mapping the Accurate Location in the City Maze
abstract
Accurate localization is a foundational capacity, required for autonomous vehicles to accomplish other tasks such as navigation or path planning. It is a common practice for vehicles to use GPS to acquire location information. However, the application of GPS can result in severe challenges when vehicles run within the inner city where different kinds of structures may shadow the GPS signal and lead to inaccurate location results. To address the localization challenges of urban settings, we propose a novel feature voting technique for visual localization. Different from the conventional front-view-based method, our approach employs views from three directions (front, left, and right) and thus significantly improves the robustness of location prediction. In our work, we craft the proposed feature voting method into three state-of-the-art visual localization networks and modify their architectures properly so that they can be applied for vehicular operation. Extensive field test results indicate that our approach can predict location robustly even in challenging inner-city settings. Our research sheds light on using the visual localization approach to help autonomous vehicles to find accurate location information in a city maze, within a desirable time constraint. The source code is available at github.com/HappyDonkey13/Visual- Localization-for- Autonomous-Driving.
Dongfang Liu, Yiming Cui 0002, Baijian Yang 0001, Victor Y. Chen
ICPR6
2020 A Large-scale Simulation Dataset: Boost the Detection Accuracy for Special Weather Conditions
abstract
Object detection is a fundamental task for autonomous driving systems. One bottleneck hindering detection accuracy is a shortage of well-annotated image data. Virtual reality has provided a feasible low-cost way to facilitate computer vision related developments. In autonomous driving area, existing public datasets from real world generally have data biases and cannot represent a wide range of weather conditions, such as rainy or snowy roads. To address this challenge, we introduce a new large-scale simulation dataset which is generated by an automated pipeline from a high realism video game. Our dataset focuses on weather conditions, which can be adopted to train networks to effectively detect objects under such conditions. We use extensive experiments to evaluate our dataset by comparing it with public datasets. The experiment results show that networks trained with our dataset outperform the networks trained by other public datasets. Our work demonstrates the effectiveness of using simulation data to address real-world challenges in the practice of object detection.
Dongfang Liu, Yiming Cui 0002, Zhiwen Cao, Victor Y. Chen
IJCNN4
2020 Indoor Navigation for Mobile Agents: A Multimodal Vision Fusion Model
abstract
Indoor navigation is a challenging task for mobile agents. The latest vision-based indoor navigation methods make remarkable progress in this field but do not fully leverage visual information for policy learning and struggle to perform well in unseen scenes. To address the existing limitations, we present a multimodal vision fusion model (MVFM). We implement a joint modality of different image recognition networks for navigation policy learning. The proposed model incorporates object detection for target searching, depth estimation for distance prediction, and semantic segmentation to depict the walkable region. In design, our model provides holistic vision knowledge for navigation. Evaluation on AI2-THOR indicates that MVFM improves on the results of a strong baseline model by 3.49% for Success weighted by Path Length (SPL) and 4% for success rate respectively. In comparison with other state-of-the-art systems, MVFM performs in the lead in terms of SPL and success rate. Extensive experiments show the effectiveness of the proposed model.
Dongfang Liu, Yiming Cui 0002, Zhiwen Cao, Victor Y. Chen
IJCNN4
2020 Video object detection for autonomous driving: Motion-aid feature calibration
Dongfang Liu, Yiming Cui 0002, Victor Y. Chen, Jiyong Zhang 0001, Bin Fan 0001
Neurocomputing3
2019 Virtual Reality Training with Passive Haptic Feedback for CryoEM Sample Preparation
abstract
We present an immersive virtual reality training system cryoVR with passive haptic feedback for training biological scientists preparing bio-sample for Cryo-Electron Microscopy (CryoEM). CryoEM requires the user to conduct careful operations on expensive delicate equipment. To minimize the risk and interruption of crucial research work, we tried to mimic the real lab using VR to let trainer practice in a virtual environment. We used 3D printed objects to provide passive haptic feedback in order to achieve a more realistic training experience. Participants are able to interact with the equipment in the virtual environment by moving or touching physical models. We developed all the necessary operations with haptic feedback, including moving, clicking, pouring, rotating, polling and pushing. By following instructions provided by our virtual reality simulator and interacting with physical objects, trainees will learn how to operate CryoEM with low cost and risk. By developing our training system, we explore the benefits, limitations and precautions of embedding haptic feedback to scientific VR training.
Jiahui Dong, Pengyu Patrick Ren, Cheryl Z. Qian, Victor Y. Chen
VR6
2019 Low Area-Overhead Low-Entropy Masking Scheme (LEMS) Against Correlation Power Analysis Attack
abstract
The low-entropy masking scheme (LEMS) is a costsecurity tradeoff solution that ensures a certain level of security with much lower overheads than a full-entropy masking scheme (FEMS). However, most existing LEMSs are based on a look-up-table (LUT) and limited to the first-order, which is vulnerable to classical higher-order correlation power analysis (CPA) attack and other special types of attack (e.g., collision attack). This paper proposes a new type of LEMS for a block cipher in which the S-box consists of power functions and an affine function. First, a low masking-complexity algorithm for evaluating S-boxes is developed by fully utilizing the property of a hybrid addition-chain (AC) named LUT-AC. Next, an LEMS for block ciphers is proposed. This LEMS provides two different masking modes to realize various cost-security tradeoff schemes. Due to the “masked invariant property” of the LUTAC, the masking complexity of the proposed LEMS is equal to O(d), whereas under FEMS it is equal to O(d2). Compared with existing LEMSs, the proposed LEMS has following advantages: higher security in terms of the masking entropy; resistance against collision attacks; and scalability to higher-order schemes. Per the proposed algorithm, an architecture without any nonlinear multiplication for evaluating AES is developed by replacing the LUT with seven scalar multiplications. The different LEMSs based on this architecture are developed. Their area overheads are evaluated by implementing different schemes in 65 nm CMOS process. The security of the first-order LEMS with rotation mode is verified by performing CPA on the SAKURA-G FPGA board. From the experimental success rates, it shows that the proposed first-order LEMS can resist CPA without revealing the correct subkey for up to 100 000 power traces, whereas the unprotected scheme is broken at 1100 traces.
Leibo Liu, Qihuan Huang, Victor Y. Chen, Shouyi Yin, Shaojun Wei
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2017 Detecting on-body devices through creeping wave propagation
abstract
The ability to detect which wearables and smartphones are on the same body has the potential to support a wealth of applications, including user authentication, automatic data synchronization, and personalized profile loading. This paper brings this feature to commercial off-the-shelf (COTS) wearables and smartphones, by creating a virtual “on-body detection sensor” based on devices' inherent wireless capabilities. We investigate using the peculiar propagation characteristics of creeping waves to discern on-body wearables. To this end, we decompose signals into multiple independent components to exploit the variation features of creeping waves. We implement our system on COTS wearables and a smartphone. Extensive experiments are conducted in a lab, apartments, malls, and outdoor areas, involving 12 volunteer subjects of different age groups, to demonstrate the robustness of our system. Results show that our system can identify on-body devices at 92.3% average true positive rate and 5% average false positive rate.
Wei Wang 0050, Victor Y. Chen, Lin Yang 0009, Qian Zhang 0001
INFOCOM2
2017 Implementation of in-loop filter for HEVC decoder on reconfigurable processor
abstract
The in‐loop filter comprises deblocking filter and sample adaptive offset filter, which is an important module for improving image quality in a high‐efficiency video coding (HEVC) decoder. The in‐loop filter has a high computational complexity that accounts for ∼20% of the HEVC decoding computing load. Furthermore, it is difficult to implement a high‐performing in‐loop filter due to its large conditional processing requirement. First, this study presents a novel reconfigurable HEVC in‐loop filter implementation on a coarse‐grained dynamically reconfigurable processing unit. Next, a repartition scheme is presented that allows the in‐loop filter implementation at a coding tree unit along with the other decoding modules in the HEVC decoder, which satisfies requirements of low latency applications. Finally, a hierarchised‐pipeline and synchronised‐parallel technique is used to improve performance by eliminating data hazards in pipeline techniques and synchronisation problems in parallel techniques. Implementation results show that the presented HEVC in‐loop filter performs up to 1920 × 1080@52 frames per second at 250 MHz. The throughput is 67.5 × 9 × more than solutions based on digital signal processor and general‐purpose processor, respectively.
Leibo Liu, Victor Y. Chen, Chenchen Deng, Shouyi Yin, Shaojun Wei
IET Image Process.2
2017 Wideband Spectrum Adaptation Without Coordination
abstract
Fixed channelization configuration in today's wireless devices falls inefficient in the presence of growing data traffic and heterogeneous devices. In this regard, a number of fairly recent studies have provided spectrum adaptation capabilities for current wireless devices, however, they are limited to inband adaptation or incur substantial coordination overhead. The target of this paper is to fill the gaps in spectrum adaptation by overcoming these limitations. We propose SEER, a frame-level wideband spectrum adaptation solution which consists of two major components: i) a specially-constructed preamble that can be detected by receivers with arbitrary RF bands, and ii) a spectrum detection algorithm that identifies the desired transmission band in the context of multiple asynchronous senders by exploiting the preamble's temporal and spectral properties. SEER can be realized on commodity radios, and can be easily integrated into devices running different PHY/MAC protocols. We have prototyped SEER on the GNURadio/USRP platform to demonstrate its feasibility. Furthermore, using 1.6GHz channel measurements and trace-driven simulations, we have evaluated the merits of SEER over state-of-the-art approaches.
Wei Wang 0050, Victor Y. Chen, Zeyu Wang 0001, Jin Zhang 0001, Kaishun Wu, Qian Zhang 0001
IEEE Trans. Mob. Comput.2
2017 Sampleless Wi-Fi: Bringing Low Power to Wi-Fi Communications
abstract
The high sampling rate in Wi-Fi is set to support bandwidth-hungry applications. It becomes energy inefficient in the post-PC era in which the emerging low-end smart devices increase the disparity in workloads. Recent advances scale down the receiver's sampling rates by leveraging the redundancy in the physical layer, which, however, requires packet modifications or very high signal-to-noise ratio. To overcome these limitations, we propose Sampleless Wi-Fi, a standard compatible solution that allows energy-constrained devices to scale down their sampling rates regardless of channel conditions. Inspired by rateless codes, Sampleless Wi-Fi recovers under-sampled packets by accumulating redundancy in packet retransmissions. To harvest the diversity gain as rateless codes without modifying legacy packets, Sampleless Wi-Fi creates new constellation diversity by exploiting the time shift effect at receivers. Our evaluation using GNURadio/USRP platform and real Wi-Fi traces has demonstrated that Sampleless Wi-Fi significantly outperforms the state-of-the-art downclocking technique in both decoding performance and energy efficiency.
Wei Wang 0050, Victor Y. Chen, Lu Wang 0002, Qian Zhang 0001
IEEE/ACM Trans. Netw.2
2016 From rateless to sampleless: Wi-Fi connectivity made energy efficient
abstract
The high sampling rate in Wi-Fi is set to support bandwidth-hungry applications. It becomes energy inefficient in the post-PC era in which the emerging low-end smart devices increase the disparity in workloads. Recent advances scale down the receiver's sampling rates by leveraging the redundancy in the physical layer (PHY), which, however, requires packet modifications or very high signal-to-noise ratio (SNR). To overcome these limitations, we propose Sampleless Wi-Fi, a standard compatible solution that allows energy-constrained devices to scale down their sampling rates regardless of channel conditions. Inspired by rateless codes, Sampleless Wi-Fi recovers under-sampled packets by accumulating redundancy in packet retransmissions. To harvest the diversity gain as rateless codes without modifying legacy packets, Sampleless Wi-Fi creates new constellation diversity by exploiting the time shift effect at receivers. Our evaluation using GNURadio/USRP platform and real Wi-Fi traces have demonstrated that Sampleless Wi-Fi significantly outperforms the state-of-the-art downclocking technique in both decoding performance and energy efficiency.
Wei Wang 0050, Victor Y. Chen, Lu Wang 0002, Qian Zhang 0001
INFOCOM2
2016 Comparing bare-hand-in-air Gesture and Object-in-hand Tangible User Interaction for Navigation of 3D Objects in Modeling
abstract
3D modeling is used in Computer Graphics in various fields. Since the growth of gestures, virtual reality and embodied cognition, there have been various new technologies developed to either improve the modeling efficiency, or to provide more nature intuitive experience to the users. In this paper, from the user experience perspective, we try to compare these methods for navigation of 3D objects in the virtual modeling environment including: simple bare hand gestures, tangible user interfaces (TUI) with object in hand, as well as mouse/keyboard as the primary input. Based on embodied cognition theory, we hypothesis that the object-in-hand method might bring better user experience since the interaction between the object and hand can enhance the user's cognition while navigating a model. We present a conceptual design, with two approaches and three design models which demonstrate differences in user interaction with 3D modeling software.
Sanmathi Dangeti, Victor Y. Chen, Chunhui Zheng
TEI2
2016 The Role of Aesthetics and Perception in Raising Situation Awareness: Lessons from SpringRain
abstract
In the face of increased cyber risks, we present our iterative design process and the resulting principles of SpringRain, an information visualization (infoVis) display design concept for large screens in network control rooms (NOCs). It aims to raise team situation awareness by visualizing large-scale multidimensional computer network data sets as a live “rainfall.” We used aesthetic principles and theories of perception to prototype this ambient, yet data-dense, visualization. By applying multiple data dimensions to different properties of a line segment, such as length, motion, and color, the two-dimensional visualization offers analytical affordances (i.e., graphic qualities that make it clear how the display should be “read”) that can be processed pre-attentively (i.e., under 200 ms). This grants that time-sensitive anomalies in the network can be noticed, processed, and addressed in a timely manner. Our design approach was driven by theories rather than existing design works in the hope of encouraging more user-centered theories-driven infoVis designs that are better suited to user needs and requirements. Expert reviewers’ feedback from the VAST 2013 Challenge committee confirmed the novelty of the design and pointed to opportunities for improvements. We discuss three redesign attempts to address some of the identified limitations.
Marlen Promann, Cheryl Z. Qian, Victor Y. Chen
Int. J. Hum. Comput. Interact.4
2016 Less Transmissions, More Throughput: Bringing Carpool to Public WLANs
abstract
A typical scenario for public WLANs is large audience environment where Wi-Fi hotspots serve scores of mobile devices. The performance of those Wi-Fi hotspots is extremely poor in terms of low goodput and severe delay due to heavy contention and MAC inefficiency. After carefully investigating the traffic patterns in public WLANs, we proposeCarpool, a practical design that facilitates transmission sharing among multiple receivers, to tackle this problem. The key idea is to reduce contention by feeding frames for multiple destinations into one transmission at physical layer (PHY). As such, each downlink transmission carries payloads for multiple receivers, which reduces contention overhead and enables in-time response to concurrent requests from multiple users. To achieve efficient and reliable transmission in Carpool, we propose i) a lightweight frame structure to support multiple receivers, and ii) a real-time channel estimation scheme to continuously calibrate channel estimation during the transmission of a Carpool frame. We have implemented the entire PHY of Carpool on the GNURadio/USRP platform and tested it in various indoor environments. Furthermore, our trace-driven MAC evaluation shows that Carpool achieves up to$3.2 \times$goodput gain and reduces up to$75$percent delay compared to the IEEE 802.11n MAC frame aggregation scheme.
Wei Wang 0050, Victor Y. Chen, Qian Zhang 0001, Kaishun Wu, Jin Zhang 0001
IEEE Trans. Mob. Comput.2
2016 Privacy-Preserving Location Authentication in Wi-Fi Networks Using Fine-Grained Physical Layer Signatures
abstract
A recent measurement reveals that a large portion of the reported locations is either forged or superfluous, which raises security issues such as bogus alibis and illegal usage of restricted resources. However, most prior approaches leak users’ location information or rely on external devices. To overcome these limitations, we proposePriLA, a privacy-preserving location authentication system that verifies users’ location information based on physical layer (PHY) information available in legacy Wi-Fi preambles. The crux of PriLA is to turn detrimental features in wireless systems, namely carrier frequency offset (CFO) and multipath, into useful signatures for privacy protection and authentication. In particular, PriLA exploits CFO and channel state information (CSI) to secure wireless transmissions starting from the handshake phase between mobile users and the access point (AP), and meanwhile verify the truthfulness of users’ reported locations based on users’ multipath profiles. We have implemented PriLA on GNURadio/USRP platform and commercial off-the-shelf Intel 5300 NICs, and the experimental results show that PriLA achieves the authentication accuracy of 93.2% on average, while leaking merely 45.7% information in comparison with the state-of-the-art approach.
Wei Wang 0050, Victor Y. Chen, Qian Zhang 0001
IEEE Trans. Wirel. Commun.2
2015 A Mixed-Grained Reconfigurable Computing Platform for Multiple-Standard Video Decoding (Abstract Only)
abstract
A mixed-grained reconfigurable computing platform targeting multiple-standard video decoding is proposed in this paper. The platform integrates eight coarse-grained Reconfigurable Processing Units (RPUs), each of which consists of 16×16 multi-functional Processing Elements (PEs) and are implemented in TSMC 65 nm technology and two Altera Stratix IV EP4SE820 FPGAs. By exploiting dynamic reconfiguration of the RPUs and static reconfiguration of the FPGAs, the proposed platform achieves scalable performances and cost trade-offs to support a variety of video coding standards, including H.264, MPEG-2, AVS and HEVC. Two types of platform configuration are tested in this work. One configuration utilizes two RPUs and targets multiple-standard high-definition (HD) video decoding, while the other utilizes only one RPU, which works under a lower frequency and targets at standard resolution (SD) decoding. The HD configuration can decode 1920×1080 H.264 video streams at 30 frames per second (fps) under 200 MHz and 1920×1080 HEVC video streams at 30 fps under 236 MHz. It achieves a 25% performance gain over an industrial coarse-grained reconfigurable processor for H.264 decoding, and a 3.85× performance boosts over the Intel i5 general-purpose CPU for HEVC decoding.
Leibo Liu, Victor Y. Chen, Dong Wang 0040, Min Zhu 0001, Shouyi Yin, Shaojun Wei
FPGA2
2015 Less Transmissions, More Throughput: Bringing Carpool to Public WLANs
abstract
The proliferation of WiFi hotspots in public places enables ubiquitous Internet access. These public WiFi hotspots usually serve scores of mobile devices and suffer from extremely poor performance in terms of low good put and severe delay. In this paper, we first study the traffic characteristics in public WiFi networks, and demonstrate that the main causes of such poor performance are media access control (MAC) inefficiency and downlink-uplink traffic asymmetry. To cope with these issues, we call attention to transmission carpool, which facilitates an access point (AP) to send multiple frames for different mobile stations (STAs) in a single transmission. It reduces contention and conveys more frames in each channel access. As such, each downlink transmission carries more payload and thus improves efficiency and solves traffic asymmetry simultaneously.
Wei Wang 0050, Victor Y. Chen, Qian Zhang 0001, Kaishun Wu, Jin Zhang 0001
ICDCS2
2015 Changing channel without strings: Coordination-free wideband spectrum adaptation
abstract
Fixed channelization configuration in today's wireless devices falls inefficient in the presence of growing data traffic and heterogeneous devices. In this regard, a number of fairly recent studies have provided spectrum adaptation capabilities for current wireless devices, however, they are limited to inband adaptation or incur substantial coordination overhead. The target of this paper is to fill the gaps in spectrum adaptation by overcoming these limitations. We propose Seer, a frame-level wideband spectrum adaptation system which consists of two major components: i) a specially-constructed preamble that can be detected by receivers with arbitrary RF bands, and ii) a spectrum detection algorithm that identifies the intended transmission band in the context of multiple asynchronous senders by exploiting the preamble's temporal and spectral properties. Seer can be realized on commodity radios, and can be easily integrated into devices running different PHY/MAC protocols. We have prototyped Seer on the GNURadio/USRP platform and tested it under various environments. Furthermore, our evaluation using 1.6GHz spectrum measurements shows that Seer largely improves system throughput over fixed channel configuration and state-of-the-art spectrum adaptation approaches.
Wei Wang 0050, Victor Y. Chen, Zeyu Wang 0001, Jin Zhang 0001, Kaishun Wu, Qian Zhang 0001
INFOCOM2
2015 A visual analytics approach to detecting server redirections and data exfiltration
abstract
How to better find potential cyberattacks is a challenging question for security researchers and practitioners. In recent years, visualization has been applied in the field of analyzing cybersecurity issues, but most work has not been able to provide better than non-visualization based techniques. In this paper, we innovatively designed a visual analytics system to allow analysts to overview network traffic and identify such suspicious such activities as server redirection attack and data exfiltration. Because of the nature of the problem, the overview design must be scalable, accurate, and fast. Through aggregating traffic data along the two dimensions of duration and payload, the system reveals key network traffic characteristics for the analyst to identify security events. The system is evaluated with the test data sets from VAST 2013 mini-challenge 3. The results are very encouraging and shed a more positive light on applying visual analytics in information security.
Baijian Yang 0001, Victor Y. Chen
ISI3
2014 Privacy-preserving location authentication in WiFi with fine-grained physical layer information
abstract
The surging deployment of WiFi hotspots in public places drives the blossoming of location-based services (LBSs) available. A recent measurement reveals that a large portion of the reported locations are either forged or superfluous, which calls attention to location authentication. However, existing authentication approaches breach user's location privacy, which is of wide concern of both individuals and governments. In this paper, we propose PriLA, a privacy-preserving location authentication protocol that facilitates location authentication without compromising user's location privacy in WiFi networks. PriLA exploits physical layer information, namely carrier frequency offset (CFO) and multipath profile, from user's frames. In particular, PriLA leverages CFO to secure wireless transmission between the mobile user and the access point (AP), and meanwhile authenticate the reported locations without leaking the exact location information based on the coarse-grained location proximity being extracted from user's multipath profile. Existing privacy preservation techniques on upper layers can be applied on top of PriLA to enable various applications. We have implemented PriLa on GNURadio/USRP platform and off-the-shelf Intel 5300 NIC. The experimental results demonstrate the practicality of CFO injection and accuracy of multipath profile based location authentication in a real-world environment.
Victor Y. Chen, Wei Wang 0050, Qian Zhang 0001
GLOBECOM1
2014 Implementation of multi-standard video decoder on a heterogeneous coarse-grained reconfigurable processor
Leibo Liu, Victor Y. Chen, Dong Wang 0040, Shouyi Yin, Peng Cao 0002, Shaojun Wei
Sci. China Inf. Sci.2
2014 Implementation of AVS Jizhun decoder with HW/SW partitioning on a coarse-grained reconfigurable multimedia system
Leibo Liu, Victor Y. Chen, Shouyi Yin, Li Zhou 0015, Shaojun Wei
Sci. China Inf. Sci.2
2014 Employing a Parametric Model for Analytic Provenance
abstract
We introduce a propagation-based parametric symbolic model approach to supporting analytic provenance. This approach combines a script language to capture and encode the analytic process and a parametrically controlled symbolic model to represent and reuse the logic of the analysis process. Our approach first appeared in a visual analytics system called CZSaw. Using a script to capture the analyst’s interactions at a meaningful system action level allows the creation of a parametrically controlled symbolic model in the form of a Directed Acyclic Graph (DAG). Using the DAG allows propagating changes. Graph nodes correspond to variables in CZSaw scripts, which are results (data and data visualizations) generated from user interactions. The user interacts with variables representing entities or relations to create the next step’s results. Graph edges represent dependency relationships among nodes. Any change to a variable triggers the propagation mechanism to update downstream dependent variables and in turn updates data views to reflect the change. The analyst can reuse parts of the analysis process by assigning new values to a node in the graph. We evaluated this symbolic model approach by solving three IEEE VAST Challenge contest problems (from IEEE VAST 2008, 2009, and 2010). In each of these challenges, the analyst first created a symbolic model to explore, understand, analyze, and solve a particular subproblem and then reused the model via its dependency graph propagation mechanism to solve similar subproblems. With the script and model, CZSaw supports the analytic provenance by capturing, encoding, and reusing the analysis process. The analyst can recall the chronological states of the analysis process with the CZSaw script and may interpret the underlying rationale of the analysis with the symbolic model.
Victor Y. Chen, Cheryl Z. Qian, Robert F. Woodbury, John Dill, Chris Shaw 0002
ACM Trans. Interact. Intell. Syst.1
2014 SimRPU: A Simulation Environment for Reconfigurable Architecture Exploration
abstract
To assist the system architects with fast exploration and performance evaluation of the reconfigurable software/hardware architectures, this paper presents a system-level simulator, named after SimRPU, for the reconfigurable processing unit (RPU), which is the major computing engine in reconfigurable processor. The proposed simulator consists of a simulation kernel, a software compiler, a system profiler providing performance, area and power information for the desired architectures, and a system debugger supporting inspecting and modification of the internal state of the RPU. Object-oriented hierarchical and parameterized architecture modeling techniques are proposed to satisfy the requirements for a fast and comprehensive evaluation. Cycle-accurate simulation mechanisms are developed to improve the accuracy of the profiled performance data. Compared with the traditional register transfer level (RTL) based simulation scheme, the proposed simulator could achieve an average speedup of 18.5× with only 3.5% reduction on performance estimation accuracy. One reconfigurable processor targeted at high-definition multimedia decoding applications (such as H.264, MPEG2, AVS, etc.) is implemented with Taiwan Semiconductor Manufacturing Company 65-nm process using the proposed exploration and design flow. The measured results show that the implemented architecture has obvious advantages in terms of both performance and power consumption than the reference designs in multimedia decoding applications.
Leibo Liu, Dong Wang 0040, Shouyi Yin, Victor Y. Chen, Min Zhu 0001, Shaojun Wei
IEEE Trans. Very Large Scale Integr. Syst.4
2013 Implementation of multi-standard video decoding algorithms on a coarse-grained reconfigurable multimedia processor
abstract
This paper proposed a THPHP (Task-based Hybrid Parallels and Hybrid Pipelines) scheme to implement multistandard video decoding algorithms, i.e. MPEG-2, H.264 and AVS (Audio Video coding Standard), on a heterogeneous coarsegrained reconfigurable multimedia processor called REMUS (REconfigurable MUltimedia System). Multiple level parallelism and multiple level pipeline techniques are proposed in this scheme. Simulation results show that the video decoder can support H.264 HP (High Profile) 1920×1080@30fps (frame per second) streams, AVS JP (Jizhun Profile) 1920×1080@39fps streams, and MPEG-2 MP (Main Profile) 1920×1080@41fps streams when exploiting a 200MHz working frequency.
Leibo Liu, Victor Y. Chen, Shouyi Yin, Dong Wang 0040, Shaojun Wei, Li Zhou 0015, Peng Cao 0002
ISCAS2
2013 From when and what to where: Linking spatio-temporal visualizations in visual analytics
abstract
This paper proposes an autolinking approach to help analysts investigate spatial details of suspicious sections from an overview temporal visualization. Analysis of spatial-temporal network security data takes place both conditionally and in sequence. Many systems use time-series curves to visualize the temporal perspectives of the data and maps to show the spatial information. To identify anomalies, the analysts frequently shift across different visualizations. In essence, time-series curves provide a temporal overview of data, and the map anchors the locations for the users to drill down for details. Anomalies may be reflected in a time-series curve as a jump, a dive, a peak, or a valley. With the autolinking mechanism, after the analyst selects a segment of a curve, the system can automatically highlight the related area on the map for further investigation. This approach adopts the slicing operation of the Online Analytical Process (OLAP) to find the basic granularities that contribute to the overall value change. This approach is implemented in our award-winning visual analytics system SemanticPrism. In this paper, we describe its structure and demonstrate three examples of use with the VAST 2012 Minichallenge 1 data.
Victor Y. Chen, Cheryl Z. Qian
ISI1
2007 Visualizing Collaborative Filtering in Digital Collections
abstract
The NEAR (navigating exhibitions, annotations and resources) panel is a method of managing digital collections and user preferences through collaborative filtering and graphically revealing implicit data relations such as sharing, reference and similarity. It is implemented on AldrVIldrRE, an online multimedia repository. AldrVIldrRE supports semi-structured collections (exhibitions) which containing various resources and annotations. Its users are encouraged to contribute, share, annotate and interpret resources. Similar to the act of adding items into shopping carts in the e-commence applications, a user's activities of searching, organizing and interpreting data in AldrVIldrRE are considered as evidence of user's preferences. The design process of NEAR was guided by several principles from the visualization literature. It implements new navigation and communication approaches that support discovery of relations. Having tested NEAR with several users, we further analyze the design, report the evaluation and consider its use in other applications.
Victor Y. Chen, Cheryl Z. Qian, Robert F. Woodbury
IV1